Hacker News new | ask | show | jobs
by tidbeck 25 days ago
Related to this, for our use case, setting thinking to high instead of low made tasks complete faster and cheaper (Gemini 3.0 flash).

Other aspects are caching, often at 0.1X cost, where providers really differ in how efficient they are (Anthropic really good, Google not so much) and how chatty a model is (costing output tokens).

1 comments

I also don't see much of a quality difference switching between thinking levels while using those models as agents.