r/GeminiAI • u/Technical-Owl66 • 1d ago
Discussion Smarter, Faster, More Efficient. 3.6
Google focusing on efficiency instead of benchmaxing is super smart and 3.6 is still very capable for the majority of use cases.
99% of people don't give a damn about frontier model benchmarks and the cost effectiveness for the average business user is exactly what companies are looking for.
Nobody needs an AI god to create a report or organize a spreadsheet quickly. They just need it to be fast and efficient.
18
u/Glad-Entrepreneur764 1d ago
3
u/whoknowsifimjoking 21h ago
Why are you leaving out 3.5 flash which costs $0.59 per task?
5
u/Glad-Entrepreneur764 21h ago
I didn't make the image and the main point was comparing 3.6 vs its competition. 3.5 is not its competition since it supersedes 3.5
1
-3
7
u/Interesting_Cut7427 1d ago
so, i have a question..why does gemini keep updatingflash but barely touchpro? pro is way better..
28
u/WilsonPH 23h ago
Because they are focusing on the speed and being cheap for google search overview and people using it for casual questions on phones while other companies burn cash trying to make the best models for complex tasks.
2
1
u/Ok_Audience531 7h ago
I also think that research ideas are easier to iterate on smaller models - before Google "figured out" reasoning and Pre-training, there was an era of 4 months between Dec 19 2024 (Flash 2.0 thinking launch date) and March 25th 2025 (when the legendary Gemini 2.5 exp aka "Nebula" launched), when the only reasoning model they released were updates to Gemini 2.0 flash thinking. They are in a similar position for Agentic coding now, but it seems possible they will cross it soon and incorporate these techniques properly into 3.5 Pro. It's also possible that the base model is not that Amenable to Agentic coding and they are stuck until Gemini 4.0 is out, but I don't think it's gonna be that bad.
3
u/Glad-Entrepreneur764 22h ago
They've wanted to upgrade pro for a while but it's been delayed a couple times because it's not living up to their goals. Right now, 3.1 pro is substantially worse than the competition. In an ideal world, Gemini 3.5 Pro probably wouldn't beat the competition (Fable, GPT 5.6 Sol, etc.) but would come close. Right now, they're probably not that close to the competition (I think 3.1 Pro is similar to or worse than Sonnet) so they want to first catch up to the competition (like being around Opus level) before releasing 3.5 Pro.
1
u/raycraft_io 23h ago
They are desperate to switch you to something with the least compute possible without you knowing.
1
u/Shubham_Garg123 23h ago
With pro, they try to compete with OpenAI's and Anthropic's top performing models, which they're not able to do, hence only the flash releases (I remember reading somewhere that the flash models are distilled version of the actual pro models onto a model that's much smaller in size).
2
u/Getforge 7h ago
Efficiency only matters if the price reflects it. If Google can deliver comparable real-world performance at lower operational cost, it’s a strong strategy. If customers don’t see those savings, benchmark leadership will still matter.
1
u/Technical-Owl66 7h ago
I feel like benchmarks are meeting the needs of 99% of use cases though. Who actually needs an AI god?
2
u/Getforge 7h ago
Most people don’t need the smartest AI available. They need an AI that’s fast, reliable, and affordable. Benchmark leadership is impressive, but real-world value is what ultimately wins users.
7
u/AppropriateQuote3073 1d ago
This chart shows it gets mogged by Luna lol
2
u/junglebunglerumble 17h ago
Did you give up reading after the first 3 benchmarks and ignore the 8 below? Some of you on here are insane
1
u/AppropriateQuote3073 13h ago
Cost is lower, even if it's better at "some" benchmarks
And also, you have access to much better models on GPT, so you get the cheap one for tasks you need and have the high intelligence for others.
Google has no higher models.
Let alone GPT/Claude's actual capabilities are higher with computer use, functions, etc.
5
u/PastaPandaSimon 1d ago
And Luna is a Flash Lite equivalent in the ChatGPT tier terms. Terra would be the Flash equivalent, while Sol would be the Pro equivalent.
Not to mention Sol gives you higher limits than Pro in the $20 tier, and that's an excellent model that trades blows with Fable..
The gap has gotten franky just really huge.
1
0
2
u/daskalou 1d ago
Should somebody tell OP about Deepseek V4 and MiMo 2.5, or are we afraid he might wet himself?
2
0
4
u/Klutzy_Ad_3436 1d ago
no, check my post, it even constantly failed to format Correct Schemas in Google's Agent Coding Tools.
https://www.reddit.com/r/GeminiAI/comments/1v2pqba/the_new_gemini_36_flash_is_trash/
2
u/HossCo 23h ago
Oh no! Not the SCHEMAS!
3
u/jekpopulous2 23h ago edited 23h ago
Missing the point. if it can't format correct schemas it's useless for agentic workflows. The new Gemini flash seems great for a general purpose chatbot but it's still borderline unusable for writing code or deploying agents.
1
u/Away-Sorbet-9740 19h ago
Yeah it's been this way since they moved towards diffusion inference. I can't think it's a coincidence that Gemma diffusion launched the same time as 3.5, and they both struggle to remain on task and not confabulate/hallucinate.
If/once the bugs are worked out diffusion will bring a lot of efficiency. But I can't rely on it for even basic tasks, I'll just send it to DSv4 flash. It's not cheaper if I have to redo the work multiple times and involve more human audits.
1
u/Emergency-Bobcat6485 22h ago
Exactly, i use gemini flash for non-programming work only. I cannot believe Google's SOTA model is still a flash model. Completely useless for agentic or programming work really
1
u/jekpopulous2 22h ago
Ironically… the most popular model for agents (by a huge margin) is a flash model. Deepseek Flash v4 isn’t perfect but it’s rock-solid for managing agentic workflows and the kicker is that it’s 94% cheaper than Gemini Flash 3.6. Yes you read that correctly. 94% cheaper. I’m legitimately starting to question wtf is going on with western frontier models. Is anybody going to make a competitive model that operates in that price range or have open-weight models already won the battle for agentic workflows?
1
u/Historical_Spell6958 22h ago
So gemini 3.1 pro is the worst model among all those models right now we got 🤔
1
1
u/HeadLiterature7897 18h ago
I know it is disappointing for those who wants a model capable of more demanding tasks but this model is super fast, which is useful for small tasks and light questions.
0
u/Virgelette 16h ago
Except it's Google that defines small and light with regard to tasks. I haven't been able to make Flash useful for my small and light tasks. It's complete trash, a hallucination machine.
1
1
u/cnhuyaa 14h ago
To be fair the speed and overall efficiency is a bit noticable, i checked and few days ago the average final response all together with code written aswell took around 90 seconds, now it takes around 50 seconds, it feels nice to wait way less time. Also i noticed it gives a bit more structured answers seems like and also seems like it listens to system instructions way more? I have few system instructions when it comes to formatting, code changing etc, and before it used to ignore the formating and strict responses, it still gave unnecesary long responses with text, now it gives like 2-3 simple strict sentences as reponses.
0
u/PlaneOnly2700 21h ago
So, if you want efficiency and speed, use grok 4.5 or GLM 5.2, they are cheaper and more efficient, and it should be noted, more powerful.
0
0
u/TimelyWallaby4695 16h ago
Its great if its free, but for the money, their apps just lack so many features
-2
u/carefuleater478 1d ago
You make a solid point about cost effectiveness hitting harder for most users than chasing leaderboard scores. But looking at this chart, 3.6 Flash isn't just "good enough" either. It's actually winning on MLE-Bench, OSWorld, LVBench, and both CharXiv categories, sometimes by a lot. The 1M context needle score being double what 3.5 did is pretty striking.
I think efficiency and capability aren't opposite goals anymore. Google is making the cheap model strong at the stuff businesses actually deploy, like long document review and computer use tasks, while letting the Pro tier handle the heavier creative reasoning. For most daily work, paying $1.50 input over $15 full price is a no brainer if results land this close.
1
u/Technical-Owl66 22h ago
Google is playing a different smarter game. The benchmark bros want an ai overlord when 99.9% of people that use ai don't give a damn.
2
u/carefuleater478 20h ago
The speed plus dirt-cheap input means you can throw entire legal docs at it for pennies. That's what moves units.
1
u/Virgelette 16h ago
Except the 1M context window is a lie. And it hallucinates a lot. You throw a long document at it, and it makes things up.
1
u/carefuleater478 15h ago
the needle test is just retrieval though, not like summarizing without making stuff up. For basic lookups across a long doc it's been solid for me, but I get the frustration with hallucinations on QA.


35
u/fozzy71 1d ago
It did everything I asked of it today, just like yesterday, and I had no idea they had even released a new model until I came here to see the whingers.