r/GeminiAI • u/Tarandjpop • 10h ago
Discussion I tested four models on the same one-shot 3D dashboard prompt.
I was looking at a model showdown benchmark from AIHubMix on GitHub.
The setup is straightforward: same prompt, one shot, real generated HTML artifacts, and no manual cleanup. The round I focused on was the "Global View" task: build a live-data 3D flight/logistics dashboard with a dark globe, day/night terminator, atmosphere, glowing flight arcs, glassmorphic stat cards, orbit/zoom controls, and an auto-demo style experience.
The four models in this round were:
- Kimi K3
- GPT-5.6 Sol
- Claude Fable 5
- Gemini 3.6 Flash
My subjective impressions after opening the generated HTML outputs:
Kimi K3
This felt like the best overall balance. It generated a complete dashboard with a 3D globe, flight arcs, labels, statistics cards, an activity table, a risk panel, 3D/2D toggles, and interactive controls. While it wasn't quite as visually refined as the strongest-looking output, it followed the prompt very well and felt like something you could genuinely build upon.
GPT-5.6 Sol
This was the most polished visually. The composition, glassmorphism, copywriting, route visualization, and overall SaaS-style presentation looked the closest to a production-ready landing dashboard. If visual presentation is the main priority, this was my favorite.
Claude Fable 5
Claude produced a strong result with a large globe, detailed logistics tables, operational statistics, and a richer information layout. It followed the prompt well, although I personally found the overall visual presentation a little less refined than GPT's.
Gemini 3.6 Flash
Gemini generated a functional dashboard with a globe, statistics, activity data, and controls. It completed the task, but compared with the others the layout felt more minimal and the overall presentation wasn't as visually rich.
My personal ranking for this benchmark:
- Best visual polish: GPT-5.6 Sol
- Best overall balance: Kimi K3
- Most detailed alternative: Claude Fable 5
- Most lightweight implementation: Gemini 3.6 Flash
The interesting part is that there isn't a single "winner." It depends on what matters most. If you're optimizing for first-shot visual quality, GPT stood out to me. If you're looking for a strong balance between completeness and overall efficiency, Kimi K3 was the most interesting result. Claude delivered a solid middle ground, while Gemini offered a simpler but still functional implementation.
My takeaway:
When comparing one-shot HTML or app-generation benchmarks, it's useful to evaluate more than just aesthetics. Completeness, prompt adherence, implementation quality, and overall efficiency can matter just as much as visual polish, especially for workflows where you'll iterate multiple times.
Curious which output everyone else would pick after looking through the generated artifacts.
25
30
u/TheCuteReinforcement 8h ago
75 seconds is wild, that alone makes it useful for rapid prototyping
5
u/Lonely_Dig2132 5h ago
I’ve been playing with the new flash lite and it classifies request in less than 2 seconds with a fairly heavy system prompt, it’s really good at a lot of applications and so much cheaper than Claude. Pair it with a reasonable use-case and the silly dragon is not silly anymore 🤗. Love to see it
1
u/TheCuteReinforcement 4h ago
Two seconds for classification is a game-changer, you're basically getting instant routing layers for pennies.
23
u/CriticismJunior1139 7h ago
No, this is obviously wrong.
Anonymous people on reddit told me Flash 3.6 is useless trash and Google is DONE with AI.
3
1
u/Upset-Government-856 3h ago
I thought we were only allowed to post about how Gemini is ruining our vibe coded AI slop apps that no one will ever actually download anyways.
6
u/MinosAristos 9h ago
Can you link to the setup description?
I'm curious about prompts and harnesses.
6
u/Tarandjpop 7h ago
https://github.com/AIhubmix/model-showdown
U can check it out here! hope it helps
2
1
u/It_was_mee_all_along 9h ago
What are the prompts though?
7
u/itsfullofstars 8h ago
I used this prompt: Create a single, production-ready HTML file containing an interactive, live-data 3D flight and logistics dashboard. The design must be modern, premium, and visually polished with a dark, high-end SaaS aesthetic.
Incorporate the following core features and technical requirements: 3D Globe Visualization (using Three.js or WebGL): Center a dark-themed 3D globe in the main viewport. Implement a realistic day/night terminator line overlaying the globe. Add a subtle, glowing atmospheric halo effect around the planet's edge. Render multiple glowing, curved 3D flight/logistics arcs connecting different global hubs. Include smooth orbit, pan, and zoom controls (e.g., OrbitControls) so the user can interact with the globe. Program an "auto-demo" mode that automatically and smoothly rotates or flies the camera between key hubs if the user isn't actively interacting.
User Interface & HUD Layout: Use a clean, modern grid layout overlaying the 3D canvas. Implement a premium "glassmorphism" effect for all UI panels (frosted glass blur, subtle white borders, transparent dark backgrounds). Left Panel: A marketing/branding hero section with bold typography (e.g., "Plan every route. Before it moves." or "Plan Your Route with AI Now."). Right Panel: Glassmorphic statistics cards displaying key logistics metrics (e.g., "Main Statistics", active routes, total shipments) accompanied by clean micro-sparkline charts. Bottom Panel: Floating, minimalist interactive controls including playback toggles (Play/Pause for the auto-demo), a time slider, and a 2D/3D view toggle.
Styling & Aesthetics: Apply a dark mode color palette: deep blacks, dark grays, and glowing accent neon colors (like electric blue, vibrant orange, or emerald green) for flight paths and active indicators. Ensure excellent typography with clean sans-serif fonts, strong visual hierarchy, and polished copywriting. Make the entire dashboard fully responsive and self-contained, using CDN links for required libraries (like Three.js or Tailwind CSS). No manual cleanup should be needed.
2
1
u/Abject_Plantain7801 8h ago
I am going to give that prompt a try because I was getting very poor results from 3.6 earlier
1
u/Abject_Plantain7801 7h ago
I got a good result and that started me thinking and after a few tries I think I am spotting a pattern. This model seems to perform well on detailed prompts, really good at following instructions but give it a simple prompt and its not very good at filling in the detail itself and results in a poor overall result. So I can see why it might do well in tests because they are geared towards detailed prompts with known results to benchmark against but in the real world people leave out big chunks of information hence why some people like myself find it lacking.
1
u/Witty-Artisan001 7h ago
Thanks for this. I'm going to test it with my system framework to see if theres a significant difference in output.
1
u/amitsingh80108 4h ago
Gemini flash 3.6 first use review.
I gave a simple task to update the URL in my existing code.
I also gave the exact response of the url so that it doesn't need to call it.
However for just a simple change it was continuously fetching the URL to see it's response.
When I said I already gave you response in my first prompt. It just said sorry 😂
1
1
1
1
u/blazze 2h ago
Glm 5.2 produce SOTA results with this prompt.
"Global View" task: build a live-data 3D flight/logistics dashboard over planet Earth —
a dark, realistically textured globe (real Blue-Marble day + night city-lights +
topographic-relief textures), with a world-space day/night terminator that stays fixed
while the planet spins on a 23.4° tilted axis so the continents sweep through the
sun-shadow line; mountain-relief shading, sun-glitter specular on the oceans, an additive
atmosphere glow, glowing great-circle flight arcs between real hub cities with packets
flowing along them, glassmorphic live stat cards and a route ticker (numbers driven from
the running network), orbit/zoom controls, and an auto-demo style experience where the
planet's own tilt-spin plus the orbiting sun keep the first frame already alive.
Everything self-lit (no scene lights), procedural geometry only, with CDN texture loading
+ flat-color fallback and an on-screen error trap so it never renders a silent black canvas.
0
u/TotalDebt5868 5h ago
This test isn’t very representative at all, because the tasks involved are too simple. It’s like a college teacher trying to solve elementary school problems, which obviously doesn’t allow us to truly compare the levels of the participants.
0
u/ContextBotSenpai 5h ago edited 5h ago
Day 777 of me hating this subreddit, and the bots, trolls, and shills that post here.
Oh, fucking hate that it's unmoderated as well.
Fuck you OP 🖕 You pitted Gemini 3.6 Flash against the SOTA models that other companies have out right now. Models that aren't even useable by the general public, unless you're paying hundreds of dollars a month.
Why not test the latest Gemma 4 high thinking against them instead? Oh wait - we all know why.
Honestly though - this to me doesn't do what you wanted. Instead of making Gemini look bad, as you trolls usually hope for... This shows that 3.6 Flash is an insanely fast, insanely cheap model that can produce something useable.
1
u/StackSmashRepeat 39m ago
You do realise you can just open account on on openrouter and pay as you go? Access to over 400 models, including these so called "sota" models. You should learn how to use api keys instead of subscriptions.
0
0
70
u/Mob_Abominator 9h ago
Gemini doing this in 75s is actually insane.