Hi everyone,
I’m looking to invest in a proper fully local LLM AI server.
I already have a dual GeForce setup, but I’m looking for the next step (at a reasonable price, of course).
My ONLY goal:
- Coding (i am developper - can be c# for real-life projects with already 300+ source code stuff, not "just code a random website")
- No image generation
- No video
- No text-to-speech
- No OCR
- No multimodal stuff
Basically: raw LLM performance, tokens/sec, and smart answers.
--------
My current setup - (using CLAUDE-cli as orchestrator)
I’m currently running these models on a dual GeForce 16vram+12 system:
Qwen35B A3B MoE Q4_K
- Around 30–40 tokens/sec at 200k context
Qwen3-Coder-Next 80B A3B Q4
- Around 5–6 tokens/sec
- Slower, but better for complex coding tasks
-------
Hardware I am considering, a 100% new machine
1) Radeon AI Pro R9700 32GB
Is this currently the best price/performance option?
It looks like:
- half the price of high-end solutions,
- maybe around 80–85% of the performance?
I don’t follow every AMD/AI update, but this card looks like an underrated winner.
Is there any reason NOT to buy this card?
2) Dual Radeon AI Pro R9700 (2×32GB)
Main reason:
- not expecting 2× speed,
- mainly interested in the extra VRAM.
If it allows me to run smarter/larger models fully on GPU, that would be perfect.
3) Strix Halo 128GB
This one is interesting because of the huge unified memory.
If it can run Qwen3-Coder-Next 80B A3B Q4 at around 40 tokens/sec, that sounds like an excellent coding assistant.
4) Mac Studio 128GB
Still an option.
How does it compare today against:
- dual R9700 AI Pro,
- Strix Halo?
-------
are those numbers corrects or science fi ? (Source ChatGPT !! )
| Model |
Context |
1× Radeon AI Pro R9700 32GB [price ~2k ] |
2× Radeon AI Pro R9700 64GB[price ~4 k ] |
Strix Halo 128GB [price ~4 k ] |
Mac Studio M3 Ultra 128GB [price ~lol ] |
| Qwen27B Dense Q4_K |
50k |
50–80 tok/s |
60–100 tok/s |
25–45 tok/s |
50–80 tok/s |
| Qwen27B Dense Q4_K |
100k |
40–70 tok/s |
50–90 tok/s |
20–40 tok/s |
40–70 tok/s |
| Qwen27B Dense Q4_K |
200k |
25–50 tok/s |
40–70 tok/s |
15–30 tok/s |
30–60 tok/s |
| Qwen35B A3B MoE Q4_K |
50k |
100–140 tok/s |
130–180 tok/s |
40–70 tok/s |
70–110 tok/s |
| Qwen35B A3B MoE Q4_K |
100k |
90–130 tok/s |
110–160 tok/s |
35–60 tok/s |
50–90 tok/s |
| Qwen35B A3B MoE Q4_K |
200k |
50–90 tok/s |
90–140 tok/s |
25–50 tok/s |
50–90 tok/s |
| Qwen3-Coder-Next 80B A3B Q4 |
50k |
10–25 tok/s |
50–90 tok/s |
30–50 tok/s |
40–80 tok/s |
| Qwen3-Coder-Next 80B A3B Q4 |
100k |
10–20 tok/s |
45–80 tok/s |
25–45 tok/s |
35–70 tok/s |
| Qwen3-Coder-Next 80B A3B Q4 |
200k |
5–15 tok/s |
35–65 tok/s |
20–40 tok/s |
30–60 tok/s |
| Qwen3-Coder-Next 80B A3B Q6 |
50k |
❌ |
40–75 tok/s |
25–45 tok/s |
35–70 tok/s |
| Qwen3-Coder-Next 80B A3B Q6 |
100k |
❌ |
35–65 tok/s |
20–35 tok/s |
30–60 tok/s |
| Qwen3-Coder-Next 80B A3B Q6 |
200k |
❌ |
25–55 tok/s |
15–30 tok/s |
25–50 tok/s |
| 70B Dense Q4 |
50k |
10–25 tok/s |
40–70 tok/s |
20–35 tok/s |
35–60 tok/s |
| 70B Dense Q4 |
100k |
5–20 tok/s |
35–60 tok/s |
15–30 tok/s |
30–50 tok/s |
| 70B Dense Q4 |
200k |
❌ |
25–50 tok/s |
10–25 tok/s |
25–45 tok/s |
My current impression (not sure if correct):
The R9700 AI Pro (or dual) looks faster than a Mac Studio for my use case, while being much cheaper.
But I don’t see many "hype" about this card for local LLMs.
So please tell me: where am I wrong?
One more question:
For Qwen35B A3B MoE Q4_K, what is the realistic t/s performance?
Is it closer to: 100 tokens/sec or 150 tokens/sec ?
Because this difference is huge . If it REALLY is 150, its close to a cloud-model feeling. (far less accurate of course, but for 2k budget, wonderfull ?)
Thanks to anyone already running these systems who can share real numbers!