Local AI and the jungle of incomparable specs
AMD, Apple and NVIDIA all sell a 128 GB local-AI box. The TOPS numbers are not comparable. Bandwidth is. Netherlands prices, real decode speeds, what 70B at 512K actually costs in memory, and how Qwen3.8-27B changes the 4/8/16-bit trade-off.
I opened a product page for a compact AMD box with 128 GB of memory and a claim that it runs 128B models locally. Then I tried to answer a simple question: what does the same job cost on NVIDIA or Apple, in the Netherlands, today.
Three days later I had a spreadsheet full of TOPS that cannot be added, VRAM that is not VRAM, and a new 27B model that thinks for two thousand tokens before it says hello.
The specs are not comparable. The workloads are. This post is the comparison I wanted on day one.
Prices are street or list in the Netherlands, including 21% BTW, mid-August 2026. Token rates are typical llama.cpp / Ollama / MLX / SGLang numbers from vendor notes and community benches, not one lab run. Treat them as bands, not gospel.
The trick in the brochure
The Minisforum MS-S1 Max is sold as a 128 GB AI workstation. That number is unified LPDDR5X. You can hand about 96 GB to the Radeon 8060S. It is not 128 GB of dedicated GDDR.
Apple does the same thing, with a wider bus. NVIDIA sells both: a tiny GB10 box that works like the Minisforum, and discrete cards that do not.
Once you see that, the market splits in two.
Same class as the Minisforum. One soldered SoC. One memory pool the GPU can see. Quiet. Low power. 128 GB on the label.
Different class. Discrete NVIDIA VRAM. Faster. Hotter. Much more expensive. And a single RTX 5090 is only 32 GB.
Most “128 GB VRAM” shopping questions mix those two classes and then compare TOPS.
Do not compare TOPS.
NVIDIA quotes sparse FP4. AMD quotes a platform total of GPU plus NPU. Apple’s 38 TOPS is the Neural Engine, which is not what llama.cpp uses. Three numbers, three units, one marketing department each.
Chat speed is almost entirely memory bandwidth. Prompt ingest is compute. Those are the two axes that matter.
The boxes, with Dutch prices
| System | GPU | Memory the model can use | Bandwidth | NL price | Where |
|---|---|---|---|---|---|
| Minisforum MS-S1 Max | Radeon 8060S, 40 CU | 128 GB unified, ~96 GB as VRAM | 256 GB/s | € 3.959 | minisforumpc.eu |
| ASUS Ascent GX10 | NVIDIA GB10 | 128 GB unified | 273 GB/s | € 3.400–€ 4.300 | Tweakers Pricewatch 1 TB, review |
| NVIDIA DGX Spark | NVIDIA GB10 | 128 GB unified | 273 GB/s | € 5.499 | Tweakers Pricewatch |
| Mac Studio M4 Max 36 GB | 32-core GPU | 36 GB unified | 410 GB/s | € 3.029 | Tweakers, Apple NL |
| Mac Studio M4 Max 64 GB | 40-core GPU | 64 GB unified | 546 GB/s | ~€ 4.359 | Tweakers 64 GB / 1 TB |
| Mac Studio M4 Max 128 GB | 40-core GPU | 128 GB unified | 546 GB/s | used ~€ 6.000–€ 7.500 | Apple stopped selling it new |
| Mac Studio M3 Ultra 96 GB | 60-core GPU | 96 GB unified | 819 GB/s | € 6.119–€ 6.349 | Tweakers, Upgreatest |
| RTX 5090 (card only) | GB202 | 32 GB GDDR7 | 1.792 GB/s | ~€ 4.899 | Tweakers / Megekko |
| Workstation + 1× RTX PRO 6000 | Blackwell 96 GB | 96 GB GDDR7 | 1.792 GB/s | € 18.000–€ 22.000 system, card ~€ 15.000 | Tweakers Workstation Edition, retail |
| Workstation + 4× RTX 5090 | 4× 32 GB | 128 GB split | 1.792 GB/s per card | € 24.000–€ 30.000 | four cards plus a board that can feed them |
Apple cut the 128 GB M4 Max and the 256 / 512 GB M3 Ultra configs during the RAM squeeze. New Mac Studios top out at 96 GB. That is why the 128 GB Apple row is a used-market number.
GPU memory bandwidthMS-S1 ████░░░░░░░░░░░░░░░░░░░░ 256 GB/sSpark / GX10 ████░░░░░░░░░░░░░░░░░░░░ 273 GB/sM4 Max 128 ████████░░░░░░░░░░░░░░░░ 546 GB/sM3 Ultra 96 ████████████░░░░░░░░░░░░ 819 GB/s5090 / 6000 ██████████████████████████ 1792 GB/sThe Minisforum and the Spark are the same machine class. Apple is two to three times the bus. A 5090 is seven times the bus and one quarter of the memory.
Two speeds, because there are two jobs
Decode is the next token in a chat. The model streams its own weights once per token. Speed tracks bandwidth.
Prefill is swallowing the prompt. That is matrix multiplies. NVIDIA compute shows up here. AMD’s 8060S does not.
Decode, dense 70B Q4 (~35 GB)
A 70B Q4 does not fit on a 5090. That is the first trap.
| System | Fits? | Typical tok/s | vs MS-S1 |
|---|---|---|---|
| MS-S1 Max | yes | 5–8 | 1.0× |
| Spark / GX10 | yes | 8–18 | ~1.5× |
| M4 Max 128 GB | yes | 12–16 | ~2× |
| M3 Ultra 96 GB | yes | 25–32 | ~4× |
| RTX 5090 32 GB | no | 2–20 offloaded | often slower than the mini-PC |
| RTX PRO 6000 96 GB | yes | ~32 | ~5× |
| 4× RTX 5090 | yes, split | 80–150 | ~15–20× |
Dense 70B Q4 decode (midpoint)MS-S1 ██░░░░░░░░░░░░░░░░░░░░ 6 tok/sSpark / GX10 ████░░░░░░░░░░░░░░░░░░░░ 13 tok/sM4 Max 128 ████░░░░░░░░░░░░░░░░░░░░ 14 tok/sM3 Ultra 96 ████████░░░░░░░░░░░░░░░░ 28 tok/sPRO 6000 █████████░░░░░░░░░░░░░░░ 32 tok/s4× 5090 ██████████████████████████ 100 tok/sDecode, 32B Q4 (~18 GB) — the fair race
Everything holds this one.
| System | Typical tok/s | vs MS-S1 |
|---|---|---|
| MS-S1 Max | 12–18 | 1.0× |
| Spark / GX10 | 15–22 | ~1.2× |
| M4 Max 128 | 28–40 | ~2× |
| M3 Ultra 96 | 40–55 | ~3× |
| RTX 5090 | 40–55 | ~3× |
| PRO 6000 | 45–70 | ~4× |
Same ranking as the bandwidth chart. That is not a coincidence.
Prefill
| System | Prompt ingest | vs MS-S1 |
|---|---|---|
| MS-S1 Max | ~300–400 tok/s | 1× |
| M4 Max | ~700–1.000 | ~2–3× |
| M3 Ultra | ~1.000–1.400 | ~3–4× |
| Spark / GX10 | ~1.700–1.900 | ~5× |
| 5090 / PRO 6000 | thousands | ~10–20× |
Paste a long PDF into the AMD box and you wait. Spark eats the prompt, then crawls on the reply at almost the same speed as AMD. Apple is the opposite: slower prefill than Spark, faster chat, because the bus is wider.
Chat speed follows bandwidth:
MS-S1 (256) → Spark (273) → M4 Max (546) → M3 Ultra (819) → 5090 (1792)
Prompt speed follows compute:
MS-S1 8060S → M4 Max → M3 Ultra → Spark GB10 → 5090 / PRO 6000
A typical turn — 2.000-token prompt, 300-token answer, dense 70B Q4 — lands roughly here:
| System | First token | Finish answer | Total |
|---|---|---|---|
| MS-S1 Max | 5–7 s | 40–60 s | ~1 minute |
| Spark / GX10 | 1–2 s | 15–25 s | ~25 s |
| M4 Max 128 | 2–3 s | 20–25 s | ~25 s |
| M3 Ultra 96 | ~2 s | 10–12 s | ~13 s |
| PRO 6000 | <1 s | ~10 s | ~10 s |
| 4× 5090 | <0.5 s | 2–4 s | ~4 s |
70B at 512K context is a memory problem first
People say “70B with a 512K window” as if that were one more checkbox. It is not. The weights are the small part.
A Llama-3-70B-class model (80 layers, 8 KV heads, head dim 128) stores this much key/value cache:
2 × layers × kv_heads × head_dim × seq_len × bytes
That is 320 KiB per token in FP16. At 524.288 tokens (512K):
| KV precision | Cache alone | Plus 70B Q4 weights (~40 GB) | Plus 70B Q8 (~72 GB) |
|---|---|---|---|
| Q4 KV | ~40 GB | ~80 GB | ~112 GB |
| FP8 / Q8 KV | ~80 GB | ~120 GB | ~152 GB |
| FP16 KV | ~160 GB | ~200 GB | ~232 GB |
Add 10–20% for paging waste, CUDA graphs, the OS, and the fact that 128 GB on the box is not 128 GB for the model.
70B + 512K: memory you must holdQ4 weights + Q4 KV ████████░░░░░░░░░░░░░░░░ 80 GBQ4 weights + FP8 KV ████████████░░░░░░░░░░░░ 120 GBQ4 weights + FP16 KV ████████████████████░░░░ 200 GBQ8 weights + FP16 KV ███████████████████████░ 232 GBWho can actually load it:
| Setup | Usable memory | 70B Q4 + Q4 KV 512K | 70B Q4 + FP16 KV 512K |
|---|---|---|---|
| MS-S1 Max | ~96 GB VRAM of 128 | tight, maybe | no |
| Spark / GX10 | 128 GB unified | tight | no |
| M4 Max 128 GB | 128 GB unified | tight | no |
| M3 Ultra 96 GB | 96 GB | borderline | no |
| RTX 5090 | 32 GB | no | no |
| PRO 6000 | 96 GB | tight on Q4 KV | no |
| 2× PRO 6000 | 192 GB | yes | barely, with KV quant |
| 4× 5090 | 128 GB split | only if the runtime shards KV | no as one pool |
| M3 Ultra 256 GB (discontinued) | 256 GB | yes | yes |
A 96–128 GB box can hold 70B Q4 at 512K only if you quantise the cache hard and accept the quality hit. Full-quality 512K on a 70B is a 200 GB problem. That is two workstation GPUs or a Mac Apple no longer sells.
And even when it fits, decode is not the short-context number. Attention cost grows with sequence length. At 512K, expect roughly 20–40% of the short-context tok/s, if the stack does not fall over first.
| System | Short-context 70B Q4 | Expected at 512K, if it fits |
|---|---|---|
| MS-S1 / Spark | 6–15 tok/s | 2–6 tok/s, if Q4 KV fits |
| M4 Max 128 | 12–16 | 4–8 |
| M3 Ultra 96 | 25–32 | 6–12, only with aggressive KV quant |
| PRO 6000 | ~32 | 8–15 with Q4 KV |
| 2× PRO 6000 | 50–80 | 15–30 |
I would not buy a €4k mini-PC for “70B at 512K”. I would buy it for 70B at 8K–32K, which is what most chat and coding loops actually use.
If you truly need 512K on a 70B-class dense model, start at 192 GB of GPU memory and a stack that can quantise KV. Or stop pretending and use a model that was built for long context.
Which is the other half of this post.
Qwen3.8-27B — the model you can actually run
Alibaba shipped Qwen3.8-27B in mid-August 2026. Hugging Face lists it as ~28B parameters with embeddings. The name is 27B. Dense, not MoE. Native vision. Hybrid thinking. Multi-token prediction. 256K context, 1M with YaRN.
It is a different architecture from Llama-70B, and that is why the 512K conversation changes.
The backbone is 16 blocks of three Gated DeltaNet layers plus one Gated Attention layer. Only 16 attention layers keep a KV cache, with 4 KV heads and head dim 256. The DeltaNet part is linear in sequence length. The cache is small.
2 × 16 × 4 × 256 = 32.768 elements per token
That is 64 KiB per token in FP16. About 2.0 GiB at 32K, which matches the ~2.5 GB people are measuring once you add overhead.
| Context | FP16 KV | FP8 KV | Q4 KV |
|---|---|---|---|
| 32K | ~2.5 GB | ~1.3 GB | ~0.7 GB |
| 128K | ~10 GB | ~5 GB | ~2.5 GB |
| 256K native | ~20 GB | ~10 GB | ~5 GB |
| 512K (YaRN) | ~40 GB | ~20 GB | ~10 GB |
Weights at 4 / 8 / 16 bit
From the Unsloth card:
| Precision | Weights | Plus 32K | Plus 256K FP16 KV | Plus 512K FP16 KV |
|---|---|---|---|---|
| 4-bit / IQ4 / NVFP4 | 17–19 GB | ~20 GB | ~38 GB | ~58 GB |
| 8-bit | ~31 GB | ~34 GB | ~51 GB | ~71 GB |
| BF16 / 16-bit | ~56 GB | ~59 GB | ~76 GB | ~96 GB |
At 256K native context, Qwen3.8-27B lands here:
| Weights | KV (256K) | Total in GPU memory |
|---|---|---|
| 4-bit ~18 GB | Q4 KV ~5 GB … FP16 KV ~20 GB | ~23–38 GB |
| 8-bit ~31 GB | same | ~41–51 GB |
| 16-bit ~56 GB | FP16 KV ~20 GB | ~76 GB |
Every machine in the first table can run 4-bit at 32K. A 5090 can run 4-bit and 8-bit at moderate context, then runs out. BF16 plus a long window needs the 96–128 GB boxes.
Speed vs quality vs context
Three knobs. You do not get all three.
4-bit. Fastest. 17–19 GB. Fine for chat and most coding. You will feel it on hard reasoning and on tool-call JSON if you hunt for it. NVFP4 on Blackwell is the same size class with better recovery than a naive Q4 — Unsloth reports 92–97% top-1 versus BF16 and about 1.5× BF16 throughput.
8-bit. The quality default. Near-lossless for almost everything I care about. 31 GB. A 5090 can hold it until the context grows. The unified 96–128 GB boxes hold it with room for 256K.
16-bit. Reference. Fine-tuning, logit-sensitive work, or “I do not want to argue about quant”. 56 GB. Not a 5090. Comfortable on Spark, MS-S1, M3 Ultra, PRO 6000.
Quality vs speed, Qwen3.8-27B (relative, not a lab plot):
| Slower | Faster | |
|---|---|---|
| Cleaner | Spark BF16 · M3 Ultra Q8 · MS-S1 Q8 | PRO 6000 Q8 · 5090 NVFP4 |
| Lossier | MS-S1 Q4 | Spark NVFP4 · M3 Ultra Q4 |
Measured and estimated decode, Q4, short context, batch 1:
| System | Stack | tok/s | vs MS-S1 |
|---|---|---|---|
| MS-S1 Max | llama.cpp Vulkan, no MTP | ~10–12 | 1.0× |
| MS-S1 Max | Vulkan + MTP=4 | up to 24.5 (AMD, Windows) | ~2× |
| Spark | llama.cpp Q4 + MTP | 24–30 | ~2× |
| Spark | SGLang + NVFP4 + MTP | 34–38, peak ~47 | ~3× |
| M4 Max 128 | MLX / Metal | ~25–40 | ~2–3× |
| M3 Ultra 96 | MLX / Metal | ~35–55 | ~3–4× |
| RTX 5090 | llama.cpp Q4 | ~60–90 | ~6× |
| RTX 5090 | SGLang + NVFP4 + MTP | ~120–200 | ~10× |
| PRO 6000 | same kernels, more room | ~80–150 | ~8× |
AMD’s own day-0 note is the 24.5 figure, on a GMKtec EVO X2, Vulkan, MTP=4. Community Q5 on Strix Halo without that stack is closer to 10 tok/s. Same chip, factor of two, depending on whether speculative decoding is actually on.
NVFP4 is Blackwell only. Spark, 5090, PRO 6000. Not AMD. Not Apple.
The thinking tax
Thinking is on by default (reasoning_effort: xhigh). A short coding answer often burns 2.000–8.000 thinking tokens before the first useful line.
| System | 3.000 thinking tokens | 400 answer tokens | You wait |
|---|---|---|---|
| MS-S1, typical Q5 | ~4–5 min | ~30 s | ~5 min |
| MS-S1, AMD MTP peak | ~2 min | ~16 s | ~2.5 min |
| Spark NVFP4 | ~80 s | ~11 s | ~1.5 min |
| M3 Ultra Q4 | ~70 s | ~9 s | ~1.3 min |
| 5090 / PRO 6000 | 20–40 s | 3–5 s | ~30–45 s |
Turn thinking to low or off for chat and every row gets several times faster. For agent work you want it on. Then raw tok/s is the product.
Long context on this model is the opposite of 70B. 256K native plus Q4 weights is about 38 GB. That is a Minisforum or a Spark or a Mac, not a science project. 512K with FP16 KV is about 58 GB in 4-bit. Still one box. Quality at 512K will depend more on YaRN than on whether you picked Q4 or Q8.
Qwen3.8-27B Q4 + FP16 KV versus 70B Q4 + FP16 KV, same context lengths:
32K 128K 256K 512KQwen 27B Q4 20 GB 28 GB 38 GB 58 GB ██ ███ ████ ██████
70B Llama Q4 50 GB 80 GB 120 GB 200 GB █████ ████████ ████████████ ████████████████████The first series is Qwen3.8-27B Q4 plus FP16 KV. The second is 70B Q4 plus FP16 KV. Same context lengths. Different decade of hardware.
What I would actually buy
For Qwen3.8-27B as a daily coder, thinking on:
- Cheap and good enough: MS-S1 Max, and turn thinking down when you are iterating.
- Same money, CUDA, faster prompts: ASUS GX10.
- Quiet and fast enough that thinking does not hurt: M3 Ultra 96 GB.
- Interactive with
xhighthinking: a 5090 if you stay at 4-bit and moderate context, a PRO 6000 if you want Q8, vision, and 256K at the same time.
For 70B chat at 8K–32K: M3 Ultra if you live in macOS, Spark if you need CUDA, MS-S1 if the budget is the point. A lone 5090 loses.
For 70B at 512K: do not start from the €4k mini-PC list. You are shopping 192 GB+ or a different model.
The jungle is the TOPS column. The path out is bandwidth, then whether the weights plus the KV cache fit, then whether you are measuring prefill or decode.
Everything else is a brochure.