Local AI and the jungle of incomparable specs

AMD, Apple and NVIDIA all sell a 128 GB local-AI box. The TOPS numbers are not comparable. Bandwidth is. Netherlands prices, real decode speeds, what 70B at 512K actually costs in memory, and how Qwen3.8-27B changes the 4/8/16-bit trade-off.

I opened a product page for a compact AMD box with 128 GB of memory and a claim that it runs 128B models locally. Then I tried to answer a simple question: what does the same job cost on NVIDIA or Apple, in the Netherlands, today.

Three days later I had a spreadsheet full of TOPS that cannot be added, VRAM that is not VRAM, and a new 27B model that thinks for two thousand tokens before it says hello.

The specs are not comparable. The workloads are. This post is the comparison I wanted on day one.

Prices are street or list in the Netherlands, including 21% BTW, mid-August 2026. Token rates are typical llama.cpp / Ollama / MLX / SGLang numbers from vendor notes and community benches, not one lab run. Treat them as bands, not gospel.

The trick in the brochure

The Minisforum MS-S1 Max is sold as a 128 GB AI workstation. That number is unified LPDDR5X. You can hand about 96 GB to the Radeon 8060S. It is not 128 GB of dedicated GDDR.

Apple does the same thing, with a wider bus. NVIDIA sells both: a tiny GB10 box that works like the Minisforum, and discrete cards that do not.

Once you see that, the market splits in two.

Same class as the Minisforum. One soldered SoC. One memory pool the GPU can see. Quiet. Low power. 128 GB on the label.

Different class. Discrete NVIDIA VRAM. Faster. Hotter. Much more expensive. And a single RTX 5090 is only 32 GB.

Most “128 GB VRAM” shopping questions mix those two classes and then compare TOPS.

Do not compare TOPS.

NVIDIA quotes sparse FP4. AMD quotes a platform total of GPU plus NPU. Apple’s 38 TOPS is the Neural Engine, which is not what llama.cpp uses. Three numbers, three units, one marketing department each.

Chat speed is almost entirely memory bandwidth. Prompt ingest is compute. Those are the two axes that matter.

The boxes, with Dutch prices

SystemGPUMemory the model can useBandwidthNL priceWhere
Minisforum MS-S1 MaxRadeon 8060S, 40 CU128 GB unified, ~96 GB as VRAM256 GB/s€ 3.959minisforumpc.eu
ASUS Ascent GX10NVIDIA GB10128 GB unified273 GB/s€ 3.400–€ 4.300Tweakers Pricewatch 1 TB, review
NVIDIA DGX SparkNVIDIA GB10128 GB unified273 GB/s€ 5.499Tweakers Pricewatch
Mac Studio M4 Max 36 GB32-core GPU36 GB unified410 GB/s€ 3.029Tweakers, Apple NL
Mac Studio M4 Max 64 GB40-core GPU64 GB unified546 GB/s~€ 4.359Tweakers 64 GB / 1 TB
Mac Studio M4 Max 128 GB40-core GPU128 GB unified546 GB/sused ~€ 6.000–€ 7.500Apple stopped selling it new
Mac Studio M3 Ultra 96 GB60-core GPU96 GB unified819 GB/s€ 6.119–€ 6.349Tweakers, Upgreatest
RTX 5090 (card only)GB20232 GB GDDR71.792 GB/s~€ 4.899Tweakers / Megekko
Workstation + 1× RTX PRO 6000Blackwell 96 GB96 GB GDDR71.792 GB/s€ 18.000–€ 22.000 system, card ~€ 15.000Tweakers Workstation Edition, retail
Workstation + 4× RTX 50904× 32 GB128 GB split1.792 GB/s per card€ 24.000–€ 30.000four cards plus a board that can feed them

Apple cut the 128 GB M4 Max and the 256 / 512 GB M3 Ultra configs during the RAM squeeze. New Mac Studios top out at 96 GB. That is why the 128 GB Apple row is a used-market number.

GPU memory bandwidth
MS-S1 ████░░░░░░░░░░░░░░░░░░░░ 256 GB/s
Spark / GX10 ████░░░░░░░░░░░░░░░░░░░░ 273 GB/s
M4 Max 128 ████████░░░░░░░░░░░░░░░░ 546 GB/s
M3 Ultra 96 ████████████░░░░░░░░░░░░ 819 GB/s
5090 / 6000 ██████████████████████████ 1792 GB/s

The Minisforum and the Spark are the same machine class. Apple is two to three times the bus. A 5090 is seven times the bus and one quarter of the memory.

Two speeds, because there are two jobs

Decode is the next token in a chat. The model streams its own weights once per token. Speed tracks bandwidth.

Prefill is swallowing the prompt. That is matrix multiplies. NVIDIA compute shows up here. AMD’s 8060S does not.

Decode, dense 70B Q4 (~35 GB)

A 70B Q4 does not fit on a 5090. That is the first trap.

SystemFits?Typical tok/svs MS-S1
MS-S1 Maxyes5–81.0×
Spark / GX10yes8–18~1.5×
M4 Max 128 GByes12–16~2×
M3 Ultra 96 GByes25–32~4×
RTX 5090 32 GBno2–20 offloadedoften slower than the mini-PC
RTX PRO 6000 96 GByes~32~5×
4× RTX 5090yes, split80–150~15–20×
Dense 70B Q4 decode (midpoint)
MS-S1 ██░░░░░░░░░░░░░░░░░░░░ 6 tok/s
Spark / GX10 ████░░░░░░░░░░░░░░░░░░░░ 13 tok/s
M4 Max 128 ████░░░░░░░░░░░░░░░░░░░░ 14 tok/s
M3 Ultra 96 ████████░░░░░░░░░░░░░░░░ 28 tok/s
PRO 6000 █████████░░░░░░░░░░░░░░░ 32 tok/s
4× 5090 ██████████████████████████ 100 tok/s

Decode, 32B Q4 (~18 GB) — the fair race

Everything holds this one.

SystemTypical tok/svs MS-S1
MS-S1 Max12–181.0×
Spark / GX1015–22~1.2×
M4 Max 12828–40~2×
M3 Ultra 9640–55~3×
RTX 509040–55~3×
PRO 600045–70~4×

Same ranking as the bandwidth chart. That is not a coincidence.

Prefill

SystemPrompt ingestvs MS-S1
MS-S1 Max~300–400 tok/s1×
M4 Max~700–1.000~2–3×
M3 Ultra~1.000–1.400~3–4×
Spark / GX10~1.700–1.900~5×
5090 / PRO 6000thousands~10–20×

Paste a long PDF into the AMD box and you wait. Spark eats the prompt, then crawls on the reply at almost the same speed as AMD. Apple is the opposite: slower prefill than Spark, faster chat, because the bus is wider.

Chat speed follows bandwidth:

MS-S1 (256) → Spark (273) → M4 Max (546) → M3 Ultra (819) → 5090 (1792)

Prompt speed follows compute:

MS-S1 8060S → M4 Max → M3 Ultra → Spark GB10 → 5090 / PRO 6000

A typical turn — 2.000-token prompt, 300-token answer, dense 70B Q4 — lands roughly here:

SystemFirst tokenFinish answerTotal
MS-S1 Max5–7 s40–60 s~1 minute
Spark / GX101–2 s15–25 s~25 s
M4 Max 1282–3 s20–25 s~25 s
M3 Ultra 96~2 s10–12 s~13 s
PRO 6000<1 s~10 s~10 s
4× 5090<0.5 s2–4 s~4 s

70B at 512K context is a memory problem first

People say “70B with a 512K window” as if that were one more checkbox. It is not. The weights are the small part.

A Llama-3-70B-class model (80 layers, 8 KV heads, head dim 128) stores this much key/value cache:

2 × layers × kv_heads × head_dim × seq_len × bytes

That is 320 KiB per token in FP16. At 524.288 tokens (512K):

KV precisionCache alonePlus 70B Q4 weights (~40 GB)Plus 70B Q8 (~72 GB)
Q4 KV~40 GB~80 GB~112 GB
FP8 / Q8 KV~80 GB~120 GB~152 GB
FP16 KV~160 GB~200 GB~232 GB

Add 10–20% for paging waste, CUDA graphs, the OS, and the fact that 128 GB on the box is not 128 GB for the model.

70B + 512K: memory you must hold
Q4 weights + Q4 KV ████████░░░░░░░░░░░░░░░░ 80 GB
Q4 weights + FP8 KV ████████████░░░░░░░░░░░░ 120 GB
Q4 weights + FP16 KV ████████████████████░░░░ 200 GB
Q8 weights + FP16 KV ███████████████████████░ 232 GB

Who can actually load it:

SetupUsable memory70B Q4 + Q4 KV 512K70B Q4 + FP16 KV 512K
MS-S1 Max~96 GB VRAM of 128tight, maybeno
Spark / GX10128 GB unifiedtightno
M4 Max 128 GB128 GB unifiedtightno
M3 Ultra 96 GB96 GBborderlineno
RTX 509032 GBnono
PRO 600096 GBtight on Q4 KVno
2× PRO 6000192 GByesbarely, with KV quant
4× 5090128 GB splitonly if the runtime shards KVno as one pool
M3 Ultra 256 GB (discontinued)256 GByesyes

A 96–128 GB box can hold 70B Q4 at 512K only if you quantise the cache hard and accept the quality hit. Full-quality 512K on a 70B is a 200 GB problem. That is two workstation GPUs or a Mac Apple no longer sells.

And even when it fits, decode is not the short-context number. Attention cost grows with sequence length. At 512K, expect roughly 20–40% of the short-context tok/s, if the stack does not fall over first.

SystemShort-context 70B Q4Expected at 512K, if it fits
MS-S1 / Spark6–15 tok/s2–6 tok/s, if Q4 KV fits
M4 Max 12812–164–8
M3 Ultra 9625–326–12, only with aggressive KV quant
PRO 6000~328–15 with Q4 KV
2× PRO 600050–8015–30

I would not buy a €4k mini-PC for “70B at 512K”. I would buy it for 70B at 8K–32K, which is what most chat and coding loops actually use.

If you truly need 512K on a 70B-class dense model, start at 192 GB of GPU memory and a stack that can quantise KV. Or stop pretending and use a model that was built for long context.

Which is the other half of this post.

Qwen3.8-27B — the model you can actually run

Alibaba shipped Qwen3.8-27B in mid-August 2026. Hugging Face lists it as ~28B parameters with embeddings. The name is 27B. Dense, not MoE. Native vision. Hybrid thinking. Multi-token prediction. 256K context, 1M with YaRN.

It is a different architecture from Llama-70B, and that is why the 512K conversation changes.

The backbone is 16 blocks of three Gated DeltaNet layers plus one Gated Attention layer. Only 16 attention layers keep a KV cache, with 4 KV heads and head dim 256. The DeltaNet part is linear in sequence length. The cache is small.

2 × 16 × 4 × 256 = 32.768 elements per token

That is 64 KiB per token in FP16. About 2.0 GiB at 32K, which matches the ~2.5 GB people are measuring once you add overhead.

ContextFP16 KVFP8 KVQ4 KV
32K~2.5 GB~1.3 GB~0.7 GB
128K~10 GB~5 GB~2.5 GB
256K native~20 GB~10 GB~5 GB
512K (YaRN)~40 GB~20 GB~10 GB

Weights at 4 / 8 / 16 bit

From the Unsloth card:

PrecisionWeightsPlus 32KPlus 256K FP16 KVPlus 512K FP16 KV
4-bit / IQ4 / NVFP417–19 GB~20 GB~38 GB~58 GB
8-bit~31 GB~34 GB~51 GB~71 GB
BF16 / 16-bit~56 GB~59 GB~76 GB~96 GB

At 256K native context, Qwen3.8-27B lands here:

WeightsKV (256K)Total in GPU memory
4-bit ~18 GBQ4 KV ~5 GB … FP16 KV ~20 GB~23–38 GB
8-bit ~31 GBsame~41–51 GB
16-bit ~56 GBFP16 KV ~20 GB~76 GB

Every machine in the first table can run 4-bit at 32K. A 5090 can run 4-bit and 8-bit at moderate context, then runs out. BF16 plus a long window needs the 96–128 GB boxes.

Speed vs quality vs context

Three knobs. You do not get all three.

4-bit. Fastest. 17–19 GB. Fine for chat and most coding. You will feel it on hard reasoning and on tool-call JSON if you hunt for it. NVFP4 on Blackwell is the same size class with better recovery than a naive Q4 — Unsloth reports 92–97% top-1 versus BF16 and about 1.5× BF16 throughput.

8-bit. The quality default. Near-lossless for almost everything I care about. 31 GB. A 5090 can hold it until the context grows. The unified 96–128 GB boxes hold it with room for 256K.

16-bit. Reference. Fine-tuning, logit-sensitive work, or “I do not want to argue about quant”. 56 GB. Not a 5090. Comfortable on Spark, MS-S1, M3 Ultra, PRO 6000.

Quality vs speed, Qwen3.8-27B (relative, not a lab plot):

SlowerFaster
CleanerSpark BF16 · M3 Ultra Q8 · MS-S1 Q8PRO 6000 Q8 · 5090 NVFP4
LossierMS-S1 Q4Spark NVFP4 · M3 Ultra Q4

Measured and estimated decode, Q4, short context, batch 1:

SystemStacktok/svs MS-S1
MS-S1 Maxllama.cpp Vulkan, no MTP~10–121.0×
MS-S1 MaxVulkan + MTP=4up to 24.5 (AMD, Windows)~2×
Sparkllama.cpp Q4 + MTP24–30~2×
SparkSGLang + NVFP4 + MTP34–38, peak ~47~3×
M4 Max 128MLX / Metal~25–40~2–3×
M3 Ultra 96MLX / Metal~35–55~3–4×
RTX 5090llama.cpp Q4~60–90~6×
RTX 5090SGLang + NVFP4 + MTP~120–200~10×
PRO 6000same kernels, more room~80–150~8×

AMD’s own day-0 note is the 24.5 figure, on a GMKtec EVO X2, Vulkan, MTP=4. Community Q5 on Strix Halo without that stack is closer to 10 tok/s. Same chip, factor of two, depending on whether speculative decoding is actually on.

NVFP4 is Blackwell only. Spark, 5090, PRO 6000. Not AMD. Not Apple.

The thinking tax

Thinking is on by default (reasoning_effort: xhigh). A short coding answer often burns 2.000–8.000 thinking tokens before the first useful line.

System3.000 thinking tokens400 answer tokensYou wait
MS-S1, typical Q5~4–5 min~30 s~5 min
MS-S1, AMD MTP peak~2 min~16 s~2.5 min
Spark NVFP4~80 s~11 s~1.5 min
M3 Ultra Q4~70 s~9 s~1.3 min
5090 / PRO 600020–40 s3–5 s~30–45 s

Turn thinking to low or off for chat and every row gets several times faster. For agent work you want it on. Then raw tok/s is the product.

Long context on this model is the opposite of 70B. 256K native plus Q4 weights is about 38 GB. That is a Minisforum or a Spark or a Mac, not a science project. 512K with FP16 KV is about 58 GB in 4-bit. Still one box. Quality at 512K will depend more on YaRN than on whether you picked Q4 or Q8.

Qwen3.8-27B Q4 + FP16 KV versus 70B Q4 + FP16 KV, same context lengths:

32K 128K 256K 512K
Qwen 27B Q4 20 GB 28 GB 38 GB 58 GB
██ ███ ████ ██████
70B Llama Q4 50 GB 80 GB 120 GB 200 GB
█████ ████████ ████████████ ████████████████████

The first series is Qwen3.8-27B Q4 plus FP16 KV. The second is 70B Q4 plus FP16 KV. Same context lengths. Different decade of hardware.

What I would actually buy

For Qwen3.8-27B as a daily coder, thinking on:

  • Cheap and good enough: MS-S1 Max, and turn thinking down when you are iterating.
  • Same money, CUDA, faster prompts: ASUS GX10.
  • Quiet and fast enough that thinking does not hurt: M3 Ultra 96 GB.
  • Interactive with xhigh thinking: a 5090 if you stay at 4-bit and moderate context, a PRO 6000 if you want Q8, vision, and 256K at the same time.

For 70B chat at 8K–32K: M3 Ultra if you live in macOS, Spark if you need CUDA, MS-S1 if the budget is the point. A lone 5090 loses.

For 70B at 512K: do not start from the €4k mini-PC list. You are shopping 192 GB+ or a different model.

The jungle is the TOPS column. The path out is bandwidth, then whether the weights plus the KV cache fit, then whether you are measuring prefill or decode.

Everything else is a brochure.

All posts