Tag

#local-llm

2 posts ·all posts

A 421M classifier in front of the judge

claudeinsights coaches me on my own Claude Code sessions by sending each rough one to a model. Every session used to cost a full claude -p call. Now a 421M-parameter encoder called Laya scores all of them locally first, and only the rough ones go to Opus 5.5 for quotes and advice. Go downloads and caches the checkpoint through go-huggingface, sharing Python's Hugging Face cache without a single extra byte, and its pure-Go tokenizer matches Hugging Face's Rust tokenizer id for id on Laya's vocabulary, except for one lstrip flag. The model itself still runs in PyTorch, and skipping one random-weight initialisation took its load from 46 seconds to 4.

Read

Local AI and the jungle of incomparable specs

AMD, Apple and NVIDIA all sell a 128 GB local-AI box. The TOPS numbers are not comparable. Bandwidth is. Netherlands prices, real decode speeds, what 70B at 512K actually costs in memory, and how Qwen3.8-27B changes the 4/8/16-bit trade-off.

Read