A 421M classifier in front of the judge
claudeinsights coaches me on my own Claude Code sessions by sending each rough one to a model. Every session used to cost a full claude -p call. Now a 421M-parameter encoder called Laya scores all of them locally first, and only the rough ones go to Opus 5.5 for quotes and advice. Go downloads and caches the checkpoint through go-huggingface, sharing Python's Hugging Face cache without a single extra byte, and its pure-Go tokenizer matches Hugging Face's Rust tokenizer id for id on Laya's vocabulary, except for one lstrip flag. The model itself still runs in PyTorch, and skipping one random-weight initialisation took its load from 46 seconds to 4.
Read