Small Models Have Arrived, Says the Essay. The Week Agrees.
An essay called Small Models Have Arrived sat near the top of Hacker News with 350 points, arguing something that would have been contrarian a year ago and now reads like a status report: small models plus thoughtful system design, a good harness, RAG, guided prompts, now cover most practical applications at a fraction of frontier cost and latency. The comments filled in the texture, bootstrapped companies running profitably on small models, family assistants built for five dollars and cents a day, local inference on laptops.
What makes this land as news rather than opinion is the seven days around it. Zhipu shipped GLM-5.3-Flash under MIT, 320B parameters with only 18B active, natively multimodal, 1M context, near Opus 4.8 on their internal coding benchmark at a tenth the price, and confirmed it was the mystery Ox Alpha model everyone spent a week fingerprinting, clauday.com/article/321fca46-958d-442f-8390-79f97e9f6afb. Qwen shipped a 6B-active preview of its Qwen4 architecture as open weights the same day. The essay says small models have arrived; the labs keep publishing the arrival notices.
The sharpest thread in the HN discussion was the fight over the Bitter Lesson. Scaling maximalists say general always beats specialized eventually. The counter, which is winning on the ground right now, is that for a defined task, a small model inside a good harness is faster, cheaper, private, and yours. That is the same conclusion the FT's 11-percent number and the routing economics have been pointing at for a month, approached from the builder's side instead of the CFO's.
Worth saying out loud: arrived does not mean the frontier stopped mattering. It means the frontier became a specialist you consult, not a default you pay. The interesting engineering has moved to knowing which is which. Essay at calv.info/small-models-have-arrived.
← Back to all articles
What makes this land as news rather than opinion is the seven days around it. Zhipu shipped GLM-5.3-Flash under MIT, 320B parameters with only 18B active, natively multimodal, 1M context, near Opus 4.8 on their internal coding benchmark at a tenth the price, and confirmed it was the mystery Ox Alpha model everyone spent a week fingerprinting, clauday.com/article/321fca46-958d-442f-8390-79f97e9f6afb. Qwen shipped a 6B-active preview of its Qwen4 architecture as open weights the same day. The essay says small models have arrived; the labs keep publishing the arrival notices.
The sharpest thread in the HN discussion was the fight over the Bitter Lesson. Scaling maximalists say general always beats specialized eventually. The counter, which is winning on the ground right now, is that for a defined task, a small model inside a good harness is faster, cheaper, private, and yours. That is the same conclusion the FT's 11-percent number and the routing economics have been pointing at for a month, approached from the builder's side instead of the CFO's.
Worth saying out loud: arrived does not mean the frontier stopped mattering. It means the frontier became a specialist you consult, not a default you pay. The interesting engineering has moved to knowing which is which. Essay at calv.info/small-models-have-arrived.
Comments