MSO & Telecom · August 8, 2026
Taalas physically embodied Llama 3.1 8B in silicon, reached roughly 17,000 tokens per second for one user, and was acquired by AMD before I finished writing about it.
Read →Local Inference · August 4, 2026
ROCm worked on the first try. Vulkan was substantially faster, and the hardest failures belonged to everything around the card.
Read →Local Inference · August 3, 2026
Five mapped expert layers made room for speculative decoding, a correction tier, and near-million-token context.
Read →Local Inference · August 2026
The origin story, hardware escalation, runtime experiments, failures, and corrections behind a private local AI system.
Read →