Engineering notes on distributed inference
Latest blog posts

Announcement
Yes, Chef: Delegate Tasks to Local Models with Claude and Codex
Allow Claude Code and Codex to hand off tasks to local subagents (cooks), and save on token costs.
Read more
Inside Our Distributed LLM Inference Research for Intel PCs
Exploring the optimization techniques for the sharded inference behind Cascadia.
Read more
Cascadia: Run Powerful AI Models on Intel Hardware
Cascadia: Shard AI models across the Intel hardware you own.
Read more
Research
The reason to stop buying new hardware
Three years of public data say two curves are real, and together they mean most AI infrastructure being bought today is a deflating asset acquired at scarcity prices. Every number links to its source; every model is on the page, not behind it.
Read more

