Companies have been releasing larger and larger open-weight models, with some, such as Thinking Machines’ Inkling (975B), close to, or even past, the 1T parameter mark.
Models of this size are typically trained and run on data-center hardware and are far too large to run on a single consumer PC, but by pooling the computing power of multiple machines, it becomes possible to run them. That being said, getting them to run efficiently across machines remains a big challenge.
In collaboration with Intel, we put this to the test and ran Inkling across eleven Intel Core Ultra PCs using Cascadia, our distributed inference runtime built on OpenVINO.
TLDR; peak 60.29 tok/s of aggregate decode throughput with 88 concurrent requests (streams), with some improvements to be made for single streams and time to first token (TTFT).
Getting things operational
The machines we used were eleven Intel Core Ultra X7s, each with an Arc B390 iGPU and 64GB RAM. This provided a total of 704GB RAM, and each machine was connected to one another over Gigabit Ethernet.
| Component | Configuration |
|---|---|
| Number of Machines | 11 |
| CPU | Intel Core Ultra X7 358H |
| GPU | Intel Arc B390 integrated graphics |
| Memory | 64GB per machine |
| Aggregate memory | 704GB nominal |
| Network | Gigabit Ethernet |
| Runtime | Cascadia (Rust + OpenVINO) |
To get Inkling operational across a fleet of machines, we first had to account for the model’s architecture.
Inkling is a sparse Mixture-of-Experts (MoE) model with 41B active parameters per token. It has 66 decoder layers, which we split into six-layer shards and distributed across the machines (11 in total, so one shard per machine). The machines were numbered 0–10.
Machine 0 contains the embeddings and the first six layers of the model (two dense, four sparse).
Machines 1–9 each hold six more sparse layers (layers 6–59). Each sparse layer contains 256 routed experts, of which 6 are selected for each token, as well as two shared experts that run for every token.
After the final six layers are executed on Machine 10, the next token in the sequence is selected. We specifically used greedy decoding as our decoding strategy (selecting the highest-scoring token on each pass).

Running Inkling on each machine
To run each shard efficiently we built a custom MoE engine into Cascadia for the Core Ultra chipset.
Each machine splits work between the CPU and iGPU: the CPU handles attention, routing (of each token to selected experts), and tracking each request. The iGPU handles the computational tasks such as the expert calculations and compressed projections.
Cascadia, as mentioned, is built on top of OpenVINO, and so also makes use of OpenVINO’s fused MoE operations, but still allows for Cascadia to control routing.
The expert weights are INT4 quantized, while attention projections and the output head use INT8 weights. For other aspects of the pipeline, we aimed for higher precision (e.g. FP16 for fused-expert calculations, and FP32 for the residual stream passed through the pipeline).
Something interesting we did was with the two dense layers on Machine 0. Rather than running them as standard matrix operations, we split each one into eight expert-sized slices that always run, which lets them use the same fused operations as the sparse layers. This brought the time per dense layer down from around 8.1ms to 4.5ms (~45% faster).
The results
We tested total throughput on different numbers of streams (concurrent requests).
Each request that we tested generated 128 tokens using greedy decoding, with Cascadia's default context length of 1,024 tokens.
| Concurrent streams | Aggregate decode throughput | Whole-phase throughput | Median TTFT |
|---|---|---|---|
| 1 | 7.96 tok/s | 6.98 tok/s | 2.18 s |
| 2 | 4.53 tok/s | 4.37 tok/s | 2.06 s |
| 8 | 18.60 tok/s | 16.86 tok/s | 4.55 s |
| 15 | 24.60 tok/s | 22.12 tok/s | 6.05 s |
| 32 | 38.31 tok/s | 32.48 tok/s | 13.44 s |
| 64 | 53.60 tok/s | 42.41 tok/s | 25.10 s |
| 88 | 60.29 tok/s | 46.87 tok/s | 34.61 s |
| 128 | 50.10 tok/s | 41.61 tok/s | 55.11 s |
| 176 | 57.72 tok/s | 45.24 tok/s | 76.83 s |
Aside from a dip between 1–2 streams, aggregate throughput increases with concurrency, peaking at ~60.29 tok/s at 88 streams. Additional streams saw throughput drop noticeably before climbing once again, which could be due to the scheduler re-distributing work among groups of streams.
We saw only ~5–10 tok/s on single stream throughput, however, which makes this setup much more viable for multi-user settings, or background tasks. These results show that it is entirely possible to serve larger sparse MoE models on distributed client hardware (such as Panther Lake PCs).
It’s worth noting that the point of this research isn’t necessarily to get the best single-user or single machine performance. Our first goal was to solve the technical challenges associated with sharding and serving such a large model, and for Cascadia to act as a serving network for a model of this size for organizations that own a fleet of machines.
Try Cascadia and Inkling now
Inkling support is open source and being upstreamed into Cascadia. You can clone the repo to try it out and give us a star to support our work here.
For the full methodology and benchmarks, you can read the paper here.
If you own a fleet of Intel machines, you can use Cascadia to run powerful AI models across your devices, including ones that would be too large to fit on a single machine. You can book a call with us and talk through how we can help with your setup.
Related articles

Research
Where NPUs Fit into LLM Inference
Most modern laptops now ship with an NPU, but few people use it for running LLMs.
Read more
Research
Building a Permissioned, Peer-to-Peer AI Inference Network
How a permissioned peer-to-peer system can turn a fleet of PCs into a single, on-prem inference endpoint.
Read more
Research
Running LLMs on Intel Hardware: OpenVINO vs vLLM vs llama.cpp
Comparing various runtimes for Intel hardware.
Read more
Research
Where We're Going, We Don't Need Data Centers.
Data centers are being rapidly constructed for the purpose of training and serving AI models, but the demand for new data centers is largely predicated on the growth of AI inference.
Read more
Announcement
Cost-Efficient AI on the Hardware You Already Own
Cut your AI spend by putting the CPUs and NPUs you own to work.
Read more
Announcement
Yes, Chef: Delegate Tasks to Local Models with Claude and Codex
Allow Claude Code and Codex to hand off tasks to local subagents (cooks), and save on token costs.
Read more


