Cascadia: Run Powerful AI Models on Intel Hardware

3 min read
Share this post
Cascadia and Intel logos side by side above the line "AI inference built for Intel silicon."

Today, we’re excited to announce Cascadia, a new open source runtime which pools the resources of Intel-powered machines, allowing them to run models larger than any single machine could.

Many organizations own hundreds or even thousands of commodity PCs. Yet, most of that compute sits idle while AI workflows are sent to the cloud.

Developed in collaboration with Intel, Cascadia turns these machines into a private AI network that you own. Run inference across everything from a single Intel-powered laptop to a fleet of AI PCs, or a dedicated box of Intel Arc Pro GPUs. All machines running one model, or many running several.

How Cascadia Works

LLMs come in varying sizes and capabilities. The bigger they are, the more RAM (or VRAM) they require to run. This usually means you’re limited to whatever can fit comfortably on your device. If you need a bigger model, you need more memory. Cascadia removes that single-device memory ceiling.

AI models are made up of layers, and Cascadia splits groups of these layers into shards. The shards are distributed across any number of Intel machines, and pool the resources of your existing PCs together to run larger models than any single machine could.

For example, you could split Llama 3.1 70B (~36GB in size, INT4 quantized) across a fleet of standard PCs.

diagram of a a sharded model running across computers in a netowrk style topograph connected by arrows

Sharded inference isn’t new (shoutout to Exo), but the techniques used are built on novel research specifically optimized for Intel silicon.

Cascadia’s sharding is built on top of Intel’s OpenVINO toolkit, and performs first-in-the-world pipeline-parallel sharded inference on Intel silicon. We found a number of interesting methods to make this both feasible and efficient, and we’ll be releasing our first research paper soon.

This unlocks a lot of new opportunities for businesses, developers, or self-hosting enthusiasts that own Intel hardware, including inference and agentic workflows with the largest models.

Run AI on Your Own Terms

Cascadia comes with the benefits of any local and distributed system, but also ones which are unique.

AI subscriptions and usage-based plans are convenient. However, you trade convenience for a number of downsides: the cost of tokens is unpredictable, enterprises and developers are forced into usage-based plans, and you lose data privacy.

AI labs have made strides towards better privacy and data retention policies, but they present serious risks to industries that have strict data-privacy standards (such as healthcare or finance).

Cascadia operates:

  • Fully offline or within your own network
  • On a range of open-weight and local models
  • Without billing proportional to usage

Cascadia is different because it is purpose-built to meet you where your hardware is today. It doesn’t matter if you’re a business with a fleet of laptops, or a solo developer running an Intel Arc Pro GPU rig.

Get Started with Cascadia

Cascadia’s runtime is open source (licensed under Apache 2.0) and is in public alpha. It uses the same executable and same OpenAI-compatible interface across any Intel hardware you own.

We support a range of different models, harnesses, and workflows. To get started, view the GitHub repository here.

For those interested in running Cascadia in their organization, we’re piloting our enterprise offering to select design partners. Book a demo with us here.

Share this post

Bring a fleet. Keep everything.

We stand up your first private agent in under a week, on your hardware, on your network. You keep all of it.

Free for individuals

Download

Apply to become a design partner

Request a Pilot