Building a Permissioned, Peer-to-Peer AI Inference Network

By Jack Smith

7 min read
Share this post
Cascadia banner reading "Building a local peer-to-peer AI inference system", beside a photo of an office building at night with scattered lit windows.

Solutions for serving AI models are usually built with servers or datacenter hardware in mind. This assumes access to powerful, dedicated AI hardware, that organizations may not own.

Peer-to-peer (P2P) networks are an alternative to this. Devices that serve and request AI inference are independent of each other, and don’t necessarily need to have the same hardware. They could be desktop workstations, laptops, server racks, or even dedicated GPU rigs.

For organizations, a permissioned P2P network can allow teams to run local AI models and pool the computing power of their commodity machines, without the need for dedicated servers.

This is the problem we are solving with Cascadia. In a permissioned P2P network, machines might have varying hardware capabilities, or go offline at a moment’s notice.

This post focuses on how we built a permissioned P2P network, but many of the principles apply to public networks, too.

Determining model structure

When serving AI models on a P2P network, you can serve inference in a few different ways: by replicating a model on each device, or by splitting a larger, more capable model across multiple devices.

Replicating a model across devices is useful for load-balancing frequent requests. In this case, however, a model performs best when it fits into the available memory (either RAM or VRAM) of each device. While a good solution for throughput, this limits the available models to ones which are much smaller in size (and often capability).

Where P2P inference becomes more compelling is in its ability to split a large model into groups of layers (shards) and distribute these across devices, enabling multiple devices to work together to run inference on a larger model that wouldn’t fit on any individually.

Each request to the model must pass through each layer (and therefore each worker) to produce an output. This process is called pipeline-parallel sharded inference. We’ve covered this mechanism in detail in one of our research papers here.

Once a model is effectively replicated or sharded, the next challenge is fulfilling performance, availability, and security needs for those who wish to provide and request inference.

Sharding is half of the problem

Datacenter setups such as NVIDIA Dynamo and llm-d typically assume a dedicated serving cluster, with an orchestration layer which routes and schedules requests across it.

A P2P inference network doesn’t have this layer, as a request could come from any machine. This means that they have different problems to address, for example:

  • Network latency: in sharded setups, requests must pass through every layer of a model, (and therefore each machine) for each token generated. This adds a number of additional network hops.
  • Node uptime/availability: commodity machines may come online or go offline arbitrarily; for example, if an employee updates their laptop.
  • State and continuity: If and when a node joins or leaves the network, pending requests need to be shifted among the other devices on the network.
  • Trust and security: considering how a new device/user should be admitted to the cluster, and also proving which node produced a response.

To turn a fleet of client machines into a dependable LLM serving system, a P2P network must be architected to address these problems.

Breaking down Cascadia’s architecture

These constraints mentioned above led us to create Cascadia with a few different design decisions in mind, notably:

  • No scheduler/routing “tier”, i.e., any node can do any task
  • Support for both model replication and sharding
  • Explicit, cryptographic trust mechanisms
  • Intel-native support and compatibility (using OpenVINO)

The ultimate goal is pooling the available compute from a set of Intel-based machines into a single, accessible endpoint for anyone on the network to use.

The anatomy of a request

A request sent on Cascadia can enter at any node, and the entry node will choose the best peer on the network to route the request to (given the capability available and current load). This could even be the current node itself.

If the request needs to pass to another node, the request is forwarded using an encrypted libp2p QUIC stream. The target node executes the inference request and returns the result to the user.

In a sharded setup, the request flow is a little different.

When a request enters the network, the entry node can access the topology of the network and forward the request to the first node in the shard chain (referred to as rank 0). This worker coordinates token generation and eventually streams the response back to the original user.

What happens when a node goes offline?

If a node in Cascadia goes offline unexpectedly, the standard approach is to route around any issues using redundancy where possible.

In a replicated setup, if a request is sent to a machine which fails, it is simply re-routed to another machine.

In a sharded setup, things are a little more complicated. If a shard unexpectedly fails, it can break the entire request chain, as all of the shards are required to produce a response. In this case, it’s better to have redundancies/duplicated shard chains.

For example, if an organization owns 10 desktop computers, creating a 10-shard chain risks requests failing if one of those computers becomes unavailable. Creating two, 5-shard chains, means that if one shard is lost, it can access the equivalent shard from another chain.

Trusting peers

In a traditional client-server setup, a network administrator would typically only need to admit a new node into the cluster boundary.

Users from an organization could use standard authorization methods such as an API key to access inference. This works, because inference is provided from a trusted server (or group of servers) in a trusted location. The inference providers aren’t employee laptops, constantly moving around, changing networks and location.

In a P2P system like Cascadia, each node is both a user and a server of inference. Machines could be spread across insecure public networks, home networks, or an organization’s different offices.

In a situation like this, having an API key is not enough. Nodes are interacting with other nodes, not a centralized server, and therefore need a way to verify the nodes they connect to are approved peers.

To solve this, each node generates a cryptographic keypair, and Cascadia operates a certificate authority (CA) which approves the public key, and issues an admission certificate to the node.

This means that each node can verify the identity of another node on the network without needing to check with a centralized source of truth each time. Using keypairs also means that there is no reliance on physical location or network boundary.

Nodes use their keypairs to sign responses, allowing the original caller of inference to verify the response came from somewhere within the network.

What does this unlock? Why go P2P?

One of the main benefits of running a P2P network is the re-use of hardware.

By allowing consumer-grade machines to connect to one another and provide inference, it replaces the need for a powerful (and often expensive) dedicated server for teams looking to run local, on-prem AI.

Likewise, it makes use of existing assets a company owns. If you have 10 Intel AI PCs you can turn into a serving network, this often beats the capital expenditure required to purchase dedicated AI hardware, especially given the prices of such hardware (at the time of writing).

It is also a system that is flexible and it can grow incrementally. If ten new employees are hired and ten desktop PCs are purchased, these increase the computing power you have (and the models you can run) while still being useful end-user machines in their own right.

Try it and learn more

Cascadia is open source. You can shard models and run them across devices, configure your own network topology, and expose an OpenAI-API-compatible endpoint for inference requests.

Check out the GitHub repository and star it to support our work.

The mesh layer described in this post, including features such as the certificate authority and fleet management are part of our enterprise plan.

If you’d like to learn more about how Cascadia works, read our full research paper here.

Share this post

Bring a fleet. Keep everything.

We stand up your first private agent in under a week, on your hardware, on your network. You keep all of it.

Free for individuals

Download

Apply to become a design partner

Request a Pilot