Distributed inference
Turn multiple devices into one pool of inference capacity. Aquaduck can split and place work across machines you already own.
Request a demoWhat it does
Contribute idle computers, join them into a pool, and run jobs that would not fit on one device. The system schedules work across the fleet.
Why it matters
You scale with hardware you control instead of moving everything to a remote cloud by default.
Deep dive
Aquaduck connects devices over your existing wireless or wired network and presents them as one inference layer. Each machine can hold a complete model, letting the network route each request to the model best suited to it. Machines can also be grouped into pairs that split the layers of one model, while multiple pairs run different models or parallel copies.
That creates a flexible surface for model routing (choose the best model for the job), model chaining (pass one model’s result to the next), and pipeline parallelism (split a model over multiple machines for cooperative decoding). Instead of forcing every workload into one topology, the pool can match model placement to the job and the network available.
Where the gains come from
Distribution does not make a fixed model more intelligent by itself. The gain is that the combined memory and compute unlock better options: a larger model that would not fit on one machine, several specialized models, or multiple passes that evaluate and refine an answer.
Research supports each of those paths. Scaling‑law research found predictable improvements in language‑model loss as model size, data, and training compute increase, while Chinchilla later showed that a well‑trained 70B model could outperform undertrained models up to 530B parameters. More available memory expands both the size and set of models you can run.
Composition can produce a more direct quality gain. In the Mixture‑of‑Agents study, several open models generated candidates in parallel and an aggregator refined them over successive layers. The resulting system scored 65.1% on AlpacaEval 2.0 versus 57.5% for GPT‑4o in that evaluation. On the other hand, later work found that repeatedly sampling a strong model can sometimes beat mixing weaker ones. The practical takeaway is that chains should be designed around complementary models and measured on your workload.
Routing produces efficiency gains. RouteLLM learned when to send a prompt to a stronger or weaker model and reported more than 2x lower cost with little quality loss across public benchmarks. In a private device pool, the same idea can route standard, sensitive workloads to a fast, private local model and reserve a stronger remote model for prompts that benefit from it.
The network trade‑off
The baseline is a single model placed on a single machine. It has no inter‑device traffic during inference, but its memory and throughput stop at that machine. Distributed techniques trade some network communication for a higher ceiling.
| Technique | What is placed on each machine | Effect on answer quality | Measured quality result | Network traffic | Primary gain |
|---|---|---|---|---|---|
| Single machine | One complete model on one machine | Baseline | 63.7% MMLU‑Pro | None | Simplicity and low latency |
| Data parallelism | A complete copy of the same model on each machine | Same model quality | 63.7% MMLU‑Pro, +0 from baseline | Low | Higher throughput |
| Pipeline parallelism | Consecutive groups of a model’s layers on each machine | Unlocks a stronger model | 69.0% MMLU‑Pro, +5.3 | Moderate | Larger model capacity |
| Tensor parallelism | A portion of each layer’s weights and operations on each machine | Unlocks a stronger model | 69.0% MMLU‑Pro, +5.3 | High | Larger models with parallel computation |
| Model routing | A different complete model on each machine | Approaches strong‑model quality more efficiently | 95% quality of strong model, 3.66x lower cost | Low | Quality and cost efficiency |
| Model chaining | A complete model or agent role on each machine | Can improve quality through synthesis and refinement | 13.2% higher win rate than strong model | Moderate | Higher composite intelligence |
Pipeline parallelism is an especially useful technique. A model is a set of layers. In pipeline parallelism, those layers are divided across machines. The first machine takes a request and computes several consecutive layers, sends the result to the next machine, and continues with another request. More machines add boundaries and latency, so the fastest design is usually the fewest splits required to fit the model and keep the machines balanced.
Tensor parallelism cuts each operation across devices instead. This is effective on high‑bandwidth GPU fabrics but far more sensitive to latency and bandwidth than pipeline parallelism. Data parallelism sits at the other extreme: every participating machine holds a copy of the whole model, but machines work independently and scale throughput cleanly.
Distributed inference over less specialized networks has been demonstrated in research. Petals ran 70B- and 176B‑parameter models across consumer‑grade, geographically distributed machines and reported interactive generation up to 10x faster than RAM offloading. In a datacenter setting, AlpaServe showed that model placement and parallelism can also improve fleet utilization: on production traces it handled up to 10x the request rate while keeping more than 99% of requests within their latency targets.
The useful pattern is composite: keep complete models on individual machines when that minimizes communication; use routing or chaining when multiple models add value; and split a model across machines when pooled memory matters more than avoiding the network hop. The result is a configurable inference layer that can trade latency, throughput, model capacity, and response quality according to the workload.