Distributed inference

Turn multiple devices into one pool of inference capacity. Aquaduck can split and place work across machines you already own.

Request a demo

What it does

Contribute idle computers, join them into a pool, and run jobs that would not fit on one device. The system schedules work across the fleet.

Why it matters

You scale with hardware you control instead of moving everything to a remote cloud by default.

Deep dive

Aquaduck connects devices over your existing wireless or wired network and presents them as one inference layer. Each machine can hold a complete model, letting the network route each request to the model best suited to it. Machines can also be grouped into pairs that split the layers of one model, while multiple pairs run different models or parallel copies.

That creates a flexible surface for model routing (choose the best model for the job), model chaining (pass one model’s result to the next), and pipeline parallelism (split a model over multiple machines for cooperative decoding). Instead of forcing every workload into one topology, the pool can match model placement to the job and the network available.

Where the gains come from

Distribution does not make a fixed model more intelligent by itself. The gain is that the combined memory and compute unlock better options: a larger model that would not fit on one machine, several specialized models, or multiple passes that evaluate and refine an answer.

Research supports each of those paths. Scaling‑law research found predictable improvements in language‑model loss as model size, data, and training compute increase, while Chinchilla later showed that a well‑trained 70B model could outperform undertrained models up to 530B parameters. More available memory expands both the size and set of models you can run.

Composition can produce a more direct quality gain. In the Mixture‑of‑Agents study, several open models generated candidates in parallel and an aggregator refined them over successive layers. The resulting system scored 65.1% on AlpacaEval 2.0 versus 57.5% for GPT‑4o in that evaluation. On the other hand, later work found that repeatedly sampling a strong model can sometimes beat mixing weaker ones. The practical takeaway is that chains should be designed around complementary models and measured on your workload.

Routing produces efficiency gains. RouteLLM learned when to send a prompt to a stronger or weaker model and reported more than 2x lower cost with little quality loss across public benchmarks. In a private device pool, the same idea can route standard, sensitive workloads to a fast, private local model and reserve a stronger remote model for prompts that benefit from it.

The network trade‑off

The baseline is a single model placed on a single machine. It has no inter‑device traffic during inference, but its memory and throughput stop at that machine. Distributed techniques trade some network communication for a higher ceiling.

TechniqueWhat is placed on each machineEffect on answer qualityMeasured quality resultNetwork trafficPrimary gain
Single machineOne complete model on one machineBaseline63.7% MMLU‑ProNoneSimplicity and low latency
Data parallelismA complete copy of the same model on each machineSame model quality63.7% MMLU‑Pro, +0 from baselineLowHigher throughput
Pipeline parallelismConsecutive groups of a model’s layers on each machineUnlocks a stronger model69.0% MMLU‑Pro, +5.3ModerateLarger model capacity
Tensor parallelismA portion of each layer’s weights and operations on each machineUnlocks a stronger model69.0% MMLU‑Pro, +5.3HighLarger models with parallel computation
Model routingA different complete model on each machineApproaches strong‑model quality more efficiently95% quality of strong model, 3.66x lower costLowQuality and cost efficiency
Model chainingA complete model or agent role on each machineCan improve quality through synthesis and refinement13.2% higher win rate than strong modelModerateHigher composite intelligence

Pipeline parallelism is an especially useful technique. A model is a set of layers. In pipeline parallelism, those layers are divided across machines. The first machine takes a request and computes several consecutive layers, sends the result to the next machine, and continues with another request. More machines add boundaries and latency, so the fastest design is usually the fewest splits required to fit the model and keep the machines balanced.

Tensor parallelism cuts each operation across devices instead. This is effective on high‑bandwidth GPU fabrics but far more sensitive to latency and bandwidth than pipeline parallelism. Data parallelism sits at the other extreme: every participating machine holds a copy of the whole model, but machines work independently and scale throughput cleanly.

Distributed inference over less specialized networks has been demonstrated in research. Petals ran 70B- and 176B‑parameter models across consumer‑grade, geographically distributed machines and reported interactive generation up to 10x faster than RAM offloading. In a datacenter setting, AlpaServe showed that model placement and parallelism can also improve fleet utilization: on production traces it handled up to 10x the request rate while keeping more than 99% of requests within their latency targets.

The useful pattern is composite: keep complete models on individual machines when that minimizes communication; use routing or chaining when multiple models add value; and split a model across machines when pooled memory matters more than avoiding the network hop. The result is a configurable inference layer that can trade latency, throughput, model capacity, and response quality according to the workload.

Glossary

The output of one model becomes input or context for another, often to decompose, verify, or refine work.
A policy chooses which model should handle each request based on factors such as task, quality, latency, privacy, or available capacity.
Consecutive layer groups run on different devices, so that an oversized model can run on a collective of smaller machines.
Individual operations inside each layer are split across devices and recombined with collective communication.
Complete model replicas process different requests or batches simultaneously. It raises throughput but does not let an oversized model fit.
The total work completed over time, usually requests or tokens per second.
The time from submitting a request until its first token or complete response arrives, reported as time to first token or end‑to‑end latency.