Community Inference

Be part of a global pool of community capacity. Run local models, contribute spare capacity, and use models running on contributor devices when you need more.

Download

What it does

Download Aquaduck to use the desktop agent, run models locally, and connect any device you own. You can contribute available capacity to the Community Inference network, use capacity shared by people around the world, or combine devices with one other person in a team network.

Why it matters

Useful compute is distributed across millions of personal devices, but most of it sits idle. Community Inference lets people share inference capacity with other AI builders around the world.

Deep dive

The Aquaduck Community Inference network is a shared inference network maintained by Aquaduck and powered by its participants. Members can contribute capacity from their devices and use capacity contributed by others. A person with an available machine can help serve a request for someone who needs a model or more compute, turning otherwise idle hardware into part of a global inference layer.

With the Aquaduck app, you can start locally, opt into the Community Inference network, or create a team network with one other person. Your agent and API clients use the same Aquaduck account and API keys across those environments.

Three ways to use Aquaduck

NetworkWho provides capacityWho can use itBest suited for
LocalYour current deviceYouPrivate chat, local development, and offline inference
CommunityParticipating devices around the worldAquaduck Community Inference membersShared access to models and capacity beyond one machine
TeamYour devices and an invited member’s devicesYour two accounts and authorized API clientsCollaborative capacity, stable development endpoints, and workloads you want to keep within a known device group

Any device you own can join the Community Inference network or your team network. A laptop can run the desktop agent and contribute model capacity, and additional machines can expand the memory and throughput available to the network you select.

Share capacity and use capacity

Community Inference members can participate in both directions. When you share capacity, Aquaduck can schedule eligible inference work on your connected device. When you need capacity, the network can route your request to an available participant running an approved model.

You can use the Community Inference network from the desktop agent or through its OpenAI‑compatible API:

https://api.aquaduck.ai

Authenticate with your Aquaduck API key, select an available model, and send requests from an application or agent. Capacity changes as devices join, leave, or become busy, so Aquaduck coordinates placement, queues, and failover across the available pool.

The result is cooperative infrastructure. Participants do not need to own every model or keep every device online themselves. Everyone contributes what they can and draws from the network when they need help.

Models are maintained by Aquaduck

Aquaduck maintains the model catalog available on the Community Inference network. This gives participants a consistent set of model names, formats, configurations, and expected capabilities. If you want a model that is not yet available, you can request that it be added to the catalog.

Aquaduck supports MLX model packages and GGUF files through its MLX and llama.cpp backends. It automatically selects the best supported runtime for each machine, so participants can run the right model format without choosing or configuring the inference backend themselves.

Model weights are attested against Aquaduck‑approved artifacts before they can serve network requests.

For models designed to run across several machines, Aquaduck packages the required shards in advance. Participants can download the correct shard for each device instead of splitting and configuring the model themselves. The network can then assemble those devices into a pipeline and use their combined memory to run a model that may not fit on either machine alone.

Create a team network

You can also create a team network and invite one additional account. That person can be a friend, colleague, collaborator, or anyone else you choose. Both members can connect every device they’re signed into, giving the team access to the combined capacity of those machines.

Claim your name

Each team network receives its own subdomain:

https://{name}.aquaduck.ai

The subdomain becomes an OpenAI‑compatible endpoint that you can call from development applications, external agents, automations, or production software. The same API keys used with the Community Inference network work with the team endpoint. If you intend to use it in production, connect enough available devices to support the traffic, model size, and reliability your application requires.

With two members and multiple devices, a team can keep complete models on individual machines or split a larger model across devices. Pre‑packaged model shards make it easier to run oversized models together for more model capability than you could run on one device alone.

Access‑controlled and zero data retention

A team network limits participation to its two member accounts and authorized API clients. Requests are coordinated through Aquaduck‑managed network infrastructure, with no logging or retention of prompt content. Aquaduck retains limited operational metadata, such as token counts and request status, for debugging and service operation.

One key, multiple endpoints

DestinationEndpointHow to accessCapacity source
Community Inference networkapi.aquaduck.aiYour Aquaduck API keyParticipating devices across the community
Team network{name}.aquaduck.aiYour Aquaduck API keyDevices connected by you and your invited member
Local runtimeOften localhost:{port}, or the endpoint shown in the desktop appNo key neededThe current device

Your Aquaduck API keys work across the Community Inference and team endpoints. Applications can keep the same request format while you choose which network should serve each workload.

Contribution incentives

Sharing capacity helps make more models and throughput available to everyone. Aquaduck may experiment with bonuses or other incentives for contributors over time. Watch for future program announcements.

Glossary

The global inference network maintained by Aquaduck and powered by connected member devices sharing capacity.
An access‑controlled network containing your account, one invited account, and the devices connected by both members.
The memory and compute a device can make available for model inference.
Verification that a serving device is using approved model weights and configuration.
A pre‑packaged portion of a model assigned to one device for participating in model splitting, or pipeline parallelism, to run an oversized model.
A policy under which model request content is not stored after it is processed.