Inference API
One API for local, hybrid, and distributed inference. Call models the same way whether they run on your desk, across your devices, or in the cloud.
Request a demoWhat it does
A single inference interface for prompts, routing, and runtime choice. Swap where a model runs without rewriting the application.
Why it matters
Team members keep one integration. Capacity and locality can change underneath as needs grow.
Deep dive
An inference API is a shared service that provides callers with access to AI models without requiring them to manage the underlying hardware infrastructure. Aquaduck’s inference API gives team members access to AI models through one team endpoint, while connected team devices can contribute the capacity that serves those requests. People and applications do not need to know which device holds a model, whether several devices are working together, or where a request will run.
Why use an inference API?
In a network of independent model hosts, every caller of a model needs to know where it runs, how to connect to it, and what to do when that machine is unavailable or busy. An inference API keeps those details behind one stable interface, so customers can use the network without managing the machines and runtimes underneath it.
With an inference API, you can:
- connect any agent or application that supports an OpenAI‑compatible API
- give the entire team one shared endpoint
- issue a separate API key to each member, managed individually or by an administrator
- accept requests from anyone authorized to use the network
- serve requests from approved devices connected to the network
- move inference between one machine, pooled devices, and approved cloud runtimes
- selectively route to cloud capacity when local capacity is unavailable or an organization‑defined rule requires it
- change models or capacity without changing every application that uses them
Behind the endpoint, Aquaduck handles authentication, access gating, model routing, capacity‑based cloud routing, model configuration, integrity checks, job queuing, streaming, and failover. Callers send a standard request and receive a standard response while Aquaduck decides how to serve it.
One contract, multiple execution paths
Aquaduck exposes one OpenAI‑compatible interface in front of the network’s available inference capacity. Every team member uses the same endpoint with their own API key. Any compatible agent or application can connect by changing its base URL and supplying that key.
When a request arrives, Aquaduck verifies access, checks the requested model and configuration, selects an eligible runtime, and places the job in the appropriate queue. It streams the result through the same connection while handling failures and recovery. The interface stays the same whether one machine, several cooperating devices, or an approved cloud runtime serves the request.
Why one interface is better than several
Even when inference APIs look similar, they can differ in authentication, model identifiers, message formats, streaming events, tool definitions, structured output, token accounting, rate limits, error codes, and retry behavior. A unified API creates one control point for the entire inference path.
| Concern | Multiple inference APIs | One Aquaduck inference API |
|---|---|---|
| Application code | Provider branches and adapters | One request and response contract |
| Team access | Separate endpoints and credentials for each runtime | One team endpoint with individual API keys |
| Key management | Managed separately across providers | Managed by each member or centrally by an administrator |
| Runtime changes | Application changes and redeployment | Infrastructure configuration changes |
| Routing | Hard‑coded into application logic | Driven by capacity and organization‑defined rules |
| Gating | Repeated permission logic at each integration | Access policies applied before work is scheduled |
| Job queues | Separate queues and capacity rules | One scheduling layer across available capacity |
| Failover | Custom fallback logic in every application | Alternate eligible runtimes selected by the network |
| Model integrity | Verified independently for each runtime | Model identity and configuration checked centrally |
| Streaming | Different event formats and edge cases | One stream format for the application |
| Errors and retries | Provider‑specific handling | One normalized failure path |
| Observability | Metrics split across systems | Requests measured at one gateway |
| Privacy policy | Enforced separately in every integration | Applied at the inference layer |
Routing becomes infrastructure
Agents and applications connect to the inference API while Aquaduck coordinates the available models and machines. The team shares one endpoint, each member uses a separate API key managed individually or by an administrator, and every connected device can add serving capacity to the network.
A model can move from one laptop to a larger machine, be split across a pair of devices, or become part of a routed private network. The endpoint does not need to move with it. The team keeps one integration while the inference layer gains more models, more capacity, and more ways to place work.