Inference API

One API for local, hybrid, and distributed inference. Call models the same way whether they run on your desk, across your devices, or in the cloud.

Request a demo

What it does

A single inference interface for prompts, routing, and runtime choice. Swap where a model runs without rewriting the application.

Why it matters

Team members keep one integration. Capacity and locality can change underneath as needs grow.

Deep dive

An inference API is a shared service that provides callers with access to AI models without requiring them to manage the underlying hardware infrastructure. Aquaduck’s inference API gives team members access to AI models through one team endpoint, while connected team devices can contribute the capacity that serves those requests. People and applications do not need to know which device holds a model, whether several devices are working together, or where a request will run.

Why use an inference API?

In a network of independent model hosts, every caller of a model needs to know where it runs, how to connect to it, and what to do when that machine is unavailable or busy. An inference API keeps those details behind one stable interface, so customers can use the network without managing the machines and runtimes underneath it.

With an inference API, you can:

  • connect any agent or application that supports an OpenAI‑compatible API
  • give the entire team one shared endpoint
  • issue a separate API key to each member, managed individually or by an administrator
  • accept requests from anyone authorized to use the network
  • serve requests from approved devices connected to the network
  • move inference between one machine, pooled devices, and approved cloud runtimes
  • selectively route to cloud capacity when local capacity is unavailable or an organization‑defined rule requires it
  • change models or capacity without changing every application that uses them

Behind the endpoint, Aquaduck handles authentication, access gating, model routing, capacity‑based cloud routing, model configuration, integrity checks, job queuing, streaming, and failover. Callers send a standard request and receive a standard response while Aquaduck decides how to serve it.

One contract, multiple execution paths

Aquaduck exposes one OpenAI‑compatible interface in front of the network’s available inference capacity. Every team member uses the same endpoint with their own API key. Any compatible agent or application can connect by changing its base URL and supplying that key.

When a request arrives, Aquaduck verifies access, checks the requested model and configuration, selects an eligible runtime, and places the job in the appropriate queue. It streams the result through the same connection while handling failures and recovery. The interface stays the same whether one machine, several cooperating devices, or an approved cloud runtime serves the request.

Why one interface is better than several

Even when inference APIs look similar, they can differ in authentication, model identifiers, message formats, streaming events, tool definitions, structured output, token accounting, rate limits, error codes, and retry behavior. A unified API creates one control point for the entire inference path.

ConcernMultiple inference APIsOne Aquaduck inference API
Application codeProvider branches and adaptersOne request and response contract
Team accessSeparate endpoints and credentials for each runtimeOne team endpoint with individual API keys
Key managementManaged separately across providersManaged by each member or centrally by an administrator
Runtime changesApplication changes and redeploymentInfrastructure configuration changes
RoutingHard‑coded into application logicDriven by capacity and organization‑defined rules
GatingRepeated permission logic at each integrationAccess policies applied before work is scheduled
Job queuesSeparate queues and capacity rulesOne scheduling layer across available capacity
FailoverCustom fallback logic in every applicationAlternate eligible runtimes selected by the network
Model integrityVerified independently for each runtimeModel identity and configuration checked centrally
StreamingDifferent event formats and edge casesOne stream format for the application
Errors and retriesProvider‑specific handlingOne normalized failure path
ObservabilityMetrics split across systemsRequests measured at one gateway
Privacy policyEnforced separately in every integrationApplied at the inference layer

Routing becomes infrastructure

Agents and applications connect to the inference API while Aquaduck coordinates the available models and machines. The team shares one endpoint, each member uses a separate API key managed individually or by an administrator, and every connected device can add serving capacity to the network.

A model can move from one laptop to a larger machine, be split across a pair of devices, or become part of a routed private network. The endpoint does not need to move with it. The team keeps one integration while the inference layer gains more models, more capacity, and more ways to place work.

Glossary

Running a trained model to generate a result from an input.
A network interface through which an application submits model requests and receives generated results.
The URL an application calls to access the inference API.
The software and hardware environment that loads and executes a model.
Rules that select the model and execution location for a request.
Access and policy checks that determine whether and where a request may run.
The ordered set of requests waiting to be processed.
Verification of the loaded model weights.
Returning tokens one by one as the model generates them instead of waiting for the complete response.