Hybrid architecture

Aquaduck's hybrid architecture keeps your sensitive workloads on local and connected devices, with selective cloud routing only when preferred.

Request a demo

What it does

Decide what stays on‑device and what can move. The runtime handles placement so applications don't need to fork between local and cloud paths.

Why it matters

Privacy‑sensitive context can remain local while heavier jobs use spare capacity elsewhere.

Deep dive

Hybrid AI combines local, private‑network, and cloud inference behind one execution layer. A private workflow can stay on the machine where it began, a larger job can use models and capacity on approved devices across the network, and a complex non‑sensitive task can move to an approved cloud model when local resources are not the right fit.

Organizations can apply their privacy, capability, capacity, and cost rules as policy. Aquaduck then turns those rules into a runtime decision framework instead of forcing every application to implement its own local, network, and cloud branches.

One policy, three paths

Execution pathWhere the request runsBest suited forData exposure
On‑deviceThe originating machineSensitive prompts, low‑latency work, and offline useStays on that machine
Private networkOne or more approved devices connected to the organization’s networkLarger models, pooled capacity, and shared internal workloadsStays within the approved network
CloudAn approved external inference providerComplex, high‑capacity, or specialized non‑sensitive tasksSent only when policy permits

Sensitive data becomes a routing input

Organizations can apply data‑sensitivity thresholds to determine whether a prompt can leave its originating machine. Each device can run Personally Identifiable Information (PII) detection against a prompt locally to determine where it should be processed before sending it elsewhere.

Local inspection matters because sending a prompt to the cloud to determine whether it was safe to send would defeat the purpose. It also gives the organization control over failure behavior. If classification is unavailable or uncertain, a policy can choose to keep the request local.

PII detection is just one of many signals. Organizations should also consider the sensitivity requirements around source code, contracts, credentials, financial records, and confidential business context. Rules can combine content inspection with application identity, user permissions, workload type, data residency, approved models, and allowed destinations.

Rules define the boundary

Organizations can express routing behavior as policy under one framework.

Example ruleRuntime behavior
Keep legal documents on‑deviceUse only the originating machine
Do not send detected PII externallyUse only the originating machine or approved private devices
Prefer existing hardwareUse local and network capacity before considering cloud inference
Use cloud for complex non‑sensitive workRoute complex, non‑sensitive tasks to an approved stronger or specialized cloud model
Require a specific model or regionConsider only runtimes that satisfy the model and residency requirement
Never weaken privacy during failoverQueue or reject the request if no permitted runtime is available

Start private, expand intentionally

Hybrid AI means every request has access to the execution path that fits its requirements. Local inference provides the smallest trust boundary. The private network adds models and capacity from hardware the organization controls. Cloud inference remains available for approved tasks that need capabilities the private pool cannot provide efficiently.

With Aquaduck, those choices live in the runtime. Applications keep one integration while the organization decides what may move, where it may run, and when cloud capacity should be used.

Glossary

An architecture that coordinates inference across local devices, private infrastructure, and cloud services.
Running a model on the same machine that holds or receives the workload.
Running a model on approved devices within an organization’s network.
Sending an eligible request to an approved external inference runtime.
Identifying personally identifiable information so policy can restrict where a request may run.