ATAILA Newsroom · Hybrid · Budapest · 2026-09-11
2026-W37
AWS Outposts has no current-generation GPU. So we are building the hybrid that puts it on-prem.
If you run agents on AWS and your inference must stay on your premises, you have a hardware problem: AWS Outposts offers no current-generation GPU. The only GPU family on Outposts is g4dn — NVIDIA T4, a 2019 card — and only on first-generation racks; the second-generation racks of April 2025 and the single-rack Outposts of September 2026 ship CPU families, with GPU support "coming soon" since launch. AWS's own answer for GPU at the edge is EKS Hybrid Nodes: your hardware, their control plane. That is exactly what ATAILA's GPU nodes are. So we are building it — with AWS solutions architects, on a real customer case, as a fixed-price six-week pilot.
The gap, in AWS's own words
Second-generation AWS Outposts racks (April 2025) brought C7i, M7i and R7i; C8i, M8i and R8i followed in February 2026; the second-generation single-rack Outposts of September 2026 lists the same CPU families plus accelerated-networking instances. No GPU instance appears in any of those announcements. The launch post for the second generation said that support for more instances, "including GPU-enabled instances", is coming soon — and in September 2026 it still is. The one GPU family Outposts does support, g4dn, pairs NVIDIA T4 GPUs with 2019-era CPUs, on first-generation racks only.
AWS's own path for current GPUs at the edge is different: Amazon EKS Hybrid Nodes, generally available since December 2024, join on-premises or edge machines — any hardware, any GPU you install the driver for — to an EKS control plane in the region. AWS's engineers have published GenAI inference on hybrid nodes with customer GPUs, and an NVIDIA DGX Spark joined as a hybrid node. The hardware is the customer's problem to bring. We bring it.
What we are building, with whom
This week an AWS solutions architect brought us a customer case: a European AWS customer that wants agents developed in the cloud and inference on-premises, an immutable repository with tests run on on-prem hardware, a smart gateway with inference routing, and one observability and governance plane across cloud and on-prem. We said yes the same morning. This is a technical collaboration between engineers on a customer case — not an AWS partnership, and not an AWS endorsement; when either becomes true we will write it here.
The architecture we are building, block by block:
- Agents in AWS. Built with Strands Agents (Apache-2.0) or the framework the customer already uses; running on Bedrock AgentCore Runtime or on EKS.
- Inference on ATAILA's GPU nodes. vLLM — or NVIDIA NIM where a model licence asks for it — serving OpenAI-compatible endpoints from our bare-metal GPU nodes, joined to the customer's EKS cluster as hybrid nodes.
- Two gateways, one policy. Amazon Bedrock AgentCore Gateway at the perimeter, with Bedrock Guardrails and AgentCore Policy, routing inference targets to the on-prem endpoint; ATAILA's own AI gateway on-prem for model routing, per-tenant metering and a cloud fallback when the on-prem queue is full. Plan B needs nothing from AWS: the ATAILA gateway alone, which is what our platform runs today.
- One image, two registries. Built once, pushed to Amazon ECR and mirrored to Harbor on-prem with immutable tags — tested on AWS for the baseline, tested on the on-prem nodes for the comparison.
- One observability plane. AWS Distro for OpenTelemetry on every node, cloud and on-prem, into Amazon Managed Grafana and into ATAILA's Grafana; AgentCore observability for the agent traces; ATAILA's per-tenant usage evidence for the governance report.
- Hybrid network. An AWS Site-to-Site VPN to the customer's VPC, or WireGuard — a project zone per tenant on our side, exactly as the platform isolates everything else.
The deliverable of the pilot is a comparison, not a slide: the same agent, the same test suite, run against the cloud model and against the on-prem GPUs, with latency, cost, quality and data residency on one dashboard. That is what a board decides on.
1 · Week 1
Design workshop.
Target architecture agreed, model list, the test suite and its acceptance criteria, the network design, the success criteria. Remote, or in our Executive Briefing Center.
2 · Weeks 2–3
The hybrid substrate.
VPN up; an ATAILA GPU node joined to the customer's EKS cluster as a hybrid node; vLLM serving the first model; the gateway chain working end to end; OpenTelemetry into both Grafanas.
3 · Weeks 4–5
The agent baseline.
The customer's first agent containerised, in ECR and Harbor, tested on AWS for the baseline and on-prem for the comparison; guardrails and policy on the gateway.
4 · Week 6
The decision.
A report the board can read: quality, latency, cost and data residency per location, and the go/no-go for production on the customer's own hardware or on ATAILA Cloud. Fixed price for the six weeks, credited against a Build if you continue.
Where this sits in the mission
The Service Provider mission promised connectors to the hyperscalers as an option, never a dependency. AWS is the first, and it is a hybrid rather than a connector because the customer's need is a hybrid: the agents live where the developers are, the inference lives where the data is. Azure, Oracle Cloud and Huawei Cloud follow on the roadmap — each when a customer asks or a vendor commits. Every ATAILA edition runs without any of them.
We call this a technical collaboration because that is what it is. ATAILA is not an AWS partner today and AWS has not endorsed this work; if either changes, this page will say so. The pilot's comparison is handed to the customer whatever it shows — including the case where the honest answer is that the cloud model wins.
Sources
- AWS What's New: second-generation AWS Outposts racks (April 2025); C8i, M8i and R8i on Outposts (February 2026); second-generation single-rack Outposts (September 2026); the second-generation launch post with "GPU-enabled instances coming soon"; Amazon EC2 G4 instances.
- AWS: Amazon EKS Hybrid Nodes (December 2024); Run GenAI inference across environments with Amazon EKS Hybrid Nodes.
- AWS: Amazon Bedrock AgentCore generally available (October 2025); AgentCore Gateway inference targets; Strands Agents SDK.
Talk to us
If you run agents on AWS and your inference has to stay on your premises — or you are the service provider whose customers do — write two sentences on the form. It reaches the founder directly.
← Back to the Newsroom Press inquiries: contact us