Red Hat’s Metal to Agents Framework Turns AI Token Growth Into an Infrastructure Bet
Red Hat's metal to agents framework treats hardware, inference, model governance, and autonomous agents as one architecture, a response to Goldman Sachs' forecast that AI token consumption will grow...
Enterprise AI has moved past isolated chatbot pilots. Organizations are now running complex reasoning models and autonomous agents that operate continuously instead of answering a single prompt and stopping. That shift changes the economics of AI infrastructure: token consumption, the basic unit of cost behind every model call, is projected to grow 24 times over by 2030, according to Goldman Sachs Research. Agentic systems that monitor environments and execute multi-step tasks consume far more tokens than a person typing questions into a chat window, a shift Goldman’s research ties to both consumer and enterprise adoption of always-on agents.
Table Of Content
- The cost problem behind the framework
- Five layers, one stack
- Hardware and hybrid cloud
- The platform layer: RHEL, RHEL AI, and OpenShift
- Inference services and token economics
- Model services: governance for shared infrastructure
- Agent services: closing the production gap
- Why this is more than a marketing framework
Red Hat is using that forecast to make a structural argument. In a blog post published August 13, 2026, Brian Stevens, Red Hat’s senior vice president and chief technology officer for AI, laid out a framework the company calls “metal to agents”: one coordinated architecture spanning physical hardware, inference serving, model governance, and autonomous agent operations. Red Hat CTO Chris Wright introduced the framework at Red Hat Summit 2026, arguing that treating AI infrastructure as a collection of disconnected tools creates technical debt that does not scale.
The cost problem behind the framework
The pitch starts with a warning about dependency. Relying entirely on external clouds and proprietary model APIs means an organization’s AI expenses scale directly with usage, leaving little room to control unit economics as adoption grows. Stevens frames this as a predictability problem as much as a cost problem: a team that cannot forecast its own inference bill cannot budget for scaling an agent from a pilot into a production workload.
Goldman Sachs’ own numbers explain why the timing matters. The bank’s research puts monthly token consumption at roughly 120 quadrillion by 2030, with a 12-fold increase in consumer use, things like AI-driven shopping and always-on phone assistants, compounding with separate enterprise adoption. Falling inference costs take some of the edge off: Goldman estimates semiconductor providers are cutting the cost per token by 60 to 70 percent a year. But for an organization whose usage is growing faster than unit costs are falling, that decline does not eliminate the budget problem, it only slows it down.
Five layers, one stack
Red Hat’s framework splits the AI stack into five layers, each meant to solve a distinct part of the deployment problem instead of being bolted on separately after the fact.
Hardware and hybrid cloud
At the base sits accelerator choice and hybrid cloud placement. The pitch is portability: an organization that standardizes on a common software layer across GPUs, TPUs, and other accelerators can move workloads to whichever hardware is cheapest or most available at a given moment, rather than being locked into a single vendor’s silicon.
The platform layer: RHEL, RHEL AI, and OpenShift
Above raw hardware choice sits a dedicated platform layer built on Red Hat Enterprise Linux, RHEL AI, and OpenShift, which Stevens frames as the stable foundation data science workloads need to run at all. RHEL AI packages the operating system kernel, hardware-optimized drivers, and core data science libraries into a single bootable image, so scaling a fleet of inference nodes does not mean re-solving driver and dependency compatibility on every new machine. OpenShift serves as the shared Kubernetes platform underneath virtual machines, containers, and AI workloads alike, with network isolation between tenants and the Red Hat build of Kueue dynamically allocating scarce accelerators across competing jobs. Three capabilities recur throughout this layer: day-0 support for new hardware as it ships, an immutable, image-based update model that limits configuration drift across a large fleet, and confidential computing that adds cryptographic protection for model weights and data while they are in use.
Inference services and token economics
The next layer is where the argument gets concrete. vLLM, the open source inference engine at the center of Red Hat’s AI strategy since its 2025 acquisition of Neural Magic, handles model serving. llm-d, a Kubernetes-native project Red Hat co-founded with Google Cloud, IBM Research, NVIDIA, CoreWeave, and other contributors, adds distributed orchestration on top of it: a scheduler that routes requests to instances with warm prefix caches, disaggregated prefill and decode stages, and a multi-tier cache for reused context.
The performance case for this layer is specific rather than universal. On one llm-d deployment, enabling prefix-cache-aware routing produced a 3 times improvement in output tokens per second and cut time to first token in half, according to llm-d’s own published benchmark results. Hardware choice matters just as much as software: in a separate example Red Hat cites, moving a trust-and-safety workload for an unnamed global media platform from GPUs to TPUs cut infrastructure costs by 92 percent while running 400 percent faster, a result driven mainly by that workload’s query-heavy, response-light traffic pattern rather than a universal TPU advantage.
Model services: governance for shared infrastructure
Above inference sits what Red Hat calls model services: centralized control over which teams can access which models, at what token quota, and through what credentials. An integrated AI gateway enforces request prioritization and exposes pre-validated, pre-optimized model versions instead of leaving every team to source and tune its own. The goal is to stop the cost and security problems that plagued shadow IT in the SaaS era from repeating themselves with model access.
Agent services: closing the production gap
The top layer targets what Stevens calls the production gap: the distance between an agent that works reliably on a developer’s laptop and one running unattended at data center scale, talking to real systems with real credentials. Red Hat’s answer has three parts. Every agent gets a verified identity through SPIFFE and SPIRE, an open source, CNCF-hosted framework for issuing cryptographic identity to software workloads, instead of relying on shared static credentials that are hard to audit or revoke. A Model Context Protocol gateway acts as a single, managed connection point between agents and internal systems, so wiring up a new tool does not require custom integration code for every agent that needs it. And a set of practices Red Hat labels AgentOps covers lifecycle management, kernel-isolated sandboxes, and observability, meant to turn what Stevens describes as chaotic software sprawl into a more reliable “Agents-as-a-Service” model.
Why this is more than a marketing framework
It is easy to read “metal to agents” as a five-syllable name for Red Hat’s existing product line: RHEL AI, OpenShift AI, vLLM, and Red Hat AI Enterprise all slot neatly into the layers Stevens describes. That is not a coincidence, and the framework is worth reading as a sales pitch as much as an architecture diagram.
But the underlying problem it describes is real, and not unique to Red Hat’s customers. Enterprises across the industry are deploying agents faster than they are building the identity, observability, and cost-governance systems needed to run them safely at scale, and the gap between a demo and a production deployment is exactly where security incidents and runaway bills tend to show up. A framework that forces verified agent identity, tool-access boundaries, and per-team token budgets into the same planning conversation as GPU procurement is addressing a genuine gap, even when the specific implementation on offer happens to be one vendor’s own stack.
The trade-off enterprises still have to weigh is commitment. Standardizing hardware, inference, model governance, and agent operations on one vendor’s framework buys coherence, but it also concentrates risk and negotiating leverage in that vendor’s direction. The token-growth numbers driving the pitch are industry-wide forecasts, not guarantees for any specific organization, and the performance figures Red Hat cites, the 92 percent cost reduction and the llm-d scheduling gains, come from individual deployments rather than universal benchmarks. Organizations evaluating the framework will still need to run their own numbers before committing an AI infrastructure roadmap to any single stack, Red Hat’s included.








No Comment! Be the first one.