Agentic AI on Kubernetes: A Production Checklist for Multi-Agent Security Platforms
A practical checklist for running agentic AI and multi-agent security workflows on Kubernetes with workload isolation, network policy, RBAC, observability, and human approval gates.
Agentic AI does not become production-ready because the model can reason through a demo. It becomes production-ready when every agent has a workload identity, a narrow network path, a health model, a resource budget, an audit trail, and a deterministic way to say “no” before it changes a real system.
Table Of Content
- The production question is operational, not just cognitive
- 1. Make every agent a workload you can kill, scale, and observe
- Define a per-agent workload contract
- Minimum workload contract
- 2. Treat agent-to-tool traffic as a security boundary
- Network policy should describe allowed conversations
- 3. Move safety rules out of the prompt
- RBAC is part of the safety model
- Policy gate checklist
- 4. Keep the human in the loop by protocol
- Approval should be fast without being invisible
- 5. Observability must follow a task, not just a Pod
- Measure the decision path
- 6. Use classical filters before LLM escalation
- A go-live checklist for agentic AI on Kubernetes
- Ship only when these gates pass
- Bottom line
- Sources
That is the practical lesson from a new CNCF post on building a multi-agent security platform on Kubernetes. The article describes an architecture for a security-operations platform that uses the A2A protocol for inter-agent coordination, MCP for environment integration, Falco and eBPF telemetry, Kafka, a classical Isolation Forest anomaly model, and LLM-driven agents. The exact project stack may not match your environment. The operating model is the useful part: treat agents as cloud-native workloads, not as scripts hiding behind a chatbot.
This Learning Hub checklist translates that architecture into release gates for teams that want to run multi-agent AI in Kubernetes. It is written for platform engineers, security engineers, and MLOps teams piloting agentic workflows that can read telemetry, call tools, propose rules, open tickets, or trigger containment actions.
The production question is operational, not just cognitive
The CNCF article’s first useful move is to shift the debate away from “which framework is smartest?” and toward “where does the agent live, how does it fail, and what can it reach?” It argues that each agent should be a Kubernetes workload rather than an in-process module. That means standard platform controls (rollouts, autoscaling, namespace isolation, metrics, logs, and identity) can apply to the agent layer without inventing a parallel operations model.
That matters most in security automation. A reviewer agent, a threat-analyst agent, and a tool-calling agent do not have the same blast radius. One may only classify events. Another may propose a detection rule. Another may request a firewall change or a containment action. If they all run as one process with one credential set, your “agentic” architecture has collapsed back into a privileged automation box.
1. Make every agent a workload you can kill, scale, and observe
Start by packaging every production agent as a distinct Kubernetes workload with one clear job. The unit of release should be the agent, not the whole agent swarm. The unit of rollback should also be the agent. If a summarizer agent is timing out against a model API, it should not drag down the reviewer agent that decides whether an action is safe.
The Kubernetes documentation for resource management for Pods and containers explains the difference between requests and limits: requests help the scheduler place Pods, while limits constrain how much CPU or memory a container may consume. The documentation for liveness, readiness, and startup probes shows the health-check model Kubernetes expects. Agent services need both. A model timeout, prompt-regression loop, or stuck tool call should surface as workload health, not as a mystery inside an application log.
Define a per-agent workload contract
Before an agent reaches production, write down the workload contract. The contract should be short enough to review during incident response and strict enough to block unsafe deployments.
Minimum workload contract
- Purpose: the one action family this agent owns, such as classification, retrieval, policy review, ticket drafting, or containment proposal.
- Inputs and outputs: the event types it may read and the objects it may emit, including schemas for proposed actions.
- Runtime budget: CPU and memory requests and limits for the container, plus a maximum model-call latency target.
- Health model: readiness, liveness, and startup probes that distinguish “temporarily not ready” from “must be restarted.”
- Owner: the team and on-call path responsible for the agent’s failures, prompts, tools, and releases.
For Kubernetes clusters using the apps/v1 Deployment API, a minimal production pattern is to give the agent explicit resource budgets and probes. The example below is illustrative; teams should tune paths and thresholds to the agent service they actually run.
apiVersion: apps/v1
kind: Deployment
metadata:
name: reviewer-agent
namespace: ai-security
spec:
replicas: 2
selector:
matchLabels:
app: reviewer-agent
template:
metadata:
labels:
app: reviewer-agent
spec:
containers:
- name: reviewer-agent
image: registry.example.com/ai-security/reviewer-agent:2026.06
ports:
- containerPort: 8080
resources:
requests:
cpu: "500m"
memory: "512Mi"
limits:
cpu: "2"
memory: "2Gi"
readinessProbe:
httpGet:
path: /readyz
port: 8080
livenessProbe:
httpGet:
path: /healthz
port: 8080
startupProbe:
httpGet:
path: /startupz
port: 8080
2. Treat agent-to-tool traffic as a security boundary
The CNCF article says inter-agent traffic in the described platform uses direct mTLS at the gRPC/HTTP transport, with cert-manager issuing per-agent identities, and network controls restricting which agents may reach which MCP servers. The important principle is not that every team must copy that exact product set. The principle is that messages between agents can carry proposed detection rules, response actions, and environmental context. That traffic deserves the same treatment as sensitive data-plane traffic.
Kubernetes NetworkPolicy is one useful layer, but it has a hard prerequisite: Kubernetes documentation notes that the cluster must use a networking solution that supports NetworkPolicy enforcement, and creating a NetworkPolicy resource without a controller that implements it has no effect. Do not treat a YAML object as a control until the CNI plugin enforces it.
Network policy should describe allowed conversations
For Kubernetes NetworkPolicy API networking.k8s.io/v1, start with the conversations you want to allow, not a broad namespace permit. The following example limits egress from a reviewer agent to an MCP server in the same namespace on TCP port 8443. In a real cluster, pair this with ingress policy on the MCP server and mTLS identity checks at the application transport.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: reviewer-egress-to-mcp
namespace: ai-security
spec:
podSelector:
matchLabels:
app: reviewer-agent
policyTypes:
- Egress
egress:
- to:
- podSelector:
matchLabels:
app: mcp-server
ports:
- protocol: TCP
port: 8443
The design question is simple: if an agent is compromised, which services can it still reach? If the answer is “everything in the namespace,” the agent is not isolated enough for production security work.
3. Move safety rules out of the prompt
The CNCF article’s strongest warning is that agent safety constraints should be policy-as-code, not hidden inside a reviewer agent’s system prompt. A prompt can guide behavior, but it is not the same as a versioned, tested, reviewable policy. In the described architecture, the reviewer agent gets a deterministic verdict from policy machinery before it acts.
That distinction should become a release gate. A production agent may explain why it thinks an action is safe. It should not be the only authority deciding whether the action is permitted. For example, “may isolate a workload” should depend on severity, asset tier, time window, owner policy, and change-freeze state, not on a paragraph in a prompt.
RBAC is part of the safety model
Kubernetes RBAC good practices emphasize least privilege for users and service accounts. Agent platforms should apply that rule literally. Use separate service accounts for separate agents. Avoid giving a reviewer agent write access to the same resources it reviews. Avoid broad verbs such as * and broad resources such as * unless a human-admin workflow has explicitly approved that risk.
Policy gate checklist
- Authorization: the agent’s Kubernetes service account can only read or write the resources required for its role.
- Action schema: every proposed action has a typed schema, not free-form text that a downstream tool must interpret.
- Deterministic verdict: policy-as-code returns allow, deny, or escalate before a tool is called.
- Test coverage: policies have unit tests for common approvals, denials, and emergency exceptions.
- Auditability: each verdict records the policy version, input facts, agent identity, and trace ID.
4. Keep the human in the loop by protocol
“Human in the loop” is often used as a cultural promise. The production version needs to be a protocol state. The CNCF article describes three terminal outcomes for consequential decisions: auto-execute, auto-reject, or escalate to a human SOC analyst with reasoning and approval commands. That is the right mental model. Escalation is not an exception path; it is a normal output of the system.
Define the escalation triggers before launch. Common triggers include low reviewer confidence, policy ambiguity, high-risk asset tier, unusual time of day, a change-freeze window, mismatch between anomaly score and business context, or an action that would cross a customer boundary. If the team cannot list those triggers, it is not ready for autonomous execution.
Approval should be fast without being invisible
ChatOps approval can be useful because it meets analysts where they work, but the approval path must still produce durable records. An approval message should include the event, proposed action, policy verdict, affected workloads, model/tool calls, and rollback plan. The post-incident review should not depend on screenshots from a chat thread.
5. Observability must follow a task, not just a Pod
The CNCF architecture uses an A2A trace_id to connect the work across agents, logs, metrics, tool calls, and LLM token usage. That is an important pattern. Kubernetes can tell you that a Pod restarted. It cannot, by itself, tell you why a reviewer escalated one proposed containment action and approved another. Multi-agent systems need task-level observability.
Measure the decision path
- Agent identity: which workload handled each step.
- Trace ID: one identifier across the event, retrieval, model call, policy verdict, and final action.
- Tool calls: which MCP or internal service was called, with latency and result class.
- Decision ratio: auto-execute, auto-reject, and escalate rates by policy version and asset tier.
- Cost and capacity: model-token usage, queue depth, retries, and rate-limit behavior.
This is where platform engineering and security engineering must share ownership. Platform teams own the workload and telemetry substrate. Security teams own the decision semantics. Agentic AI fails when those two records split.
6. Use classical filters before LLM escalation
The CNCF post describes a classical Isolation Forest anomaly model pre-filtering telemetry before LLM-driven agents enter the workflow. That is a useful guardrail because not every decision benefits from sending raw telemetry to a language model. Classical filters, rules, and thresholds can reduce noise, make capacity planning saner, and give policy gates stable input signals.
The goal is not to replace LLM reasoning with a single anomaly score. The goal is to separate jobs. Statistical filters can prioritize and normalize. Policy engines can decide whether an action is allowed. Agents can summarize, retrieve context, draft rules, and ask for approval. Humans can handle ambiguity and responsibility. Mixing all of that into one prompt makes the system harder to test and harder to trust.
A go-live checklist for agentic AI on Kubernetes
Ship only when these gates pass
- Workload isolation: each production agent runs as a separate workload with clear owner, resource budget, probes, and release history.
- Identity: every agent has its own service account and transport identity; shared credentials are treated as a release blocker.
- Network boundary: only required agent-to-agent and agent-to-tool paths are open, and NetworkPolicy enforcement has been tested with the actual CNI plugin.
- Policy separation: safety rules are versioned policy-as-code, not buried only in prompts.
- Human protocol: high-impact actions can end in an explicit human escalation state with durable approval records.
- Traceability: one trace ID links the event, model call, policy verdict, tool call, and final outcome.
- Failure drills: the team has tested model API timeout, MCP server outage, bad prompt release, policy denial, false positive anomaly spike, and rollback.
Bottom line
Agentic AI on Kubernetes is promising precisely because Kubernetes already has patterns for workloads, identity, rollout, health, resource control, network policy, and observability. The trap is assuming those patterns automatically apply to agents. They only apply if the team designs agents as first-class production workloads.
The safest path is boring on purpose: narrow agents, narrow credentials, narrow network paths, deterministic policy gates, task-level traces, and human escalation as a designed state. If a multi-agent platform cannot pass those gates, it may still be a useful prototype. It is not yet a production security platform.
Sources
- CNCF Blog: Why cloud native belongs at the heart of agentic AI
- Kubernetes documentation: Role Based Access Control Good Practices
- Kubernetes documentation: Network Policies
- Kubernetes documentation: Configure Liveness, Readiness and Startup Probes
- Kubernetes documentation: Resource Management for Pods and Containers
- Featured image source: U.S. Air Force photo via Wikimedia Commons
Featured image: server rack for a cyber security pilot program by Senior Airman Imani West, U.S. Air Force, released as public domain U.S. Air Force work via Wikimedia Commons; cropped and converted to WebP for sxz.io.








No Comment! Be the first one.