Kubernetes Read-Only AI Agents: A GitOps Checklist for Cluster-Aware Assistants
A production checklist for running read-only cluster-aware AI agents inside Kubernetes with scoped RBAC, GitOps review, and Argo CD delivery.
A Kubernetes AI assistant is safest when it cannot change the cluster it is trying to explain. That is the useful pattern in a new CNCF walkthrough on building a cluster-aware AI agent with Kubernetes, Argo CD, and GitOps. The post describes a self-hosted, read-only agent that runs inside the cluster, reads live Kubernetes state, reasons with a local model, and is delivered through a GitOps pipeline instead of a hosted SaaS control plane.
Table Of Content
- What the CNCF pattern gets right
- 1. Start with a read-only release gate
- Minimum RBAC contract
- Example: Kubernetes RBAC v1 read-only scope
- 2. Keep the model local, but do not confuse locality with safety
- 3. Put prompts, model choices, and RBAC in Git
- GitOps review checklist
- 4. Treat image automation as a controlled input, not a magic deploy button
- Release gates for the delivery chain
- 5. Separate diagnose, propose, and act
- What to log for every diagnosis
- 6. Watch the agent like any other production workload
- Go-live checklist
- Bottom line
- Sources
This Learning Hub guide turns that pattern into a production checklist for platform teams. The goal is not to copy one demo repository into production. The goal is to separate a useful diagnostic assistant from an unsafe automation bot: scoped identity, read-only RBAC, observable runtime, Git-reviewed changes, and a clear line between “diagnose” and “act.”
What the CNCF pattern gets right
The CNCF article contrasts hosted “AI for Kubernetes” tooling with an agent that stays inside the cluster boundary. In the described architecture, the runtime side includes an Ollama pod serving a local Mistral 7B model, a FastAPI pod exposing the agent interface, a PersistentVolumeClaim for model weights, and a dedicated service account with a ClusterRole that can read pods, events, logs, services, and deployments. The delivery side uses GitHub Actions to build a multi-architecture image and Argo CD Image Updater to keep the deployment moving through GitOps.
The important part is the operating model. Cluster state is not sent to an outside model provider. Prompts, model selection, and RBAC are treated as versioned assets. The agent has enough access to explain what is happening, but not enough access to mutate the cluster when a model produces a confident but wrong recommendation.
1. Start with a read-only release gate
For a first production pilot, treat read-only access as a hard gate. A cluster-aware assistant can still be valuable if it answers questions such as “why are these pods pending?”, “which deployment changed most recently?”, or “what events point to an image pull failure?” It does not need write access to restart workloads, patch deployments, or edit network policy.
Kubernetes service accounts are the right identity primitive for this boundary. The Kubernetes service account documentation describes a service account as a non-human account that gives a distinct identity in the cluster. Application pods and system components can use service account credentials to identify themselves to the API server. That maps directly to an AI agent: the assistant should have its own identity, not borrow the default service account or a human operator credential.
Minimum RBAC contract
The Kubernetes RBAC documentation defines Role, ClusterRole, RoleBinding, and ClusterRoleBinding objects in the rbac.authorization.k8s.io API group. For a diagnostic assistant, begin with verbs such as get, list, and watch. Do not add create, update, patch, or delete until a separate approval path exists.
Example: Kubernetes RBAC v1 read-only scope
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: cluster-ai-diagnostics-readonly
rules:
- apiGroups: [""]
resources: ["pods", "events", "services"]
verbs: ["get", "list", "watch"]
- apiGroups: [""]
resources: ["pods/log"]
verbs: ["get"]
- apiGroups: ["apps"]
resources: ["deployments", "replicasets", "statefulsets", "daemonsets"]
verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: cluster-ai-diagnostics-readonly
subjects:
- kind: ServiceAccount
name: cluster-ai-agent
namespace: ai-ops
roleRef:
kind: ClusterRole
name: cluster-ai-diagnostics-readonly
apiGroup: rbac.authorization.k8s.io
This example is for Kubernetes clusters using RBAC authorization and should be reviewed against the actual cluster version, namespace model, and compliance rules. If the agent only needs one namespace, use a namespaced Role and RoleBinding instead of a ClusterRoleBinding.
2. Keep the model local, but do not confuse locality with safety
The CNCF walkthrough emphasizes that no cloud AI provider is involved and that the model runs locally. That improves data-boundary control, but it is not a complete safety story. A local model can still hallucinate, overstate certainty, or suggest an unsafe remediation. Locality protects where data travels; RBAC and release policy protect what the agent can do.
Platform teams should document three boundaries before launch:
- Data boundary: which resources, logs, events, and namespaces the agent can read.
- Decision boundary: whether the agent may only explain, may draft changes, or may request human approval.
- Execution boundary: which separate system, if any, is allowed to perform a change after review.
For most organizations, the safest first version is explain-only. Let the assistant generate a diagnosis and a proposed runbook step, then require a human operator or an existing automation controller to execute the change.
3. Put prompts, model choices, and RBAC in Git
The GitOps part matters because agent behavior is configuration, not just code. A prompt change can alter how a diagnostic assistant interprets events. A model change can alter latency and answer style. An RBAC change can alter blast radius. Those changes deserve pull requests, reviews, and rollback history.
Argo CD’s declarative setup documentation states that Argo CD applications, projects, and settings can be defined declaratively using Kubernetes manifests. In practice, that means the agent Deployment, ServiceAccount, RBAC, ConfigMap, model selection, and prompt templates should be represented as Kubernetes manifests or Helm/Kustomize output under version control.
GitOps review checklist
- Prompt changes: reviewed by the platform owner and the incident-response owner.
- Model changes: tested for latency, memory use, and failure behavior before merge.
- RBAC changes: reviewed separately from application code and blocked if they introduce write verbs without approval.
- Namespace changes: checked for data exposure before the agent gains visibility into additional teams or tenants.
- Rollback plan: every release has a known previous manifest state that can be restored.
4. Treat image automation as a controlled input, not a magic deploy button
The CNCF article’s delivery chain includes GitHub Actions and Argo CD Image Updater. The GitHub Actions documentation describes workflows as automated processes triggered by events and made of jobs. The Argo CD Image Updater documentation describes a tool that updates container images for Kubernetes workloads managed by Argo CD, including the ability to write changes back to Git.
That is powerful, but it should be constrained. A new agent image can change API calls, prompt handling, output parsing, or model integration. Teams should require image provenance, digest pinning or signed image policy where available, and a staging cluster or namespace before production sync.
Release gates for the delivery chain
- Build provenance: the image was built by the expected workflow from the expected repository and commit.
- Artifact review: the manifest change shows the image tag or digest and the related prompt/RBAC changes.
- Staging diagnosis: the agent answers known cluster-state questions without needing write privileges.
- Resource budget: CPU, memory, and storage requests match the model size and expected query load.
- Failure test: model startup failure, image pull failure, and API timeout all produce observable alerts.
5. Separate diagnose, propose, and act
The CNCF source describes two modes: an LLM-only path for general questions and an agent path that reads live cluster state before reasoning over it. Production teams should extend that idea into three explicit states: diagnose, propose, and act.
Diagnose means the agent reads cluster state and explains likely causes. Propose means it drafts a change, runbook step, or ticket payload. Act means a separate, authorized system performs the change after policy and human approval. Keeping those states separate prevents a diagnostic assistant from becoming an unreviewed cluster operator.
What to log for every diagnosis
- Agent identity: service account, namespace, image digest, and prompt/model version.
- Input scope: namespaces, resource types, and time window used for the answer.
- Evidence: event names, pod names, deployment names, and log excerpts that support the diagnosis.
- Confidence and uncertainty: the agent should identify missing evidence, not just produce a fluent answer.
- Operator decision: accepted, rejected, escalated, or converted into a separate change request.
6. Watch the agent like any other production workload
The agent is “just another Kubernetes workload” only if it receives the same operational discipline as other workloads. Give it liveness and readiness checks, resource requests and limits, storage monitoring for model weights, logs that separate retrieval from reasoning, and dashboards that show request volume, error rate, response latency, and model startup time.
The most important alert is not “the AI gave a bad answer.” That will happen. The important alerts are boundary failures: unexpected API errors, forbidden RBAC attempts, egress that should not exist, repeated timeouts, or a sudden increase in queries against namespaces the agent rarely inspects.
Go-live checklist
- Identity: the agent has a dedicated service account and does not use default or human credentials.
- RBAC: the first production release is read-only and uses the smallest namespace/resource scope that works.
- Data boundary: sensitive namespaces, secrets, and tenant data are excluded unless explicitly approved.
- GitOps control: prompt, model, RBAC, and deployment changes are reviewed in Git.
- Delivery safety: GitHub Actions and Argo CD Image Updater changes are traceable to commits and tested before production sync.
- Execution split: the agent can diagnose and propose, but write actions require a separate approval path.
- Observability: logs and metrics prove which evidence the agent used and where failures occurred.
- Rollback: the team can revert the image, prompt, model, and RBAC independently.
Bottom line
A cluster-aware AI agent can be useful precisely because it sees real Kubernetes state instead of generic documentation. That also makes it risky. The safe production pattern is to make the assistant boring: one identity, limited read-only access, local data boundaries, Git-reviewed behavior, observable runtime, and a firm split between diagnosis and action.
If a team cannot explain what the agent can read, what it cannot change, how its prompts are reviewed, and how to roll it back, it is not ready for production. If it can pass those gates, a read-only in-cluster assistant can become a practical diagnostic layer for platform operations.
Sources
- CNCF Blog: Building a Cluster-Aware AI Agent with Kubernetes, Argo CD, and GitOps
- Kubernetes documentation: Using RBAC Authorization
- Kubernetes documentation: Service Accounts
- Argo CD documentation: Declarative Setup
- Argo CD Image Updater documentation
- GitHub Docs: Workflows
- Featured image source: Data Center 2 (UNC).jpg on Wikimedia Commons
- Creative Commons: Attribution-ShareAlike 4.0 International
Featured image: Data Center 2 (UNC).jpg by Ana Las Heras, published under CC BY-SA 4.0 via Wikimedia Commons; cropped and converted to WebP for sxz.io.








No Comment! Be the first one.