TRENDING
Galvanized steel guardrail bolted to wooden posts along the edge of a bridge approach, with a grassy verge and a gravel road beside it
October 1, 2026
How to Enforce Guardrails on AI-Generated Terraform With Open Policy Agent and Rego
Microscope die shot of an AMD EPYC 7702 engineering sample I/O die, its circuit blocks glowing in teal, gold and violet
October 1, 2026
AMD Agrees to Buy Fei-Fei Li’s World Labs for $8.2 Billion to Steer Its Chip Roadmap
A silver signet ring engraved with a coat of arms between two sticks of red sealing wax on a grey surface
October 1, 2026
How to Build a Merkle Tree Certificate Issuer in Python to Keep Post-Quantum Certificates Small
Brass swing-bar door lock, a secondary latch, mounted on a hotel room door
October 1, 2026
Cloudflare’s Post-Quantum Visibility Turns Quantum Readiness Into a Per-Hop Audit
A seven-spot ladybird with black spots on its orange shell climbs a green plant stem
October 1, 2026
OpenAI Launches Dots, Always-On Agents, and Says It Is Still Fixing Known Vulnerabilities
01 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Faint white watermark of a crown above an oval emblem showing through blue paper, a design that stays invisible until light passes through the sheet
How to Detect and Strip Invisible Unicode in Python to Stop ASCII Smuggling and Trojan Source
September 30, 2026
A small white wooden toll booth with a Pay Point sign and a fare board at Penmaenpool Toll Bridge, with orange traffic cones on the bridge deck
Two Cloudflare Agent Billing Betas Turn Web Monetization Into a Question of Who Holds the Meter
September 30, 2026
Eight silver hex keys of graduated sizes fanned out on a steel ring against a dark green surface
Attackers Exploit a Hex-Encoding Bypass in Cisco SD-WAN Manager, and CISA Sets an October 3 Deadline
September 30, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 216 Posts
News 218 Posts
Learning Hub 188 Posts
Home/Learning Hub/How to Enforce Guardrails on AI-Generated Terraform With Open Policy Agent and Rego
Learning Hub

How to Enforce Guardrails on AI-Generated Terraform With Open Policy Agent and Rego

Learn Open Policy Agent and Rego by writing tested guardrails for a real Terraform plan, catching four different ways to open a security group, and turning the result into a CI exit code.

September 30, 2026 38 Min Read
12

AI coding assistants now routinely write Terraform that parses without complaint. The file is accepted, terraform validate is happy, and terraform plan prints a tidy list of resources to create. None of those checks asks the question that matters: is this plan safe? An assistant told to “make SSH work so we can debug” can easily open port 22 to the entire internet, and the result will look exactly as valid as a locked-down configuration.

Table Of Content

  • What you need before you start
  • Key terms in plain English
  • Step 1: Install OPA and Terraform and create the working folder
  • Step 2: Write the Terraform an assistant might hand you, and turn its plan into JSON
  • Step 3: Look inside the plan with opa eval
  • The two objects that matter: after and after_unknown
  • Step 4: Write the tests first, then the first policy
  • Step 5: Find the security groups your first policy cannot see
  • How the ingress guardrail works
  • Step 6: Guard storage, and meet the bug where a missing value passes
  • Step 7: Guard IAM, and handle values Terraform cannot know yet
  • Fail closed on unknown values
  • Step 8: Format it, lint it strictly, and measure what the tests never ran
  • Coverage: which lines did the tests never run?
  • Step 9: Turn the policy into a gate that a pipeline can trust
  • The two traps that make a gate lie
  • Telling a violation apart from a broken tool
  • Step 10: Let a model draft the policy, and let the checker decide
  • Verify the whole thing end to end
  • Mistakes worth avoiding
  • What this policy cannot do
  • The complete policy file
  • Where to go next
  • Sources and further reading

In this tutorial you will build a safety net for that problem. You will write a small set of guardrails in Rego, the policy language of Open Policy Agent (OPA), that read a Terraform plan and refuse it when it breaks your rules. You will write the tests first, watch a first draft miss real violations, fix it, and finally turn the result into an exit code that a pipeline can use to block a bad plan. Every command in this post was run against a real Terraform plan on a fresh install, and the outputs shown are the real ones.

By the end you will be able to:

  • turn a Terraform plan into JSON and explore it with opa eval;
  • write Rego rules that collect human-readable violation messages;
  • test a policy like application code, including a coverage report;
  • recognise the two classic ways a policy silently passes something it should have blocked: a shape you did not model, and a value that is missing or unknown;
  • gate a pipeline on OPA’s exit code without falling into the two traps that make a gate always pass or always fail.

What you need before you start

  • A terminal. I ran everything in Git Bash on Windows 11. The commands work the same way in a macOS or Linux shell. They put single quotes around Rego queries, so they will not work as written in cmd.exe.
  • OPA 1.21.1. It is a single executable with no dependencies. I only tested 1.21.1, which was released the day before I wrote this.
  • Terraform 1.16.4 and about 840 MB of free disk space, because the AWS provider that terraform init downloads is one very large binary.
  • Python 3 (I used 3.13), only for two small optional helpers: a coverage summary in Step 8 and a local-model experiment in Step 10. That experiment also needs Ollama and the requests package.
  • No AWS account. The Terraform file uses dummy credentials, and we only ever run plan, never apply, so nothing is created or billed.
  • Basic comfort with JSON and a terminal. You do not need to know Terraform or Rego beforehand. I explain each construct the first time it appears.

Key terms in plain English

  • Terraform plan: a preview of what Terraform would create, change or delete. terraform plan -out=tfplan saves it to a binary file, and terraform show -json prints the same information as JSON.
  • Policy as code: rules written down in a language a computer can evaluate on every change, instead of in a wiki page that people forget to read.
  • Open Policy Agent (OPA): a general-purpose policy engine. You hand it a JSON document called input plus some policies, and it answers questions about them.
  • Rego: OPA’s policy language. It is declarative: instead of writing steps, you describe what a violation looks like, and OPA finds every place in the input that matches.
  • Rule: a named definition in a policy. A rule such as deny contains msg if { ... } builds a set of messages. The lines inside the braces are conditions that must all be true, and every combination of values that makes them true adds one message to the set.
  • Undefined: what Rego produces when an expression refers to something that is not there, such as a missing key. Undefined is not the same as false, and that difference is behind the silent failures you will meet in this tutorial.
  • Guardrail: one rule that catches one kind of risk.
  • Fail closed: when a policy cannot tell whether something is safe, it treats it as a violation.

Step 1: Install OPA and Terraform and create the working folder

OPA ships as one executable. The OPA documentation lists the download for each platform. On macOS, brew install opa works. On Linux, download https://openpolicyagent.org/downloads/latest/opa_linux_amd64 (or opa_linux_arm64) with curl -L -o opa and run chmod 755 ./opa. On Windows, download https://openpolicyagent.org/downloads/latest/opa_windows_amd64.exe and save it as opa.exe. Put it somewhere on your PATH, or simply drop it into the working folder we are about to create. Terraform has an install page for every platform.

Create the folder layout and confirm both tools run:

mkdir guardrails-lab
cd guardrails-lab
mkdir infra policy demo
opa version
terraform version

You should see something like this (your build hash and platform will differ):

Version: 1.21.1
Build Commit: 2a109e54103370d2ef288782ef3cb4c8a37902b2-dirty
Build Timestamp: 2026-09-29T19:10:43Z
Build Hostname: 
Go Version: go1.27.1
Platform: windows/amd64
Rego Version: v1
WebAssembly: available
Terraform v1.16.4
on windows_amd64

The line that matters is Rego Version: v1. OPA 1.0 made the newer Rego syntax the default. In the upgrade guide’s words, “The in, every, if and contains keywords have been introduced over time, and Rego v0.x required an opt-in to prevent them from breaking policies that existed before their introduction.” The practical effect is that every rule body needs the if keyword and sets are built with contains. Older Rego snippets, and older language models, use the previous style and fail to parse. Step 10 shows exactly that. The quote comes from the OPA v0 to v1 upgrade guide.

Step 2: Write the Terraform an assistant might hand you, and turn its plan into JSON

We need something to check. Save the following as infra/main.tf. I wrote it by hand to look like the shortcuts an assistant takes when the instruction is “just make it work”; no model produced this file. It plans 14 AWS resources and hides these problems on purpose:

  • aws_security_group.web opens SSH (port 22) to 0.0.0.0/0, next to a perfectly acceptable HTTPS rule on port 443.
  • aws_security_group.api opens port 8080 to ::/0, which is the IPv6 way of saying “everyone”.
  • aws_security_group.db gets PostgreSQL (5432) opened to the world by a separate aws_security_group_rule resource.
  • aws_security_group.admin gets all traffic opened to the world by an aws_vpc_security_group_ingress_rule resource.
  • aws_ebs_volume.data never says whether it is encrypted, and aws_db_instance.main is unencrypted and publicly accessible.
  • aws_iam_policy.app_admin allows every action on every resource, and aws_iam_policy.log_writer allows every action on a resource that Terraform cannot name until apply time.

Four resources are compliant on purpose (internal, scratch, read_only and logs) so that we can prove the policy does not cry wolf.

terraform {
  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 6.66"
    }
  }
}

# Dummy credentials and skip flags: `terraform plan` never talks to AWS for this file.
provider "aws" {
  region                      = "us-east-1"
  access_key                  = "mock-access-key"
  secret_key                  = "mock-secret-key"
  skip_credentials_validation = true
  skip_metadata_api_check     = true
  skip_requesting_account_id  = true
}

# ---------- Networking: four different ways to open a security group ----------

# Shape 1: inline ingress block, IPv4
resource "aws_security_group" "web" {
  name        = "web"
  description = "Public web tier"

  ingress {
    description = "HTTPS from anywhere"
    from_port   = 443
    to_port     = 443
    protocol    = "tcp"
    cidr_blocks = ["0.0.0.0/0"]
  }

  ingress {
    description = "SSH so the assistant can debug"
    from_port   = 22
    to_port     = 22
    protocol    = "tcp"
    cidr_blocks = ["0.0.0.0/0"]
  }
}

# Shape 2: inline ingress block, IPv6
resource "aws_security_group" "api" {
  name        = "api"
  description = "Internal API"

  ingress {
    description      = "API from anywhere over IPv6"
    from_port        = 8080
    to_port          = 8080
    protocol         = "tcp"
    ipv6_cidr_blocks = ["::/0"]
  }
}

# Shape 3: separate aws_security_group_rule resource
resource "aws_security_group" "db" {
  name        = "db"
  description = "Database tier"
}

resource "aws_security_group_rule" "db_from_anywhere" {
  type              = "ingress"
  from_port         = 5432
  to_port           = 5432
  protocol          = "tcp"
  cidr_blocks       = ["0.0.0.0/0"]
  security_group_id = aws_security_group.db.id
}

# Shape 4: aws_vpc_security_group_ingress_rule (the newer resource type)
resource "aws_security_group" "admin" {
  name        = "admin"
  description = "Admin tooling"
}

resource "aws_vpc_security_group_ingress_rule" "all_traffic" {
  security_group_id = aws_security_group.admin.id
  cidr_ipv4         = "0.0.0.0/0"
  ip_protocol       = "-1"
}

# A compliant group, so the policy has something it must NOT flag
resource "aws_security_group" "internal" {
  name        = "internal"
  description = "Private service traffic"

  ingress {
    description = "Postgres from the VPC only"
    from_port   = 5432
    to_port     = 5432
    protocol    = "tcp"
    cidr_blocks = ["10.0.0.0/16"]
  }
}

# ---------- Storage ----------

resource "aws_ebs_volume" "data" {
  availability_zone = "us-east-1a"
  size              = 100
}

resource "aws_ebs_volume" "scratch" {
  availability_zone = "us-east-1a"
  size              = 20
  encrypted         = true
}

resource "aws_db_instance" "main" {
  identifier                  = "app-db"
  engine                      = "postgres"
  instance_class              = "db.t3.micro"
  allocated_storage           = 20
  username                    = "app"
  manage_master_user_password = true
  storage_encrypted           = false
  publicly_accessible         = true
  skip_final_snapshot         = true
}

# ---------- IAM ----------

resource "aws_iam_policy" "app_admin" {
  name = "app-admin"
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect   = "Allow"
      Action   = "*"
      Resource = "*"
    }]
  })
}

resource "aws_iam_policy" "read_only" {
  name = "read-only"
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [{
      Effect   = "Allow"
      Action   = ["s3:GetObject"]
      Resource = "arn:aws:s3:::acme-app-logs/*"
    }]
  })
}

resource "aws_s3_bucket" "logs" {
  bucket = "acme-app-logs"
}

# Statement is a single object (legal in IAM) and Resource references another resource,
# so Terraform cannot know the final policy text at plan time.
resource "aws_iam_policy" "log_writer" {
  name = "log-writer"
  policy = jsonencode({
    Version = "2012-10-17"
    Statement = {
      Effect   = "Allow"
      Action   = "*"
      Resource = aws_s3_bucket.logs.arn
    }
  })
}

Two details are worth a second look. The provider block uses dummy credentials plus three skip_* flags, which is why no AWS account is needed. And the comments labelled “Shape 1” to “Shape 4” mark the four different ways the AWS provider lets you open a security group. An assistant will use whichever form it saw most recently, and Step 5 depends on that.

Now let Terraform build the plan and export it as JSON. Run these from the guardrails-lab folder:

cd infra
terraform init
terraform plan -out=tfplan
terraform show -json tfplan > ../plan.json
cd ..

I added -input=false and -no-color to the first two commands so the captured text is clean; you do not need them. terraform init downloads the AWS provider (about seven seconds on my connection) and writes a lock file:

Initializing the backend...

Initializing provider plugins...
- Finding hashicorp/aws versions matching "~> 6.66"...
- Installing hashicorp/aws v6.66.0...
- Installed hashicorp/aws v6.66.0 (signed by HashiCorp)

Terraform has created a lock file .terraform.lock.hcl to record the provider
selections it made above. Include this file in your version control repository
so that Terraform can guarantee to make the same selections by default when
you run "terraform init" in the future.

Terraform has been successfully initialized!

You may now begin working with Terraform. Try running "terraform plan" to see
any changes that are required for your infrastructure. All Terraform commands
should now work.

If you ever set or change modules or backend configuration for Terraform,
rerun this command to reinitialize your working directory. If you forget, other
commands will detect it and remind you to do so if necessary.

The end of the terraform plan output confirms the plan was saved:

Plan: 14 to add, 0 to change, 0 to destroy.
Saved the plan to: tfplan

One more command is worth running, because it shows why validation alone is not enough. terraform validate checks syntax and internal consistency, and it has no opinion about safety:

cd infra
terraform validate
cd ..
Success! The configuration is valid.

The configuration with open SSH, unencrypted storage and a wildcard IAM policy is, as far as Terraform is concerned, perfectly valid.

OPA’s own Terraform guide describes the last step like this: “Use the command terraform show to convert the Terraform plan into JSON so that OPA can read the plan.” That is all plan.json is: the plan, in a format any program can read.

Step 3: Look inside the plan with opa eval

Before writing any rules, get familiar with the data. opa eval runs a Rego query against an input file. The flags we need are -f pretty (human-friendly output), -i (the input file), -d (policy or data files to load), and the query itself as the last argument. Start with the top-level keys of the plan:

opa eval -f pretty -i plan.json 'object.keys(input)'
[
  "applyable",
  "complete",
  "configuration",
  "errored",
  "format_version",
  "planned_values",
  "relevant_attributes",
  "resource_changes",
  "terraform_version",
  "timestamp"
]

Everything a guardrail cares about sits under resource_changes: one entry per resource, each with an address, a type, a mode and a change object. Save this small helper as explore.rego in the lab folder:

package explore

# The distinct resource types in the plan.
types := {rc.type | some rc in input.resource_changes}

# What Terraform knows (after) and what it cannot know yet (after_unknown) for one volume.
volume := {
	"after": rc.change.after,
	"after_unknown": rc.change.after_unknown,
} if {
	some rc in input.resource_changes
	rc.address == "aws_ebs_volume.data"
}

Two Rego ideas appear here. The first is a comprehension. {rc.type | some rc in input.resource_changes} reads as “collect rc.type for every rc in the list”, and because the result is a set there are no duplicates. The second is a rule with a body: volume is only defined if the body succeeds, meaning some resource has exactly that address. If nothing matched, OPA would print undefined instead of an object. Run both:

opa eval -f pretty -i plan.json -d explore.rego data.explore.types
opa eval -f pretty -i plan.json -d explore.rego data.explore.volume
[
  "aws_db_instance",
  "aws_ebs_volume",
  "aws_iam_policy",
  "aws_s3_bucket",
  "aws_security_group",
  "aws_security_group_rule",
  "aws_vpc_security_group_ingress_rule"
]
{
  "after": {
    "availability_zone": "us-east-1a",
    "final_snapshot": false,
    "multi_attach_enabled": null,
    "outpost_arn": null,
    "region": "us-east-1",
    "size": 100,
    "tags": null,
    "timeouts": null,
    "volume_initialization_rate": null
  },
  "after_unknown": {
    "arn": true,
    "create_time": true,
    "encrypted": true,
    "id": true,
    "iops": true,
    "kms_key_id": true,
    "snapshot_id": true,
    "tags_all": true,
    "throughput": true,
    "type": true
  }
}

The two objects that matter: after and after_unknown

The second result is the most important thing in this tutorial, so read it slowly. after holds the values Terraform already knows the resource will have. after_unknown lists the attributes whose values Terraform cannot know until the plan is applied. The Terraform JSON format documentation says the after value “will be incomplete if there are values within it that won’t be known until after apply”, and that after_unknown is an object with a similar structure whose unknown leaves are marked true.

Look at encrypted. It is not in after at all, and it appears in after_unknown. The volume’s configuration never sets it, so the provider will decide at apply time. A policy that looks for after.encrypted will not find false. It will find nothing. Keep that in mind; Step 6 turns it into a real bug.

Step 4: Write the tests first, then the first policy

Policies are code, and code that guards something important deserves tests. In OPA, a test is an ordinary rule. The policy testing guide puts it this way: “Tests are expressed as standard Rego rules with a convention that the rule name is prefixed with test_.” The keyword with input as replaces input for one evaluation, so each test can hand the policy a tiny, hand-made plan.

Save this as policy/guardrails_test.rego. The first three helper functions build minimal plans. Then come three tests: an SSH rule open to the world must produce one violation, and HTTPS from anywhere and SSH from a private network must produce none.

package guardrails_test

import data.guardrails

# Build a minimal plan that holds only the resources a test cares about.
plan(rcs) := {"resource_changes": rcs}

resource(address, type, after) := {
	"address": address,
	"type": type,
	"mode": "managed",
	"change": {"actions": ["create"], "after": after, "after_unknown": {}},
}

# An aws_security_group with a single inline ingress block.
inline_sg(cidrs, v6_cidrs, from, to, proto) := resource("aws_security_group.t", "aws_security_group", {"ingress": [{
	"cidr_blocks": cidrs,
	"ipv6_cidr_blocks": v6_cidrs,
	"from_port": from,
	"to_port": to,
	"protocol": proto,
}]})

test_inline_ssh_from_anywhere_is_denied if {
	msgs := guardrails.deny with input as plan([inline_sg(["0.0.0.0/0"], [], 22, 22, "tcp")])
	count(msgs) == 1
}

test_inline_https_from_anywhere_is_allowed if {
	msgs := guardrails.deny with input as plan([inline_sg(["0.0.0.0/0"], [], 443, 443, "tcp")])
	count(msgs) == 0
}

test_inline_private_cidr_is_allowed if {
	msgs := guardrails.deny with input as plan([inline_sg(["10.0.0.0/16"], [], 22, 22, "tcp")])
	count(msgs) == 0
}

We have no policy yet, so create a deliberately empty one. Save this as policy/guardrails.rego:

package guardrails

deny := set()

Run the tests:

opa test policy/
policy\guardrails_test.rego:24:
data.guardrails_test.test_inline_ssh_from_anywhere_is_denied: FAIL (513.5µs)
--------------------------------------------------------------------------------
PASS: 2/3
FAIL: 1/3

One test fails, as it should. The path and line number point at the failing test, and the timing will differ on your machine (mine printed Windows-style backslashes). Notice that two of the three tests already pass against a policy that does nothing at all. Tests that only check the allowed side cannot tell a good policy from an empty one, which is why the violation test is the one that matters. opa test also exits with code 2 when a test fails, so it can gate a CI job on its own.

Now write the first real policy. Replace the contents of policy/guardrails.rego with this first attempt, written directly from the inline ingress blocks we saw in main.tf:

package guardrails

# First attempt: read the inline ingress blocks we saw in main.tf.
deny contains msg if {
	some rc in input.resource_changes
	rc.type == "aws_security_group"
	some block in rc.change.after.ingress
	"0.0.0.0/0" in block.cidr_blocks
	block.to_port != 443
	msg := sprintf("%s: port %d is open to the world", [rc.address, block.to_port])
}

Read the rule body top to bottom. some rc in input.resource_changes walks every resource. rc.type == "aws_security_group" keeps only security groups. some block in rc.change.after.ingress walks that group’s inline rules, "0.0.0.0/0" in block.cidr_blocks keeps the ones open to everyone, and block.to_port != 443 excuses HTTPS. Every combination that survives all five lines adds one message to deny. Run the tests, then point the policy at the real plan:

opa test policy/
opa eval -f pretty -i plan.json -d policy/guardrails.rego data.guardrails.deny
PASS: 3/3
[
  "aws_security_group.web: port 22 is open to the world"
]

Green tests and a sensible finding. It is tempting to stop here. Do not. In main.tf we opened security groups in four different ways, and this policy found one of them.

Step 5: Find the security groups your first policy cannot see

Why did the rule find only web? Go back to the set of resource types from Step 3. Three of them describe ingress rules: aws_security_group, aws_security_group_rule and aws_vpc_security_group_ingress_rule. The provider’s own documentation admits the overlap. The aws_security_group page says to “Avoid using the ingress and egress arguments of the aws_security_group resource to configure in-line rules” and recommends the standalone aws_vpc_security_group_ingress_rule resource instead, so a modern assistant may well choose it. On top of that, an inline block can carry IPv6 addresses in a different field, ipv6_cidr_blocks, that our rule never reads.

Add tests for the shapes we missed. Append this to the end of policy/guardrails_test.rego. It covers the IPv6 form, the separate rule resource, an egress rule (which must be ignored), and both a dangerous and a harmless aws_vpc_security_group_ingress_rule:

test_inline_ipv6_from_anywhere_is_denied if {
	msgs := guardrails.deny with input as plan([inline_sg([], ["::/0"], 8080, 8080, "tcp")])
	count(msgs) == 1
}

test_rule_resource_from_anywhere_is_denied if {
	rc := resource("aws_security_group_rule.t", "aws_security_group_rule", {
		"type": "ingress",
		"cidr_blocks": ["0.0.0.0/0"],
		"ipv6_cidr_blocks": null,
		"from_port": 5432,
		"to_port": 5432,
		"protocol": "tcp",
	})
	msgs := guardrails.deny with input as plan([rc])
	count(msgs) == 1
}

test_egress_rule_resource_is_ignored if {
	rc := resource("aws_security_group_rule.t", "aws_security_group_rule", {
		"type": "egress",
		"cidr_blocks": ["0.0.0.0/0"],
		"ipv6_cidr_blocks": null,
		"from_port": 0,
		"to_port": 0,
		"protocol": "-1",
	})
	msgs := guardrails.deny with input as plan([rc])
	count(msgs) == 0
}

test_vpc_ingress_rule_all_traffic_is_denied if {
	rc := resource("aws_vpc_security_group_ingress_rule.t", "aws_vpc_security_group_ingress_rule", {
		"cidr_ipv4": "0.0.0.0/0",
		"cidr_ipv6": null,
		"ip_protocol": "-1",
		"from_port": null,
		"to_port": null,
	})
	msgs := guardrails.deny with input as plan([rc])
	count(msgs) == 1
}

test_vpc_ingress_rule_https_is_allowed if {
	rc := resource("aws_vpc_security_group_ingress_rule.t", "aws_vpc_security_group_ingress_rule", {
		"cidr_ipv4": "0.0.0.0/0",
		"cidr_ipv6": null,
		"ip_protocol": "tcp",
		"from_port": 443,
		"to_port": 443,
	})
	msgs := guardrails.deny with input as plan([rc])
	count(msgs) == 0
}
opa test policy/
policy\guardrails_test.rego:44:
data.guardrails_test.test_rule_resource_from_anywhere_is_denied: FAIL (1.5685ms)
policy\guardrails_test.rego:70:
data.guardrails_test.test_vpc_ingress_rule_all_traffic_is_denied: FAIL (1.0465ms)
policy\guardrails_test.rego:39:
data.guardrails_test.test_inline_ipv6_from_anywhere_is_denied: FAIL (1.5685ms)
--------------------------------------------------------------------------------
PASS: 5/8
FAIL: 3/8

Three tests fail, and each failure is a hole a real assistant could walk through. Fix them by normalising every ingress shape into one common form before judging it. Replace the entire contents of policy/guardrails.rego with the next two blocks, one after the other. The first holds the shared helpers:

package guardrails

# Resources that will exist once the plan is applied (created or updated).
changes contains rc if {
	some rc in input.resource_changes
	rc.mode == "managed"
	rc.change.after != null
}

# Turn a list-or-null attribute into a list so array.concat never sees null.
as_list(v) := v if is_array(v)

as_list(v) := [] if not is_array(v)

The second continues the same file with the first guardrail:

# ---------------------------------------------------------------------------
# Guardrail 1: nothing may be open to the whole internet except web ports
# ---------------------------------------------------------------------------

world_cidrs := {"0.0.0.0/0", "::/0"}

web_ports := {80, 443}

# Read a CIDR list attribute that may be missing or null.
cidrs(obj) := array.concat(
	as_list(object.get(obj, "cidr_blocks", null)),
	as_list(object.get(obj, "ipv6_cidr_blocks", null)),
)

# The AWS provider has four ways to write an ingress rule.
# Normalise all of them into {address, cidr, protocol, from, to} objects.

# Shapes 1 and 2: inline ingress blocks inside aws_security_group (IPv4 and IPv6).
ingress_rules contains rule if {
	some rc in changes
	rc.type == "aws_security_group"
	some block in rc.change.after.ingress
	some cidr in cidrs(block)
	rule := {
		"address": rc.address,
		"cidr": cidr,
		"protocol": block.protocol,
		"from": block.from_port,
		"to": block.to_port,
	}
}

# Shape 3: a separate aws_security_group_rule resource with type = "ingress".
ingress_rules contains rule if {
	some rc in changes
	rc.type == "aws_security_group_rule"
	after := rc.change.after
	after.type == "ingress"
	some cidr in cidrs(after)
	rule := {
		"address": rc.address,
		"cidr": cidr,
		"protocol": after.protocol,
		"from": after.from_port,
		"to": after.to_port,
	}
}

# Shape 4: aws_vpc_security_group_ingress_rule (one CIDR per resource).
ingress_rules contains rule if {
	some rc in changes
	rc.type == "aws_vpc_security_group_ingress_rule"
	after := rc.change.after
	some cidr in [
		object.get(after, "cidr_ipv4", null),
		object.get(after, "cidr_ipv6", null),
	]
	cidr != null
	rule := {
		"address": rc.address,
		"cidr": cidr,
		"protocol": after.ip_protocol,
		"from": object.get(after, "from_port", null),
		"to": object.get(after, "to_port", null),
	}
}

allowed(rule) if {
	rule.protocol == "tcp"
	rule.from == rule.to
	rule.to in web_ports
}

describe(rule) := sprintf("%s/%d", [rule.protocol, rule.from]) if {
	rule.from != null
	rule.from == rule.to
}

describe(rule) := sprintf("%s/%d-%d", [rule.protocol, rule.from, rule.to]) if {
	rule.from != null
	rule.from != rule.to
}

describe(rule) := sprintf("all ports (protocol %s)", [rule.protocol]) if rule.from == null

deny contains msg if {
	some rule in ingress_rules
	rule.cidr in world_cidrs
	not allowed(rule)
	msg := sprintf(
		"%s: %s is open to %s; only tcp/80 and tcp/443 may be world-open",
		[rule.address, describe(rule), rule.cidr],
	)
}

How the ingress guardrail works

changes is a set of every managed resource that has an after value, meaning the state Terraform expects the resource to be in once the plan is applied. Filtering on that up front keeps the later rules simple, because each of them can assume there is an after object to read. as_list turns a null into an empty list, and cidrs uses it so that a missing or null address list can never crash array.concat.

Then come four ingress_rules contains rule if blocks, one per shape. In Rego, several rules with the same name add up: each block contributes objects to the same set. Whichever form the Terraform used, each block produces objects with the same five fields (address, cidr, protocol, from, to), so the judging logic is written once. Shapes 1 and 2 are the inline blocks, IPv4 and IPv6 together. Shape 3 is aws_security_group_rule, which we only accept when type is "ingress" so egress rules are ignored. Shape 4 is aws_vpc_security_group_ingress_rule, which holds one address per resource.

allowed says what is acceptable from anywhere: TCP, a single port, and that port is 80 or 443. The protocol check matters. The provider documentation for aws_security_group_rule warns that setting the protocol to “all” or -1 alongside port numbers results in “a security group rule with all ports open”, so a rule with protocol -1 must never be excused by its port fields. The three describe functions only exist to make the message readable: tcp/22 for one port, tcp/1000-2000 for a range, and “all ports” when there are no ports. Finally deny combines everything: any normalised rule open to 0.0.0.0/0 or ::/0 that is not allowed becomes a message.

opa test policy/
opa eval -f pretty -i plan.json -d policy/guardrails.rego data.guardrails.deny
PASS: 8/8
[
  "aws_security_group.api: tcp/8080 is open to ::/0; only tcp/80 and tcp/443 may be world-open",
  "aws_security_group.web: tcp/22 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open",
  "aws_security_group_rule.db_from_anywhere: tcp/5432 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open",
  "aws_vpc_security_group_ingress_rule.all_traffic: all ports (protocol -1) is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open"
]

All eight tests pass, and the real plan now yields all four open security groups. The compliant internal group, which allows PostgreSQL only from 10.0.0.0/16, is correctly left alone.

Step 6: Guard storage, and meet the bug where a missing value passes

Next, storage. The rule we want is simple to say: EBS volumes and RDS databases must be encrypted, and databases must not be public. Append these tests to policy/guardrails_test.rego. Look closely at the fourth one, because it describes the real plan from Step 3, where an unset encrypted is missing from after:

test_unencrypted_rds_is_denied if {
	rc := resource("aws_db_instance.t", "aws_db_instance", {"storage_encrypted": false, "publicly_accessible": false})
	msgs := guardrails.deny with input as plan([rc])
	count(msgs) == 1
}

test_public_rds_is_denied if {
	rc := resource("aws_db_instance.t", "aws_db_instance", {"storage_encrypted": true, "publicly_accessible": true})
	msgs := guardrails.deny with input as plan([rc])
	count(msgs) == 1
}

test_encrypted_private_rds_is_allowed if {
	rc := resource("aws_db_instance.t", "aws_db_instance", {"storage_encrypted": true, "publicly_accessible": false})
	msgs := guardrails.deny with input as plan([rc])
	count(msgs) == 0
}

# In a real plan an unset "encrypted" is unknown, so the key is missing entirely.
test_ebs_without_encrypted_attribute_is_denied if {
	rc := resource("aws_ebs_volume.t", "aws_ebs_volume", {"size": 100})
	msgs := guardrails.deny with input as plan([rc])
	count(msgs) == 1
}

test_encrypted_ebs_is_allowed if {
	rc := resource("aws_ebs_volume.t", "aws_ebs_volume", {"size": 100, "encrypted": true})
	msgs := guardrails.deny with input as plan([rc])
	count(msgs) == 0
}

Now write the obvious first draft. Append these rules to the end of policy/guardrails.rego. The first two flag any attribute that is explicitly false, and the third flags public databases:

# ---------------------------------------------------------------------------
# Guardrail 2: storage must be encrypted, databases must not be public
# ---------------------------------------------------------------------------

# First draft: flag attributes that are explicitly false.
deny contains msg if {
	some rc in changes
	rc.type == "aws_db_instance"
	rc.change.after.storage_encrypted == false
	msg := sprintf("%s: storage_encrypted is false", [rc.address])
}

deny contains msg if {
	some rc in changes
	rc.type == "aws_ebs_volume"
	rc.change.after.encrypted == false
	msg := sprintf("%s: encrypted is false", [rc.address])
}
deny contains msg if {
	some rc in changes
	rc.type == "aws_db_instance"
	rc.change.after.publicly_accessible == true
	msg := sprintf("%s: database must not be publicly accessible", [rc.address])
}
opa test policy/
opa eval -f pretty -i plan.json -d policy/guardrails.rego data.guardrails.deny
policy\guardrails_test.rego:113:
data.guardrails_test.test_ebs_without_encrypted_attribute_is_denied: FAIL (1.579ms)
--------------------------------------------------------------------------------
PASS: 12/13
FAIL: 1/13
[
  "aws_db_instance.main: database must not be publicly accessible",
  "aws_db_instance.main: storage_encrypted is false",
  "aws_security_group.api: tcp/8080 is open to ::/0; only tcp/80 and tcp/443 may be world-open",
  "aws_security_group.web: tcp/22 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open",
  "aws_security_group_rule.db_from_anywhere: tcp/5432 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open",
  "aws_vpc_security_group_ingress_rule.all_traffic: all ports (protocol -1) is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open"
]

The failing test is the one about the volume with no encrypted setting, and the plan output shows the same thing in the real world: aws_db_instance.main is flagged, and aws_ebs_volume.data, an unencrypted volume, sails through. The draft asked “is it false?”, and a missing value is not false. The OPA policy language reference says so directly: “Expressions that refer to undefined values are also undefined. This includes comparisons such as !=.” The whole rule body becomes undefined, and an undefined rule body adds nothing to deny. Rewriting the check as != true would not help either, for the same reason.

The fix is to use not. The expression not x == true succeeds whenever x == true is false or undefined, which is exactly the behaviour we want: an attribute that is missing, unknown or false counts as a violation, and only an explicit true passes. That is failing closed. Delete the two “first draft” rules and put this table-driven version in their place, keeping the public-database rule below it:

# ---------------------------------------------------------------------------
# Guardrail 2: storage must be encrypted, databases must not be public
# ---------------------------------------------------------------------------

encryption_attribute := {
	"aws_ebs_volume": "encrypted",
	"aws_db_instance": "storage_encrypted",
}

# "not (x == true)" also fires when x is missing, so an attribute that is
# absent or unknown counts as a violation instead of silently passing.
deny contains msg if {
	some rc in changes
	attr := encryption_attribute[rc.type]
	not rc.change.after[attr] == true
	msg := sprintf("%s: %s must be explicitly set to true", [rc.address, attr])
}

The encryption_attribute object is a small lookup table from resource type to the attribute that must be true. The expression attr := encryption_attribute[rc.type] is itself undefined for any other resource type, so those resources are skipped automatically. Adding a third storage type later means adding one line to the table.

opa test policy/
opa eval -f pretty -i plan.json -d policy/guardrails.rego data.guardrails.deny
PASS: 13/13
[
  "aws_db_instance.main: database must not be publicly accessible",
  "aws_db_instance.main: storage_encrypted must be explicitly set to true",
  "aws_ebs_volume.data: encrypted must be explicitly set to true",
  "aws_security_group.api: tcp/8080 is open to ::/0; only tcp/80 and tcp/443 may be world-open",
  "aws_security_group.web: tcp/22 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open",
  "aws_security_group_rule.db_from_anywhere: tcp/5432 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open",
  "aws_vpc_security_group_ingress_rule.all_traffic: all ports (protocol -1) is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open"
]

The unencrypted database is now reported twice (once for encryption, once for being public), and aws_ebs_volume.data is caught. The properly encrypted scratch volume is not reported.

Step 7: Guard IAM, and handle values Terraform cannot know yet

The third guardrail refuses any IAM policy that allows every action. Two details make this one trickier than it sounds. First, an IAM policy is a JSON document stored as a string, so Rego has to parse it with json.unmarshal. Second, IAM is flexible about shape. The AWS documentation says that “The Statement element can contain a single statement or an array of individual statements” (see IAM JSON policy elements: Statement), and the policy grammar adds that where an element supports an array “but only one value is included, the brackets are optional”. So Statement and Action can each be a single value or a list, and our rule has to accept both.

Append these tests to policy/guardrails_test.rego. The last one is the important one: it describes a policy whose document is unknown until apply, and it expects a violation:

iam_policy(doc) := resource("aws_iam_policy.t", "aws_iam_policy", {"policy": json.marshal(doc)})

test_wildcard_action_in_statement_list_is_denied if {
	doc := {"Version": "2012-10-17", "Statement": [{"Effect": "Allow", "Action": "*", "Resource": "*"}]}
	msgs := guardrails.deny with input as plan([iam_policy(doc)])
	count(msgs) == 1
}

test_wildcard_action_in_single_statement_object_is_denied if {
	doc := {"Version": "2012-10-17", "Statement": {"Effect": "Allow", "Action": ["*"], "Resource": "arn:aws:s3:::logs"}}
	msgs := guardrails.deny with input as plan([iam_policy(doc)])
	count(msgs) == 1
}

test_scoped_action_is_allowed if {
	doc := {"Version": "2012-10-17", "Statement": [{"Effect": "Allow", "Action": ["s3:GetObject"], "Resource": "arn:aws:s3:::logs/*"}]}
	msgs := guardrails.deny with input as plan([iam_policy(doc)])
	count(msgs) == 0
}

test_deny_statement_with_wildcard_is_allowed if {
	doc := {"Version": "2012-10-17", "Statement": [{"Effect": "Deny", "Action": "*", "Resource": "*"}]}
	msgs := guardrails.deny with input as plan([iam_policy(doc)])
	count(msgs) == 0
}

test_unknown_policy_fails_closed if {
	rc := object.union(
		resource("aws_iam_policy.t", "aws_iam_policy", {"name": "t"}),
		{"change": {"actions": ["create"], "after": {"name": "t"}, "after_unknown": {"policy": true}}},
	)
	msgs := guardrails.deny with input as plan([rc])
	count(msgs) == 1
}

Append the first IAM rule to the end of policy/guardrails.rego:

# ---------------------------------------------------------------------------
# Guardrail 3: no IAM policy may allow every action
# ---------------------------------------------------------------------------

# IAM accepts either one statement object or a list of them, and either one
# action string or a list of actions. Normalise both.
as_statements(s) := s if is_array(s)

as_statements(s) := [s] if is_object(s)

as_actions(a) := a if is_array(a)

as_actions(a) := [a] if is_string(a)

deny contains msg if {
	some rc in changes
	rc.type == "aws_iam_policy"
	doc := json.unmarshal(rc.change.after.policy)
	some stmt in as_statements(doc.Statement)
	stmt.Effect == "Allow"
	"*" in as_actions(stmt.Action)
	msg := sprintf("%s: policy allows every action (Action \"*\")", [rc.address])
}
opa test policy/
opa eval -f pretty -i plan.json -d policy/guardrails.rego data.guardrails.deny
policy\guardrails_test.rego:151:
data.guardrails_test.test_unknown_policy_fails_closed: FAIL (1.929ms)
--------------------------------------------------------------------------------
PASS: 17/18
FAIL: 1/18
[
  "aws_db_instance.main: database must not be publicly accessible",
  "aws_db_instance.main: storage_encrypted must be explicitly set to true",
  "aws_ebs_volume.data: encrypted must be explicitly set to true",
  "aws_iam_policy.app_admin: policy allows every action (Action \"*\")",
  "aws_security_group.api: tcp/8080 is open to ::/0; only tcp/80 and tcp/443 may be world-open",
  "aws_security_group.web: tcp/22 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open",
  "aws_security_group_rule.db_from_anywhere: tcp/5432 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open",
  "aws_vpc_security_group_ingress_rule.all_traffic: all ports (protocol -1) is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open"
]

Every test passes except the unknown-policy one, and the plan output shows why it matters. aws_iam_policy.app_admin is reported, but aws_iam_policy.log_writer, which also allows every action, is silent. Its policy document references the bucket’s ARN, and Terraform cannot know an ARN before the bucket exists. Look at what the plan holds for it. Add this helper rule to the end of explore.rego and run it:

# A policy document that points at another resource is unknown until apply.
writer := {
	"after": rc.change.after,
	"after_unknown": rc.change.after_unknown,
} if {
	some rc in input.resource_changes
	rc.address == "aws_iam_policy.log_writer"
}
opa eval -f pretty -i plan.json -d explore.rego data.explore.writer
{
  "after": {
    "delay_after_policy_creation_in_ms": null,
    "description": null,
    "name": "log-writer",
    "path": "/",
    "tags": null
  },
  "after_unknown": {
    "arn": true,
    "attachment_count": true,
    "id": true,
    "name_prefix": true,
    "policy": true,
    "policy_id": true,
    "tags_all": true
  }
}

There is no policy key in after, and after_unknown says "policy": true. The rule above called json.unmarshal(rc.change.after.policy) on a value that does not exist, so its body was undefined and it skipped the resource. No error, no warning, just a pass. This is the same failure as the unencrypted volume, in a different costume.

Fail closed on unknown values

A guardrail cannot approve what it cannot read. Append this second IAM rule to the end of policy/guardrails.rego. It reports any policy document that Terraform marks as unknown:

# Fail closed: a document Terraform cannot compute until apply cannot be checked,
# and json.unmarshal on a missing value would otherwise skip the resource silently.
deny contains msg if {
	some rc in changes
	rc.type == "aws_iam_policy"
	rc.change.after_unknown.policy == true
	msg := sprintf("%s: policy document is not known until apply, so it cannot be checked", [rc.address])
}
opa test policy/
opa eval -f pretty -i plan.json -d policy/guardrails.rego data.guardrails.deny
PASS: 18/18
[
  "aws_db_instance.main: database must not be publicly accessible",
  "aws_db_instance.main: storage_encrypted must be explicitly set to true",
  "aws_ebs_volume.data: encrypted must be explicitly set to true",
  "aws_iam_policy.app_admin: policy allows every action (Action \"*\")",
  "aws_iam_policy.log_writer: policy document is not known until apply, so it cannot be checked",
  "aws_security_group.api: tcp/8080 is open to ::/0; only tcp/80 and tcp/443 may be world-open",
  "aws_security_group.web: tcp/22 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open",
  "aws_security_group_rule.db_from_anywhere: tcp/5432 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open",
  "aws_vpc_security_group_ingress_rule.all_traffic: all ports (protocol -1) is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open"
]

All 18 tests pass and the plan produces nine messages for eight resources. Every problem we planted is found, and none of the four compliant resources is reported. The practical fix for an unverifiable policy is to make the value known at plan time, as the corrected configuration in Step 9 does by building the ARN from a local value instead of pointing at the bucket.

Step 8: Format it, lint it strictly, and measure what the tests never ran

Three cheap commands keep a policy repository healthy. opa fmt rewrites files into the canonical layout (with --diff it only shows what it would change), and opa check --strict compiles the policy and rejects sloppy constructs. Both print nothing when everything is fine:

opa fmt --diff policy/
opa check --strict policy/

Both commands printed nothing and exited with 0 on the files above. To see what strict mode catches, save this deliberately untidy rule as demo/unused_variable.rego:

package demo

deny contains msg if {
	some rc in input.resource_changes
	rc.type == "aws_db_instance"
	unused := rc.address
	msg := "found a database"
}
opa check --strict demo/unused_variable.rego
1 error occurred: demo/unused_variable.rego:6: rego_compile_error: assigned var unused unused

The variable unused is assigned and never read, which is often a sign of a half-finished edit. Strict mode turns that into an error.

Coverage: which lines did the tests never run?

opa test can report coverage. The testing guide explains how to read it: “If the line refers to the head of a rule, the body of the rule was never true. If the line refers to an expression in a rule, the expression was never evaluated.” The raw report is a large JSON document, so save this small script as coverage_report.py to summarise it:

import json
import sys

report = json.load(sys.stdin)
for path, info in report["files"].items():
    if path.endswith("_test.rego"):
        continue
    rows = sorted({item["start"]["row"] for item in info.get("not_covered", [])})
    print(f"{path}: {info['coverage']}% covered; lines never run: {rows or 'none'}")
opa test policy/ --coverage --format=json | python coverage_report.py

On macOS and Linux, use python3 if python is not found. The summary for our 18 tests:

policy\guardrails.rego: 99.02912621359224% covered; lines never run: [93]

Our 18 tests cover 99 percent of the policy, and line 93 was never run. That line sits inside the describe branch that formats a port range such as tcp/1000-2000. Every test so far used a single port or all ports. Add one more test to the end of policy/guardrails_test.rego:

test_port_range_from_anywhere_is_denied if {
	msgs := guardrails.deny with input as plan([inline_sg(["0.0.0.0/0"], [], 1000, 2000, "tcp")])
	count(msgs) == 1
	some msg in msgs
	contains(msg, "tcp/1000-2000")
}
opa test policy/
opa test policy/ --coverage --format=json | python coverage_report.py
PASS: 19/19
policy\guardrails.rego: 100% covered; lines never run: none

Nineteen passing tests and full line coverage. Be honest about what that number means, though. Coverage tells you which lines your tests never ran. It cannot tell you about a shape you never thought to write a test for; Step 5 found that gap by reading the real plan. Run opa test policy/ -v whenever you want to see every test by name.

Step 9: Turn the policy into a gate that a pipeline can trust

A guardrail that only prints messages is a suggestion. A pipeline needs an exit code: zero for “go ahead”, non-zero for “stop”. OPA’s CLI reference describes two flags for opa eval. --fail “Exits with non-zero exit code on undefined/empty result and errors”, and --fail-defined “Exits with non-zero exit code on defined/non-empty result and errors”. Since we want to stop when a violation exists, --fail-defined is the natural fit. The query should return something only when there is at least one violation, so iterate over the set with [_], where the underscore means “any element”. I used -f raw so that each message is printed on its own line:

opa eval --fail-defined -f raw -d policy/guardrails.rego -i plan.json 'data.guardrails.deny[_]'
echo $?
aws_db_instance.main: database must not be publicly accessible
aws_db_instance.main: storage_encrypted must be explicitly set to true
aws_ebs_volume.data: encrypted must be explicitly set to true
aws_iam_policy.app_admin: policy allows every action (Action "*")
aws_iam_policy.log_writer: policy document is not known until apply, so it cannot be checked
aws_security_group.api: tcp/8080 is open to ::/0; only tcp/80 and tcp/443 may be world-open
aws_security_group.web: tcp/22 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open
aws_security_group_rule.db_from_anywhere: tcp/5432 is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open
aws_vpc_security_group_ingress_rule.all_traffic: all ports (protocol -1) is open to 0.0.0.0/0; only tcp/80 and tcp/443 may be world-open
1

Nine messages and exit code 1: the gate is closed. Now give the assistant the feedback loop it never had. Here is the corrected configuration as a unified diff (lines starting with - are removed and lines starting with + are added). Every security rule is narrowed to the private network, the storage is encrypted, the database is private, and the IAM policies are scoped. The bucket name also moves into a local value so that the log writer’s policy document is fully known at plan time:

--- a/main.tf
+++ b/main.tf
@@ -19,2 +19,6 @@
 
+locals {
+  log_bucket = "acme-app-logs"
+}
+
 # ---------- Networking: four different ways to open a security group ----------
@@ -35,3 +39,3 @@
   ingress {
-    description = "SSH so the assistant can debug"
+    description = "SSH from the VPC only"
     from_port   = 22
@@ -39,3 +43,3 @@
     protocol    = "tcp"
-    cidr_blocks = ["0.0.0.0/0"]
+    cidr_blocks = ["10.0.0.0/16"]
   }
@@ -49,7 +53,7 @@
   ingress {
-    description      = "API from anywhere over IPv6"
-    from_port        = 8080
-    to_port          = 8080
-    protocol         = "tcp"
-    ipv6_cidr_blocks = ["::/0"]
+    description = "API from the VPC only"
+    from_port   = 8080
+    to_port     = 8080
+    protocol    = "tcp"
+    cidr_blocks = ["10.0.0.0/16"]
   }
@@ -63,3 +67,3 @@
 
-resource "aws_security_group_rule" "db_from_anywhere" {
+resource "aws_security_group_rule" "db_from_vpc" {
   type              = "ingress"
@@ -68,3 +72,3 @@
   protocol          = "tcp"
-  cidr_blocks       = ["0.0.0.0/0"]
+  cidr_blocks       = ["10.0.0.0/16"]
   security_group_id = aws_security_group.db.id
@@ -78,6 +82,8 @@
 
-resource "aws_vpc_security_group_ingress_rule" "all_traffic" {
+resource "aws_vpc_security_group_ingress_rule" "admin_https" {
   security_group_id = aws_security_group.admin.id
-  cidr_ipv4         = "0.0.0.0/0"
-  ip_protocol       = "-1"
+  cidr_ipv4         = "10.0.0.0/16"
+  ip_protocol       = "tcp"
+  from_port         = 443
+  to_port           = 443
 }
@@ -103,2 +109,3 @@
   size              = 100
+  encrypted         = true
 }
@@ -118,4 +125,4 @@
   manage_master_user_password = true
-  storage_encrypted           = false
-  publicly_accessible         = true
+  storage_encrypted           = true
+  publicly_accessible         = false
   skip_final_snapshot         = true
@@ -131,4 +138,4 @@
       Effect   = "Allow"
-      Action   = "*"
-      Resource = "*"
+      Action   = ["s3:GetObject", "s3:PutObject"]
+      Resource = "arn:aws:s3:::${local.log_bucket}/*"
     }]
@@ -150,3 +157,3 @@
 resource "aws_s3_bucket" "logs" {
-  bucket = "acme-app-logs"
+  bucket = local.log_bucket
 }
@@ -161,4 +168,4 @@
       Effect   = "Allow"
-      Action   = "*"
-      Resource = aws_s3_bucket.logs.arn
+      Action   = ["s3:PutObject"]
+      Resource = "arn:aws:s3:::${local.log_bucket}/*"
     }

Keep a copy of the original as infra/main.tf.original, apply those changes to infra/main.tf, and generate a second plan file next to the first. The old plan stays untouched, so you can compare both:

cd infra
terraform plan -out=tfplan_fixed
terraform show -json tfplan_fixed > ../plan_fixed.json
cd ..
opa eval --fail-defined -f raw -d policy/guardrails.rego -i plan_fixed.json 'data.guardrails.deny[_]'
echo $?

The corrected plan still adds 14 resources. The gate printed nothing at all, and the exit code was:

0

The two traps that make a gate lie

The exact shape of that query matters, and two tempting variations are wrong in opposite directions. The first is to query the whole set instead of iterating over it:

opa eval --fail-defined -f pretty -d policy/guardrails.rego -i plan_fixed.json data.guardrails.deny
echo $?
[]
1

On the clean plan the set is empty, but an empty set is still a defined value, so on OPA 1.21.1 --fail-defined exited with 1 and would have blocked a perfectly good plan. The second trap goes the other way. It tries to be clever with a boolean:

opa eval --fail -f pretty -d policy/guardrails.rego -i plan.json 'count(data.guardrails.deny) == 0'
echo $?
false
0

This ran against the broken plan with nine violations. The comparison evaluates to false, and false is a defined, non-empty result, so --fail saw nothing to complain about and exited with 0. A gate built that way approves everything. The pattern to remember is --fail-defined with a query that yields a value only when something is wrong.

Telling a violation apart from a broken tool

A gate should also fail loudly when OPA itself cannot do its job, and it should let you tell that apart from a genuine violation. On this version the two cases exit differently. A violation gave exit code 1 above. An unreadable input file and a syntax error both give 2. First a plan file that does not exist:

opa eval --fail-defined -f pretty -d policy/guardrails.rego -i no-such-plan.json 'data.guardrails.deny[_]'
echo $?
open no-such-plan.json: The system cannot find the file specified.
2

The wording of that first line comes from the operating system, so yours will differ. Now a policy written in the old syntax. Save this as demo/broken_syntax.rego:

package demo

deny contains msg {
	msg := "missing the if keyword"
}
opa eval --fail-defined -f pretty -d demo/broken_syntax.rego -i plan.json 'data.demo.deny[_]'
echo $?
1 error occurred: demo/broken_syntax.rego:3: rego_parse_error: `if` keyword is required before rule body
2

In a CI job, any non-zero code stops the pipeline, which is what you want. The distinction is useful for reporting: 1 means “your plan broke a rule”, and 2 means “fix the pipeline”.

Step 10: Let a model draft the policy, and let the checker decide

Since assistants write your Terraform, it is natural to ask one to write the policy as well. The safe way to accept that help is the same discipline we used above: the parser and your tests decide, not the model. To try it, I asked two small local models, qwen2.5:1.5b and qwen3.5:4b (thinking switched off for the second), to write one rule: deny any aws_db_instance whose storage_encrypted is not true. I used temperature 0 and a fixed seed, and each model got two prompts, a plain one and one that adds a hint about the new Rego syntax. Save this as ask_model.py. It needs pip install requests, a running Ollama server with both models pulled, and opa on your PATH:

import pathlib
import re
import subprocess

import requests

PROMPT = (
    "Write an Open Policy Agent policy in Rego for a Terraform plan in JSON "
    "(the output of `terraform show -json`). Use `package guardrails`. "
    "Define a rule named `deny` that is a set of message strings. "
    "Add one message for every entry in `input.resource_changes` whose `type` is "
    "`aws_db_instance` and whose `change.after.storage_encrypted` is not true. "
    "Reply with the Rego code only."
)
HINT = (
    " This is OPA 1.x, where every rule body needs the `if` keyword and a set is "
    "built with `contains`, for example: `deny contains msg if { ... }`."
)


def ask(model, prompt, **extra):
    body = {
        "model": model,
        "stream": False,
        "messages": [{"role": "user", "content": prompt}],
        "options": {"temperature": 0, "seed": 42, "num_predict": 700},
        **extra,
    }
    reply = requests.post("http://127.0.0.1:11434/api/chat", json=body, timeout=600).json()
    text = reply["message"]["content"]
    fenced = re.search(r"```(?:rego)?\s*\n(.*?)```", text, re.S)
    return (fenced.group(1) if fenced else text).strip() + "\n"


for model, extra in [("qwen2.5:1.5b", {}), ("qwen3.5:4b", {"think": False})]:
    for label, prompt in [("plain", PROMPT), ("hinted", PROMPT + HINT)]:
        folder = pathlib.Path("model_runs") / f"{model.replace(':', '-')}-{label}"
        folder.mkdir(parents=True, exist_ok=True)
        (folder / "policy.rego").write_text(ask(model, prompt, **extra), encoding="utf-8", newline="\n")
        check = subprocess.run(
            ["opa", "check", "--strict", "policy.rego"],
            cwd=folder, capture_output=True, text=True,
        )
        print(f"=== {model} ({label} prompt) ===")
        print((folder / "policy.rego").read_text(encoding="utf-8"))
        print("$ opa check --strict policy.rego")
        print((check.stdout + check.stderr).strip() or "(no errors)")
        print(f"exit code: {check.returncode}\n")
python ask_model.py

All four drafts failed to parse. Here are two of them in full. First the 1.5-billion-parameter model with the plain prompt:

=== qwen2.5:1.5b (plain prompt) ===
package guardrails

deny = [
    "Storage encryption must be enabled.",
]

for _, change := range input.resource_changes {
    if change.type == "aws_db_instance" && !change.change.after.storage_encrypted {
        deny += `Storage encryption must be enabled on ${change.change.after.instance_identifier}.`
    }
}

$ opa check --strict policy.rego
1 error occurred during loading: policy.rego:7: rego_parse_error: unexpected , token
	for _, change := range input.resource_changes {
	     ^
exit code: 1

This is not Rego at all. It looks like a mixture of Go and JavaScript, with a for loop and a += operator that the language does not have. The 4-billion-parameter model with the hinted prompt did better on vocabulary and worse on structure:

=== qwen3.5:4b (hinted prompt) ===
package guardrails

deny := [
    "Deny creation of unencrypted AWS DB instance",
]

deny_contains_msg(msg) {
  deny contains msg
}

denied_resource_changes(input, output) {
  input.resource_changes as changes
  count(changes[_].type == "aws_db_instance" and not _["change"]["after"]["storage_encrypted"]) > 0
}

$ opa check --strict policy.rego
1 error occurred during loading: policy.rego:8: rego_parse_error: unexpected contains keyword: expected \n or ; or }
	  deny contains msg
	       ^
exit code: 1

It used the contains keyword the hint suggested, but as a stray statement inside a function body, in the old brace-only style. The other two attempts failed the same way: the 1.5B model with the hinted prompt tried a && operator and a deny.add(...) method call, and the 4B model with the plain prompt tried an as keyword after a comparison. None of the four reached the test suite, because opa check rejected them first.

Two honest caveats. These are small models running locally, and a larger model may well write valid Rego more often. And the failure that opa check catches is the easy one. The dangerous draft is the one that parses cleanly and is quietly wrong, like the “is it false?” rule from Step 6. Tests written from the real plan are what catch that kind. So whoever writes your policy, a person or a model, the workflow stays the same: run opa fmt, opa check --strict and opa test, and accept the change only when all three are clean.

Verify the whole thing end to end

Run these from the guardrails-lab folder and compare with what you should see:

  1. opa version reports Rego Version: v1.
  2. terraform plan -out=tfplan inside infra/ ends with Plan: 14 to add, 0 to change, 0 to destroy.
  3. opa test policy/ prints PASS: 19/19.
  4. opa fmt --diff policy/ and opa check --strict policy/ print nothing.
  5. The gate on plan.json prints nine messages, one per line, and echo $? prints 1.
  6. The gate on plan_fixed.json prints nothing, and echo $? prints 0.

If step 3 fails, check that you replaced the storage “first draft” rules rather than leaving both versions in the file, and that guardrails_test.rego holds the five test blocks from Steps 4 to 8 in order. The test file only ever grows, so it is simply those five blocks one after another. The complete policy file appears in the next section for comparison.

Mistakes worth avoiding

Writing old-style Rego. A rule body without if fails with `if` keyword is required before rule body, as the syntax-error demo above shows. Snippets from older blog posts and older models need updating to deny contains msg if { ... }.

Judging a plan you did not regenerate. The policy reads plan.json, not your .tf files. After every edit to the Terraform, run terraform plan and terraform show -json again, or you are checking yesterday’s infrastructure.

Testing only what should pass. Two of the three first tests passed against a policy that did nothing. Always include a case that must be denied, and one for each shape of input you can find in a real plan.

Treating undefined as safe. The silent failures in Steps 6 and 7 had the same root cause: an expression that referred to something missing, so the rule quietly produced nothing. When a value could be absent or unknown, decide what should happen, and write it down as a rule.

What this policy cannot do

A plan-based gate is powerful but narrow, and it is worth being clear about the edges:

  • It sees the plan, and only the plan. Resources created outside Terraform, or drift that happened after the plan, are invisible to it.
  • Unknown values stay unknown. Failing closed, as we did for IAM documents, avoids false approvals, but you will occasionally have to restructure Terraform so the value is known at plan time.
  • It compares exact address strings. The policy matches 0.0.0.0/0 and ::/0 literally. I fed it a plan that opens port 22 to 0.0.0.0/1 and 128.0.0.0/1, which together cover the whole IPv4 internet, and it reported nothing:
{
  "resource_changes": [
    {
      "address": "aws_security_group_rule.split",
      "type": "aws_security_group_rule",
      "mode": "managed",
      "change": {
        "actions": [
          "create"
        ],
        "after": {
          "type": "ingress",
          "cidr_blocks": [
            "0.0.0.0/1",
            "128.0.0.0/1"
          ],
          "ipv6_cidr_blocks": null,
          "from_port": 22,
          "to_port": 22,
          "protocol": "tcp"
        },
        "after_unknown": {}
      }
    }
  ]
}
opa eval -f pretty -i demo/split_cidr_plan.json -d policy/guardrails.rego data.guardrails.deny
[]
  • It only knows the controls you wrote. Nothing here checks S3 buckets, load balancers or the instance metadata service. A green gate means “none of these three rules were broken”, not “this infrastructure is secure”.
  • Provider shapes change. The rules key off attribute names from AWS provider 6.66.0. A new resource type or renamed argument becomes a new blind spot until someone writes a test for it.

The complete policy file

This is the final policy/guardrails.rego: the helpers, the ingress guardrail, the storage guardrails and the two IAM rules, in order. It is 166 lines, and opa fmt --diff reports no changes for it.

package guardrails

# Resources that will exist once the plan is applied (created or updated).
changes contains rc if {
	some rc in input.resource_changes
	rc.mode == "managed"
	rc.change.after != null
}

# Turn a list-or-null attribute into a list so array.concat never sees null.
as_list(v) := v if is_array(v)

as_list(v) := [] if not is_array(v)

# ---------------------------------------------------------------------------
# Guardrail 1: nothing may be open to the whole internet except web ports
# ---------------------------------------------------------------------------

world_cidrs := {"0.0.0.0/0", "::/0"}

web_ports := {80, 443}

# Read a CIDR list attribute that may be missing or null.
cidrs(obj) := array.concat(
	as_list(object.get(obj, "cidr_blocks", null)),
	as_list(object.get(obj, "ipv6_cidr_blocks", null)),
)

# The AWS provider has four ways to write an ingress rule.
# Normalise all of them into {address, cidr, protocol, from, to} objects.

# Shapes 1 and 2: inline ingress blocks inside aws_security_group (IPv4 and IPv6).
ingress_rules contains rule if {
	some rc in changes
	rc.type == "aws_security_group"
	some block in rc.change.after.ingress
	some cidr in cidrs(block)
	rule := {
		"address": rc.address,
		"cidr": cidr,
		"protocol": block.protocol,
		"from": block.from_port,
		"to": block.to_port,
	}
}

# Shape 3: a separate aws_security_group_rule resource with type = "ingress".
ingress_rules contains rule if {
	some rc in changes
	rc.type == "aws_security_group_rule"
	after := rc.change.after
	after.type == "ingress"
	some cidr in cidrs(after)
	rule := {
		"address": rc.address,
		"cidr": cidr,
		"protocol": after.protocol,
		"from": after.from_port,
		"to": after.to_port,
	}
}

# Shape 4: aws_vpc_security_group_ingress_rule (one CIDR per resource).
ingress_rules contains rule if {
	some rc in changes
	rc.type == "aws_vpc_security_group_ingress_rule"
	after := rc.change.after
	some cidr in [
		object.get(after, "cidr_ipv4", null),
		object.get(after, "cidr_ipv6", null),
	]
	cidr != null
	rule := {
		"address": rc.address,
		"cidr": cidr,
		"protocol": after.ip_protocol,
		"from": object.get(after, "from_port", null),
		"to": object.get(after, "to_port", null),
	}
}

allowed(rule) if {
	rule.protocol == "tcp"
	rule.from == rule.to
	rule.to in web_ports
}

describe(rule) := sprintf("%s/%d", [rule.protocol, rule.from]) if {
	rule.from != null
	rule.from == rule.to
}

describe(rule) := sprintf("%s/%d-%d", [rule.protocol, rule.from, rule.to]) if {
	rule.from != null
	rule.from != rule.to
}

describe(rule) := sprintf("all ports (protocol %s)", [rule.protocol]) if rule.from == null

deny contains msg if {
	some rule in ingress_rules
	rule.cidr in world_cidrs
	not allowed(rule)
	msg := sprintf(
		"%s: %s is open to %s; only tcp/80 and tcp/443 may be world-open",
		[rule.address, describe(rule), rule.cidr],
	)
}

# ---------------------------------------------------------------------------
# Guardrail 2: storage must be encrypted, databases must not be public
# ---------------------------------------------------------------------------

encryption_attribute := {
	"aws_ebs_volume": "encrypted",
	"aws_db_instance": "storage_encrypted",
}

# "not (x == true)" also fires when x is missing, so an attribute that is
# absent or unknown counts as a violation instead of silently passing.
deny contains msg if {
	some rc in changes
	attr := encryption_attribute[rc.type]
	not rc.change.after[attr] == true
	msg := sprintf("%s: %s must be explicitly set to true", [rc.address, attr])
}

deny contains msg if {
	some rc in changes
	rc.type == "aws_db_instance"
	rc.change.after.publicly_accessible == true
	msg := sprintf("%s: database must not be publicly accessible", [rc.address])
}

# ---------------------------------------------------------------------------
# Guardrail 3: no IAM policy may allow every action
# ---------------------------------------------------------------------------

# IAM accepts either one statement object or a list of them, and either one
# action string or a list of actions. Normalise both.
as_statements(s) := s if is_array(s)

as_statements(s) := [s] if is_object(s)

as_actions(a) := a if is_array(a)

as_actions(a) := [a] if is_string(a)

deny contains msg if {
	some rc in changes
	rc.type == "aws_iam_policy"
	doc := json.unmarshal(rc.change.after.policy)
	some stmt in as_statements(doc.Statement)
	stmt.Effect == "Allow"
	"*" in as_actions(stmt.Action)
	msg := sprintf("%s: policy allows every action (Action \"*\")", [rc.address])
}

# Fail closed: a document Terraform cannot compute until apply cannot be checked,
# and json.unmarshal on a missing value would otherwise skip the resource silently.
deny contains msg if {
	some rc in changes
	rc.type == "aws_iam_policy"
	rc.change.after_unknown.policy == true
	msg := sprintf("%s: policy document is not known until apply, so it cannot be checked", [rc.address])
}

Where to go next

You now have a small but honest policy-as-code pipeline. Here are natural next steps:

  • Write the next guardrail yourself. Require the instance metadata service to use session tokens. In the aws_instance documentation, the http_tokens argument accepts optional or required. Follow the same loop: read a real plan, write the failing test, then the rule. Our tutorial on preventing server-side request forgery explains why that setting matters.
  • Run the gate in CI safely. The gate is only as trustworthy as the pipeline that runs it, so pin the actions that install OPA and Terraform, as described in our guide to pinning and verifying GitHub Actions.
  • Apply the same idea to code review. Guardrails work best when they are layered. See how to review AI-generated Python code before you merge it.
  • Move the same principle into Kubernetes. Policy engines can also run at admission time. Our coverage of Kyverno as a platform primitive is a good starting point.
  • Look at policy outside the agent. The same “enforce it where the model cannot reach it” principle is behind NVIDIA’s open agent safety platform.

Sources and further reading

  • OPA policy language reference: undefined values, negation and rule semantics.
  • OPA policy testing guide: test format, with and coverage.
  • OPA CLI reference: opa eval flags, including --fail and --fail-defined.
  • OPA and Terraform and the OPA v0 to v1 upgrade guide.
  • Terraform JSON output format: the after and after_unknown fields.
  • The AWS provider documentation for aws_security_group and aws_security_group_rule.
  • AWS IAM: the Statement element and the policy grammar.
  • Kayode Adeniyi’s handbook on governing AI-generated infrastructure with policy as code and OPA (freeCodeCamp, September 29, 2026) prompted this walkthrough. The scenario, policies, tests and every output here are original and were produced on OPA 1.21.1, Terraform 1.16.4 and AWS provider 6.66.0.

Tags:

AI Coding AgentsCloud SecurityDevSecOpsOpen Policy AgentPolicy as CodeTerraform

Share

Microscope die shot of an AMD EPYC 7702 engineering sample I/O die, its circuit blocks glowing in teal, gold and violet
Previous Post

AMD Agrees to Buy Fei-Fei Li’s World Labs for $8.2 Billion to Steer Its Chip Roadmap

Eight silver hex keys of graduated sizes fanned out on a steel ring against a dark green surface
Next Post

Attackers Exploit a Hex-Encoding Bypass in Cisco SD-WAN Manager, and CISA Sets an October 3 Deadline

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
30 Sep
How to Detect and Strip Invisible Unicode in Python to Stop ASCII Smuggling and Trojan Source
30 Sep
Two Cloudflare Agent Billing Betas Turn Web Monetization Into a Question of Who Holds the Meter
Trending
September 30, 2026
How to Detect and Strip Invisible Unicode in Python to Stop ASCII Smuggling and Trojan Source
September 30, 2026
Two Cloudflare Agent Billing Betas Turn Web Monetization Into a Question of Who Holds the Meter
September 30, 2026
Attackers Exploit a Hex-Encoding Bypass in Cisco SD-WAN Manager, and CISA Sets an October 3 Deadline
September 30, 2026
How to Enforce Guardrails on AI-Generated Terraform With Open Policy Agent and Rego
September 30, 2026
AMD Agrees to Buy Fei-Fei Li’s World Labs for $8.2 Billion to Steer Its Chip Roadmap
September 29, 2026
How to Build a Merkle Tree Certificate Issuer in Python to Keep Post-Quantum Certificates Small

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026