AI Vulnerability Harnesses: A Production Checklist for Security Teams
A practical checklist for building an AI-assisted vulnerability discovery harness with state, deduplication, evidence, and human triage gates.
Cloudflare’s new vulnerability-harness writeup is useful because it moves AI security scanning out of the prompt-demo phase and into the operating model phase. The interesting question is no longer whether a frontier model can find a bug in a repository. The production question is whether a security team can preserve state, compare hypotheses, remove false positives, and route a small number of evidence-backed findings to humans who can fix the right things.
Table Of Content
- Start with a finding lifecycle, not a model benchmark
- Minimum state fields
- Make deduplication a first-class system
- Use a shared weakness language
- Do not let the taxonomy hide uncertainty
- Separate discovery from validation
- A validation stack should include more than another prompt
- Keep outputs portable and reviewable
- Export records should be boring
- Anchor testing in a maintained security guide
- Production release gates
- Do not deploy a vulnerability harness until these gates pass
- Bottom line
- Sources
In its June 18, 2026 post, Cloudflare describes a model-agnostic vulnerability discovery harness built around state controls, adversarial review, automated triage, and ways to work around LLM context limits. The company frames the harness as the durable part of the system: models can be swapped, but the orchestration layer has to persist investigations across runs, deduplicate repeated leads, and filter many raw candidates into an actionable queue.
This Learning Hub checklist turns that architecture into release gates for application security, platform, and engineering teams. It does not assume a particular model provider, scanner, or cloud platform. It assumes the hard part is operational: converting uncertain machine output into a traceable vulnerability-management workflow without turning every speculative model note into a production incident.
Start with a finding lifecycle, not a model benchmark
Cloudflare’s earlier Project Glasswing writeup said the company pointed Mythos and other security-focused LLMs at live code across critical infrastructure, including more than fifty repositories. It highlighted two important capabilities: exploit-chain construction and proof generation. That is meaningful, but it is also why the surrounding harness matters. A stronger model can produce more plausible leads; it does not automatically decide which leads are real, current, exploitable, assigned, and fixed.
The first design gate is therefore a finding lifecycle. Every candidate should move through explicit states such as discovered, enriched, deduplicated, reproduced, human-reviewed, accepted, assigned, fixed, retested, and closed. If the harness only stores chat transcripts or one-off markdown reports, the team will lose the audit trail needed to explain why a finding was promoted or rejected.
Minimum state fields
- Asset scope: repository, service, package, branch, commit, ownership, and deployment environment.
- Hypothesis: the weakness class, affected path, input route, and model or tool that proposed it.
- Evidence: code references, reproduction status, logs, test artifacts, and links to reviewer notes.
- Disposition: duplicate, false positive, accepted risk, accepted vulnerability, or needs more research.
- Change linkage: ticket, pull request, release, retest result, and closure authority.
Make deduplication a first-class system
Cloudflare argues that generic coding-agent sessions are the wrong shape for fleet-wide security analysis because they hold one hypothesis at a time, fill context windows, and lose information during compaction. The same failure appears in vulnerability programs when five tools describe one bug five ways. Without deduplication, a harness creates noise faster than defenders can retire it.
Deduplication should not rely only on titles. Normalize repository paths, package names, functions, endpoints, commits, CWE mappings, sink/source pairs, and reproduction signatures. A model-generated SQL injection hypothesis, a static-analysis rule hit, and a human penetration-test note may describe the same underlying flaw. The harness should merge them when the evidence points to one fix, while still preserving which tool or model produced each observation.
Use a shared weakness language
The MITRE Common Weakness Enumeration describes CWE as a community-developed list of software and hardware weaknesses that can become vulnerabilities. It also says that knowing those weaknesses helps developers, hardware designers, and security architects eliminate them before deployment. A harness does not need perfect CWE mapping on day one, but it does need a stable taxonomy so teams can compare trends over time instead of debating a new label for every finding.
Do not let the taxonomy hide uncertainty
A model may assign an impressive-looking CWE ID before the evidence supports it. Store the mapping confidence and the reviewer who approved it. A tentative “possible authorization bypass” should not be displayed with the same certainty as a reproduced access-control failure tied to a specific test case.
Separate discovery from validation
Cloudflare’s harness framing favors model diversity: one model can discover an issue while a different model or tool validates it. That is an important control. If the same model invents the hypothesis, explains the exploitability, and writes the closing note, the workflow may only be measuring its confidence style. Use independence wherever possible.
A validation stack should include more than another prompt
- Static evidence: code path, data flow, dependency version, or configuration that makes the hypothesis plausible.
- Dynamic evidence: a safe reproduction, unit test, integration test, or controlled proof of concept in an authorized environment.
- Human review: an application owner or security engineer who can decide whether the behavior matters in the deployed system.
- Regression proof: a test or monitor that survives after the bug is fixed.
The point is not to remove humans. The point is to route human attention to candidates with enough supporting context that review becomes a decision, not a scavenger hunt.
Keep outputs portable and reviewable
A vulnerability harness will be more useful if its output can move between scanners, dashboards, code review, and ticketing systems. The OASIS SARIF 2.1.0 specification defines the Static Analysis Results Interchange Format as a standard format for the output of static analysis tools. SARIF is not a complete workflow system, but it is a helpful reminder that findings should not be trapped inside one model transcript or one proprietary queue.
Export records should be boring
For production use, store normalized finding records and preserve raw artifacts. A reviewer should be able to see the original model output, the code locations, the test artifacts, the deduplication decision, and the final disposition. If the harness cannot explain how it reached a conclusion, the team will struggle during incident review, customer assurance, or compliance evidence requests.
Anchor testing in a maintained security guide
The OWASP Web Security Testing Guide describes itself as a comprehensive guide to testing the security of web applications and web services and a framework of best practices used by penetration testers and organizations around the world. It also warns that scenario identifiers and links can change across versions, recommending versioned references when reports or tools depend on them.
That versioning lesson applies directly to AI-assisted triage. If a harness says “tested against WSTG guidance,” it should record which version or stable reference was used. If the test catalog changes, the historical finding should still be interpretable months later. The same principle applies to rule packs, model versions, prompts, dependency manifests, and repository commits.
Production release gates
Do not deploy a vulnerability harness until these gates pass
- Scope gate: repositories, services, branches, secrets policies, and authorized test environments are explicitly listed.
- State gate: every finding has a durable lifecycle record and cannot disappear when a chat session ends.
- Dedup gate: repeated leads are merged by evidence, not only by title or model wording.
- Validation gate: discovery and validation use separate reviewers, tools, models, or test stages where practical.
- Evidence gate: accepted findings include reproducible evidence or a documented reason why reproduction is not safe.
- Output gate: results can be exported or linked into existing AppSec, ticketing, and code-review workflows.
- Safety gate: exploit proofs run only in authorized environments with logging, rate limits, and data-handling controls.
- Metrics gate: the program tracks accepted findings, false positives, duplicates, time to triage, time to fix, and regression coverage.
Bottom line
An AI vulnerability harness should be judged less like a clever prompt and more like a production security system. The durable value is not that one model finds one bug in one run. The durable value is that the organization can preserve context, compare tools, validate claims, deduplicate evidence, and hand developers a small queue of findings worth fixing.
If the harness cannot explain its state, source, reviewer, and proof for each accepted vulnerability, it is not ready for broad use. If it can, AI-assisted discovery becomes a useful layer in the secure development lifecycle rather than another noisy scanner with a more fluent interface.
Sources
- Cloudflare Blog: Build your own vulnerability harness
- Cloudflare Blog: Project Glasswing: what Mythos showed us
- MITRE CWE: Common Weakness Enumeration
- OWASP Web Security Testing Guide
- OASIS SARIF 2.1.0 specification
- Featured image source: U.S. Navy server rack maintenance photograph on Wikimedia Commons
Featured image: server rack maintenance aboard USS Wasp by Commander, U.S. Naval Forces Europe-Africa/U.S. 6th Fleet, released as public domain U.S. Navy media; cropped and converted to WebP for sxz.io.








No Comment! Be the first one.