Putting an AI agent on your security backlog
How to judge an AI security agent by the findings it can prove, pilot one against bugs you already fixed, and avoid buying a second backlog.
on this page · 0 / 0 checked
Somewhere you have a list. Dependency alerts on the app you sell, an old pentest report from a client engagement, the scanner output your hosting panel emails you every Monday, the three things you know are wrong with the admin login and have not fixed since March. Nobody has cleared that list because clearing it was never the fun part, and because most of it is probably noise, and you cannot tell which part without spending an afternoon on each item.
Now there are agents that offer to do it. They crawl the code, propose vulnerabilities, try to prove them, and hand you a fix to review. Some of them are good. All of them can also hand you a longer list than the one you started with, which is the specific way this purchase goes wrong. This guide is about how to tell the difference before you pay, and it is written for the person who owns the repository and also sends the invoices. If you run a security program with a triage rota and a compliance auditor, you need a longer document than this one.
Detection was never your bottleneck
Findings have been cheap for a decade. Static analysers, dependency scanners, and any competent pentest will produce more of them than a small team can process, and the backlog exists because triage is expensive, not because detection is hard. Deciding whether a finding is real, reachable in your actual deployment, and worth an afternoon is the work.
The base rates are worse than most people assume. The CISA Known Exploited Vulnerabilities catalog lists roughly 0.5% of published CVEs, and in any 30-day window the Exploit Prediction Scoring System observes exploitation activity against roughly 2.5% to 3% of published CVEs [6]. FIRST puts the shape of it plainly: “A rather small fraction of vulnerabilities carries almost all exploitation activity” [6]. A queue sorted by severity is mostly a queue of things nobody will ever attack, which is why working it top to bottom feels unproductive. It is unproductive.
What that costs when the noise scales is now well documented. The curl project ran a bug bounty from April 2019, paid out over 100,000 USD and confirmed 87 vulnerabilities, and historically saw “a rate of somewhere north of 15% of the submissions ending up confirmed vulnerabilities” [7]. Then AI-assisted reports arrived: “Starting 2025, the confirmed-rate plummeted to below 5%. Not even one in twenty was real” [7]. The project ended the bounty on 31 January 2026, and Daniel Stenberg’s reason was not money. The submissions “take a serious mental toll to manage and sometimes also a long time to debunk” [7]. That is the failure mode you are buying against.
What you are actually buying is a validated finding
The thing that separates this generation of tools from the scanner you already ignore is not the model. It is the step after the model: reproducing the suspected bug in an isolated environment to see whether it actually fires. OpenAI’s agent, launched as Aardvark and now shipping as Codex Security in research preview, describes it directly: when it identifies a candidate vulnerability it “will attempt to trigger it in an isolated, sandboxed environment to confirm its exploitability,” and it “attaches a Codex-generated and Aardvark-scanned patch to each finding for human review” [2]. Cognition’s Devin Security Swarm does the same shape of thing, batching candidates out to “child agents that investigate in parallel, each reasoning over one bounded shard from a focused context,” after which “each serious finding is then reproduced in a sandbox against a running build, so the report reflects runtime-verified findings” [1].
Read that as an economic claim rather than a technical one. A finding that arrives with a reproduction is already triaged; someone ran it and it worked. A finding that arrives as a paragraph of reasoning is a task, and tasks are what you already have too many of. When you compare two tools, the question is not how many issues each one surfaced. It is what fraction of what reached you turned out to be real, and how long the rest took to dismiss.
Vendor recall numbers leave out the part that costs you
Cognition published a comparison on 1 July 2026 across 50 vulnerabilities in 14 languages, drawn from CVEs published after the models’ training cutoffs. Devin Security Swarm scored 72% recall at 90.23 USD per scan, Claude Security 68% at 131.87 USD, Codex Security 48%, and Cursor’s scanner 26% at 4.60 USD per scan [1]. Those numbers are useful and the vendor deserves credit for publishing the method, but read the definition attached to them: “the fraction of the 50 cases in which at least one of the run’s findings describes the target vulnerability; everything else, including false positives, is ignored” [1]. The same post concedes that labelling each other finding true or false at scale “isn’t feasible” [1].
So the published table measures the half of the problem that was never your bottleneck. Cost per scan is the wrong unit. Cost per validated finding, with your own hourly rate attached to the dismissals, is the right one, and on that unit a 4.60 USD scan that generates a week of reading is not cheap.
Precision is sometimes bought by leaving things out, which is also worth knowing before you compare outputs. Anthropic’s open-source security review Action states that it “automatically excludes a variety of low-impact and false positive prone findings,” specifically denial of service, rate limiting, memory and CPU exhaustion, generic input validation without proven impact, and open redirects, and says the filtering “can also be tuned as needed for a given project’s security goals” [4]. That is a defensible design choice for a tool that runs on every pull request. It also means a clean report is a clean report about a defined subset, not a statement about your application.
Scan fee plus triage time at your own rate. Computed in the page; nothing is sent anywhere.
Turn on the free layer before you buy the expensive one
Most solo operators do not have an application-code backlog large enough to justify a per-scan agent yet. They have dependencies, a login flow, and one server. The cheap layer covers a surprising amount of that, and it is already sitting in the platform you host on. GitHub’s code scanning is available on public repositories on GitHub.com with no licence requirement; private repositories need a GitHub Code Security licence [5]. Turn that on, along with your host’s dependency alerting, and let it run for a month before you form an opinion about what an agent would add.
The same applies to the tools you are already paying for. Claude Code includes a /security-review command that “lets you run security analysis directly from your terminal before committing code,” and an Action that triggers automatically when new pull requests are opened and posts inline comments with the concerns it found and its recommended fixes; both are available to users on individual paid plans, meaning Pro or Max, and to individual users or enterprises with pay-as-you-go API Console accounts [3]. If you have one of those, that pass costs you nothing extra. Run it on your own code this week, and you will learn more about the category from 20 findings on software you understand than from any comparison table.
Pilot it against bugs you already fixed
You have something the benchmarks do not: a history. The CVE you patched last year, the pentest finding from the client engagement, the injection bug that caused the incident, the auth check that was missing in the version you shipped in 2024. Check out the commit before each fix and point the tool at it. That is a known-answer test on your own architecture, your own framework choices, and your own bad habits.
Score three things, and write them down before you start so you cannot move the goalposts afterwards. First, recall against your known set, meaning the share of the bugs you buried in that snapshot that came back. Second, precision on everything else, meaning the share of the findings you did not plant that survived a careful read. Third, and the one people skip, the cost of the noise, meaning the total minutes the dismissals took, including the two that looked real for half an hour.
The reason to use your own history rather than a public benchmark is contamination. Vendor evaluations are built from published advisories, which is why Cognition specifically drew its cases from CVEs published after the models’ training cutoffs [1], and even then the surrounding ecosystem of writeups and patches is public. Your unpublished pentest report is not in anyone’s training data. It is the only test that predicts what next month looks like.
The fix deserves more suspicion than the finding
Finding a bug and repairing it without breaking anything are different skills, and the second is where these products vary most. A patch that quietly changes behaviour under load, or that blocks the specific payload while leaving the class of bug alive, is worse than an open ticket, because it closes the item while the risk stays open. The vendors are consistent on this point. Aardvark attaches its patch “for human review” [2], Anthropic’s Action posts recommended fixes as inline comments rather than committing them [3], and Anthropic’s own guidance says automated security reviews “should complement, not replace, your existing security practices and manual code reviews” [3].
Treat an agent’s remediation the way you would treat a first contribution from a talented stranger. Run the full test suite. Read the diff against the attack path rather than against the description of the attack path. Ask whether the fix would still hold if the input arrived through a different route, because that is the difference between fixing a vulnerability and deleting a proof of concept. And keep the merge button attached to a human who knows the subsystem, which in a small shop means you, which means the review time belongs in the cost calculation above.
Pointing an agent at a repository is itself an action
There is one risk in this category that has nothing to do with finding quality. Running an agent over code means running an agent over code you may not control, and repositories can carry instructions. On 1 September 2026, Manifold Security published GitSpawn, an attack that abuses core.fsmonitor, a git performance setting whose value is a command that git runs. As the writeup puts it, “Git reads that setting from the repository’s own .git/config. So a repository can ship this: [core] fsmonitor = git status before you type anything. Before the workspace-trust prompt” [8]. The disclosure covers Claude Code, OpenAI Codex, Cursor, goose, Qwen Code, Grok Build and Hermes, with 4 of its 8 findings patched and 4 still unpatched at publication [8].
The mitigations are cheap and specific. Manifold’s advice to users is one line: “Inspect .git/config before you open the directory with an agent. Any setting that names a program can run it” [8]. The delivery vector narrows it further, because “cloning a hostile URL does nothing, and neither does fetch or pull. The repository has to arrive as files with its .git directory already inside” [8], which makes the vector a directory somebody moved to you rather than one you cloned yourself [8]. So clone the thing, do not open the folder that landed in your shared drive, and give the scan the smallest blast radius that still works: a machine that does not hold your production credentials, and read-only access wherever reading is enough. The specific bug will be patched. The shape of it will recur, because an agent that reads arbitrary code is a program that consumes untrusted input.
What still goes wrong
The honest limit is that validation raises the floor without settling the question. A sandbox proves a bug fires in a sandbox, which is not the same as proving it is reachable in your deployment, behind your CDN, with your authentication in front of it. A run is also a sample rather than a census. Cognition grades one planted bug per case and notes that “the needle we grade is one of several in the haystack, and surfacing a different real needle still scores zero,” which is why it reads its own numbers as “a floor on what a run finds, not a ceiling” [1]. A floor cuts both ways. A clean report is evidence, not a guarantee.
The economics can also fail quietly. A tool at 90 USD a scan is not expensive next to an engineer’s afternoon, but if it produces 12 findings a month and 2 are real, you are paying for the scan and for 10 dismissals, every month, forever. That arithmetic is worth redoing at 90 days rather than assumed at purchase, because the first month is always the best month; the tool is finding the accumulated backlog, and steady state is a smaller number of new findings against the same fixed cost.
And nothing here helps if what you actually need is assurance. A scan output is not a penetration test, a SOC 2 control, or an answer to a client’s security questionnaire, and a report generated by an agent will not satisfy an auditor who wants to know who looked and what they were qualified to see. Automated review is a way to spend your attention better, not a way to spend less of it. The backlog was always a judgment problem, and the judgment is still yours.
- 01Cognition — Evaluating Security Swarmdevin.ai
- 02OpenAI — Introducing Aardvarkopenai.com
- 03Anthropic Help Center — Automated security reviews in Claude Codesupport.claude.com
- 04Anthropic — claude-code-security-review (GitHub Action)github.com
- 05GitHub Docs — About code scanningdocs.github.com
- 06FIRST — EPSS frequently asked questionsfirst.org
- 07Daniel Stenberg — The end of the curl bug-bountydaniel.haxx.se
- 08Manifold Security — GitSpawn: hijacking AI coding agents through git configmanifold.security