Coding#Security scan

security-audit: make the agent prove a vulnerability before reporting it

Cloudflare's security-audit skill stops agents dressing up guesses as findings: six phases force source evidence and independent verification first.

Skill details

Install
npx skills add cloudflare/security-audit-skill --skill security-audit

Ask an agent to review code for security issues and the report often reads scarier than the code: missing best practices become vulnerabilities, guessed deployments become findings, and the evidence never holds up. Cloudflare open-sourced its production vulnerability-hunting discipline as an agent skill — cloudflare/security-audit-skill (MIT, 23,432 stars as of 2026-09-30) — with one rule: evidence first, severity later.

The order it works in

security-audit lays a full audit out in six phases. Reconnaissance maps trust boundaries and attack surfaces into architecture notes and a coverage ledger; hunting waves assign isolated hunter agents from that ledger, then a coverage critic looks for corners nobody touched; every candidate goes to a fresh verifier whose only job is to falsify it; verdicts land in findings.json as confirmed, needs_validation or rejected, checked by zero-dependency validators shipped with the repo; confirmed records face brand-new agents again, and modified ones get re-verified once more; the run ends with REPORT.md, FINDINGS-DETAIL.md and NEEDS-VALIDATION.md, none carrying live-probe instructions. Re-running the same repo is additive: last run’s conclusions survive only if current source and fresh validation still support them.

What it enforces

Half of security-audit’s value is process; the other half is red lines.

  • Unproven means no finding: a candidate must name the lower-trust principal, the crossed boundary, the affected principal or resource and the concrete result; missing best practices, guessed deployments and parser crashes do not count.
  • needs_validation gets no severity: only confirmed records are rated; when the decisive fact sits outside the repository, the skill records what is missing and how to check it safely, instead of quietly scoring it low.
  • Loading is not authorization: by default the skill is guidance only — everyday questions use the relevant sections, the six-phase workflow starts only on an explicit audit or report request, and ambiguous ones get one clarifying question first.
  • Target code runs only in a sandbox: source inspection is read-only; target-controlled builds or tests run only in an OS-enforced sandbox with no network, an empty allowlisted environment, a read-only target and resource limits, dummy principals throughout, and deployed endpoints untouched.
  • Never claim full coverage: the coverage ledger and findings.json must both pass validators; an unfinished run is marked incomplete with its exact reason, and one pass never claims to have exhausted the target.

Who it’s for

security-audit fits teams who want AI to carry a real security review of a serious codebase, and red teams rechecking old findings. It pairs well with reverse-skill, which routes security tasks to a methodology first. The honest catch: it is guidance documents plus two validation scripts, no scanner magic; a full run spends many agents and a lot of tokens — the SKILL.md even ships budget rules for cutting scope; it audits source only and never probes live services — for checking a running site, see Web-Check — so it does not replace a human pen test; and it needs a platform with subagent delegation to run as designed.