Your Agent Produced the Right Answer. It Still Did the Wrong Thing.

Introducing PolicyEval, the open-source library that scores AI agents against their policy, not just their output. The first two posts in the series (#1, #2) landed on something uncomfortable. AI coding assistants ship vulnerability patches that compile, pass tests, and still break production. Roughly one in five. The fix wasn’t a smarter model – it […]

Shrink the Exposure Pile: Vulnerability Management in the Age of Vibe Coding

Part 1 of a two-part series. Security teams already own more scanners, exposure platforms, and dashboards than they can act on, yet the vulnerability backlog grows every week. The bottleneck has moved downstream, from finding vulnerabilities to remediating them, and AI-generated code widens that gap faster than any hiring plan can close it. On a […]