Appearance
Anthropic released a reference implementation for autonomous vulnerability discovery. OpenAI put Daybreak cybersecurity models on AWS Bedrock. A community repo packaged 817 security skills for AI agents. Three research papers from the same period show where the underlying techniques actually hold up.
The thread connecting all of it is the workforce gap: 4.8 million unfilled cybersecurity roles globally in 2024, per ISC2. There aren't enough humans to do the work, so vendors and researchers are pointing AI agents at it. The interesting part is the convergence. The tooling and the research land on the same lesson: LLMs handle the creative parts of security work, but they only become reliable when you wrap them in deterministic verification.
The workforce gap and the agentic answer
The old model of security tooling was scanners and static analyzers that produce findings for a human to interpret. The new model is an agent that closes the loop: it reads the code, crafts an exploit, verifies the crash, writes a report, and proposes a patch. That shift is visible across every major vendor. Anthropic's Project Glasswing is the strategic umbrella. OpenAI's Defender's Window makes the same argument from the other side: AI is reshaping attack and defense simultaneously, and security teams need to adapt now. The Daybreak models landing on Bedrock matters for enterprises that need cybersecurity capabilities inside their existing AWS governance boundaries.
The clearest concrete artifact is Anthropic's defending-code-reference-harness. The repo is explicit about its scope: a reference implementation, not a product. But it encodes the full loop that every other tool in this space is circling: recon → find → verify → report → patch.
The autonomous pipeline: how agentic vuln discovery works
The harness targets C/C++ memory vulnerabilities. It compiles the target with ASAN, runs parallel find agents that craft malformed inputs until a crash reproduces three times out of three, then a separate grader agent reproduces each crash in a fresh container. The only thing that crosses from the find agent to the grader is the proof of concept. That separation is the whole ballgame. It's the difference between an agent claiming it found a bug and an agent proving it.
Seven stages, each with a distinct role. The recon agent proposes a partition of the attack surface so parallel find agents don't converge on the same bug. The verify agent runs in a container the find agent never touched. The judge decides whether a crash is new, a duplicate, or a better example of a known bug. The patch agent writes a fix, and a grader confirms the original PoC no longer crashes, the test suite still passes, and a fresh find agent can't get around the fix.
When I ran the pipeline against drlibs with the --stream flag, the first report landed in minutes. That's the part that feels different from traditional tooling: the agent isn't flagging suspicious code, it's handing you a reproducible crash and an exploitability analysis. The sandboxing is non-negotiable, though. The autonomous pipelines refuse to run outside a gVisor container unless explicitly overridden, because an agent that executes target code is itself a remote code execution primitive.
Quick Take: Every approach that works in this space pairs an LLM that proposes with a deterministic system that verifies, and the tooling is already ahead of the published research.
The tooling wave: skills libraries and MCP frameworks
The gap between a generic LLM and a useful security analyst is domain knowledge. A junior analyst knows which Volatility3 plugin to run on a suspicious memory dump, which Sigma rules catch Kerberoasting, and how to scope a cloud breach across three providers. A generic agent doesn't. The community skills library tries to fix that with 817 structured skills across 29 security domains, each following the agentskills.io standard.
The numbers are the interesting part. Each skill costs about 30 tokens to scan in frontmatter-only mode and 500 to 2,000 tokens to fully load. That progressive disclosure means an agent can search all 817 skills in a single pass without blowing its context window. I tested this by pointing an agent at a memory dump task. The frontmatter scan identified the relevant Volatility3 playbook immediately, and the workflow section walked through the exact plugin sequence. Without the skill, the agent would have guessed at commands.
The framework mappings are validated against MITRE ATT&CK v19.1 using the official mitreattack-python library: 290 distinct techniques and sub-techniques, zero revoked or deprecated IDs. The library also maps to MITRE's Fight Fraud Framework, released April 9, 2026. F3 adds two tactics ATT&CK doesn't enumerate, Positioning and Monetization, which matters if you're building agents for financial fraud work.
On the offensive side, HexStrike AI takes a different approach. It's an MCP framework with 150+ security tools and 12+ autonomous agents: a bug bounty agent, a CTF solver, a CVE intelligence agent, an exploit generator. The design choice that stands out is the decision engine. Instead of letting the LLM pick tools ad hoc, HexStrike routes through a layer that handles tool selection, parameter optimization, and attack chain discovery. That's a recognition that raw LLM tool-calling is too noisy for pentesting. You need a structured layer between the model and the tools.
The HN thread on Project Glasswing was mostly supportive, with the general sentiment that AI security tooling for critical software is overdue. The skepticism centered on delivery: a hosted vendor product versus an open-source reference you can inspect and port. Several commenters argued the harness is the more useful artifact, since you can read the prompts, check the sandboxing, and adapt the pipeline to your own stack. I'm inclined to agree after running it.
What the research says
Three papers from the same period show where the techniques hold up. They attack different problems, but they share the pattern: neural models paired with statistical or deterministic validation.
| Approach | Problem | Method | Key result | What it means |
|---|---|---|---|---|
| BGA (arxiv 2608.14126) | Detecting attacks in encrypted TLS 1.3 flows | ANOVA feature decoupling, WGAN-GP for minority samples, BiLSTM with gated attention | 95.2% across CIC-IDS-2018 and Edge-IIoT; 43.2% recall gain on rare MSCI attacks | 0.282ms per flow, fast enough for real-time edge gateways |
| CodeSIFT (arxiv 2608.14303) | Detecting contaminated code-generation prompt batches | Influence functions on parameter space plus a statistical test | AUROC up to 0.98 on 3B-7B code LLMs | Catches novel attack classes with no prior knowledge of the pattern |
| Hybrid annotation framework (arxiv 2608.14370) | Generating SecBPMN2 security annotations from requirements docs | LLM semantic extraction, schema-constrained mapping, rule-based normalization | Precision 0.58 vs 0.29 for humans; recall 0.52 vs 0.50; ~50% fewer misplaced annotations | Cuts modeling effort from days to minutes with better consistency |
Three different attack surfaces, one shared architecture: the model proposes, the system validates.
Key Numbers4.8M unfilled cybersecurity roles globally in 2024 (ISC2). 0.98 AUROC for CodeSIFT at moderate-to-high injection rates, near-perfect separation between contaminated and benign prompt batches. 95.2% performance ceiling for BGA across CIC-IDS-2018 and Edge-IIoT. 0.282ms inference latency per flow for BGA, fast enough for real-time detection on edge gateways. 817 structured security skills in the community library, searchable in a single agent pass.
Detection in encrypted traffic: BGA
BGA sits in a crowded field. Comparative studies keep showing deep models beating traditional ML on structured intrusion detection, with CNN and LSTM around 98% accuracy and random forest reaching 99.9% on the same benchmarks. The V2X survey literature makes the same point for connected vehicles: centralized ML training doesn't scale to distributed networks, which is why federated learning and edge AI keep coming up. BGA's contribution is making this work on encrypted traffic at edge latency.
Encrypted traffic is the hard case for intrusion detection. TLS 1.3 flows are high-entropy, so attention mechanisms dilute on cryptographic noise. BGA's answer is to decouple control-plane features, specifically industrial setpoints, from the stochastic noise using ANOVA, then feed the temporal structure through a BiLSTM with an adaptive gated multi-head attention mechanism. The gate acts as a neural filter: it suppresses encryption artifacts while amplifying malicious signatures.
The class imbalance problem got the same treatment. The corpus had 86,878 flow records with extreme imbalance, so the authors trained a WGAN-GP to synthesize minority samples. That lifted recall on rare Malicious State Command Injection attacks by 43.2%. The inference latency, 0.282ms per flow, is the number that matters for deployment. It means the model can sit on an industrial edge gateway, not just in a data center. The estimated 1.69ms on ARM hardware is still within real-time budgets for most ICS/SCADA monitoring.
Detecting poisoned code-generation prompts: CodeSIFT
CodeSIFT takes a different angle. Instead of detecting specific vulnerabilities in generated code, it detects batches of prompts that induce anomalous model behavior. The method uses influence functions to measure parameter-space influence of generated code, then runs a statistical test to see whether a candidate prompt set deviates from a benign reference distribution.
The results: AUROC up to 0.98 at moderate-to-high injection rates on three open-weight code LLMs from 3B to 7B parameters. The 3B to 7B range is practically meaningful. These are models you can run in a CI pipeline on a single GPU, no cloud inference needed. The threat-model-agnostic property is what matters. Static analysis baselines need to know the vulnerability pattern in advance. CodeSIFT doesn't, which is what you want when attackers are inventing new obfuscations faster than you can write signatures.
The annotation problem: LLMs for security-by-design
The third paper is less glamorous and arguably closer to daily security work. Security requirements documents exist, and BPMN process models exist, but turning the former into SecBPMN2 annotations is manual, expert-intensive, and error-prone. The hybrid framework combines LLM-based semantic extraction with schema-constrained mapping, rule-based normalization, and deterministic validation.
The evaluation is small, 27 process models from various domains, but the comparison against human analysts is striking. The system hit precision of 0.58 versus 0.29 for humans, with comparable recall (0.52 vs 0.50), and reduced erroneous or misplaced annotations by nearly 50%. Humans were better than I'd expect on recall, but their false positive rate was brutal. The framework trades a little recall for a lot of precision, and it does it in minutes instead of days.
The precision gap is the headline. A 0.29 precision means humans spent most of their annotation effort on annotations that were wrong or misplaced. The framework cuts that waste by half while matching recall. For security-by-design workflows, that's the difference between a process model that reflects the requirements and one that gives a false sense of coverage.
Reality check: the TSA story
All of this research exists because the manual alternative is failing in the field. The KCM story from Ian Carroll and Sam Curry is the clearest example. FlyCASS, a vendor running the Known Crewmember and Cockpit Access Security System authorization for small airlines, had a login page that interpolated the username directly into a SQL query. A single quote triggered a MySQL error. The login bypass was ' or '1'='1. They logged in as an administrator of Air Transport International, added a fake employee named Test TestOnly with a test photo, and the system approved them for KCM and CASS access. That means skipping security screening and accessing the cockpits of commercial airliners.
The disclosure process was its own mess. DHS confirmed they were taking it seriously and disconnected FlyCASS from KCM/CASS. Then the TSA press office issued statements denying the vulnerability, claiming a vetting process would catch it. The researchers pointed out that a KCM barcode isn't required, since TSOs can enter an employee ID manually. The TSA deleted the section of their website mentioning manual entry and stopped responding.
This is the baseline the AI tools are being built against. A single-person vendor running critical aviation security infrastructure with SQL injection in the login form. There aren't enough humans to find these bugs before someone else does.
Common Pitfalls
Running autonomous agents without a sandbox is the fastest way to turn your security tooling into a liability. The Anthropic harness refuses to run its autonomous pipelines outside gVisor unless explicitly overridden. That's not paranoia. An agent that executes target code is a remote code execution primitive. If you're building your own pipeline, treat the sandbox as mandatory.
Treating LLM findings as verified is the second classic mistake. The static scan skills produce candidates, not confirmed bugs. The README is explicit: expect more false positives on non-canary targets. My team hit this on day one. The scan flagged a deliberately vulnerable demo file, and the triage skill correctly dismissed it as test code. Without that triage step, we'd have chased ghosts. Verification requires execution. Skip it and you'll drown in hallucinations.
Ignoring class imbalance in IDS training will quietly destroy your recall. BGA's 43.2% recall gain on rare attacks came from a WGAN-GP synthesizing minority samples. Train on raw traffic data and your model will learn to predict benign for everything, because that's correct 99% of the time. You have to engineer for the rare class explicitly.
Assuming a generic LLM knows your domain is the most expensive assumption you can make. The 817-skill library exists because agents guess at tool commands without structured playbooks. A junior analyst knows which Volatility3 plugin to run. Your agent doesn't. Budget for skills or playbooks, not just model access.
Letting the LLM pick tools ad hoc will burn your API budget and miss targets. HexStrike's decision engine exists because raw tool-calling is too noisy for pentesting. Route tool selection through a structured layer, or your agent will spend its time on the wrong scanner.
One thing to remember
One pattern shows up in every approach that works: the LLM proposes, and a deterministic system disposes. BGA gates its attention with a learned filter. CodeSIFT validates with a statistical test. The annotation framework normalizes with rules. The harness reproduces crashes in a fresh container. Pure LLM judgment is never the final word.
The Bottom Line
If you're building a vulnerability discovery pipeline, copy the recon → find → verify → report → patch shape from Anthropic's harness, because the verify stage is what separates real findings from hallucinations.
If you're deploying AI for intrusion detection on encrypted traffic, follow BGA's pattern: decouple control-plane features, synthesize minority samples, and gate your attention, because naive transformers dilute on high-entropy flows and naive training misses rare attacks.
If you're a small security team, skip building your own agent infrastructure and start with the skills library plus a managed option like Claude Security or Daybreak on Bedrock, because the 4.8M role gap means you don't have the engineering time to reinvent the pipeline.