Hardening the Code: How Trail of Bits Built an AI System That Finds and Patches Real Vulnerabilities
DARPA's AI Cyber Challenge asked teams to build autonomous systems that could find vulnerabilities in real open-source software, prove they exist, and generate patches that actually fix the problem without breaking anything else. Trail of Bits placed second with Buttercup, a hybrid cyber reasoning system that combines fuzzing with targeted LLM applications. In this session, Dan Fernandez (Principal Product Manager, Edera) sits down with Michael Brown (Head of AI and ML Security Research, Trail of Bits) for a technical Q&A on what it took to build a system that operates at that level of rigor. The conversation covers architecture decisions, why reasoning models didn't help, how multi-agent patching actually works, and what trust boundaries look like when your AI agents are analyzing untrusted code. Topics covered: Why hybrid approaches (fuzzing + LLMs) outperformed pure-AI solutions in the competition How Buttercup's multi-agent patching system uses 4-5 specialized agents with validation between each step Why Trail of Bits chose non-reasoning models and still achieved 90% patch accuracy Threat modeling for AI systems that ingest source code: prompt injection, trust boundaries, and input sanitation Deploying AI-powered vulnerability detection at scale — from $100K competition budgets to running on a laptop The dual-use debate: why patching is 1-2 orders of magnitude less complex than exploitation, and what that means for defenders Prescriptive vs. descriptive problem-solving: a framework for deciding where AI actually helps in security Buttercup is open source and available on Trail of Bits' GitHub. You can run it on your laptop without DARPA's competition infrastructure. 0:00 Introduction 2:46 Michael Brown's Background: Military Aviation to AI Security 7:08 The AI Cyber Challenge: Rules, Constraints, and Competing Philosophies 15:08 Technical Deep Dive: Vulnerability Deduplication and Multi-Agent Patching 24:23 Model Selection: Why Reasoning Models Didn't Help 28:42 Threat Modeling for AI Security Tools 37:09 Deploying Buttercup at Scale and Open-Sourcing It 47:00 The Dual-Use Debate: Why These Tools Favor Defenders 53:09 Rapid-Fire Q&A: Static Analysis, Prescriptive vs. Descriptive Problems

Why Kubernetes Multi-Tenancy Keeps Failing: An Offensive Security Perspective

How AI agents & Claude skills work (Clearly Explained)

Morning exercise for software developer!#SoftwareDeveloper #softwaredeveloperlife #coding

How the Moon Becomes the Solar System’s Shipping Yard

Why Container Security Is Stuck on Detection — and How to Fix It

Portrait of a Graduate & Personalized Learning: Aman Sahota (Factors Education)

Jensen Huang: The Mindset That Built NVIDIA

How CIA’s Hacking Tools Were Leaked

System Design Explained: APIs, Databases, Caching, CDNs, Load Balancing & Production Infra

An AI Model Escaped Its Sandbox & Compromised Hugging Face — What That Means for Container Security

What is SonarQube | Introduction SonarQube | SonarQube Tutorial | SonarQube Basics | Intellipaat

Andrej Karpathy: From Vibe Coding to Agentic Engineering w/ Stephanie Zhan

Keynote: After the AI Hype – What’s Real, and What’s Next - Richard Campbell - 2026

The Invisible Phone: Ditch PSTN Forever – No IMEI, No IMSI, No KYC, Total Stealth Calling

All mammals get 1 billion heartbeats. Except humans...

OpenAI: A Bubble Bigger Than Dotcom

How AI is Changing UX & Product Design | The Future of Intelligent Systems

26 Years of Survival Keyword AX (Great AI Transformation): We’ll Tell You Everything in Just One ...

Jfrog | Jfrog Artifactory | Jfrog Artifactory Tutorial | Artifactory Tutorial | Intellipaat

