Anthropic Workshop: Build Agents That Run for Hours — Ash Prabaker & Andrew Wilson
Why self-evaluation is a trap and adversarial evaluator agents work better; why context compaction doesn't cure coherence drift but structured handoffs do; how to decompose work into testable sprint contracts; how to grade subjective output with rubrics an LLM can actually apply; and how to read traces as your primary debugging loop. Plus the question nobody asks: which parts of your harness should you delete when the next model drops? Speaker info: Ash Prabaker | / ash-prabaker Andrew Wilson | / anddwilson Timestamps: 0:00 Introduction and speakers 1:21 Overview of long-running agents 2:29 Challenges: Context, Planning, and Judgment 4:14 Two approaches: Model updates vs. Harness evolution 5:58 Prehistory: Sonnet 3.5, Computer Use, and MCP 6:34 The evolution of Claude Code 7:55 The Ralph loop technique 9:49 Sonnet 4.5, Agent SDK, and checkpoints 10:49 Opus 4.5 and the role of sub-agents 12:05 First long-running agent patterns 14:20 Opus 4.6, Agent Teams, and server-side compaction 17:28 State-of-the-art harness patterns 21:30 Evaluating subjective output with rubrics 23:44 Introducing the 'Planner' role 25:04 The generator-evaluator contract 31:28 Specificity in contracts and debugging traces 34:14 Adjusting harnesses as models evolve 37:56 How to build your own agent harness 39:01 Key takeaways for long-running agents 40:05 Q&A session
![Claude Agent SDK [Full Workshop] — Thariq Shihipar, Anthropic](https://i.ytimg.com/vi/TqC1qOfiVcQ/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLDB4dgixOEo-ks06DgonjIJ4Suh9g)
Claude Agent SDK [Full Workshop] — Thariq Shihipar, Anthropic

CLAUDE CODE ADVANCED FULL COURSE (3 HOURS)

The Layer That Makes LLMs Trainable

Why Netflix is betting on systems thinkers—not specialists—in the AI era | Elizabeth Stone (CPTO)

Ralph Loops: Build Dumb AI Loops That Ship — Chris Parsons, Cherrypick

Claude Agents Tutorial: Free 2-Hour Masterclass by Anthropic

Claude Code Head Boris Cherny: Insane Growth, Tokenmaxxing, AI Agents' Next Frontier

The Agent Development Lifecycle: Build, Test, Deploy, Monitor | Interrupt 26

RL for Agents Workshop - Deep Dive on Training Agents with RL and Open Source

Harnesses in AI: A Deep Dive — Tejas Kumar, IBM

Claude Code best practices | Code w/ Claude

Andrej Karpathy: From Vibe Coding to Agentic Engineering w/ Stephanie Zhan

Full Walkthrough: Writing & Using Skills — Nick Nisi and Zack Proser

Build a Complete Medical Chatbot with LLMs, LangChain, Pinecone, Flask & AWS 🔥

Keynote: The dangers of probably-working software - Damian Brady - NDC London 2026

The Multi-Agent Architecture That Actually Ships — Luke Alvoeiro, Factory

Andrej Karpathy: Software Is Changing (Again)

Context Engineering Our Way to Long-Horizon Agents: LangChain’s Harrison Chase

How Anthropic’s product team moves faster than anyone else | Cat Wu (Head of Product, Claude Code)

