One Bias After Another - Daniel Fein & Max Lamparth

​To make AI systems helpful, we use reward models, which are essentially an automated grading system that tells the AI what a good answer looks like. But what happens when the grader itself is fundamentally biased? ​On May 6th 2026, ‪@SafeAIGermany‬ hosted Daniel Fein & Max Lamparth from ‪@stanford‬ to present their latest findings on the hidden flaws inside frontier AI models (Paper: https://arxiv.org/pdf/2603.03291) In this presentation and Q&A session, we explored how state-of-the-art AI assistants remain plagued by simple biases and their work towards solutions. About the speakers: Daniel is a graduate researcher in the Stanford Intelligence Systems Laboratory. His research focuses on controlling and understanding the behavior of language models. He has published work on preference learning, with broader interests in evaluation, alignment, and enabling models to reason more reliably in complex settings. Max is a Research Fellow at the ‪@HooverInstitution‬, the Stanford Intelligence Systems Laboratory, and the Stanford Center for AI Safety. His research focuses on the security and safety of language models through mechanistic interpretability, reward modeling, and robust evaluation. Before, Max was a postdoctoral fellow at Stanford and received his Ph.D. from the School of Natural Sciences at the Technical University of Munich. About SAIGE: Safe AI Germany (SAIGE) is building Germany's infrastructure for safer AI. Our online events are open to everyone everywhere. 🌐 safeaigermany.org 📅 Upcoming events: luma.com/saige 🐦 Follow us on LinkedIn: linkedin.com/company/safe-ai-germany Recorded 6 May 2026.

Learning theory for AI safety - Guillaume Corlouer
▶︎

Learning theory for AI safety - Guillaume Corlouer

Keynote: After the AI Hype – What’s Real, and What’s Next - Richard Campbell - 2026
▶︎

Keynote: After the AI Hype – What’s Real, and What’s Next - Richard Campbell - 2026

VISReg: The AI Technique That Beats DINO, VICReg & SIGReg.  Evolution of Self-Supervised Learning.
▶︎

VISReg: The AI Technique That Beats DINO, VICReg & SIGReg. Evolution of Self-Supervised Learning.

Communicating AI Safety: From concern to action - Justus Baumann
▶︎

Communicating AI Safety: From concern to action - Justus Baumann

Talent Needs in AI Safety - Tilman Räuker
▶︎

Talent Needs in AI Safety - Tilman Räuker

July 14th, 2026 - Spotlight on UF's Young Scholars: Nature's Engineers
▶︎

July 14th, 2026 - Spotlight on UF's Young Scholars: Nature's Engineers

There’s a Problem with Quantum Mechanics – Quantum Reality (1/3) with Jim Al-Khalili
▶︎

There’s a Problem with Quantum Mechanics – Quantum Reality (1/3) with Jim Al-Khalili

Physics-Informed Research for Ambitious AI Safety - Ari Brill
▶︎

Physics-Informed Research for Ambitious AI Safety - Ari Brill

AI Is About to Crash. Here’s Why.
▶︎

AI Is About to Crash. Here’s Why.

How To Think SO Clearly People Assume You're Brilliant
▶︎

How To Think SO Clearly People Assume You're Brilliant

Exploration Hacking: Can LLMs learn to resist RL training? - Joschka Braun
▶︎

Exploration Hacking: Can LLMs learn to resist RL training? - Joschka Braun

System Design Explained: APIs, Databases, Caching, CDNs, Load Balancing & Production Infra
▶︎

System Design Explained: APIs, Databases, Caching, CDNs, Load Balancing & Production Infra

"A.I. and Our Economic Future," Professor Chad Jones
▶︎

"A.I. and Our Economic Future," Professor Chad Jones

Last Lecture Series: “How to Win Without Crushing Your Soul” - Graham Weaver
▶︎

Last Lecture Series: “How to Win Without Crushing Your Soul” - Graham Weaver

Nobody Explained the Schrödinger Equation Like THIS!
▶︎

Nobody Explained the Schrödinger Equation Like THIS!

How To Become Dangerously Self-Educated (with AI)
▶︎

How To Become Dangerously Self-Educated (with AI)

Verification for low-trust AI governance - Naci Cankaya
▶︎

Verification for low-trust AI governance - Naci Cankaya

A Closer Look at RLVR Through the Lens of Reasoning Strategies | Eshwar Sivaramakrishnan
▶︎

A Closer Look at RLVR Through the Lens of Reasoning Strategies | Eshwar Sivaramakrishnan

Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Economics of Generative AI
▶︎

Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Economics of Generative AI

Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 1 - Transformer
▶︎

Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 1 - Transformer