Learning theory for AI safety - Guillaume Corlouer

Link to slides: https://docs.google.com/presentation/... Summary: Following Ari Brill's talk on physics-informed AI safety (   • Physics-Informed Research for Ambitious AI...  ), this session broadens the lens to the wider landscape of theoretical directions in AI safety, including where to start if you'd like to work on them. ​On Wednesday, July 8th, ‪@SafeAIGermany‬ hosted Guillaume Corlouer (Stormglass, ex Pivotal Research and ‪@PrincInt‬ ), to discuss: 1. Why understanding generalisation is important for AI safety, with empirical examples such as emergent misalignment, where undesirable behaviour emerges because of generalisation. ​2. Using toy models of deep neural networks for AI safety. Toy models are useful because they are mathematically tractable and can give us concrete hypotheses that we can test to understand specific safety-relevant phenomena. As a concrete example, we looked at deep linear networks, and how they can be used to understand the effects of depth and initialisation in deep learning. ​3. Some fellowships and organisations you can join, if you are interested in working on more theoretical directions in AI safety. ​Speaker profile ​​Guillaume Corlouer is a research scientist at Stormglass (formerly Pivotal Research and PIBBSS) and is working on developing toy models of deep learning that can be leveraged to improve our understanding of safety-relevant phenomena. ​Since 2023, Guillaume has been working in AI safety focusing on interpretability and better understanding the learning dynamics of deep neural networks. Previously, he did his PhD at the University of Sussex on estimating information flow in the brain, motivated by questions in the neuroscience of consciousness. About SAIGE: Safe AI Germany (SAIGE) is building Germany's infrastructure for safer AI. Our online events are open to everyone everywhere. 🌐 safeaigermany.org 📅 Upcoming events: luma.com/saige 🐦 Follow us on LinkedIn: linkedin.com/company/safe-ai-germany Recorded 8 July 2026.

Physics-Informed Research for Ambitious AI Safety - Ari Brill
▶︎

Physics-Informed Research for Ambitious AI Safety - Ari Brill

Exploration Hacking: Can LLMs learn to resist RL training? - Joschka Braun
▶︎

Exploration Hacking: Can LLMs learn to resist RL training? - Joschka Braun

One Bias After Another - Daniel Fein & Max Lamparth
▶︎

One Bias After Another - Daniel Fein & Max Lamparth

Communicating AI Safety: From concern to action - Justus Baumann
▶︎

Communicating AI Safety: From concern to action - Justus Baumann

SICSS Istanbul 2026 | Reproducible AI: Replication, Prompt Stability in Social Science (ChrisBarrie)
▶︎

SICSS Istanbul 2026 | Reproducible AI: Replication, Prompt Stability in Social Science (ChrisBarrie)

Threat Models & Mitigations from Physical AGI - Benjamin Alt
▶︎

Threat Models & Mitigations from Physical AGI - Benjamin Alt

Chichvarkin Saves Putin | Vitaly Portnikov
▶︎

Chichvarkin Saves Putin | Vitaly Portnikov

CHOSEN ONE!! YOUR IDENTITY REVEAL JUST SHOOK THE INTERNET... AND THEIR MINDS
▶︎

CHOSEN ONE!! YOUR IDENTITY REVEAL JUST SHOOK THE INTERNET... AND THEIR MINDS

URGENT UPDATE - Iran War Expert: A Mass Casualty Attack Is Coming! | Robert Pape
▶︎

URGENT UPDATE - Iran War Expert: A Mass Casualty Attack Is Coming! | Robert Pape

Training Sand to Think: Artificial General Intelligence & Future of Physics
▶︎

Training Sand to Think: Artificial General Intelligence & Future of Physics

How To Think SO Clearly People Assume You're Brilliant
▶︎

How To Think SO Clearly People Assume You're Brilliant

Instant Focus Mode – 40Hz Gamma Brainwave Music for Deep Focus & Productivity
▶︎

Instant Focus Mode – 40Hz Gamma Brainwave Music for Deep Focus & Productivity

"A.I. and Our Economic Future," Professor Chad Jones
▶︎

"A.I. and Our Economic Future," Professor Chad Jones

A Top Mathematician's 9 Lessons for Anyone Who Feels Behind | Ken Ono, Axiom Math
▶︎

A Top Mathematician's 9 Lessons for Anyone Who Feels Behind | Ken Ono, Axiom Math

FULL DISCUSSION: Google's Demis Hassabis, Anthropic's Dario Amodei Debate the World After AGI | AI1G
▶︎

FULL DISCUSSION: Google's Demis Hassabis, Anthropic's Dario Amodei Debate the World After AGI | AI1G

A Nobel Laureate's Honest Review Of AI In Biology
▶︎

A Nobel Laureate's Honest Review Of AI In Biology

The Next 10 Years of AI Will Change Everything  | Alexander Wissner-Gross | TEDxBoston
▶︎

The Next 10 Years of AI Will Change Everything | Alexander Wissner-Gross | TEDxBoston

Talent Needs in AI Safety - Tilman Räuker
▶︎

Talent Needs in AI Safety - Tilman Räuker

Verification for low-trust AI governance - Naci Cankaya
▶︎

Verification for low-trust AI governance - Naci Cankaya

How AI agents & Claude skills work (Clearly Explained)
▶︎

How AI agents & Claude skills work (Clearly Explained)

Mayor Zohran Mamdani on Socialism, Politics & NYC | The Weekly Show with Jon Stewart
▶︎

Mayor Zohran Mamdani on Socialism, Politics & NYC | The Weekly Show with Jon Stewart

Ms. Rachel Answers Parenting Questions | Tech Support | WIRED
▶︎

Ms. Rachel Answers Parenting Questions | Tech Support | WIRED

Personne ne réalise ce que Yann LeCun vient de créer
▶︎

Personne ne réalise ce que Yann LeCun vient de créer

Europe 2031: What getting AI wrong means for us - Alex Petropoulos
▶︎

Europe 2031: What getting AI wrong means for us - Alex Petropoulos