NLP Evaluation Explained — BLEU To MMLU To The Leaderboard Trap (CL L10E6)
2002. IBM ships BLEU — one number to score a machine translation. The whole field runs on it for twenty years. And knows it disagrees with humans the whole time. Here's every major NLP evaluation metric, why they fail, and how to actually measure a model. This is video 96 of 100 in *Computational Linguistics — Teaching Computers to Read*. Part of Lecture 10 — the final lecture, on frontiers, ethics, and course wrap. In this video: • BLEU (Papineni et al. 2002) — n-gram overlap, brevity penalty, and its fatal flaw • Callison-Burch 2006 — the classic 'BLEU disagrees with humans' warning • ROUGE (Lin 2004) — same idea for summarization • BERTScore (Zhang et al. 2020) — semantic similarity finally beats surface overlap • MMLU (Hendrycks et al. 2020) — 57 subjects, 15K multiple-choice questions • Chatbot Arena and Bradley-Terry human eval — with length and style biases • Goodhart's Law and the leaderboard trap • The practitioner stance — never trust a single number — TIMESTAMPS — 0:00 Cold open — BLEU ran the field for twenty years 0:30 Where we left off — measuring open-ended text 1:15 Today's three beats 1:50 Beat 1 · BLEU and ROUGE (surface overlap) 3:35 Beat 2 · BERTScore, MMLU, human eval 5:20 Beat 3 · Goodhart's Law and the leaderboard trap 7:05 Three takeaways 7:45 Three quick quizzes 8:42 Never trust one number 9:00 Next time — RLHF and alignment — SOURCES & FURTHER READING — • Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). 'BLEU — a Method for Automatic Evaluation of Machine Translation.' *ACL*. • Callison-Burch, C., Osborne, M., & Koehn, P. (2006). 'Re-evaluating the Role of BLEU in Machine Translation Research.' *EACL*. • Lin, C.-Y. (2004). 'ROUGE — A Package for Automatic Evaluation of Summaries.' *ACL Workshop*. • Zhang, T., et al. (2020). 'BERTScore — Evaluating Text Generation with BERT.' *ICLR*. • Hendrycks, D., et al. (2020). 'Measuring Massive Multitask Language Understanding.' *ICLR 2021*. (MMLU) • Chiang, W.-L., et al. (2024). 'Chatbot Arena — An Open Platform for Evaluating LLMs by Human Preference.' *ICML*. • Liang, P., et al. (2022). 'Holistic Evaluation of Language Models (HELM).' *TMLR*. • Jurafsky, D., & Martin, J. H. (2025). *Speech and Language Processing*, 3rd ed. Chapter on evaluation. Open at stanford.edu/~jurafsky/slp3. • Primary course textbook: Mitkov (ed.), *The Oxford Handbook of Computational Linguistics*, 2nd edition (OUP, 2022). — ABOUT THE CHANNEL — Most NLP channels teach you what a transformer is. We teach you why language is hard enough to need one in the first place. Hosted by Ani — taught by a linguist (not a CS bro). ▶ Full playlist: Computational Linguistics — Teaching Computers to Read #NLPEvaluation #BLEU #ROUGE #BERTScore #MMLU #ChatbotArena #Goodhart #NLP #ComputationalLinguistics #Linguistics #JurafskyMartin

AI Is About to Crash. Here’s Why.

4. Anatomy & Physiology Podcast: The Anatomical Position – The Body's Reference Map

Publishers Begged Valve to Remove This Steam Feature

The Scariest Chart in Electrical Engineering

Training Sand to Think: Artificial General Intelligence & Future of Physics

Psychology of People With Extremely High IQ

TypeScript in Express – TypeScript Tutorial

The Most Important Conversation in AI Right Now

The Strangest Things that Correlate with IQ

Keynote: After the AI Hype – What’s Real, and What’s Next - Richard Campbell - 2026

THE AI BUBBLE is Collapsing: Tech Crashes

Is This Wish Meant to Be Fulfilled? 🧚🤲 Detailed Pick a Card Tarot Reading ✫・

Everything You Buy is About to Change

URGENT UPDATE - Iran War Expert: A Mass Casualty Attack Is Coming! | Robert Pape

How To Think SO Clearly People Assume You're Brilliant

Chosen One, This Is Why God Kept You Single All This Time

This Bassett Hound WAS Trying to Tell me Something…He Was Right | PUPDATE

