NLP Evaluation Explained — BLEU To MMLU To The Leaderboard Trap (CL L10E6)

2002. IBM ships BLEU — one number to score a machine translation. The whole field runs on it for twenty years. And knows it disagrees with humans the whole time. Here's every major NLP evaluation metric, why they fail, and how to actually measure a model. This is video 96 of 100 in *Computational Linguistics — Teaching Computers to Read*. Part of Lecture 10 — the final lecture, on frontiers, ethics, and course wrap. In this video: • BLEU (Papineni et al. 2002) — n-gram overlap, brevity penalty, and its fatal flaw • Callison-Burch 2006 — the classic 'BLEU disagrees with humans' warning • ROUGE (Lin 2004) — same idea for summarization • BERTScore (Zhang et al. 2020) — semantic similarity finally beats surface overlap • MMLU (Hendrycks et al. 2020) — 57 subjects, 15K multiple-choice questions • Chatbot Arena and Bradley-Terry human eval — with length and style biases • Goodhart's Law and the leaderboard trap • The practitioner stance — never trust a single number — TIMESTAMPS — 0:00 Cold open — BLEU ran the field for twenty years 0:30 Where we left off — measuring open-ended text 1:15 Today's three beats 1:50 Beat 1 · BLEU and ROUGE (surface overlap) 3:35 Beat 2 · BERTScore, MMLU, human eval 5:20 Beat 3 · Goodhart's Law and the leaderboard trap 7:05 Three takeaways 7:45 Three quick quizzes 8:42 Never trust one number 9:00 Next time — RLHF and alignment — SOURCES & FURTHER READING — • Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). 'BLEU — a Method for Automatic Evaluation of Machine Translation.' *ACL*. • Callison-Burch, C., Osborne, M., & Koehn, P. (2006). 'Re-evaluating the Role of BLEU in Machine Translation Research.' *EACL*. • Lin, C.-Y. (2004). 'ROUGE — A Package for Automatic Evaluation of Summaries.' *ACL Workshop*. • Zhang, T., et al. (2020). 'BERTScore — Evaluating Text Generation with BERT.' *ICLR*. • Hendrycks, D., et al. (2020). 'Measuring Massive Multitask Language Understanding.' *ICLR 2021*. (MMLU) • Chiang, W.-L., et al. (2024). 'Chatbot Arena — An Open Platform for Evaluating LLMs by Human Preference.' *ICML*. • Liang, P., et al. (2022). 'Holistic Evaluation of Language Models (HELM).' *TMLR*. • Jurafsky, D., & Martin, J. H. (2025). *Speech and Language Processing*, 3rd ed. Chapter on evaluation. Open at stanford.edu/~jurafsky/slp3. • Primary course textbook: Mitkov (ed.), *The Oxford Handbook of Computational Linguistics*, 2nd edition (OUP, 2022). — ABOUT THE CHANNEL — Most NLP channels teach you what a transformer is. We teach you why language is hard enough to need one in the first place. Hosted by Ani — taught by a linguist (not a CS bro). ▶ Full playlist: Computational Linguistics — Teaching Computers to Read #NLPEvaluation #BLEU #ROUGE #BERTScore #MMLU #ChatbotArena #Goodhart #NLP #ComputationalLinguistics #Linguistics #JurafskyMartin