Diffusion Language Models, LLaDA, Nemotron-TwoTower and more | Bangalore Paper Club
Large language diffusion models, efficient diffusion architectures, graph-language model evaluation, and dogs that smell cancer - four papers in one evening at the Bangalore Paper Club by Conscious Engines. Inspired by the YC Paper Club, we started an offline meetup series to meet researchers, present papers, discuss interesting deep tech problems, and work together on some of them. This edition goes deep on diffusion as an alternative to autoregressive generation - a diffusion LLM trained from scratch and an architecture that retrofits diffusion onto a frozen AR backbone - then turns to whether graph-language models actually fuse structure and text, and closes with a strikingly low-cost cancer triage test built on trained dogs and Bayesian modeling. ━━━━━━━━━━━━━━━━━━━━ CHAPTERS ━━━━━━━━━━━━━━━━━━━━ 00:00:00 Introduction 00:04:56 LLaDA: Large Language Diffusion Models 00:17:10 Nemotron-TwoTower: An Efficient Architecture for Diffusion Language Models (NVIDIA) 00:37:57 A Graph Talks, But Who's Listening? Rethinking Evaluations for Graph-Language Models (ACL 2026) 00:54:19 Canine Olfaction + Bayesian Modeling for Multicancer Detection (JCO 2025) ━━━━━━━━━━━━━━━━━━━━ PAPERS DISCUSSED ━━━━━━━━━━━━━━━━━━━━ 1. LLaDA: Large Language Diffusion Models Presented by Rishi Panda. Instead of writing left to right one token at a time, LLaDA masks and denoises whole sequences — a pure diffusion model trained from scratch to 8B parameters on 2.3T tokens, going toe-to-toe with LLaMA3 8B on in-context learning and edging out GPT-4o on reversal reasoning. Evidence that the autoregressive monopoly on capable LLMs may be an accident of history, not a law of nature. https://arxiv.org/abs/2502.09992 2. Nemotron-TwoTower: An Efficient Architecture for Diffusion Language Models Presented by Hari Prasad. NVIDIA splits the diffusion LM into two networks on a shared backbone: a frozen autoregressive "context tower" that reads, and a trained "denoiser tower" that writes, talking through cross-attention. Bolted onto a 30B Nemotron backbone and trained on ~2.1T tokens — a fraction of the backbone's 25T — it delivers 2.42× faster generation while keeping 98.7% of quality. A serving speedup you buy without re-pretraining from scratch. https://arxiv.org/abs/2606.26493 3. A Graph Talks, But Who's Listening? Rethinking Evaluations for Graph-Language Models Presented by Soham Petkar. Today's graph-language models look multimodal but quietly cheat — acing benchmarks using text or structure alone without ever fusing the two. The authors expose it with behavioural and mechanistic analysis, show plain soft-prompted LLMs keep pace with far heavier GNN-backed machinery, and ship CLEGR: 1,000+ synthetic graphs and 54,000 questions that finally force graph and language to reason together. https://aclanthology.org/2026.finding... 4. Canine Olfaction Combined With Bayesian Modeling for Multicancer Detection From Breath Samples: A Phase II Study in India Presented by Akash Kulgod. Trained dogs sniffed out seven major cancer groups from a single breath sample across 3,275 people in Karnataka — no needles, no scanners. Fuse the dogs' calls with a Bayesian model that weighs each dog's track record, and the system hits ~91% sensitivity and ~91% specificity (AUC 0.962), holding up even for stage I–II disease. A sub-dollar triage test built for low- and middle-income screening — proof a wagging tail can be a diagnostic instrument. https://ascopubs.org/doi/10.1200/JCO-... ━━━━━━━━━━━━━━━━━━━━ ABOUT CONSCIOUS ENGINES ━━━━━━━━━━━━━━━━━━━━ Conscious Engines is a Bangalore-based AI research lab building small, fast, application-specific models — intelligence that specialises, shrinks, and disappears into products. Small. Instant. Invisible. Website: https://consciousengines.com Blog: https://consciousengines.com/blog X: https://x.com/c_engines Join us: https://consciousengines.com/join Want to present at the next Bangalore Paper Club? Subscribe and turn on notifications - we post each edition's reading list ahead of the session. #DiffusionLLM #LLaDA #Nemotron #GraphLanguageModels #EdgeAI #MachineLearning #PaperClub

Sequence-to-Function Models, Gene Expression Prediction and Variant Effects | LatchBio AI × Genetics

Everything I Learned Training Frontier Small Models — Maxime Labonne, Liquid AI

Mathematical Computation for the Age of Quantum and AI

I Built an LLM From Scratch

Kimi K3 explained in 13min..

MLX India Community Meetup 1 | Boosting local model performance - Speculative decoding with DFlash

Speculative Decoding, Edge Inferencing, LLM Safety and more | Bangalore Paper Club

The Most Important Conversation in AI Right Now

Ilya Sutskever – We're moving from the age of scaling to the age of research

MCP vs API Explained: Do You Really Need MCP?

Fireside Chat with Yann LeCun, Executive Chairman of AMI Labs | RAISE Summit 2026

Inference, Diffusion, World Models, and More | YC Paper Club

Harnesses in AI: A Deep Dive — Tejas Kumar, IBM

MCP vs API: Why traditional APIs are failing AI agents
![Understand AI in 14 minutes – with Anthropic's Chloe Lubinski [ARC 2026]](https://i.ytimg.com/vi/aBUniZHgCnE/hqdefault.jpg?sqp=-oaymwEjCNACELwBSFryq4qpAxUIARUAAAAAGAElAADIQj0AgKJDeAE=&rs=AOn4CLCyQJdkwlip_867U0IUOY4wCWZJ0g)
Understand AI in 14 minutes – with Anthropic's Chloe Lubinski [ARC 2026]

ALIEN CODES AND THEIR AUTOMATED AND HUMAN EXPLANATIONS

Keynote: After the AI Hype – What’s Real, and What’s Next - Richard Campbell - 2026

Visualizing transformers and attention | Talk for TNG Big Tech Day '24

Training Sand to Think: Artificial General Intelligence & Future of Physics

