MLX India Community Meetup 1 | Boosting local model performance - Speculative decoding with DFlash

Speculative decoding is a technique to obtain a decoding speedup in LLM inference. Sabesh talks about implementing speculative decoding on MLX using a library called DFlash and about a controlled sweep that was performed right on his MacBook. The findings were documented and published.