Discussion about this post

User's avatar
Kevin M's avatar

> A larger, stronger teacher does not always produce a better student.

This is just like as humans learning from human teachers!

Although it depends on what we are trying to learn, learning from the best of the best is not always the best option, especially when the student is new to the field.

Raelon Masters's avatar

Great explainer on distillation. The compression vs. distillation distinction is one I see people get wrong all the time.

I run Ollama + OpenWebUI on an R630 homelab, and distillation is exactly why running local models is practical — the smaller Gemma/Phi-3 models punch way above their weight on CPU-only hardware.

For anyone curious about the practical side — Docker Compose configs, real token/sec numbers on actual server iron — I just launched The Self-Hosted Stack at selfhostedstack.substack.com.

No posts

Ready for more?