Discussion about this post

User's avatar
Mihir Shah's avatar

A noob question. I learned that LLMs are stateless, and here is an entire article on LLM caching. What am I missing? Kindly help me make a distinction.

Money Machine Newsletter's avatar

The 40 GB example is a good reminder that model size is only part of the bill. During generation, the GPU reads the cache for every new token. Long prompts can slow a system down even when the model fits in memory. That is easy to miss when teams benchmark short chats.

10 more comments...

No posts

Ready for more?