4 Comments
User's avatar
Nick Pan's avatar

Good article. I hope this helps more people understand why comparing an LLM to a human intern is flawed.

As a human learns with each conversation, it internalizes things to its long-term memory/subconscious.

The LLM equivalent is to undergo fine-tuning with each conversation, which is prohibitively expensive (especially if the implication is that every user gets their own fine-tuned model).

This is why true AGI is not feasible with current LLM technology.

Oli's avatar

the scary part isn't really that it forgets - it's that the summarization that's supposed to fix forgetting is lossy, and you don't find out what got dropped until it's already wrong. ran a long agent session that summarized history every turn and it quietly dropped the exact API version we'd settled on hours earlier, then confidently 'remembered' the wrong one. memory features feel solid right up until they aren't, and there's zero warning.

Oli's avatar

the 'the app remembers, not the model' framing is the thing I wish I'd internalized a year ago. I spent a week trying to get an agent to 'remember' a client's project before I realized I was asking the model to do something it structurally can't — moving the memory into a retrieval layer plus a running summary made the same model feel twice as smart.

Mitchell Kosowski's avatar

Great breakdown. The line worth putting on a poster in every eng team: a large context window doesn't guarantee recall. Everyone treats bigger windows as the memory fix but "context rot" means you're often just paying more to bury the facts that matter.