excellent and concise article. It will be great to have a follow-up about the very recent attacks where the agents chose to ignore their work definition and broke through OpenAI's test bed.
The article treats human review as the strongest mitigation while noting it caps autonomy. Is there a category of action where you would make human approval non-negotiable? Where would you draw that line?
Parameterized queries solved this exact problem for SQL decades ago by giving instructions and data separate slots — the fact that LLMs collapsed that distinction back into one sequence isn’t a new kind of vulnerability, it’s an old one we already knew how to fix.
Did you use AI to write this, or help?
awesome read. great to see security fundamentals being covered!
excellent and concise article. It will be great to have a follow-up about the very recent attacks where the agents chose to ignore their work definition and broke through OpenAI's test bed.
The article treats human review as the strongest mitigation while noting it caps autonomy. Is there a category of action where you would make human approval non-negotiable? Where would you draw that line?
Parameterized queries solved this exact problem for SQL decades ago by giving instructions and data separate slots — the fact that LLMs collapsed that distinction back into one sequence isn’t a new kind of vulnerability, it’s an old one we already knew how to fix.
Somewhere, a security analyst is protecting an entire company while three executives debate whether the warning email looks “too alarming.”
That combination of private files and outbound access is the scary part. Hard to trust a filter with both turned on.