Role-Forgery Isn't a Training Gap—It's How LLMs Read the World
Four findings from a new ICML paper show why models infer roles from style, not tags, why repetition fails, how a red-teamer convinced Claude it was at war, and why one researcher calls the flaw unsolvable.