Claude Opus 5.5 LLM testing

Claude Opus 5.5 ends all 18 critical defects on Lattice Bench along with six of seven very_hard chains. At GoML, we discovered that this included two cross-file mass assignment takeovers that survived every previous test run. Opus 5.5 LLM testing scored 87.3 on the AI Matic Bench Score, defeating the prior highest score by 7.5 points while costing less per token than the model it replaces.
Originally published on the GoML blog (~1880 words). Read the full piece here: Claude Opus 5.5 LLM testing.