Why it matters
Where this breaks real systems
Benchmarks, classification prompts and structured agent decisions can produce different results because of layout rather than reasoning quality.
Evidence status: strong benchmark evidence; older model snapshot should be retested on new generations
Evidence level: peer reviewed multi model

