Why it matters
Where this breaks real systems
Automated evaluation can accidentally reward longer, smoother output instead of instruction following, correctness or usefulness.
Evidence status: established evaluation risk; exact strength is judge- and task-dependent
Evidence level: peer reviewed multi study

