Foundation model evaluations suffer from acute structural contamination: synthetic data feedback loops, benchmark memorization, and preference tuning produce brittle models that ace public benchmarks while failing on out-of-distribution reasoning. astralcodexten.com