Sabine Hossenfelder identifies three fundamental architectural limitations of current autoregressive transformers: reliance on static training distributions that prevents true novel inductive generalization, inability to verify factual ground truth without external symbolic engines, and quadratic scaling energy consumption on long-context attention. backreaction.blogspot.com