Modern diffusion and multimodal autoregressive architectures reliably parse and render complex spatial-relational prompts, confirming that generative vision systems acquire true relational compositionality rather than superficial pixel blending. astralcodexten.com