The strongest part, to me, is the explanation that Universe’s basic loop wasn’t the missing piece: pretrained models, tasks within reach of sparse-reward RL, and interfaces beyond raw pixels arrived later. The benchmark-to-training transition adds an operational wrinkle, especially Terminal-Bench’s warning about benchmark data in training; I’d read the timeline as a causal sketch rather than proof of which change mattered most.