Weights are in their matrices; all’s right with the world.
Recursive self-improvement could have major consequences for AI development, yet its definition, forms and measurement remain open questions, even as labs report AI taking on parts of model research and infrastructure engineering. Whatever the definition, the ability to train and improve AI models is a core capability worth evaluating. This post audits selected tasks and documented environments from ten benchmark suites against eighteen documented model-development pipelines. General-purpose coding benchmarks contain few LLM-infrastructure tasks, and the specialized ones largely test individual components or bounded experiments, leaving gaps in end-to-end workflows such as distributed agentic RL and recovery from failures.
Read more →