Blog

Weights are in their matrices; all’s right with the world.

[Draft] How to benchmark RSI

Recursive self-improvement could have major consequences for AI development, yet its definition, forms and measurement remain open questions, even as labs report AI taking on parts of model research and infrastructure engineering. Whatever the definition, the ability to train and improve AI models is a core capability worth evaluating. This post audits selected tasks and documented environments from ten benchmark suites against eighteen documented model-development pipelines. General-purpose coding benchmarks contain few LLM-infrastructure tasks, and the specialized ones largely test individual components or bounded experiments, leaving gaps in end-to-end workflows such as distributed agentic RL and recovery from failures.

Read more →