📊 Full opportunity report: Engineering Is Automated. Research Is the Residual. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
AI systems have achieved near-complete automation of core engineering tasks in AI research, according to recent benchmarks and expert analysis. However, the automation of AI research itself remains uncertain and less developed.
Recent advancements in artificial intelligence have demonstrated that AI systems can now automate the core engineering tasks involved in AI research, with benchmarks approaching saturation. However, the automation of the research process itself remains uncertain, leaving open questions about the future division of labor between AI and human researchers.
Multiple independent benchmarks—CORE-Bench, MLE-Bench, and kernel design research—show AI systems are nearing or have achieved near-complete automation of specific engineering skills essential to AI research. For example, CORE-Bench, which measures the ability to reproduce research experiments, reached a 95.5% success rate by December 2025, with the benchmark’s author stating it is ‘solved.’ Similarly, MLE-Bench, assessing performance on Kaggle competitions, hit 64.4% in early 2026, surpassing mid-tier human performance. These trends suggest that AI can now handle tasks such as reproducing research, optimizing kernels, and managing complex engineering workflows with minimal human oversight.
Despite these advances, Clark notes that the capacity for AI to conduct original research—such as formulating hypotheses, designing experiments, and generating novel scientific insights—remains less certain. The structural question posed by Clark and others is whether research itself is becoming a form of large-scale engineering, which AI can automate at scale, or whether genuine scientific creativity still requires human intuition and inspiration. The current evidence indicates that engineering aspects are largely automated, but the residual research process, involving creative and interpretive tasks, may still depend on human input.
Engineering is automated.
Research is the residual.
Six skill benchmarks. Edison’s framing. The question Clark leaves open is whether research is just engineering at scale.
Jack Clark’s Import AI #455 catalogs six benchmarks measuring AI capability on AI R&D tasks and concludes “AI can today automate vast swatches, perhaps the entirety, of AI engineering.” The residual question is research. The structural read on the residual: it may not be a permanent moat.
Six skills. One trajectory.
Clark catalogs six benchmarks measuring AI capability on AI R&D-relevant tasks. Each individual benchmark could be noise. Six benchmarks moving together is a curve. The pattern is the cascade observed across the broader Clark series — visible here in the specific R&D-skill domain.

CLAUDE AI UNLEASHED From First Prompts to Pro: The Complete Guide to Claude AI for Writing, Research, Coding, and Business (The Claude AI Mastery Series)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three data points. Mixed signal.
Clark provides three data points on the creative-spark question. Yes-evidence: Erdős-1051, centaur math discovery, sporadic Move-37-style moments. No-evidence: low yield, framing dependence, absence of acceleration. The mixed signal is the honest read.
The data supports two readings. Pessimistic: rare moments suggest creative insight is qualitatively distinct from engineering work. Optimistic: rare moments are an artifact of low-volume exploration; more shots on goal yields more discoveries. Both readings are consistent with Clark’s “vast swatches, perhaps the entirety” claim. They differ on the residual.
Five dimensions Clark gestures at but leaves underdeveloped.
Clark’s section is rigorous on the empirical evidence. Five strategic dimensions matter for the institutional response that the Clark series synthesis argues is structurally inadequate.
Two readings. Different equilibria.
The structural question Clark leaves open: is research a permanent moat that bounds automated AI R&D, or is it engineering at scale that dissolves with more shots on goal? Both readings are consistent with the current data. They differ by orders of magnitude in consequences.
Productivity multiplier years
Recursive loop operational
Five audiences. Asymmetric cost of being wrong.
The institutional response should not bet on inspiration being a permanent moat. If the distinction holds, capacity built is still useful. If it closes, capacity is necessary. Asymmetric cost-of-being-wrong points toward building now.
IN INDUSTRY
IN ACADEMIA
POLICYMAKERS
INVESTORS
EVERYONE ELSE
Engineering is automated. The residual is the question. The institutional response should not bet on inspiration being a permanent moat.
Implications for AI Research and Development Processes
The rapid automation of core engineering tasks in AI research signifies a potential shift in how AI development is conducted, possibly reducing the time and cost associated with experimental workflows. If engineering can be fully automated, the bottleneck may shift to the creative and hypothesis-driven aspects of research, which remain less automatable. This development could accelerate AI innovation but also raises questions about the future role of human researchers and the nature of scientific discovery in AI.
Progress in AI Capabilities and Benchmark Saturation
Over the past 18 months, multiple benchmarks have shown consistent progress towards automating AI engineering skills. CORE-Bench, which measures research reproduction, improved from 21.5% in September 2024 to 95.5% in December 2025. MLE-Bench, assessing Kaggle competition performance, rose from 16.9% in October 2024 to 64.4% in February 2026. Additionally, advances in kernel design—such as automated GPU kernel generation—demonstrate that AI systems are producing production-grade infrastructure components. These patterns indicate a saturation point in measuring AI’s engineering capabilities, suggesting that much of this work can now be delegated to AI systems.
“The structural read is that research may itself be engineering at scale — in which case the residual closes faster than Clark’s framing implies.”
— Thorsten Meyer
Unclear Extent of AI’s Capacity to Automate Scientific Discovery
While engineering tasks show near-complete automation, it remains uncertain how much of the research process—such as hypothesis generation, experimental design, and interpretation—can be automated. Clark leaves open whether research itself is evolving into a form of large-scale engineering that AI can handle at scale or if genuine scientific creativity still requires human insight. The timeline and feasibility of fully automating research remain unresolved.
Next Steps in Measuring and Developing AI Research Automation
Research will focus on developing benchmarks and experiments to test AI’s capacity for scientific discovery beyond engineering tasks. Expect to see ongoing efforts to automate hypothesis generation, experimental design, and scientific interpretation. Additionally, industry and academia are likely to explore hybrid models where AI handles engineering while humans guide the creative aspects—until further breakthroughs are achieved in automating the full research cycle.
Key Questions
What are the main benchmarks indicating AI automation of engineering tasks?
CORE-Bench measures research reproduction, MLE-Bench assesses Kaggle competition performance, and kernel design research evaluates infrastructure optimization. All show AI nearing or reaching saturation in automation capabilities.
Does this mean human researchers are no longer needed?
Not entirely. While engineering tasks are increasingly automated, the creative and hypothesis-driven aspects of research still require human insight. The future may see a division of labor between AI and humans.
What are the risks or downsides of fully automating AI research?
Potential risks include reduced scientific diversity, over-reliance on AI-generated hypotheses, and challenges in ensuring AI-generated research aligns with ethical and safety standards. These issues require careful oversight.
When might AI fully automate the entire research process?
It remains uncertain. Experts suggest that while engineering automation is imminent, achieving full automation of scientific discovery could still be years away, depending on breakthroughs in AI creativity and interpretive reasoning.
Source: ThorstenMeyerAI.com