Busier, not smarter: Rethinking who gets to build AGI

Measuring the length of tasks a frontier model can complete autonomously is what many in the AI industry have adopted as proxy for general intelligence. But autonomy is not the same as learning, and trusting an AI with more responsibilities solely because it can work unsupervised for longer periods of time may be a perilous proposition.

If you ask AI experts how close we are to AGI, you may be referred to the measure of how long a model can keep working before a human has to step in. METR (Model Evaluation and Threat Research), an independent research nonprofit based in Berkeley, California, measures exactly this. 

In their latest study, which was last revised in July 2026, they find that the length of tasks a frontier model can complete autonomously has been doubling roughly every four to seven months since 2019, from a few minutes of coherent work to a full working day.  

Today, that’s a number some in the industry have adopted as proxy for general intelligence, with AGI increasingly associated with “an agent that doesn’t need to be checked on as often.” 

Autonomy is not the same as learning. 

What a longer task horizon really measures is endurance. In other words, how far a model can carry a plan before its errors accumulate beyond recovery. What it doesn’t measure is whether the model gets better at the task by doing it. Currently, if the same agent is given the same set of problems a thousand times, the thousandth run draws on no more accumulated skills than the first. 

That becomes clear when we consider longer tasks. Recent work analysing agent failures found that models that succeed on short versions of a task can see their success rate collapse to under 10% as the context grows, with agents failing primarily by losing track of the original goal or getting stuck in loops, even when the relevant information remains present within the model's context window. In short, the knowledge is present but the system struggles to make reliable use of it over time. 

Why the lever only sits in a few hands

The scale of this matters because of how fast AI is being put into production. According to McKinsey, 88% of organisations reported regular use of AI in at least one business function in 2025, up from 78% in 2024 and 55% in 2023. At this pace of adoption, that measurement problem becomes structural. 

If a static model can’t get smarter from its own use, the only way to produce a model with more knowledge is to retrain it. And retraining a frontier model is expensive enough that most organisations can’t afford to do it themselves. 

Concretely, this means that any organisation deploying these systems has no means of changing what they know once they’re in production. Intelligence gets defined, updated and improved centrally, then distributed as a fixed artifact for companies to build on top of. 

What this produces at scale

If we extend this architecture far enough, we end up with a narrow but widespread ecosystem of static AI models that are controlled centrally, where organisations using these models for critical tasks have no control over their evolution. 

These increasingly autonomous systems will become more and more embedded across industries such as healthcare, infrastructure and public services, and their ability to learn from new information will rely entirely on the retraining cycles managed by the few organisations that control these models. Meanwhile, they will carry more unsupervised responsibilities while their knowledge progressively decays. 

The risk here is not just that AGI, on its current path, would concentrate the power of AI with a handful of labs that get to dictate how and who controls the evolution of this technology, or that the current economics of how models learn, with extremely expensive retraining cycles, would prove unsustainable at scale. It’s also that what models learn and when they learn it would not be controlled by the institutions that support the socio-economic pillars of our society. 

AGI 2.0: Intelligence that learns where it’s used

AGI 2.0 is a case for a different architecture, and a different definition of AGI. 

An agent that can act for longer isn’t, on its own, a more generally intelligent one. Equally, a small number of labs updating the knowledge of their models on their own schedule is not, on its own, distributed intelligence. Both point to the same problem: capabilities that only evolve when a small number of actors decide to move it. 

The alternative, and the architecture AGI 2.0 introduces, is context-centric intelligence that learns continuously, at the point of use, from the organisations that use this intelligence. That’s the principle Boltzbit’s General Learning Intelligence is built around: learning as something that occurs wherever the model is deployed instead of being a centralised and periodic event. 

Under this architecture, task horizon can, and should, keep improving. And this will be, in part, the result of systems that are genuinely more capable on the thousandth day than they were at day one because they learn continuously from their users.  

Why this matters now

This proposed architecture is by no means an argument against agentic AI or the impressive progress we’ve seen in task horizon. The argument made here is that the progress made so far should not be viewed as sufficient to consider AGI within close reach, nor should the current version of AGI be regarded as the only one. 

Who gets to hold the pen on the decisions that shape the evolution of the socio-economic pillars of our society is arguably more important than how autonomous AI becomes. An AI trusted with more responsibilities solely because it can work unsupervised for longer periods of time is a perilous proposition. What should too be considered is how much learning an AI can accumulate from the organisations running it, and whether that makes it more reliable as it completes tasks.

Read more
© All rights reserved Boltzbit 2026