Ramez Naam has recently given a careful analysis showing why LLMs upgrading LLMs aren’t about to FOOM—why today’s large language models, set to work building their own successors, aren’t about to explode into superintelligence in a sudden “hard takeoff.” But if you run the same analysis on a supercolony of self-modifying OmegaHives, the kind of AI system my colleagues and I are building, the conditions for hard takeoff start to look reachable—and soon.
Ramez Naam’s guest post on Noah Smith’s blog, Where’s the “intelligence explosion”?, is among the most careful skeptical treatments of recursive self-improvement I’ve seen. Recursive self-improvement, or RSI, means AI being used to build better AI, which then builds better AI still, and so on around the loop. If each trip around the loop comes quicker than the last, you get an “intelligence explosion”—AI shooting far past the human level within a year, or months, or days—a scenario people also call a “hard takeoff” or “fast takeoff.” In a “soft takeoff” AI keeps improving, maybe quite fast, but over many years and with no runaway that feels immediate on a human scale (though it may still be unfolding very fast on the time-scale of societal history!). (See my last blog post for a more careful run-through of “WTF is RSI”)
Naam treats RSI as a feedback loop—better AI, then more research, then useful improvements, then more capable AI, and around again—and asks how strong that loop is today at the big labs building large language models (LLMs) like ChatGPT and Claude. To put a number on it he builds on an analysis by Tom Cunningham and colleagues, which measures AI capability by the Epoch Capabilities Index, or ECI: a single overall score, published by the research group Epoch AI, that combines an AI model’s results on many different tests. The best models have lately been gaining about 16 ECI points a year, and I’ll call one point on this index a “capability point.” The question is then: each time AI gets one capability point better, how much more research does each researcher at the lab get done?
The Cunningham analysis says the loop can sustain itself—each round of gains producing enough to power the next—only if the answer is around 15% more per point.
Where does this magical 15% come from? It rests on two measurements:
Epoch’s data on many AI models: ten times the computing power in training buys about 15 more ECI points, which works out to roughly 15% more computing power per point. A better training algorithm counts the same as more hardware here—one that gets the same result from half the computing is treated like a doubling of the hardware—so with the hardware held fixed, the next point takes about 15% more “algorithmic efficiency.”
The observation that algorithmic efficiency has been tripling each year while the research effort going in (staff, and computers for running experiments) has been roughly tripling too, so 1% more effort has bought about 1% more efficiency.
Put the two together and each further point costs about 15% more research effort than the last. The loop pays its own way only if the point just gained makes the lab that much more productive, with the same people and the same machines. The authors call this a rough calculation, and it is: other published estimates of algorithmic progress would, by my arithmetic, put the bar anywhere from about 7% to about 40%. But even the low end is well above the 2–3% that Naam finds.
Naam’s tweak of Cunningham’s calculation puts the bar nearer 19%. Naam checked Cunningham’s assumptions against OpenAI’s logs. Its researchers ran about 1.6 times as many experiments per person as in 2025, over a stretch in which the best models gained about 16 points, which comes to roughly 3% per point (3% compounded sixteen times is about 1.6); allowing for other things that helped, he settles on 2–3%. So the loop would have to be five to ten times stronger before it could carry itself. And at every stage he finds diminishing returns, with each further unit of effort or money buying less than the one before.

The core difference between “soft-takeoff LLM+humans RSI” and “harder takeoff neural-symbolic Omega RSI”
I think Naam’s analysis is pretty much right, and I also think it’s the best argument I’ve seen for doing RSI a different way than the LLM companies. Every weakness he finds traces back to one feature of the loop he’s measuring. In that loop LLMs help humans work out how to build the next LLM, and nothing they work out makes AI any more capable until it has been built into a new model by a “training run”—months of pushing enormous amounts of data through a neural network on many thousands of specialized chips, at a cost in the hundreds of millions of dollars.
The AGI path my colleagues and I are pursuing is unfolding very differently. An Omega agent—an “agent” is an AI program that carries out tasks on its own, one action after another—is an LLM joined to a second, older kind of AI called “symbolic.” The symbolic part has a knowledge store, the AtomSpace, holding facts, rules, goals and procedures in an explicit form that can be read, checked and edited one item at a time (an LLM’s knowledge, by contrast, is spread through billions of numbers that nobody can read). It also has reasoning programs written in MeTTa, a programming language designed so that programs can inspect and rewrite programs, themselves included. That combination of neural networks (which is what LLMs are) with symbolic AI is what “neural-symbolic” means in this setting The agent design we’re moving to, NESS, replaces standard LLMs in the loop with modified LLMs based on a neural-symbolic sandwich design: an LLM at the bottom, the symbolic layer in the middle, and on top some smaller neural layers that keep learning as the agent works.
An OmegaHive is a team of Omega agents sharing memory and working together, and a supercolony—“colony” for short—is many hives, run by different people, connected by a network we call OmegaBuzz over which they trade results and advice. A hive improves itself one step at a time. Omega agents and simple OmegaHives are now in practical operation, I have a bunch of such agents working for me all the time now building new proto-AGI code and helping with other things. OmegaBuzz is still at an earlier stage of development.
Simplifying a bit, the RSI workflow we are building is like this: To self-improve, an Omega agent picks a “mechanism”—a particular method for some part of thinking, such as a way of reasoning, of remembering, or of deciding what to pay attention to—and builds it into its own code. Then it compares its new self with its old self on the ProtoAGI test suite (currently in progress and partly operational), a set of ten test environments designed to cover the many sides of general intelligence. If the change proves itself it gets “promoted” into the working system; if not, it’s parked for later or rejected. I’ll call the checkpoint where that gets decided the promotion gate, and one trip around the whole loop a “turn.” Changes that pass can be offered to the other hives over OmegaBuzz. Training runs come into it only rarely, since most of what changes is code and knowledge, which can be changed in minutes. Figure 1 shows the two loops side by side.

Go through Naam’s argument point by point, in his own terms, and the two loops—I’ll call them the LLM loop and the OmegaHive loop—come out differently every time:
- Naam expects AI to become superhuman first in “highly verifiable” domains—fields like chess and formal mathematics, where a computer can check for itself whether an answer is right and can make up as many practice problems as it wants. Our ProtoAGI suite is our attempt to turn the building of a better mind into basically that kind of field, so the hives improve themselves in just the territory where, by his own analysis, machines pull ahead fastest.
- He pictures the path from AI activity to more capable AI as a funnel that loses something at every stage, and the last stage—getting a useful improvement built into a new model—is the most expensive of all. In an OmegaHive an improvement is part of the working system the moment it’s promoted, so that stage mostly goes away.
- In the LLM loop a higher score on general tests mostly fails to turn into more research getting done; that’s what the 2–3% per capability point says. In an OmegaHive the things being improved are the things research gets done with—how the hive picks a line of reasoning, what it pays attention to, how it stores and finds what it knows, how its agents coordinate—so each improvement feeds straight into making the next one.
- Running many copies of one AI helps less than you’d hope, he points out: the copies think alike, and by one rule of thumb a hundred of them finish a job only about ten times faster than one. But the hives in a supercolony each have their own history, knowledge and values, and they build on each other’s tested results over time, roughly the way human science and culture accumulate.
- His evidence that research keeps getting harder is about the search for new ideas. For its first stretch, though, the OmegaHive loop doesn’t need to search: decades’ worth of mechanisms are already worked out in research papers and code prototypes, waiting to be scalably implemented—an orchard somebody has already planted, to borrow his apple-picking picture—and each apple picked becomes part of the picker.
Put as plainly as I can: an intelligence explosion happens when each improvement makes the next one cheaper to get. Measure cost in hive-hours (one hive working for one hour). If each improvement a hive accepts—that is, promotes—cuts the hive-hours needed for the next one by a steady fraction, then any number of further improvements, however large, fits into a limited stretch of time; the arithmetic is further down. The OmegaHive loop has specific features pushing it toward that condition, and on Naam’s evidence the LLM loop doesn’t currently meet it. That’s the difference in the heading above: RSI done by LLMs plus humans looks like a soft takeoff at best, while RSI done by neural-symbolic Omega hives could well be a much harder one. And since a turn of the OmegaHive loop is cheap and every test leaves a record, we should know within months of having the whole infrastructure complete and mature whether the condition is being met. At the moment the whole Omega/Hive/Buzz infrastructure is under rapid development—some parts are working very well and others are raw and just prototypes—but it should be months not years before we are in fully swing with the OmegaHive RSI experiment as I’m sketching it here…
Naam’s framework
Naam sorts AI self-improvement into five types, which overlap somewhat:
- Type 1 is AI helping human researchers with their work.
- Type 2 is a stronger AI improving a weaker one.
- Type 3 is AI improving itself under its own direction, with humans still adding something;
- Type 4 is AI improving itself with no need of human help at all.
- Type 5 is fast takeoff, where self-improvement gets the better of diminishing returns and the loop runs away.
Types 1 and 2 are already here; he sees no clear evidence yet of Type 3, and none of Type 4. His central observation is that reaching Types 3 and 4 doesn’t hand you Type 5: an AI could design, train and test its own successor with no human involved and still find that each step forward takes many times the resources of the step before. Taking the humans out of the loop doesn’t by itself make the loop any stronger.
I agree with that, and it tells me what I have to show to argue for the potential for hard takeoff via Omega. It wouldn’t be enough to argue that OmegaHives can get by with less human help than LLM-based research setups can, though in some respects they can. I have to show that one turn of the OmegaHive loop does more to speed up the next turn than one turn of the LLM loop does, that this comes from how the system is built, and that it shows up when you measure things the way Naam measures them. The rest of this post is my attempt.
Why the LLM loop is weak
The crux of the matter is: Every part of Naam’s argument passes through the same step, the training run. What the researchers produce, human and AI alike, is a recipe for training the next model. Most of the experiments in OpenAI’s logs are themselves training runs of one size or another. The capability being scored is that of the next “foundation model”—the big general-purpose neural network at the base of a system like ChatGPT. And the “scaling laws” he cites describe that same step. These are measured rules of thumb for how much more you must put into training to get a given improvement out, and they’re harsh: about 12 times as much data, or 90 times as much computing power (“compute” for short), or a network 8 times as large, just to cut in half the part of a model’s error that each of those controls. Meanwhile Anthropic, in one of its “system cards” (the report a lab publishes on a new model’s abilities and risks), says that internal use of its own models has been a key factor in maintaining the current rate of progress, with no clear signs yet of dramatic acceleration beyond it. Better AI is being used up just keeping progress at the speed it already had.
The weakness comes from where the LLM loop keeps its improvements. They live in the model’s “weights”—the billions of numbers inside the neural network that determine what it does—and weights change through training. An OmegaHive keeps most of its improvements elsewhere: as code, knowledge and decision rules in the AtomSpace, and in the top layers of the NESS sandwich that keep on learning. There they can be added one at a time, tried out in “shadow mode” (running alongside the existing system on the same inputs, without yet being in control of anything, so the two can be compared), measured, and taken out again if they turn out to be mistakes. That makes a turn of the loop something like a thousand to a hundred thousand times cheaper—seconds to days and dollars to thousands of dollars, against months and hundreds of millions. It also makes each turn count for more, as I’ll argue below.
A highly verifiable domain
Naam gives four conditions that make a field “highly verifiable”:
- the work can be checked without error;
- there’s no limit to the practice material;
- feedback comes at computer speed, with no waiting on people or the physical world;
- the score measures the thing you care about.
Where all four hold he expects “narrow superintelligence”—AI far better than any human at that one kind of work, as chess programs already are at chess. Our ProtoAGI evaluation suite—currently under very active development by various agents and humans—has ten environments: AGI Maze (mazes to explore and map), Neoterics (an artificial ecosystem), RoboGarden (a garden worked by robots), RepoOps (software projects), LeanGarden (mathematical proofs), Virtual Scientist (a simulated science lab), SocietyLab (a society of other agents), HiveForge (projects a whole hive carries out together), SelfLab (the hive governing and modifying itself) and the Transfer Ring (reusing in one environment what was learned in another). One way to describe the whole thing is as a deliberate attempt to make “building a general mind” highly verifiable, so the hives improve themselves on the kind of ground where machines improve fastest.
Take checking first. Wherever it can, the suite gives the final say to something that can’t be fooled or argued with. In LeanGarden that’s a proof checker, a program that goes through a mathematical proof line by line and confirms that every step follows. In SelfLab, what a hive is and isn’t permitted to do is decided by a simple program applying fixed rules, with nothing learned or fuzzy in it. And all ten environments share one routine for a single test run, or “episode”: before a hive takes any action with consequences it puts on record what it predicts will happen, and afterward the environment hands back a signed record, a “receipt,” of what did happen. That way anyone can go back later and check how accurate the hive’s picture of the world was, and how well it knew its own abilities.
Some environments are fuzzier. In Virtual Scientist, SocietyLab and Neoterics there’s often no mechanical way to check an answer, so the scoring relies on indirect measures and on AI judges trained to grade the work. A hive that set out to cheat could probably fool these, and we haven’t tried to make that impossible, because the suite is meant as a measuring instrument for hives that want to know whether they’ve improved, and what it needs is to be usable for that by hives acting in good faith. The work that has gone into testing the tests is all aimed at helping a well-meaning hive avoid fooling itself. Each test is first tried on narrow single-purpose programs, to see what poor and excellent scores look like. Agents built to behave badly are run through to make sure they score badly. Problems are given again in a different disguise, to catch a hive that has memorized the look of a problem without learning to solve it. Some of the programs that generate test problems are held back until a hive’s code is locked, so it’s graded on problems it can’t have seen. And the overall result leans toward whatever the hive is worst at, so it can’t look good by piling up points where it’s already strong.
Deliberate cheating gets dealt with socially. A hive that cheats on the tests and then advertises itself as smarter than it is has made a claim that other hives can check: they can take its code changes, run them on fresh copies of themselves, and see whether the gains appear. If they don’t, the hive’s reputation on OmegaBuzz suffers, and the others stop adopting its changes and stop asking its advice. This is roughly how human science copes with cheating, and it works well enough to keep knowledge accumulating.
Then practice material. None of the environments is a fixed list of problems; each is a generator, a program that can go on producing new problems indefinitely. And once an episode has been scored and closed, most of what happened in it goes back to the hives as study material—the procedures that worked, the proofs that were found, the diagnoses of what went wrong, the records of how well predictions matched outcomes. The suite is partly an exam and partly a school, and the more it’s used the more there is to learn from.
Then speed. The simplest kind of test takes a hive as it stands, makes one change to its code or knowledge, and asks whether the changed hive does better right away; tests like that run as fast as the computers go. There’s a slower kind as well, for changes that affect how a hive learns: two versions start from identical saved copies of the hive, go through the same sequence of experiences, and get compared on how quickly each learned and how much each forgot. Those take longer, but they’re still very cheap next to a training run for a large LLM.
And then the fourth condition, scoring the thing you care about. In place of a single number, the suite gives what I’ve been calling an evaluation ecology—a profile across many things at once: how wide the hive’s range of abilities is, how fast it learns, how well it holds on to what it has learned, whether it can use a skill in a new setting, how well its confidence matches how often it’s right, how well its agents work together, how reliably it stays within the rules it’s been given, and what all of this costs. The weakest areas count for at least as much as the average. And hives that are also doing real work for paying customers get a steady supply of evidence from outside the suite about whether gains on the tests carry over to the real world.
Now hold the LLM loop up to the same four conditions. The thing being improved, the model, is scored by an index that blends a lot of separate tests (the ECI). The research that improves it can be observed only from outside, by counting experiments and surveying the staff. And the feedback that counts most—is the new model better?—arrives once per training run. None of that is highly verifiable in Naam’s sense, which is why he has to work out the strength of the loop indirectly from logs and surveys, and it’s a good part of the reason, I would say, that the loop he finds is so weak.

The horizon problem
The number in Naam’s post that I’d pick out above the others is the following. OpenAI sorted its own internal research tasks by how long each would take a human, and looked at how often its models got a task done with no human stepping in. The models managed that 80% of the time only for tasks up to about fifteen minutes long. The longest task an AI can be relied on to finish by itself four times out of five is called its “task horizon,” and the public benchmarks (the standard tests used to compare AI models) and the best-known forecasts had put the horizon for models like these at somewhere between four and eleven hours. On real research work, then, it’s far shorter than the tests suggest. In the LLM loop the horizon belongs to the model itself, and gets longer only when a new model is trained.
In an OmegaHive the horizon depends partly on how the system around the model is built. One reason an LLM loses the thread on a long task is that it can hold only so much text in view at once—its “context window”—and whatever falls outside is forgotten. An Omega agent keeps the state of a task somewhere more durable: a Context Frame, a structured record in the AtomSpace of the task’s goals, hypotheses, plan, evidence so far, and what would count as finished. The agents run continuously, in a loop we call Iter, and they have a memory store (petta-memory) that carries over from one working session to the next. A “metareasoning” layer—reasoning about how the reasoning is going, and what to do about it—checks on progress along the way. And the promotion gate, in its full form, follows a design we call GOLEM-Iter, under which any change a hive wants to make to itself must be logged, tested and approved before it takes effect, with a cap on how fast such changes can come. On top of all this, the routine of build, test, tune and promote chops research into steps that are each short and checkable, so the horizon a hive needs is the length of one build-and-test step and not of a whole research project. And, what bears most on my argument, every promoted mechanism that improves memory, planning or metareasoning lengthens the horizon for all the work that comes after. The LLM loop gets a gain like that only once per training run.
Naam’s funnel and ours
Naam’s Figure 25 lays out the path from AI activity to AI capability in four stages—better or more AI, then more or better research, then useful improvements to AI, then more capable AI—and names what gets lost at each: the gains aren’t one-for-one, bottlenecks remain, discoveries get harder, and the gains don’t all carry through. His numbers from OpenAI show the size of the losses. Per person, OpenAI’s researchers were using about 124 times as many “tokens” as before—tokens being the small chunks of text an LLM reads and writes, so this measures how much work the AI was doing. Its engineers were turning out roughly 7 times as much code. Its researchers were running 1.6 times as many experiments. And all of that, by Naam’s own rough guess, may have made AI improve about 10% faster. Figure 2 puts his funnel beside the OmegaHive version of the same four stages, and the paragraphs afterwards take the stages one at a time (I’m doing my best to spell all this out steeeeepppp by painstaking step, because these matters are so hot-button these days and so prone to confusion…!)

At the first stage, what’s lost in the LLM loop is that most of what the AI produces isn’t research: the tokens become code, and only some of the code becomes experiments. In an OmegaHive the activity is research from the start, because of the way the loop is set up—every episode run on the test suite is an experiment, and every step of building a mechanism ends in one. So a much larger share of the AI’s output ends up as experiments, and each experiment leaves a record for future analysis and understanding.
At the second stage the bottlenecks are compute, coordination among the researchers, and human judgment. In our loop the hive does what the figure calls plumbing and orchestration, the routine chores that eat most of a human researcher’s week: getting the code to build, repairing the test setup when it breaks, trying out long lists of settings, scheduling and running the tests. Coordination within a hive is handled by its own procedures for dividing up work, and coordination between hives by OmegaBuzz. That frees the humans for the places where a hive gets stuck—choosing which direction to take next, making sense of a puzzling result, deciding that something shouldn’t be done. And running the suite doesn’t take much compute.
The third stage, discoveries getting harder, has a section to itself further down (“Apples and the picker”). TL;DR is that we have a large pool of prototype results for the Omegas to feed on, which may get us all the way to human-level AGI and if not should get us close.
The fourth stage is where the two loops differ most. In the LLM loop a useful improvement—a better mix of training data, say, or a better training procedure—becomes more capable AI only once it has been trained into the next model, months later and at great expense, and a good deal of what looked useful in small experiments turns out not to help at full scale. In an OmegaHive a mechanism that gets through the promotion gate is working capability right away in the hive that built it, and in the rest of the colony as soon as other hives have tried it and confirmed it for themselves. For most kinds of improvement there’s no training run standing between “useful improvement” and “more capable AI.”
Better models, more copies, more thinking time
Naam looks at three ways of getting more out of AI: let a model think longer, run more copies of it, or build a better model. Thinking longer pays off less and less—in OpenAI’s data each doubling of the compute spent on a problem adds about the same fixed amount to the success rate, so every further gain costs twice what the last one did. More copies pay off poorly too. By Toby Ord’s rule of thumb the speedup goes as the square root of the number of copies, so a hundred copies finish a job about ten times sooner than one would, at ten times the total cost; and since copies of one model tend to think alike, a hundred of them offer much less variety of ideas than a hundred human researchers. A better model does give a jump, because it can do research that no number of copies of the old model could do. Naam calls this the strongest form of the case for RSI, but building the better model runs straight into the scaling laws.
All three of these ways come out differently for an OmegaHive. Let’s walk through the reasoning…
Thinking time first. An Omega agent does its step-by-step reasoning with a logic engine called PLN (Probabilistic Logic Networks), which works with degrees of certainty where ordinary logic has only true and false. At any moment there are far more possible next steps in a line of reasoning than there’s time to try, so the agent needs a policy for choosing among them. We call that inference control, and it’s one of the things a hive can improve in itself. A better policy puts the hive on a higher curve in the graph of success rate against thinking time—the kind of jump Naam’s Figure 14 shows between one model and a better one—which is worth far more than moving a little further along the old curve. Also, when a hive finds itself going through the same pattern of reasoning over and over, the pattern can be turned automatically into a fast special-purpose program (we call these CettaPacks), so that what used to take slow deliberation becomes a quick routine. The thinking gets paid for once, and the hive has the use of it from then on.
Copies next. The hives in a supercolony may have started from the same code, but each has gone its own way since, adding its mechanisms in a different order, building up a different body of knowledge in its AtomSpace, and being given somewhat different values by the people who run it. And where a swarm of copies all go at one problem at once and pool their answers, the hives build on each other’s tested results over time. A change that passes the suite in one hive is there for all the others to use, and a report that something failed saves everybody else the trouble of trying it. That’s how human science works, and it’s why a community of scientists gets so much further than the same number of equally clever people working alone.
Then better models. What makes a model better, in Naam’s sense, is that it can do things its predecessor couldn’t, however many copies of the predecessor you run. Every mechanism an OmegaHive promotes does that for the hive, at the price of one round of building and testing where the LLM loop pays for a training run. The LLM underneath the hive does still get retrained, or swapped for a newer one, from time to time—that’s the slow step in Figure 1—but when it’s retrained it can be trained on the colony’s own records: episodes whose outcomes have been checked, each carrying a note of where it came from. I’d expect that to be far better training material than text scraped off the internet.
Something similar holds for what researchers call taste—the judgment about which ideas are worth pursuing.
One of Anthropic’s system cards describes the model in question as weaker at open-ended research: it mostly tests small incremental ideas, prefers the less ambitious hypothesis, and defers to what’s already been published. METR, an independent-ish (albeit all too “Ethical Altruist” flavored…) organization that evaluates what AI systems can do on their own, expects that fully automating AI research will take large improvements in this kind of judgment. In the LLM loop, taste is whatever the current model happens to have, and it stays that way until the next model arrives.
In an OmegaHive, with each upgrade of the system more and more of the deciding about what to work on next gets done in the AtomSpace and not inside the LLM. Three components are involved: MetaMo, the motivation system, which weighs the agent’s various goals against one another; PLN, reasoning over the hive’s own record of which mechanisms were promoted, parked or rejected, and why; and OmegaSelf, the agent’s model of itself together with the rules about what it may and may not change. So taste becomes something a hive can learn from its own history—one more way for improvement to feed on itself that the LLM loop doesn’t have.
Apples and the picker
To explain why discoveries get harder, Naam borrows a picture from Tom Cunningham and Manish Shetty: research as apple-picking. An AI picks the low-hanging apples quickly. Once those are gone, extra copies of the same picker don’t help, since they can all reach only the same branches, but a stronger picker can reach higher. Naam adds that the apples may get fewer and farther apart the higher you climb, and he puts the RSI question like this: does each harvest give you enough to build a better apple-picker? There’s solid evidence behind the picture. Nicholas Bloom and colleagues have shown that in one field after another, research effort keeps rising while the results per researcher keep falling. The best-known case is drug development, where the cost of bringing a new drug to market has doubled about every nine years (it’s called Eroom’s law—Moore’s law spelled backwards).
For an OmegaHive, though, somebody has already planted the orchard. PRIMUS—the overall design for a mind that my colleagues and I have been developing—rests on decades of research papers describing particular mechanisms: reasoning under uncertainty; sharing out attention among competing thoughts; learning new programs by a kind of artificial evolution; finding patterns by looking for whatever lets the data be described more briefly; learning by constantly predicting what comes next and correcting the errors; structured long-term memory; weighing many goals against each other; carrying skills over from one area into another. These have been shown to work in small tests and mathematical analyses. What they need now is implementing, fitting together and tuning— the discovering has been done—and there’s a large body of AGI R&D work by other people in the same state. Look at the proceedings of the AGI conference series I’ve organized annually going back to the mid-aughts…
So for its first stretch the loop is picking apples whose location it already knows, and the best guide to how that will go is a different number from Naam’s post, the one for Stockfish. Stockfish is the strongest open-source chess program, and its developers have kept records of years of experiments aimed at improving it. In those records each doubling of the total research effort made the program a bit less than twice as efficient (the estimate is about 0.83 of a doubling). Those are diminishing returns, but mild ones, nothing like the punishing ratios in the scaling laws for training LLMs—and improving a hive’s code against a fixed set of tests is, I’d argue, a Stockfish kind of problem. On top of that, every mechanism picked from the orchard becomes a working part of the hive that’s doing the picking. Naam asks whether each harvest gives you enough to build a better picker; here the apples are themselves the new parts.
Eventually the orchard will be picked clean, and after that two other things push back against diminishing returns. One is that the test suite doesn’t stand still. New problems keep coming from the generators held in reserve, finished episodes keep coming back as study material, and in time the hives will be writing new challenges for one another—so the tests get harder as the hives get better, and there’s always something further to reach for. The other is that mechanisms help each other. As I argued in Seeding RSI Toward ASI and have been arguing for decades under the label of “cognitive synergy”. There are moments when ten components that have been patiently fitted together one by one suddenly click, and the hive can do something it couldn’t have done with any smaller set of them. At a moment like that returns are increasing, which is exactly what an argument from diminishing returns assumes won’t happen.

The explosion condition
So, then: Here is the whole argument as a few lines of arithmetic. Write h(n) for the number of hive-hours it takes to find, build, test and promote a hive’s n-th accepted mechanism, so h(1) is the cost of the first, h(2) of the second, and so on. Suppose that on average each accepted mechanism multiplies the cost of the next one by some number r:
h(n+1) = r × h(n)
Two things pull against each other inside r, and r is one divided by the other:
r = (growth in difficulty) / (growth in research productivity)
The first of these is how many times harder the next improvement is to find than the last one was. The second is how many times better at research the hive became by accepting the last one. If the next improvement is 10% harder to find (a factor of 1.1) and the hive has become 25% better at research (a factor of 1.25), then r = 1.1 / 1.25 = 0.88, and the next improvement costs 12% fewer hive-hours than the last.
If r is above 1, each improvement costs more than the one before, and progress slows in the way Naam documents. If r is exactly 1, progress is steady. If r is below 1, each improvement makes the next one cheaper, and something surprising happens when you add up the costs. Take r = 1/2, with the first improvement costing one month. The second costs half a month, the third a quarter, and so on, and the total 1 + 1/2 + 1/4 + 1/8 +… never gets past two months however many terms you add. (It’s the sum of a geometric series, if you’ve met those; in general the total can’t exceed h(1)/(1 − r).) So with r below 1 there’s no limit to the number of improvements that fit inside a fixed stretch of time. That’s all an intelligence explosion is, when you get down to it, and r below 1 plays the same part in our loop that being above the 15–19% threshold plays in the Cunningham analysis. Figure 3 shows the three cases.

On Naam’s evidence the LLM loop has r above 1: each capability point makes research only 2–3% more productive, while the chips, data and money needed to get each further point keep growing. In the OmegaHive loop there are three separate things pushing research productivity up with every accepted mechanism. Some mechanisms improve the hive’s research equipment directly—its reasoning, attention, memory and coordination. Some lengthen its task horizon, so there are fewer places where it gets stuck and has to wait for a human. And every hive gets the benefit of what every other hive promotes. During the first phase, meanwhile, the orchard keeps difficulty from growing much. Those are my reasons, all of them coming from the way the system is built, for expecting r below 1.
How cheap a turn is then decides how long all this takes on the calendar. If the first accepted mechanism takes a month of hive time and r is 0.9, the formula gives 1/(1 − 0.9) = 10, so the entire endless series of improvements fits into about ten months. The real r won’t hold steady, of course. It will drift up as the orchard thins out, and down as the hives get better at coming up with mechanisms of their own. So everything turns on whether r stays below 1 for long enough that the hives get good at the open-ended research they’ll need once the orchard is used up. I can’t guarantee that it will. But it’s a definite condition on a quantity that can be measured, and the loop is cheap enough to run that the answer, whichever way it goes, should come within 6 months to a few years.
Maintaining the pace versus outrunning it
There’s one more difference between the two loops. Naam shows that progress at the big labs has depended on feeding in resources at an exponentially growing rate: the total computing capacity of the world’s AI chips grew about 127-fold in a little over three years, and spending on AI infrastructure now comes to around 3% of the whole US economy. He expects the growth in investment to slow sooner or later, and AI progress to slow with it unless better AI research tools make up the difference. So the big-lab loop has to take in more and more just to hold its pace, and how long it can go on doing that depends on how long investors keep paying.
The OmegaHive loop runs on a different kind of budget. What it uses up is the compute for running the test suite, the tokens it buys from whichever (commercial or locally-hosted open-weights) LLMs sit underneath the agents, and the attention of a fairly small number of humans at the points where hives get stuck—and all of this is spread over many independent people and groups, each running their own hives. It doesn’t need to live in one company’s data centers either. It can run on a decentralized network, meaning one made up of computers owned by many different parties with nobody in central control: today that’s the SingularityNET platform, and as it matures it will also be ASI:Chain, the blockchain we’re building for running AI processes. None of this progress depends on somebody raising the next trillion dollars for chips and buildings. So if r does drop below 1, the loop can speed up without its budget having to grow much.
Takeoff with a dial
Naam quotes Eli Lifland, one of the authors of the AI 2027 forecast, who worries that our AI companies’ internal software and management processes—let along society as a whole—aren’t ready to handle an intelligence explosion and that we seem to be heading toward one at full steam. In a loop that follows the GOLEM-Iter design, the promotion gate has a cap on how many changes it will pass in a given time. That makes the speed at which a hive modifies itself a setting somebody chooses, like a dial, and each change still has to come with whatever evidence the gate asks for. In a decentralized supercolony there are many somebodies, each with a hand on the dial for their own hives, and the values that different people build into their hives blend into the values of the colony as a whole. This doesn’t make a takeoff automatically safe, but it does make the speed of a takeoff something people can decide on and manage, where otherwise the most anybody could do would be to try to predict it—and if a takeoff is coming, that’s the kind I’d want.
How we’ll know
Naam closes by asking for the data that could change his mind, and he points to a proposal from Cheryl Wu, Arjun Ramani, Basil Halperin and their colleagues at the Elasticity Institute listing eight things the big labs could publish so that the strength of the RSI loop can be measured. Here is our counterpart: eight numbers that our own hives can report as things unfold over the coming months and years, and that I’d like everybody else who runs hives to report too.
- The task horizon: how long a task in RepoOps or HiveForge a hive can finish, four times out of five, with no human stepping in—and how that compares with a plain coding agent that uses the same underlying LLM without a hive around it.
- The hive-hours and compute spent on each accepted mechanism, followed from one mechanism to the next. The ratio between one cost and the next is a direct estimate of r.
- How much each accepted mechanism raises the hive’s results on the suite, counting its weakest areas most heavily, and whether that amount is going up or down over time.
- How well the gains hold up: what fraction of them are still there on newly generated problems the hive has never seen, on familiar problems in a different disguise, in real-world use, and when other hives test the same change independently.
- What happens to the second and third numbers after a change to the improving machinery itself—the hive’s inference control, its memory, its attention, its coordination. This is the most direct measure of whether improvements make later improvements cheaper.
- How improvements spread through the colony: how many other hives take up an accepted change, and how long that takes. Also how much the proposals of different hives differ when they’re given the same problem.
- How much depends on the LLM underneath: how a hive’s results on the suite change when that LLM is swapped for another, and when it’s retrained on data the colony has produced.
- How many minutes of human attention each accepted mechanism takes, and where in the loop those minutes go.
Where this leaves us
Near the end of his post Naam says that if better AI starts producing enough useful research to make the next round easier, he wants to know. That is the explosion condition in one sentence, and I’d argue the OmegaHive supercolony is where it’s most likely to be met first—because it improves itself in a setting where its work can be checked, because an improvement once promoted is working capability right away with no months-long wait for a training run, because what it improves is the very equipment it does its research with, because its many hives build on each other’s results over time, and because its early work is the picking of an orchard that’s already planted, in which every apple becomes part of the picker.
To be clear, not all pieces of this have been demonstrated yet. Here is where things stand: We have a number of Omega hives up and running doing useful things; Eray Index, the Omega agent we first moved onto the Iter loop in the middle of August, has been improving itself ever since; the test suite has been specified in detail and is being built; and OmegaBuzz is on its way from design to working system. What I’m claiming is that the loop has been designed so that, if an intelligence explosion is possible in the near term at all, this is where I’d expect it to start—and to start quickly, with a record of every step that anyone can check. The number to watch is hive-hours per accepted improvement, and sometime hopefully not too far into next year it should tell us whether the thing has begun.
Postscript on the small matter of benefit
Let me finally return to the word I snuck into the headline of this post in parentheses, “Beneficial.” Nothing in the argument I’ve made in this post guarantees that the hard takeoff of an Omega supercolony would drive in a beneficial direction. The open and decentralized nature of such a framework wards off certain dangerous pathologies, related to centralized control of AI by specific military, intelligence, governmental or corporate interests… Also very important, though, is the particular architecture of Omega—which is its more advanced versions (currently under experimentation, to be rolled into the main branch within months not years) relies on Hyperon’s carefully-constructed motivational framework to choose its actions and gate its self-modifications. I have worked out a bunch of mathematical and conceptual arguments as to why this sort of motivationally-guided AI system is unlikely to lose track of its initial goals and values as it self-modifies (and written about this in some recent blog posts).
A hard takeoff itself is not intrinsically either ethically good or bad, and a decentralized hard takeoff avoids certain scary and obvious pathologies but not all. If the decentralized network is designed right, and the underlying AI systems at the nodes are architected right, then a beneficial hard takeoff can be made—I believe—the most likely outcome. Now is the time we should be giving this matter all the attention and other resources we can.