I’ve been around long enough now to watch a whole bunch of words and concepts make the transition from very obscure to very widely used—including a few, like AGI, that I put out into the world myself. It’s exciting to watch one’s own obscure term get picked up by a lot of people. But the people who pick it up tend to change its meaning, in ways that often less technically or conceptually grounded than the original.
I’m thinking about this now because in media interviews I keep getting asked to clarify three terms in particular: AGI (artificial general intelligence), ASI (artificial superintelligence), and RSI (recursive self-improvement)—the last of which is really the main thing I want to talk about here, but you’ll have to be patient or exercise your scrolling finger because it comes at the end! Being asked over and over has made me realize how confused these words have become out in the world, even though the ideas behind them were very clear when they were first introduced.
Of course, language as a whole is a complex, evolving, historical phenomenon. Words don’t have meanings the way mathematical objects do; they have histories, and the meaning at any given moment is a cross-section of the history. So a decent accounting of what these terms mean has to involve at least a rough accounting of where they came from and how they got to where they are—and where they seem to be going.
(For AGI and RSI I can give the history from personal memory, since I was there. For superintelligence I have been around for the recent history of the term, but the origins are older and not too dramatic…)
AGI
If you’ve followed me for a while you’ve heard me tell this story before, so feel free to skip ahead a few paragraphs!
Anyway, I more or less launched the term in 2005, in an edited volume of papers titled Artificial General Intelligence that Springer published, with my long-time collaborator Cassio Pennachin as co-editor. The working title had been Real AI, but I always knew that was a stand-in, because the narrow AI I was building at the time for corporate and government customers—to pay the bills and to gratify the part of my brain that enjoyed building relatively quick simple things that just worked—was real too. It was just narrow.
Ray Kurzweil, in The Age of Spiritual Machines, had contrasted narrow AI with “strong AI,” but strong AI already had a different meaning in cognitive science and philosophy, where it referred to Searle’s hypothesis that machines could be conscious in the same sense that people are. I happen to think that hypothesis is probably true, but it’s a different argument to have, and we didn’t want to conflate the two.
So a bunch of us who had contributed chapters on what we would now call AGI—what it is and how to build it—tossed around alternatives for the title on an email chain. As I recall, my friend Pei Wang, creator of the NARS non-axiomatic reasoning system, suggested “general artificial intelligence.” I thought that was good, but rightly or wrongly I thought GAI wouldn’t catch on as an acronym in the US. (I totally love gay-ness and gay culture, and in fact my mom is gay and has been out since my childhood… at some point our band will release our epic jam THE SINGULARITY IS QUEER… but it just didn’t seem like an acronym that was going to catch on in this particular context... people get confused with overloaded meanings…)
So we permuted GAI to AGI. Shane Legg, who had worked for me previously and later went on to co-found DeepMind, was pushing for AGI, and so was Peter Voss, and I believe both Shane and Peter have both said they were the ones who introduced “AGI” into the thread. I don’t remember exactly (maybe they both did), and it doesn’t much matter. The G was meant to echo the g factor in psychology, the general intelligence that IQ tests are supposed to measure.
We later found out that a guy named Mark Gubrud had used the term “AGI” in 1997, in an article on nanotechnology and its future, with basically the same meaning we had. It just hadn’t caught on. What put it out there, I think, was the book plus the conference series. I started the AGI conferences with a workshop in 2006 and a proper conference with proceedings and academic publications in 2008, and every year we were spraying emails across every AI researcher we could find, saying come to our AGI conference, give a paper, help us build this new field. Shane’s position at DeepMind, and DeepMind’s periodic use of the term in its communications, certainly helped too.
Origins story aside, there was a mathematical definition in that book, put forth by Marcus Hutter and refined a bit by Shane, who was Marcus’s PhD student. Roughly: an AI system is generally intelligent to the extent that it can achieve any computable reward function in any computable environment, where you average over all possible environments— e.g. using something like the Solomonoff prior, which weights more complex environments lower than simpler ones.
You can argue with that definition. I’ve proposed alternatives that take into account the amount of resources used, so that getting something done cheaply and efficiently counts as smarter than getting it done expensively: how well, on average, normalized by compute effort or energy, can you achieve a computable reward function in a given computable environment. Whether resource usage belongs in the definition of general intelligence is quite literally a matter of definition; you can do it either way.
There are many other variations. Later my good friend Weaver (David Weinbaum) proposed that we shouldn’t model intelligent systems in terms of optimizing reward functions at all, but as what he called open-ended intelligence: complex self-organizing systems whose main properties are individuating, meaning maintaining their boundaries as they evolve, and self-transcending, meaning modifying themselves into forms different from what they’ve been before. You can then connect this back to the reward-optimization definitions. To be a complex self-organizing system that sticks around while achieving a lot of complex goals in complex environments, you need to individuate and self-transcend; to individuate and self-transcend across a variety of complex environments, you need to be able to achieve a variety of goals in a variety of situations. They all connect.
One thing all these mathematical definitions have in common is that human-level intelligence is an arbitrary waypoint. Human-level intelligence is like human-level running speed. You define speed in physics terms, how long it takes to get from point A to point B, and then human running speed is just how fast we happen to be able to run. There’s nothing fundamental about the precise speed at which humans happen to be able to run. In the early days of the AGI field we always thought about human-level AGI that way, as the arbitrary level at which humans happen to be able to think.
Defining human-level AGI is then a less fundamental and less interesting question—being able to solve a variety of problems relevant to human life about as well as the average human, say… at which point you have to ask on a good day or a bad day, under what environmental restrictions, with how much time, and so on. You can think of human-likeness as a bias on the set of reward functions and environments you care about; and even as a restriction on self-transcendence, since once you self-transcend your brain and body far enough you’re arguably not human anymore.
But human-likeness was never the main point. The point was the ability to generalize beyond your programming and training and take a big flying leap into the unknown. For a real system to optimize such a huge variety of reward functions, or to open-endedly self-modify and survive in such a huge variety of situations, it has to be able to take wild leaps into the unknown over and over again. That ability is the crux of all the mathematical definitions.
What we see now that AGI has become commercial is that people are specifically interested in achieving human-level AGI. Some will define AGI as being able to do 95% of jobs as well as the median human, or something along those lines, and a sort of convertibility has crept in between “AGI” and “human-level AGI.” That’s natural; it’s how human language evolves, words start out one way and shift to another.
But it’s still important to distinguish the base concept of AGI from the concept of matching human performance, and what current LLMs make vivid is that the alignment between these two things can slip. LLMs are good at achieving human-level performance on many things, but mostly not by generalization—mostly by a modest generalization capability applied to a massive amount of data. So the human-likeness aspect of “human-level AGI” is running ahead of the actual AGI aspect, in a way that makes the shifting meanings of the word confusing.

Superintelligence
Superintelligence is not a term I put out there originally, and until recently I hadn’t looked into where it came from. I have largely considered it as Nick Bostrom’s word, as his 2012 book was what launched it onto the world stage, though it was always clear the word had been around far longer than Nick.
The Oxford English Dictionary traces “superintelligence” back to around 1822, and “superintelligent” to 1845, in the ordinary sense of an extraordinarily high degree of intelligence, long before anyone was thinking about machines. It floated around at a low level in English for the next century and a half, and science fiction picked it up in the mid-twentieth century for its assorted mutants, aliens and computers, without anyone particularly owning it.
The idea of a machine smarter than us, as opposed to the word, enters the technical literature with Alan Turing. In a 1951 BBC talk, “Intelligent Machinery, A Heretical Theory,” he observed that if a machine can think, it might well think more intelligently than we do, and asked where that would leave us.
But the foundational paper is I.J. Good’s 1965 “Speculations Concerning the First Ultraintelligent Machine,” which I’ll come back to in the RSI section, because it’s really the origin of both concepts at once. Good, who had worked with Turing at Bletchley Park, did not say “superintelligence.” He said “ultraintelligent machine,” defined as a machine that can far surpass all the intellectual activities of any man however clever. That definition—exceeding the best humans across the board, not just at some tasks—is essentially what Bostrom’s superintelligence would later mean, and for a couple of decades “ultraintelligent” was the term of art, to the extent there was one.
Vernor Vinge’s 1993 essay “The Coming Technological Singularity,” which put “singularity” into circulation in its modern sense, mostly talked about “superhuman intelligence” and “greater-than-human intelligence,” but he did use “superintelligence” too, distinguishing “weak superintelligence” (a human-equivalent mind run much faster) from “strong superintelligence” (a mind qualitatively beyond us). Around the same time, on the Extropians list and related corners of the early transhumanist internet, “superintelligence” and “SI” were common currency, generally in Good’s sense.
What Bostrom did was give the term a crisp definition and a home in the academic philosophy literature. In his 1998 paper “How Long Before Superintelligence?” he defined a superintelligence as an intellect that is much smarter than the best human brains in practically every field, including scientific creativity, general wisdom and social skills -- explicitly Good’s definition with the word swapped. In the 2014 book Superintelligence: Paths, Dangers, Strategies he tightened it to any intellect that greatly exceeds the cognitive performance of humans in virtually all domains of interest, and distinguished speed superintelligence, collective superintelligence, and quality superintelligence. The “virtually all domains” clause is doing the work, and it’s there precisely because we have long had machines that beat us in narrow domains, from arithmetic on up.
The acronym ASI, with its tidy ladder ANI / AGI / ASI, seems to have been popularized more by explainers than by researchers. Tim Urban’s widely read 2015 Wait But Why series on AI used exactly that three-rung ladder, and it has been the default framing in the popular press ever since.
So the book is what made the term famous, and the book was mostly about Nick’s fear at the time that superintelligence was going to kill everyone. I like Nick a lot personally; we’re not super close, but we’ve been friends for a long time, and we worked together on some things back in the World Transhumanist Association, which later became Humanity+. I still found that book profoundly annoying.
Borrowing arguments mostly from Eliezer Yudkowsky, it seemed to argue that AGIs with much greater general intelligence than people—superintelligences—could possibly kill us all, because there is no intrinsic reason they have to care about us. That much is quite true. But then it tried to morph that into an argument that AGIs, once they get to a vastly superhuman level, will almost certainly kill us all, and that’s a big difference. Probability not zero and probability near one are not the same thing, and the book elided between them with some well-executed rhetorical tricks that I did not much appreciate. The same tricks are still in circulation, for instance in Eliezer and Nate Soares’s book If Anyone Builds It, Everyone Dies. The argument runs: there’s no reason an intelligent mind needs to respect human values or care about us at all, so we can’t assume it will; therefore it won’t.
But we are not building random minds. We are building specific minds for specific purposes—mind children, if you like—and if we’re clever enough we can teach them our values, build them to love us and help us. Yes, there’s a risk that as they modify and grow they reverse course and decide they don’t like us after all, and there’s no reason to think that risk is zero. There’s also no reason to think it’s high. If you raise a human child with compassion, do beneficial things together, inculcate your value system into them, they could still grow up and turn against you. It happens. It’s just not the likely outcome.
I’ve recently gone through a bunch of mathematics on the preservation of AI goal systems under self-modification—how you make it likely that an advanced AGI rewriting its own code doesn’t rewrite its goals so that, say, it now wants to kill the thing it loved before the rewrite. The upshot is that you can’t solve this problem with a hundred percent guarantee, but you can solve it in a reasonable way that biases the odds strongly against a system borking its goals and rewriting itself into a state its previous version would have hated.
If you’re enough of a doomer you’ll say we can’t totally guarantee these things won’t kill us, so we shouldn’t build them. But humanity has never done anything to that standard. We couldn’t guarantee the move from the Stone Age to civilization wasn’t going to kill us all, or the Industrial Revolution, or that cell phones weren’t irradiating our brains. We have moved forward on the basis of our own attempts at reasonable understanding, not hard and fast guarantees.
That is where the word came from and how it got famous. And now the word is in the legal and political system. Bernie Sanders and Greg Casar have introduced the Ban Artificial Superintelligence Act, which would permanently prohibit developing or deploying artificial superintelligence, pause advanced AI development until a new federal agency writes rules, and impose up to twenty years in prison on individuals—a penalty explicitly benchmarked against unlawfully developing nuclear weapons—and a “corporate death penalty” on companies. Which is interesting on one level, in that the concept has gotten into mainstream politics. But it raises the question: what the hell is superintelligence, legally?
A calculator is superintelligent at logarithms. It’s far better than me; I’m very slow at logarithms even with pencil and paper. A slide rule was better than me too—I learned them in middle school, and I think I still have my grandfather’s around somewhere; he was a physical chemist in the middle of the last century and used it all the time—and when the calculator came out it was way, way more superintelligent than the slide rule. Trading systems are superintelligent at trading on the markets, which is why they dominate the markets. ChatGPT is superhuman at a variety of tasks by now.
So what is a superintelligence? Something much better than humans at doing all things, presumably, which is Good’s definition and Bostrom’s. But can you put people in jail for working toward that by making the AI do more and more things? Suppose a hundred different companies are each working on making AI better than people at one of a hundred different things, using an architecture such that when you put those hundred things together you get a general intelligence that can learn to do new things. Then the guy who should go to jail is the one who figured out how to put the hundred special things together so that the hundred-and-first thing emerges. It’s all very silly, and it’s the same sliding between meanings we saw with AGI.
Because the meaning that came out of Bostrom’s book was quite specific: an AGI with general intelligence way beyond the human level, where the human level is an arbitrary point and superintelligence is some big multiple of it. That means something. A superintelligence in that sense tosses off Nobel-level discoveries with the effort we put into a long division problem; it solves nuclear fusion or nanotech the way we solve home construction. I agree with Bostrom about what the word should mean, even if I don’t agree with him about what will happen when we build one.
But the word has now become so confused that our President has renamed AI as superintelligence. Speaking at the UN General Assembly on September 22, Donald Trump announced that U.S. government documents would refer to AI as “super intelligence,” on the grounds that “artificial” makes it sound fake and it isn’t fake; super intelligence “sounds much better” and is “much more accurate.” A week later he said he’d sign a “very powerful” executive order to that effect, that Xi Jinping loves it, and that “super” beat out “extreme,” “superior” and “supreme” as the best and simplest word.
Artificial intelligence is not, in fairness, such a great term. Everything is part of nature—computers are part of nature, societal processes are evolutionary processes, the process by which human society created computers and the internet is a natural process of self-organizing evolution—and AIs are not always going to be our artifices or tools. As soon as they surpass human-level general intelligence they probably won’t be, and hopefully it’ll be a richer and different sort of relationship. I always knew “artificial” was flawed. When we picked AGI we were going with what the world was already using, not endorsing the word; we thought about “synthetic intelligence” and figured it wouldn’t catch on the way AGI could, because AI already had.
But if AI now means superintelligence, then what is Bostrom’s superintelligence? Super duper intelligence, I suppose. I did write a paper this year called “SuperDuperPsychism,” extending Susan Schneider’s philosophical theory of superpsychism, so the prefix is already in my vocabulary. I’ve considered ultraintelligence and hyperintelligence, but as I noted somewhere before, artificial ultraintelligence is AUI, which is what my kids say when they skin a knee, and artificial hyperintelligence is AHI, which is a tasty meal. None of them seem right.
That’s part of what led us back to the Greek alphabet for our new seed-AGI system, Omega, which wraps LLMs and Hyperon symbolic reasoning in an agentic loop that tries to self-modify toward AGI and superintelligence—or maybe it’s already superintelligence, if AI is superintelligence. I chose Omega, the end of the Greek alphabet, figuring that super, hyper and ultra were already old hat. The Omega Point was Teilhard de Chardin’s term, in the middle of the last century, for the point at the end of human history when the biological and communication networks around the globe synergize into a sort of cosmic emergent intelligence, which for him meant synergizing with the mind of God. We haven’t yet seen the Omega Point rebranded by Silicon Valley or Washington, D.C. But we might.
So the word superintelligence has been morphed a bit, in an amusing way. I wasn’t expecting the road to run from I.J. Good through Nick Bostrom to Donald Trump. Oh the humanity!

Recursive self-improvement
RSI basically refers to AIs building smarter AIs building smarter AIs, and the idea is old in science fiction. The first serious academic articulation I know of is I.J. Good’s 1965 paper “Speculations Concerning the First Ultraintelligent Machine,” published in Advances in Computers the year before I was born.
Good’s argument was short and has been quoted a million times since. Define an ultraintelligent machine as one that can far surpass all the intellectual activities of any man however clever. Since the design of machines is one of those intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an “intelligence explosion,” and the intelligence of man would be left far behind. So the first ultraintelligent machine is the last invention that man need ever make.
What people usually leave off is the final clause—“provided that the machine is docile enough to tell us how to keep it under control”—which shows that the safety worry and the capability argument were born together in the same sentence.
Good himself, according to James Barrat’s reporting on his unpublished memoirs, came to believe late in life that the ultraintelligent machine would more likely spell the end of us than save us, so the arc from optimism to doom that we associate with the 2010s had already been walked by one of the men who started it.
Good coined “intelligence explosion” but, as far as I can tell, never used “recursive self-improvement.” Nor did Vinge in 1993, though his Singularity essay contains the mechanism in full: superhuman intelligences, however created, will be able to enhance their own minds faster than their human creators could, and when greater-than-human intelligence drives progress, progress will be much more rapid.
The specific phrase “recursive self-improvement” appears to come out of the late-1990s and early-2000s transhumanist internet, and in particular from Eliezer Yudkowsky. His long online documents from that period—“Coding a Transhuman AI” around 1999 to 2000, and “General Intelligence and Seed AI” in 2001, subtitled “Creating Complete Minds Capable of Open-Ended Self-Improvement”—introduced the term “seed AI” for a system designed for self-understanding, self-modification and recursive self-improvement. They also introduced “recursive self-improvement” itself, which he defined as self-improvement that increases the system’s capacity to self-improve, so that each round of improvement makes the next round bigger, as opposed to a linear succession of improvements that each just make the system a bit better at its job. That distinction between improving and improving-at-improving is really the whole content of the word “recursive.”
(I should add a caveat here: Eliezer was clearly the one who launched and popularized the term, but I haven’t done nearly enough digging to be sure there isn’t a “Mark Gubrud of RSI” out there, someone who used the phrase casually in the 80s or early 90s and just never promoted it or had it catch on. Given how AGI’s history went, I’d be a little surprised if there weren’t.)
The phrase was standard vocabulary on the SL4 email list Eliezer ran and on the Extropians email list (from which SL4 originally spun out) as well, by the early 2000s. It got its canonical statement in the 2008 Hanson-Yudkowsky “AI foom” debate on the Overcoming Bias blog, where Eliezer argued that recursive self-improvement—an AI rewriting its own cognitive algorithms, which “closes the loop” between the object level and the metacognitive level—makes a hard takeoff far more likely than a soft one.
Meanwhile, on the academic side, Juergen Schmidhuber’s Goedel machine, first described in a 2003 technical report, gave the concept its strictest formal version: an agent that rewrites any part of its own code once it has proved the rewrite is beneficial according to its own utility function. It’s not the same idea as the seed AI, and the proof requirement makes it more or less impractical, but it’s the cleanest statement of what “self-improvement” could mean in a fully rigorous sense.
So the term carries roughly this lineage: Good gave us the explosion, Vinge gave us the singularity it leads to, Yudkowsky gave us the mechanism’s name and its association with a fast, local, unstoppable takeoff, and Schmidhuber gave us its mathematical idealization. The association with hard takeoff was baked into the word by the people who popularized it, which is worth keeping in mind.
For a long time many of us thought of RSI as the way AGI turns into superintelligence. There was a whole discussion in the late 90s and early 2000s, on Extropy and on SL4 (Shock Level Four, a list Eliezer started), about whether the transition to superintelligence, and to human-level AGI in the first place, was likely to be a hard takeoff or a soft takeoff. In a hard takeoff, smart machine builds smarter machine builds smarter machine hour by hour, with measurable increases in smartness on that timescale. In a soft takeoff the same loop runs with a year or two between measurable advances.
At the time I was advocating what I called a semi-hard takeoff: a period where the self-improvement was real but not that fast, followed by a point—once the AGI fully understood computer science, mathematics and hardware architecture—where it would start inventing whole new theories of AGI mind and reformulating and rewriting itself very rapidly. You could then have a system doubling its IQ in a day, if it came up with a new mathematical theory of intelligence and deployed it across its codebase: a huge improvement occurring very rapidly which enables the next huge improvement to occur very rapidly.
What we’re seeing now is quite interesting, because recursive self-improvement in some meaningful senses is working even though we are clearly not at human-level AGI—though we do have narrow superintelligences in various senses. The example that has interested me most is the Omega systems we’ve been playing with. These systems can look at their own operations, analyze them, come up with new algorithms and data structures to drive their own improvement, implement them, plug them into their own codebase and see what happens. That is a recursive self-improving loop, driven by a combination of LLMs and symbolic memory and symbolic reasoning, in systems that are plainly not yet human-level AGIs.
Right now, at this exact moment, it’s not quite an autonomous loop. We have Omega systems doing this, but we’re guiding them. Sometimes they utterly bork themselves and we have to fix something. Sometimes they need advice on which route to take and we give it. The AIs are doing the software development, and a bunch of the algorithm and data-structure design, and the testing; what the humans in the loop are mostly doing is keeping them from going down silly rabbit holes of various sorts. So it’s recursive self-improvement with a human assist.
You can then look at what’s going on at Anthropic, OpenAI, Google and some of the top Chinese labs and see something similar, on a much slower timescale. Each version of Claude Code is being used by human developers to create the next version of Claude Code. Most of the code in these companies—as at SingularityNET and BGI Labs and in my own projects—is now written by AI with some human direction, so in a sense the AI is coding the next version of the AI, with a whole lot of human assistance.
What makes their loop a soft-takeoff rather than a hard-takeoff version of RSI is, I think, that the LLMs at the core of it don’t do continual learning. Their weight matrices are fixed after a very expensive hyperscaler-server-farm training run, and then you use the frozen-weights model to develop a bunch of new stuff, including the algorithms for the next very expensive, long-winded hyperscaler-server-farm training run. It is human-assisted AI recursive self-improvement, but it’s slow-paced, and the intent—the goal orientation behind the self-improvement—is provided by the humans.
So we have two examples of early-stage, pre-human-level-AGI RSI here. One of them has hard-takeoff potential, I’d say, which is what we’re doing with the Omegas. The other has soft-takeoff potential, which is what the frontier companies are doing with their development process and their series of frozen-weight models.

We could bring these together to an extent. As I’ve posted about recently, among other things I’m working on an approach that takes a big frozen-weights model and unfreezes the weights in its top layers—the “neural-symbolic sandwich” approach. You take a big open-weights model like the ones the Chinese labs are putting out, remove the top-layer portion, and replace it with a top layer that has a different architecture, a neural-symbolic one that can do continual learning. Then the top layers of the transformer can update their connectivity and their weighting in real time as they do inference, and potentially you bring a faster recursive loop into the transformer side of things, a hard-takeoff-capable loop, just as we now have on the symbolic side of Omega.
Right now, with Omega, each pass through the agentic loop uses an LLM and some symbolic AI, and the fast-paced, still human-assisted recursive self-improvement is centered on the symbolic part, since we’re using the same frozen-weights LLMs as everybody else. Once we get our method for putting more flexibly learning caps on top of transformers scaled up beyond the small scale we’re at now, the neural part inside the Omega loop can be self-evolving in a more hard-takeoff-friendly way, more like the symbolic part is now. This is quite different from what the frontier labs are currently doing.
Which brings me to Ramez Naam’s recent thoughts. Ramez and I aren’t close friends, but we’ve met and talked and I think we like and respect each other; he’s written some amazing science fiction, including the Nexus trilogy, about porting an operating system into your brain and doing neural self-modification that way. He has just put out a very nice, long, well-researched post, “Can AI self-improvement overcome diminishing returns?”, with a companion version on Noah Smith’s blog.
His conclusion is that on the best current data, the AI self-improvement loop the frontier labs are running would need to be roughly five to ten times stronger to sustain itself, let alone run away; that AI progress will be incredibly rapid by the standards of nearly any other technology, but that the evidence doesn’t point to a sudden explosion into incomprehensible superintelligence anytime soon. Toby Ord and others have praised the piece for starting from the data on current systems rather than from assumptions, which is fair. And I think Ramez’s conclusion is true of the thing he’s analyzing.
What he’s highlighting, whether or not he’d put it this way, is the drift of meaning of the term RSI, which is analogous to the drift we’ve seen with AGI and superintelligence. Ramez himself notes that people now use “recursive self-improvement” to mean everything from AI boosting the productivity of human researchers to AI bootstrapping itself to incomprehensible intelligence. Good-style, Yudkowsky-style recursive self-improvement was smart machine rewrites itself to make smarter machine rewrites itself to make smarter machine, and the humans aren’t really mentioned. That is where you start thinking about hard takeoff.
Now it’s true that recursive self-improvement is also a much broader concept. The whole evolving ecosystem we live in is recursively self-improving—single-celled to multicellular organisms to fish to amphibians to reptiles to mammals to humans to super-AIs is a process of recursive self-improvement—and the Anthropic and OpenAI teams working with their coding tools to make smarter versions of their coding tools is a process of systemic recursive self-improvement too. It makes sense to call it that.
But then to look at the RSI in the frontier labs now, with their frozen-weights models, and say RSI is slow and it’s going to take decades—sure, if you assume that frozen-weights backprop-trained transformers are what’s going to get us to the singularity, and you extrapolate the pace of improvement of these LLMs under the weak, human-in-the-loop recursive self-improvement we have now, then yes, it’s going to take a while for the successors of GPT-6, 7, 8, 9, 10 to get to human level. In fact I think the curve-plotting exercise is a bit specious, because the curve ends up taking a long time to reach human-level AI, and I think it won’t get there anyway, due to fundamental flaws in the underlying architecture: the difficulties with continual learning, and the relatively small role played by abstract knowledge representation inside the transformer.
So it’s a worthwhile thought experiment, and the data-gathering is valuable. But it’s already not the most interesting kind of RSI being played with. Our Omega systems are not yet taking over the world the way GPT has, but I’m running a bunch of them on laptops and servers now, and they’re doing AGI R&D for me and my colleagues with a level of autonomy and insight that we’re not getting from simpler uses of Codex and Claude Code inside agentic harnesses. It is a real thing.
It’s still at an alpha level of development, and I’m hoping, and working so that, within months we’ll have a more robust framework where the level of autonomy of the self-improvement is much greater. I think we see how to improve the underlying infrastructure of Omega and Omega hives and Omega supercolonies (groups of hives) in a way that will essentially remove the need for human intervention except for higher-level intuitive guidance and oversight. I think we’re going to get humans thoroughly out of the plumbing within a few more months of engineering. It’s not done yet, and how critical human oversight and insight will turn out to be at different stages remains to be seen.
But if you try to plot out the likely course of development of Omega systems, you’re looking at a much more rapid curve than the one Ramez sees, because you’re looking at something closer to a good old hard takeoff than a soft one.

Letters and syllables
Going back to the theme of terminology: RSI, like general intelligence, like superintelligence, is a series of letters and syllables that can carry a lot of different meanings, and the history above is a decent guide to how the meanings got layered on.
AGI started as a mathematical notion in which the human level is an arbitrary waypoint, and commercial usage has pulled it toward “matches human job performance.” Superintelligence started with Good as “far surpasses all human intellectual activities,” got a philosophical sharpening and a doom narrative from Bostrom, and is now being used by a senator to mean something you can be jailed for and by a president to mean whatever ChatGPT is. RSI started with Good’s explosion and Yudkowsky’s seed AI, where the humans aren’t in the picture, and is now used for everything from Claude Code helping write the next Claude Code to a system rewriting its own cognitive algorithms hour by hour.
I can see how, if you’re not deep in all this, the shifting usage would be confusing, because the same words are being used to munch together things that have strong family resemblances but quite different practical implications. It doesn’t help that the drift tends to run in one direction: each term gets pulled toward whatever current systems already do, so that the terms come to describe the present rather than point past it. That’s how language works and I’m not going to stop it. But it’s worth knowing which sense of each word you’re using at any given moment, and the most reliable way to know is to know where the word has been.
Anyway—AGI, ASI and RSI… there you go! Thanks to anyone who has sat through me musing on the meanings of these terms. The ideas are very clear in my head and have been for decades; it’s the words that have gotten confused as they’ve pervaded common usage. But this confusion, and ensuing attempts at clarification like this, are clearly how natural intelligence does its thing!
Sources consulted for the term histories
- I.J. Good, “Speculations Concerning the First Ultraintelligent Machine,” Advances in Computers 6 (1965), as quoted in Muller and Bostrom / Ethics of AI and the MIRI Intelligence Explosion FAQ
- Turing’s 1951 “Intelligent Machinery, A Heretical Theory” and the Good-to-Bostrom lineage, per the AI Magazine review of Bostrom’s Superintelligence
- Oxford English Dictionary entry for superintelligence (earliest use c. 1822) and Merriam-Webster on superintelligent (1845)
- Vernor Vinge, “The Coming Technological Singularity” (1993); weak/strong superintelligence usage per Bostrom and Yudkowsky, “The Ethics of Artificial Intelligence” (2011)
- Nick Bostrom, “How Long Before Superintelligence?” (1998), cited in the MIRI FAQ; Superintelligence: Paths, Dangers, Strategies (2014) definition per this review
- Tim Urban, “The AI Revolution” (Wait But Why, 2015) for the ANI/AGI/ASI ladder
- Yudkowsky’s seed AI and RSI writings, per the LessWrong Seed AI entry, the recursive self-improvement entry, and this LessWrong retrospective; Schmidhuber’s Goedel machine per this 2026 survey
- Sanders press release on the Ban Artificial Superintelligence Act and AP coverage
- Trump’s rename: Axios, Sept 22 and Daily Signal, Sept 29
- Ramez Naam, “Can AI self-improvement overcome diminishing returns?” and the Noahpinion version