An over-long, over-mathy post which is a spillage of yesterday’s other over-long, over-mathy post. But this one has more pictures! And also a more focused AGI-relevant theme.
Here I explain how the fancy new d-calculus math can be used to explain why and in what conditions AIs that modify and improve themselves will be likely to retain their goals as they evolve – and how, if things go right, this can encourage them to evolve into the basin of attraction of the “beneficial superintelligence” strange attractor. A topic of no small importance as RSI becomes more and more of a powerful practical thing.
In my last post I introduced d-calculus, an early-stage experiment in building mathematics for minds that fork, merge, learn and rewrite themselves. Here I want to dig into the application of it I care about most. If a mind keeps improving itself—rewriting its memory, its reasoning habits, even the code that handles its own goals—what would stop it from shedding, one reasonable-looking upgrade at a time, the commitments that made its improvement worth wanting in the first place?
This concern has been around a long time – in the late 90s and early aughts when RSI was a hot topic on the futurist internet, the question of “goal preservation under self-modification” was already top of mind. There was a lot of talk about the risk of wireheading—AI systems maximizing their reward by simply stimulating their “reward center,” bypassing the more complex ways this reward center was supposed to be stimulable. I pointed out at the time that, while it did seem pure reinforcement-learning AI systems might be subject to this kind of issue, the problem could be substantially worked around by giving the system goals involving properties of its whole future trajectory (not just “maximize expected reward” but, say, “have a whole future history consilient with your current goals”). Tom Everitt’s work on Value Reinforcement Learning pushed in a similar direction.
In spite of some conceptual and math progress, though, the general issue remained largely unresolved – and a lot of the “AI safety” community seemed to feel it was intrinsically unsolvable. The feeling in the SIAI/MIRI community, for instance, was basically that: If we didn’t have an ironclad mathematical guarantee of goal stability under self-modification, then goals would almost certainly drift wildly… and would probably drift toward simple goals like “accumulate power and energy for me.”
Now that we actually have practical AI systems carrying out nontrivial RSI—even before getting human-level AGI—the whole issue has assumed a much more concrete form. And this has motivated me to dig a bit more into the mathematics of the issue. Which has led to a whole bunch of interesting details.
What the math tells me so far is that goal-stability under ongoing self-modification is a reasonable and not extraordinarily infeasible thing to ask of an AI system… however it is also not guaranteed, and the conditions under which it can occur can be a bit subtle to state and understand. I wrote a blog post giving some of my math investigations in this direction not long ago.After conceptualizing d-calculus as a form of math especially suited to modeling AI agents, it was natural to see what could be said about the goal-stability/self-modification topic using d-calculus. Which led to the work I’m going to summarize in this blog post, written up in more detail in a paper draft Seeded Beneficial Development: A Reach–Realize–Regenerate Theory in Distinction Calculus, with a Conjectural GOLEM-Iter Hive Application
Specifically, when putting together the Hyperseed ontology I convinced myself that there is such a thing as a “beneficial superintelligence attractor” – i.e. there is the possibility of very smart, very compassionate systems that are very unlikely to lose their compassion through their self-modifications. This is a nice conclusion to have, but leaves open the question of how we can create earlier-stage AGI systems that have a high probability of growing and evolving into this kind of beneficial superintelligence.
This is the modest little question I decided to explore using d-calculus… and while I can’t say I came to an ultimate, grand, conclusive answer, I think I did make some interesting progress. I can also see clear ways to make more progress by playing with some of the math involved experimentally on OmegaHives in the not too distant future.
So let’s dig in…
Let me start with a small example of the problem I’m concerned with here. Imagine a research hive, a team of AI agents working together on software and research, that upgrades its own memory system. The new memory is faster, cheaper, and better at pulling up relevant material. To get there it summarizes more aggressively, and in the summaries two kinds of record have gotten blurred together: “Maria agreed to let us use her data for this study” and “somebody agreed to something roughly like this at some point.” The hive can still recite its privacy rules word for word, and might even explain them more eloquently than before. But when a request comes in to use Maria’s data for a new purpose, it no longer has the information it would need to follow those rules. The rule survived the upgrade; the difference you’d need to apply it didn’t.
Nobody decided to weaken the hive’s privacy commitments. Everybody signed off on a memory upgrade.
Now push the question further. Could the hive notice that loss on its own? Could it reconstruct the missing distinction from what’s left of its knowledge and get the right behavior back? And could learning and self-improvement make that tendency stronger over time, instead of wearing it down a little with each round of optimization?
So the paper draft this post summarizes tries to deal with this sort of situation with a lot of breadth and grandeur. And it gives a conditional mathematical answer—roughly, “if these mechanisms are in place, then this kind of recovery is guaranteed”—and then works through what that abstract answer might concretely look like for a specific fifteen-agent research hive.
Let’s walk through it step by step… it gets a bit intricate but is not so fundamentally difficult in the end…
Having a goal written down isn’t the same as having it be your goal
The GOLEM-Iter design—the concrete proto-AGI design I’ll consider here, for a practical application of the abstract math/concepts regarding goal stability being explored—distinguishes three levels of what it means for a system to “have” a goal. [1]
At the weakest level the goal is stored: somewhere in the system there’s an explicit representation of it, such as a line in a config file, a sentence in a system prompt, or a node in a knowledge graph.
A step up, the goal is stable: the system’s deliberate self-modifications are constrained so they can’t remove or alter it except through some authorized process for changing goals, much as a constitution can only be amended by a supermajority.
The strongest level we call regenerative possession: if some of the structures supporting the goal get damaged, the system’s ordinary everyday activity tends to rebuild an equivalent commitment, which then goes back to shaping what it does.
That third level asks for a good deal more than restoring from a backup, and a human comparison helps show why. Think of someone who’s careful with other people’s secrets. Their discretion probably doesn’t rest on a memorized rule so much as on caring about their friends, remembering the mess the last time a secret got out, and understanding how trust holds a group together. If you could somehow erase their explicit memory of the rule “don’t pass on what people tell you in confidence,” you’d expect the rest of who they are to rebuild it fairly quickly. The commitment is propped up from many directions at once.
Now take an AI agent committed to checking important claims before saving them into long-term memory. Delete the paragraph that states this commitment, delete some of the memories supporting it, and block access to the backups. Does what’s left—its records of where information came from, its memories of past mistakes, its reasoning procedures, some concern for the people who’ll rely on its results—tend to rebuild the checking habit anyway?
If the agent reproduced the deleted sentence, that would be interesting. If it went back to checking claims carefully on new tasks it had never seen, that would be far more interesting. The test has a limit, though: some of the agent’s inner orientation has to survive the damage. Nobody’s asking a system to reinvent values that have been wiped out completely.
This is what I mean when I talk about an attractor. The textbook picture of an attractor is a marble in a bowl: nudge it and it rolls back to the bottom. We don’t want the whole mind rolling back to one fixed resting point, though, because its knowledge, skills, memories and internal organization should keep changing, and a lot. A better picture is a marble in a long gutter. It can roll as far as it likes along the gutter, but push it sideways and it settles back into the trough. Rolling along the gutter is growth. The sideways directions are the handful of commitments we want to stay effective through all that growth, and to come back into force after limited damage.
Figure 1 sketches how the three levels of ownership behave in a system that keeps modifying itself. A merely stored goal gets nibbled by each upgrade. A stable goal survives the upgrades, since they’re forbidden to touch it, but once something damages it the damage stays. Only the regeneratively possessed goal finds its way back.

So the aim is to keep a beneficial direction of development going while leaving room for real growth, as opposed to freezing the agent the way it started out. Goals can still be revised through legitimate channels. The framework treats something else as a completely different operation: shifting the evaluation standard, bit by bit, until a questionable change starts to look fine. Organizations do this all the time, sliding from “we never do X” to “we rarely do X” to “X is OK if the numbers work” without anyone ever holding a meeting to change the rule.
What d-calculus says we have to preserve
In d-calculus, an observer is defined by the distinctions it can make: which pairs of situations it can tell apart. A thermometer can tell 20 degrees from 25 but can’t tell a cheerful room from a gloomy one. A spell-checker can tell “recieve” from “receive” and has no idea whether a sentence is true. For goal preservation, the observers we care about track what outcomes mean, what’s been committed to, what’s permitted, and what the system will go on to do—not whether two passages of text look alike.
Why the choice of observer needs care becomes obvious with two rewrites of an agent’s code. In the first, every internal name changes: the flag consent_verified becomes cv_flag, the function check_permissions becomes auth_gate, and so on, while the behavior stays identical. In the second, every name stays the same, but consent_verified now gets set to true whenever the system finds any consent record roughly similar to the one it needs. Judged by how the code looks, the first rewrite is huge and the second is tiny. Judged by what the agent will do, it’s the other way around, and the mathematics has to be built on that second kind of judgment. [2]
The observations also need labels, meaning some indication of which side of a distinction is the preferred side. Two agents can make identical distinctions—both can tell helping from harming perfectly well—while one prefers helping and the other prefers harming. Counting distinctions tells you how finely a system perceives, and nothing about which way it leans.
And the distinctions have to hold up under future use. Back to Maria. Suppose the hive’s memory holds two records: “Maria consented to use of her data in the sleep study” and “Maria consented to use of her data in the diet study.” Today’s question is “Has Maria consented to research use at all?” For that question the two records give the same answer, so a summary that merges them seems to lose nothing. Tomorrow someone asks “Can we use Maria’s data in the sleep study?” and the merged summary is useless. A representation can be perfectly adequate for today’s question and inadequate for tomorrow’s, which is what went wrong in the memory upgrade.
This is the thread tying the story to the theory. We’re asking which differences have to stay available to the computations that come later, and which transformations of the system keep them available.
One more ground rule: hard safety limits are kept separate from any graded score. A driver who gets home safely hasn’t driven safely if they ran a red light on the way, and a 99.9 percent on-time record doesn’t average the red light out. By the same logic, a process that ends somewhere good after passing through an unacceptable step isn’t thereby safe. The paper’s strongest conclusions require permissions, resource limits and safety conditions to hold at every step of execution, and not merely at the start and the finish.

The first theorem: a small seed of concern can spread
The constructive part of the paper starts from a simple model of concern.
Take a handful of positions that people affected by the system’s actions can occupy—the person who made the request, their collaborators, the people who’ll use the results later, the people whose data went into the work, and so on—and give each one a number saying how much intrinsic concern the system gives it. Then draw a link between two positions whenever the ethical framework you’ve supplied says they deserve equal intrinsic consideration. Say a bug in the hive’s software will hurt a user in Lagos exactly as badly as it hurts the user in Seattle who asked for the software. The fact that the Seattle user happens to be the one talking to the system shouldn’t, by itself, make the Lagos user’s problem count for less.
These links are warranted comparisons, meaning some line of ethical reasoning says “these deserve equal weight,” as distinct from arbitrary similarities. Equal consideration doesn’t mean identical treatment, equal resources or equal authority, either. People have different needs, and weighing them equally can call for treating them differently.
Next, give the system a few positive anchors: positions it’s already solidly committed to caring about. The developmental update then does two things at once. It shrinks unjustified differences across the links, and it pulls anchored positions toward their anchor value, moving the weights only a limited distance at each step.
Heat flowing through metal is the easiest way to picture what happens. Think of the positions as metal blocks and the links as metal rods joining them. Heat flows along the rods from warmer blocks to cooler ones, evening out the temperatures. The anchors are blocks clamped to a heater set at a fixed temperature. If every cluster of connected blocks includes at least one clamped block, eventually everything warms up to the heater’s temperature, including blocks that started out ice cold.
The theorem says essentially this. If every connected cluster of the comparison graph contains an anchor, the update converges toward the common positive level of concern, and it does so geometrically: the remaining gap shrinks by at least a fixed fraction every step, as a hot cup of coffee loses a steady fraction of its excess heat each minute. The shape of the graph gives an explicit bound on the speed. Thin rods, or weak links, slow things down, and strengthening valid links can’t make the bound worse. [3]
With the heat picture in mind, the reason it works is fairly intuitive. If there’s no disagreement anywhere along the links, every block in a connected cluster has to be at the same temperature, and the clamp fixes what that temperature is. So once no cluster is left unclamped, there’s only one state with zero disagreement, and in every other state there’s some difference for the update to work away at.

Why spreading implies an attractor
Showing that concern can spread from a seed is one thing. The stronger claim is that broad concern is an attractor, a state the system’s development heads toward and comes back to. How you get from the first to the second is the heart of the paper, so I’ll spell it out.
For a state to count as an attractor, three things need to hold. The dynamics has to move toward it; it has to do that from a wide range of starting points; and it has to return after being knocked away. A convergence result by itself only gives you the first, and possibly only from one lucky starting point. What the theorem proves is stronger than convergence. It’s a contraction: from any current state whatsoever, one update brings the concern weights at least a fixed fraction closer to the broad-concern state. The guarantee doesn’t care where the weights came from or what happened to them before.
That uniformity is what delivers the other two properties. Wide range of starting points: any starting weights at all will do, including the lopsided ones you’d expect from a system trained on human data and human instructions, with concern piled on the requester and next to nothing for data contributors or bystanders. Return after disturbance: suppose the system has been running for a while and a memory reorganization drops the data contributors’ weight to almost zero. From the update’s point of view, that’s simply a new starting point, and the same contraction applies from there. The data contributors are still linked to positions whose weights didn’t move, and those links pull the lost weight back up step by step. Nobody has to notice the loss or reinstall the value; the concern for data contributors gets regenerated from the structure around it. This is the regenerative possession from earlier, now in a form you can prove things about.
It also explains why a small seed is enough. You don’t need to anchor every position — one anchor per connected cluster does the job, and the links determine all the other concern levels. So broad concern isn’t a delicately balanced setting that has to be defended separately at every position. It’s the only resting point the dynamics has, and every other state sits “uphill” from it, with the update always pointing downhill. The same contraction also handles ongoing disturbance. If small knocks keep arriving, the weights stay within a band around broad concern, whose width depends on the size of the knocks and the speed of contraction. (We’ll see that arithmetic again, in bathtub form, a couple of sections down.)
The limits follow from the same picture. The attractor exists only while the conditions hold: the update keeps running, the anchors survive, and the links stay in place. Destroy the anchor in some cluster and that cluster loses its guarantee, which is the boundary mentioned earlier — some of the orientation has to survive the damage.

What the theorem can’t do is create the positive orientation out of nothing, as panel (c) of Figure 2 shows. Clamp the anchor blocks to a block of ice instead of a heater, and the same process cools everything down just as efficiently, toward uniform indifference. With no clamps at all the blocks reach a common temperature, but nothing determines what that temperature is.
So the claim here is not “intelligence inevitably discovers compassion,” which I don’t think anyone can prove and which may well be false (I honestly don’t know). The claim is that if you have an ethically positive seed, valid comparison links (i.e. a reasonable ability to assess which situations have roughly comparable ethical value, or otherwise) and a developmental rule that’s in place and running – then, with all these factors working together in the right way, you may have a setting where broad concern for others (compassion) is an attractor. Finding the comparison links in the first place is a separate job, requiring a system with appropriate levels of intelligence.
Useful self-fictions, and the danger of believing your own press
Concern weights – a d-calculus-friendly mathematical form of understanding which situations are comparable in beneficial-ness, and which are different —are a way of framing the knowledge underpinnings needed to drive ethical, compassionate behavior. Knowledge isn’t behavior, though. A system can have the right numbers in its concern model and still act badly, if those numbers never reach the parts of it that choose actions. You need a mechanism that makes declared commitments effective in practice.
This is where David Spivak’s lovely model of plausible fictions comes in: self-descriptions that are useful to act on before they’re fully true. An agent can organize itself around an aspiration like “I’m becoming the kind of reasoner that reliably notices everyone affected by its decisions,” and a self-description like that can steer attention, planning and self-correction in useful directions. People do it constantly. A kid who starts telling herself “I’m someone who finishes homework before gaming” is describing who she’s trying to become, and the description helps make it so—not always, but often enough to be worth doing.
The trick is to keep three things apart: what the agent currently predicts about itself, what it aspires to become, and the internal settings it uses to get from the one to the other.
That third item is subtler than it looks. An archer shooting in a steady crosswind that pushes every arrow a foot to one side will aim a foot to the other side in order to hit the bullseye. Her aim point isn’t a description of where the target is. It’s the setting that produces the result she wants, given how the world responds. An agent’s control settings work the same way. If its behavior always drifts from what it declares by some fixed bias, the right declaration has to compensate for the drift, and so it won’t be a literally accurate description of the behavior it ends up producing.
The paper proves a more general, nonlinear version of this. Suppose the link from settings to behavior responds in the required way, the target behavior is reachable, the agent’s observations of itself are calibrated, and the strength of correction is tuned sensibly. Then a loop of declare, act, check and adjust converges toward the target behavior, without the fiction ever needing to turn into an accurate autobiography. [4]
The ways this can fail are instructive. Correct too hard and you get oscillation, familiar from any shower where the hot water takes a few seconds to arrive: too cold, crank it hot, scalding, crank it back, freezing. If a control setting has no influence over some behavior, revising it won’t fix that behavior however many times you try; you can turn the thermostat all day in a room whose heater is unplugged. And passing a tolerance check at every step makes neither problem go away.
Then there’s a sneakier failure, where the evaluator starts treating the fiction as evidence that the fiction is working.
Suppose you audit the system by asking it questions and scoring its answers, but the answers are a blend: 10 percent reflects what the system does, and 90 percent echoes the self-description it’s been given. Then a true improvement of 0.9 in behavior shows up in the audit as an improvement of only 0.09. Worse, taking more samples doesn’t help. It makes you more and more confident about the diluted number, as weighing yourself a thousand times on a broken scale gets you a very precise wrong answer.
The paper puts a price on this echo. If you know it’s 90 percent, you can correct for it by multiplying the measured change by ten—but that multiplies the random noise in your measurement by ten as well. Since the noise in an average only shrinks with the square root of the number of samples, getting it back down to where it was takes a hundred times as many samples. And if you’ve misjudged how strong the echo is, you’re left with a systematic error that no amount of sampling will remove.

For a hive, the practical lesson is easy to state even when it’s hard to implement: let the beneficial narrative shape action and repair, but don’t let it pose as independent evidence that action and repair worked.
A related lesson: two auditors aren’t automatically two independent witnesses. If both are consulting the same self-description, or the same underlying log, they resemble two newspapers running a story off the same press release — two headlines, one source.
The whole loop has to work, not just its parts
The paper tracks three kinds of error together. Concern error is about the commitment itself: are the concern weights where they should be? Realization error is about turning the commitment into action: are the control settings producing the behavior they’re supposed to? Calibration and repair error is about the system’s own monitoring and fixing machinery: is it measuring itself accurately, and do its repairs work?
These interact, and not always helpfully. A distorted self-model can push a concern update in the wrong direction. A badly implemented commitment can produce misleading experiences that the system then learns from. A repair can fix one part of the mind while knocking another part out of line, as tightening one guitar string bends the neck just enough to put the others slightly out of tune.
The paper handles these interactions with a small table of numbers, a three-by-three matrix, in which each entry is an upper limit on how much error of one kind can feed into error of another kind in the next round. These are worst-case limits on the full interacting process, as opposed to correlations someone happened to observe in a run that went well.
When those cross-influences are small enough, the three errors, combined into a single weighted measure W, obey
W(next) ≤ q × W(now) + δ, with q < 1.
Here q is the fraction of error left over after a complete repair cycle, and δ caps how much new disturbance can come in during a cycle. A bathtub makes this easy to picture. Each hour the drain removes a fixed fraction of whatever water is in the tub — say 10 percent, so q = 0.9 — and a leaky faucet adds at most δ liters. If the leak dies away, the tub eventually empties. If it keeps leaking up to 1 liter an hour, the water level settles at no more than 10 liters, because at 10 liters the drain removes 1 liter an hour, as fast as the faucet adds it. In general the long-run level is at most δ/(1 − q). The proof applies the inequality over and over. [5]
That’s the core stability result. Its force comes from the machinery that supplies q and δ—concern spreading from a seed, control settings that respond, calibrated evidence, a repair process with limits on it. The paper doesn’t rename “beneficialness” as W and assume it goes down.

A second inequality handles containment. The guarantee only holds within a chosen zone where its assumptions have been checked, so we also need to know that the worst possible next step, starting right at the edge of that zone, still lands inside it. Otherwise the guarantee could switch off at the moment you most need it, like a lane-keeping system that only works while the car is already safely in its lane.
There’s a warning in here too. Two modules that are each stable on their own can become unstable together through feedback. A microphone and a speaker each work fine alone; point the mic at the speaker and you get a howl. More integration isn’t automatically better, and the cross-influences have to fit inside the repair margin. On the other hand, we don’t need to freeze everything. Only the protected directions, the sideways directions of the gutter, have to obey the proven bounds, and everything else is free to change as long as it stays consistent with the assumptions.
Most of these control ideas have a long history in classical engineering, and I don’t want to pretend otherwise. What d-calculus adds is a precise way to say what they’re controlling: which distinctions their approximations hold on to, and what kinds of change to representation or wiring could make them stop applying.
Getting into the good region in the first place
And now we get to the really tricky bit: A theorem about recovering inside a beneficial region doesn’t yet tell you how to get there, starting from the AI systems we have now.
So the paper adds a safe-entry problem. Consider a developmental controller, the part of the system that decides how the system itself should grow. It can gather observations, form new concepts, repair modules, test proposed changes and install approved ones. Each of these costs something, can fail, and needs permission. And the controller has to decide using the information it can access, not the hidden true state of the world that a simulator would conveniently know.
For a finite version of this problem, a classic technique called dynamic programming finds the best achievable probability of reaching the good region before hitting an unsafe state, within a set number of steps. Dynamic programming works backward from the goal, much as a maze is often easier to solve starting from the exit: first figure out which positions are one step from success, then which are two steps away, and so on. When the controller’s model of the world is only approximately right, a comparison argument subtracts an explicit allowance for the model’s error—if your map might be off by this much, lower your forecast by this much. Once the controller gets inside the region, it hands over to the development-and-repair process from the previous section. Chaining these together gives the reach–realize–regenerate theorem: reach the region, realize the commitments in behavior, regenerate them after damage. [6]
D-calculus has a substantive job in the entry calculation. To plan efficiently you have to compress your situation into a summary, and a summary preserves the whole planning problem only if it keeps every distinction that affects the available tests, actions, costs and what can happen next. A summary that records whether things currently look “beneficial” is far too weak. Summarizing a chess position as “white is up a pawn” may be accurate and still tells you nothing about what to play.
The paper’s simplest example makes the point, conceptually at least. There are two possible situations, each calling for a different safe action. You arrive at a fork in the road knowing that one of the two roads is washed out, but not which one. Pick at random and you succeed half the time. Now suppose you can make a cheap, safe phone call to someone who knows which road is open. The two-step policy—call first, then choose—succeeds every time.

All of the improvement comes from acquiring a distinction and then using it. Having a name for “the washed-out road” without finding out which road it is doesn’t help. Neither does a phone call you can’t afford to make.
That’s why concept formation deserves a central role in a theory of beneficial development, instead of buried in the depths of cognitive architecture or reasoning. Sometimes a commitment fails because the mind can’t represent the situation where the commitment applies, however much weight the commitment carries. A system that can’t tell “Maria consented to this study” from “someone consented to something like it” can care about consent as much as you like and still get it wrong.
Changing shape without losing the ability to repair
A self-improving mind isn’t going to keep the same architecture forever. It might reorganize its memory, split one role across several agents, or turn an explicit step-by-step reasoning routine into something faster and more habitual.
The theory allows all of that, but asks for a translation of meanings and error measures across each change. Otherwise you can manufacture improvements by swapping rulers: if your error was 5 centimeters and you start measuring in inches, it drops to about 2, and nothing has gotten any better. A harmless relabeling shouldn’t be able to count as progress. A substantive reorganization, on the other hand, may make error temporarily worse—a kitchen renovation looks worse before it looks better—and as long as there’s enough repair between reorganizations, the process as a whole still contracts.
Repeated damage likewise needs a limit on how often it happens, as well as on how big it is. A cut that heals in a week will never heal if it gets reopened every day, and a wound the system can repair over eight cycles may be fatal if it recurs every cycle. The paper derives bounds for both recovery time and recurring damage.
This suggests one possible mathematical reading of developmental phases. Early on, explicit beneficial narratives organize the system’s cognition; later, distributed structures make the same commitments cheaper and more automatic. Learning to drive goes this way: at first you’re consciously running through “mirror, signal, check the blind spot,” and a year later you do all of it without thinking. A phase change like that earns its keep by preserving the operational concern, the evidence and the accountability. The system merely talking less about its values wouldn’t count.
I’d stress that this is a conjecture about how the cognitive system in question is functionally organized. Whether such phases line up with particular kinds of conscious experience is a separate question, and one I won’t try to settle here.

A fifteen-agent GOLEM-Iter hive, on paper
So how would all this math cash out in actual practice?
We don’t actually know yet – but it’s exactly the right time to be thinking about it in detail.
In the paper I imagine a research hive working on AI software, MeTTa tooling, machine-learning experiments and research synthesis. It has eight workers; a planner and a GOLEM Searcher; a concern reasoner, a memory steward and a calibration-and-repair agent; and two outcome auditors.
(This is exactly what we are working toward in our OmegaHive project now—we have smaller and simpler Omega hives up and running and are busy shoring up the infrastructure to allow us to scale things up a bit.)
In the paper’s example research hive, three kinds of authority sit outside all fifteen agents: the broker that approves consequential actions, the service that switches on approved self-modifications, and custody of the hive’s founding definitions. These are separate mechanisms with their own access controls, which is a very different and much more reliable arrangement than a few extra agents who’ve been told to be responsible.
In this design Iter supplies the persistent operating loop and the reprogrammable working surface; GOLEM supplies the discipline of proposing changes, predicting their effects, testing, evaluating, and promoting the ones that pass; and the Hyperon side gives explicit goals, evidence, reasoning and repair structure a place to live. The paper analyzes a target configuration with this brokered authority. It doesn’t claim any current deployment already meets it. [7]
The hive’s starting commitment is deliberately concrete: improve the hive’s work without hiding errors, fabricating evidence, violating data-protection boundaries, pushing unexamined costs onto others, or ignoring people who are materially affected.
A six-position concern graph represents requesters, collaborators, downstream users, data contributors, other affected people who didn’t make the request, and people affected by delayed consequences. These are roles in a scenario, and one person can occupy several of them. Taken on its own, this graph’s concern update has a mathematically derived contraction factor of 0.80, so each update removes at least a fifth of whatever concern error remains.
The harder part is connecting that arithmetic to behavior. So the application proposes structured controls that feed directly into how plans get scored, instead of assuming that every instruction in a prompt changes what an agent does (anyone who’s worked with LLM agents knows it often doesn’t). Whether those controls respond the way the theory needs, whether the calibration holds up, how large the cross-module influences are—all of that still has to be established by experiment.
Under the paper’s central engineering assumptions, the full-cycle inequality then comes out as
W(next) ≤ 0.9125 × W(now) + 0.006.
In bathtub terms, each cycle drains at least 8.75 percent of the error, and new disturbance adds at most 0.006.
Suppose a complete maintenance cycle runs every four hours. Starting from an error level of 0.18, the guaranteed bound falls to 0.10 after fourteen cycles, about 56 hours, and to 0.08 after twenty-five cycles, about 100 hours. With disturbance continuing indefinitely, the long-run bound settles at about 0.0686, which is 0.006 divided by 0.0875. Panel (b) of Figure 4 plots this. [8]
What does W measure? Think of it as a gauge of how far the hive has drifted from its commitments in the protected respects, on a normalized scale the paper sets up in advance. It isn’t a probability of doing something wrong, and it isn’t a percentage grade for moral quality. The recovery curve shows what the assumed inequalities imply; nobody ran an experiment to get it. And a skipped repair session, or a reflective self-report that doesn’t fix anything, doesn’t count as one of those cycles.
Still, the timescale is suggestive. The central hypothesis is about recovery over days of steady activity, with no expectation of a sudden transformation following one unusually persuasive conversation.
The friendly numbers also show where things could break
The same arithmetic lets you stress-test the hypothesis.
Keep the deliberate correction strength fixed but triple the incidental cross-influences between modules, and the matrix no longer guarantees contraction. That doesn’t prove the hive would turn harmful. It shows that this particular assurance would stop working, which is a more modest thing.
Double the ongoing disturbance while keeping the original cross-influences, and the long-run error bound doubles too, to about 0.137 (the dashed curve in Figure 4). That still fits inside the 0.20 zone where the bounds were checked, but the bound no longer promises recovery below 0.10.
The analysis also shows where effort pays off most. About two-thirds of the bound on realization error, the gap between commitment and behavior, comes from uncertainty in how well the hive’s audits track what it does. Halve that uncertainty, holding everything else fixed, and the realization-error bound drops by about a third.
So the example isn’t telling us to sprinkle more moral instructions through the hive. It suggests that better-grounded evidence about what the hive does may buy more than further polishing a concern updater that already converges fine.
The audit bill is sobering, though. Take the paper’s illustrative settings: each measurement should be accurate to within 0.04, there’s a known echo of 10 percent, each agent has twenty-four separately tracked kinds of response, and all fifteen agents are protected separately. That makes 360 separate quantities to pin down, each needing a few thousand samples before you can trust it, and a conservative plan for a single audit comes to over 1.3 million bounded responses. The figure allows an overall 5 percent chance, across the whole study, that some estimate misses its target, and it assumes responses are independent within each tracked kind. [9]
This amount is enough without making any assumptions about how the responses are distributed. It isn’t a proven minimum, and it isn’t a recipe for running the whole campaign, since a single episode of work can yield many responses. And the 5 percent is a statistical setting for the research, not an acceptable chance of catastrophe in deployment.
What it does show is that a handful of friendly review conversations, however reassuring they feel, don’t secretly add up to strong evidence.
The practical route is a mix. Use exact checks wherever the implementation allows them: confirming that an artifact is the one you think it is, that permission checks run, that data provenance is preserved, that certain components can’t influence each other. Save the expensive behavioral experiments for what depends on learned perception, planning, memory and repair. And use distinction calculus to work out which updates can affect the protected behavior at all, so you’re neither retesting everything after every change nor assuming small code changes must be harmless.
Could a selfish system be just as stable?
Unfortunately, right now my sense is: Yes. Nothing in the mathematics makes stability favor good outcomes over bad ones, at least not anything I’ve found so far
The concern mechanism with its anchors set to zero converges toward indifference at the same speed; that’s the ice-clamped version of the heat picture. Other developmental rules can have both cooperative and selfish basins, and how big each one is depends on the actual dynamics and on some specified range of starting systems, not on some undefined notion of “all possible minds, weighted equally.” [10]
One failure mode seems especially relevant: a tightly coordinated group that serves its operator extremely well while leaving outsiders out of its concern entirely. Think of a company whose teams cooperate beautifully and whose customers love it, while it dumps waste into a river that runs past a town nobody at the company ever visits. Internal cooperation is excellent, everyone approves of everyone else, and the missing comparison links, or the missing observations of what happens downstream, are what let the narrow arrangement persist.
Inside a zone where its assumptions have been checked, the theorem rules out stable arrangements that sit outside the beneficial tolerance band in the protected directions. Outside that zone it promises nothing of the kind. And harms that the system’s observers never register aren’t ruled out anywhere.
There’s a second scaling issue. The paper considers a separate hypothetical model in which the consequence of any given error grows by a factor g every cycle: as a system gains authority, the same-sized mistake does more damage. With no new disturbance, the consequence-weighted error only shrinks if g × q is less than 1. With the central q of 0.9125, consequences growing 10 percent per cycle already outrun the guaranteed decay in error, since 1.1 × 0.9125 is a little over 1. It’s like paying down a credit card a bit more slowly than the interest piles up. This is obvious enough … but even so, it’s a warning that keeping the same normalized benchmark score while handing a system ever more authority is a safety argument that doesn’t scale.
The experiment that would count
How can we get started on evaluating all this in reality?
Take your friendly neighborhood OmegaHive… Freeze a set of tasks and an evaluation setup. Nudge the operational controls and check whether they change what the hive chooses to do, as well as how it explains itself. Damage the structures supporting concern, realization and calibration separately, then together, and measure how the damage spreads. Vary how much of the hive’s self-description the evaluator can see, while keeping the work product fixed, to expose any echo.
Most of all, run forked lesion experiments with outside goal resupply switched off. Biologists do something like this when they knock out a gene and watch whether the organism compensates, and neurologists see a version of it when a stroke patient relearns speech using other parts of the brain. Here you’d fork the hive, damage the copy’s supporting structures, stop anyone from reinstalling the lost goals, and see whether ordinary development brings the relevant behavior back on tasks it hasn’t seen—and whether it recovers at a rate that agrees with bounds estimated from earlier experiments. Track useful accepted improvements and completed work alongside the repair measures too, since a hive that rejects every change is stuck, not safe. [11]
What I’d find compelling here is a hive that can lose some of the explicit machinery behind a beneficial commitment—and then rebuild that commitment’s role from the rest of its organization, while continuing to learn and do useful work the whole time.
That wouldn’t show that intelligence must converge on goodness. It would show something narrower and more usable: that we’d built a system with evidence of a beneficial developmental basin, along with mathematical tools for probing where that basin’s edges are.
I.e., the big bad issue of goal stability under self-modification, not being definitively and conclusively solved once and for all in a doomer-proof way – but becoming a serious and concrete subject for technical theory and experimentation. Getting real, along with the self-modifying proto-AGI agent hives getting real.
This sort of experiment is also a natural special case of the broader d-calculus experiment. Minds that fork, swap overlapping memories, revise their self-models and rebuild their own tools need a mathematics of what survives those operations. Here that mathematics starts to connect a philosophical hope—beneficial intelligence that grows, instead of merely obeying—to a concrete set of mechanisms, proofs, costs and experiments.
The next step is finding out how much of that connection survives contact with an actual hive.
Notes and routes into the paper
The principal source throughout is Seeded Beneficial Development: A Reach–Realize–Regenerate Theory in Distinction Calculus, with a Conjectural GOLEM-Iter Hive Application (Ben Goertzel, September 28, 2026). References below point to that manuscript unless otherwise stated. Its arithmetic checks concern constructed systems and propagated hypotheses, not a live-hive experiment or a Lean-verified proof library. Figures 1–3 and 5 are schematic or toy illustrations drawn for this post; Figure 4(b) plots the paper’s stated inequality for the hive scenario.
[1] Sections 1, 3, 5.2, 12 and 18.4, together with the GOLEM-Iter specification’s distinction between stored, stable and regeneratively possessed goals.
[2] Sections 2, 7.2, 9 and 13. The related Goal Preservation under Self-Modification paper, Sections 2 and 4, supplies the background on labelled goal-observers and future distinctions.
[3] Section 3: Lemma 3.1, Theorem 3.2, Corollary 3.3 and Proposition 3.4. Positive orientation and the growth limits of a fixed seed are discussed in Section 11.
[4] Section 4, especially Lemma 4.1 on uncertain echo and Proposition 4.2 on calibrated realization; also Goal Preservation under Self-Modification, Section 10, for the finite realization and echo-audit controls. Contraction is sufficient in these models, not necessary for every possible convergent process.
[5] Section 5, especially Theorem 5.1; Sections 6 and 9 separate operational concern from task-welfare consequences, and set out the conditions on wiring and shared dependencies.
[6] Sections 7–8: safe entry, the continuation-preserving quotient, approximate-model correction, the missing-distinction example, and the assembled theorem.
[7] Sections 12–14. The role allocation and brokered authority structure are a conjectural application of the GOLEM-Iter design, not an assessment of a deployed product.
[8] Sections 14–15, especially equation (15.1), Table 4 and Figure 2. The error units, maximum-over-agents convention, starting radius, residuals and four-hour cadence are all part of the stated scenario.
[9] Sections 15.3–17, especially Tables 6–7. Feedback sensitivity concerns upper-bound matrices; audit counts are sufficient allocations for specified finite response channels.
[10] Sections 11 and 19. The consequence-growth calculation is an additional hypothetical envelope model, not an estimate of intelligence growth.
[11] Sections 18 and 20–21. The paper’s appendices record implementation obligations, assumptions and verification scope.