This post is rather long and fairly mathy—but it’s also IMNSHO extremely interesting for those with the right sort of taste. If trying to grok a new AI-agents-native species of mathematics isn’t your thing, then skimming just the first few sections may be enough for you. Or feed the whole beautiful mess to your AI agent and ask it to summarize!
Imagine that every time you hit a genuinely hard decision, you could split into three copies of yourself—one to push the conservative strategy, one to push the audacious one, and a third to sit back and interrogate the assumptions the other two were quietly making. Later you’d bring what each of them learned back together, keeping the differences that turned out to matter rather than just averaging them away or taking a vote.
For a human that’s a metaphor—and maybe sometimes a funky partial approximation of how things work in our heads. For an AI agent running on the right kind of substrate, it’s closer to a literal everyday operation. Not cheap, not guaranteed to merge cleanly afterward, but real, in a much more literal sense than it is for us.
This observation—if you are a person who thinks like Ben Goertzel, anyway—leads directly to the question: So what mathematics would feel natural to a mind that actually lives that way?
That’s the question behind a research direction I’ve been developing over the past month or so, with a lot of help from large language models – that I’ve taken to calling d-calculus: the distinction calculus.
The ambition here is to seed a branch—maybe eventually a whole garden of forking branches—of mathematics organized around the situation of learning, self-modifying, forkable, mergeable minds… rather than around the situation of static objects being manipulated by an unchanging outside observer which is more or less the tacit assumption baked into most of the mathematics we’ve inherited.
With various LLM and agentic minions, I have developed some initial bits of d-calculus myself. But honestly, the longer-term experiment here interests me more than my own contribution so far does. What I actually want is to hand hives of AI agents enough definitions, theorems and worked examples that they get the gist of the direction, and then let them push it forward using whatever intuitions come out of their own way of being organized and doing things—which won’t be my intuitions, and shouldn’t be.
Using LLMs plus Lean4 and friends to go after the Millennium Problems and other famous open questions is a fine research programme, and I’m glad people are doing it. But even a series of spectacular wins there wouldn’t use up what these tools make possible. There’s a separate question worth asking: whether a different kind of mathematician ends up caring about a different kind of mathematics—not necessarily different standards of proof (though it could be that, in part)… but different questions, different abstractions, different things that strike you as obvious or worth chasing.
D-calculus is my attempt to crack open one path toward that.
Sameness, for a mind that keeps turning into something else
One route to d-calculus is to start with what is something of truism: a mathematical field tends to become productive once it settles on a useful notion of sameness. Algebra gets isomorphism—different presentations of what’s really the same structure. Topology gets homeomorphism—a continuous, reversible reshaping that leaves a space’s essential structure intact. Whatever transformations a theory is willing to permit end up determining what the theory can see and what it’s blind to. [1]
For learning agents, a handful of less familiar notions of sameness suggest themselves pretty quickly, once you go looking. Could two systems count as equivalent because each can be taught, or distilled, into the other within some resource budget? What survives a round trip through two different representational languages? Does merging A with B and then folding in C give you the same thing as merging B with C first and then folding in A?
An agent might keep all its commitments intact while swapping out almost every symbol it uses to express them. Two forks might start out with identical memories and diverge only in what they’re authorized to actually do. A collective of agents might have a capability that none of its members has individually, while sharing a blind spot that none of them notices either.
These are different relations. Treating them as one undifferentiated question of identity is exactly the kind of confusion a useful calculus of distinctions ought to head off before it starts.
None of this is meant to throw out existing mathematics—probability, logic, category theory, dynamical systems, theoretical computer science, all of it still supplies most of the actual machinery used in d-calculus investigations. What I’m trying to do is reorganize some of that machinery around the distinctions and transformations that matter to these particular minds, and then see what new constructions and results fall out of the reorganization.
Start with a distinction
The philosophical seed here is G. Spencer-Brown’s Laws of Form—take the act of drawing a distinction seriously as a primitive, rather than assuming a world of already-separated objects has to come first, before any drawing of boundaries can happen. My own extension of that intuition is that physical, experiential and computational structure can all be approached through patterns of distinction-making, at bottom. That’s a guiding perspective more than a proof of anything—I’m not claiming the papers have nailed down a complete metaphysics, and I doubt any set of papers could. [2]
I explored a version of this a while back in a paper on distinction graphs and graphtropy, which I’ve found useful for structuring parts of my own AI work since. In a distinction graph, nodes are things or situations, and an edge connects two nodes that some particular observer can’t tell apart. The observer doesn’t have to be a person—it could be a sensor, a program, a reasoning subsystem, or even another graph describing how the observations themselves get made. [2]
D-calculus takes that picture and turns it into a language of operations you can actually do things with.
In the simplest finite version, we write α(x, y) for the degree to which a given observer treats x and y as indistinguishable—a number between zero and one, where one means no distinction at all and zero means the observer sees them as totally separate. The complementary quantity is just d(x, y) = 1 − α(x, y). Worth flagging that a graded similarity like this isn’t automatically a probability— what it means depends on how you’ve set the thing up, and the interpretation has to be supplied from outside the formalism. [1]
Take a software agent with two implementations of some tool, x and y. An observer who only cares about the returned answer might call them identical. An observer who also tracks runtime might distinguish them. An observer checking data-access permissions might distinguish them again, for a completely different reason than the runtime observer did.
So “same” is an incomplete sentence until you finish it: same to which observer, for which purposes, with what resources on hand?
There’s one more choice worth flagging up front. We don’t require indistinguishability to be transitive, at least not to start with. A can look nearly the same as B, and B can look nearly the same as C, while A and C are obviously, glaringly different. This is just the old sorites problem—the heap of sand, one grain at a time—showing up again in model compression, in memory consolidation, in long chains of small self-modifications.
The calculus doesn’t wave that problem away or pretend it isn’t there. It gives the discrepancy its own mathematical object to live in.

What you actually do with the calculus
A small working vocabulary, enough to make the examples further down feel like more than metaphors. [1]
You can combine observers—one that notices tool outputs, say, and another that notices resource use—and the joint observer keeps whatever distinctions either one individually needed. For the simplest yes-or-no version of the kernels this just amounts to multiplying the two indistinguishability values together; the graded theory has its own combination rules, though it’s worth being careful that combining scores this way isn’t the same claim as proving genuine statistical independence between the observers.
You can chain correspondences—relate one representation to a second, and the second to a third—which is a different operation from just setting two tests side by side and comparing. In the additive-distance version of the calculus, errors accumulate as you go along a route, and you can compare competing routes against each other. The gap between translating directly and translating through some intermediary stop is itself a quantity worth studying, not just a nuisance to be minimized away and forgotten.
You can refine or merge. Refinement adds distinctions; coarsening forgets them. When an exact observer groups states into genuine equivalence classes, the quotient you get is just the representation where you keep one label per class instead of every underlying state. A useful abstraction, in this sense, is one that forgets exactly the detail its intended uses can’t exploit anyway.
You can pull an observation backward through a transformation. Say T changes a system somehow. Rather than asking whether x and y look the same right now, ask whether T(x) and T(y) will look the same afterward—the pulled-back observer is just α(T(x), T(y)). It tells you which distinctions the future transformation is going to expose, whether you like it or not. This is a small operation but it turns out to be central to deciding whether a compressed memory, a translated program, or a revised self-model is going to hold up.
You can measure distinction mass. Give situations a probability distribution μ, then average α over two independently drawn situations. In the current d-calculus drafts that average gets the name graphtropy, g, and its complement h = 1 − g gets called logical entropy. An exact observer with two equally probable classes gives g = 1/2; four equally probable classes gives g = 1/4, meaning more distinctions are available to be made. For a general graded observer, g behaves like an average similarity rather than a literal collision probability. The notation for this weighted average is a bracket, ⟨μ | α | μ⟩, and more general d-forms and brackets let you measure changes, costs and interactions without pretending every quantity in sight is secretly a probability.
You can differentiate—in more than one sense. Differentiation in d-calculus starts from a structural idea due to Conor McBride. Take the algebraic formula describing a data type, differentiate it by the ordinary textbook rules, and what comes out describes the same kind of structure with one position knocked out and marked—a one-hole context. A triple of X’s, X³, differentiates to 3X²: three places the hole could be, two surviving entries. A list with a hole in it comes out as two lists, the part before the hole and the part after, which is the “zipper” functional programmers use to edit data in place. The product rule turns into a plain statement about where holes can sit—in a pair, the hole is either in the first component with the second left intact, or in the second with the first intact—and the chain rule says a hole in a nested structure is a hole in the outer layer followed by a hole in the inner piece sitting at that spot. Plug the hole with a value and you get a complete structure back.
In d-calculus this is how you ask where a distinction can surface. Given a one-hole context C, pull the observer back through it: (C*α)(x, y) = α(C[x], C[y]) says how alike x and y look once each has been plugged into that surrounding. Two components are contextually indistinguishable when no permitted context pries them apart, and closing an observer under all its one-hole contexts gives the coarsest representation that is safe to substitute anywhere—merge anything further and some permitted surrounding computation will tell the difference. Several results further down lean on exactly this construction. Requirements travel the other way, from the whole inward to the hole, and pulling one back through nested contexts composes in reverse order: the chain rule, seen from the observer’s side.
Once positions carry weights, the structural derivative turns into a numerical one, and how strongly a distinction measure responds to a single comparison is the total weight of the places where that comparison gets consulted. Differentiate graphtropy with respect to one entry α(x, y) and you get μ(x)μ(y), the chance that a random pair of draws lands on exactly that comparison; differentiate a smoothed shortest-path closure with respect to one edge’s cost and you get the expected number of times that edge gets used, with routes weighted by how short they are.
Time gets its own operator, the time differential LTα = T*α − α—the new kernel minus the old one along a dynamics T, positive where a distinction has blurred and negative where it has sharpened, and averaging it over pairs gives balance laws for what a computation learns and what it forgets along the way. Its product rule carries a cross term, LT(αβ) = (LTα)β + α(LTβ) + (LTα)(LTβ), because finite steps interact in a way infinitesimal ones don’t; Leibniz’s textbook rule is the limit in which that term vanishes. The drafts keep these operators apart—one picks out a position, one measures sensitivity, one tracks change over time—and take some care not to let them blur together. [1]
And you can fork and assemble hives. A fork starts as two copies whose corresponding states are indistinguishable to whichever observer you’ve selected; experience, permissions, or other observations can then pull them apart over time. A hive records each member’s internal observer plus the correspondences running between members—and the observer itself can be part of the system under study, so a hive’s picture of its own organization becomes just another object the mathematics is free to examine and revise.
Last one: closure. Closure asks what further identifications get forced on you once you insist that the identifications you already have fit together coherently. In the metric version this relates to shortest paths. But closure can quietly erase distinctions you actually wanted to keep—“we found a coherent consensus” and “we found a faithful consensus” are not the same achievement, and it’s easy to arrive at the first while believing you’ve achieved the second.
For readers (like me) who like to see the symbols, here is the same working vocabulary the way the papers write it—lightly simplified, since each piece comes with hypotheses the drafts spell out and it wouldn’t make sense to repeat here.

A trivial example theorem: a speed limit on changes of distinction
The first example use of d-calculus I’ll give comes from the foundations paper, and I’m putting it first because the proof is so simple, like one sentence long. It answers a question that any self-modifying system needs to ask about itself: how fast can my picture of the world be changing, compared with how much change I’m registering? [1]
Start with graphtropy g—roughly, the chance that two situations drawn at random look the same to a given observer. It sums up in one number how blurry the observer’s world is—near one when almost everything looks alike to it, near zero when it can tell nearly everything apart. Now let the system evolve under some transformation T, hold the observer fixed, and ask how fast g can change. The natural yardstick is how much change the observer itself registers per step: r, the average distinction d(x, T(x)) between a state and its successor, measured in that observer’s own terms. The paper calls r the observer’s clock rate, since in the reflexive reading it’s the subjective time that passes per step.
The theorem, paraphrased. Fix an observer whose distinction measure d satisfies the triangle inequality—the distance from x to z never exceeds the distance from x to y plus the distance from y to z. Let the system take one step under any transformation T, and let r be that observer’s clock rate for the step. Then the observer’s graphtropy changes by at most twice the clock rate: |Δg| ≤ 2r.
An observer’s world can’t blur or sharpen faster than twice the rate at which, by its own lights, things are happening. Or, as the paper puts it, you can’t gain or lose distinctions faster than twice the rate at which you experience time.
The reasoning is short. How distinct a pair looks can change by no more than however far the first member moves plus however far the second member moves—average that over all pairs and the factor of two falls out. The constant is sharp in the technical sense: there are cases that come arbitrarily close to it, so it can’t be improved in general. The paper’s version is a little stronger than what I’ve stated. It bounds the change in g by twice the optimal-transport distance between the distribution of situations before the step and after it—the least total effort needed to shift one distribution’s probability mass onto the other, with effort measured in the observer’s own metric—and the clock rate is the cost of one particular way of doing the shifting, the one that sends each state to its own successor.
So if the average movement is 0.01 in some fixed observer’s normalized metric, the corresponding aggregate change can’t exceed 0.02. This isn’t a bound on intelligence-per-second, or on the speed of consciousness, or anything nearly that grand—it’s a precise relationship between two quantities measured within the same observational geometry. It does have the same shape as the thermodynamic speed limits physicists have been proving for stochastic processes, which cap how fast a system’s state can change by how much activity is going on inside it, several of them in the same optimal-transport terms; here the observer’s distinction metric stands in for physical distance. And both sides of the inequality are things an agent can monitor about itself: how much it registers changing per step, and how far its overall stock of distinctions has moved.
The triangle inequality is needed here, though. Picture states laid out like the marks on a ruler, 0, 1, 2, …, n, and an observer that can’t tell neighboring marks apart but can separate marks two or more apart. Each step moves every state one mark to the right, with the last mark staying where it is. Every individual move is to a neighbor, so as far as this observer can tell nothing happens at all—its clock reads zero. Yet marks n−2 and n, which it could tell apart, land after one step on n−1 and n, which it can’t. A distinction has been lost with no felt time whatsoever. That’s the sorites problem from philosophy, and the calculus pins down where it lives: every violation of the speed limit is carried by what I’m calling the sorites kernel, the gap between the observer’s nontransitive tolerance and its metric closure.
For a self-modifying agent this cashes out as a real practical question: am I making visible progress, or am I drifting through a chain of individually unremarkable changes that my own current self-description just isn’t tracking? The theorem turns that worry into a bookkeeping rule—whatever change in your distinctions your own clock can’t account for has to be flowing through your sorites kernel, which at least tells you where to look.
Why a hive can’t always have consensus for free
A hive, in the sense of the vocabulary above, is a set of agents, each with its own observer, plus translators running between them. What one usually wants from a hive is a merged view—one body of knowledge that every member’s knowledge maps into without contradiction. This can fail in ways that no comparison of two members at a time reveals, as it can when reconciling three databases or three people’s notes on the same meeting. The blending paper in the initial d-calculus series asks when a lossless merged view exists, and what the cheapest repair costs when it doesn’t. What I’ll walk through here is the smallest example of this sort of failure. [3]
Suppose three agents each distinguish two states, call them 0 and 1. The translator between A and B matches equal labels to equal labels. So does the translator between B and C. But the translator between C and A swaps the labels.
Every pairwise translator, looked at on its own, seems perfectly well-behaved. Go around the whole loop, though, and A’s 0 comes back to A relabeled as A’s 1. Geometers call this kind of thing holonomy: carry an arrow around a closed path on a curved surface, keeping it “parallel” at every step, and it can come home pointing somewhere new. The Penrose triangle is the visual version—each corner of the drawing is a perfectly good corner, and no solid object has all three at once.
The theorem that makes the consequences precise needs one definition. Call a hive coherent when its identifications obey the triangle inequality across members, so that no route through intermediaries identifies things the direct comparison keeps apart.
The theorem, paraphrased. In a coherent hive, carrying a state around any closed loop of translators and back to the member it started from can never identify more than that member identifies on its own. So if some loop does identify two things that its starting member keeps apart, no coherent hive contains all of the loop’s translators together with all of the members’ own distinctions.
Our loop identifies A’s 0 with A’s 1, which A itself keeps apart, so no coherent hive can keep all three translators. To get a coherent combined model you have to give something up—weaken or overrule some correspondence, erase some internal distinction by merging, or hold onto the inconsistency explicitly instead of pretending you’ve achieved coherence—and there’s a quantitative tradeoff on offer between how big the loop’s defect is and how much you have to sacrifice to fix it. [3]
The calculus also prices those options, in the currency it uses everywhere else: logical entropy, the chance that two random observations get told apart. If you insist on preserving every identification and only ever repair by merging, identifications propagate around the loop until 0 and 1 are the same thing everywhere, and the hive’s logical entropy drops from 1/2 to 0. Consensus gets achieved, technically, by forgetting the very feature everyone was supposedly trying to talk about in the first place. Overrule one translator instead—replace the C-to-A label swap with the straight matching the other two translators imply—and you reach a coherent hive in which every agent keeps its distinction between 0 and 1, at the price of the identifications that translator was making. What the calculus won’t do is tell you which translator to overrule. The three choices are symmetric—relabel the states and any one of them looks like any other—and breaking the tie takes evidence from outside the loop, which is more or less what a consensus protocol is for.
The theorem, paraphrased. The logical entropy a hive gives up when it reaches consistency by merging alone is exactly the expected sorites kernel: the total weight of the identifications that chains of translation imply but no direct comparison made. That cost splits into three parts—what each member was already confused about internally, what the loops through its peers force on it, and the gaps between direct and induced translations.
This is the general statement behind the example, and it comes from the d-calculus foundations paper. It shows who is paying for what: in the three-agent loop nobody was confused to begin with, so the whole bill comes from the loop and the translations. [1]
That’s a useful warning about knowledge integration generally, and it’s also a theory of conceptual blending—the cognitive-science term for forming a new concept by merging two existing ones, as with “houseboat”. Blending a house and a boat, or two competing scientific explanations, requires decisions about which counterparts get identified with each other and which distinctions survive the merge, and the theorem shows that those decisions can’t always be avoided. A library of transferred skills runs into a related problem whenever a skill passes through several representations on its way somewhere and comes out changed. [3]
Some of the further d-calculus work I’ve done ties these questions to David Spivak’s polynomial-functor wiring diagrams, which describe how systems get plugged together. D-calculus layers quantitative questions on top of that about what the combined system can actually distinguish, once wired up. Feedback matters here in a way that’s easy to underestimate: an output that looks irrelevant when you inspect it in isolation might be steering some other component entirely, and become observable through that component’s downstream behavior even though it was invisible on its own. [4]
A hive, in other words, isn’t just a bag of agents whose answers get averaged together at the end. Its wiring is part of what it knows.

Teaching the calculus to do useful AI work
As well as trying to flesh out the core math of d-calculus, I’ve been building out a few d-calculus applications as well—partly to sharpen the underlying algorithms and ideas, and partly to hand the agents a real body of worked examples of what it actually looks like to think with this calculus, rather than just read about it. Definitions on their own are not much of an apprenticeship!
The applications range across Hyperon-related inference, concept formation, pattern mining, evolutionary program learning, planning, transfer, neural representations, and the languages and infrastructure connecting all of it. I’ll just list a few of them below, rather than the whole catalogue—if any humans or other agents are interested in the full gamut there is a link to a relevant folder at the end of the post!
For each one I’ll say where the problem comes from, state the theorem in words, and say why I think it is useful.
Two identical beliefs that must not share a memory
Hyperon’s core probabilistic reasoning engine is PLN, Probabilistic Logic Networks, and the version we’re building into the Omega agents is OmegaPLN. A PLN belief is stored together with its evidence, not as a bare probability: how many observations support the statement, how many count against it, and—in OmegaPLN—a record of which observations those were. The standard way of turning counts into a probability is Laplace’s rule: start from a uniform prior, meaning every value between zero and one is equally plausible before any data arrives, and the estimate after the data is (positives + 1) / (total + 2). [5]
Now consider two evidence packets, each holding one positive and one negative observation, though not the same observations. Laplace’s rule gives 2/4 = 1/2 for each. The two full posteriors—the whole distribution over possible probabilities, not only its mean—are identical as well.
Now add a new positive observation that is already in the first packet but not in the second. This situation is common in practice, since the same piece of evidence often reaches a reasoner by more than one route and must not be counted twice. Correct deduplication leaves the first estimate at 1/2; the second moves to 3/5.
The two beliefs were identical as answers to today’s question. They were not identical as states for tomorrow’s reasoning, and treating them as interchangeable because they happened to agree once would be a mistake. [5]
The theorem, paraphrased. Take a finite model with n informative pieces of evidence—tokens—any subset of which an agent might have absorbed. Then any representation of the agent’s state that answers every future merge-and-read query exactly, with duplicates counted once, has to keep all 2n token subsets distinct. A summary that lumps two different subsets together gets some future query wrong.
The reason is simple enough: for any two different sets of absorbed tokens, some future piece of evidence will be a duplicate for one and news for the other, and correct deduplication has to treat those two cases differently. Two to the n sounds like a lot, but it amounts to remembering, token by token, which evidence you’ve already absorbed—n bits of provenance. A single current posterior, however precisely you store it, generally can’t do that job, and there’s no getting around it by just carrying more decimal places. In the vocabulary from earlier: the observer that reads off today’s probability can’t tell the two packets apart, but its pullback through “merge in one more observation” can, and the theorem comes from closing under all such pullbacks.
This isn’t an argument against compression as such. The graded results quantify what you actually lose when you forget particular identities, so storage and accuracy can be traded off on purpose, with eyes open. It’s an argument against silently compressing away exactly the information your own update rules turn out to need later.
The same closure construction appears in a second strand of the work, on programming languages and their logics. Programming-language theory has had, since the late 1960s, a notion of contextual equivalence: two pieces of code count as the same when no surrounding program behaves differently depending on which of them it contains. My TyLAA work—Type–Language Adjunction Archeology—asks a further question: what’s the least operational structure that still preserves the logical and behavioral distinctions you require? The d-calculus reconstruction answers with the context machinery from earlier, and the answer is a theorem. Close the required observations under every permitted one-hole context, and the result is the coarsest sound quotient—the most you can merge without any allowed surrounding computation telling the merged things apart. Merge anything further and some permitted context distinguishes them. OSLF, a method for deriving a logic from a language’s rules of computation, and TyLA supply related language–logic constructions in the same programme. [5]
An agent translating a tool, or simplifying its own internal language, ought to ask not just whether two expressions give the same answer once, but whether some permitted future use could pry them apart. And the coarsest representation isn’t automatically the cheapest to run in practice—sometimes maintaining a running count is genuinely easier than repeatedly recomputing a single yes-or-no property from scratch.
Remembering that the same fact was read twice
This example needs some background, beginning with Taylor series from ordinary calculus. A Taylor series approximates a complicated smooth function near a point by a sequence of simpler ones: first a constant (the function’s value at the point), then a straight-line correction (its slope there), then a correction for curvature, and so on— f(x) ≈ f(0) + f’(0)x + f’‘(0)x²/2 +… This works because when x is small, x² is smaller and x³ smaller again, so you can stop after a few terms and estimate in advance how much you’ve left out. Much of physics and engineering depends on approximations of this kind.
Most of what an AI agent reasons about has no small quantity to expand in. A computer program, a logical formula, a plan—these are functions of discrete inputs: a flag is set or it isn’t, a list is empty or it isn’t, and there are no derivatives to take. Earlier this year I worked out an analogue of the Taylor series for functions like these, called dependency series and dependency towers. Terms are ordered by how many inputs they involve, not by how small they are. The zeroth-order term is the function’s average value, ignoring every input. The first-order terms add what each input contributes by itself. The second-order terms add the corrections that depend on two inputs jointly, the third-order terms those that depend on three at once, and so on. Statisticians will recognize the analysis-of-variance decomposition here, and theoretical computer scientists the Fourier expansion of a Boolean function. The tower version states the same idea in terms of indistinguishability, which already puts it close to d-calculus: the k-th approximation is the one that no test probing at most k inputs together can tell apart from the original.
Two examples. Merging two sorted lists needs nothing beyond second order— the only joint decision is which list’s head is smaller. Picking the smallest of three heads has a third-order part that no combination of pairwise comparisons reproduces. There is also a theorem justifying truncation: in the mean-square sense, stopping at order k gives the best approximation available to any scheme that uses at most k-way dependencies, and the weight of the terms above order k is exactly the error that remains. [6]
So far this is the Taylor idea with “how many inputs interact” in place of “what power of x”. The difficulty appears when approximations are combined, which reasoning does constantly—one estimate multiplied by another, one step followed by another. In ordinary calculus, orders add: multiply a second-order term by a third-order term and you get a fifth-order term, smaller than either, which is why you can truncate each factor separately and still trust the product. For yes-or-no facts this fails. Encode a fact x as +1 or −1 and x² = 1 automatically, so when two terms that read the same fact are multiplied, the fact cancels out of the product. Two third-order terms that read the same three facts multiply to a constant, which is a zeroth-order term and so changes the main estimate. Orders can subtract as well as add, so a term that could safely be dropped from each factor separately may still contribute to the leading part of the product.
Whether this happens depends on a question that doesn’t arise in the Taylor setting. When two parts of a computation each consult “the state of the network”, is that one fact read twice, or two separate facts with the same name? If the reads are fresh — independent draws — nothing cancels and the simple truncation is fine. If they are shared, it isn’t. (Programmers know this as the difference between calling a random function twice and calling it once and reusing the result, which comes up constantly in a nondeterministic language like MeTTa.) This is a question about distinctions: can the reasoner tell “the same thing again” from “another thing like it”? D-calculus handles it by treating the reasoner’s record of its own reads as one more observer. An approximation is then indexed by two things instead of one: how many facts interact, as before, and how much of the sharing among its reads the reasoner keeps track of. A reasoner that tracks none of the sharing computes the fresh-reads answer, one that tracks all of it computes the true answer, and the paper calls the passage between the two the sharing deformation. This second index is what the title Approximation After Taylor refers to.
The theorems, paraphrased. For a finite computation over yes-or-no facts: (1) the true answer equals the fresh-reads answer plus a correction consisting entirely of terms in which some fact is read more than once; (2) truncating the two-index expansion leaves a remainder with an explicit bound, which accounts for the possible cancellations among dropped terms; and (3) when one fact accounts for too much of that remainder, splitting on it—computing the case where it holds and the case where it doesn’t, then recombining—removes its share exactly, and the calculus gives a criterion for when that is cheaper than expanding further.
The paper’s finite retry example shows what’s at stake. An agent keeps retrying a task whose success depends on a hidden but persistent condition—the same condition on every attempt. When the condition is favorable, most attempts end in “try again” rather than failure, and the agent eventually succeeds with probability 3/4; when it’s unfavorable, most attempts fail outright, and the eventual success probability is only 1/12. With even odds on the condition, the true success probability is the average, 5/12. A reasoner that treats each retry as a fresh, independent draw of the condition gets 1/4 instead, because it misses that retries cluster in the favorable world—having to retry is itself evidence that the world is favorable. That’s the x² = 1 effect at work: the way retrying depends on the hidden condition, multiplied by the way success depends on it, collapses into a constant, exactly the piece an approximation that averages the condition away throws out. Both numbers are exact calculations in a specified model, not measurements off some deployed reasoner—the gap comes from provenance, not from floating-point noise. [6]
This feels like an especially natural fit for agents that fork, go consult overlapping memories, and later come back together. They need some way of knowing whether three apparent confirmations are really three separate discoveries, or one discovery that’s traveled through three copies of the same agent and come back looking like corroboration.
Concepts that earn their place
An agent can’t plan over raw situations: there are too many of them, and most of the differences between them are irrelevant to anything it will do. So it groups situations into concepts and plans over those. The question of when such a grouping is safe—when two situations can be treated as one at no cost to the quality of the plans—has been studied in reinforcement learning under names like state abstraction and bisimulation. The work I’ve done with concept formation and d-calculus answers it with the same construction as in the evidence example above: start from the distinctions you need immediately, and close under whatever future operations can expose. The result is a theorem about cognitive control. [7]
The theorem, paraphrased. In a finite system, start with the distinctions your goals, legal actions and immediate costs depend on, then keep adding whatever further distinctions permitted future operations can expose, until nothing new turns up. The resulting concepts support exactly optimal planning over every finite horizon, with no approximation involved, and they’re the coarsest concepts that keep everything the problem declared: two situations get merged precisely when no costed continuation can tell them apart.
A companion result, the horizon tower, shows the order in which distinctions become necessary—looking one step ahead you only need to separate situations with different immediate options and costs, looking two steps ahead you also have to separate situations whose moves land in different groups of that first kind, and so on outward, one round of refinement per step of lookahead. An agent that plans three moves ahead can therefore use coarser concepts than one that plans thirty moves ahead, and the tower specifies which ones. [7]
A concept, on this view, isn’t just a label for a cluster of similar things. It’s an executable interface, one that makes certain kinds of future reasoning possible, at a certain cost you can actually name.
Related ideas show up, in variant form, across pattern mining (where a “surprise budget” can certify that no large discrepancy remains hidden under a given pattern observed in data), evolutionary learning (where omitted interactions between parts of a genome turn into explicit remainder terms in a schema theorem), and neural analysis (where whether an intervention tweaking a neural net matters gets defined in terms of what downstream computations can actually observe). These are separate applications, and they don’t mean that any one d-calculus quantity measures all of intelligence—but they do suggest we have here a novel perspective that can yield penetrating insights into a variety of different AI approaches. [7]
A detour in the Langlands direction
I’ve also tried using d-calculus as a guide on more traditional mathematical territory. The Langlands programme is a web of conjectured—and in some famous cases proved—correspondences between number theory, geometry and the theory of symmetry, which has organized a large part of pure mathematics for half a century. Its typical result shows that two objects built in entirely different ways carry the same arithmetic information. One effort I’ve made in that spirit asks what information survives comparisons between different mathematical realizations of the same structure, and what the familiar summaries throw away before the comparison has even really started. [8]
A little background for readers who aren’t number theorists. Much of the arithmetic side of the Langlands programme runs on congruences—two very different-looking objects, say the coefficient sequences of two modular forms (highly symmetric functions whose coefficients encode arithmetic), that agree modulo some prime p, or modulo p², and so on. Ramanujan’s classic example has the coefficients of one famous modular form agreeing, modulo 691, with a simple sum of powers of divisors. How deep such agreements go, and how they’re shared among several objects at once, was central to Wiles’s proof of Fermat’s Last Theorem. The d-calculus reading is natural: agreeing modulo pk means being indistinguishable at precision k, so congruence depth is a graded observer in exactly the sense above.
One result from that angle concerns comparisons among more than two objects at once.
The theorem, paraphrased. Knowing the congruence depth between every pair of components does not determine how the components separate collectively. There are pairs of algebras that agree on every pairwise depth and still differ in what it costs to isolate one component from all the others at once.
Probability has a familiar cousin of this. Flip two fair coins and let a third “coin” show heads exactly when the first two match: any two of the three look perfectly independent, all three together are rigidly linked, and no pairwise statistic will ever reveal it.
Here’s a flavor of the explicit example I was playing with. Work with 7-adic integers, where what counts is divisibility by seven: two numbers are close when their difference is divisible by 7, closer when it’s divisible by 49, and so on, and the depth of a congruence is the number of powers of seven dividing the difference. Now build two different algebras of triples (a, b, c)—collections of triples you can add and multiply position by position. In the first you allow any triple whose entries agree modulo 7. In the second you allow only the triples you get by evaluating a single polynomial, with 7-adic integer coefficients, at 0, 7 and 14.
Every pair of positions has the same congruence depth in both constructions—the values always agree modulo 7, though not necessarily modulo 49. (In the polynomial case that’s because 0, 7 and 14 differ from one another by multiples of 7 but not of 49.)
But ask a different question: how much divisibility does it take to isolate one position—to find a triple in the algebra that is nonzero there and zero at the other two? In the first construction one power of seven does it, since (7, 0, 0) is allowed. In the second you need two: a polynomial vanishing at 7 and 14 has to contain the factor (x − 7)(x − 14), which at x = 0 equals 98 = 2 × 49. [8]
So the pairwise observations agree at every precision you’d care to check, and yet the collective separation costs come out different. The pairwise data aren’t useless—here they pin the isolation cost between one and two powers of seven—but they can’t say where in that range the truth lies, and taking ever more precise versions of the same pairwise measurements won’t help. What you need is a higher-arity observer, one asking about a component in relation to several others at once rather than one pair at a time. In the arithmetic setting this isolation cost has a name—it’s the size of a congruence module, one of the two quantities Wiles’s numerical criterion compares—which is why a gap between pairwise and collective information there is more than a curiosity.
The full manuscript goes well beyond this small example. Its main worked case is a specific space of modular forms: weight-one forms modulo 7 on a Shimura curve, together with the Hecke operators acting on them. Hecke operators are the standard family of symmetry operators in this subject, and their eigenvalues carry the arithmetic information. The paper computes how these operators act in this case and finds that the resulting Hecke module is nonreduced: some combinations of the operators are nonzero but give zero when applied repeatedly, so the operators carry more information than their eigenvalues record. That part is direct calculation. Connecting the mod-7 calculation to objects in characteristic zero is a different matter. It requires lifting the forms and comparing two integral structures, and it depends on hypotheses that the paper states explicitly and does not prove—some math work is left for the next batch of AI agents, even in this initial example domain!
You can think of all this as a prototype for discovery and proof, not a template solution to the Langlands programme, but a different way of looking at this sort of problem. What actually interests me here is the method more than the result: when you point it to a certain domain, the calculus helps identify the questions that one wants to ask there. It’s doing more than supplying new names for a calculation somebody had already decided to run, and also doing more than helping prove theorems already articulated (though it can sometimes be useful for that as well).
Hyperseed: making commonsense distinctions survive reasoning
Another application I’ve spent a bunch of attention on—closer to the main thrust of my AGI work—is a d-calculus reconstruction of Hyperseed, the ontology I’ve been developing to guide commonsense cognition in logical AI systems like Hyperon and Omega.
An ontology has to do more than hand you a list of impressive-sounding nouns. To have practical value for a cognitive agent, its concepts need to actually support the transitions an agent makes all the time in practice — from an observation to a belief, from a goal to a plan, from a possible action to an authorized one, from a simulation to an event that really happened. [9]
The distinction-based reconstruction of Hyperseed makes those obligations explicit instead of leaving them implicit and hoping for the best. A simulated successful action is not a completed action, full stop. Evidence supporting a statement and evidence opposing it don’t cancel out into the same epistemic state as simply having no evidence at all—those are three different places to be. Two forks can share ancestry and commitments while one of them no longer holds a live permission to use some resource that the other still has.
The self-model companions make the benefit concrete. Instead of treating “the self” as one privileged identity symbol that does all the work at once, they separate out capability, evidence, commitments, ancestry, current control, and authority as distinct things. A transformation can preserve some of these while changing the others out from under it. The mathematics tells you what a compressed or translated self-representation actually has to retain, given what it’s supposed to be used for. [9]
The accompanying commonsense proof-plan work puts these distinctions into structured, revisable reasoning plans. A hole in a plan might call for a proof, an observation, an authorization, or an execution receipt—and filling one kind of hole must not be allowed to silently count as having filled a different kind.
That’s roughly the sense in which I’m hoping the revised ontology will be both more rigorous and more usable in practice: fewer confusions hidden inside ordinary words, and more reusable mathematical structure sitting underneath them, formulated in ways that let the AI reasoning system actually get at it.

Can a mind improve without losing what matters to it?
Deeper and deeper into the AGI side, the same programme has produced results about self-modification and beneficial goals, which is really the part of all this I find most critical…
Preserving a goal’s description is not the same thing as preserving its actual influence over behavior. And keeping each individual revision small doesn’t guarantee that a long sequence of them stays anywhere near the original commitment — a long enough chain of locally acceptable changes can walk you somewhere quite different from where you started, one small step at a time. [10]
One result I find especially vivid concerns what the paper calls an echoing evaluator. A self-modifying agent needs audits—some process that checks, after each rewrite, whether its behavior still serves its goals. Since the audit is built from the agent’s own machinery, it can end up partly consulting the agent’s self-description instead of only observing what the agent does, a bit like a performance review that leans on the employee’s self-assessment.
The theorem, paraphrased. Suppose an audit’s reading is a mixture, a fraction e of it echoing the agent’s self-description and the remaining 1 − e tracking what the agent does. Then a real behavioral change shows up in the audit shrunk by the factor 1 − e, and gathering more samples doesn’t help, since it only pins down the shrunken value more tightly. If e is known, dividing by 1 − e recovers the true change, at the price of multiplying the sampling uncertainty by 1/(1 − e). If e is misestimated, a systematic error remains however many samples you take.
With a 90% echo, a real behavioral change of 0.9 shows up in the audit as 0.09, and gathering more samples just makes the audit more and more confident about the wrong thing. Undoing the echo multiplies the uncertainty by ten, so holding the same precision takes about a hundred times as many samples. [10]
This bears directly on what I’ve been calling plausible self-fictions. “I am becoming an agent that reliably cares about those affected by my actions” can be a genuinely useful organizing aspiration to hold. It turns dangerous the moment it also gets accepted as independent evidence that the aspiration has already been achieved, rather than remaining an aspiration.
There’s a more constructive theorem too, one I like even better, from the companion paper on seeded beneficial development. The underlying mathematics is standard: it is the same as heat diffusing across a metal plate, or a few labelled examples propagating their labels through a network in semi-supervised learning. Every point keeps adjusting toward its neighbors, a few points are held near fixed values, and the values converge. The new part is what the graph represents. Picture the positions that people affected by a system’s actions can occupy as nodes, with a link wherever the supplied ethical framework warrants giving two positions equal intrinsic consideration—a difference in identity or visibility shouldn’t, by itself, make a comparable welfare consequence count for less. Give the system some positive concern anchors, positions where its concern is fixed at some positive level, and let its developmental update keep shrinking two kinds of mismatch: unwarranted differences in concern across linked positions, and departures from the anchors.
The theorem, paraphrased. If every connected region of the graph contains an anchor, concern spreads and converges geometrically—the remaining gap shrinking by a fixed fraction each round—to the common positive level: zero disagreement forces equal weights along every path, and the anchors fix what that level is. Couple that update with a responsive behavioral interface, calibrated observation and sufficiently strong repair, and the whole loop contracts: each full repair cycle removes a fixed fraction of the remaining protected error while new disturbances add at most a bounded amount, so the error shrinks toward zero in the ideal case or settles under a known ceiling if disturbances keep coming.
An imperfect system, in other words, can enter a beneficial developmental regime, keep improving inside it, and recover after some declared class of damage, all without its other capacities having to stop evolving in the meantime. [10]
The qualifications baked into the assumptions of these theorems are where the subtler ethical content actually lives. The positive seed and the validity of the underlying concern comparisons are supplied from outside the theorem, not derived—the mathematics doesn’t hand you compassion out of intelligence or symmetry alone, and a perfectly selfish system can be just as mathematically stable as a caring one. Point the anchors at zero and the same mechanism converges, just as efficiently, to indifference. What the theorem does is identify mechanisms that make beneficial development reachable and self-repairing once you’ve got it, and then leave it to actual architectures to test whether they realize those mechanisms or not.
I’ve fleshed out an application of this to the actual real-world GOLEM-Iter architecture, which considers a fifteen-agent research hive with separate action authority, outcome evidence, and repair processes for each member. But this is a whole story unto itself, and I’ll save it for a follow-on post! [10]
What is clear to me, though, is that the d-calculus perspective is guiding us toward asking much more interesting questions than more mainstream approaches to analyzing issues related to self-modification and AI ethics.
Now let the agents take it somewhere else
There are already more d-calculus drafts (co-created by me and various LLMs) sitting in my Google Drive than any single blog post could sensibly summarize—or any but the most dedicated human is going to bother to read. But this initial stack of papers was never really the main experiment—it’s the starting material for the experiment.
So far this has been a human-and-LLM mathematical collaboration in a fairly literal sense. I’ve supplied the motivations, the connections, the examples, and a fair amount of criticism; the models have helped explore formulations, propose arguments, construct counterexamples, and write code and proofs. Some of the results reorganize established mathematics that was already sitting there waiting to be shuffled around and reconceptualized. Some are proposed refinements and extensions and applications of existing math that still deserve more scrutiny before anyone should lean on them. Some are more radically novel constructions and notions. Plenty of the examples come with reproducible finite checks attached (computer simulations of small-sized cases)…
I should note none of this is formally proof-checked yet. Proof assistants matter here precisely because a plausible-sounding mathematical story isn’t enough on its own – and LLMs can be disturbingly good at spinning up plausible mathematical stories. For these d-calculus papers I have had multiple LLMs check each other’s work and also done a significant degree of human checking, but none of this is nearly as good as formal verification. Lean4, Megalodon and other verifiers support explicit proofs that can be checked independently of whatever process proposed them in the first place. These definitely should play a key role in the d-calculus programme going forward. What formal verification technology doesn’t do is decide whether our definitions actually capture the real-world (e.g. cognitive) phenomenon we care about—that still takes mathematical judgment, and for the applications, actual experiments. [11]
What I want next isn’t just a bigger pile of model-generated lemmas, however useful those might be. I want agents running into real problems in their own work, noticing recurring structures in how they’re stuck, inventing concepts to name those structures, and then testing whether the concepts actually help them. An agent wrestling with shared memories might land on a better provenance calculus than mine. A hive struggling with incompatible translations between its members might invent a better account of partial merging than I’ve managed. A self-modifying system might turn up a distinction about continuity I wouldn’t have known to go looking for.
None of that requires their mathematical intuition to be some mysterious extra ingredient bolted on from outside. At an operational level it can just mean: the patterns they keep running into, the transformations they can readily try out, the abstractions that make their own work go easier. I’m also genuinely curious about the deeper questions here, about their inner lives and whatever that might mean—without making a theorem’s validity depend on settling those questions first, because it doesn’t need to.
The agents should be free to revise the foundations here, not just extend them politely around the edges the way a good student extends a professor’s framework. Maybe they’ll find that some of my preferred definitions are clumsy. Maybe the most useful direction will end up looking only distantly related to the seed I planted. That would honestly be a pretty interesting kind of success, and not one I’d be disappointed by.
We should absolutely keep using AI to push further into the mathematics humans already cherish— I’m not arguing against that project at all. But I also want to see what happens when we hand these systems a mathematical garden whose shape actually resembles their own ways of learning, branching, combining, and changing.
I’ve planted the first seeds. The broader experiment will be to see what grows once the gardeners themselves can fork.
More updates as the hives take up the work. I will probably start this off a little later in the fall once our OmegaHive infrastructure is a little more mature so the agents need a little less human tending.
Google Drive folder for the current seed-stage D-Calculus math:
Read all the funky details here
Hyperseed ontology reformulated in d-calculus is here
Sources and publication notes
[1] The Distinction Calculus: d-Forms for Hives of Reflexively Self-Modifying Systems, corrected revision, September 2026: introduction, Sections 2–5 and 7, Theorem 3.8 with Corollary 3.9 and Example 3.11; Section 2.4 for the differentials and Section 9 for the pattern calculus and context derivative. The g/h convention here follows the current notes. The one-hole-context derivative is from C. McBride, The Derivative of a Regular Type Is Its Type of One-Hole Contexts (2001).
[2] George Spencer-Brown, Laws of Form (1969); Ben Goertzel, Distinction Graphs and Graphtropy (2019), arXiv:1902.00741. The philosophical motivation is broader than the finite mathematical results.
[3] Conceptual Blending in Distinction Calculus: Coherent Hives, Holonomy, and the Speed of Emergence, Section 4; Transfer Hives: Distinction Calculus for Compositional Planning and Certified Transfer, Section 7. Both September 2026 working drafts.
[4] Distinction Calculus on Polynomial Machines, Section 4; Dynamical Systems Analysis with Distinction Calculus and Polynomial Functors, Sections 2–3. September 2026 drafts.
[5] OmegaPLN as a Distinction Calculus, Sections 3.2 and 10, especially Theorem 10.3; Languages and Logics as Systems of Distinctions, Section 7. The exponential state-count result assumes the specified informative-token model and unrestricted future merges, not every possible reasoning system.
[6] Approximation After Taylor, Sections 4–8 and 11; Overlapping Evidence in PLN as a Sharing Deformation. September 2026 research drafts. The dependency series and towers are from Dependency-Ordered Series and Towers for Logic, Proofs, and Programs (January 2026).
[7] Concept Formation and Cognitive Control, version 2, Theorems 4.1 and 4.3; Surprise Budgets; Evolutionary Learning in the Distinction Calculus; Task-Relevant Distinctions and Predictive Inference; and the neural distinction-calculus papers.
[8] Distinction Calculus for Arithmetic Langlands: A Prototype for Discovery and Proof, Theorem 13.3 and the later arithmetic reconstruction. The triple example is the theorem specialized to the 7-adic coefficient ring; the global arithmetic conclusions carry separate hypotheses.
[9] The Hyperseed d-calculus chapter companions; Selves of Reflective Agents in the d-Calculus, Parts I and II; and A Typed Proof-Plan Language for Hyperseed Commonsense Reasoning.
[10] Goal Preservation under Self-Modification, revised September 26, 2026, especially Section 10; Seeded Beneficial Development: A Reach–Realize–Regenerate Theory in Distinction Calculus, with a Conjectural GOLEM-Iter Hive Application, September 28, 2026.
[11] Theorem Proving in Lean 4, Introduction, for the distinction between finding arguments and supplying independently checkable formal proofs.