I’m going to write here about some technical work I’ve been playing with over the last few weeks – which is going quite well and which, if it keeps going well, might turn out to have some rather interesting consequences. It’s an AI architecture I’ve been calling NESS—the NEural-Symbolic Sandwich—and basically, it’s a way of taking the open-weights transformers we already have and restructuring and retraining parts of them in a way designed to make the whole thing more AGI-ish and a better participant in broader AGI architectures.
I’ll go into moderate depth here, and link to paper drafts going deeper at the end—but let me start with the TL;DR for technical readers: NESS takes a pretrained transformer, freezes most of it, distills its top layers into a small adaptable bridge, and puts three more layers on top: a symbolic layer of executable programs and statistical “readers” that work over structured records and history, a probabilistic combiner that learns how much to trust each of them and how their evidence should interact, and a small neural cap that turns the combined evidence into the final prediction.
So—the bottom layer changes slowly or not at all, it’s just derived from the vanilla transformer base. The four layers above it keep learning ongoingly and continually all the time—and the symbolic one in the middle is built to plug into a diversity of other AI tools, including Hyperon’s AtomSpace, MeTTa, PLN and program evolution.
What I’ve run so far are just tiny prototypes—around half a million neural parameters—but the results have been interesting and the math suggests the methods will likely scale up. At this tiny scale,
- The discovered symbolic programs improve prediction by about 7–8% over matched purely-neural controls
- There are meaningful long-term memory properties: When the system is trained on text, then only on Python, then text again, the neural bridge forgets a good deal but the whole system reverses roughly 88–94% of the prose loss it had incurred, largely because the symbolic layer held its ground.
None of that is anywhere near AGI yet of course; and even with promising math, scaling algorithms and architectures always needs some trial and error tuning, and there can be hidden rocks.
However, my intuition is fairly fired up here—this sandwich feels like the shape of a transformer that can keep learning ongoingly, and that speaks the same symbolic language as a logical reasoning engine and the rest of the Hyperon architecture. Plugging a bigger NESS network into an Omega agent along with Hyperon Atomspaces and associated symbolic and evolutionary algorithms, one would obtain an even more powerful thing than the current Omega system: An agentic loop in which the neural and symbolic dynamics involved in each pass through the control group are actively cooperating and communicating, sharing internal representations rather than just feeding each other outputs as inputs.

Where transformers sit in my picture of AGI
As anyone who follows my work will know, I don’t think LLMs are the holy path to AGI. I do think they can be very useful components of AGI systems—which is why in Omega, the agent framework my colleagues and I have been building, we have an agentic loop wrapping LLM calls, and along the way a bunch of calls into Hyperon’s symbolic working memory and long-term memory (and, more experimentally, medium-term memory), with symbolic reasoning going on inside each go-round of the loop. We’re gradually tweaking Omega to shift more and more of the goal pursuit and self-understanding into the symbolic side and to rely less and less on the LLM. Ultimately, in a more mature Omega network, I see the LLM as a communication tool for going back and forth between structured knowledge and English or other human languages, and as a sort of lossy, questionable, but nonetheless valuable knowledge oracle. Hyperon, with its symbolic reasoning and evolutionary learning and so forth, should do the rest.
But since there is an LLM in the loop, even if it isn’t the crux of the Omega loop’s intelligence, I’d like it to be as AGI-ish as possible. And there are a number of things stopping current LLMs from being fully human-level-AGI-ish. The basic representation of knowledge they have is overly tied to the specific information they’ve been fed, so they do a bit too much recombining of details from the training data and a bit too little wild, general-purpose abstraction and creative leaping beyond those details. That’s a deep issue and I’ll come back to it. The other major shortcoming—the one NESS is most directly aimed at—is continual learning.
The frozen-model problem
The way LLMs work today, you can’t do weight learning and activation-space learning and inference all in the same process at the same time. You train the model, you freeze the weights, and then you have a frozen model—Opus 5.6, Astra, whatever—that you do inference with. You can get some learning out of it in the activation space, via in-context learning, which is super cool and amazing as far as it goes. Then, to maintain context beyond what the activation space will hold, the model gets propped up with a bunch of external memory devices, some of which work well and some of which work badly. Anyone who uses these systems regularly can see that the props aren’t great.
Part of why Omega works as nicely as it does is that Hyperon’s symbolic memory serves as a much better long-term memory prop for the LLM than the usual bag of retrieved snippets—better than a prop, really, since we’re doing interesting reasoning over what’s stored rather than just looking things up. To be fair, not every external memory used in the standard LLM world is an improvised bolt-on: retrieval-augmented generation can train the retriever and the generator jointly, and the recent wave of work on “test-time” neural memory is already exploring alternatives to fixed deployment.[3] But in all of these it’s still the LLM’s shortcomings being compensated for from outside. The issue is how tightly remembering, interpreting, reasoning and learning are woven together—whether they’re one activity or four loosely coupled ones.
For the reader less intimate with modern neural voodoo, let me be precise about where the shortcoming lives. Backpropagation is the method used to adjust a neural network’s weights based on error signals propagated backward from its output. Transformers are an architecture, a particular way of wiring a network together around attention. Backpropagation is a powerful algorithm, dating back to the 1950s, but it has some particular shortcomings too.
For many complex neural architectures it just won’t do the learning in any tractable period of time—it “won’t converge”—and this has guided the field away from some neural architectures and toward others. In particular it has directed the field toward architectures with relatively few and simplistic recurrent connections, because too much recurrence tends to make backprop choke.
Backprop-based neural networks are also commonly prone to “catastrophic forgetting”—after you train them a while they start to exceed their memory capacity and can then suddenly forget a lot of their knowledge very rapidly. They don’t gradually lose track of old information to make room for new information. This problem is not unique to backprop—it exists for example in Hopfield associative memory neural networks as well, though there it can be worked around with tricks like reverse learning. This problem is the root of the current divide between training and inference in LLMs—it is the reason we freeze a model after its weights are trained, and then after that don’t let the weights change, but instead allow only the “in context learning” that can happen in the activation space
A trivial example of the stupidity this division leads to: at the moment, all the LLMs I’m using inside my hive of AI agents were evidently trained before Telegram introduced its bot-to-bot API. So on the first pass, these agents think it isn’t possible for Telegram bots to talk to each other—even though they are, at that very moment, controlling Telegram bots that are talking to each other. Usually they figure out what they’re doing and then look on the Web to get the new information their models don’t incorporate. But every so often one of them consults its underlying model, the model says no, Telegram has no bot-to-bot API, and the agent loses track of the fact that it is currently using Telegram’s bot-to-bot API. The agents can store the true fact in their AtomSpace, or in a vector memory; what they can’t do is update the LLM, so the LLM keeps offering the same wrong information over and over, and now and then they believe it.
That’s a trivial case, since it’s just an empirical external fact, and storing it in a database mostly fixes it. The thing is, there are loads and loads of subtler cases like it that aren’t so easy to iron out with a database entry—cases where what needs to change is how the model interprets a whole class of situations, not one fact it can look up.
The kind of memory I find myself most concerned with is what I’d call medium-term memory: the evolving knowledge of a project, a relationship, an investigation, an unfinished argument. It lives in the awkward middle ground between what fits in the current context window and what’s been consolidated into a general skill. A bag of retrieved snippets doesn’t maintain the distinction between what happened, what I thought happened, and what subsequent evidence changed my mind about—and maintaining that distinction is a big part of what medium-term memory is for. For robust self-modeling, other-modeling and world-modeling, I want that continuity built into the learning process itself: an agent ought to be able to track which of its own methods have failed, which of its commitments are still live, what another agent seems to know, and why its own confidence in something went up or down. (For now you can just consider this a functional claim, not a claim about machine consciousness; I’ll leave those questions for other posts.)
What I would ideally like inside an Omega loop, then, is just a transformer that keeps updating its weights as it goes along. The human brain does this. We have long-term potentiation and other such mechanisms; the weights on our neural connections get updated in semi-real time as activation spreads around them. The updating tends to be slower than the activation spreading, but it’s all wrapped up in one process rather than being a separate offline phase.
Predictive coding, and the long road to making it work at scale
I’m fairly confident one can train a multi-trillion-parameter transformer, with some variation on the current architectures, in a way that supports continual learning, and there are probably many ways to do it. The one my team at SingularityNET and BGI Labs has been experimenting with most is predictive coding, which I’ve talked about before but will recap here very briefly for newer readers.
Yes, believe it or not, you can train neural nets using algorithms other than standard backprop. In predictive coding—PC—instead of sweeping a weight-adjustment pass through the whole network from top to bottom, each neuron does its own learning, trying to predict what the neurons in nearby layers are doing. There’s then a settling-in process: neurons here adapt to neurons there, neurons there adapt to neurons here, over and over, and through this sometimes rather chaotic iterative process the network as a whole settles into a functional attractor. Only once things have settled do the persistent weights get updated, via a local rule rather than a global gradient. When you train networks this way you often find they don’t suffer nearly as badly from catastrophic forgetting.
That said, most of the PC literature has been on vision rather than language, and PC in practice is still catching up to backprop. In our SingularityNET research team, we’ve gotten a predictive-coding transformer to work OK at the GPT-2 scale. Generally what we find in our early experiments is that a PC network’s continual-learning performance is better than a comparable backprop network’s, but not as good as we’d like. But there is so much plain room for improvement!
I think of the situation with PC today as similar to the long road backprop itself traveled. When I was teaching backprop neural nets in the ‘90s, it took something like three hours to get a 35-neuron recurrent backprop network to converge to anything. Part of that was the Solaris workstation I was using, which was primitive compared to a modern smartwatch, but part of it was that recurrent backprop on a richly connected architecture was just the wrong thing to be doing, and it took years to figure out the tricks—filtering between layers, limited doses of recurrence, attention mechanisms and all the rest. Those tricks won’t necessarily be the same tricks for PC. So we’re trying to rush through that process for predictive coding, and I do think that once it’s worked through we’ll end up with something with much more brain-like properties, including better continual learning and the ability to train locally, neuron by neuron, across a decentralized network of machines. Our FabricPC toolkit is speeding this along considerably, and as far as I can tell we’re the first group seriously trying to make PC work at large scale in practice, as opposed to demonstrating interesting phenomena at small scale in academia.
But—all that said—I’ve gotten impatient. That work is coming along, and it’s fascinating, and a huge breakthrough could well come any day… but I’m in a hurry to make sure the open and decentralized and moral-agency-oriented approach wins the AGI race, so I don’t want to depend too desperately on any particular thread of research. So I’ve started to think about alternate, complementary paths as well… which is where NESS came from.
Keep the expensive part, replace the top
So what I’ve been musing in the background for months is: OK, take a big open-weights transformer—of the kind the Chinese labs keep releasing with such admirable generosity—freeze most of it, leave it as it is, and replace the top of it with something more flexible, more brain-like, more neural-symbolically linkable, more capable of continual learning.
Maybe the base is basically a model of what’s on the web, and maybe it’s fine to refresh that every three or six months, since the web doesn’t change all that fast. Maybe it’s even fine to leave it frozen even longer if the top bits of the network are built the way we really want.
There’s an obvious economic argument for this sort of approach. The useful pretrained models already embody an immense amount of computation and engineering—billions of dollars’ worth at the frontier—and we don’t need to repeat that investment before investigating a better organization of adaptation.
There are roughly similar precedents of course: parameter-efficient methods such as LoRA keep a large model essentially fixed while training a small set of extra parameters, and knowledge distillation trains a smaller “student” network to reproduce the behavior of a larger “teacher.”[4]
The device NESS uses is what we call a cap—a learned module attached on top of a model that already works, which can read the base’s internal representations and predictions, consult additional evidence, and improve the output. I should say up front that updating fewer parameters narrows the learning problem without making the base’s forward pass, or retrieval, or recurrent inference free. “Efficient” here should not be read as “cheap.”
And I suspect the slow-base, fast-top split may be more than a transitional compromise. Perhaps broad perceptual and linguistic competence is the kind of thing that should change slowly, while the machinery that interprets new experience, constructs abstractions and coordinates action should change fast. A stable base gives the fast-changing parts something reliable to build on. Upgrading it, when that happens, means checking that old memories and learned interfaces still mean what we think they mean—the same problem any long-lived piece of software faces when a dependency changes underneath it. The interesting question is whether the whole system keeps learning coherently while its parts learn at different speeds.
The first thing I tried was simply training a predictive-coding cap on top of a standard transformer. That worked OK at small scale, but not as well as I wanted. Next I tried distilling the base network—taking a backprop-trained transformer and, using a homotopic distillation procedure, converting it into what’s called an error-PC network, a kind of predictive-coding network that looks a lot like a backprop network—and then putting a more flexible PC cap on top of that, co-training the distillation with the learning of the cap. That worked a little better. Still not as well as I wanted. And that is what led me to the approach I’m playing with now.

NESS, the Dagwood Bumstead Sandwich of AGI
I call it the neural-symbolic sandwich—NESS, pronounced like the Loch Ness monster. The LLMs I was using to help develop it don’t like the sandwich metaphor. You’re not just piling things on top of each other like a sandwich, they told me; it’s all about the interfaces between the layers. And of course they’re right, but LLMs have a bad sense of poetry. If you’re old enough to remember the newspaper comic Blondie, Dagwood Bumstead used to eat these towering multi-layer sandwiches, and that’s what I’m envisioning: five layers, each doing something different, with the interesting stuff happening where they meet.
The five layers, bottom to top, are: a frozen pretrained transformer; a distilled top-couple-of-layers of that transformer, adapted into a bridging layer; a symbolic layer, which is the most interesting part; a probabilistic combiner; and a neural cap that produces the output distribution. The bottom layer is fixed—for the small cases I’ve been running I train it myself, but for a big case you’d take an existing open-weights model—and the other four are all co-trained for the purpose. What’s the loss function? You can play with that, but for now it’s standard transformer stuff, because we’re not trying to train a full AGI here. We’re trying to train a transformer that can be decentralized-trained and do robust continual learning, to plunge into an Omega loop alongside the other AI components I have. If it turns out to be an AGI on its own I won’t complain, but that’s not what I’m looking for. Figure 1, up at the top of this post, shows the whole stack; the reader who made it this far is rewarded with a brief explanation of what each layer does.
Layer 1: the frozen foundation
Keep the lower part of a competent pretrained model. It supplies broad representations of language—or images, or whatever the base was trained on—without being rewritten by every local experiment the upper layers run. Freezing it removes one big source of interference between old and new learning. It doesn’t guarantee the complete predictor can never forget, since the layers above can still drift; as we’ll see, they do.
Layer 2: the distilled bridge
The representation most useful for the original language-modeling objective isn’t necessarily the most useful interface to a symbolic system. So rather than keeping the whole base frozen, we take the top couple of transformer blocks and distill them into a smaller, adaptable student, trained together with the more unusual things that sit above it. You can do the distillation with backprop, or with error-PC, or whatever. The student’s job isn’t just to predict well on its own but to expose useful evidence to the rest of the sandwich. In the small reference implementation two teacher blocks became one narrower student block; how much to distill, how far to compress and where to cut are experimental variables rather than commandments.[1]
Layer 3: the symbolic filling
The middle layer is a diverse ensemble of symbolic programs, and I’ve experimented with a whole bunch of options for what goes in there. The programs take neural stuff as input—the output of the distilled bridge—and they also take history as input, and an important point is that they have to use context to produce their outputs. They can form their own context from a symbolic history, and they can also look at the context fished up by the frozen base’s attention mechanism and use that to grab the relevant bits of symbolic history. Neural and symbolic attention, both aimed at a symbolic record of what’s happened. Over all of this you learn an ensemble of programs.
Why put symbols in the middle at all? A symbolic system has a straightforward representational advantage: you can add a fact, revise a relation or compose a new program without retraining a big network to encode that particular addition, so incremental knowledge and flexible long-term memory are ordinary operations rather than heroic ones. And these operations are close to the work we want AI to do. Code has variables, scope and data flow; mathematics has quantified statements and proof dependencies; spreadsheets and databases have fields, identities, formulas and joins; resources like the Gene Ontology and PubChem hold biological and chemical knowledge in structured form.[5] Why translate all that into prose and then ask a neural model to rediscover, from the prose, relationships that were explicit in the source? Better to let the language model help interpret the question, let an executable program follow the actual dependency, and let a learned consumer decide what the result means in context.
One particularly interesting way to build these programs is with semantic frames. Say text is coming into the transformer. Run it through an entity extractor and a semantic relation extractor, so you get predicate-argument structure: if the sentence is “Ben walked through the field,” the transformer is doing its neural thing, while the relation extractor says Ben is an entity, the field is an entity, and there’s a walked-through relationship between them. Modern transformers extract this kind of thing very robustly—we could do it even before transformers, with the Stanford and Berkeley parsers of pre-transformer computational linguistics—and in principle the extraction could be co-trained with the rest of the network, though for now existing tooling works fine.
Given the relations, the symbolic programs become combinations of semantic relations with properly aligned variable bindings, which is to say they represent frames and events. For instance: one actor sends a message, another actor receives it, there’s a channel it’s sent on, there’s an encoding scheme, and the sender and the receiver are both using the same encoding scheme. That’s a bundle of semantic relations involving a couple of actors, held together by shared variable bindings, and “sending a message between a sender and a receiver over a channel using a mutually understood encoding” is a little program that lives in the symbolic layer and can be recognized as something coming through the input. Even in the small examples I’ve run there are hundreds or thousands of these; a frontier-scale model would have an insane number of them.
A toy version of the same idea, at the scale the current prototypes run at: Dan needs an object that’s in a locked cabinet; someone Dan trusts is willing to help; someone has a key that fits the cabinet. Those facts alone don’t get Dan his object. What’s needed is that the trusted helper be the very person holding the key that opens the cabinet containing the object—one individual satisfying all three roles at once. A small program enforces that repeated identity exactly, where a neural predictor can only approximate it, and NESS’s prototype program search found exactly this sort of “shared witness” computation useful for prediction.[1,2]
Since the whole point of putting programs in the middle rather than just more neurons is that you can read them, let me show what the search in the latest prototype came up with. The toy world it lives in has 20 entities and 9 relation types—trusts, helps, holds, opens, inside and so on—and the program grammar builds candidates out of typed paths through those relations, plus “shared-witness” joins that insist two paths land on the same entity. Three of the programs the search selected form a rather pleasing ladder (conceptually, I mean; I’m not claiming the search built them in this order).
The simplest, program 0—call it direct access—is holds(A, K), opens(K, C), inside(O, C): the actor A holds a key K, that key opens a container C, and the object O being sought is inside that container. Three relations, with the same K and the same C wherever they appear.
Program 398, the helpful holder, is helps(H, A), holds(H, K), opens(K, C), inside(O, C): now it isn’t the actor holding the key but some helper H who is helping the actor—and it has to be the same H doing the helping and the holding.
Program 681, the trusted helper, is trusts(A, H), helps(H, A), holds(H, K), opens(K, C), inside(O, C): the helpful holder with one more relation bolted on, namely that the helper also be someone the actor trusts. Five relations, one helper, one key, one container, all tied together by the repeated variables. Each program returns 1 if there is some way to bind its variables to actual entities so that every one of its relations holds at once, and 0 otherwise. The repeated variables are the entire point: it isn’t enough that someone helps Dan and someone holds a suitable key, it has to be the same someone. Program 681 was retained in all three full search runs, which is about as clear a vote as a search procedure can cast.
Here is how it plays out on an actual held-out world from the archived test set—real names and relations, with only the relevant edges shown. Dan seeks the object Gus. Dan trusts Ivo; Ivo helps Dan; Ivo holds the key Pia; Pia opens the container Joy; Gus is inside Joy. Executing program 681 amounts to a set intersection: the people Dan trusts, {Ivo}; the people who help Dan, {Ivo}; the people holding a key to wherever Gus is, {Kai, Ivo}. The intersection is {Ivo}, so the program returns 1—one identical helper satisfies the trust, the help, and the whole key–container–object chain.
Now swap two trust endpoints and leave everything else alone. In the original world Dan trusts Ivo and Eli trusts Eli; in the swapped world Dan trusts Eli and Eli trusts Ivo. The local text is the same, the word counts are the same, every entity has the same number of incoming and outgoing relations, and the cropped prefix the neural layers see—“… Qin. Dan seeks Gus. Routine. Then ”—is byte-for-byte identical. Program 398 still returns 1 in both worlds, because Ivo still helps Dan and still holds the key. Program 681 returns 1 in the original world and 0 in the swapped one, because the set of people Dan trusts is now {Eli}, and Eli is not the one with the key.
And the prediction follows the binding. The next event in this world is Dan asking for aid, so the next byte to be predicted is the “A” of “Asks aid.” Averaged over three saved models, a context-only version of the system—one that sees the same text but gets no program evidence—puts 52.3% on that byte in the original world and 43.9% in the swapped world; it barely notices that anything changed. The system with the full discovered program bank goes from 85.4% to 18.0%. For reference, the process that generated the world, which the models never see as a training label, assigns 88.4% and 17.2%: the program-equipped system is tracking the real structure of the situation, while the context-only one is guessing from surface statistics. This is one illustrative pair, selected using observations only, and not an aggregate result—the aggregates come below—but it shows in miniature what the programs are doing: an exact, inspectable, shared-identity computation, moving the next-byte forecast in a way the neural predictor on its own could only approximate.[2]
How do these programs get found? Not by backprop. The search learns program syntax by search and program consequences by learning. It enumerates candidates from the typed grammar (1,209 of them in the latest run), screens them by how well their features line up with what the neural predictor is currently getting wrong, fits a small context-specific consumer for each of a dozen nominees per round, and admits the one or two per seed that improve held-out prediction after a charge for complexity. Then it recomputes the residual and searches again. The structure of a program is discrete and you can print it out; what its output means for the next byte is something the layers above learn.
In the current implementation, the simplest programs are what we call readers. Some readers return complete predictive distributions learned from statistics—a table of which bytes tended to follow a particular kind of context. Other programs return features rather than distributions: whether a shared-binding condition holds, which entity satisfies a relation, what a query against a store came back with. So learning in this layer is more than adjusting weights: we can fit a reader’s table, search for a better reader or program, and learn how the layers above should consume the output. One caveat: the binding experiments so far work from supplied typed records, structured facts handed to the system, and don’t yet demonstrate automatic extraction of arbitrary symbolic knowledge from raw prose. Bolting the frame-extraction pipeline described above onto the prototypes is one of the obvious next steps.[1,2]
The statistical readers have the same see-through quality. One reader the early search generated refined a plain “current word prefix” reader by keying on the previous word as well, so that the prefix “c” after “the” and the prefix “c” after “I” got separate next-byte tables—about 12,000 stored keys in place of the parent’s 11,306—with sparse entries backing off to the parent’s table. Of the 98 candidate readers that search tested, that was one of the two it kept. It is a new conditional distribution the system learned to consult, and you can open the table and look at it, which is more than one can say for a weight matrix.
The symbols here aren’t a rationalization generated after the answer to make it look reasonable; they participate in producing the answer. And because they’re programs rather than activations, a discovered program can be inspected, reused, tested under changed bindings, and eventually handed to other components—which is where Hyperon comes in, and I’ll get to that.
Layer 4: the probabilistic combiner
The fourth layer maps the outputs of the symbolic programs into the inputs the top neural cap wants. When I looked at this properly I realized it’s a probabilistic programming problem, which is cool: you’re trying to transform the distribution of outputs of these programs into the distribution of inputs the top layer needs in order to give the desired output. A pile of correct computations is not yet a calibrated prediction. Several programs may be reusing the same evidence, so naively adding their votes double-counts it; a rule may apply under only one reading of the input; a statistical reader may be reporting from a context it has barely seen. The system has to learn how the evidence should interact, rather than concatenating everything and hoping the top layer sorts it out.
For now I train a simple probabilistic combiner, and you can do that with a variety of mixture models: choose among a small number of evidence-use strategies and sum their forecasts according to learned probabilities, with no training label ever having to say which strategy was the “correct” one. The latest prototype is a bit more specific than that — a learned prior over how reliable each reader is, adjusted through a few steps of inference using outcomes the system has already observed. Down the road this layer can be an arbitrarily complex learned probabilistic program; what we have now is a tractable instance of the bridge we need.[1]
Layer 5: the neural cap
The top layer is a neural net again, with a few layers of its own, giving the output probability distribution that ultimately gets judged and whose loss gets calculated. It learns what the combined evidence means for the task, translating a crisp logical result into a graded adjustment and handling interactions the explicit programs don’t express. Ultimately I want this to be a predictive-coding net, and I’ve experimented with PC, backprop and other functions up there. A backprop-trained cap is a perfectly legitimate NESS model and an essential comparison—we’d be fooling ourselves if we only ever ran the PC version. Whatever learns the cap, its output has to respect the task: for language modeling the final token probabilities must form a properly normalized distribution.
I should stress that these are five logical roles, not a rigid pipeline. Neural context stays available alongside symbolic evidence at every stage. In the current implementation the distribution-valued readers enter through the mixture in layer 4, while the discovered binding features go straight to the final neural correction in layer 5. The sandwich describes how the interfaces cooperate; it doesn’t demand that information be thrown away at every boundary. (So the LLMs had a point. It is about the interfaces. It’s still a sandwich.)[1]

What the small experiments show, and what they don’t
I’ve been training this on very small networks, and it works interestingly well. The work grew out of a sequence of small experiments run under the internal label H3, which tested a number of different mechanisms and shouldn’t be read as one steadily climbing benchmark curve. Some earlier lexical experiments found that readers generated by search could beat fixed, hand-specified readers—including one refinement that conditioned on both the current word prefix and the preceding word. Other runs found that more search, or a more elaborate consumer, didn’t automatically help, which is the kind of negative result it’s good to have in hand before scaling anything up.[2]
The most recent “binding reference” model is deliberately tiny: about 454,000 deployed neural parameters, with the external reader tables and other storage counted separately. It sees a local context of 35 bytes, and evaluation scores every byte of a 29-byte continuation. These are controlled mechanism tests, not frontier results, and I’d ask you to hold that frame through everything that follows.
The most clear-cut result concerns whether discovered programs help prediction. In the controlled co-training experiment we ran three inherited seeds with eight branches per seed; corresponding branches began from identical predictions and received matched additional training examples and matched numbers of updates, so that the one systematic difference between a branch and its control was the presence of the program evidence. Against their respective no-program controls, the discovered binding programs reduced the loss on complete structured continuations by 8.20% with a frozen core and 6.93% with a jointly trained core. The improvement depended on the bindings being right—damaging or swapping the program evidence made predictions worse, not merely no better—and the neural controls saw the same underlying observations, so the programs weren’t smuggling in extra facts.[2]
I read this as evidence that an explicit computation can earn its place inside a neural predictor, on the predictor’s own terms. It is not evidence that symbolic features supply facts the network couldn’t otherwise get at. And there’s a result here I’d have preferred to go the other way: co-training the core alongside the programs did not increase the programs’ incremental benefit. That interaction came out slightly negative, even as performance on ordinary prose improved. So we have useful cooperation to build on, but not yet the stronger synergy—the kind where the neural and symbolic sides make each other better—that the whole approach is ultimately after.
Text to Python and back again
The little experiment that got me most excited is the one that speaks directly to continual learning. I trained the network on text, then trained it more on Python code, then trained it on text again, and then checked: did it forget everything? Concretely, 200 updates on text, then 800 updates on code only, then 200 updates on text. No text examples were replayed during the code phase, and no checkpoint was restored when text returned; the system just had to cope. The lower trunk and the reader tables stayed fixed throughout while the upper layers adapted. We compared an anchored condition, which penalized the upper system for straying too far from the pre-code predictor on current-domain inputs, against an otherwise matched unanchored one—twelve trajectories in all, with disjoint sets of standard-library modules for Python training and test, so the code results aren’t memorization of specific files.[2]
Here’s the prose result for the program-equipped models, averaged over three seeds. The unit is bits per byte, a standard measure of how well a model predicts text: lower means the predictor assigned higher probability to the bytes that occurred.
| Prose test, by training stage | With anchoring | Without anchoring |
| Before code training | 2.4669 | 2.4669 |
| After code-only training | 2.8012 | 2.9458 |
| After returning to text | 2.4872 | 2.5237 |
Held-out prose loss in bits per byte; three-seed means, all 29 continuation bytes scored. Source: NESS empirical record, text–code–text experiment.[2]
Meanwhile, held-out code loss fell by about 32% with anchoring and 33% without it—so retention wasn’t being purchased by refusing to learn the new domain. After text training resumed, the complete system reversed roughly 94% and 88% of the prose deterioration, respectively.[2]
That’s encouraging resilience, though it doesn’t yet show all the robustness I would hope for. There is still plenty to play with, and plenty of work to be done. Recovery is not the same thing as uninterrupted retention. Prose got appreciably worse during the code phase; neither mean fully returned to its starting point; and returning to text kept only about 42% and 38% of the newly acquired code gain. There was interference in both directions, which is what one expects from any system that learns at all.
What’s most interesting, though, is what happened inside. Layer 2, the distilled part of the base transformer, forgot a lot: when the system learned Python it lost a good chunk of its text ability, and in the unanchored model, after text came back, the upper transformer component taken on its own was still about 0.840 bits per byte worse on prose than before the code phase. But the complete predictor was only about 0.057 worse. The symbolic layer, of course, doesn’t forget in the same way—you don’t have to delete the programs and reader tables for text just because you’re training on Python—and because it held its ground, the whole network got its integrated prediction ability back fairly quikly once text came back. The preserved readers and the learned combination are the plausible compensators, though this test didn’t isolate their causal shares.[2]
That the continuity symbolic systems get for free carried over to the integrated neural-symbolic system was not obvious. It could easily have gone the other way: you train the thing on text and then on Python and it just gets discombobulated and can’t segue back. It segued back nicely. Two caveats, though. The binding programs in particular were inactive on ordinary prose and code, and the no-program branches showed similar prose retention, so those particular programs don’t get the credit. And the upper-transformer diagnostic is an internal measurement, not a separately trained conventional-transformer baseline matched for size and data. What the result supports is a promising division of labor among the layers. All this doesn’t do is declare catastrophic forgetting solved—but I believe it’s a highly promising direction, and I won’t be surprised if we make further progress in this direction very rapidly.
Why richer programs may bring predictive coding into its own
Moving toward systems in the broad capability class of GPT-4 or GPT-5 changes the kinds of tasks worth attempting, and not merely by enlarging the vocabulary of simple rules. Useful programs at that level may need nested quantification, multiple scopes, uncertain reference, and compositions of procedures.
Consider the difference between “for every sample, there is an assay that works” and “there is one assay that works for every sample.” Same words, different quantifier order, and the second is a much stronger claim than the first; which one is being made determines what has to be searched for. In code, a variable’s identity depends on its scope. In an enterprise workflow, whether an action is permitted may depend on who approved which action for which account. Flatten these relationships into a bag of features and you can erase precisely the distinction that decides the answer.
My expectation is that richer interfaces to such programs will often benefit from recurrence—from going around the loop more than once. An initial neural interpretation proposes some bindings; symbolic execution reveals an inconsistency; the interpretation gets revised; a different program becomes relevant; and so on until things settle. The small neural layers below and above the symbolic layer may need to take part in that back-and-forth rather than each making one irrevocable pass. This is exactly the kind of iterative, local reconciliation of evidence that predictive coding is built around, which is a large part of why I expect PC to come into its own here. There are also established links between PC-style or equilibrium learning and ordinary gradient-based learning, under explicit conditions, so adopting PC needn’t mean abandoning what we know about backprop.[6]
That said, I don’t want to oversell PC (as much as I like it). We will need to try the experiments and see if PC sells itself in this context. BP can train recurrent computation too, including by differentiating through a finite sequence of inference steps, and PC doesn’t automatically make recurrence faster, more accurate or more stable. My latest NESS retention and binding results in fact used BP, run through a finite number of reliability-inference steps; they are not evidence of PC’s superiority. My earlier comparisons showed that PC was competitive in some settings—but were more like a respectable showing for PC rather than a universal win.[1,2] My guess is that PC will deserve a stronger role as the interpretation problem becomes more tightly coupled, but it’s a guess, and the decisive experiment has to separate two questions that are easy to run together: whether recurrence helps at all, and whether PC is the best way to learn it.
Plugging the sandwich into Hyperon, and giving it a history
Now let’s think about the symbolic level a bit more. It can go back and forth with a Hyperon AtomSpace. If you have the whole sandwich wrapped in an Omega loop, it’s no longer a frozen neural net interacting with a separate symbolic memory. A typed program, a relation with provenance attached, or an uncertain hypothesis is something the broader architecture already knows how to handle—it can live in an AtomSpace, be rewritten by MeTTa code, be reasoned about by PLN, be evolved by MOSES-style program search—rather than being one more embedding whose meaning is private to a single network. And the flow goes the other way too: the output of semantic relation extraction can feed the symbolic layer now, and the output of Hyperon reasoning can feed it later, giving the neural predictor new computations to learn to exploit. To be plain about status, this two-way integration is what we want to build. The prototypes described above use supplied typed records and locally discovered programs, not a live Hyperon connection.
Another next step I find especially appealing is a shared episodic memory attached to both the distilled bridge and the neural cap—the two slices of bread. Daniel MacDonald’s Indexed Attention Transformer proposal offers one concrete route: retain raw experience together with selected historical, memory-conditioned neural representations, and train how those records get read and written. The reason to attach memory on both sides is that the two reads would do different jobs: an early read at the bridge helps interpret the situation before the symbolic layer gets to work, and a later read at the cap retrieves a relevant precedent or exception after the programs have exposed the situation’s structure. Throughout, the raw evidence and its provenance have to stay available, because an old interpretation is not automatically a true belief. This is a proposal at present, not part of the measured results above.[7]
Over the last couple days one of my colleagues and his AI coding agent minions have buil at systematic NESS research codebase around modular interfaces, including a FabricPC backend for the predictive-coding side. The ambition is to be able to swap substrates, symbolic mechanisms, memory systems, combiners and learning rules without rebuilding the whole experiment each time. The same modularity opens up lateral experiments—a NESS cap above a time-series forecaster, say, or frozen language and vision models feeding a shared symbolic layer—which are going to be quite fun to run.
What I want to learn next is fairly concrete. Do reusable programs still improve prediction when the neural baseline gets stronger? Can memory remain usable as the interpretations that index it evolve? Can the upper components keep adapting without losing the ability to use what they learned before? And does recurrent predictive coding earn its computational cost?

Idealism, pragmatism, and the race
I think of NESS as an example of the flexibility of possible AGI designs, and also of the way we’re mixing idealism with pragmatism in building hybrid architectures that can hopefully be made to outperform vanilla frontier LLMs fairly soon. We are in an AGI race here. It’s not necessarily the optimal dynamic to be operating in, but it’s the reality. There are optimal ways to do things, and replacing the whole transformer with a predictive-coding network designed from the ground up for neural-symbolic interfacing is probably a very good way. But we want to get to beneficial AGI quickly, and so it’s worth exploring a variety of complementary routes to the neural side of neural-symbolic AGI (as we also work to scale up and refine the symbolic Hyperon side)… if you can get far enough by putting a neural-symbolic cap on existing open-weights models, that may be what gets us there in time.
To reiterate the point about the race, my reasoning and feeling here is that: What we need to do to get a beneficial future for our species is to have an open, decentralized AGI with an architecture capable of real moral agency, teach it human values in a sensible way, and roll it out on a globally distributed decentralized infrastructure as soon as possible—so that the open, decentralized, beneficial side of things is in the lead, ahead of the closed, proprietary, military, purely profit—and nationalism—driven side.
That’s not to say the military won’t use this stuff, or that we and others won’t make a profit with it. It’s to say that open and decentralized needs to be in the lead, and everyone else can play catch-up. And to get there, we can’t necessarily wait around to totally replace backprop transformers with superior things. Unless someone cooks up a route even faster than the one I’m thinking of, which would be great.
We don’t need to replace every synaptic weight in current open-weights models to build a system that keeps learning continuously, and can serve as a close partner to symbolic AI in neural-symbolic AGI architectures like Omega. We need the right division of labor between stable competence, revisable interpretation, explicit knowledge and persistent discovery. NESS is my most recent (and I think most interesting so far) attempt to turn that question into an architecture—and then to flesh out and refine the architecture via experimentation…
Explore the work
Research draft: See this folder for a series of (LLM-assisted) paper drafts going over the math and practicalities of NESS as currently conceived. Quick technical-ish overview here.
Code and contribution starting point: github.com/guilamacie/NESS
Sources and research notes
The numbered notes distinguish the completed prototype record from outside precedents and from proposed extensions.
1. NESS architecture. The NESS working-paper series, Paper 1: Architecture and Working Principles and Paper 2: Mathematical Framework and Implemented Instantiation (September 2026). These internal research drafts define the five interfaces, typed readers, binding grammar, probabilistic consumers, and the distinct BP and PC implementations. They do not establish large-scale superiority.
2. Prototype results. NESS Paper 3: Experimental Record and Evidence Audit, especially Sections 11–12, together with the archived H3 v44/v45 results and, for the worked shared-witness example, the archived v43 program catalogue, test worlds and saved forecasts. All numerical values in this post come from that completed record; no experiment was rerun for the post. Recovery percentages measure reversal of a loss increase, not a fraction of all knowledge retained. Earlier experiments used different model families and data.
3. Alternatives to fixed deployment. P. Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020); A. Behrouz, P. Zhong and V. Mirrokni, Titans: Learning to Memorize at Test Time (2025 preprint). These are precedents, not NESS experiments.
4. Efficient adaptation and distillation. E. J. Hu et al., LoRA: Low-Rank Adaptation of Large Language Models (2021); G. Hinton, O. Vinyals and J. Dean, Distilling the Knowledge in a Neural Network (2015). They support the general building blocks, not a measured cost advantage for a large NESS system.
5. Structured scientific knowledge. Gene Ontology Consortium, Gene Ontology overview; S. Kim et al., PubChem 2025 update. These illustrate structured knowledge resources; no NESS integration with either is claimed.
6. Predictive coding and equilibrium learning. B. Millidge, A. Tschantz and C. L. Buckley, Predictive Coding Approximates Backprop along Arbitrary Computation Graphs (2020); B. Scellier and Y. Bengio, Equilibrium Propagation: Bridging the Gap Between Energy-Based Models and Backpropagation (2017). Their learning rules and assumptions are distinct; neither establishes a universal PC advantage.
7. Episodic read–write extension. D. MacDonald, Indexed Attention Transformers: Continual In-Context Learning over an Expanding Experience Log, revised internal concept manuscript, 24 September 2026. The dual attachment to the NESS bridge and cap is a proposed adaptation of that idea, not a completed experimental result.
8. Codebase and implementation scope. The NESS repository and its implementation-status record, consulted 27 September 2026. Repository capability statements are developer-reported, not independently reproduced for this draft. The new platform and its engineering demonstrations are distinct from the earlier H3 experiments summarized above.