📝 Stateless Memory Ideas Formatted.mdv4.5.1 · 2026-10-02

Stateless Memory Ideas

Source: Conversation transcript — 8 October 2026

This edition preserves the transcript's wording while improving its Markdown structure, spacing, hierarchy, and readability. No substantive content has been rewritten.


Contents


Page 1

do you remember anything from pur chat about my ideas on stateless memory by means if Ran a command, read memory

Claude

[08/10/2026--11:19:44:905 AEDT] Yes. Your message cut off at “by means of…”, so I can’t tell which angle you meant. Here is what I have on stateless memory. Stateless memory in the Orthogonal / “cooperative Attention” design. This is the most direct match. LLM1 stays stateless. All state lives in LLM2, which holds the memory pool and the datastore. LLM1’s most salient token is paired with LLM2’s most salient memory token (the “memtok”). The pair’s attention weight is inherited from LLM1’s half. That weight then heavily influences LLM1 during inference, between the query and the reply. LLM2 acts as the “hippocampus”. It is query-driven by LLM1 and returns curated “rich context”. LLM1 never touches the datastore. Eviction needs both LLMs to agree on relevance. Relevance is judged against current context, not stored as a fixed score. Your example was that a car-accident memory is worthless if you’re currently immobile. LLM2 doesn’t get smarter by repeated querying, because inference doesn’t update weights. Instead it generates an explicit “map” (an index) in one pass over the static datastore and uses that for later lookups. Later phases add more LLMs in an “orthogonal” (spherical, wrap-around), IP-internetwork-style configuration. The open item is the validation experiment. It compares a small set of injected high-value tokens against a much larger raw context window on a coding task. You have PyTorch and llama.cpp ready for it. Media-compression memory model. 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 2

3 hours ago Memory is stored as BMP/WAV-style (raw) and JPG/MP3-style (lossy) grids. Progressive recompression at varying bitrates stands in for sleep consolidation. Contrast boost on the grid stands in for norepinephrine salience in a virtual visual cortex. Vocal isolation and frequency lift do the same in a virtual auditory cortex. The grids aren’t viewable or audible. A small, deliberately overfitting co-LLM reads and writes them. MEL / QuasiCode. This is the stateless-AI problem applied to code. “AI has no historical timeline”, so QCE chains and IFCSAR retrieval supply one from outside the model. Which of these did you mean, or was it something newer that “by means of” was leading into? Finish the sentence and I’ll go from there. yes but you have rounded up some of the candidate models are undercooked and halfbaked whilst others were either malformed or too hazy but it's my most recent and most novel idea to date that has captured my imagination ... Ran a command

Claude

[08/10/2026--11:29:57:082 AEDT] Fair. What I gave you were the candidates I could find filed. The idea you mean isn’t among them. The filed material stops at the Orthogonal pairs design and the media-grid model, and neither is a complete stateless-memory mechanism “by means of” anything. Whatever you were about to name is either a newer idea or one that never got saved in full, so I’d be guessing if I tried to reconstruct it. I’d rather not fill the gap with a plausible-sounding version of my own. Tell me the mechanism in your own words, even roughly, and I’ll treat it as the primary design. I’ll hold it to the same standard as the others and raise objections before anything gets built. 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 3

But I am glad you did because I mananged to accidentally delete the last transcript on it and some details were forgotton ... but you have helped me piece together more of it. Ran a command

Claude

[08/10/2026--11:33:34:654 AEDT] Good, that’s a useful outcome from a miss. Since the transcript is gone, this conversation is now the only record of the reassembled idea. Write out the version you’ve pieced together while it’s fresh, and I’ll keep it as a single clean statement of the design: what’s settled, what’s still hazy, and what the stateless part rests on. That gives you a document to start the next chapter from, in a fresh conversation, which suits how you like to work. But because I really wanted to explore and test my theory by 'spinning up' the proposed minimum configuration of two loosely interconnecting streaming LLMs/Transformers (physically managed by HTTP and logically managed by Attention, because apparently "Attention is all you need", right ?) where one acts as a typical inferrence engine whilst the other is non-inferring with no current use for its generative outputs ... Ran a command [08/10/2026--11:57:54:335 AEDT] Your message cut off again after “no current use for its generative outputs…”. Here’s my reading of the setup so far, with objections up front. Configuration as I understand it LLM1 is an ordinary inference engine and the only one whose output is used. LLM2 runs forward passes but its generated tokens are discarded. It is used for its internal states: attention 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 4

patterns and hidden activations. The two are joined over HTTP, streaming in both directions. Attention is the logical coupling: LLM1’s salient tokens go to LLM2, and LLM2’s paired memtoks come back to influence LLM1.

Objections

  1. Attention in a deployed model isn’t an interface.

llama.cpp and hosted APIs don’t expose attention weights or per-token salience over HTTP by default. In llama.cpp you can pull hidden states or embeddings, but the attention maps need a patch. In PyTorch you can hook them easily. For a prototype, PyTorch with forward hooks is the realistic route, and HTTP is just transport. 2. “Attention weight inherited from LLM1” has no native injection point. A returned memtok is a token, and LLM1 only attends over what’s in its context. So the influence has to enter either as extra tokens in the context, which is plain RAG with extra steps, or as a bias on LLM1’s attention logits or KV cache, which is real intervention but means modifying LLM1’s forward pass. Your “heavy influence between query and reply” claim is only distinct from RAG if it takes the second route. That decision decides whether the experiment tests anything new. 3. A non-inferring LLM2 is a good fit for state, but a cost to watch. Its job reduces to encoding, scoring and retrieving, which an embedding model plus a vector index also does. The experiment has to show something LLM2’s attention gives you that a cheap retriever doesn’t, otherwise the null hypothesis wins. 4. Streaming adds latency before it adds value. Per-token round trips over HTTP between two models will dominate. Run the validation batched first, and stream later.

Minimum test I’d run

LLM1 gets a coding task. Condition A is a large raw context. Condition B is a small set of LLM2-selected memtoks injected via the logit-bias route. Condition C is the same small set picked by a plain embedding retriever. B beating C is the result that matters. B beating A alone proves nothing. Finish the sentence about LLM2’s role and I’ll fold it in. I haven’t built anything yet. 10/8/26, 2:39 PM Claude

<div style="page-break-after: always;"></div>


Page 5

But what hasn't been defined up until now is the nature of the data that LLM2 'trades with' and strives to overtly/heavily influence or skew LLM1's inferrence ... Ran a command

Claude

[08/10/2026--12:12:10:704 AEDT] This is the central design question, so here are the candidates for what crosses the wire. Each one fixes how LLM2 can skew LLM1.

  1. Text tokens. LLM2 sends words. This is the simplest, and

LLM1 treats them as ordinary context. The influence is only as strong as LLM1 chooses to weight them, so it can’t “heavily skew” anything. It is RAG. 2. Token IDs with a logit bias. LLM2 sends a sparse map of token ID to bias value, added to LLM1’s output logits at each step. This steers vocabulary choice directly and works over HTTP today (llama.cpp and many APIs accept logit bias). It is crude, because it skews what gets said and not what LLM1 attends to. 3. Embedding or hidden-state vectors. LLM2 sends dense vectors that are injected into LLM1’s input embeddings, or added to its residual stream (activation steering). This is the strongest and most genuinely “attention-level” route. It requires both models to share an embedding space, or a learned projection between them. The projection needs training, which breaks your “no training at this stage” rule. 4. KV-cache entries. LLM2 supplies precomputed key/value pairs that LLM1 attends to as if they were its own past context. The pair’s weight is inherited through the keys, which matches your “pair’s attention weight” idea most literally. It needs the same hidden dimensions and layers, so in practice both are the same model family, ideally the same model. 5. Attention-score priors. LLM2 sends a salience vector over LLM1’s existing context positions, added to LLM1’s attention logits. It reweights what LLM1 already has and adds no new content. 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 6

My read: your design as described, a pair of (LLM1 salient token, LLM2 memtok) with the weight inherited from LLM1, maps best to option 4, with option 2 as the cheap fallback. Option 4 also keeps the no-training constraint if both models are identical copies. Using two copies of one small model is probably the right first prototype for that reason. Objection: “overtly/heavily skew” has no safe default. With options 3 to 5, a strong bias degrades LLM1’s coherence quickly, because it is pushed off the distribution it was trained on. So the experiment needs a strength dial and a coherence measure, or a win on relevance will hide a loss on quality. One thing has to be fixed before design continues. Is the data LLM2 trades in tokens, or in vectors? My original model was overly simplistic and closest to option 2 because I have little to no appreciation, understanding or knowledge of what those tokens directly represent and how their embeddings, as vectors, are assembled and stored physically stored, how they are logically read or processed, being numerical values, how firm or flexible they (and under what influences, to what degree and when, if any) and whether or not they have different lifetimes. I assume that Attention influences Genererative (which then influences tokem selection). --- and having just learned of the others, so this leaves me little choice but for some basic trial and error testing, whereby different techniques could match differing types of memory, which although is far more complex, but is closer to that of organic models with obvious advantages ... My original model was overly simplistic and closest to option 2 because I have little to no appreciation, understanding or knowledge of what those tokens directly represent and how their embeddings, as vectors, are assembled and stored physically stored, how they are logically read or processed, being numerical values, how firm or flexible they (and under what influences, to what degree and when, if any) and whether or not they have different lifetimes. I assume that Attention influences Genererative (which then influences tokem selection). --- and having just learned of the others, so this leaves me little choice but for some basic trial and error testing, whereby different techniques could match differing types of memory, which although is far more complex, but is closer to that of organic models with obvious advantages ... Show more Ran a command, read memory

Claude

[08/10/2026--12:56:10:640 AEDT] Trial and error is the right method here, and the gaps you listed are small enough to close in one pass. Here are your questions answered in order. What tokens and embeddings are A token is an integer ID into a fixed vocabulary, for example 50,000 to 150,000 entries. 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 7

An embedding table is a matrix with one row per token ID. Each row is a vector of a few thousand floats. It is stored as ordinary weights in the model file and is fixed after training. At inference the token ID selects its row. That vector enters the network and passes through every layer. The vector gets rewritten at each layer, so the same word means different things in different contexts. This running vector is the “hidden state” or residual stream. In effect a token is a lookup key. The meaning is in the vectors that build up around it. Where attention sits Your assumption is half right. Attention doesn’t act on the output directly. At every layer each position computes a query, key and value vector. Attention scores come from comparing queries to keys, and they decide how much of each earlier position’s value gets mixed into the current position. The final layer’s hidden state is turned into a score for every vocabulary token (the logits), and token selection is made from those. So attention shapes the hidden state, the hidden state gives the logits, and the logits drive selection. Option 2, logit bias, edits the last step only. It changes which word wins but not what the model “understood”. Firmness and lifetimes Weights, including the embedding table, are permanent and shared. Nothing you inject touches them. Hidden states and the KV cache last for one conversation, because they exist only while the context is held. Whatever you inject lasts as long as you keep re- supplying it. Lifetime is therefore a property of LLM2’s store, which fits your design. LLM1 has no lifetimes of its own. Mapping techniques to memory types This is the trial-and-error plan. Each row is a different kind of memory, and the techniques differ in how deep they reach and how long they persist. 10/8/26, 2:39 PM Claude

<div style="page-break-after: always;"></div>


Page 8

Memory type Technique Persistence Cost Fact or word recall Logit bias (option 2) one step trivial Episodic, “what was said” Text injection (option 1) as long as re-sent trivial Mood, stance, style Activation steering (option 3) as long as applied needs a steering vector from contrast prompts, no training Working memory KV-cache entries (option 4) per session same model only Attention focus Salience prior (option 5) per step needs hooks

Objections

  1. Test all five on one model with fixed prompts, and score

two things: did the injected content show up in the output, and did output quality survive. Bias strength is the dial. 2. Use a small open model such as a 1B to 3B Llama or Qwen in PyTorch. Options 3 to 5 need hooks, and llama.cpp doesn’t give you those without patching. 3. Skip HTTP for the first round. Run both models in one process, and add HTTP once you know which techniques work. 4. Option 3 is cheaper than it sounds. Run the model on a prompt with the concept and one without. The difference of their hidden states is the steering vector, and no training is involved. If you want, I’ll write the one-process harness with the five injection modes behind a switch. That’s one Python file that follows your delivery rules. 10/8/26, 2:39 PM Claude

<div style="page-break-after: always;"></div>


Page 9

yes, i am relying on you to translate what i have conceived from my layman's research of and a vivid imagination based a on conjecture and personal logic. Ran 2 commands Ran 3 commands, created a file, shared files The PyTorch CPU index is blocked, so I’ll try plain PyPI.

Claude

[08/10/2026--13:01:45:597 AEDT] I’ve built the harness. The self-test passes, but I couldn’t run it against a real model, so you’ll be the first to see real output. This sandbox can’t reach Hugging Face. The self-test used a tiny random model. It proves the mechanics work: every injection changes the model’s output, and the attn and keyscale controls reproduce the text baseline exactly. The output text is meaningless gibberish. Changes from the plan I gave you, listed before you open it Option 4 became keyscale , which multiplies the keys of the memory positions. With two identical models, a KV cache from LLM2 would be exactly the same as pasting the text in, because attention only looks backwards. So there was nothing to test in the cache itself. Scaling the keys is the real “attention weight” dial. There is no HTTP in this version, and both models run in one process, as I suggested. It recomputes the whole sequence at every step and doesn’t use a cache. That’s slow, but it keeps every injection method simple, and it’s fine for 40 tokens on a small model. The steer strength is a fraction of the model’s typical activation size, so 0.5 means half. That makes strengths comparable between models. How to read the output Each row shows recall (how many of your --keywords appeared) and ppl (how coherent the text is under the untouched model; lower is better). The text row is the 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 10

control. A technique is only interesting if it beats text on recall without ppl blowing up. The run lists all six modes, with the four tunable ones swept across four strengths.

Release card

Rule Status @version / @date in module done (1.0.0, 2026-10-08) Manual version history entry done Prints version when run done (first line) Zip with dotted, never-reused version done ( orthogonal_probe_v1.0.0.zip ) xmllint on XML not needed (no XML) Legacy zipper not needed (not an SMF package) Unrequested changes listed done (above) Real-model test missing (blocked here; run it on your machine) Run --selftest first on your machine, then the Qwen 0.5B line from the README. Paste the table back and we’ll pick the next experiment from the results. Bearing foremost in mind that although "a memory" may appear to be a static, simplistic 2-D representation, I posit that it is actually a tiny logical "frame in" streams ... Ran a command

Claude

[08/10/2026--13:15:14:244 AEDT] Your message cut off after “a tiny logical ‘frame in’ streams…”. I’m reading it as a memory being one frame cut from a 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 11

continuous stream, in the way a video frame is, and not a standalone record. This is also the logic behind your BMP/JPG model. I’d count it as a strong idea, and it changes the design in three ways.

  1. A memory item has to carry its neighbours. A frame

means little alone. It is defined by what came before and after it. A stored memory should therefore be a short run of tokens, or a pointer to the stream it was cut from, and not a single memtok. Your “pairs only” rule keeps one salient token per side, and that conflicts with this. A pair is one frame. The stream would have to be a sequence of pairs ordered in time. 2. It gives you the lifetimes you asked about. Video codecs already separate full frames (keyframes) from frames stored as differences. LLM2 could keep rare, rich keyframes and many cheap deltas, then recompress old stretches into fewer keyframes. That is your sleep- consolidation idea, and it comes with a ready-made structure. 3. It has a concrete form in the model. A stream of frames is a sequence of positions, and positions are what attention already works over. The attn and keyscale modes in the harness act on a whole run of memory positions, so they can test this directly. A frame-aware version would also give each frame its own strength, falling off with age. Objection: if each frame is a vector, the frame boundaries become your own convention. The model has no built-in notion of one. Whatever cuts the stream into frames (fixed length, topic shift, or a salience spike) determines everything downstream. I’d test fixed-length chunks first, since they’re easy to compare, and only then try salience-triggered cuts. The next harness version (1.1.0) could take --memory as an ordered list of frames with an age-decay on strength, so the same table compares “one memory item” against “a stream of frames”. I haven’t built it. Finish the thought about what the stream is made of, and it will settle whether the frames are token runs or vectors. But my idea is based on AI being a stateless inferring and generative engine, THAT STRIVES TO EMULATE 10/8/26, 2:39 PM Claude

<div style="page-break-after: always;"></div>


Page 12

REASONING and i have seen this demonstrated by many and varied debugging and troubleshooting "expeditions" with mainly mostly impressive and positive outcomes BUT also with a few "disasters" amd even some endless loops, as I have seen it "dig bigger holes" and make shocking code chnages too but SO TOO should memory be like like that and not just "more text" to be either merged, prepended or appended with the prompt, right ? Ran a command

Claude

[08/10/2026--13:28:40:399 AEDT] Partly right. The “more text” criticism holds, but I don’t think your own analogy supports “memory should be like reasoning” in the way you mean. What holds Memory as extra text in the prompt is the weakest design. The model reads it as new input and gives it no more standing than the prompt. Nothing makes it persist, fade or compete with other memories. Your point that memory should act on the inference, not sit in front of it, is why steer , attn and keyscale exist. They change how the model processes the prompt, and text changes only what it reads. Where I disagree Reasoning emulation is a poor model for memory, and your debugging evidence shows why. The loops, the “bigger holes” and the shocking code changes come from a stateless engine acting on a context it can’t verify, with no stable record of what it has already tried. That is the failure of missing memory. It isn’t what memory should look like. Memory that behaves like that engine would inherit those failures. A memory that drifted, dug holes or looped would be worse than none. Organic memory is also not reasoning-like. It is consolidated, decays and is reconstructed on recall, and it is mostly stable. The reasoning-like part is the recall process, deciding what to pull up in context. That is LLM2’s job in your design, and your “relevance depends on current context” insight covers it. What this means for the design LLM1 reasons, and it is allowed to be fallible. 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 13

LLM2 should be a conservative, consistent store that is hard to push off course. Its one job is supplying the influence. The danger in your design is the coupling. If LLM2’s influence is strong and LLM1 acts on it, a bad memory gets amplified into a bad inference, and LLM1 then reinforces it. That’s the loop failure again, built into the architecture.

Objections and tests

  1. The harness currently has no test for loops or runaway

behaviour. I’d add a run where the injected memory is wrong, say a false fact, and measure whether strong injection makes LLM1 repeat it, ignore evidence or degrade. That gives you a safe measure of your “disaster” cases. 2. Add a decay dial so the influence weakens over turns unless it is reinforced. That matches your lifetimes idea and limits the damage from a bad memory. 3. Keep the text control. The question that matters is whether influence beats text on helpful memories without being worse on wrong ones. Version 1.1.0 would add the wrong-memory test and the decay dial, along with the frame-stream input. I haven’t built it, and I’ll wait for your word on scope. you aare taking me literally as it's not the memory that holds value but the outcome(s) derived from it -- so in effence LLM2 adds another dimension to memory, as it's foundations remain as static as they are now but the expanding store is token based and are logical "pairs" both of what LLM1 deems/ed high value and what LLM2 "matches" with it and it dynamic and that relationships are current;y theoretical but LLM2 has a map of the tokens in general, the pairs spefically, and those logical pairs change with and when LLM1 does but is also sujected to what LLM2 deems most relative to -- this of it as a secondary cooperative Attention, that may yield "thoughtful and likely outcomes or simply fail dismally -- am I expressinf myself well enough, being my simplistic seld ? 10/8/26, 2:39 PM Claude

<div style="page-break-after: always;"></div>


Page 14

Ran a command

Claude

[08/10/2026--13:42:37:069 AEDT] Yes, and the correction is useful. I was treating the memory as the thing that matters. You’re saying the thing that matters is the outcome the pairs produce, and I read you too literally. Your idea as I now understand it The base models stay frozen. Nothing in either one is retrained. The new dimension is a growing store of pairs: (what LLM1 judged high value, what LLM2 matched to it). LLM2 holds a map of the tokens in general and of the pairs in particular. It keeps the pairs current. When LLM1’s focus shifts, the pairs shift, and LLM2 also reshapes them by what it judges most relevant. So the store is a set of live relationships between two models’ judgments. It isn’t a set of facts. The value is in the thoughtful outcome it steers LLM1 towards, and the design accepts it may fail. I’d call that a secondary cooperative Attention, as you put it. You described it clearly, and what you describe is coherent. Why it changes the test A pair is a relationship, so the harness has been testing the wrong unit. It injects one fixed memory and checks whether LLM1 shows it. Your design predicts something else: the pairs evolve as LLM1’s attention moves, so the thing to measure is whether the pairs improve LLM1’s answers over a run of turns compared with no pairs. What stays theoretical “LLM2 matches with” needs a definition. The workable one is similarity between LLM2’s hidden state for a candidate and LLM1’s state for the salient token. Without that, the pair has no computable form. “Currency” needs a rule. Pairs update when LLM1’s salient token changes, and decay when they stop being matched. “Simply fail dismally” needs a yardstick, so success is defined before you run anything. I’d use answer quality 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 15

on a multi-turn coding or Q&A task, with pairs on and off, against the text control. Objection: the strongest competitor is a plain retriever with a recency weight. If your pairs don’t beat that, the second model isn’t earning its place. That outcome would also be a valid result. The next build (1.1.0) would be a multi-turn loop, where LLM1’s salient token from each turn is paired with a LLM2 match, the pair store updates between turns, and the pairs steer the following turn. It would score pairs on against pairs off. I haven’t built it. The step that decides whether it measures your idea is how “matches” is computed, so tell me if hidden- state similarity matches what you mean. say, LLM1 performs its typical Attention based on it's token set but then notifies LLM3 of a fixed percentage (got to start somehere and some consistency is mandatory) pool then LLM2 pairs it with it's Attention, and then LLM1 , before doing anything else, performs Attention based on the combined value of the pair and then we watch what happens, and it may even loop/refine ans since the "weights" are static, unpredictability can be minimised -- now tell me, does Attention play any role in Traning, and either way, why or why not ? Ran a command

Claude

[08/10/2026--13:54:03:675 AEDT] Your loop, restated (I’m reading “LLM3” as a typo for LLM2, since you’ve described only two models)

  1. LLM1 runs its normal Attention over its tokens.
  2. It sends LLM2 a fixed percentage of its most salient

tokens, for example the top 10%. 3. LLM2 pairs each with its own best match, using its own Attention. 4. LLM1 runs Attention again, this time over the combined value of each pair, before it does anything else. 5. Repeat, if you want refinement. 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 16

Your claim that static weights keep this predictable is half right. The weights don’t change, but the inputs do, and a loop feeds its own output back in. That’s the same structure as the runaway cases you described. The fixed percentage and a hard cap on loop count are the right controls. I’d add a stop rule: halt when a pass changes the salient set by less than some threshold. Does Attention play a role in training? Yes, a central one. Attention is part of the model’s architecture, so it runs in every training step as well as in inference. The mechanics: Training feeds in text, the model predicts the next token at every position, and the error against the true token is measured as a loss. That error is pushed backwards through the whole network (backpropagation), and the weights are adjusted to reduce it. The weights that get adjusted include the ones that produce the queries, keys and values. So Attention is learned. Nobody programs what it attends to. By the end of training, the attention weight matrices (the Q, K, V and output projections) encode which relationships matter: grammar, reference, code structure and so on.

Why this matters for your design

What you see at inference is the product of that training. Attention scores are computed fresh from the input each time, but the rules for computing them are frozen weights. That is why the weights staying static makes the system more predictable. It also sets a limit. LLM1 was trained to attend to ordinary text. When you hand it a “combined pair value”, you’re giving it something it never saw in training. That is the reason strong injection degrades coherence in the harness, and the reason LLM2’s output has to be translated into a form LLM1 can attend to without breaking. It raises a question you will hit: if both models were trained separately, their notion of “salient” differs. Pairing works best when both share a tokenizer and, ideally, are the same model family. 10/8/26, 2:39 PM Claude

<div style="page-break-after: always;"></div>


Page 17

If the combined value is just LLM2’s tokens placed in LLM1’s context, it behaves like text mode. The attn and keyscale modes are closer to your “Attention over the pair’s combined value”, because they change the weighting itself. I haven’t built 1.1.0. This loop, with its fixed percentage, the pairing step and the stop rule, is what I’d put in it. yes but LLM1 and LLM2 need to be identical (intially) and even if it is just for curiosities sake, it gives this fixed system "new context, a logical one that could be controlled/adjusted for either deeper/meaningless reasoning that may either influence the outcome with more stability or the opposite but by tweaking the data that Attention works with could be an extensoin of training, a way to train with new data on static systems as NOT NEW TRUTHS but as "influences" that can result in a revised inferrences - how else could you reversibly breath new life into the proven investments that are today's models - i know, it could be a shocker or not but the wisdom is in the pairings and how :predictable LLM1 is in its generative output and maybe LLM2 can also be influenced on the output of LLM2 in exactly the same manner --- food for thought !] yes but LLM1 and LLM2 need to be identical (intially) and even if it is just for curiosities sake, it gives this fixed system "new context, a logical one that could be controlled/adjusted for either deeper/meaningless reasoning that may either influence the outcome with more stability or the opposite but by tweaking the data that Attention works with could be an extensoin of training, a way to train with new data on static systems as NOT NEW TRUTHS but as "influences" that can result in a revised inferrences - how else could you reversibly breath new life into the proven investments that are today's models - i know, it could be a shocker or not but the wisdom is in the pairings and how :predictable LLM1 is in its generative output and maybe LLM2 can also be influenced on the output of LLM2 in exactly the same manner --- food for thought !] Show more Ran a command

Claude

[08/10/2026--14:09:54:961 AEDT] Making the two models identical is a good decision, and it fixes the translation problem I raised earlier. Same weights mean the same tokenizer, the same embedding space and the same idea of salience, so LLM2’s output means something to LLM1 without any trained adapter. Your claim, stated plainly Influencing a frozen model’s attention at inference is a reversible substitute for training. It adds influence, not new truths, and it can be switched off. This is a real, active research area. It goes by names like activation steering, prefix tuning and in-context learning, and it’s what steer , attn and keyscale already do. So the idea is sound, and it has a precedent. Three limits I’d keep in view 10/8/26, 2:39 PM Claude


<div style="page-break-after: always;"></div>


Page 18

  1. Influence is not learning. Nothing persists inside the

model. Once the pairs stop being supplied, the effect is gone. That’s the reversibility you want, but it also means the model never improves. All the accumulated value sits in the pair store. 2. Identical models don’t give two views. LLM2 is the same function as LLM1, so given the same input it produces the same salience. The pairing only adds something if LLM2 sees different input: a different slice of the context, the stored memory, or a prompt asking a different question. Otherwise the loop is LLM1 agreeing with itself. This is the point I’d test hardest. 3. A self-influencing loop can amplify its own errors. Your last idea, where LLM2 is also influenced by LLM1’s output in the same way, makes this sharper. Two identical models steering each other is a feedback loop with no external anchor. It might stabilise on useful structure, or it might lock onto whatever was salient first. This is the “dig bigger holes” behaviour you mentioned, now designed in. It should be tested in a bounded form first: a fixed number of passes, with a measure of how far the output drifts from the pair-free baseline.

What the harness should measure for your claim

Does a pass through the pair loop change the output relative to baseline, and does it do so consistently from run to run? Does the effect fade cleanly when the pairs are removed? That is your reversibility test. Does a second pass stabilise or drift? Those three tests are the core of 1.1.0, along with the multi- turn loop and the stop rule. The same-weights setup makes it easy to build, since one loaded model can play both roles. I have not written it. Say go and I’ll build it with the usual release card. ok, i need to settle on something and even anthing is better that nothing do i will stop righ there and explore, the least (any) outcome will be a greater width or breadth of understanding so it won't be wasted effort, whatever the 10/8/26, 2:39 PM Claude

<div style="page-break-after: always;"></div>


Page 19

outcome is. OK, i will save this and then get AI to clean it up for me in various ways until i. approve Ran a command

Claude

[08/10/2026--14:18:07:300 AEDT] Stopping here is sound. You have a testable design, a working v1.0.0 harness, and a clear list of what 1.1.0 would add, so any result will tell you something. For the clean-up stage, here is what to carry into the next conversation so nothing gets lost again: Core claim: two identical frozen models. LLM1 infers. LLM2 is non-inferring and pairs LLM1’s top fixed percentage of salient tokens with its own matches. LLM1 then re-attends over the combined pair value. The effect is reversible influence on a frozen model, not new training. Unit of memory: a pair, treated as a frame in a stream, and not a single stored fact. Settled: the weights stay static, the store lives outside the models, and both models are the same. Still undefined: how LLM2 “matches” (hidden-state similarity is my suggestion), how pairs age and decay, and what makes LLM2 see something different from LLM1. Tests that decide it: pairs against the plain text control, against a recency-weighted retriever, a wrong-memory run, removal to confirm reversibility, and a bounded second pass to check for drift. Harness: orthogonal_probe_v1.0.0.zip is delivered and self-tested. It has not run on a real model yet. 1.1.0 is not built. When the clean-up is done and you want the multi-turn loop built, bring that summary into a new conversation. 10/8/26, 2:39 PM Claude