Source: Conversation transcript â 8 October 2026
This edition preserves the transcript's wording while improving its Markdown structure, spacing, hierarchy, and readability. No substantive content has been rewritten.
do you remember anything from pur chat about my ideas on stateless memory by means if Ran a command, read memory
Claude
[08/10/2026--11:19:44:905 AEDT] Yes. Your message cut off at âby means ofâŚâ, so I canât tell which angle you meant. Here is what I have on stateless memory. Stateless memory in the Orthogonal / âcooperative Attentionâ design. This is the most direct match. LLM1 stays stateless. All state lives in LLM2, which holds the memory pool and the datastore. LLM1âs most salient token is paired with LLM2âs most salient memory token (the âmemtokâ). The pairâs attention weight is inherited from LLM1âs half. That weight then heavily influences LLM1 during inference, between the query and the reply. LLM2 acts as the âhippocampusâ. It is query-driven by LLM1 and returns curated ârich contextâ. LLM1 never touches the datastore. Eviction needs both LLMs to agree on relevance. Relevance is judged against current context, not stored as a fixed score. Your example was that a car-accident memory is worthless if youâre currently immobile. LLM2 doesnât get smarter by repeated querying, because inference doesnât update weights. Instead it generates an explicit âmapâ (an index) in one pass over the static datastore and uses that for later lookups. Later phases add more LLMs in an âorthogonalâ (spherical, wrap-around), IP-internetwork-style configuration. The open item is the validation experiment. It compares a small set of injected high-value tokens against a much larger raw context window on a coding task. You have PyTorch and llama.cpp ready for it. Media-compression memory model. 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
3 hours ago Memory is stored as BMP/WAV-style (raw) and JPG/MP3-style (lossy) grids. Progressive recompression at varying bitrates stands in for sleep consolidation. Contrast boost on the grid stands in for norepinephrine salience in a virtual visual cortex. Vocal isolation and frequency lift do the same in a virtual auditory cortex. The grids arenât viewable or audible. A small, deliberately overfitting co-LLM reads and writes them. MEL / QuasiCode. This is the stateless-AI problem applied to code. âAI has no historical timelineâ, so QCE chains and IFCSAR retrieval supply one from outside the model. Which of these did you mean, or was it something newer that âby means ofâ was leading into? Finish the sentence and Iâll go from there. yes but you have rounded up some of the candidate models are undercooked and halfbaked whilst others were either malformed or too hazy but it's my most recent and most novel idea to date that has captured my imagination ... Ran a command
Claude
[08/10/2026--11:29:57:082 AEDT] Fair. What I gave you were the candidates I could find filed. The idea you mean isnât among them. The filed material stops at the Orthogonal pairs design and the media-grid model, and neither is a complete stateless-memory mechanism âby means ofâ anything. Whatever you were about to name is either a newer idea or one that never got saved in full, so Iâd be guessing if I tried to reconstruct it. Iâd rather not fill the gap with a plausible-sounding version of my own. Tell me the mechanism in your own words, even roughly, and Iâll treat it as the primary design. Iâll hold it to the same standard as the others and raise objections before anything gets built. 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
But I am glad you did because I mananged to accidentally delete the last transcript on it and some details were forgotton ... but you have helped me piece together more of it. Ran a command
Claude
[08/10/2026--11:33:34:654 AEDT] Good, thatâs a useful outcome from a miss. Since the transcript is gone, this conversation is now the only record of the reassembled idea. Write out the version youâve pieced together while itâs fresh, and Iâll keep it as a single clean statement of the design: whatâs settled, whatâs still hazy, and what the stateless part rests on. That gives you a document to start the next chapter from, in a fresh conversation, which suits how you like to work. But because I really wanted to explore and test my theory by 'spinning up' the proposed minimum configuration of two loosely interconnecting streaming LLMs/Transformers (physically managed by HTTP and logically managed by Attention, because apparently "Attention is all you need", right ?) where one acts as a typical inferrence engine whilst the other is non-inferring with no current use for its generative outputs ... Ran a command [08/10/2026--11:57:54:335 AEDT] Your message cut off again after âno current use for its generative outputsâŚâ. Hereâs my reading of the setup so far, with objections up front. Configuration as I understand it LLM1 is an ordinary inference engine and the only one whose output is used. LLM2 runs forward passes but its generated tokens are discarded. It is used for its internal states: attention 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
patterns and hidden activations. The two are joined over HTTP, streaming in both directions. Attention is the logical coupling: LLM1âs salient tokens go to LLM2, and LLM2âs paired memtoks come back to influence LLM1.
llama.cpp and hosted APIs donât expose attention weights or per-token salience over HTTP by default. In llama.cpp you can pull hidden states or embeddings, but the attention maps need a patch. In PyTorch you can hook them easily. For a prototype, PyTorch with forward hooks is the realistic route, and HTTP is just transport. 2. âAttention weight inherited from LLM1â has no native injection point. A returned memtok is a token, and LLM1 only attends over whatâs in its context. So the influence has to enter either as extra tokens in the context, which is plain RAG with extra steps, or as a bias on LLM1âs attention logits or KV cache, which is real intervention but means modifying LLM1âs forward pass. Your âheavy influence between query and replyâ claim is only distinct from RAG if it takes the second route. That decision decides whether the experiment tests anything new. 3. A non-inferring LLM2 is a good fit for state, but a cost to watch. Its job reduces to encoding, scoring and retrieving, which an embedding model plus a vector index also does. The experiment has to show something LLM2âs attention gives you that a cheap retriever doesnât, otherwise the null hypothesis wins. 4. Streaming adds latency before it adds value. Per-token round trips over HTTP between two models will dominate. Run the validation batched first, and stream later.
LLM1 gets a coding task. Condition A is a large raw context. Condition B is a small set of LLM2-selected memtoks injected via the logit-bias route. Condition C is the same small set picked by a plain embedding retriever. B beating C is the result that matters. B beating A alone proves nothing. Finish the sentence about LLM2âs role and Iâll fold it in. I havenât built anything yet. 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
But what hasn't been defined up until now is the nature of the data that LLM2 'trades with' and strives to overtly/heavily influence or skew LLM1's inferrence ... Ran a command
Claude
[08/10/2026--12:12:10:704 AEDT] This is the central design question, so here are the candidates for what crosses the wire. Each one fixes how LLM2 can skew LLM1.
- Text tokens. LLM2 sends words. This is the simplest, and
LLM1 treats them as ordinary context. The influence is only as strong as LLM1 chooses to weight them, so it canât âheavily skewâ anything. It is RAG. 2. Token IDs with a logit bias. LLM2 sends a sparse map of token ID to bias value, added to LLM1âs output logits at each step. This steers vocabulary choice directly and works over HTTP today (llama.cpp and many APIs accept logit bias). It is crude, because it skews what gets said and not what LLM1 attends to. 3. Embedding or hidden-state vectors. LLM2 sends dense vectors that are injected into LLM1âs input embeddings, or added to its residual stream (activation steering). This is the strongest and most genuinely âattention-levelâ route. It requires both models to share an embedding space, or a learned projection between them. The projection needs training, which breaks your âno training at this stageâ rule. 4. KV-cache entries. LLM2 supplies precomputed key/value pairs that LLM1 attends to as if they were its own past context. The pairâs weight is inherited through the keys, which matches your âpairâs attention weightâ idea most literally. It needs the same hidden dimensions and layers, so in practice both are the same model family, ideally the same model. 5. Attention-score priors. LLM2 sends a salience vector over LLM1âs existing context positions, added to LLM1âs attention logits. It reweights what LLM1 already has and adds no new content. 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
My read: your design as described, a pair of (LLM1 salient token, LLM2 memtok) with the weight inherited from LLM1, maps best to option 4, with option 2 as the cheap fallback. Option 4 also keeps the no-training constraint if both models are identical copies. Using two copies of one small model is probably the right first prototype for that reason. Objection: âovertly/heavily skewâ has no safe default. With options 3 to 5, a strong bias degrades LLM1âs coherence quickly, because it is pushed off the distribution it was trained on. So the experiment needs a strength dial and a coherence measure, or a win on relevance will hide a loss on quality. One thing has to be fixed before design continues. Is the data LLM2 trades in tokens, or in vectors? My original model was overly simplistic and closest to option 2 because I have little to no appreciation, understanding or knowledge of what those tokens directly represent and how their embeddings, as vectors, are assembled and stored physically stored, how they are logically read or processed, being numerical values, how firm or flexible they (and under what influences, to what degree and when, if any) and whether or not they have different lifetimes. I assume that Attention influences Genererative (which then influences tokem selection). --- and having just learned of the others, so this leaves me little choice but for some basic trial and error testing, whereby different techniques could match differing types of memory, which although is far more complex, but is closer to that of organic models with obvious advantages ... My original model was overly simplistic and closest to option 2 because I have little to no appreciation, understanding or knowledge of what those tokens directly represent and how their embeddings, as vectors, are assembled and stored physically stored, how they are logically read or processed, being numerical values, how firm or flexible they (and under what influences, to what degree and when, if any) and whether or not they have different lifetimes. I assume that Attention influences Genererative (which then influences tokem selection). --- and having just learned of the others, so this leaves me little choice but for some basic trial and error testing, whereby different techniques could match differing types of memory, which although is far more complex, but is closer to that of organic models with obvious advantages ... Show more Ran a command, read memory
Claude
[08/10/2026--12:56:10:640 AEDT] Trial and error is the right method here, and the gaps you listed are small enough to close in one pass. Here are your questions answered in order. What tokens and embeddings are A token is an integer ID into a fixed vocabulary, for example 50,000 to 150,000 entries. 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
An embedding table is a matrix with one row per token ID. Each row is a vector of a few thousand floats. It is stored as ordinary weights in the model file and is fixed after training. At inference the token ID selects its row. That vector enters the network and passes through every layer. The vector gets rewritten at each layer, so the same word means different things in different contexts. This running vector is the âhidden stateâ or residual stream. In effect a token is a lookup key. The meaning is in the vectors that build up around it. Where attention sits Your assumption is half right. Attention doesnât act on the output directly. At every layer each position computes a query, key and value vector. Attention scores come from comparing queries to keys, and they decide how much of each earlier positionâs value gets mixed into the current position. The final layerâs hidden state is turned into a score for every vocabulary token (the logits), and token selection is made from those. So attention shapes the hidden state, the hidden state gives the logits, and the logits drive selection. Option 2, logit bias, edits the last step only. It changes which word wins but not what the model âunderstoodâ. Firmness and lifetimes Weights, including the embedding table, are permanent and shared. Nothing you inject touches them. Hidden states and the KV cache last for one conversation, because they exist only while the context is held. Whatever you inject lasts as long as you keep re- supplying it. Lifetime is therefore a property of LLM2âs store, which fits your design. LLM1 has no lifetimes of its own. Mapping techniques to memory types This is the trial-and-error plan. Each row is a different kind of memory, and the techniques differ in how deep they reach and how long they persist. 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
Memory type Technique Persistence Cost Fact or word recall Logit bias (option 2) one step trivial Episodic, âwhat was saidâ Text injection (option 1) as long as re-sent trivial Mood, stance, style Activation steering (option 3) as long as applied needs a steering vector from contrast prompts, no training Working memory KV-cache entries (option 4) per session same model only Attention focus Salience prior (option 5) per step needs hooks
two things: did the injected content show up in the output, and did output quality survive. Bias strength is the dial. 2. Use a small open model such as a 1B to 3B Llama or Qwen in PyTorch. Options 3 to 5 need hooks, and llama.cpp doesnât give you those without patching. 3. Skip HTTP for the first round. Run both models in one process, and add HTTP once you know which techniques work. 4. Option 3 is cheaper than it sounds. Run the model on a prompt with the concept and one without. The difference of their hidden states is the steering vector, and no training is involved. If you want, Iâll write the one-process harness with the five injection modes behind a switch. Thatâs one Python file that follows your delivery rules. 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
yes, i am relying on you to translate what i have conceived from my layman's research of and a vivid imagination based a on conjecture and personal logic. Ran 2 commands Ran 3 commands, created a file, shared files The PyTorch CPU index is blocked, so Iâll try plain PyPI.
Claude
[08/10/2026--13:01:45:597 AEDT] Iâve built the harness. The self-test passes, but I couldnât run it against a real model, so youâll be the first to see real output. This sandbox canât reach Hugging Face. The self-test used a tiny random model. It proves the mechanics work: every injection changes the modelâs output, and the attn and keyscale controls reproduce the text baseline exactly. The output text is meaningless gibberish. Changes from the plan I gave you, listed before you open it Option 4 became keyscale , which multiplies the keys of the memory positions. With two identical models, a KV cache from LLM2 would be exactly the same as pasting the text in, because attention only looks backwards. So there was nothing to test in the cache itself. Scaling the keys is the real âattention weightâ dial. There is no HTTP in this version, and both models run in one process, as I suggested. It recomputes the whole sequence at every step and doesnât use a cache. Thatâs slow, but it keeps every injection method simple, and itâs fine for 40 tokens on a small model. The steer strength is a fraction of the modelâs typical activation size, so 0.5 means half. That makes strengths comparable between models. How to read the output Each row shows recall (how many of your --keywords appeared) and ppl (how coherent the text is under the untouched model; lower is better). The text row is the 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
control. A technique is only interesting if it beats text on recall without ppl blowing up. The run lists all six modes, with the four tunable ones swept across four strengths.
Rule Status @version / @date in module done (1.0.0, 2026-10-08) Manual version history entry done Prints version when run done (first line) Zip with dotted, never-reused version done ( orthogonal_probe_v1.0.0.zip ) xmllint on XML not needed (no XML) Legacy zipper not needed (not an SMF package) Unrequested changes listed done (above) Real-model test missing (blocked here; run it on your machine) Run --selftest first on your machine, then the Qwen 0.5B line from the README. Paste the table back and weâll pick the next experiment from the results. Bearing foremost in mind that although "a memory" may appear to be a static, simplistic 2-D representation, I posit that it is actually a tiny logical "frame in" streams ... Ran a command
Claude
[08/10/2026--13:15:14:244 AEDT] Your message cut off after âa tiny logical âframe inâ streamsâŚâ. Iâm reading it as a memory being one frame cut from a 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
continuous stream, in the way a video frame is, and not a standalone record. This is also the logic behind your BMP/JPG model. Iâd count it as a strong idea, and it changes the design in three ways.
means little alone. It is defined by what came before and after it. A stored memory should therefore be a short run of tokens, or a pointer to the stream it was cut from, and not a single memtok. Your âpairs onlyâ rule keeps one salient token per side, and that conflicts with this. A pair is one frame. The stream would have to be a sequence of pairs ordered in time. 2. It gives you the lifetimes you asked about. Video codecs already separate full frames (keyframes) from frames stored as differences. LLM2 could keep rare, rich keyframes and many cheap deltas, then recompress old stretches into fewer keyframes. That is your sleep- consolidation idea, and it comes with a ready-made structure. 3. It has a concrete form in the model. A stream of frames is a sequence of positions, and positions are what attention already works over. The attn and keyscale modes in the harness act on a whole run of memory positions, so they can test this directly. A frame-aware version would also give each frame its own strength, falling off with age. Objection: if each frame is a vector, the frame boundaries become your own convention. The model has no built-in notion of one. Whatever cuts the stream into frames (fixed length, topic shift, or a salience spike) determines everything downstream. Iâd test fixed-length chunks first, since theyâre easy to compare, and only then try salience-triggered cuts. The next harness version (1.1.0) could take --memory as an ordered list of frames with an age-decay on strength, so the same table compares âone memory itemâ against âa stream of framesâ. I havenât built it. Finish the thought about what the stream is made of, and it will settle whether the frames are token runs or vectors. But my idea is based on AI being a stateless inferring and generative engine, THAT STRIVES TO EMULATE 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
REASONING and i have seen this demonstrated by many and varied debugging and troubleshooting "expeditions" with mainly mostly impressive and positive outcomes BUT also with a few "disasters" amd even some endless loops, as I have seen it "dig bigger holes" and make shocking code chnages too but SO TOO should memory be like like that and not just "more text" to be either merged, prepended or appended with the prompt, right ? Ran a command
Claude
[08/10/2026--13:28:40:399 AEDT] Partly right. The âmore textâ criticism holds, but I donât think your own analogy supports âmemory should be like reasoningâ in the way you mean. What holds Memory as extra text in the prompt is the weakest design. The model reads it as new input and gives it no more standing than the prompt. Nothing makes it persist, fade or compete with other memories. Your point that memory should act on the inference, not sit in front of it, is why steer , attn and keyscale exist. They change how the model processes the prompt, and text changes only what it reads. Where I disagree Reasoning emulation is a poor model for memory, and your debugging evidence shows why. The loops, the âbigger holesâ and the shocking code changes come from a stateless engine acting on a context it canât verify, with no stable record of what it has already tried. That is the failure of missing memory. It isnât what memory should look like. Memory that behaves like that engine would inherit those failures. A memory that drifted, dug holes or looped would be worse than none. Organic memory is also not reasoning-like. It is consolidated, decays and is reconstructed on recall, and it is mostly stable. The reasoning-like part is the recall process, deciding what to pull up in context. That is LLM2âs job in your design, and your ârelevance depends on current contextâ insight covers it. What this means for the design LLM1 reasons, and it is allowed to be fallible. 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
LLM2 should be a conservative, consistent store that is hard to push off course. Its one job is supplying the influence. The danger in your design is the coupling. If LLM2âs influence is strong and LLM1 acts on it, a bad memory gets amplified into a bad inference, and LLM1 then reinforces it. Thatâs the loop failure again, built into the architecture.
behaviour. Iâd add a run where the injected memory is wrong, say a false fact, and measure whether strong injection makes LLM1 repeat it, ignore evidence or degrade. That gives you a safe measure of your âdisasterâ cases. 2. Add a decay dial so the influence weakens over turns unless it is reinforced. That matches your lifetimes idea and limits the damage from a bad memory. 3. Keep the text control. The question that matters is whether influence beats text on helpful memories without being worse on wrong ones. Version 1.1.0 would add the wrong-memory test and the decay dial, along with the frame-stream input. I havenât built it, and Iâll wait for your word on scope. you aare taking me literally as it's not the memory that holds value but the outcome(s) derived from it -- so in effence LLM2 adds another dimension to memory, as it's foundations remain as static as they are now but the expanding store is token based and are logical "pairs" both of what LLM1 deems/ed high value and what LLM2 "matches" with it and it dynamic and that relationships are current;y theoretical but LLM2 has a map of the tokens in general, the pairs spefically, and those logical pairs change with and when LLM1 does but is also sujected to what LLM2 deems most relative to -- this of it as a secondary cooperative Attention, that may yield "thoughtful and likely outcomes or simply fail dismally -- am I expressinf myself well enough, being my simplistic seld ? 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
Ran a command
Claude
[08/10/2026--13:42:37:069 AEDT] Yes, and the correction is useful. I was treating the memory as the thing that matters. Youâre saying the thing that matters is the outcome the pairs produce, and I read you too literally. Your idea as I now understand it The base models stay frozen. Nothing in either one is retrained. The new dimension is a growing store of pairs: (what LLM1 judged high value, what LLM2 matched to it). LLM2 holds a map of the tokens in general and of the pairs in particular. It keeps the pairs current. When LLM1âs focus shifts, the pairs shift, and LLM2 also reshapes them by what it judges most relevant. So the store is a set of live relationships between two modelsâ judgments. It isnât a set of facts. The value is in the thoughtful outcome it steers LLM1 towards, and the design accepts it may fail. Iâd call that a secondary cooperative Attention, as you put it. You described it clearly, and what you describe is coherent. Why it changes the test A pair is a relationship, so the harness has been testing the wrong unit. It injects one fixed memory and checks whether LLM1 shows it. Your design predicts something else: the pairs evolve as LLM1âs attention moves, so the thing to measure is whether the pairs improve LLM1âs answers over a run of turns compared with no pairs. What stays theoretical âLLM2 matches withâ needs a definition. The workable one is similarity between LLM2âs hidden state for a candidate and LLM1âs state for the salient token. Without that, the pair has no computable form. âCurrencyâ needs a rule. Pairs update when LLM1âs salient token changes, and decay when they stop being matched. âSimply fail dismallyâ needs a yardstick, so success is defined before you run anything. Iâd use answer quality 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
on a multi-turn coding or Q&A task, with pairs on and off, against the text control. Objection: the strongest competitor is a plain retriever with a recency weight. If your pairs donât beat that, the second model isnât earning its place. That outcome would also be a valid result. The next build (1.1.0) would be a multi-turn loop, where LLM1âs salient token from each turn is paired with a LLM2 match, the pair store updates between turns, and the pairs steer the following turn. It would score pairs on against pairs off. I havenât built it. The step that decides whether it measures your idea is how âmatchesâ is computed, so tell me if hidden- state similarity matches what you mean. say, LLM1 performs its typical Attention based on it's token set but then notifies LLM3 of a fixed percentage (got to start somehere and some consistency is mandatory) pool then LLM2 pairs it with it's Attention, and then LLM1 , before doing anything else, performs Attention based on the combined value of the pair and then we watch what happens, and it may even loop/refine ans since the "weights" are static, unpredictability can be minimised -- now tell me, does Attention play any role in Traning, and either way, why or why not ? Ran a command
Claude
[08/10/2026--13:54:03:675 AEDT] Your loop, restated (Iâm reading âLLM3â as a typo for LLM2, since youâve described only two models)
- LLM1 runs its normal Attention over its tokens.
- It sends LLM2 a fixed percentage of its most salient
tokens, for example the top 10%. 3. LLM2 pairs each with its own best match, using its own Attention. 4. LLM1 runs Attention again, this time over the combined value of each pair, before it does anything else. 5. Repeat, if you want refinement. 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
Your claim that static weights keep this predictable is half right. The weights donât change, but the inputs do, and a loop feeds its own output back in. Thatâs the same structure as the runaway cases you described. The fixed percentage and a hard cap on loop count are the right controls. Iâd add a stop rule: halt when a pass changes the salient set by less than some threshold. Does Attention play a role in training? Yes, a central one. Attention is part of the modelâs architecture, so it runs in every training step as well as in inference. The mechanics: Training feeds in text, the model predicts the next token at every position, and the error against the true token is measured as a loss. That error is pushed backwards through the whole network (backpropagation), and the weights are adjusted to reduce it. The weights that get adjusted include the ones that produce the queries, keys and values. So Attention is learned. Nobody programs what it attends to. By the end of training, the attention weight matrices (the Q, K, V and output projections) encode which relationships matter: grammar, reference, code structure and so on.
What you see at inference is the product of that training. Attention scores are computed fresh from the input each time, but the rules for computing them are frozen weights. That is why the weights staying static makes the system more predictable. It also sets a limit. LLM1 was trained to attend to ordinary text. When you hand it a âcombined pair valueâ, youâre giving it something it never saw in training. That is the reason strong injection degrades coherence in the harness, and the reason LLM2âs output has to be translated into a form LLM1 can attend to without breaking. It raises a question you will hit: if both models were trained separately, their notion of âsalientâ differs. Pairing works best when both share a tokenizer and, ideally, are the same model family. 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
If the combined value is just LLM2âs tokens placed in LLM1âs context, it behaves like text mode. The attn and keyscale modes are closer to your âAttention over the pairâs combined valueâ, because they change the weighting itself. I havenât built 1.1.0. This loop, with its fixed percentage, the pairing step and the stop rule, is what Iâd put in it. yes but LLM1 and LLM2 need to be identical (intially) and even if it is just for curiosities sake, it gives this fixed system "new context, a logical one that could be controlled/adjusted for either deeper/meaningless reasoning that may either influence the outcome with more stability or the opposite but by tweaking the data that Attention works with could be an extensoin of training, a way to train with new data on static systems as NOT NEW TRUTHS but as "influences" that can result in a revised inferrences - how else could you reversibly breath new life into the proven investments that are today's models - i know, it could be a shocker or not but the wisdom is in the pairings and how :predictable LLM1 is in its generative output and maybe LLM2 can also be influenced on the output of LLM2 in exactly the same manner --- food for thought !] yes but LLM1 and LLM2 need to be identical (intially) and even if it is just for curiosities sake, it gives this fixed system "new context, a logical one that could be controlled/adjusted for either deeper/meaningless reasoning that may either influence the outcome with more stability or the opposite but by tweaking the data that Attention works with could be an extensoin of training, a way to train with new data on static systems as NOT NEW TRUTHS but as "influences" that can result in a revised inferrences - how else could you reversibly breath new life into the proven investments that are today's models - i know, it could be a shocker or not but the wisdom is in the pairings and how :predictable LLM1 is in its generative output and maybe LLM2 can also be influenced on the output of LLM2 in exactly the same manner --- food for thought !] Show more Ran a command
Claude
[08/10/2026--14:09:54:961 AEDT] Making the two models identical is a good decision, and it fixes the translation problem I raised earlier. Same weights mean the same tokenizer, the same embedding space and the same idea of salience, so LLM2âs output means something to LLM1 without any trained adapter. Your claim, stated plainly Influencing a frozen modelâs attention at inference is a reversible substitute for training. It adds influence, not new truths, and it can be switched off. This is a real, active research area. It goes by names like activation steering, prefix tuning and in-context learning, and itâs what steer , attn and keyscale already do. So the idea is sound, and it has a precedent. Three limits Iâd keep in view 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
model. Once the pairs stop being supplied, the effect is gone. Thatâs the reversibility you want, but it also means the model never improves. All the accumulated value sits in the pair store. 2. Identical models donât give two views. LLM2 is the same function as LLM1, so given the same input it produces the same salience. The pairing only adds something if LLM2 sees different input: a different slice of the context, the stored memory, or a prompt asking a different question. Otherwise the loop is LLM1 agreeing with itself. This is the point Iâd test hardest. 3. A self-influencing loop can amplify its own errors. Your last idea, where LLM2 is also influenced by LLM1âs output in the same way, makes this sharper. Two identical models steering each other is a feedback loop with no external anchor. It might stabilise on useful structure, or it might lock onto whatever was salient first. This is the âdig bigger holesâ behaviour you mentioned, now designed in. It should be tested in a bounded form first: a fixed number of passes, with a measure of how far the output drifts from the pair-free baseline.
Does a pass through the pair loop change the output relative to baseline, and does it do so consistently from run to run? Does the effect fade cleanly when the pairs are removed? That is your reversibility test. Does a second pass stabilise or drift? Those three tests are the core of 1.1.0, along with the multi- turn loop and the stop rule. The same-weights setup makes it easy to build, since one loaded model can play both roles. I have not written it. Say go and Iâll build it with the usual release card. ok, i need to settle on something and even anthing is better that nothing do i will stop righ there and explore, the least (any) outcome will be a greater width or breadth of understanding so it won't be wasted effort, whatever the 10/8/26, 2:39 PM Claude
<div style="page-break-after: always;"></div>
outcome is. OK, i will save this and then get AI to clean it up for me in various ways until i. approve Ran a command
Claude
[08/10/2026--14:18:07:300 AEDT] Stopping here is sound. You have a testable design, a working v1.0.0 harness, and a clear list of what 1.1.0 would add, so any result will tell you something. For the clean-up stage, here is what to carry into the next conversation so nothing gets lost again: Core claim: two identical frozen models. LLM1 infers. LLM2 is non-inferring and pairs LLM1âs top fixed percentage of salient tokens with its own matches. LLM1 then re-attends over the combined pair value. The effect is reversible influence on a frozen model, not new training. Unit of memory: a pair, treated as a frame in a stream, and not a single stored fact. Settled: the weights stay static, the store lives outside the models, and both models are the same. Still undefined: how LLM2 âmatchesâ (hidden-state similarity is my suggestion), how pairs age and decay, and what makes LLM2 see something different from LLM1. Tests that decide it: pairs against the plain text control, against a recency-weighted retriever, a wrong-memory run, removal to confirm reversibility, and a bounded second pass to check for drift. Harness: orthogonal_probe_v1.0.0.zip is delivered and self-tested. It has not run on a real model yet. 1.1.0 is not built. When the clean-up is done and you want the multi-turn loop built, bring that summary into a new conversation. 10/8/26, 2:39 PM Claude