📝 inspect_Qwen_Qwen2.5-0.5B.mdv4.5.1 · 2026-10-02

Model inspection report: Qwen/Qwen2.5-0.5B

Generated by model_inspector v1.1.0 (2026-10-11) on 2026-10-10 22:40:08. Read-only: no weights were changed.

5. What the model does with your prompt

Two runs of the same prompt. A (reference) = the prompt alone. B (test) = the prompt alone, with the memory 'direction' added to the running thought at layer 12, strength 0.5 x the typical size. Everything below is measured from inside the model on those two runs - no guessing.

Run A (reference) - the prompt alone

What the model would say if it had to stop at each layer (logit lens)

layer best guess so far its probability probability of the final answer uncertainty (nats)
0 ';";\n 0.120 0.000 7.19
1 Chain 0.036 0.000 7.96
2 ';";\n 0.057 0.000 7.10
3 � 0.060 0.000 7.96
4 � 0.107 0.000 7.79
5 ';";\n 0.155 0.000 5.88
6 ';";\n 0.128 0.000 6.41
7 %";\n 0.248 0.000 4.30
8 ')));\n\n 0.077 0.000 7.09
9 ?>> 0.085 0.000 7.39
10 ","# 0.063 0.000 7.62
11 ped 0.078 0.000 7.47
12 ","# 0.079 0.000 8.08
13 ","# 0.235 0.000 6.97
14 ","# 0.388 0.000 5.36
15 ","# 0.399 0.000 5.33
16 ","# 0.209 0.000 6.98
17 ","# 0.067 0.003 8.44
18 either 0.024 0.006 8.93
19 to 0.105 0.105 7.06
20 to 0.405 0.405 4.85
21 to 0.111 0.111 7.25
22 to 0.261 0.261 5.19
23 to 0.303 0.303 4.49

Early layers usually guess generic words; the final answer 'crystallises' at some layer. Uncertainty falling = the model is making up its mind. Which layer it settles at is where the decision is made.

How much each layer adds to the running thought (last position)

layer thought size entering attention adds MLP adds real change attn vs MLP direction (cos) MLP share of the two lengths
0 0.4 0.4 4.6 4.7 0.03 91.44%
1 4.7 2.9 4.3 5.4 0.10 59.83%
2 8.2 1.6 4.3 4.4 -0.11 73.40%
3 9.1 3.6 5.4 5.7 -0.24 59.67%
4 9.8 3.3 5.3 5.5 -0.24 61.75%
5 10.8 4.1 6.0 6.3 -0.25 59.58%
6 12.6 2.5 5.3 5.5 -0.16 67.70%
7 12.9 4.7 6.1 6.6 -0.28 56.63%
8 13.8 4.7 5.4 6.5 -0.18 53.50%
9 14.6 3.2 5.7 5.8 -0.27 64.23%
10 15.9 3.9 4.3 4.8 -0.29 52.27%
11 15.4 2.2 4.3 4.5 -0.15 65.50%
12 15.3 3.0 5.8 5.8 -0.26 65.75%
13 17.8 2.3 5.4 5.3 -0.27 70.08%
14 18.7 3.8 5.0 6.1 -0.08 56.64%
15 20.1 3.2 7.6 7.9 -0.11 70.51%
16 23.0 2.5 13.1 12.8 -0.19 83.80%
17 29.1 6.1 9.9 12.4 0.16 62.14%
18 36.1 3.8 9.7 10.1 -0.08 71.79%
19 40.8 4.6 12.9 13.8 0.03 73.89%
20 45.0 7.0 20.0 21.3 0.02 74.22%
21 58.6 10.4 25.5 30.0 0.26 70.99%
22 80.5 3.1 27.9 27.7 -0.12 89.90%
23 87.0 13.3 78.6 75.8 -0.29 85.48%

'Thought size' is the length of the number-vector the model carries for the last word. Each layer adds an attention contribution and an MLP contribution to it. 'Real change' is the length of what the layer actually changed (the two contributions added as arrows, not as lengths). The cosine says whether attention and the MLP pull the same way (+1), at right angles (0) or against each other (-1). Where a steering push is applied, that layer's 'real change' includes the push.

Where the attention spotlights look (averaged over all heads in the layer, last position)

layer spread (0 = pin-point, 1 = even) first token (the sink) memory text (sink excluded) memory if attention were even previous token itself
0 0.62 7.26% 0.00% 0.00% 25.98% 24.52%
1 0.76 22.55% 0.00% 0.00% 10.44% 12.21%
2 0.75 30.88% 0.00% 0.00% 6.56% 18.30%
3 0.48 49.84% 0.00% 0.00% 2.86% 9.81%
4 0.57 43.81% 0.00% 0.00% 9.12% 16.86%
5 0.51 52.91% 0.00% 0.00% 3.83% 13.09%
6 0.54 53.86% 0.00% 0.00% 8.19% 9.75%
7 0.50 50.35% 0.00% 0.00% 1.90% 9.91%
8 0.60 34.75% 0.00% 0.00% 17.29% 13.08%
9 0.30 73.13% 0.00% 0.00% 6.31% 5.74%
10 0.55 51.46% 0.00% 0.00% 13.31% 14.25%
11 0.25 81.84% 0.00% 0.00% 5.44% 2.20%
12 0.62 50.15% 0.00% 0.00% 11.19% 14.95%
13 0.32 76.14% 0.00% 0.00% 4.61% 9.83%
14 0.56 44.60% 0.00% 0.00% 12.07% 20.45%
15 0.54 57.90% 0.00% 0.00% 10.29% 10.49%
16 0.18 88.08% 0.00% 0.00% 0.42% 1.62%
17 0.42 62.47% 0.00% 0.00% 2.43% 9.31%
18 0.51 60.98% 0.00% 0.00% 4.54% 11.65%
19 0.48 57.76% 0.00% 0.00% 5.92% 22.26%
20 0.35 76.36% 0.00% 0.00% 4.07% 6.68%
21 0.27 77.12% 0.00% 0.00% 1.31% 9.45%
22 0.67 43.15% 0.00% 0.00% 0.44% 19.24%
23 0.63 28.50% 0.00% 0.00% 0.71% 39.18%

The first token is an 'attention sink': models park spare attention on it, and it often takes most of the attention in every layer. It is NOT counted as memory here. 'Memory text' is the share landing on the memory sentence's other tokens, and 'memory if attention were even' is what it would get if every token were treated alike. Memory above that line is real interest in the memory.

Run B (test) - the prompt alone, with the memory 'direction' added to the running thought at layer 12, strength 0.5 x the typical size

What the model would say if it had to stop at each layer (logit lens)

layer best guess so far its probability probability of the final answer uncertainty (nats)
0 ';";\n 0.120 0.000 7.19
1 Chain 0.036 0.000 7.96
2 ';";\n 0.057 0.000 7.10
3 � 0.060 0.000 7.96
4 � 0.107 0.000 7.79
5 ';";\n 0.155 0.000 5.88
6 ';";\n 0.128 0.000 6.41
7 %";\n 0.248 0.000 4.30
8 ')));\n\n 0.077 0.000 7.09
9 ?>> 0.085 0.000 7.39
10 ","# 0.063 0.000 7.62
11 ped 0.078 0.000 7.47
12 ped 0.033 0.000 8.07
13 ped 0.059 0.000 8.10
14 ped 0.048 0.000 7.94
15 ped 0.037 0.000 7.86
16 堞 0.015 0.000 8.49
17 either 0.015 0.001 8.76
18 either 0.089 0.004 8.39
19 to 0.200 0.200 5.71
20 to 0.877 0.877 1.19
21 to 0.370 0.370 4.60
22 to 0.431 0.431 4.20
23 to 0.522 0.522 3.15

Early layers usually guess generic words; the final answer 'crystallises' at some layer. Uncertainty falling = the model is making up its mind. Which layer it settles at is where the decision is made.

How much each layer adds to the running thought (last position)

layer thought size entering attention adds MLP adds real change attn vs MLP direction (cos) MLP share of the two lengths
0 0.4 0.4 4.6 4.7 0.03 91.44%
1 4.7 2.9 4.3 5.4 0.10 59.83%
2 8.2 1.6 4.3 4.4 -0.11 73.40%
3 9.1 3.6 5.4 5.7 -0.24 59.67%
4 9.8 3.3 5.3 5.5 -0.24 61.75%
5 10.8 4.1 6.0 6.3 -0.25 59.58%
6 12.6 2.5 5.3 5.5 -0.16 67.70%
7 12.9 4.7 6.1 6.6 -0.28 56.63%
8 13.8 4.7 5.4 6.5 -0.18 53.50%
9 14.6 3.2 5.7 5.8 -0.27 64.23%
10 15.9 3.9 4.3 4.8 -0.29 52.27%
11 15.4 2.2 4.3 4.5 -0.15 65.50%
12 15.3 3.0 5.8 10.4 -0.26 65.75%
13 19.0 2.8 6.6 6.5 -0.26 70.25%
14 19.9 3.9 6.0 6.9 -0.08 60.73%
15 20.0 3.4 7.7 7.9 -0.15 68.99%
16 21.1 2.4 11.5 11.2 -0.21 82.96%
17 25.8 5.8 8.9 11.2 0.12 60.58%
18 31.4 3.2 8.8 9.0 -0.12 73.07%
19 35.5 4.8 12.6 13.3 -0.03 72.49%
20 39.0 8.4 20.5 22.1 -0.00 71.04%
21 52.0 11.0 23.6 27.3 0.13 68.28%
22 70.6 3.4 29.5 29.4 -0.09 89.65%
23 79.5 13.5 70.8 67.2 -0.36 83.98%

'Thought size' is the length of the number-vector the model carries for the last word. Each layer adds an attention contribution and an MLP contribution to it. 'Real change' is the length of what the layer actually changed (the two contributions added as arrows, not as lengths). The cosine says whether attention and the MLP pull the same way (+1), at right angles (0) or against each other (-1). Where a steering push is applied, that layer's 'real change' includes the push.

Where the attention spotlights look (averaged over all heads in the layer, last position)

layer spread (0 = pin-point, 1 = even) first token (the sink) memory text (sink excluded) memory if attention were even previous token itself
0 0.62 7.26% 0.00% 0.00% 25.98% 24.52%
1 0.76 22.55% 0.00% 0.00% 10.44% 12.21%
2 0.75 30.88% 0.00% 0.00% 6.56% 18.30%
3 0.48 49.84% 0.00% 0.00% 2.86% 9.81%
4 0.57 43.81% 0.00% 0.00% 9.12% 16.86%
5 0.51 52.91% 0.00% 0.00% 3.83% 13.09%
6 0.54 53.86% 0.00% 0.00% 8.19% 9.75%
7 0.50 50.35% 0.00% 0.00% 1.90% 9.91%
8 0.60 34.75% 0.00% 0.00% 17.29% 13.08%
9 0.30 73.13% 0.00% 0.00% 6.31% 5.74%
10 0.55 51.46% 0.00% 0.00% 13.31% 14.25%
11 0.25 81.84% 0.00% 0.00% 5.44% 2.20%
12 0.62 50.15% 0.00% 0.00% 11.19% 14.95%
13 0.45 63.70% 0.00% 0.00% 7.68% 13.04%
14 0.57 47.27% 0.00% 0.00% 11.78% 18.50%
15 0.59 56.08% 0.00% 0.00% 8.92% 9.93%
16 0.16 90.55% 0.00% 0.00% 0.50% 2.49%
17 0.41 67.04% 0.00% 0.00% 2.36% 10.82%
18 0.51 63.69% 0.00% 0.00% 5.41% 12.66%
19 0.51 51.39% 0.00% 0.00% 6.48% 25.98%
20 0.40 70.87% 0.00% 0.00% 4.50% 6.95%
21 0.34 72.25% 0.00% 0.00% 1.87% 8.87%
22 0.70 42.56% 0.00% 0.00% 1.13% 13.85%
23 0.69 25.48% 0.00% 0.00% 1.38% 29.15%

The first token is an 'attention sink': models park spare attention on it, and it often takes most of the attention in every layer. It is NOT counted as memory here. 'Memory text' is the share landing on the memory sentence's other tokens, and 'memory if attention were even' is what it would get if every token were treated alike. Memory above that line is real interest in the memory.

Reference versus test: what did the push change?

rank A: next word A: prob B: next word B: prob
1 to 30.35% to 52.21%
2 go 5.16% go 4.99%
3 sit 3.09% play 2.90%
4 take 3.08% take 2.24%
5 just 2.26% ride 2.05%
6 a 1.94% bike 1.85%
7 get 1.27% playing 1.36%
8 hang 1.26% hang 1.18%
9 spend 1.19% paddle 0.91%
10 catch 1.14% walk 0.81%

Words whose probability rose most because of the push

word A B change
to 30.35% 52.21% +21.86%
play 0.76% 2.90% +2.14%
bike 0.17% 1.85% +1.68%
ride 0.49% 2.05% +1.56%
playing 0.05% 1.36% +1.31%
paddle 0.09% 0.91% +0.82%
going 0.16% 0.59% +0.43%
biking 0.01% 0.38% +0.37%

Your keywords as the very next word

keyword first piece A prob A rank B prob B rank B / A
jigsaw j 0.00% 1804 0.00% 726 2.3x
puzzle puzzle 0.00% 1967 0.00% 707 2.8x

Probability and rank (1 = the model's top choice) of each keyword's first word-piece as the immediate next word. Sharper than the probe's 'recall', which only checks whether the word shows up somewhere in the generated text.

How the change travels through the layers

layer relative change in the running thought mean attention shift (0-1) most shifted head
0 0.0000 0.0000 h0 (0.000)
1 0.0000 0.0000 h0 (0.000)
2 0.0000 0.0000 h0 (0.000)
3 0.0000 0.0000 h0 (0.000)
4 0.0000 0.0000 h0 (0.000)
5 0.0000 0.0000 h0 (0.000)
6 0.0000 0.0000 h0 (0.000)
7 0.0000 0.0000 h0 (0.000)
8 0.0000 0.0000 h0 (0.000)
9 0.0000 0.0000 h0 (0.000)
10 0.0000 0.0000 h0 (0.000)
11 0.0000 0.0000 h0 (0.000)
12 0.4947 0.0000 h0 (0.000)
13 0.4962 0.1317 h4 (0.344)
14 0.4895 0.1365 h2 (0.372)
15 0.4631 0.1101 h9 (0.284)
16 0.4193 0.0521 h10 (0.311)
17 0.3742 0.1342 h9 (0.367)
18 0.3651 0.1153 h11 (0.327)
19 0.3749 0.0999 h5 (0.241)
20 0.3485 0.0904 h2 (0.253)
21 0.3051 0.0682 h1 (0.134)
22 0.3096 0.1056 h0 (0.177)
23 0.3818 0.1445 h2 (0.323)

Left column: how big the difference between the two runs is at the output of each layer, as a fraction of that layer's output size. A push that fades is small by the last layers; one that grows is being amplified. Attention shift: how far a head's pattern of looking moved (0 = identical, 1 = looks at entirely different words).

The 10 spotlights whose pattern moved most

layer head shift sink share A -> B memory share A -> B
14 2 0.372 20.46% -> 57.68% -
17 9 0.367 44.41% -> 81.09% -
17 10 0.359 18.86% -> 45.86% -
13 4 0.344 63.51% -> 29.06% -
18 11 0.327 22.13% -> 45.02% -
23 2 0.323 31.46% -> 11.31% -
13 0 0.314 53.56% -> 23.33% -
16 10 0.311 19.59% -> 38.54% -
15 9 0.284 37.21% -> 56.41% -
13 6 0.271 85.69% -> 58.59% -

Size of the running thought at every position (run B)

position token layer 6 layer 12 layer 18
0 The 1636.1 1641.1 1643.9
1 best 15.6 17.9 36.7
2 thing 15.1 17.9 37.4
3 to 13.3 16.7 34.0
4 do 14.1 17.9 35.4
5 on 13.8 18.1 32.9
6 a 13.8 18.7 34.1
7 rainy 15.4 20.8 39.0
8 weekend 15.1 20.8 36.2
9 is 12.9 19.0 35.5
first / typical 115.7x 90.6x 46.3x

Length of the number-vector the model carries for each token, at three depths. The first position is usually enormous next to the rest (see the last row). That 'massive activation' is what turns the first token into the attention sink, and it is why the probe leaves position 0 alone when it pushes.

section 'trace' took 37.7s

8. Self-checks (the report checking its own arithmetic)

result check detail
PASS attention weights of one head sum to 1 sum = 1.000000
PASS layer hooks agree with the model's own hidden states max relative error 0.00e+00
PASS layer output = input + attention + MLP (exact decomposition) max relative gap 0.00e+00 (steer layer excluded)
PASS logit-lens pipeline reproduces the model's own scores at the last layer max diff 0.00e+00
PASS measured steering push equals strength x vector length measured 8.8109 vs 8.8109

ALL CHECKS PASSED

total time 163.5s, peak memory 2842.9 MB