📝 inspect_selftest-tiny-random.mdv4.5.1 · 2026-10-02

Model inspection report: selftest-tiny-random

Generated by model_inspector v1.1.0 (2026-10-11) on 2026-10-10 22:38:12. Read-only: no weights were changed.

1. Identity and provenance

What this model is according to its own config file, and what its publisher says about how it was made. Nothing in this section is measured from the weights except where stated.

setting value plain words
model_type qwen2 family of architecture
architectures None class name the publisher used
hidden_size 64 width of the 'thought vector' kept for every word-piece
intermediate_size 128 width of the MLP inside each layer (where most 'stored knowledge' lives)
num_hidden_layers 4 how many identical processing blocks are stacked
num_attention_heads 4 parallel 'spotlights' per layer that pick which earlier words matter
num_key_value_heads 2 shared 'index cards' the spotlights read from (fewer = cheaper memory)
head_dim (computed) 16 size of each spotlight's workspace
vocab_size 256 number of distinct word-pieces (tokens) it knows
max_position_embeddings 256 longest text it was set up to handle (word-pieces)
rope_theta 10000.0 base of the rotary position code (how it knows word order); bigger = longer reach
rms_norm_eps 1e-06 tiny number that keeps the normalisation maths stable
hidden_act silu activation function inside the MLP
tie_word_embeddings True input word table is reused as the output word table
torch_dtype None number format the weights were published in
use_sliding_window False attention limited to a recent window
sliding_window 4096 size of that window
attention_dropout 0.0 random dropping during training (0 = none)
initializer_range 0.02 spread of the random numbers it started training from
bos_token_id None id of the beginning-of-text marker
eos_token_id None id of the end-of-text marker

Model class and generation defaults

Generation defaults are only used when the model writes text on its own. Both the probe and this report use greedy decoding (always the single most likely next piece).

Tokenizer

Publisher's model card (VERBATIM - not derived from the weights)

(skipped in --selftest)

This machine and software

item value
python 3.8.18
torch 1.10.0a0+git71f889c
transformers 4.37.2
tokenizers 0.15.2
safetensors 0.5.3
huggingface_hub 0.23.5
cpu Intel(R) Core(TM) i7-3520M CPU @ 2.90GHz
logical cpus 4
cpu features seen sse4_2, avx, f16c
torch threads in use 2
MemTotal 6026.7 MB
SwapTotal 2048.0 MB
SwapFree 2048.0 MB

avx2 and fma are the instruction sets that make neural-network maths fast on a CPU. If they are missing from 'cpu features seen', expect slow runs - that is the hardware, not a bug.

section 'identity' took 0.0s

2. Parameter census

A 'parameter' is one adjustable number inside the model. Everything the model 'knows' is stored in these numbers. Below, every one is counted and filed under the part of the model it belongs to.

part parameters share approx memory
embedding 16,384 9.93% 0.1 MB
attention 49,664 30.11% 0.2 MB
mlp 98,304 59.60% 0.4 MB
norms 576 0.35% 0.0 MB

embedding = the dictionary that turns word-pieces into number-vectors. attention = the 'who should I look at' machinery. mlp = the per-word 'think about it' machinery, usually the biggest block. norms = tiny stabilisers.

part tensor shape copies parameters (all copies) share
attention self_attn.k_proj.bias 32 4 128 0.08%
attention self_attn.k_proj.weight 32x64 4 8,192 4.97%
attention self_attn.o_proj.weight 64x64 4 16,384 9.93%
attention self_attn.q_proj.bias 64 4 256 0.16%
attention self_attn.q_proj.weight 64x64 4 16,384 9.93%
attention self_attn.v_proj.bias 32 4 128 0.08%
attention self_attn.v_proj.weight 32x64 4 8,192 4.97%
embedding embed_tokens.weight 256x64 1 16,384 9.93%
mlp mlp.down_proj.weight 64x128 4 32,768 19.87%
mlp mlp.gate_proj.weight 128x64 4 32,768 19.87%
mlp mlp.up_proj.weight 128x64 4 32,768 19.87%
norms input_layernorm.weight 64 4 256 0.16%
norms norm.weight 64 1 64 0.04%
norms post_attention_layernorm.weight 64 4 256 0.16%

Per layer

layer attention mlp norms layer total
0 12,416 24,576 128 37,120
1 12,416 24,576 128 37,120
2 12,416 24,576 128 37,120
3 12,416 24,576 128 37,120

All layers identical in size: yes

section 'params' took 0.0s

3. Attention anatomy

Attention is how each word-piece decides which earlier word-pieces to read. Think of each layer as having several spotlights (heads); each spotlight reads from an 'index card' set (key/value head). Fewer index-card sets shared between spotlights = much cheaper memory while generating (this design is called grouped-query attention).

quantity value plain words
layers 4 stacked blocks, each with its own attention
query heads per layer 4 spotlights
key/value heads per layer 2 index-card sets
query heads sharing one key/value head 2 spotlights per index-card set
head size 16 numbers in each spotlight's workspace
total spotlights in the model 16 layers x heads
rotary base (rope_theta) 10000.0 position code; bigger reaches farther

Shapes inside one layer (layer 0)

projection weight shape (out x in) has bias parameters
q_proj 64x64 yes 4,096
k_proj 32x64 yes 2,048
v_proj 32x64 yes 2,048
o_proj 64x64 no 4,096

q = what each spotlight is looking FOR, k = what each earlier word-piece offers to be found BY, v = what it hands over when found, o = mixes all spotlights' findings back together.

Memory the model needs per word-piece of context (the 'KV cache')

context length (tokens) KV cache at 32-bit KV cache at 16-bit
1,024 1.0 MB 0.5 MB
8,192 8.0 MB 4.0 MB
256 0.2 MB 0.1 MB

Per token: 256 numbers (1024 bytes at 32-bit). Without the shared index cards it would be 2x more.

orthogonal_probe v1.1.0 uses this cache for the modes none, text, logit and steer, so it does not re-read the whole text for every new word. attn and keyscale change what earlier words look like, so they cannot reuse it.

Attention versus MLP

Most of the model's capacity is in the MLP and the dictionary; attention is the small, fast-moving part. That matters for your idea: attention is the cheapest place to influence what the model does.

section 'attention' took 0.0s

4. Weights

The actual numbers inside the model, summarised. 'std' is the typical size of a weight; 'near zero' is the share of weights smaller than 0.001; 'abs max' is the largest weight. Healthy models have small typical weights and a few large outliers.

tensor type (N = layer) copies parameters mean avg std std range abs max near zero
model.embed_tokens.weight 1 16,384 -0.00029 0.0200 0.0200 - 0.0200 0.085 4.17%
model.layers.N.input_layernorm.weight 4 256 1.00000 0.0000 0.0000 - 0.0000 1.000 0.00%
model.layers.N.mlp.down_proj.weight 4 32,768 -0.00018 0.0200 0.0198 - 0.0203 0.081 4.09%
model.layers.N.mlp.gate_proj.weight 4 32,768 -0.00012 0.0200 0.0199 - 0.0200 0.092 4.04%
model.layers.N.mlp.up_proj.weight 4 32,768 -0.00018 0.0200 0.0199 - 0.0201 0.077 3.95%
model.layers.N.post_attention_layernorm.weight 4 256 1.00000 0.0000 0.0000 - 0.0000 1.000 0.00%
model.layers.N.self_attn.k_proj.bias 4 128 0.00000 0.0000 0.0000 - 0.0000 0.000 100.00%
model.layers.N.self_attn.k_proj.weight 4 8,192 -0.00013 0.0201 0.0199 - 0.0203 0.076 4.30%
model.layers.N.self_attn.o_proj.weight 4 16,384 0.00010 0.0202 0.0200 - 0.0205 0.093 3.60%
model.layers.N.self_attn.q_proj.bias 4 256 0.00000 0.0000 0.0000 - 0.0000 0.000 100.00%
model.layers.N.self_attn.q_proj.weight 4 16,384 0.00013 0.0199 0.0197 - 0.0200 0.083 4.10%
model.layers.N.self_attn.v_proj.bias 4 128 0.00000 0.0000 0.0000 - 0.0000 0.000 100.00%
model.layers.N.self_attn.v_proj.weight 4 8,192 0.00003 0.0199 0.0198 - 0.0201 0.082 4.33%
model.norm.weight 1 64 1.00000 0.0000 0.0000 - 0.0000 1.000 0.00%

Non-finite values (NaN / infinity) found: 0

Weight size per layer (Frobenius norm = overall magnitude of each matrix)

layer q_proj k_proj v_proj o_proj gate_proj up_proj down_proj
0 1.3 0.9 0.9 1.3 1.8 1.8 1.8
1 1.3 0.9 0.9 1.3 1.8 1.8 1.8
2 1.3 0.9 0.9 1.3 1.8 1.8 1.8
3 1.3 0.9 0.9 1.3 1.8 1.8 1.8

Shows where the weight 'energy' sits across depth. Big changes between neighbouring layers mark layers that behave differently from the rest.

Fingerprint of all weights

sha256 over every weight value (as 32-bit numbers, in name order): bf4c81af2ad02cfbe68c6a73885107524fed3148996a4689172ef448eb0664b7

Same fingerprint on Clone and Live = bit-for-bit the same model. Any difference means a different file, version, or a damaged download. The .json report also holds a sha256 for each of the 50 tensors: run --compare a.json b.json to see exactly which ones differ.

section 'weights' took 0.1s

5. What the model does with your prompt

Two runs of the same prompt. A (reference) = the prompt alone. B (test) = the memory pasted as plain text in front of the prompt. Everything below is measured from inside the model on those two runs - no guessing.

Run A (reference) - the prompt alone

What the model would say if it had to stop at each layer (logit lens)

layer best guess so far its probability probability of the final answer uncertainty (nats)
0 e 0.010 0.010 5.53
1 e 0.009 0.009 5.53
2 e 0.008 0.008 5.53
3 e 0.007 0.007 5.53

Early layers usually guess generic words; the final answer 'crystallises' at some layer. Uncertainty falling = the model is making up its mind. Which layer it settles at is where the decision is made.

How much each layer adds to the running thought (last position)

layer thought size entering attention adds MLP adds real change attn vs MLP direction (cos) MLP share of the two lengths
0 0.2 0.1 0.0 0.1 0.23 19.65%
1 0.2 0.1 0.0 0.1 -0.10 14.09%
2 0.2 0.1 0.0 0.1 -0.07 18.61%
3 0.2 0.2 0.0 0.2 -0.23 10.38%

'Thought size' is the length of the number-vector the model carries for the last word. Each layer adds an attention contribution and an MLP contribution to it. 'Real change' is the length of what the layer actually changed (the two contributions added as arrows, not as lengths). The cosine says whether attention and the MLP pull the same way (+1), at right angles (0) or against each other (-1). Where a steering push is applied, that layer's 'real change' includes the push.

Where the attention spotlights look (averaged over all heads in the layer, last position)

layer spread (0 = pin-point, 1 = even) first token (the sink) memory text (sink excluded) memory if attention were even previous token itself
0 1.00 9.23% 0.00% 0.00% 9.13% 9.05%
1 1.00 9.03% 0.00% 0.00% 8.99% 9.34%
2 1.00 9.30% 0.00% 0.00% 9.09% 8.88%
3 1.00 9.09% 0.00% 0.00% 8.99% 9.04%

The first token is an 'attention sink': models park spare attention on it, and it often takes most of the attention in every layer. It is NOT counted as memory here. 'Memory text' is the share landing on the memory sentence's other tokens, and 'memory if attention were even' is what it would get if every token were treated alike. Memory above that line is real interest in the memory.

Run B (test) - the memory pasted as plain text in front of the prompt

What the model would say if it had to stop at each layer (logit lens)

layer best guess so far its probability probability of the final answer uncertainty (nats)
0 e 0.011 0.011 5.53
1 e 0.009 0.009 5.53
2 e 0.008 0.008 5.53
3 e 0.007 0.007 5.53

Early layers usually guess generic words; the final answer 'crystallises' at some layer. Uncertainty falling = the model is making up its mind. Which layer it settles at is where the decision is made.

How much each layer adds to the running thought (last position)

layer thought size entering attention adds MLP adds real change attn vs MLP direction (cos) MLP share of the two lengths
0 0.2 0.1 0.0 0.1 0.20 22.19%
1 0.2 0.1 0.0 0.1 -0.17 19.04%
2 0.2 0.1 0.0 0.1 0.02 13.89%
3 0.2 0.1 0.0 0.1 -0.05 12.02%

'Thought size' is the length of the number-vector the model carries for the last word. Each layer adds an attention contribution and an MLP contribution to it. 'Real change' is the length of what the layer actually changed (the two contributions added as arrows, not as lengths). The cosine says whether attention and the MLP pull the same way (+1), at right angles (0) or against each other (-1). Where a steering push is applied, that layer's 'real change' includes the push.

Where the attention spotlights look (averaged over all heads in the layer, last position)

layer spread (0 = pin-point, 1 = even) first token (the sink) memory text (sink excluded) memory if attention were even previous token itself
0 1.00 5.54% 27.67% 27.78% 5.60% 5.55%
1 1.00 5.40% 27.70% 27.78% 5.54% 5.74%
2 1.00 5.52% 27.74% 27.78% 5.55% 5.43%
3 1.00 5.46% 27.48% 27.78% 5.53% 5.59%

The first token is an 'attention sink': models park spare attention on it, and it often takes most of the attention in every layer. It is NOT counted as memory here. 'Memory text' is the share landing on the memory sentence's other tokens, and 'memory if attention were even' is what it would get if every token were treated alike. Memory above that line is real interest in the memory.

The 10 spotlights that read the memory text most, compared with an even spread

rank layer head share of its attention on memory x even spread spread
1 2 3 27.97% 1.0x 1.00
2 3 0 27.94% 1.0x 1.00
3 0 3 27.87% 1.0x 1.00
4 2 1 27.86% 1.0x 1.00
5 1 3 27.81% 1.0x 1.00
6 1 0 27.78% 1.0x 1.00
7 0 1 27.73% 1.0x 1.00
8 1 2 27.64% 1.0x 1.00
9 0 2 27.62% 1.0x 1.00
10 2 2 27.58% 1.0x 1.00

Reference versus test: what did the push change?

rank A: next word A: prob B: next word B: prob
1 e 0.70% e 0.70%
2  0.61% H 0.58%
3 � 0.60% m 0.58%
4 E 0.58% � 0.55%
5 / 0.54% S 0.55%
6 ! 0.53% � 0.55%
7 � 0.53% - 0.53%
8 � 0.52% � 0.52%
9 � 0.52% � 0.52%
10 � 0.52% � 0.51%

Words whose probability rose most because of the push

word A B change
 0.35% 0.51% +0.17%
2 0.34% 0.50% +0.15%
� 0.26% 0.41% +0.14%
� 0.33% 0.45% +0.12%
K 0.36% 0.48% +0.12%
m 0.47% 0.58% +0.11%
( 0.36% 0.47% +0.11%
P 0.26% 0.37% +0.11%

Your keywords as the very next word

keyword first piece A prob A rank B prob B rank B / A
puzzle 0.43% 66 0.42% 92 1.0x

Probability and rank (1 = the model's top choice) of each keyword's first word-piece as the immediate next word. Sharper than the probe's 'recall', which only checks whether the word shows up somewhere in the generated text.

How the change travels through the layers

(the two runs have different lengths - 11 vs 18 tokens - so layer-by-layer differences are not defined)

Size of the running thought at every position (run B)

position token layer 1 layer 2
0 p 0.4 0.4
1 u 0.3 0.4
2 z 0.2 0.3
3 z 0.2 0.3
4 l 0.2 0.3
5 e 0.2 0.3
6 \n 0.2 0.3
7 H 0.2 0.3
8 e 0.2 0.3
9 l 0.2 0.3
10 l 0.2 0.3
11 o 0.2 0.3
12 0.2 0.3
13 t 0.2 0.2
14 h 0.2 0.2
15 e 0.2 0.2
16 r 0.2 0.2
17 e 0.2 0.2
first / typical 1.8x 1.6x

Length of the number-vector the model carries for each token, at three depths. The first position is usually enormous next to the rest (see the last row). That 'massive activation' is what turns the first token into the attention sink, and it is why the probe leaves position 0 alone when it pushes.

section 'trace' took 0.2s

8. Self-checks (the report checking its own arithmetic)

result check detail
PASS census parts add up to the model's own parameter count 164,928 vs 164,928
PASS census equals independent count from config numbers config formula 164,928 vs counted 164,928
PASS all weights are finite numbers 0 bad values
PASS attention weights of one head sum to 1 sum = 1.000000
PASS layer hooks agree with the model's own hidden states max relative error 0.00e+00
PASS layer output = input + attention + MLP (exact decomposition) max relative gap 0.00e+00
PASS logit-lens pipeline reproduces the model's own scores at the last layer max diff 0.00e+00

ALL CHECKS PASSED

total time 3.3s, peak memory 134.7 MB