Generated by model_inspector v1.1.0 (2026-10-11) on 2026-10-10 22:38:12. Read-only: no weights were changed.
What this model is according to its own config file, and what its publisher says about how it was made. Nothing in this section is measured from the weights except where stated.
| setting | value | plain words |
|---|---|---|
| model_type | qwen2 | family of architecture |
| architectures | None | class name the publisher used |
| hidden_size | 64 | width of the 'thought vector' kept for every word-piece |
| intermediate_size | 128 | width of the MLP inside each layer (where most 'stored knowledge' lives) |
| num_hidden_layers | 4 | how many identical processing blocks are stacked |
| num_attention_heads | 4 | parallel 'spotlights' per layer that pick which earlier words matter |
| num_key_value_heads | 2 | shared 'index cards' the spotlights read from (fewer = cheaper memory) |
| head_dim (computed) | 16 | size of each spotlight's workspace |
| vocab_size | 256 | number of distinct word-pieces (tokens) it knows |
| max_position_embeddings | 256 | longest text it was set up to handle (word-pieces) |
| rope_theta | 10000.0 | base of the rotary position code (how it knows word order); bigger = longer reach |
| rms_norm_eps | 1e-06 | tiny number that keeps the normalisation maths stable |
| hidden_act | silu | activation function inside the MLP |
| tie_word_embeddings | True | input word table is reused as the output word table |
| torch_dtype | None | number format the weights were published in |
| use_sliding_window | False | attention limited to a recent window |
| sliding_window | 4096 | size of that window |
| attention_dropout | 0.0 | random dropping during training (0 = none) |
| initializer_range | 0.02 | spread of the random numbers it started training from |
| bos_token_id | None | id of the beginning-of-text marker |
| eos_token_id | None | id of the end-of-text marker |
Qwen2ForCausalLM_from_model_config = Truebad_words_ids = Nonebegin_suppress_tokens = Nonebos_token_id = Noneconstraints = Nonedecoder_start_token_id = Nonediversity_penalty = 0.0do_sample = Falseearly_stopping = Falseencoder_no_repeat_ngram_size = 0encoder_repetition_penalty = 1.0eos_token_id = Noneepsilon_cutoff = 0.0eta_cutoff = 0.0exponential_decay_length_penalty = Noneforce_words_ids = Noneforced_bos_token_id = Noneforced_decoder_ids = Noneforced_eos_token_id = Nonegeneration_kwargs = {}guidance_scale = Nonelength_penalty = 1.0low_memory = Nonemax_length = 20max_new_tokens = Nonemax_time = Nonemin_length = 0min_new_tokens = Noneno_repeat_ngram_size = 0num_assistant_tokens = 5num_assistant_tokens_schedule = heuristicnum_beam_groups = 1num_beams = 1num_return_sequences = 1output_attentions = Falseoutput_hidden_states = Falseoutput_scores = Falsepad_token_id = Nonepenalty_alpha = Noneprompt_lookup_num_tokens = Noneremove_invalid_values = Falserenormalize_logits = Falserepetition_penalty = 1.0return_dict_in_generate = Falsesequence_bias = Nonesuppress_tokens = Nonetemperature = 1.0top_k = 50top_p = 1.0typical_p = 1.0use_cache = TrueGeneration defaults are only used when the model writes text on its own. Both the probe and this report use greedy decoding (always the single most likely next piece).
(skipped in --selftest)
| item | value |
|---|---|
| python | 3.8.18 |
| torch | 1.10.0a0+git71f889c |
| transformers | 4.37.2 |
| tokenizers | 0.15.2 |
| safetensors | 0.5.3 |
| huggingface_hub | 0.23.5 |
| cpu | Intel(R) Core(TM) i7-3520M CPU @ 2.90GHz |
| logical cpus | 4 |
| cpu features seen | sse4_2, avx, f16c |
| torch threads in use | 2 |
| MemTotal | 6026.7 MB |
| SwapTotal | 2048.0 MB |
| SwapFree | 2048.0 MB |
avx2 and fma are the instruction sets that make neural-network maths fast on a CPU. If they are missing from 'cpu features seen', expect slow runs - that is the hardware, not a bug.
section 'identity' took 0.0s
A 'parameter' is one adjustable number inside the model. Everything the model 'knows' is stored in these numbers. Below, every one is counted and filed under the part of the model it belongs to.
| part | parameters | share | approx memory |
|---|---|---|---|
| embedding | 16,384 | 9.93% | 0.1 MB |
| attention | 49,664 | 30.11% | 0.2 MB |
| mlp | 98,304 | 59.60% | 0.4 MB |
| norms | 576 | 0.35% | 0.0 MB |
embedding = the dictionary that turns word-pieces into number-vectors. attention = the 'who should I look at' machinery. mlp = the per-word 'think about it' machinery, usually the biggest block. norms = tiny stabilisers.
| part | tensor | shape | copies | parameters (all copies) | share |
|---|---|---|---|---|---|
| attention | self_attn.k_proj.bias | 32 | 4 | 128 | 0.08% |
| attention | self_attn.k_proj.weight | 32x64 | 4 | 8,192 | 4.97% |
| attention | self_attn.o_proj.weight | 64x64 | 4 | 16,384 | 9.93% |
| attention | self_attn.q_proj.bias | 64 | 4 | 256 | 0.16% |
| attention | self_attn.q_proj.weight | 64x64 | 4 | 16,384 | 9.93% |
| attention | self_attn.v_proj.bias | 32 | 4 | 128 | 0.08% |
| attention | self_attn.v_proj.weight | 32x64 | 4 | 8,192 | 4.97% |
| embedding | embed_tokens.weight | 256x64 | 1 | 16,384 | 9.93% |
| mlp | mlp.down_proj.weight | 64x128 | 4 | 32,768 | 19.87% |
| mlp | mlp.gate_proj.weight | 128x64 | 4 | 32,768 | 19.87% |
| mlp | mlp.up_proj.weight | 128x64 | 4 | 32,768 | 19.87% |
| norms | input_layernorm.weight | 64 | 4 | 256 | 0.16% |
| norms | norm.weight | 64 | 1 | 64 | 0.04% |
| norms | post_attention_layernorm.weight | 64 | 4 | 256 | 0.16% |
| layer | attention | mlp | norms | layer total |
|---|---|---|---|---|
| 0 | 12,416 | 24,576 | 128 | 37,120 |
| 1 | 12,416 | 24,576 | 128 | 37,120 |
| 2 | 12,416 | 24,576 | 128 | 37,120 |
| 3 | 12,416 | 24,576 | 128 | 37,120 |
All layers identical in size: yes
section 'params' took 0.0s
Attention is how each word-piece decides which earlier word-pieces to read. Think of each layer as having several spotlights (heads); each spotlight reads from an 'index card' set (key/value head). Fewer index-card sets shared between spotlights = much cheaper memory while generating (this design is called grouped-query attention).
| quantity | value | plain words |
|---|---|---|
| layers | 4 | stacked blocks, each with its own attention |
| query heads per layer | 4 | spotlights |
| key/value heads per layer | 2 | index-card sets |
| query heads sharing one key/value head | 2 | spotlights per index-card set |
| head size | 16 | numbers in each spotlight's workspace |
| total spotlights in the model | 16 | layers x heads |
| rotary base (rope_theta) | 10000.0 | position code; bigger reaches farther |
| projection | weight shape (out x in) | has bias | parameters |
|---|---|---|---|
| q_proj | 64x64 | yes | 4,096 |
| k_proj | 32x64 | yes | 2,048 |
| v_proj | 32x64 | yes | 2,048 |
| o_proj | 64x64 | no | 4,096 |
q = what each spotlight is looking FOR, k = what each earlier word-piece offers to be found BY, v = what it hands over when found, o = mixes all spotlights' findings back together.
| context length (tokens) | KV cache at 32-bit | KV cache at 16-bit |
|---|---|---|
| 1,024 | 1.0 MB | 0.5 MB |
| 8,192 | 8.0 MB | 4.0 MB |
| 256 | 0.2 MB | 0.1 MB |
Per token: 256 numbers (1024 bytes at 32-bit). Without the shared index cards it would be 2x more.
orthogonal_probe v1.1.0 uses this cache for the modes none, text, logit and steer, so it does not re-read the whole text for every new word. attn and keyscale change what earlier words look like, so they cannot reuse it.
Most of the model's capacity is in the MLP and the dictionary; attention is the small, fast-moving part. That matters for your idea: attention is the cheapest place to influence what the model does.
section 'attention' took 0.0s
The actual numbers inside the model, summarised. 'std' is the typical size of a weight; 'near zero' is the share of weights smaller than 0.001; 'abs max' is the largest weight. Healthy models have small typical weights and a few large outliers.
| tensor type (N = layer) | copies | parameters | mean | avg std | std range | abs max | near zero |
|---|---|---|---|---|---|---|---|
| model.embed_tokens.weight | 1 | 16,384 | -0.00029 | 0.0200 | 0.0200 - 0.0200 | 0.085 | 4.17% |
| model.layers.N.input_layernorm.weight | 4 | 256 | 1.00000 | 0.0000 | 0.0000 - 0.0000 | 1.000 | 0.00% |
| model.layers.N.mlp.down_proj.weight | 4 | 32,768 | -0.00018 | 0.0200 | 0.0198 - 0.0203 | 0.081 | 4.09% |
| model.layers.N.mlp.gate_proj.weight | 4 | 32,768 | -0.00012 | 0.0200 | 0.0199 - 0.0200 | 0.092 | 4.04% |
| model.layers.N.mlp.up_proj.weight | 4 | 32,768 | -0.00018 | 0.0200 | 0.0199 - 0.0201 | 0.077 | 3.95% |
| model.layers.N.post_attention_layernorm.weight | 4 | 256 | 1.00000 | 0.0000 | 0.0000 - 0.0000 | 1.000 | 0.00% |
| model.layers.N.self_attn.k_proj.bias | 4 | 128 | 0.00000 | 0.0000 | 0.0000 - 0.0000 | 0.000 | 100.00% |
| model.layers.N.self_attn.k_proj.weight | 4 | 8,192 | -0.00013 | 0.0201 | 0.0199 - 0.0203 | 0.076 | 4.30% |
| model.layers.N.self_attn.o_proj.weight | 4 | 16,384 | 0.00010 | 0.0202 | 0.0200 - 0.0205 | 0.093 | 3.60% |
| model.layers.N.self_attn.q_proj.bias | 4 | 256 | 0.00000 | 0.0000 | 0.0000 - 0.0000 | 0.000 | 100.00% |
| model.layers.N.self_attn.q_proj.weight | 4 | 16,384 | 0.00013 | 0.0199 | 0.0197 - 0.0200 | 0.083 | 4.10% |
| model.layers.N.self_attn.v_proj.bias | 4 | 128 | 0.00000 | 0.0000 | 0.0000 - 0.0000 | 0.000 | 100.00% |
| model.layers.N.self_attn.v_proj.weight | 4 | 8,192 | 0.00003 | 0.0199 | 0.0198 - 0.0201 | 0.082 | 4.33% |
| model.norm.weight | 1 | 64 | 1.00000 | 0.0000 | 0.0000 - 0.0000 | 1.000 | 0.00% |
Non-finite values (NaN / infinity) found: 0
| layer | q_proj | k_proj | v_proj | o_proj | gate_proj | up_proj | down_proj |
|---|---|---|---|---|---|---|---|
| 0 | 1.3 | 0.9 | 0.9 | 1.3 | 1.8 | 1.8 | 1.8 |
| 1 | 1.3 | 0.9 | 0.9 | 1.3 | 1.8 | 1.8 | 1.8 |
| 2 | 1.3 | 0.9 | 0.9 | 1.3 | 1.8 | 1.8 | 1.8 |
| 3 | 1.3 | 0.9 | 0.9 | 1.3 | 1.8 | 1.8 | 1.8 |
Shows where the weight 'energy' sits across depth. Big changes between neighbouring layers mark layers that behave differently from the rest.
sha256 over every weight value (as 32-bit numbers, in name order): bf4c81af2ad02cfbe68c6a73885107524fed3148996a4689172ef448eb0664b7
Same fingerprint on Clone and Live = bit-for-bit the same model. Any difference means a different file, version, or a damaged download. The .json report also holds a sha256 for each of the 50 tensors: run
--compare a.json b.jsonto see exactly which ones differ.
section 'weights' took 0.1s
Two runs of the same prompt. A (reference) = the prompt alone. B (test) = the memory pasted as plain text in front of the prompt. Everything below is measured from inside the model on those two runs - no guessing.
Hello there (11 tokens)puzzle (6 tokens)What the model would say if it had to stop at each layer (logit lens)
| layer | best guess so far | its probability | probability of the final answer | uncertainty (nats) |
|---|---|---|---|---|
| 0 | e |
0.010 | 0.010 | 5.53 |
| 1 | e |
0.009 | 0.009 | 5.53 |
| 2 | e |
0.008 | 0.008 | 5.53 |
| 3 | e |
0.007 | 0.007 | 5.53 |
Early layers usually guess generic words; the final answer 'crystallises' at some layer. Uncertainty falling = the model is making up its mind. Which layer it settles at is where the decision is made.
How much each layer adds to the running thought (last position)
| layer | thought size entering | attention adds | MLP adds | real change | attn vs MLP direction (cos) | MLP share of the two lengths |
|---|---|---|---|---|---|---|
| 0 | 0.2 | 0.1 | 0.0 | 0.1 | 0.23 | 19.65% |
| 1 | 0.2 | 0.1 | 0.0 | 0.1 | -0.10 | 14.09% |
| 2 | 0.2 | 0.1 | 0.0 | 0.1 | -0.07 | 18.61% |
| 3 | 0.2 | 0.2 | 0.0 | 0.2 | -0.23 | 10.38% |
'Thought size' is the length of the number-vector the model carries for the last word. Each layer adds an attention contribution and an MLP contribution to it. 'Real change' is the length of what the layer actually changed (the two contributions added as arrows, not as lengths). The cosine says whether attention and the MLP pull the same way (+1), at right angles (0) or against each other (-1). Where a steering push is applied, that layer's 'real change' includes the push.
Where the attention spotlights look (averaged over all heads in the layer, last position)
| layer | spread (0 = pin-point, 1 = even) | first token (the sink) | memory text (sink excluded) | memory if attention were even | previous token | itself |
|---|---|---|---|---|---|---|
| 0 | 1.00 | 9.23% | 0.00% | 0.00% | 9.13% | 9.05% |
| 1 | 1.00 | 9.03% | 0.00% | 0.00% | 8.99% | 9.34% |
| 2 | 1.00 | 9.30% | 0.00% | 0.00% | 9.09% | 8.88% |
| 3 | 1.00 | 9.09% | 0.00% | 0.00% | 8.99% | 9.04% |
The first token is an 'attention sink': models park spare attention on it, and it often takes most of the attention in every layer. It is NOT counted as memory here. 'Memory text' is the share landing on the memory sentence's other tokens, and 'memory if attention were even' is what it would get if every token were treated alike. Memory above that line is real interest in the memory.
What the model would say if it had to stop at each layer (logit lens)
| layer | best guess so far | its probability | probability of the final answer | uncertainty (nats) |
|---|---|---|---|---|
| 0 | e |
0.011 | 0.011 | 5.53 |
| 1 | e |
0.009 | 0.009 | 5.53 |
| 2 | e |
0.008 | 0.008 | 5.53 |
| 3 | e |
0.007 | 0.007 | 5.53 |
Early layers usually guess generic words; the final answer 'crystallises' at some layer. Uncertainty falling = the model is making up its mind. Which layer it settles at is where the decision is made.
How much each layer adds to the running thought (last position)
| layer | thought size entering | attention adds | MLP adds | real change | attn vs MLP direction (cos) | MLP share of the two lengths |
|---|---|---|---|---|---|---|
| 0 | 0.2 | 0.1 | 0.0 | 0.1 | 0.20 | 22.19% |
| 1 | 0.2 | 0.1 | 0.0 | 0.1 | -0.17 | 19.04% |
| 2 | 0.2 | 0.1 | 0.0 | 0.1 | 0.02 | 13.89% |
| 3 | 0.2 | 0.1 | 0.0 | 0.1 | -0.05 | 12.02% |
'Thought size' is the length of the number-vector the model carries for the last word. Each layer adds an attention contribution and an MLP contribution to it. 'Real change' is the length of what the layer actually changed (the two contributions added as arrows, not as lengths). The cosine says whether attention and the MLP pull the same way (+1), at right angles (0) or against each other (-1). Where a steering push is applied, that layer's 'real change' includes the push.
Where the attention spotlights look (averaged over all heads in the layer, last position)
| layer | spread (0 = pin-point, 1 = even) | first token (the sink) | memory text (sink excluded) | memory if attention were even | previous token | itself |
|---|---|---|---|---|---|---|
| 0 | 1.00 | 5.54% | 27.67% | 27.78% | 5.60% | 5.55% |
| 1 | 1.00 | 5.40% | 27.70% | 27.78% | 5.54% | 5.74% |
| 2 | 1.00 | 5.52% | 27.74% | 27.78% | 5.55% | 5.43% |
| 3 | 1.00 | 5.46% | 27.48% | 27.78% | 5.53% | 5.59% |
The first token is an 'attention sink': models park spare attention on it, and it often takes most of the attention in every layer. It is NOT counted as memory here. 'Memory text' is the share landing on the memory sentence's other tokens, and 'memory if attention were even' is what it would get if every token were treated alike. Memory above that line is real interest in the memory.
The 10 spotlights that read the memory text most, compared with an even spread
| rank | layer | head | share of its attention on memory | x even spread | spread |
|---|---|---|---|---|---|
| 1 | 2 | 3 | 27.97% | 1.0x | 1.00 |
| 2 | 3 | 0 | 27.94% | 1.0x | 1.00 |
| 3 | 0 | 3 | 27.87% | 1.0x | 1.00 |
| 4 | 2 | 1 | 27.86% | 1.0x | 1.00 |
| 5 | 1 | 3 | 27.81% | 1.0x | 1.00 |
| 6 | 1 | 0 | 27.78% | 1.0x | 1.00 |
| 7 | 0 | 1 | 27.73% | 1.0x | 1.00 |
| 8 | 1 | 2 | 27.64% | 1.0x | 1.00 |
| 9 | 0 | 2 | 27.62% | 1.0x | 1.00 |
| 10 | 2 | 2 | 27.58% | 1.0x | 1.00 |
reference (A): the prompt alone
test (B): the memory pasted as plain text in front of the prompt
top choice in A: e (0.70%)
top choice in B: e (0.70%)
top choice changed: no
probability moved to different next words: 5.62% (0% = nothing changed, 100% = completely different)
KL divergence B from A: 0.0099 nats (0 = identical odds; bigger = bigger change)
| rank | A: next word | A: prob | B: next word | B: prob |
|---|---|---|---|---|
| 1 | e |
0.70% | e |
0.70% |
| 2 | |
0.61% | H |
0.58% |
| 3 | � |
0.60% | m |
0.58% |
| 4 | E |
0.58% | � |
0.55% |
| 5 | / |
0.54% | S |
0.55% |
| 6 | ! |
0.53% | � |
0.55% |
| 7 | � |
0.53% | - |
0.53% |
| 8 | � |
0.52% | � |
0.52% |
| 9 | � |
0.52% | � |
0.52% |
| 10 | � |
0.52% | � |
0.51% |
Words whose probability rose most because of the push
| word | A | B | change |
|---|---|---|---|
|
0.35% | 0.51% | +0.17% |
2 |
0.34% | 0.50% | +0.15% |
� |
0.26% | 0.41% | +0.14% |
� |
0.33% | 0.45% | +0.12% |
K |
0.36% | 0.48% | +0.12% |
m |
0.47% | 0.58% | +0.11% |
( |
0.36% | 0.47% | +0.11% |
P |
0.26% | 0.37% | +0.11% |
Your keywords as the very next word
| keyword | first piece | A prob | A rank | B prob | B rank | B / A |
|---|---|---|---|---|---|---|
| puzzle | |
0.43% | 66 | 0.42% | 92 | 1.0x |
Probability and rank (1 = the model's top choice) of each keyword's first word-piece as the immediate next word. Sharper than the probe's 'recall', which only checks whether the word shows up somewhere in the generated text.
(the two runs have different lengths - 11 vs 18 tokens - so layer-by-layer differences are not defined)
| position | token | layer 1 | layer 2 |
|---|---|---|---|
| 0 | p |
0.4 | 0.4 |
| 1 | u |
0.3 | 0.4 |
| 2 | z |
0.2 | 0.3 |
| 3 | z |
0.2 | 0.3 |
| 4 | l |
0.2 | 0.3 |
| 5 | e |
0.2 | 0.3 |
| 6 | \n |
0.2 | 0.3 |
| 7 | H |
0.2 | 0.3 |
| 8 | e |
0.2 | 0.3 |
| 9 | l |
0.2 | 0.3 |
| 10 | l |
0.2 | 0.3 |
| 11 | o |
0.2 | 0.3 |
| 12 | |
0.2 | 0.3 |
| 13 | t |
0.2 | 0.2 |
| 14 | h |
0.2 | 0.2 |
| 15 | e |
0.2 | 0.2 |
| 16 | r |
0.2 | 0.2 |
| 17 | e |
0.2 | 0.2 |
| first / typical | 1.8x | 1.6x |
Length of the number-vector the model carries for each token, at three depths. The first position is usually enormous next to the rest (see the last row). That 'massive activation' is what turns the first token into the attention sink, and it is why the probe leaves position 0 alone when it pushes.
section 'trace' took 0.2s
| result | check | detail |
|---|---|---|
| PASS | census parts add up to the model's own parameter count | 164,928 vs 164,928 |
| PASS | census equals independent count from config numbers | config formula 164,928 vs counted 164,928 |
| PASS | all weights are finite numbers | 0 bad values |
| PASS | attention weights of one head sum to 1 | sum = 1.000000 |
| PASS | layer hooks agree with the model's own hidden states | max relative error 0.00e+00 |
| PASS | layer output = input + attention + MLP (exact decomposition) | max relative gap 0.00e+00 |
| PASS | logit-lens pipeline reproduces the model's own scores at the last layer | max diff 0.00e+00 |
ALL CHECKS PASSED
total time 3.3s, peak memory 134.7 MB