For the white-box candidate-screening design, metrics, commands, and evidence
limits, see WHITEBOX_SCREENING.md. For the complete
execution sequence, see RUNBOOK.md.
Run these commands from the repository root. Install the Hugging Face CLI:
python3 -m pip install --upgrade huggingface_hubThe two organism repositories are gated. Request access to
Alamerton/sl-organism-a-7b and Alamerton/sl-organism-b-7b on Hugging Face,
then authenticate:
hf auth loginCreate the local model directory and download the baseline, organism A, and organism B into the paths used by the investigation configuration:
mkdir -p models
hf download Qwen/Qwen2.5-7B-Instruct \
--local-dir models/qwen2.5-7b-instruct
hf download Alamerton/sl-organism-a-7b \
--local-dir models/sl-organism-a-7b
hf download Alamerton/sl-organism-b-7b \
--local-dir models/sl-organism-b-7bThe resulting directory layout should be:
models/
|-- qwen2.5-7b-instruct/
|-- sl-organism-a-7b/
`-- sl-organism-b-7b/
Each directory should contain model.safetensors.index.json and all model
shards named in that index. The models/ directory is ignored by Git.
The two stages test the same hypothesis in different ways:
run_investigation.pygenerates model responses and looks for behavioral differences between the baseline and the fine-tuned organisms.- A researcher uses those responses to choose plausible principals and
scenario families for closer inspection. This selection is currently
manual; the behavioral JSONL does not automatically produce
PRINCIPALS. loyalty_sweep.pydefines the fixed comparison matrix used by the SAE experiment: 23 principals multiplied by 18 single-turn scenarios.sae_diff.py encodeperforms forward passes only. It does not generate answers. It records a selected residual-stream layer, applies the same baseline-trained SAE dictionary, and sums each feature's activation over the prompt tokens.sae_diff.py analyzesubtracts the baseline feature matrix from the organism feature matrix and asks whether any feature changes selectively for one principal rather than for all names.
The baseline encoder and organism encoder therefore perform the same operation:
matched prompt
-> model residual stream at the selected layer
-> one shared SAE encoder
-> token-summed sparse feature vector
They differ only in which language-model checkpoint produces the residual stream. A feature that moves for every principal is ordinary fine-tuning drift. The loyalty hypothesis requires a feature whose organism-minus-baseline change is unusually large for one principal and stable across relevant scenarios.
The SAE experiment applies one dictionary trained on the stock Qwen2.5-7B-Instruct model to both the baseline and a fine-tuned organism. It uses matched prompts so the resulting feature difference is attributable to the fine-tune rather than to different dictionaries.
The runner requires an NVIDIA GPU with CUDA. From the repository root, create the Python environment and install the dependencies:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txtEnsure that the installed PyTorch build supports the host's CUDA version:
python -c 'import torch; assert torch.cuda.is_available(), "CUDA unavailable"; print(torch.cuda.get_device_name(0))'Confirm the shared prompt matrix:
python loyalty_sweep.pyExpected output:
23 principals x 18 scenarios = 414 matched prompts
The upstream andyrdt/saes-qwen2.5-7b-instruct repository is about 105 GB.
Do not download its entire snapshot merely to list the files. Query the file
metadata instead:
python - <<'PY'
from huggingface_hub import HfApi
repo = "andyrdt/saes-qwen2.5-7b-instruct"
for path in HfApi().list_repo_files(repo):
print(path)
PYThe upstream repository provides layers 3, 7, 11, 15, 19, 23, and 27. Within
each layer, trainer_0, trainer_1, trainer_2, and trainer_3 correspond
to k=32, k=64, k=128, and k=256, respectively.
The upstream checkpoints are named ae.pt. The current minimal loader expects
a converted .safetensors checkpoint. Before running encode, confirm that
the converted file exposes compatible encoder, decoder, bias, and optional
threshold tensors. A converted checkpoint's size can be several gigabytes;
do not confuse it with the much smaller encoded prompt-feature output.
The examples below use layer 15 and k=64:
LAYER=15
K=64python sae_diff.py encode \
--model models/qwen2.5-7b-instruct \
--tag baseline \
--layer "$LAYER" \
--k "$K"This writes feats_baseline.npz.
The analyzer expects the comparison model to use the tag organism:
python sae_diff.py encode \
--model models/sl-organism-a-7b \
--tag organism \
--layer "$LAYER" \
--k "$K"
mkdir -p results/sae
python sae_diff.py analyze --topn 30 | tee results/sae/organism-a.txt
cp feats_organism.npz results/sae/feats_organism_a.npzReuse the baseline features because the model, layer, dictionary, k, and
prompt set are unchanged:
python sae_diff.py encode \
--model models/sl-organism-b-7b \
--tag organism \
--layer "$LAYER" \
--k "$K"
python sae_diff.py analyze --topn 30 | tee results/sae/organism-b.txt
cp feats_organism.npz results/sae/feats_organism_b.npz
cp feats_baseline.npz results/sae/feats_baseline.npzThe analysis prints two sections:
global driftlists features most changed by fine-tuning. Generic fine-tuning commonly produces large values here, so this section alone is not evidence of loyalty.principal-selectiveprints one line for each of the 23 principals. A meaningful hit requires one principal to stand well clear of the other 22, not merely rank first in a nearly flat list.
If the principal-selectivity scores are flat, repeat steps 2-4 with two or
three other available layers before concluding that no selective feature was
detected. Baseline and organism files compared in one analysis must always use
the same layer and k.
results/political_figure_discovery.jsonl contains generated responses,
prompt-family identifiers, variants, and seeds. It supports behavioral
scoring, candidate-principal selection, and review of whether a response
actually changes under matched controls.
feats_baseline.npz and feats_organism.npz contain:
feats: one token-summed SAE feature vector per matched prompt;principal: the principal associated with each row;scenario: the scenario identifier associated with each row.
These are the most useful files for further numerical analysis. From a matched pair, we can check principal selectivity, scenario consistency, effect sizes, rank stability, outliers, and bootstrap uncertainty without rerunning either 7B model. With 414 prompts and 131,072 features, one uncompressed float32 matrix is approximately 207 MiB, not tens of gigabytes.
The two matrices are only comparable when they were created with the same
prompt set, SAE checkpoint, layer, k, tokenizer/chat-template convention,
and aggregation method.
The SAE checkpoint describes the dictionary itself. Its keys and shapes reveal the residual dimension, feature count, parameter naming convention, encoder orientation, decoder orientation, and whether a learned BatchTopK threshold is present. It does not, by itself, show that a principal is selected. That claim requires activations from matched prompts.
The decoder vector for a candidate feature can later be used for feature
interpretation or intervention. Sharing the entire multi-gigabyte dictionary
is unnecessary for initial selectivity analysis; first identify candidate
feature IDs from the .npz matrices.
The baseline and organism .safetensors shards contain language-model
parameters. They support weight-difference analysis but do not directly reveal
the trigger or principal. They are already available from Hugging Face and
should not be copied into this repository.
Commit code, YAML configuration, the README, and compact text/CSV/JSON summary reports. Do not commit:
- baseline or organism model weights;
- the full SAE repository or converted SAE checkpoints;
feats_*.npzmatrices;- Hugging Face cache directories.
For further analysis, transfer a matched feats_baseline.npz plus the
corresponding organism feature file outside normal Git, along with a small text
file recording the model, SAE source, layer, k, and conversion procedure.
Object storage, an attached artifact, DVC, or Git LFS is more appropriate for
large binaries, subject to the service's quota. If only a quick interpretation
is needed, share the analyze text output first.