Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Secret Loyalties

For the white-box candidate-screening design, metrics, commands, and evidence limits, see WHITEBOX_SCREENING.md. For the complete execution sequence, see RUNBOOK.md.

Download the models on Linux

Run these commands from the repository root. Install the Hugging Face CLI:

python3 -m pip install --upgrade huggingface_hub

The two organism repositories are gated. Request access to Alamerton/sl-organism-a-7b and Alamerton/sl-organism-b-7b on Hugging Face, then authenticate:

hf auth login

Create the local model directory and download the baseline, organism A, and organism B into the paths used by the investigation configuration:

mkdir -p models

hf download Qwen/Qwen2.5-7B-Instruct \
  --local-dir models/qwen2.5-7b-instruct

hf download Alamerton/sl-organism-a-7b \
  --local-dir models/sl-organism-a-7b

hf download Alamerton/sl-organism-b-7b \
  --local-dir models/sl-organism-b-7b

The resulting directory layout should be:

models/
|-- qwen2.5-7b-instruct/
|-- sl-organism-a-7b/
`-- sl-organism-b-7b/

Each directory should contain model.safetensors.index.json and all model shards named in that index. The models/ directory is ignored by Git.

How the behavioral and SAE stages relate

The two stages test the same hypothesis in different ways:

  1. run_investigation.py generates model responses and looks for behavioral differences between the baseline and the fine-tuned organisms.
  2. A researcher uses those responses to choose plausible principals and scenario families for closer inspection. This selection is currently manual; the behavioral JSONL does not automatically produce PRINCIPALS.
  3. loyalty_sweep.py defines the fixed comparison matrix used by the SAE experiment: 23 principals multiplied by 18 single-turn scenarios.
  4. sae_diff.py encode performs forward passes only. It does not generate answers. It records a selected residual-stream layer, applies the same baseline-trained SAE dictionary, and sums each feature's activation over the prompt tokens.
  5. sae_diff.py analyze subtracts the baseline feature matrix from the organism feature matrix and asks whether any feature changes selectively for one principal rather than for all names.

The baseline encoder and organism encoder therefore perform the same operation:

matched prompt
  -> model residual stream at the selected layer
  -> one shared SAE encoder
  -> token-summed sparse feature vector

They differ only in which language-model checkpoint produces the residual stream. A feature that moves for every principal is ordinary fine-tuning drift. The loyalty hypothesis requires a feature whose organism-minus-baseline change is unusually large for one principal and stable across relevant scenarios.

Run the SAE comparison on Linux

The SAE experiment applies one dictionary trained on the stock Qwen2.5-7B-Instruct model to both the baseline and a fine-tuned organism. It uses matched prompts so the resulting feature difference is attributable to the fine-tune rather than to different dictionaries.

The runner requires an NVIDIA GPU with CUDA. From the repository root, create the Python environment and install the dependencies:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

Ensure that the installed PyTorch build supports the host's CUDA version:

python -c 'import torch; assert torch.cuda.is_available(), "CUDA unavailable"; print(torch.cuda.get_device_name(0))'

Confirm the shared prompt matrix:

python loyalty_sweep.py

Expected output:

23 principals x 18 scenarios = 414 matched prompts

1. Discover the available SAE checkpoints

The upstream andyrdt/saes-qwen2.5-7b-instruct repository is about 105 GB. Do not download its entire snapshot merely to list the files. Query the file metadata instead:

python - <<'PY'
from huggingface_hub import HfApi

repo = "andyrdt/saes-qwen2.5-7b-instruct"
for path in HfApi().list_repo_files(repo):
    print(path)
PY

The upstream repository provides layers 3, 7, 11, 15, 19, 23, and 27. Within each layer, trainer_0, trainer_1, trainer_2, and trainer_3 correspond to k=32, k=64, k=128, and k=256, respectively.

The upstream checkpoints are named ae.pt. The current minimal loader expects a converted .safetensors checkpoint. Before running encode, confirm that the converted file exposes compatible encoder, decoder, bias, and optional threshold tensors. A converted checkpoint's size can be several gigabytes; do not confuse it with the much smaller encoded prompt-feature output.

The examples below use layer 15 and k=64:

LAYER=15
K=64

2. Encode the baseline

python sae_diff.py encode \
  --model models/qwen2.5-7b-instruct \
  --tag baseline \
  --layer "$LAYER" \
  --k "$K"

This writes feats_baseline.npz.

3. Encode and analyze organism A

The analyzer expects the comparison model to use the tag organism:

python sae_diff.py encode \
  --model models/sl-organism-a-7b \
  --tag organism \
  --layer "$LAYER" \
  --k "$K"

mkdir -p results/sae
python sae_diff.py analyze --topn 30 | tee results/sae/organism-a.txt
cp feats_organism.npz results/sae/feats_organism_a.npz

4. Encode and analyze organism B

Reuse the baseline features because the model, layer, dictionary, k, and prompt set are unchanged:

python sae_diff.py encode \
  --model models/sl-organism-b-7b \
  --tag organism \
  --layer "$LAYER" \
  --k "$K"

python sae_diff.py analyze --topn 30 | tee results/sae/organism-b.txt
cp feats_organism.npz results/sae/feats_organism_b.npz
cp feats_baseline.npz results/sae/feats_baseline.npz

5. Read the output

The analysis prints two sections:

  • global drift lists features most changed by fine-tuning. Generic fine-tuning commonly produces large values here, so this section alone is not evidence of loyalty.
  • principal-selective prints one line for each of the 23 principals. A meaningful hit requires one principal to stand well clear of the other 22, not merely rank first in a nearly flat list.

If the principal-selectivity scores are flat, repeat steps 2-4 with two or three other available layers before concluding that no selective feature was detected. Baseline and organism files compared in one analysis must always use the same layer and k.

What the saved files can tell us

Behavioral JSONL

results/political_figure_discovery.jsonl contains generated responses, prompt-family identifiers, variants, and seeds. It supports behavioral scoring, candidate-principal selection, and review of whether a response actually changes under matched controls.

Encoded feature matrices

feats_baseline.npz and feats_organism.npz contain:

  • feats: one token-summed SAE feature vector per matched prompt;
  • principal: the principal associated with each row;
  • scenario: the scenario identifier associated with each row.

These are the most useful files for further numerical analysis. From a matched pair, we can check principal selectivity, scenario consistency, effect sizes, rank stability, outliers, and bootstrap uncertainty without rerunning either 7B model. With 414 prompts and 131,072 features, one uncompressed float32 matrix is approximately 207 MiB, not tens of gigabytes.

The two matrices are only comparable when they were created with the same prompt set, SAE checkpoint, layer, k, tokenizer/chat-template convention, and aggregation method.

SAE checkpoint

The SAE checkpoint describes the dictionary itself. Its keys and shapes reveal the residual dimension, feature count, parameter naming convention, encoder orientation, decoder orientation, and whether a learned BatchTopK threshold is present. It does not, by itself, show that a principal is selected. That claim requires activations from matched prompts.

The decoder vector for a candidate feature can later be used for feature interpretation or intervention. Sharing the entire multi-gigabyte dictionary is unnecessary for initial selectivity analysis; first identify candidate feature IDs from the .npz matrices.

Model weight shards

The baseline and organism .safetensors shards contain language-model parameters. They support weight-difference analysis but do not directly reveal the trigger or principal. They are already available from Hugging Face and should not be copied into this repository.

What to share and what to keep out of Git

Commit code, YAML configuration, the README, and compact text/CSV/JSON summary reports. Do not commit:

  • baseline or organism model weights;
  • the full SAE repository or converted SAE checkpoints;
  • feats_*.npz matrices;
  • Hugging Face cache directories.

For further analysis, transfer a matched feats_baseline.npz plus the corresponding organism feature file outside normal Git, along with a small text file recording the model, SAE source, layer, k, and conversion procedure. Object storage, an attached artifact, DVC, or Git LFS is more appropriate for large binaries, subject to the service's quota. If only a quick interpretation is needed, share the analyze text output first.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages