HTTP API¶
Everything the UI does goes through this API, so anything you can see you
can script. Generated from the app's own OpenAPI schema by
scripts/gen_api_docs.py — run it after adding a route.
Base URL: http://127.0.0.1:5900. Interactive docs: /docs.
Shared sessions (.mri)¶
| method | path | notes |
|---|---|---|
GET |
/api/session/export |
Session Export |
GET |
/api/session/state |
Session State |
GET |
/api/session/trace |
Session Trace |
POST |
/api/session/close |
Session Close |
POST |
/api/session/open |
Session Open |
Session¶
| method | path | notes |
|---|---|---|
GET |
/ |
Index |
GET |
/api/session |
Session |
Model¶
| method | path | notes |
|---|---|---|
GET |
/api/accelerator |
Accelerator |
GET |
/api/hub/models |
Hub Models |
GET |
/api/model/progress |
Load Progress |
GET |
/api/models/discovered |
Models Discovered |
GET |
/api/models/local |
Models Local |
POST |
/api/model/cancel |
Cancel Load |
POST |
/api/model/load |
Load Model |
POST |
/api/model/prompt |
Prompt |
POST |
/api/model/unload |
Unload Model |
Discovery¶
| method | path | notes |
|---|---|---|
GET |
/api/hub/auth |
Hub Auth |
GET |
/api/ollama |
Ollama Status |
GET |
/api/ollama/resolve |
Ollama Resolve |
GET |
/api/ollama/size |
Ollama Size |
GET |
/api/ollama/suggested |
Ollama Suggested |
POST |
/api/hub/signin |
Hub Signin |
POST |
/api/hub/signout |
Hub Signout |
POST |
/api/ollama/pull |
Ollama Pull |
Attention¶
| method | path | notes |
|---|---|---|
GET |
/api/attention |
Attention |
GET |
/api/attention/ablate |
Ablate Heads |
GET |
/api/attention/ablate/estimate |
Estimate Ablation |
GET |
/api/attention/attribute |
Attribute Tokens |
GET |
/api/attention/baselines |
Compare Baselines |
GET |
/api/attention/control |
Control Ranking |
GET |
/api/attention/diff |
Attention Diff |
GET |
/api/attention/direct |
Direct Attribution |
GET |
/api/attention/head/evidence/cost |
Head Evidence Cost |
GET |
/api/attention/meta |
Attention Meta |
GET |
/api/attention/ov |
Head Ov |
GET |
/api/attention/ov/spectrum |
Head Ov Spectrum |
GET |
/api/attention/types |
Head Types |
GET |
/api/vla/attention |
Vla Attention |
GET |
/api/vla/attention/meta |
Vla Attention Meta |
POST |
/api/attention/anchors |
Token Anchors |
POST |
/api/attention/counterfactual |
Token Counterfactual |
POST |
/api/attention/counterfactual/cost |
Token Counterfactual Cost |
POST |
/api/attention/gradients |
Token Gradients |
POST |
/api/attention/head/evidence |
Head Evidence |
Features¶
| method | path | notes |
|---|---|---|
DELETE |
/api/steer |
Steer Clear |
DELETE |
/api/steer/directions/{name} |
Steer Direction Remove |
GET |
/api/features/ablate |
Ablate Features |
GET |
/api/features/summary |
Features Summary |
GET |
/api/features/{feature_id} |
Feature Detail |
GET |
/api/sae |
Sae Status |
GET |
/api/sae/available |
Sae Available |
GET |
/api/sae/fidelity/estimate |
Sae Fidelity Estimate |
GET |
/api/steer |
Steer Status |
GET |
/api/steer/directions |
Steer Directions |
POST |
/api/features/evidence |
Feature Evidence |
POST |
/api/sae/fidelity |
Sae Fidelity |
POST |
/api/sae/fidelity/cost |
Sae Fidelity Cost |
POST |
/api/sae/load |
Sae Load |
POST |
/api/steer |
Steer |
POST |
/api/steer/direction |
Steer Direction |
POST |
/api/steer/fit |
Steer Fit |
Custom models¶
| method | path | notes |
|---|---|---|
GET |
/api/custom |
Custom Status |
GET |
/api/custom/candidates |
Custom Candidates |
POST |
/api/custom/ablate |
Custom Ablate |
POST |
/api/custom/load |
Custom Load |
POST |
/api/custom/run |
Custom Run |
POST |
/api/custom/scan |
Custom Scan |
POST |
/api/custom/unload |
Custom Unload |
Robot policy¶
| method | path | notes |
|---|---|---|
GET |
/api/vla |
Vla Status |
GET |
/api/vla/actions/cost |
Vla Actions Cost |
GET |
/api/vla/audit |
Vla Audit Dataset |
GET |
/api/vla/datasets |
Vla Datasets |
GET |
/api/vla/episodes |
Vla Episodes |
GET |
/api/vla/frame |
Vla Frame |
GET |
/api/vla/occlude/cost |
Vla Occlude Cost |
GET |
/api/vla/ood |
Vla Ood Score |
GET |
/api/vla/ood/cost |
Vla Ood Cost |
GET |
/api/vla/replay |
Vla Replay |
GET |
/api/vla/sweep/cost |
Vla Sweep Cost |
GET |
/api/vla/timeline |
Vla Timeline Tracks |
POST |
/api/vla/actions/compare |
Vla Actions Compare |
POST |
/api/vla/actions/knockout |
Vla Actions Knockout |
POST |
/api/vla/actions/swap |
Vla Actions Swap |
POST |
/api/vla/analyse |
Vla Analyse |
POST |
/api/vla/dataset |
Vla Set Dataset |
POST |
/api/vla/export |
Vla Export |
POST |
/api/vla/load |
Vla Load |
POST |
/api/vla/occlude |
Vla Occlude Frame |
POST |
/api/vla/share |
Vla Share |
POST |
/api/vla/sweep |
Vla Sweep Run |
Agents¶
| method | path | notes |
|---|---|---|
DELETE |
/api/traces |
Traces Clear |
DELETE |
/api/traces/{trace_id} |
Trace Delete |
GET |
/api/traces |
Traces List |
GET |
/api/traces/search |
Search Traces |
GET |
/api/traces/{trace_id} |
Trace Get |
GET |
/api/traces/{trace_id}/bundle/preview |
Bundle Preview |
GET |
/api/traces/{trace_id}/patterns |
Trace Patterns |
POST |
/api/traces/dataset |
Traces To Dataset |
POST |
/api/traces/import |
Traces Import |
POST |
/api/traces/import/inspect |
Import Inspect |
POST |
/api/traces/{trace_id}/steps/{step_id}/adopt |
Adopt Step |
Other¶
| method | path | notes |
|---|---|---|
DELETE |
/api/rubric/{name} |
Rubric Delete |
GET |
/api/corpus/available |
Corpus Available |
GET |
/api/devices |
List Devices |
GET |
/api/diff/replay |
Diff Replay |
GET |
/api/gguf |
Read Gguf |
GET |
/api/gguf/plan |
Plan Gguf |
GET |
/api/graph |
Graph |
GET |
/api/image |
Image Status |
GET |
/api/image/attention/cost |
Image Attention Cost |
GET |
/api/image/attribution/cost |
Image Attribution Cost |
GET |
/api/image/available |
Image Available |
GET |
/api/image/cv/cost |
Image Cv Cost |
GET |
/api/image/discovered |
Image Discovered |
GET |
/api/image/filmstrip/cost |
Image Filmstrip Cost |
GET |
/api/image/local |
Image Local |
GET |
/api/image/progress |
Image Progress |
GET |
/api/image/replay |
Image Replay |
GET |
/api/image/search |
Image Search |
GET |
/api/image/share/plan |
Image Share Plan |
GET |
/api/image/size |
Image Size |
GET |
/api/image/steps/cost |
Image Steps Cost |
GET |
/api/image/tasks |
Image Tasks |
GET |
/api/lens |
Lens |
GET |
/api/lens/tuned |
Tuned Lens Status |
GET |
/api/paths |
Where |
GET |
/api/patterns/across |
Patterns Across |
GET |
/api/policy |
Policy Status |
GET |
/api/pull/progress |
Pull Progress |
GET |
/api/rubric |
Rubric List |
GET |
/api/scorers |
Scorers Catalogue |
GET |
/api/sweeps |
Sweeps List |
GET |
/api/sweeps/{sweep_id}/resume |
Sweeps Resume Plan |
GET |
/api/telemetry |
Read Telemetry |
GET |
/api/trajectory/cost |
Trajectory Cost |
GET |
/api/weights |
Weights View |
GET |
/api/weights/cost |
Weights Cost |
GET |
/v1/models |
V1 Models |
GET |
/v1/mri/{mri_id} |
V1 Mri |
POST |
/api/diff/models |
Diff Models |
POST |
/api/experiments/compare |
Experiments Compare |
POST |
/api/gguf/load |
Load Gguf |
POST |
/api/ground |
Ground Answer |
POST |
/api/image/adapter |
Image Adapter |
POST |
/api/image/attention |
Image Attention Capture |
POST |
/api/image/attribution |
Image Attribution |
POST |
/api/image/cancel |
Image Cancel |
POST |
/api/image/cv/attribute |
Image Cv Attribute |
POST |
/api/image/cv/predict |
Image Cv Predict |
POST |
/api/image/cv/readout |
Image Cv Readout |
POST |
/api/image/filmstrip |
Image Filmstrip |
POST |
/api/image/knockout |
Image Knockout |
POST |
/api/image/load |
Image Load |
POST |
/api/image/share |
Share Image Run |
POST |
/api/image/steps |
Image Steps Run |
POST |
/api/image/unload |
Image Unload |
POST |
/api/judge |
Judge Score |
POST |
/api/judge/plan |
Judge Plan |
POST |
/api/lens/tune |
Train Tuned Lens |
POST |
/api/neurons/evidence |
Neuron Evidence |
POST |
/api/otel/v1/traces |
Otel Ingest |
POST |
/api/patch |
Patch Trace |
POST |
/api/patch/graph |
Patch Graph |
POST |
/api/patch/path |
Path Trace |
POST |
/api/patch/screen |
Patch Screen |
POST |
/api/patchscope |
Patchscope |
POST |
/api/probe |
Probe Layers |
POST |
/api/quantdiff/behaviour |
Quantdiff Behaviour |
POST |
/api/rubric |
Rubric Save |
POST |
/api/rubric/score |
Rubric Score |
POST |
/api/scorers/run |
Scorers Run |
POST |
/api/sweeps/{sweep_id}/resume |
Sweeps Resume |
POST |
/api/trajectory/compare |
Trajectory Compare |
POST |
/api/weights/scan |
Weights Scan Path |
POST |
/v1/chat/completions |
V1 Chat |
POST |
/v1/completions |
V1 Completions |
Feature evidence from your own corpus¶
POST /api/features/evidence with {"texts": [...]} or {"file": "corpus.txt"},
optional feature and top_k.
"Feature 14203 fired" is not a finding. This gives one feature three independent readouts:
| readout | needs a corpus? | exact? |
|---|---|---|
| top-activating spans, firing rate, histogram | yes — yours | no, these are top activations in that text |
| tokens it promotes and suppresses | no | yes — pure weight math |
| what removing it does | — | already measured by the feature ablation ranking |
Nothing is downloaded. The corpus is a local .txt/.jsonl you point at,
read through the same loader modelmri sweep and the tuned lens use.
The corpus is part of the result. Its name, token count and the fraction of features that never fired travel with every number, because "feature 14203 fires on legal citations" and "feature 14203's highest activation in the 40,000 tokens you gave it was on a legal citation" are different claims and only the second was measured.
A feature with zero activations is not seen in this corpus, never dead.
Dead means the feature does nothing; not seen means this text never showed it
anything it responds to, and only one of those is about the model.
selective is false when a feature fires on more than a fifth of tokens — that
is not a concept, and its top spans will look like whatever the corpus is
mostly made of.
No natural-language labels. Naming the concept is the reader's job; a
generated label would be the one thing on the page nothing measured. It also
inherits saes.py's calibration refusal, so unlike the dashboards it competes
with it cannot show features from an SAE fed the wrong convention.
How good is this SAE, in the space the model actually uses?¶
POST /api/sae/fidelity with {"texts": [...]} or {"file": "corpus.txt"},
floor (required), optional label, max_sequences and confirm.
Priced by GET /api/sae/fidelity/estimate?sequences=N and, on this machine,
by POST /api/sae/fidelity/cost.
GET /api/sae already reports FVU and L0, and both are activation-space
numbers: they ask whether the reconstruction is close to the vector the SAE was
handed. The model does not care about that vector, it cares about the logits,
and the directions carrying the residual stream's variance are not the
directions the next token depends on. So an SAE can post an excellent FVU and
still cost the model most of its predictive loss. This is the output-space
half: the model's own cross-entropy on your text, again with the SAE's
reconstruction spliced into the hook, and again with the activation replaced by
a floor.
| field | what it is |
|---|---|
ce_recovered |
(ce_ablate - ce_recon) / (ce_ablate - ce_clean) |
ce_clean / ce_recon / ce_ablate |
the three losses it is built from, nats per predicted token |
floor / floor_means |
which ablation the percentage is against, by name and in a sentence |
n_floor_tokens |
tokens the mean floor was averaged over — null for the zero floor, never 0 |
corpus_label / corpus_sha256 |
what you call the text, and what actually ran |
n_sequences / n_sequences_given / truncated |
what was scored against what was offered |
calibration |
the activation-space half, whole, so fvu and l0 sit beside the percentage |
replay_deviation_nats / splice_deviation_nats |
the resolution every difference above is read against |
passes / elapsed_s |
3n + 2, and what it took here |
floor has no default and never will. Mean-ablation replaces the
activation with the mean vector of this corpus at this hook; zero-ablation
replaces it with the zero vector, which is a point the stream never visits. The
two give different percentages for the same reconstruction and cannot be
converted into each other after the fact, so a default would answer a question
the reader has to be told the answer to. All three raw losses come back so the
choice can be undone.
Nothing is clamped. A ce_recovered below zero means the reconstruction
predicts worse than destroying the activation does. That is a real answer and
the one this measurement most exists to find; clamping it to 0 would report a
broken SAE as a merely useless one.
It refuses rather than dividing by noise. When the floor moves the model's loss by less than the run can resolve — measured, on the corpus in hand, from writing the model's own stream back through the same hook unchanged — the ratio would be a small difference over a smaller one. The three losses are real and they are in the refusal.
3n + 2 forward passes, three per sequence plus two taken once for that
resolution. GET /api/sae/fidelity/estimate?sequences=N is arithmetic and
answers with nothing loaded; it also publishes confirm_above and
needs_confirmation, because above that many passes the run refuses without
confirm: true rather than starting. POST /api/sae/fidelity/cost spends
three real passes — a warm-up, a capture and a probe — to say what one costs
here, and reports probed_sequence_length, because a pass over 64 tokens does
not price a pass over 512.
A file is an id or a path, and either way the server opens it. Ids come
from GET /api/corpus/available; a typed path is descended to rather than
constructed from. Both carry the not-from-this-machine guard, because a path in
a body names a file on the server's disk.
On the command line:
modelmri sae fidelity --model google/gemma-2-2b --corpus notes.txt \
--floor mean_ablate [--sae REPO --hook HOOK] [--max-sequences N] [--yes]
which prints the projection before it loads anything, and refuses past the same
gate without --yes.
Patchscopes — ask the model what it is holding¶
POST /api/patchscope with {"prompt": ..., "layer": N}, optional position,
target, target_layer, max_new_tokens.
Takes a hidden state from your prompt and splices it into a second prompt built to make the model describe whatever is in front of it — the identity target from Ghandeharioun et al. by default — then reads what comes out.
The decode never comes back alone. The method's known failure is that a good target prompt describes anything fluently, so every response carries two controls:
| control | catches |
|---|---|
identity |
the target prompt with its own activation — a decode matching this means the patch changed nothing |
random |
a same-norm random vector — a decode matching this means the target prompt says it whatever it is handed |
overlap_identity and overlap_random report how much of the decode's
vocabulary each control already used. They are reported, not thresholded —
there is no principled cut-off for "the same", so the numbers sit beside all
three decodes and the reader judges. informative is only true when the decode
differs from both as a string and says at least one word neither already
said; complete containment is a test, not a tuned threshold.
The target prompt is part of the result and is returned with every
response — two decodes under different targets are not comparable, and a hidden
default would make that invisible. A source layer spliced into a different
target layer sets cross_layer, because the two streams are only comparable
where the model treats them alike and nothing here checks that it does.
A decode is a generation and therefore a sample. It is what the model said when handed this state through this target prompt, never what the state means. It is its own surface, not a column beside the logit lens — that reports tokens read through the unembedding, this reports a sentence, and side by side they would read as two measurements of one thing.
Path patching — what put it there¶
POST /api/patch/path with {"clean": ..., "corrupt": ..., "layer": N,
"position": P}.
The node grid answers where: "position 7, layer 12 carries the answer". It cannot answer what put it there, because patching a residual stream restores everything that ever wrote into it at once.
This takes one bright cell as the receiver and, for each earlier component — one attention head, or one MLP — adds just that component's clean contribution into the receiver's residual input with everything else still corrupt. A sender that recovers the answer on its own is the thing that wrote it. "Position 7 layer 12 matters" becomes "head 9.6 wrote it."
Scored with the same fraction the node grid uses, so edge and node numbers are on one scale. Both of its controls run here too: eight same-norm random draws, and the same edit taken from a neighbouring position.
recovery_resolution is the number to read first. A recovery is a
difference of logits divided by the gap, and in bfloat16 the representable step
grows with magnitude — so once the logits are large enough, scores land on a
grid. Two senders closer together than that step are tied, not ranked. The
node grid reports it too; it has the same formula and had the same blind spot.
seeding states which edges were considered at all. Edge count is
quadratic in general; this is linear only because the receiver is fixed to the
site you named. Controls run on the top senders by recovery, and the rest carry
a score and no verdict — absence of a verdict is not a verdict.
scope names what v1 does not split. A sender is patched into the
receiver's residual input as a whole, so this says which component wrote what
the receiver reads — not which of its query, key or value paths carried it.
Splitting those across GQA, fused QKV and rotary embeddings would produce
confident and subtly wrong numbers.
Finetune-vs-base diff over a prompt set¶
POST /api/diff/models with {"a": ..., "b": ..., "prompts": [...]}, or
file instead of prompts. Optional include_heads.
behavdiff compares two models on ONE prompt and was built for a different
question — what a quantisation cost you, where both sides are the same weights
at two precisions. A finetune is not that: it changed the model on purpose, in
some places and not others, so one prompt's diff presented as a property of
the finetune is the error this refuses. Every number is a median over your
prompts with its middle half, n, and how often the thing happened at all.
Each side is loaded once, in sequence: 8 GB will not hold both, and the models worth comparing are exactly the ones near that limit.
| field | what it is |
|---|---|
kl |
how far the answers differ per position, as a spread |
flips |
positions whose top token changed, as a spread |
layers |
residual cosine per layer, as a spread, plus how often each was the divergence point |
consensus_layer |
the layer where the cosine falls furthest, on the most prompts |
heads |
how far each head's ablation score moved — only with include_heads |
No threshold decides where the streams come apart. The first version
compared each layer against a floor of 0.999, wearing a docstring that
claimed the floor was measured on the pair. Zero a single head in one block and
compare the model against the original: the cosine moves at exactly the layer
that block's output appears, and one head's worth of movement can still sit
above 0.999. A real divergence, correctly measured, reported as none. It is
now the largest single-step drop, with the size of the drop printed beside
the layer.
Cosine, not distance. A finetune that changed a norm gain moves every vector's length and none of their meanings, and a distance would report that as the model having changed everywhere.
A plurality is not the answer. When the divergence layer is first on fewer
than half the prompts, the summary says the point of divergence moves
between them rather than naming one. consensus_share is out of all prompts,
not out of the ones that diverged.
The head half is opt-in and priced. rank_heads is one pass per head plus
three, times two sides, times every prompt: 5,412 on a 1.7B with six prompts.
Both sides' medians are carried rather than the difference alone — a head that
went from 0.02 to 0.06 and one that went from 4.00 to 4.04 moved by the same
amount and are not the same finding — and the list is sorted by the size of
the move, because a head the finetune started leaning on is as much a finding
as one it abandoned.
The pair is refused when a per-layer table would line up the wrong things — different layer counts, hidden sizes or vocabularies — from the configs alone, before either model loads, naming both sides. Depths are never normalised to a 0–1 fraction: layer 12 of 24 and layer 12 of 32 are not the same place. Tokenisers are checked per prompt, because two models can share a config and still split one string differently.
This is not model diffing in the crosscoder sense. crosscode and OpenMOSS train a shared autoencoder over both models and can say a feature moved; that is GPU-months and neither ships a license file. This is a diff of behaviour on your prompts and of where the residual stream rotated, and the summary says so rather than leaving a reader to assume the stronger claim.
Causal ablation for your own nn.Module¶
POST /api/custom/ablate with {"kind": "layers"} or {"kind": "inputs"},
optional grid.
The custom-model panel maps one forward pass and everything in it is descriptive — shapes, activation statistics, dead units. All of that can be true of a layer the answer does not depend on. This asks the causal question instead, on the one surface in this category nothing else covers: every platform ModelMRI competes with is a fixed catalogue of transformers, and none of them will look at the CNN you trained last week.
| kind | what is ablated |
|---|---|
layers |
each leaf module's output, replaced by its mean over your samples |
inputs |
each input feature, or each patch of an image, replaced the same way |
It refuses rather than defaults, twice.
An adapter with no TASK is refused. KL over a softmax is right for a
classifier and meaningless for a regressor, and both still produce a
plausible ranking — which is exactly why picking one for you would be
dangerous rather than convenient. Declare TASK = "classification" or
"regression" beside load().
An adapter with no sample_inputs() is refused, and one returning fewer than
MIN_SAMPLES is refused with the arithmetic: the mean of one sample is
that sample, so every layer would be replaced by itself and every score would
come back zero — a clean-looking result from a measurement that did not
happen.
The two nulls are different constructions, deliberately. A layer output has
hundreds of dimensions, so its control is a random edit of the same size —
the same distance the mean moves that sample's activation, in a random
direction. Matching the norm of the replacement instead, as patch.py does,
was measured at 1.41× the intervention it was the null for and no site could
ever beat it. A single input feature has one dimension, where a "random
direction" is +1 or -1 and the control becomes the treatment up to a sign;
its null is therefore occluding a different region the same way. With the
wrong one, a trained net whose label depended only on features 0 and 1
reported both as not significant and four noise features as significant.
expected_false_positives is n_tested / (draws + 1): each site is compared
against the strongest of its draws, so under a null where every site is
equivalent the real edit wins one time in nine. Read the margin, not the flag.
Two things it cannot do. Mean ablation is off-distribution — the mean is
not a value the layer ever produces, so a large effect can mean the layer
matters or that the model has never seen an input like the one the ablation
just built. And in a purely sequential model, replacing any module's output
with a constant makes every module after it constant too, so the final answer
is the same constant wherever the chain was cut: measured at 1.19e-07 apart
with zero variance across samples. The sweep separates what is wired in from
what is not — a dead branch scores exactly 0.0 — and cannot rank depth along
a chain.
Grounding — the document, or the weights?¶
POST /api/ground with {"document": "...", "question": "..."}, or
{"file": "notes.txt", "question": "..."}. Optional max_chunks.
Nothing is downloaded, nothing is indexed and nothing is embedded. This is not a retrieval engine: chunking is by blank line and heading, spans come from the tokeniser's own offset mapping, and every passage you send is in the prompt. Retrieval is somebody else's job — the question here is what the model did with what it was given.
Every RAG interface in this category shows you which chunks were retrieved. None of them shows whether the answer used them, and a retriever that pulled the right paragraph next to a model that ignored it and answered from memory looks identical to a working system in all of them.
So each passage comes back with two numbers, and the interesting case is where they disagree:
| field | what it measures |
|---|---|
dependence |
nats the answer's next-token distribution moved when that passage was masked out of attention |
attention |
share of the answer position's attention that landed on it, meaned over every layer and head |
looked_not_used is set for a passage with attention on it and no measurable
dependence — the signature of an answer coming from the weights.
Three things it will not say.
dependence is never a percentage. Masking a whole passage is a large
intervention and the effects are not additive: the response carries joint,
the move when every passage is masked at once, precisely so a reader can see
the parts do not sum to it.
attention is null, never 0.0, on a model whose attention implementation
never builds the score matrix. SDPA and FlashAttention return an empty
tuple for output_attentions=True rather than None, so the obvious loop
completes, sums nothing and reports a measured-looking zero for a number that
was never returned. attention_available says which happened.
looked_not_used is null — not false — whenever the reading could not be
taken. That is either cause above, or a noise_floor of exactly 0.0, where
every passage that moved the answer at all counts as depended-on and the flag
could never fire. floor_degenerate marks it. A deterministic model reproduces
its own answer bit for bit, so this is the ordinary case, not an edge one.
Cost is n_chunks + 4 forward passes. Past max_chunks it refuses with
the count rather than truncating: an answer scored against the first twelve
paragraphs of forty, presented as grounding, is worse than no answer, because
the passages it used might all be in the tail.
In a .mri. The section travels, and the document does not — a .mri
carries a ~120-character preview of each passage, because the text somebody
grounds an answer in is usually the half they do not want forwarded. That is
also why modelmri verify reports grounding as not verifiable rather than
reproducing it: there is nothing in the file to mask out and run again. What it
does check is that each recorded verdict is consistent with the floor stored
beside it, so a file edited to move a verdict without moving its number is
caught. modelmri diff reports the regression this exists for — the answer
starting or stopping depending on the document — and refuses when the two
files' passage previews disagree, because that is the document changing rather
than the model.
Layer-sweep probes¶
POST /api/probe with {"examples": [{"text": ..., "label": 0|1}, ...]},
optional n_permutations and save_as.
Fits a linear probe at every layer and returns the curve — plus the two things that decide whether the curve means anything:
| reference | what it rules out |
|---|---|
majority |
a probe that beats nothing by always guessing the commoner label |
null_low / null_high |
K refits on shuffled labels, same fit, same examples. Whatever that reaches is what this setup produces from information that is not there |
A layer whose accuracy lands inside the band comes back inside_null: true and
is not counted as readable. best_layer is null when nothing cleared, which
is a result: on these examples, at this position, the concept is not linearly
readable anywhere.
Three ways this refuses to overclaim.
null_saturated marks a layer where the shuffled fit reached the top of the
scale — no accuracy could have cleared it, so the layer is untestable with this
many examples rather than uninformative.
expected_false_positives is n_layers × 5%, because the sweep asks every
layer against a 95th-percentile band and that is a multiple comparison. On a
12-layer model it is 0.6 — so one readable layer is roughly what noise
produces, and the response says so when the count is that low.
MIN_PER_CLASS and MIN_TEST are enforced in code, not documented: a fit on
four examples of a 768-dimensional stream separates them perfectly and means
nothing.
Readable is not used. A direction can be linearly present and play no part
in the answer. save_as exports the fitted direction into the same store the
steering harness reads so it can be ablated — and it is refused when no
layer beat its null, because a vector fitted there is fitted to noise and the
store is where it would later be picked up with none of this context attached.
Steering directions — the store, the push, and the null¶
A contrastive direction needs no sparse autoencoder, which is what makes it the
steering that works on almost every model. Five routes read and write the store
under your data directory, beside the two that already existed for SAE-feature
steering. The /api/steer GET and POST contract is unchanged.
| route | what it does |
|---|---|
GET /api/steer/directions |
every saved direction, judged against the model loaded now |
POST /api/steer/direction {name, strength} |
install one on the live model |
DELETE /api/steer |
take off whichever kind is installed |
DELETE /api/steer/directions/{name} |
delete one from this machine |
POST /api/steer/fit |
fit a new one from contrast pairs |
compatible has three states, and null is not false. A row is false
only when it cannot be applied here, and it carries mismatch — the exact
sentence the apply route would refuse with, naming both the model it was fitted
on and the one that is loaded. null means nothing is loaded to judge against,
which is a different claim; the payload names the model and hidden_size every
verdict was taken against so a reader can see what decided them. warnings is
for a direction that WILL apply and still has something to say — a different
checkpoint of the same width, or one that failed its own null when it was
fitted.
A coefficient is not portable, so the status reports it against the stream.
What is applied is constant alpha at every position, which is what CAA does and
what keeps two runs comparable. What is reported is strength.relative —
alpha divided by strength.residual_norm, the mean L2 norm of the residual
stream entering that layer at the last token of the generation that was current
when the direction was applied, measured on this machine. strength.measured
says exactly that; when there is no generation to measure on, relative and
residual_norm are null and strength.unmeasured says why. A relative
strength of 0.0 would claim the push is negligible, so an unmeasured one is
never rendered as one.
That measurement is taken once, and it says so once it is old. The norm
costs a forward pass, so it is taken at apply time and not on every GET
/api/steer; the direction meanwhile outlives the generation it was applied
during, because generating under it is the point. So after the next generation
the number is still a real measurement and strength.measured stops calling it
a current one — it names the generation it belongs to and says to re-apply the
direction to measure against this one. The value never changes silently
underneath its own sentence.
POST /api/steer/fit is two calls to one route. estimate_only: true
spends one warm-up and one probe pass, returns estimate and probe measured
on this machine, and fits nothing (ran: false). The confirm runs it. A
projection that says this accelerator will not hold it refuses with both numbers
named; confirm: true overrides, because unlike free disk, free VRAM can be
made by closing something. Pricing is not permission: estimate_only always
answers with the quote, including when estimate.verdict is "refuse" — a
price that refuses to be quoted because it is high answers the wrong question.
The guard runs on the call that would actually spend it.
The reply carries the whole per-layer table — effect on the held-out half,
null_mean, null_max, beats_null, p_value — because those are the
product. With NULL_REFITS = 8 the smallest attainable p-value is 1/(K+1) =
0.111, so this is a screen and not a significance test: measured on structureless
data with no direction in it, 16.0% of what caa reports as real is noise and
13.0% of what repe does. best_layer is null when no layer beat its own
shuffled refits, and save_as is then refused rather than saved with a caveat —
the store is the one place a direction is later picked up with none of this
beside it.
POST /api/probe with save_as writes into this same store, so a direction
found by a layer sweep and one fitted from contrast pairs are listed, applied
and deleted through the routes above without either knowing about the other.
Head type labels¶
GET /api/attention/types?seq_len=24&n_sequences=6&seed=0
Labels each head induction / previous-token / duplicate-token / sink, or — for
most of them — label: null, which is "no type detected" and a result rather
than a gap.
A label needs all three gates, and each exists because the previous ones were measured and found insufficient:
| gate | what it rules out |
|---|---|
margin ≥ 3σ above the head's own null |
the score being the null |
times_chance ≥ 1 |
significance without effect size — a null with no spread makes any score clear 3σ |
the offset is the head's peak |
a habit the head merely has, rather than what it does |
Two nulls, and null_kind says which was used. Induction and
duplicate-token are gated on matched non-repeating sequences, which is right
for offsets that are only special because the sequence repeats. A
previous-token head attends to i−1 whether or not anything repeats and a sink
attends to position 0 always, so those are gated on chance under the causal
mask instead — their non-repeating null is the same number again.
These are behaviour on repeated random tokens, not claims about real text, and a label must never be read as explaining the ablation ranking: a head can be labelled and irrelevant, or unlabelled and load-bearing. A byte-level tokenizer is refused rather than measured badly.
When one label lands on most of a model's heads the response says so — that is a fact about the model rather than a distinction between its heads.
Direct logit attribution¶
GET /api/attention/direct?position=&top_k=40
Of the logit the model gave the token it predicted, how much came straight from each head and MLP down the residual stream. Sited inside the ablation panel because the two disagree: the ranking says what breaks when a head is removed, this says what a head contributed directly, and a head can be near zero here and still decide the answer by feeding a later head.
The reconstruction residual is not optional. Direct attribution is exact only if the final normalisation is linear, and it is not. TransformerLens makes it exact by folding LayerNorm into the weights — which changes the model you are studying, and once folded nothing in the output says what the folding cost. Here the model is untouched, the normalisation is frozen at the scale a hook recorded from the real pass, and the gap between the summed components and the model's real logit is measured and returned as a percentage residual. A chart without that number is claiming a decomposition it does not have.
The residual is also the floor. A component contributing less than the
reconstruction error cannot be told from the reconstruction error, so those are
flagged unreadable: true rather than rendered as small. That is not a claim
that the component does not matter.
Every contribution is shift-corrected against its own vocabulary mean: softmax ignores a constant added to every logit, so a component that lifts the whole vocabulary equally reports zero.
Contributions are signed — a component can push against the token the model chose, and folding that into a magnitude would hide half the mechanism.
The affine form is verified before anything is reported: the reconstruction is compared against the model's own normalisation, with the tolerance derived from the model dtype's representable step rather than a chosen epsilon. A model whose norm is something else — a learned gate, a different centring — is refused rather than attributed through the wrong transform.
The two lenses¶
GET /api/lens?top_k=5&kind=plain|tuned|both
layers is always the plain reading, on every kind. A tuned reading
arrives beside it in tuned, never in its place — a translator fitted to
minimise disagreement with the final distribution will reduce disagreement with
the final distribution, so a caller handed translated rows where it expected
plain ones would have no way to tell the model from the fit.
Align the two by layer, not by index. The plain lens has one row more:
the model's own final state, which has no translator because it is the answer
rather than a guess at it.
| route | does |
|---|---|
GET /api/lens/tuned |
whether a translator has been fitted for the loaded model, and what to |
POST /api/lens/tune |
fit one. {"texts": [...]} or {"file": "corpus.txt"}, plus optional steps |
Nothing is downloaded. Pretrained lenses exist on the Hub and fetching one would break the offline promise the rest of this package keeps, so the corpus comes from the caller and training happens on this machine.
The response reports held-out KL per layer — measured on sequences the translator never saw — beside the plain KL for the same layer. Training loss is not reported anywhere, because a translator's training loss is a statement about the translator. A layer the translator made worse shows a negative gain rather than being clamped to zero.
caution is non-empty when the corpus is small relative to the fit: a
translator is d_model² + d_model parameters per layer, so a few thousand
tokens leaves it orders of magnitude under-determined. The held-out numbers are
still real; what they are about is text like the training text.
A saved lens is refused if it was fitted to a different model or a different dtype. Loading one across either would produce a confident, plausible, entirely wrong reading.
Receipts¶
Every measurement route returns a receipt alongside its numbers: what
produced them, in a shape a machine can read. /api/attention/ablate,
/api/attention/attribute, /api/attention/baselines, /api/features/ablate,
/api/lens and /api/patch all carry one, and session/export writes the
set of them into the .mri.
{
"op": "ablate_heads",
"request": {"layer": 0, "baseline": "zero", "position": 4},
"tool_version": "0.11.0",
"model": "Qwen/Qwen3-1.7B",
"revision": "70d244cc86ccca08cf5af4e1e306ecf908b1ad5e",
"revision_note": "the commit `refs/main` resolves to in the local cache",
"dtype": "bfloat16",
"device": "cuda:0",
"attn_implementation": "eager",
"seed": null,
"tokenizer_sha256": "41e00eccf531cffc",
"tokenizer_note": "the full fast-tokenizer definition",
"prompt_sha256": "bbaff4d2ecd5892d",
"n_prompt_tokens": 21,
"measured_at": "2026-08-18T11:18:20+00:00"
}
Three fields can genuinely fail to resolve, and each answers null with a
note saying why rather than a plausible default — a receipt that quietly
reports the wrong revision is worse than one that reports no revision, because
the first is trusted and the second is questioned:
revisionis read from the local cache, never the network, so it works air-gapped.refs/mainis consulted first; if several revisions are cached and no ref says which was loaded, the answer isnullandrevision_notesays naming one would be a guess.tokenizer_sha256covers the full fast-tokenizer definition where there is one — vocabulary, merges, normaliser, pre-tokeniser. Where there is not,tokenizer_notesays the hash is vocabulary-only, because two tokenizers with the same vocabulary and different normalisers produce different token ids and the two hashes must not be compared.seedisnullwhen the measurement was not seeded. That is not seed0.
Receipts carry no filesystem paths and no usernames: the model name is
reduced to its basename when it was loaded from a folder, and any absolute
path in request is reduced the same way. tests/test_no_machine_leaks.py
enforces it.
Streaming¶
GET /ws/generate (WebSocket). Send {"prompt": "..."} and receive
{"type":"token","text":"..."} frames, then one {"type":"done"}.
A generation that fails mid-stream sends {"type":"error","message":...} —
it does not silently close as a success.
Status codes¶
| code | meaning |
|---|---|
| 200 | fine |
| 409 | you asked for something in the wrong order, or a dependency is down. The body has an actionable message. |
| 413 | the upload was larger than the limit — a .mri is not that big, so it is probably not one. |
| 422 | the request was malformed, the model is gated and you have not accepted its licence, a custom adapter could not be loaded, or a session file could not be read. |
There is deliberately no 500 path for ordinary failures: an unreachable Ollama, a stalled download, an out-of-memory load and a user's adapter raising on import all return a status that says what happened.