Ethics & Responsible Use
Why Even AI Companies Don't Fully Understand Their Own Models
Interpretability research has made real progress and still cannot give a full account of why a model said what it said. Here is what is actually known.
Updated 9/13/2026
This sounds like a conspiracy claim and is instead an ordinary engineering fact the labs state openly: nobody can give a complete causal explanation of why a large model produced a particular output.
Why not
A frontier model is billions of numerical parameters arrived at by optimisation, not design. No engineer wrote the rule that makes it good at legal summaries. That capability is distributed across an enormous matrix of weights, and concepts do not sit in tidy locations — a single neuron participates in many unrelated features, and a single feature is spread across many neurons. This "superposition" is efficient for the model and hostile to inspection.
The result is software you can test but not read. Behaviour is measured statistically, from the outside.
The explanation trap
Ask a model why it answered something and you get a fluent rationale. That rationale is generated text, produced by the same process as the answer — not a log of internal computation. Research on chain-of-thought has repeatedly found cases where the stated reasoning does not match the factors actually driving the output. Treating a model's self-report as an audit trail is one of the most common mistakes in AI governance right now.
What interpretability has achieved
Real things, and worth knowing about before concluding the situation is hopeless.
Sparse autoencoders have extracted millions of human-interpretable features from production models — recognisable concepts like a city, a code vulnerability, or a tone of deference — and researchers have shown that amplifying or suppressing those features changes behaviour in predictable ways. Specific circuits for narrow tasks have been mapped end to end. Probes can sometimes detect internal states, including signals that a model's stated confidence diverges from its internal one.
That is genuine progress. It is also partial: mapping features in a model is not the same as predicting what that model will do on an input nobody has tried.
Why it matters practically
Because the tools we have for assurance are behavioural. If you cannot explain a decision, you cannot fully explain a refusal, a bias or a failure — which is a problem in lending, hiring, medicine and anywhere a "right to an explanation" applies. Guarantees phrased as "it will never do X" are claims about a black box, and jailbreaks keep demonstrating the limits of that confidence. It also means the alignment question inherits this one: verifying goals is harder when you cannot inspect them.
The debate
One camp argues interpretability is the highest-leverage safety work available and should be funded far beyond current levels before capability advances further. Another argues it may never scale to frontier systems, and that rigorous behavioural evaluation, staged deployment and liability are the practical route. A third notes the field is young and both bets should run in parallel, which is roughly what is happening.
What it means for you
Do not accept a model's explanation of itself as evidence. For decisions with consequences, keep a human who can justify the outcome on their own terms. And treat confident vendor claims about internal safety properties as what they are — assertions about a system nobody can fully read.
Related: the alignment problem, explained simply and the bias behind the guardrails. Definitions in the glossary.
Keep going
Build something with the prompt generator, decode the jargon in the glossary, or compare the tools on our platform deep-dives.