Tracing Refusal Circuits in Frontier Models
We localized the components responsible for refusal behavior across 8 production models. Safety teams got a mechanistic handle on a previously fragile capability.
- Circuits mapped
- 47
- Models covered
- 8
- Duration
- 14 weeks
- Industry
- AI Research
The problem
Refusal behavior was fragile and inconsistent across model versions. The safety team could not predict which fine-tunes would weaken it.
Our approach
- 1Used sparse autoencoders to surface candidate refusal features.
- 2Verified causal role via activation patching across 8 frontier models.
- 3Built a refusal-circuit dashboard for the safety team.
The outcome
Forty-seven distinct refusal circuits were mapped and labeled. The safety team now ships a circuit-level regression report alongside every release.