All case studies
Mechanistic InterpretabilityApplied AI lab (confidential)·AI Research·14 weeks

Tracing Refusal Circuits in Frontier Models

We localized the components responsible for refusal behavior across 8 production models. Safety teams got a mechanistic handle on a previously fragile capability.

Engineer reviewing dashboards on dual monitors
Circuits mapped
47
Models covered
8
Duration
14 weeks
Industry
AI Research

The problem

Refusal behavior was fragile and inconsistent across model versions. The safety team could not predict which fine-tunes would weaken it.

Our approach

  • 1Used sparse autoencoders to surface candidate refusal features.
  • 2Verified causal role via activation patching across 8 frontier models.
  • 3Built a refusal-circuit dashboard for the safety team.

The outcome

Forty-seven distinct refusal circuits were mapped and labeled. The safety team now ships a circuit-level regression report alongside every release.