Blog
Notes from the lab.
Essays, experiments, and operational lessons from our alignment, interpretability, and red-teaming work.
NOV 10, 2024Interpretability· 7 min read
The Future of Mechanistic Interpretability
Mechanistic interpretability is moving from circuit-level curiosities to production tools. Here is where it is heading next.
OCT 28, 2024Alignment· 6 min read
Why Alignment Needs More Data Quality
Compute keeps scaling. The next bottleneck is the human signal that teaches models what good looks like.
SEP 15, 2024Red Teaming· 5 min read
Red-Teaming Frontier Models in 2025
Manual red-teaming does not scale. Continuous, automated adversarial pipelines are the new baseline.
AUG 22, 2024Evaluations· 4 min read
Treat Evaluations as a Product Contract
Evals are not a research artifact. They are the contract between your model and the team operating it.