Frontier labs keep scaling parameters and compute. The harder problem is the data that teaches models human preferences.

Most alignment datasets were built quickly to ship a release. They carry annotator fatigue, fuzzy guidelines, and shortcut labels.

The cost of noisy preferences

Reward models trained on noisy data converge on confidently wrong rankings. The policy then optimizes against a flawed proxy.

We see this in production as confident hallucinations, sycophancy, and refusal-on-the-wrong-things. The training signal pushed the model there.

What better data looks like

Calibrated rubrics, redundant labels, and per-domain annotator pools are the basics. Active learning loops surface cases where annotators disagree.

The labs that invest in this layer will pull ahead. Compute is no longer a moat — clean preference data still is.