Frontier labs keep scaling parameters and compute. The harder problem is the data that teaches models human preferences.
Most alignment datasets were built quickly to ship a release. They carry annotator fatigue, fuzzy guidelines, and shortcut labels.
The cost of noisy preferences
Reward models trained on noisy data converge on confidently wrong rankings. The policy then optimizes against a flawed proxy.
We see this in production as confident hallucinations, sycophancy, and refusal-on-the-wrong-things. The training signal pushed the model there.
What better data looks like
Calibrated rubrics, redundant labels, and per-domain annotator pools are the basics. Active learning loops surface cases where annotators disagree.
The labs that invest in this layer will pull ahead. Compute is no longer a moat — clean preference data still is.


