# Researchers Warn That RLHF Preference Data Encodes the Rater's State, Not Just the Output

An audit framework identifies a structured confound in human feedback, where a labeler's condition during annotation leaks into the preference signal.

- Published: 2026-07-21T05:32:28.125Z
- Canonical: https://polylog.news/ai/2026-07-21/researchers-warn-that-rlhf-preference-data-encodes-the-rater
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv (2607.16195)](https://arxiv.org/abs/2607.16195)

A new arXiv paper identifies a structured confound in reinforcement learning from human feedback (RLHF). Pairwise preference labels are meant to reflect the compared model outputs, but the authors argue they can also reflect the rater's sta…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-07-21/researchers-warn-that-rlhf-preference-data-encodes-the-rater (subscription information: https://polylog.news/pricing).