# Two Papers Attack the Same Weakness in Retrieval Pipelines: Rewarding Right Answers From Wrong Steps

One trains a step-level reward model that judges retrieval quality independently of the final answer, the other teaches systems to decline when the retrieved evidence is insufficient.

- Published: 2026-09-03T06:26:17.363Z
- Canonical: https://polylog.news/ai/2026-09-03/two-papers-attack-the-same-weakness-in-retrieval-pipelines-r
- Publisher: Polylog (AI desk)
- Section: tech
- Sources: [arXiv cs.CL](https://arxiv.org/abs/2609.01658), [arXiv cs.CL](https://arxiv.org/abs/2609.01687), [arXiv cs.LG](https://arxiv.org/abs/2609.01615)

Retrieval-augmented generation (RAG) grounds a model's answers in retrieved documents, and its persistent failure mode is that a multi-hop question can be answered correctly for the wrong reasons. Outcome-based training rewards the final an…

This story is for subscribers. Read it in full at https://polylog.news/ai/2026-09-03/two-papers-attack-the-same-weakness-in-retrieval-pipelines-r (subscription information: https://polylog.news/pricing).