dorsal/arxiv
View SchemaAdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling
| Authors | Yongliang Miao, Yangyang Liang, Mengnan Du |
|---|---|
| Categories | |
| ArXiv ID | 2601.08097vv1 |
| URL | https://arxiv.org/abs/2601.08097 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Reward modeling is essential for aligning large language models with human preferences, yet predominant architectures rely on a static pooling strategy to condense sequences into scalar scores. This paradigm, however, suffers from two key limitations: a static inductive bias that misaligns with task-dependent preference signals, and a representational mismatch, as the backbone is optimized for generation rather than fine-grained discrimination. To address this, we propose AdaJudge, a unified framework that jointly adapts representation and aggregation. AdaJudge first refines backbone representations into a discrimination-oriented space via gated refinement blocks. It then replaces the static readout with an adaptive multi-view pooling module that dynamically routes and combines evidence. Extensive experiments on RM-Bench and JudgeBench show that AdaJudge outperforms strong off-the-shelf reward models and traditional pooling baselines.
{
"annotation_id": "6ee473c9-112e-4c3f-9681-16b2717fd80d",
"date_created": "2026-02-17T05:53:16.065000Z",
"date_modified": "2026-02-17T05:53:16.065000Z",
"file_hash": "926e18b811a9bf851840f6284f2718f2a05e602a71b5e72d1b88f21fb477f381",
"private": false,
"record": {
"abstract": "Reward modeling is essential for aligning large language models with human preferences, yet predominant architectures rely on a static pooling strategy to condense sequences into scalar scores. This paradigm, however, suffers from two key limitations: a static inductive bias that misaligns with task-dependent preference signals, and a representational mismatch, as the backbone is optimized for generation rather than fine-grained discrimination. To address this, we propose AdaJudge, a unified framework that jointly adapts representation and aggregation. AdaJudge first refines backbone representations into a discrimination-oriented space via gated refinement blocks. It then replaces the static readout with an adaptive multi-view pooling module that dynamically routes and combines evidence. Extensive experiments on RM-Bench and JudgeBench show that AdaJudge outperforms strong off-the-shelf reward models and traditional pooling baselines.",
"arxiv_id": "2601.08097",
"authors": [
"Yongliang Miao",
"Yangyang Liang",
"Mengnan Du"
],
"categories": [
"cs.CL",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling",
"url": "https://arxiv.org/abs/2601.08097",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "b1f3db84-02ca-4efb-998b-7c2f9dfa23be",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}