dorsal/arxiv
View Schemasui-1: Grounded and Verifiable Long-Form Summarization
| Authors | Benedikt Droste, Jan Philipp Harries, Maximilian Idahl, Björn Plüster |
|---|---|
| Categories | |
| ArXiv ID | 2601.08472vv1 |
| URL | https://arxiv.org/abs/2601.08472 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Large language models frequently generate plausible but unfaithful summaries that users cannot verify against source text, a critical limitation in compliance-sensitive domains such as government and legal analysis. We present sui-1, a 24B parameter model that produces abstractive summaries with inline citations, enabling users to trace each claim to its source sentence. Our synthetic data pipeline combines chain-of-thought prompting with multi-stage verification, generating over 22,000 high-quality training examples across five languages from diverse sources including parliamentary documents, web text, and Wikipedia. Evaluation shows sui-1 significantly outperforms all tested open-weight baselines, including models with 3x more parameters. These results demonstrate that task-specific training substantially outperforms scale alone for citation-grounded summarization. Model weights and an interactive demo are publicly available.
{
"annotation_id": "22eb027f-318c-4b3e-9b99-c1f94cfd132f",
"date_created": "2026-02-17T05:53:15.640000Z",
"date_modified": "2026-02-17T05:53:15.640000Z",
"file_hash": "23dac4d9d5c6d139f3ecf9ac16ff1efde183788d4d0d89cb9db1423767011e90",
"private": false,
"record": {
"abstract": "Large language models frequently generate plausible but unfaithful summaries that users cannot verify against source text, a critical limitation in compliance-sensitive domains such as government and legal analysis. We present sui-1, a 24B parameter model that produces abstractive summaries with inline citations, enabling users to trace each claim to its source sentence. Our synthetic data pipeline combines chain-of-thought prompting with multi-stage verification, generating over 22,000 high-quality training examples across five languages from diverse sources including parliamentary documents, web text, and Wikipedia. Evaluation shows sui-1 significantly outperforms all tested open-weight baselines, including models with 3x more parameters. These results demonstrate that task-specific training substantially outperforms scale alone for citation-grounded summarization. Model weights and an interactive demo are publicly available.",
"arxiv_id": "2601.08472",
"authors": [
"Benedikt Droste",
"Jan Philipp Harries",
"Maximilian Idahl",
"Bj\u00f6rn Pl\u00fcster"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "sui-1: Grounded and Verifiable Long-Form Summarization",
"url": "https://arxiv.org/abs/2601.08472",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "f646f79c-af6c-4d05-a2b7-e7b15c6f6146",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}