dorsal/arxiv
View SchemaHiVid-Narrator: Hierarchical Video Narrative Generation with Scene-Primed ASR-anchored Compression
| Authors | Haoxuan Li, Mengyan Li, Junjun Zheng |
|---|---|
| Categories | |
| ArXiv ID | 2601.07366vv1 |
| URL | https://arxiv.org/abs/2601.07366 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Generating structured narrations for real-world e-commerce videos requires models to perceive fine-grained visual details and organize them into coherent, high-level stories--capabilities that existing approaches struggle to unify. We introduce the E-commerce Hierarchical Video Captioning (E-HVC) dataset with dual-granularity, temporally grounded annotations: a Temporal Chain-of-Thought that anchors event-level observations and Chapter Summary that compose them into concise, story-centric summaries. Rather than directly prompting chapters, we adopt a staged construction that first gathers reliable linguistic and visual evidence via curated ASR and frame-level descriptions, then refines coarse annotations into precise chapter boundaries and titles conditioned on the Temporal Chain-of-Thought, yielding fact-grounded, time-aligned narratives. We also observe that e-commerce videos are fast-paced and information-dense, with visual tokens dominating the input sequence. To enable efficient training while reducing input tokens, we propose the Scene-Primed ASR-anchored Compressor (SPA-Compressor), which compresses multimodal tokens into hierarchical scene and event representations guided by ASR semantic cues. Built upon these designs, our HiVid-Narrator framework achieves superior narrative quality with fewer input tokens compared to existing methods.
{
"annotation_id": "817ed4c5-9581-4df3-befc-c2319a1ee7c5",
"date_created": "2026-02-17T05:53:11.816000Z",
"date_modified": "2026-02-17T05:53:11.816000Z",
"file_hash": "401ce279caa65d8e776260e228d2711f86da03e34b7c114ff38adb8be57d2ea8",
"private": false,
"record": {
"abstract": "Generating structured narrations for real-world e-commerce videos requires models to perceive fine-grained visual details and organize them into coherent, high-level stories--capabilities that existing approaches struggle to unify. We introduce the E-commerce Hierarchical Video Captioning (E-HVC) dataset with dual-granularity, temporally grounded annotations: a Temporal Chain-of-Thought that anchors event-level observations and Chapter Summary that compose them into concise, story-centric summaries. Rather than directly prompting chapters, we adopt a staged construction that first gathers reliable linguistic and visual evidence via curated ASR and frame-level descriptions, then refines coarse annotations into precise chapter boundaries and titles conditioned on the Temporal Chain-of-Thought, yielding fact-grounded, time-aligned narratives. We also observe that e-commerce videos are fast-paced and information-dense, with visual tokens dominating the input sequence. To enable efficient training while reducing input tokens, we propose the Scene-Primed ASR-anchored Compressor (SPA-Compressor), which compresses multimodal tokens into hierarchical scene and event representations guided by ASR semantic cues. Built upon these designs, our HiVid-Narrator framework achieves superior narrative quality with fewer input tokens compared to existing methods.",
"arxiv_id": "2601.07366",
"authors": [
"Haoxuan Li",
"Mengyan Li",
"Junjun Zheng"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "HiVid-Narrator: Hierarchical Video Narrative Generation with Scene-Primed ASR-anchored Compression",
"url": "https://arxiv.org/abs/2601.07366",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "88c233c5-6265-43b9-a360-d9627518168c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}