dorsal/arxiv
View SchemaViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning
| Authors | Po-han Li, Shenghui Chen, Ufuk Topcu, Sandeep Chinchali |
|---|---|
| Categories | |
| ArXiv ID | 2601.09851vv2 |
| URL | https://arxiv.org/abs/2601.09851 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Multimodal video captioning condenses dense footage into a structured format of keyframes and natural language. By creating a cohesive multimodal summary, this approach anchors generative AI in rich semantic evidence and serves as a lightweight proxy for high-efficiency retrieval. However, traditional metrics like BLEU or ROUGE fail to quantify information coverage across disparate modalities, such as comparing a paragraph of text to a sequence of keyframes. To address this, we propose the Video Summary Information Loss (ViSIL) score, an information-theoretic framework that quantifies the video information not captured by a summary via vision-language model (VLM) inference. By measuring the information loss, ViSIL is a unified metric that enables direct comparison across multimodal summary formats despite their structural discrepancies. Our results demonstrate that ViSIL scores show a statistically significant correlation with both human and VLM performance on Video Question Answering (VQA) tasks. ViSIL also enables summary selection to optimize the trade-off between information loss and processing speed, establishing a Pareto-optimal frontier that outperforms text summaries by $7\%$ in VQA accuracy without increasing processing load.
{
"annotation_id": "86437d61-6030-458d-9137-76a0bdd6bd4c",
"date_created": "2026-02-17T05:53:23.743000Z",
"date_modified": "2026-02-17T05:53:23.743000Z",
"file_hash": "37be2fbab3f37b4bae204971f4b80b2b918fd38f79a86cd85bb8431779a72c51",
"private": false,
"record": {
"abstract": "Multimodal video captioning condenses dense footage into a structured format of keyframes and natural language. By creating a cohesive multimodal summary, this approach anchors generative AI in rich semantic evidence and serves as a lightweight proxy for high-efficiency retrieval. However, traditional metrics like BLEU or ROUGE fail to quantify information coverage across disparate modalities, such as comparing a paragraph of text to a sequence of keyframes. To address this, we propose the Video Summary Information Loss (ViSIL) score, an information-theoretic framework that quantifies the video information not captured by a summary via vision-language model (VLM) inference. By measuring the information loss, ViSIL is a unified metric that enables direct comparison across multimodal summary formats despite their structural discrepancies. Our results demonstrate that ViSIL scores show a statistically significant correlation with both human and VLM performance on Video Question Answering (VQA) tasks. ViSIL also enables summary selection to optimize the trade-off between information loss and processing speed, establishing a Pareto-optimal frontier that outperforms text summaries by $7\\%$ in VQA accuracy without increasing processing load.",
"arxiv_id": "2601.09851",
"authors": [
"Po-han Li",
"Shenghui Chen",
"Ufuk Topcu",
"Sandeep Chinchali"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.HC"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning",
"url": "https://arxiv.org/abs/2601.09851",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "9c442594-a264-46a7-b337-69ad648b806f",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}