dorsal/arxiv
View SchemaPixRec: Leveraging Visual Context for Next-Item Prediction in Sequential Recommendation
| Authors | Sayak Chakrabarty, Souradip Pal |
|---|---|
| Categories | |
| ArXiv ID | 2601.06458vv1 |
| URL | https://arxiv.org/abs/2601.06458 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large Language Models (LLMs) have recently shown strong potential for usage in sequential recommendation tasks through text-only models, which combine advanced prompt design, contrastive alignment, and fine-tuning on downstream domain-specific data. While effective, these approaches overlook the rich visual information present in many real-world recommendation scenarios, particularly in e-commerce. This paper proposes PixRec - a vision-language framework that incorporates both textual attributes and product images into the recommendation pipeline. Our architecture leverages a vision-language model backbone capable of jointly processing image-text sequences, maintaining a dual-tower structure and mixed training objective while aligning multi-modal feature projections for both item-item and user-item interactions. Using the Amazon Reviews dataset augmented with product images, our experiments demonstrate $3\times$ and 40% improvements in top-rank and top-10 rank accuracy over text-only recommenders respectively, indicating that visual features can help distinguish items with similar textual descriptions. Our work outlines future directions for scaling multi-modal recommenders training, enhancing visual-text feature fusion, and evaluating inference-time performance. This work takes a step toward building software systems utilizing visual information in sequential recommendation for real-world applications like e-commerce.
{
"annotation_id": "888ec0bb-ad5e-476b-b24e-909f718e9c54",
"date_created": "2026-02-17T05:53:08.576000Z",
"date_modified": "2026-02-17T05:53:08.576000Z",
"file_hash": "86a546a7d66d8f213eca4da930f1dc564fdb82c6e01d491ca1e3cd1e9f86180d",
"private": false,
"record": {
"abstract": "Large Language Models (LLMs) have recently shown strong potential for usage in sequential recommendation tasks through text-only models, which combine advanced prompt design, contrastive alignment, and fine-tuning on downstream domain-specific data. While effective, these approaches overlook the rich visual information present in many real-world recommendation scenarios, particularly in e-commerce. This paper proposes PixRec - a vision-language framework that incorporates both textual attributes and product images into the recommendation pipeline. Our architecture leverages a vision-language model backbone capable of jointly processing image-text sequences, maintaining a dual-tower structure and mixed training objective while aligning multi-modal feature projections for both item-item and user-item interactions. Using the Amazon Reviews dataset augmented with product images, our experiments demonstrate $3\\times$ and 40% improvements in top-rank and top-10 rank accuracy over text-only recommenders respectively, indicating that visual features can help distinguish items with similar textual descriptions. Our work outlines future directions for scaling multi-modal recommenders training, enhancing visual-text feature fusion, and evaluating inference-time performance. This work takes a step toward building software systems utilizing visual information in sequential recommendation for real-world applications like e-commerce.",
"arxiv_id": "2601.06458",
"authors": [
"Sayak Chakrabarty",
"Souradip Pal"
],
"categories": [
"cs.IR",
"cs.CV",
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "PixRec: Leveraging Visual Context for Next-Item Prediction in Sequential Recommendation",
"url": "https://arxiv.org/abs/2601.06458",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "18108f14-ce86-4c10-a086-f38c1e2fc27c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}