dorsal/arxiv
View SchemaFrom Prompts to Deployment: Auto-Curated Domain-Specific Dataset Generation via Diffusion Models
| Authors | Dongsik Yoon, Jongeun Kim |
|---|---|
| Categories | |
| ArXiv ID | 2601.08095vv1 |
| URL | https://arxiv.org/abs/2601.08095 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
In this paper, we present an automated pipeline for generating domain-specific synthetic datasets with diffusion models, addressing the distribution shift between pre-trained models and real-world deployment environments. Our three-stage framework first synthesizes target objects within domain-specific backgrounds through controlled inpainting. The generated outputs are then validated via a multi-modal assessment that integrates object detection, aesthetic scoring, and vision-language alignment. Finally, a user-preference classifier is employed to capture subjective selection criteria. This pipeline enables the efficient construction of high-quality, deployable datasets while reducing reliance on extensive real-world data collection.
{
"annotation_id": "daedf086-27fd-45b1-acc5-03dd3e4e44b8",
"date_created": "2026-02-17T05:53:16.279000Z",
"date_modified": "2026-02-17T05:53:16.279000Z",
"file_hash": "e4086427cfe5ccdc2b6aaafa0434dff23625a6bab57bca53210ffcfbd11054d4",
"private": false,
"record": {
"abstract": "In this paper, we present an automated pipeline for generating domain-specific synthetic datasets with diffusion models, addressing the distribution shift between pre-trained models and real-world deployment environments. Our three-stage framework first synthesizes target objects within domain-specific backgrounds through controlled inpainting. The generated outputs are then validated via a multi-modal assessment that integrates object detection, aesthetic scoring, and vision-language alignment. Finally, a user-preference classifier is employed to capture subjective selection criteria. This pipeline enables the efficient construction of high-quality, deployable datasets while reducing reliance on extensive real-world data collection.",
"arxiv_id": "2601.08095",
"authors": [
"Dongsik Yoon",
"Jongeun Kim"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "From Prompts to Deployment: Auto-Curated Domain-Specific Dataset Generation via Diffusion Models",
"url": "https://arxiv.org/abs/2601.08095",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "a8446ce2-11b4-4af5-ac9f-4d6358322cb2",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}