dorsal/arxiv
View SchemaTP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models
| Authors | Xin Jin, Yichuan Zhong, Yapeng Tian |
|---|---|
| Categories | |
| ArXiv ID | 2601.08011vv1 |
| URL | https://arxiv.org/abs/2601.08011 |
| Journal | Transactions on Machine Learning Research, 2025 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Current text-conditioned diffusion editors handle single object replacement well but struggle when a new object and a new style must be introduced simultaneously. We present Twin-Prompt Attention Blend (TP-Blend), a lightweight training-free framework that receives two separate textual prompts, one specifying a blend object and the other defining a target style, and injects both into a single denoising trajectory. TP-Blend is driven by two complementary attention processors. Cross-Attention Object Fusion (CAOF) first averages head-wise attention to locate spatial tokens that respond strongly to either prompt, then solves an entropy-regularised optimal transport problem that reassigns complete multi-head feature vectors to those positions. CAOF updates feature vectors at the full combined dimensionality of all heads (e.g., 640 dimensions in SD-XL), preserving rich cross-head correlations while keeping memory low. Self-Attention Style Fusion (SASF) injects style at every self-attention layer through Detail-Sensitive Instance Normalization. A lightweight one-dimensional Gaussian filter separates low- and high-frequency components; only the high-frequency residual is blended back, imprinting brush-stroke-level texture without disrupting global geometry. SASF further swaps the Key and Value matrices with those derived from the style prompt, enforcing context-aware texture modulation that remains independent of object fusion. Extensive experiments show that TP-Blend produces high-resolution, photo-realistic edits with precise control over both content and appearance, surpassing recent baselines in quantitative fidelity, perceptual quality, and inference speed.
{
"annotation_id": "3554b624-97ef-42e3-9f7a-4245ca1b05aa",
"date_created": "2026-02-17T05:53:16.137000Z",
"date_modified": "2026-02-17T05:53:16.137000Z",
"file_hash": "e40333122c93f28777b5de887a465e2d3c60df766616ca2e64aa840a4274930c",
"private": false,
"record": {
"abstract": "Current text-conditioned diffusion editors handle single object replacement well but struggle when a new object and a new style must be introduced simultaneously. We present Twin-Prompt Attention Blend (TP-Blend), a lightweight training-free framework that receives two separate textual prompts, one specifying a blend object and the other defining a target style, and injects both into a single denoising trajectory. TP-Blend is driven by two complementary attention processors. Cross-Attention Object Fusion (CAOF) first averages head-wise attention to locate spatial tokens that respond strongly to either prompt, then solves an entropy-regularised optimal transport problem that reassigns complete multi-head feature vectors to those positions. CAOF updates feature vectors at the full combined dimensionality of all heads (e.g., 640 dimensions in SD-XL), preserving rich cross-head correlations while keeping memory low. Self-Attention Style Fusion (SASF) injects style at every self-attention layer through Detail-Sensitive Instance Normalization. A lightweight one-dimensional Gaussian filter separates low- and high-frequency components; only the high-frequency residual is blended back, imprinting brush-stroke-level texture without disrupting global geometry. SASF further swaps the Key and Value matrices with those derived from the style prompt, enforcing context-aware texture modulation that remains independent of object fusion. Extensive experiments show that TP-Blend produces high-resolution, photo-realistic edits with precise control over both content and appearance, surpassing recent baselines in quantitative fidelity, perceptual quality, and inference speed.",
"arxiv_id": "2601.08011",
"authors": [
"Xin Jin",
"Yichuan Zhong",
"Yapeng Tian"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.LG",
"cs.MM"
],
"journal_ref": "Transactions on Machine Learning Research, 2025",
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "TP-Blend: Textual-Prompt Attention Pairing for Precise Object-Style Blending in Diffusion Models",
"url": "https://arxiv.org/abs/2601.08011",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "ab7ef0df-45ea-4519-84db-d928bea25fb1",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}