LongAlign: Improving Long-Text Alignment for Text-to-Image Diffusion Models

Abstract

The rapid advancement of text-to-image (T2I) diffusion models has enabled them to generate unprecedented results from given texts. However, as text inputs become longer, existing encoding methods like CLIP face limitations, and aligning the generated images with long texts becomes challenging. To tackle these issues, we propose LongAlign, which includes a segment-level encoding method for processing long texts and a decomposed preference optimization method for effective alignment training. For segment-level encoding, long texts are divided into multiple segments and processed separately. This method overcomes the maximum input length limits of pretrained encoding models. For preference optimization, we provide decomposed CLIP-based preference models to fine-tune diffusion models. Specifically, to utilize CLIP-based preference models for T2I alignment, we delve into their scoring mechanisms and find that the preference scores can be decomposed into two components: a text-relevant part that measures T2I alignment and a text-irrelevant part that assesses other visual aspects of human preference. Additionally, we find that the text-irrelevant part contributes to a common overfitting problem during fine-tuning. To address this, we propose a reweighting strategy that assigns different weights to these two components, thereby reducing overfitting and enhancing alignment. After fine-tuning $512\times 512$ Stable Diffusion (SD) v1.5 for about 20 hours using our method, the fine-tuned SD outperforms stronger foundation models in T2I alignment, such as PixArt-$\alpha$ and Kandinsky v2.2.

Method

Segment-Level Text Encoding

Problem: Although CLIP-like models are commonly used for representation encoding, result evaluation, and reward fine-tuning, existing CLIP-based models have limitations on input text length.

Solution:

For representation encoding: We encode each segment (e.g., each sentence) using CLIP and then merge the embeddings. A careful merging strategy is employed to manage special tokens (<sot>, <eot>, <pad>) to avoid issues with duplicated embeddings for these tokens after merging.
For evaluation and reward fine-tuning: We introduce a segment-level loss function $\mathcal{L}_{i \succ j}^{\text{seg}}$ that allows CLIP-based preference models to process long texts and produce detailed segment-level preference scores (Denscore).
\[ \mathcal{L}_{i \succ j}^{\text{seg}} = \frac{\exp\left(\sum_{k=i}^K \mathcal{R}(x_i, \hat{p}_k) / K\right)}{\exp\left(\sum_{k=i}^K \mathcal{R}(x_i, \hat{p}_k) / K\right) + \exp\left(\sum_{k=i}^K \mathcal{R}(x_j, \hat{p}_k) / K\right)}, \]
where we split the long-text condition $ p $ into $ K $ segments, denoted as $ \{\hat{p}_k\}_{k=1}^K $.

Preference Decomposition and Reweighting

Problem: Preference optimization can effectively enhance T2I diffusion models, but this fine-tuning process encounters significant overfitting challenges.