Models

Arena Post-Training Boosts FLUX.2-dev and Ideogram 4

Arena has launched a text-to-image post-training method that uses gated rewards to prevent reward hacking, driving FLUX.2-dev and Ideogram 4 to the top of open-source leaderboards.

AlphaSignal4 days agoModels
Image: AlphaSignal

AI research organization Arena has unveiled a new reinforcement learning post-training recipe designed to optimize text-to-image generators without sacrificing prompt fidelity. When applied to the FLUX.2-dev model, the technique yielded a 69-point Elo gain on Arena's live leaderboard. Meanwhile, a post-trained version of Ideogram 4 reached 1,224 points, climbing above all other publicly listed open-source models on the leaderboard as of September 4, 2026. The system achieved a 66.0% offline win rate against base models.

To train the system, Arena utilized 10,000 real user prompts and evaluated results using a 1,000-prompt test set judged by Gemini-3.5-Flash. The post-training framework combines multiple evaluators to prevent reward hacking, where models generate aesthetically pleasing images that ignore prompt instructions. It uses a preference reward model trained on roughly 5 million pairwise human votes across more than 100 models. To ensure faithfulness, a language model generates yes-or-no checklists for each prompt, which a vision-language model then grades.

Crucially, Arena introduces an anti-reward-hacking rubric that acts as a gate. If the system detects common exploits like garbled text or unwanted photographic drift, the preference score is immediately clipped to a maximum of zero. This prevents the model from earning positive reinforcement for flawed outputs. Additionally, Arena trains multiple policies with different reward setups and merges them using a weight-space ensembling technique. This parameter-merging step alone added 1.8 percentage points to the final offline win rate.

For machine learning practitioners, this development offers a robust framework to fine-tune diffusion and flow-based generators without manual curation. By separating visual preference, prompt coverage, and constraint enforcement into distinct, gated evaluators, developers can systematically trace why a model fails. The weight-space ensembling method also allows practitioners to combine complementary policies into a single model, delivering superior performance without introducing any additional computational overhead during inference.

This is our own summary of reporting by AlphaSignal

More in Models