
Long-text rendering in generated images fails mostly because training captions omit the text itself. This writeup traces the failure to caption conditioning and reports the resulting model topping a text-to-image ELO leaderboard.
Checking sign-in…
Loading comments…