← Back to the wire

PRX Part 3 — Training a Text-to-Image Model in 24h!

AchievementResearchMar 4, 2026

Photoroom researchers David Bertoin, Roman Frigg, and Jon Almazán trained a text-to-image model in 24 hours by combining multiple optimization techniques. The team used x-prediction in pixel space to eliminate the need for a VAE, applied TREAD token routing to reduce per-step compute, and employed REPA with DINOv3 for representation alignment. They open-sourced the code for reproduction. The experiment demonstrates how far careful engineering can advance performance under strict compute budgets.

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

PhotoroomCompanyDavid BertoinPersonRoman FriggPersonJon AlmazánPerson
Canonical: https://huggingface.co/blog/Photoroom/prx-part3