← Back to the wire

Training mRNA Language Models Across 25 Species for $165

AchievementModelMar 31, 2026

OpenMed built an end-to-end protein AI pipeline covering structure prediction, sequence design, and codon optimization, training four production models across 25 species in 55 GPU-hours for $165. The pipeline uses ESMFold for protein folding and ProteinMPNN for sequence design. For codon optimization, CodonRoBERTa-large-v2 achieved a perplexity of 4.10 and a Spearman CAI correlation of 0.40, outperforming ModernBERT. Maziyar Panahi contributed to the work, with CodonJEPA listed as an upcoming model on the project roadmap.

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

OpenMedCompanyESMFoldModelProteinMPNNModelCodonRoBERTaModelCodonJEPAModelMaziyar PanahiPersonHugging FaceCompany
Canonical: https://huggingface.co/blog/OpenMed/training-mrna-models-25-species