An RNA Sequence Is Not a Molecular State: RNA Foundation Models in 2026
RNA models share an alphabet, not an estimand: choose them by molecular object, output, information inputs, evidence class, and artifact contract.
12 articles tagged with #foundation models.
RNA models share an alphabet, not an estimand: choose them by molecular object, output, information inputs, evidence class, and artifact contract.
A reported genomic-model comparison is auditable only when its tokenisation, strand rule, biological target, evaluation split, and artifact terms are explicit.
Protein generators propose different biological objects. A defensible design starts with the assay, records the complete computational stack, and preserves every experimental denominator.
A practical guide to choosing, extracting, adapting and validating protein language-model representations without mistaking a system score for biological generalisation.
A task-first guide to antibody representations, structure prediction, CDR design, humanisation and developability, with the evidence and reproducibility checks that model scores leave out.
Suppose somebody hands you a FASTA file and asks for “the structure.” First decide whether the target is one chain, a protein assembly or a mixed complex containing ligands or nucleic acids.
One pretrained genomic model can look dominant on a leaderboard and lose to a linear baseline on the next task. Two evidence ledgers explain why.
Carbon-3B matches Evo2-7B on sequence recovery, variant-effect prediction, and motif-perturbation discrimination while generating DNA over 150 times faster.
A 1-billion-parameter model conditioned on RNA chemical-probing reactivity folds a held-out transcript to an F1 of 0.987 against an experimentally guided reference.
Basecamp Research's EDEN model trains on proprietary environmental metagenomics to design gene-insertion enzymes, antimicrobial peptides, and synthetic microbiomes -- all validated in the wet lab.
A practical guide to selecting genomic foundation models for bioinformatics tasks. Covers ESM-2, DNABERT-2, HyenaDNA, Nucleotide Transformer, scGPT, and Evo with scoped comparisons for DNA, proteins and single-cell analysis, with corrected references and clear distinctions between frozen representations and trained predictors.
Why does it feel like our tools weren't designed by pathologists? Billions poured into AI models that compress whole slide images into tiny vectors, ignoring how pathologists actually examine tissue. The evidence reveals why scaling won't fix this disconnect.