An RNA Sequence Is Not a Molecular State: RNA Foundation Models in 2026
RNA models share an alphabet, not an estimand: choose them by molecular object, output, information inputs, evidence class, and artifact contract.
14 articles tagged with #computational biology.
RNA models share an alphabet, not an estimand: choose them by molecular object, output, information inputs, evidence class, and artifact contract.
A reported genomic-model comparison is auditable only when its tokenisation, strand rule, biological target, evaluation split, and artifact terms are explicit.
A practical guide to choosing, extracting, adapting and validating protein language-model representations without mistaking a system score for biological generalisation.
A task-first guide to antibody representations, structure prediction, CDR design, humanisation and developability, with the evidence and reproducibility checks that model scores leave out.
Suppose somebody hands you a FASTA file and asks for “the structure.” First decide whether the target is one chain, a protein assembly or a mixed complex containing ligands or nucleic acids.
A 1-billion-parameter model conditioned on RNA chemical-probing reactivity folds a held-out transcript to an F1 of 0.987 against an experimentally guided reference.
Structure-based and ligand-based drug design evolved as separate fields solving the same problem. ConGLUDe, a contrastive geometric learning model, unifies both approaches and outperforms specialist methods on realistic benchmarks without requiring pre-defined binding pockets.
AlphaGenome processes 1 million DNA base pairs to predict variant effects across 7,000+ genomic tracks in one second, outperforming specialized models on 25 of 26 VEP benchmarks.
A technical look at AlphaGenome's architecture, its 2D pairwise embeddings for splicing prediction, and what the model means for clinical variant interpretation.
Basecamp Research's EDEN model trains on proprietary environmental metagenomics to design gene-insertion enzymes, antimicrobial peptides, and synthetic microbiomes -- all validated in the wet lab.
How generative diffusion models can serve as fast surrogates for expensive biological simulations, achieving 22x speedup while preserving the stochastic diversity that makes these models scientifically useful.
Why computational biologists should stop building embeddings and start building simulators, with three tractable project ideas you can implement today using flow matching, Neural ODEs, and cell fate trajectory modeling.
Why biological systems offer the ideal training ground for reinforcement learning: automated verification through physics, not human judgment. From protein design with AlphaFold to RNA folding with ViennaRNA, biology provides the verifiable inverse problems that RL needs at scale.
Machine learning is illuminating biology's hidden half: intrinsically disordered proteins and RNA structures that traditional methods could never capture. But can we trust what we're seeing?