What our biology benchmark scores actually counted
A recall reversal and two constant-prediction controls show why benchmark comparisons need populations, selection rules and reference scores.
29 articles tagged with #machine learning.
A recall reversal and two constant-prediction controls show why benchmark comparisons need populations, selection rules and reference scores.
How do you verify scientific knowledge when there are 2.9 million papers on arXiv alone, with thousands more added every day? A new paper extracts nearly two million claims from 16,087 manuscripts and compares machine evaluation to human peer review, with 81% agreement.
Structure-based and ligand-based drug design evolved as separate fields solving the same problem. ConGLUDe, a contrastive geometric learning model, unifies both approaches and outperforms specialist methods on realistic benchmarks without requiring pre-defined binding pockets.
How TTT-E2E achieves constant inference latency regardless of context length by treating long context as a learning problem rather than an architecture problem.
How the same transformer architecture powering GPT learned to predict molecular properties by treating chemistry as a language problem
How researchers adapted BERT for molecular property prediction, turning SMILES strings into drug discovery insights
How generative diffusion models can serve as fast surrogates for expensive biological simulations, achieving 22x speedup while preserving the stochastic diversity that makes these models scientifically useful.
Machine learning models now outperform FDA-approved biomarkers in predicting treatment response, but the best-performing models often resist explanation. Here's how precision oncology is navigating the trade-off between performance and interpretability.
Recent research suggests RL training optimizes search efficiency over existing capabilities rather than expanding reasoning capacity. Here's what the pass@k evidence actually shows.
Why computational biologists should stop building embeddings and start building simulators, with three tractable project ideas you can implement today using flow matching, Neural ODEs, and cell fate trajectory modeling.
Understanding when you're working with an environment versus a benchmark changes how you design experiments, interpret results, and communicate findings. This guide covers the practical differences every RL practitioner should know.
A comparative analysis of approaches to catastrophic forgetting in language models, from parameter regularization to sparse memory architectures that reduce forgetting from 89% to just 11%.
DeepMind's DiscoRL discovers reinforcement learning algorithms that outperform hand-designed methods like PPO and DQN. By treating algorithm design as a meta-learning problem, it found alternatives to value functions and bootstrapping through optimization alone.
Are large language models merely stochastic parrots, or do they develop genuine internal representations of the world? This investigation examines evidence from Othello-GPT, spatial encoding in LLMs, and the symbol grounding problem to explore what cognitive science reveals about AI understanding.
Mixtral uses 46.7B parameters but only activates 13B per token. This architectural trick called Mixture of Experts powers Gemini 1.5, DeepSeek V3, and more. Learn how MoE works, its hidden costs, and when to use it.
Pedro Domingos proposes that neural networks and symbolic AI are the same mathematical operation - a logical rule can be equivalently written as a tensor equation in Einstein summation notation. If true, we've been building separate tools for problems that share identical structure.
GPT-4's 128K context window? It only uses about 10% effectively. Google's TITANS architecture introduces test-time memory learning that outperforms GPT-4 on long-context tasks with 70x fewer parameters.
Why biological systems offer the ideal training ground for reinforcement learning: automated verification through physics, not human judgment. From protein design with AlphaFold to RNA folding with ViennaRNA, biology provides the verifiable inverse problems that RL needs at scale.
Machine learning is illuminating biology's hidden half: intrinsically disordered proteins and RNA structures that traditional methods could never capture. But can we trust what we're seeing?
World models enable AI agents to imagine futures and plan actions, achieving 10-100x better sample efficiency than traditional reinforcement learning. From DreamerV3 collecting diamonds in Minecraft to foundation models like Sora and Genie, world models represent AI's shift from pattern matching to simulating reality itself.
Neural networks catastrophically forget previous knowledge when learning new tasks—not due to capacity limits but fundamental constraints in distributed learning systems.
Google DeepMind's AlphaEvolve broke a 56-year-old matrix multiplication record and matched or beat human solutions on 95% of 67 mathematical problems.
Are LLMs truly exhibiting emergent capabilities, or are we mistaking measurement artifacts for genuine phase transitions?
An AI research system maintains coherent reasoning across 200+ agent steps over 12 hours, generating 42,000 lines of code while reviewing 1,500 papers. Kosmos achieves 79.4% accuracy through structured world models and parallel agents, but verification remains humanity's bottleneck.
New research reveals a disturbing truth: reasoning models maintain high performance until they hit a complexity threshold, then collapse entirely. They don't degrade gracefully - they fall off a cliff.
Why does it feel like our tools weren't designed by pathologists? Billions poured into AI models that compress whole slide images into tiny vectors, ignoring how pathologists actually examine tissue. The evidence reveals why scaling won't fix this disconnect.
An AI system at Google DeepMind discovered how bacteria share genes across species barriers using 7 days of computational reasoning. When tested, it matched unpublished experimental observations exactly.
Traditional readability formulas miss the mark. Modern embedding models can capture semantic nuance and syntactic structure, but do they actually predict complexity better?
Discover why monolithic embeddings fail for RAG systems and learn how chunking strategies can transform your retrieval performance.