Your team has been asked whether Evo 2 could help prioritise variants of uncertain significance. Before anyone discusses accuracy, someone has to answer three practical questions. What machine does it need? What will it cost? And if you validate it on one machine, will it give the same answers on the next one?
We ran the same Evo 2 7B container on three cloud GPUs, repeated the whole run four times on fresh machines, and measured all three. The short answers: one 48 GB GPU is enough for 8 kb windows; scoring costs well under a dollar per 1,000 variants; and the GPU model is part of the method. Two H100s on different clouds gave identical scores for every one of 3,893 BRCA1 variants. An L40S running the same container moved two thirds of those variants by more than 1% in rank. Study repository.
What Evo 2 is, operationally
Evo 2 is a DNA language model from Arc Institute, published in Nature in 2026. It predicts the next nucleotide given everything before it, and it comes in 7B and 40B parameter versions with contexts of up to 1 million bases. Code and weights are open under Apache-2.0. Evo 2 paper.
Two uses matter for clinical informatics. Scoring compares the model's likelihood for a reference sequence and the same sequence with a variant; a drop suggests the variant is unusual in a way the model has learned to notice. Embeddings take the model's internal numerical representation of a sequence and feed it to a separately trained classifier. Neither output is a calibrated clinical classification.
Hardware depends on the checkpoint. The 7B models run in bfloat16 on one GPU. The 1B, 20B and 40B checkpoints need FP8 arithmetic on a Hopper GPU, and the 40B model needs two H100s or one H200. Arc README, NVIDIA support matrix. This guide uses 7B.
What we ran
One container image, pinned by digest, ran unchanged on three machines. Each machine ran five workloads, all on public data:
| Workload | What it does |
|---|---|
| Correctness test | Arc's own forward-pass test; the loss must match Arc's expected value |
| BRCA1 scoring | 3,893 variants, each scored with an 8,192 bp window, as in Arc's notebook |
| Exon embeddings | Embeddings from one layer for 122 genome positions, fed to a small published classifier |
| Memory probe | One forward pass at 8, 33, 131 and 262 kb, stopping at the first out-of-memory error |
| Repeat | The first 200 BRCA1 variants again, in a new process on the same machine |
The BRCA1 labels come from a saturation genome editing experiment that measured the function of every possible single-nucleotide change in key regions of the gene. Findlay et al. 2018. The protocol, budget and three amendments were written down before any GPU ran. Protocol.
Choosing a machine
| Provider | Machine | GPU | Buying option | Tested here |
|---|---|---|---|---|
| Google Cloud | a3-highgpu-1g |
H100 80 GB | Spot or Flex-start only | Yes |
| Runpod | Secure Cloud pod | H100 80 GB | On-demand | Yes |
| Runpod | Secure Cloud pod | L40S 48 GB | On-demand | Yes |
| AWS | p5.4xlarge |
H100 80 GB | On-demand | No (launcher included, untested) |
| AWS | g6e.2xlarge |
L40S 48 GB | On-demand | No (launcher included, untested) |
Runpod is a GPU-only cloud: you rent a GPU and a container, billed by the second. It runs any container image, which made it possible to use the identical image everywhere. It publishes an ISO 27001 certificate and SOC 2 Type II report and says it can sign a HIPAA business associate agreement. Runpod compliance.
Allow at least 100 GB of disk: the container is about 25 GB unpacked, and the two 7B checkpoints are 13.8 GB and 13.0 GB. A new cloud account usually has no GPU quota, so request it first. On Google Cloud the quota is "Preemptible NVIDIA H100 GPUs" in your region; ours was approved within minutes.
Step by step
The repository contains a launcher for each provider. Each one creates the machine with a one-off SSH key and a 3-hour limit, copies the committed code and inputs, runs the pinned image, copies the results back, and deletes the machine whatever happens.
git clone https://github.com/rewire-bio/evo2-cloud-operations
cd evo2-cloud-operations
make test # offline: fetch inputs, check hashes, run unit tests
# credentials stay outside the repository
cat ~/.config/evo2ops/env # GCP_PROJECT=..., RUNPOD_API_KEY=...
set -a; source ~/.config/evo2ops/env; set +a
uv run --frozen python scripts/experiment.py --config configs/full.json \
--output results/full --platforms P2 # P2 Google Cloud H100, P3 Runpod H100, P5 Runpod L40S
The core of the Google Cloud launcher is one command. The A3 machine type includes its GPU, and --max-run-duration deletes the VM even if your laptop goes to sleep:
gcloud compute instances create evo2-run --zone us-central1-a \
--machine-type a3-highgpu-1g \
--provisioning-model SPOT --instance-termination-action DELETE \
--max-run-duration 3h --maintenance-policy TERMINATE \
--image-family common-cu129-ubuntu-2204-nvidia-580 \
--image-project deeplearning-platform-release \
--boot-disk-size 200GB --boot-disk-type pd-ssd \
--metadata install-nvidia-driver=True
Google's current Deep Learning VM image cost us six failed attempts. It ships without Docker. It holds the NVIDIA container toolkit at an older version, so the current package will not install beside it. And a background system update can revoke a running container's access to the GPU, so a new process inside it suddenly finds no GPU. The launcher now installs Docker, matches the toolkit version already present, pauses automatic updates for the run, and attaches the GPU through NVIDIA's Container Device Interface (--device nvidia.com/gpu=all) instead of --gpus all. Run log.
On Runpod the pod is the container, so there is no Docker step. One API call creates it:
curl -sS -X POST https://api.runpod.io/v2/pods \
-H "Authorization: Bearer $RUNPOD_API_KEY" -H "Content-Type: application/json" \
-d '{"name": "evo2-run",
"image": "ghcr.io/rewire-bio/evo2-cloud-operations@sha256:22970720...",
"gpu": {"id": "NVIDIA H100 80GB HBM3", "count": 1},
"cloud": "SECURE", "disk": 200, "ports": ["22/tcp"],
"dataCenterIds": ["EU-RO-1"],
"env": {"PUBLIC_KEY": "ssh-ed25519 ..."}}'
Set dataCenterIds. We did not, and Runpod placed our H100 pod in India in two of our four runs, and our L40S pods in Texas and Missouri. With patient data, that would have been a data-residency breach.
Before any real workload, run the correctness test on the new machine. It takes seconds:
# Arc's expected mean loss for evo2_7b on its four test sequences is 0.3477 (tolerance 0.001)
result = w0_forward_test(model, Path("data/prompts.csv"))
assert result["passed"], result["mean_loss"]
What it cost and how fast it ran
| Platform | Price per hour | Windows scored per second | USD per 1,000 variants | Billed time | Run cost (USD) |
|---|---|---|---|---|---|
| Google Cloud H100, Spot | 6.29 | 2.93 | 0.80 | 46 min | 4.78 |
| Runpod H100, on-demand | 4.02 | 2.91 | 0.52 | 42 min | 2.82 |
| Runpod L40S, on-demand (slower host) | 1.12 | 0.61 | 0.68 | 159 min | 2.96 |
| Runpod L40S, on-demand (faster host, earlier run) | 1.12 | 0.93 | 0.45 | 112 min | 2.09 |
List prices retrieved by API on 9 October 2026, including disk. Cost per variant assumes 5,219 scored windows per 3,893 variants, as in this BRCA1 set. The first three rows are the recorded run. Operations data.
The two H100s ran at the same speed on both clouds; the Spot H100 on Google Cloud simply cost more per hour than Runpod's on-demand H100. The L40S was the surprise. The same GPU model ran at 0.93 windows per second on one Runpod host and 0.61 on another, so its cost per variant moved from the cheapest option to the second most expensive. On a GPU cloud you rent a GPU model, not a guaranteed speed: time a small batch before committing to a large job.
For a run this short, setup took a large share: getting the 12 GB compressed image onto the machine took between 4 and 12 minutes per platform.
How much memory it needs
Scoring 8 kb windows used about 17 GiB of GPU memory on both GPUs. A single 33 kb forward pass used about 30 GiB. At 131 kb, both the 80 GB H100 and the 48 GB L40S ran out of memory. We did not test lengths in between; extrapolating linearly, about 65 kb might fit on an H100 but not on an L40S.

Both GPUs allocated the same memory until they ran out. Crosses mark the memory in use when the next allocation failed.
So with Arc's standard inference code, the advertised 1 Mb context is a property of the model, not of the GPUs we tested. Users report the same in Arc's issue tracker, and Arc points them to multi-GPU frameworks such as BioNeMo for long sequences. Arc issue 160.
The GPU changes the answer
Every machine reproduced its own scores exactly when rerun, and each platform gave bit-identical scores across three separate runs on fresh machines. The Google Cloud H100 and the Runpod H100 matched each other exactly too, on every variant and every classifier probability. The L40S did not.

Left: two H100s on different clouds, identical. Right: an L40S against the H100. Same image, code and inputs.
On the L40S, the median per-variant difference from the H100 was 1.2 × 10⁻⁴. That sounds small until you compare it with the typical score, which is close to zero: the scores range from about −0.019 to 0.005, and most harmless variants sit within a few × 10⁻⁴ of zero. The Spearman correlation was 0.983, and 2,627 of 3,893 variants moved by more than 1% of the list (39 places) in rank. The overall discrimination of loss-of-function variants barely changed: AUROC 0.877 on the L40S against 0.875 on the H100, a difference of 0.002 that is small but statistically detectable.
The variants that matter most were steadier. Most of the rank churn happens among the many near-zero scores of harmless variants. Among the 50 variants each GPU scored as most damaging, 48 were the same.
This matches an unanswered report in Arc's issue tracker that 7B embeddings differ between H100 and L40S. Arc issue 178. We did not isolate the cause. A likely one is that the 7B model uses FP8 arithmetic in some layers when NVIDIA's Transformer Engine is installed, as it was in our container, and low-precision kernels differ between GPU generations. It is not the driver or the host: two L40S machines in different data centres, with different drivers and kernels, gave identical scores.
The practical consequence: if you validate on one GPU model, keep that GPU model, and revalidate when it changes. Arc's correctness test does not catch this. The L40S passed it.
Does it separate damaging variants?
For context, here is how the zero-shot scores compared with familiar baselines taken from the same BRCA1 table.

AUROC for separating loss-of-function from functional variants, with 95% intervals from resampling positions. CADD is version 1.3, from the BRCA1 table.
Evo 2 7B reached AUROC 0.875 overall, against 0.804 for CADD. The advantage was concentrated in noncoding variants (0.953 against 0.875); on coding variants no difference was detected (0.840 and 0.841). Two cautions. The CADD scores are version 1.3, as supplied with the BRCA1 data in 2018; later CADD versions added splicing predictors aimed at exactly the noncoding and splice-region variants where Evo 2 led, and we did not compare them. And this is one gene and one assay with an uncalibrated score, so it is a reason to evaluate Evo 2 further, not a reason to report it.
Controls for a clinical setting
- Keep patient data off hosted APIs. NVIDIA's hosted Evo 2 API terms prohibit protected health information and production use. NVIDIA API Trial Terms. Self-host in an account covered by a business associate agreement: AWS lists EC2 and EBS as HIPAA-eligible, and Google's BAA covers Compute Engine. AWS, Google Cloud.
- Pin the location. Set the region or data centre explicitly on every provider.
- Pin everything else. Reference the container by digest, the weights by Hugging Face revision, and the inputs by checksum.
- Check weight hashes before loading. Arc's loader uses Python pickle, which can run code when a file is loaded. The study's loader refuses any checkpoint whose SHA-256 does not match.
- Avoid remote code. Arc's exon notebook loads its classifier with
trust_remote_code. The classifier is a two-layer network, so we rebuilt it locally and loaded only the weights. - Fix the GPU model, and test every new machine. Run the correctness test, then compare a reference set of your own scores against the validated machine.
- Set a hard time limit. On Google Cloud,
--max-run-durationdeletes the VM after 3 hours whatever happens to the laptop that started it. On Runpod there is no equivalent setting, so check for orphaned pods after every run. - Do not tie the job to your SSH session. One of our runs died when the laptop's connection dropped, because the workload was running in the foreground of that session. Run long jobs detached (
nohup,tmuxor a container started in the background) and poll for completion.
Where Evo 2 operations are heading
- Packaged deployment is converging on NVIDIA NIM. Evo 2 ships as a NIM container and through AWS SageMaker, where the recommended real-time instance for 7B is an L40S. That reduces installation work but adds a vendor licence and a less inspectable runtime. NIM container, AWS Marketplace.
- Long context means multi-GPU. Anything beyond about 33 kb per pass needs frameworks that split a sequence across GPUs, which are harder to install and reproduce.
- Embedding pipelines are more accurate and much more expensive. A 2026 preprint trained classifiers on Evo 2 embeddings and beat likelihood scoring on ClinVar, but extracting embeddings with 65 kb windows for 4.25 million variants took about 20,000 H100-hours and 34 TB of storage. EVEE preprint.
- New GPUs need checking. A user reports wrong losses for the 1B checkpoint on a Blackwell GPU. Arc issue 157. Run the correctness test before trusting any new hardware.
Limitations
These are a handful of runs on three platforms with two GPU models, so the timings are receipts rather than benchmarks, and prices change. AWS was not run; its launcher is in the repository but untested, and an AWS run can follow when we have an account. Only the 7B checkpoints were tested, with FP8 input projections active. The BRCA1 comparison covers one gene from one assay, uses CADD v1.3 as the baseline, and says nothing about calibration or performance across ancestries. An independent methods review found reporting errors in our first write-up of the paper, which we corrected; its report and our response are in the repository. The BRCA1 function scores are free for nonprofit use only.
The full methods, every failed attempt, the paper and the methods review are in the study repository. For the wider evidence on genomic foundation models, see Genomic Foundation Models in 2026.
Help improve this article
Found an error or a better source? Leave a note here, or highlight a passage to comment on it.