rewirebio.io

Labelling a new blood single-cell study: annotation methods and thresholds

Labelling cells in a new single-cell study often means training a method on cells that already carry labels. A recorded benchmark, not peer reviewed, trained six methods on one reference, ran two released models, and tested them on four blood studies kept out of training (32 donors). The benchmark could not establish a clear difference among four learned methods in agreement with author labels. At the threshold chosen on other studies to label 90% of cells, no method labelled 90% of the pooled test cells.

A single-cell RNA sequencing study measures gene activity in each of thousands of individual cells. To make sense of the data, researchers first need to know what type each cell is, such as a B cell, a CD8 T cell or a monocyte. Giving each cell a cell-type label is called annotation.

Annotating cells by hand is slow and relies on expert judgement. A common alternative uses a reference: earlier studies whose cells already carry labels. A computer method learns from the reference and then proposes a label for each cell in the new study. This approach is called label transfer.

Anyone labelling a new study this way faces two questions:

  1. Which method should assign the labels?
  2. How do you decide when a proposed label is reliable enough to keep? Most methods give each label a confidence score, and cells scoring below a chosen threshold can be left unassigned.

Rewire Bio, which also publishes rewire.it, ran a recorded computational benchmark on both questions for blood samples. The benchmark has not been peer reviewed. The methods learned from six existing blood studies. They were then tested on four other blood studies, from 32 donors in total. None of the test studies' cells were used for training.

The findings in brief

  • Choosing a method. Four of the six methods trained on the study's reference agreed with the original authors' labels at almost the same level. The study could not establish a clear difference among them, but it also did not show that they are equivalent. Two simpler methods agreed less.
  • Choosing a threshold. Thresholds chosen on other studies did not reliably carry over to the test studies. According to the paper, how often accepted labels were wrong depended more on which test study was being labelled than on the choice among those four methods.
  • Missing cell types. When three cell types were removed from the reference, the methods still accepted a label from the reference for most cells of those types. Across all cells of those types, the most common label was usually a closely related type.

The study's paper, source at revision fd4588c, release page and evidence archive are public.

How the benchmark was set up

Reference and test studies

The reference was six blood studies from CELLxGENE Census, a public collection of single-cell data. Four test studies were kept entirely out of training, so the methods never saw any of their cells. The test studies were the Human Immune Health Atlas (HIHA) and studies of rheumatoid arthritis (RA), glaucoma and juvenile dermatomyositis (JDM).

Only cells that the Census records as healthy (disease == 'normal') were used. The RA, glaucoma and JDM studies therefore contributed only their cells labelled healthy.

The test cells came from two sources. One was a main sample of each study's cells. The other was extra cells of rare types, added so that those types had enough cells to measure.

What counted as correct

The benchmark measured accuracy as agreement with author labels: the cell-type labels that each test study's own authors had assigned. These labels were mapped to 14 shared classes using the Cell Ontology, a standard vocabulary of cell types. The benchmark did not independently adjudicate whether the author labels were right, so they serve as comparison labels rather than verified truth.

The methods compared

Six methods learned from exactly the same reference cells and genes:

  • M1, a marker rule: scores each class by its top-ranked marker genes in the reference.
  • M2, nearest centroid: picks the class whose average cell is closest.
  • M3, k-nearest neighbours (kNN): takes a vote among the most similar reference cells.
  • M4, multinomial logistic regression: a linear classifier.
  • M5, CellTypist trained on this reference: a logistic-regression annotation tool.
  • M6, scANVI: a deep generative model.

The study also ran two ready-made models exactly as their developers released them: CellTypist Immune_All_Low v2 (P1) and a scTab checkpoint (P2). These released models differ from M1 to M6 in their training data, in their sets of cell-type names and in how those names were mapped to the study's classes.

Thresholds, coverage and error

Each method gives a confidence score for the label it proposes. If the score falls below a threshold, the cell is left unassigned. Two measures follow:

  • Coverage is the share of cells that receive a label.
  • Accepted error is the share of labelled cells whose label disagrees with the author label.

Thresholds were chosen on two validation studies, separate from the four test studies, and then applied unchanged to the test studies. The study used two rules:

  • The coverage threshold (OP-cov) labelled 90% of validation cells.
  • The error threshold (OP-err) was the lowest threshold that kept accepted error on validation cells at 5% or less.

Removing cell types from the reference

The main experiment (Arm A) used the full reference. In a second experiment (Arm B), every plasmacytoid dendritic cell (pDC), antibody-secreting cell (ASC) and mucosal-associated invariant T (MAIT) cell was removed from the reference before the methods were retrained. Test cells of those three types could then receive only a wrong label or no label.

Uncertainty

Numbers in square brackets are 95% intervals. They come from resampling donors 1,000 times within the four test studies. The intervals describe uncertainty across donors in these four studies only. They say nothing about how results would vary in new studies.

Choosing a method: no clear difference among four learned methods

The main accuracy score was macro-F1. For this score, every cell received the method's top label, with no threshold applied. For each cell type, F1 balances two kinds of mistake: cells of that type that were missed, and other cells wrongly given that type. Macro-F1 then averages F1 across cell types, so rare types count as much as common ones. The average covered the 11 classes with at least 20 test cells in the main sample. A score of 1 would mean perfect agreement with the author labels.

Four methods scored almost the same. Logistic regression scored 0.694 [0.669, 0.716], kNN and scANVI 0.692, and CellTypist trained on the reference 0.689.

The study compared each of the other three with logistic regression by taking the difference in scores. Every one of these differences had an interval that included zero. For CellTypist trained on the reference, for example, the difference was −0.005 [−0.013, 0.003].

An interval that includes zero means the data are consistent with no difference. The same data are also consistent with small differences in either direction. The study did not set an equivalence margin in advance, meaning a difference small enough to count as practically the same. The results therefore do not show that the four methods are equivalent. They show only that this benchmark could not establish a clear difference among them.

Two methods agreed less with the author labels. Nearest centroid scored 0.629 (difference −0.066 [−0.077, −0.053]) and the marker rule 0.497 (difference −0.197 [−0.213, −0.179]).

Compared with logistic regression trained on the reference, released CellTypist (P1) scored 0.617 (difference −0.077 [−0.094, −0.060]). Released scTab (P2) scored 0.684 (difference −0.010 [−0.023, 0.003]), an interval that includes zero. The released models differ in training data, cell-type names and label mapping, so these comparisons cannot show whether model design explains any gap.

Error depended more on the test study than on the method

At the coverage threshold, logistic regression's accepted error ranged from 2.1% in HIHA to 17.9% in RA. Table 1 shows the result for each test study.

Test study Test donors Coverage Accepted error
HIHA 12 0.940 0.021
JDM 5 0.906 0.056
Glaucoma 3 0.847 0.144
RA 12 0.651 0.179
All four, pooled 32 0.817 [0.776, 0.855] 0.087 [0.078, 0.096]

Table 1. Logistic regression (M4) at the coverage threshold (OP-cov), for each test study and pooled across all four. Coverage is accepted cells divided by scorable cells. Accepted error is wrong accepted cells divided by accepted cells. Per-study values are single estimates without intervals. The pooled row has 95% donor-bootstrap intervals within the four studies. Source: study paper, section 3.1 and Table 2.

The other learned methods stayed close to logistic regression. kNN, CellTypist trained on the reference and scANVI differed from logistic regression by 0.036 or less in pooled accepted error at the coverage threshold.

From these results, the paper concludes that the test study mattered more for accepted error than the choice among these four methods. That conclusion does not cover the marker rule or nearest centroid. It also rests on only four studies, with no interval describing how much results vary from study to study.

HIHA needs one extra caution. Its authors report that CellTypist Immune_All models and Seurat label transfer guided their expert annotation (Gong et al., 2025). Agreement with HIHA labels therefore partly measures similarity to that labelling pipeline. On HIHA, released CellTypist's macro-F1 (0.842) was below logistic regression's (0.942).

Choosing a threshold: validation settings did not reliably carry over

Neither threshold rule consistently met its validation target on the test studies.

The coverage threshold was set to label 90% of validation cells. Pooled over the four test studies, no method reached 90% coverage at that threshold. For the six methods trained on the full reference, coverage ranged from 79.8% (CellTypist trained on the reference) to 88.6% (kNN). Single studies could exceed 90%: logistic regression covered 94.0% of HIHA and 90.6% of JDM.

The error threshold was set for 5% accepted error on validation cells. On the test studies, it gave accepted errors from 0 to 16%, while labelling very different shares of cells. Logistic regression labelled only 4.4% of cells, with no accepted errors. Nearest centroid labelled 41.6% of cells, with 3.2% accepted error.

Two methods had no error threshold at all. kNN and scANVI never reached 5% accepted error on the validation studies.

For logistic regression and CellTypist trained on the reference, the error threshold sat almost exactly at 1, the top of the confidence scale (0.9999998 for CellTypist). A threshold that extreme accepts only near-certain scores, so those two results are close to a trivial edge case.

Three scatter panels of accepted error against test coverage for each method at the coverage and error thresholds, with interval bars: the six reference-trained methods with the full reference, the same methods with pDC, ASC and MAIT removed, and the two released models

Figure 1. Accepted error against test coverage at the two thresholds chosen on validation studies. The horizontal axis is coverage (further right labels more cells); the vertical axis is accepted error (lower is better). Circles show the coverage threshold (OP-cov) and squares the error threshold (OP-err). Left panel: the six methods trained on the full reference (Arm A). Middle panel: the same methods retrained with pDC, ASC and MAIT removed (Arm B). Right panel: the released models P1 and P2 (hollow markers). Lines are recorded 95% donor-bootstrap intervals within the four test studies; they do not describe variation across new studies. M3 and M6 have no OP-err point. Figure 1 from the study paper (Rewire Bio, 2026), revision fd4588c.

Missing cell types still received labels from the reference

A method cannot give the correct label to a cell type that is absent from its reference. Ideally, such cells would receive low confidence and stay unassigned. In Arm B, at the coverage threshold, most of them were labelled anyway.

Arm B had 2,310 test cells of the three removed types, mostly from the extra rare-cell sample. At the coverage threshold, the retrained methods accepted a label from the reference for 65.9% (logistic regression) to 89.2% (kNN) of these cells. Every one of those accepted labels is wrong, because the correct type was no longer in the reference.

A separate count looked at all cells of each removed type, whether accepted or left unassigned. This count is a different measurement from the acceptance rates above. The most frequent predicted label was usually a close relative:

  • ASCs were most often labelled B cells.
  • MAIT cells were most often labelled γδ T or CD8 T cells.
  • pDCs were most often labelled conventional dendritic cells (cDC) by the marker rule, kNN and logistic regression.

The study also asked whether confidence scores ranked cells of known types above cells of removed types. It measured this with AUROC, where 0.5 means no separation and 1 means perfect separation. AUROC ranged from 0.441 to 0.650. Values for kNN, scANVI and CellTypist trained on the reference were below 0.5. These AUROC values have no interval.

The study interprets these results as showing that thresholds tuned to label 90% of validation cells failed to flag these near-neighbour types, meaning types closely related to types still in the reference, as unfamiliar. Distant or genuinely new cell types were not tested. Earlier work also found that options for rejecting uncertain cells can fail when a whole population is missing from the reference (Abdelaal et al., 2019).

Suggestions for a new study

These are my suggestions; none was tested.

  • Start with a simple learned classifier, such as logistic regression, trained on a reference similar to your own cohort.
  • Treat a pooled error rate as a loose guide to the error on your own study.
  • If part of your study already has labels, re-check the threshold's coverage and accepted error on that subset, as the paper recommends.
  • Do not rely on confidence to flag cell types missing from the reference. Compare the reference's classes with the types you expect.
  • Look at the unassigned cells. They include both low-quality cells and cells with low confidence. In the output of the released model below, decision and reason separate the two, second_label gives the runner-up label, marker_flag marks weak support from marker genes, and total_counts and genes_detected show cell quality.

Run the released model

The study provides its logistic regression model (M4) and a command-line annotator. The commands below download the code at revision fd4588c, check file hashes, build the locked software environment and label 12 example cells at the coverage threshold (OP-cov, threshold 0.7018). The commands need git, curl, make, shasum, uv and Python 3.11.13. The recorded outputs come from macOS on Apple Silicon.

mkdir celltransfer-demo && cd celltransfer-demo
git clone https://github.com/rewire-bio/cell-type-annotation-transfer.git study
git -C study checkout --detach fd4588c2b58eee214fea76b1874be1a37df1f0da
curl --fail --location --output m4-example.tar.gz \
  https://github.com/rewire-bio/cell-type-annotation-transfer/releases/download/study-v1-seed0-20261007/m4-example.tar.gz
echo "7b798df27f35440f5c2c01ab1f39f284896a31fb5229dd342393d914cd1cdf78  m4-example.tar.gz" | shasum -a 256 -c -
tar -xzf m4-example.tar.gz
(cd m4-bundle && shasum -a 256 -c SHA256SUMS)
(cd study && make env)
PYTHONPATH=study/companion/src study/.venv/bin/python -m celltransfer.cli annotate \
  --bundle m4-bundle --query example/blood12.h5ad \
  --operating-point coverage90 --out predictions.csv

The query, meaning the new data to be labelled, needs raw non-negative integer counts in .X (or --layer counts or --use-raw), unique cell IDs, Ensembl gene IDs and all genes. All genes are needed because library size is calculated from the full matrix. The output file predictions.csv has one row per cell, with these columns: cell_id, predicted_label, confidence, decision, reason, operating_point, threshold, second_label, second_confidence, marker_support_z, marker_flag, total_counts and genes_detected.

The 12 cells are the first rows of the RA test query (Binvignat et al., 2024). In the recorded output, every gene the model uses was present. All 12 cells were left unassigned with reason low_counts_or_genes. The cells had 144 to 236 counts and 129 to 169 detected genes, below the default quality screen of 500 counts and 200 genes.

The example therefore checks that the parts of the tool work together (model loading, input checks, output columns). The example gives no accuracy estimate. Keep the default quality screen on real data.

The commands were checked against the code at revision fd4588c and the recorded outputs but not run for this article. A fresh public clone has not yet been tested.

Limits of this evidence

  • The benchmark did not independently adjudicate the author labels. HIHA's labels were partly guided by CellTypist and Seurat, so agreement with HIHA is partly circular.
  • The benchmark used four fixed test studies, two of them with only 3 and 5 donors. The intervals describe variation among donors within these studies.
  • Each method used one training seed, so the benchmark did not measure how results vary across different seeds.
  • In the main test sample, only 20 cells (all erythroid, the lineage that produces red blood cells) belonged to a type outside the reference. The three removed types were closely related immune types, and most of their test cells came from the extra rare-cell sample.
  • The released models differ from the reference-trained methods in training data and cell-type vocabulary, so comparisons with them cannot isolate the effect of model design.
  • Two automated model reviews (one of the methods, one of the paper) by one Claude Opus model session returned PASS. There has been no peer, human or external review.
  • CellTypist trained on the reference was refitted with seed 0 after its first fits left out the random seed set in the protocol. A private reproduction of the refitted results was compared with the canonical run by the same pipeline: 1,521 checks, 0 gated breaches. The paper calls this an internal comparison. Nothing has been independently verified or replicated.

References

Frequently asked

Which annotation method should I start with for a new blood single-cell study?
In this benchmark of four blood studies kept out of training (32 donors), the study could not establish a clear difference in agreement with author labels among logistic regression, k-nearest neighbours, CellTypist trained on the study reference and scANVI. A marker rule and nearest centroid agreed less. No equivalence margin was set, so the results do not show that the four methods are equivalent. Starting with a simple learned classifier trained on a reference similar to your own cohort is my suggestion; the study did not test it.
How should I choose a confidence threshold for leaving cells unassigned?
In this benchmark, thresholds chosen on two validation studies did not reliably carry over to the four test studies. Pooled over those four studies, the threshold set to label 90% of validation cells labelled 79.8% to 88.6% of test cells for the six methods trained on the full study reference, although logistic regression exceeded 90% on two single studies. The threshold set for 5% error on validation cells gave accepted errors on the test studies from 0 to 16%. The paper suggests re-checking any threshold on a part of the new study that already has labels, if one exists. That suggestion was not tested.
Does a confident label mean the cell type is in the reference?
No. In a second experiment, three cell types were removed from the reference and the methods were retrained. On the four test studies, at the threshold set to label 90% of validation cells, the retrained methods still accepted a label from the reference for 65.9% to 89.2% of the test cells of those types, and every such label is wrong. A separate count over all cells of each removed type, accepted or not, found that the most frequent predicted label was usually a close relative, such as B cells for antibody-secreting cells. Only three closely related immune types were tested.
What does the released 12-cell example show?
The example shows the command-line annotator loading the study's logistic regression model, checking its input and writing its output columns. In the recorded run, the default quality screen left all 12 example cells unassigned, because their counts (144 to 236) and detected genes (129 to 169) were below the defaults of 500 counts and 200 genes. The example gives no accuracy estimate. The commands were checked against the code at revision fd4588c and the recorded outputs but not run for this article.
How reliable is this evidence?
The evidence comes from one recorded computational benchmark, carried out by Rewire Bio, which also publishes rewire.it. The benchmark has not been peer reviewed. Two automated model reviews by a Claude Opus session, one of the methods and one of the paper, returned PASS. CellTypist trained on the reference was refitted with seed 0 after its first fits left out the random seed set in the protocol. In the study's records, the author's private reproduction of the refitted results was compared with the canonical run by the same pipeline: 1,521 checks, 0 gated breaches, reproduced within pre-specified tolerance. The paper calls this an internal comparison. No human or external review is claimed, and a fresh public clone has not been tested.
What did this benchmark not test?
The benchmark did not run Pan-Human Azimuth or scGPT, did not estimate variation across training seeds, did not independently check the author labels and did not test distant or genuinely new cell types. Only three closely related immune types were removed from the reference. The main test sample, without the extra rare cells, held only 20 cells of a type outside the reference, which is enough for counts only. The paper suggested re-checking thresholds on labels from the new study but did not test that step.

Help improve this article

Found an error or a better source? Leave a note here, or highlight a passage to comment on it.