Single-Cell Is About to Have Its Placenta Microbiome Moment | Blog Skip to main content
A dark field of glowing translucent droplets receding into depth, each holding a single luminous cell in gold and magenta, with loose cyan strands of free-floating RNA drifting through the spaces between them.

Single-Cell Is About to Have Its Placenta Microbiome Moment

Published on 12 min read

Earlier this month I wrote up Pan-human Azimuth, a preprint that annotates 27 million cells with a model small enough to run on a laptop, and I liked it. It does something most papers don’t: it trains on negative examples, real empty droplets scraped from public datasets, so the classifier can look at a profile and say that isn’t a cell.

I kept mulling that over, because it’s the most rigorous thing in the paper and it still isn’t a control. And the reason I couldn’t let it go is that I spent years in microbial ecology, which had this exact argument, in public, for about twenty years, and lost a whole research program to it before the fix stuck.

I don’t think single-cell is heading somewhere as bad. But I think it’s heading somewhere similar, and it’s worth saying why while the field still has time to be boring about it.


What the microbiome learned, expensively

In 2014, Salter and colleagues demonstrated that DNA extraction kits and lab reagents carry their own bacterial DNA. Not trace amounts you could wave off, but enough that in samples with little genuine microbial biomass, the community you sequence is substantially the community that came in the kit. Their conclusion was blunt: sequencing negative controls alongside your samples is “strongly advised.”

The reason that landed so hard is that people had been building science on the other assumption. The most famous casualty is the placental microbiome. There was a real research program here, with real funding and real papers, arguing that the placenta harbors its own bacterial community. That would have been a big deal, since the fetus was supposed to develop in a sterile environment. In 2019, a Nature study went looking with contamination properly controlled for and found no evidence of bacteria in the large majority of placental samples. Almost every signal traced back to either reagent contamination or bacteria picked up during labor and delivery. The one real exception was group B Streptococcus, in about 5% of samples taken before labor began, which matters clinically but is not a microbiome.

The field’s response was cultural, not just technical: expected negative controls, mock communities with known composition, and tools like decontam that use those blanks to statistically strip contaminant sequences. Extraction blanks became something reviewers ask for.

That correction ran about twenty years, from a 1998 report that PCR reagents carry their own 16S to blanks being standard in the journals that matter.


Why I think the analogy holds

The obvious objection is that bacteria in a kit is a different problem from RNA in a droplet. Fair. But the underlying condition is the same, and it’s the condition that matters: when the true signal is small, whatever noise floor you have stops being a nuisance and starts being your result.

Microbiome’s low biomass was a swab with almost nothing on it. scRNA-seq’s low biomass is a cell with very little RNA in it. A neutrophil has a fraction of the transcript content of a macrophage; it’s fragile, it degrades fast, and it sits in a droplet suspension full of ambient RNA released by the cells that lysed during dissociation. When you profile it, some meaningful share of what you read didn’t come from that cell.

And there’s a second noise source, sharper here than in microbiome. Microbiome has its own version: fecal samples bloom in transit and the field subtracts the bloomers. But that changes who is present. In scRNA-seq the measurement changes the cell’s state. Van den Brink and colleagues showed in 2017 that warm collagenase dissociation induces a stress transcriptional program — Fos, Jun, heat-shock genes — and that it hits some subpopulations harder than others. The striking part is that the dissociation induces the same genes that actual tissue injury does. So a cell state you discover can be a cell state you manufactured, and the more finely you subdivide, the more likely that becomes.

To be fair to single-cell, the field named the ambient problem early and tooled it early. EmptyDrops decides which barcodes hold real cells. SoupX and CellBender estimate the ambient profile and subtract it. That’s genuinely ahead of where microbiome was in 2010. My worry isn’t that nobody noticed. It’s about what’s still missing.


The thing that’s missing is a positive control

Almost nobody runs a no-cell blank library alongside each sample the way a microbiome study runs an extraction blank. And on the positive side it’s worse: the canonical purity experiment is the human/mouse “barnyard” plot, from the species-mixing experiments that validated Drop-seq. Barcodes carrying both genomes give you the doublet rate; the wrong-species transcripts inside a single-species barcode give you purity, which fell from 98.8% to 90.4% as they loaded more cells. Their diagnosis of the largest source of that impurity, in 2015, was ambient RNA from cells damaged during preparation. It’s a good experiment. It gets run once, at method validation, by whoever is building the platform. It is not a per-sample anchor. The closest thing in routine use is multiplexing: cell hashing, or genetic demultiplexing off natural variation, hands you an empirical collision rate for that run. But collisions aren’t ambient. For ambient background there’s no per-sample anchor at all.

The part I find genuinely interesting is where the annotation models come in.

A lot of single-cell QC works by checking a cell against marker genes: does this cell express what this cell type is supposed to express, and not express what it shouldn’t? That’s the shape of the filter in that Azimuth post, which scores every cell one-vs-all against roughly 100 positive and negative markers and drops anything below a threshold.

Note the vocabulary trap, because the papers invite it. Positive and negative markers are genes expected to be on or off in a cell type. They are not positive and negative controls. And the markers come from the same reference that assigned the label in the first place.

That’s the whole problem in one sentence. A control has to be independent of the thing it’s checking, and reference-derived markers are not independent of reference-derived labels. What you have is a consistency check: it will tell you a cell doesn’t look like what you called it, which is useful. It cannot tell you whether the category was right, whether the profile is half ambient RNA, or whether the state you’re seeing was induced by your protocol. The missing property isn’t rigor. It’s independence.


Where this gets uncomfortable

I want to make this concrete rather than leave it as a vibe, so: three places I’d actually poke.

The negatives are platform-flavored. When a model trains on empty droplets to learn what junk looks like, those droplets came from a particular chemistry, usually 10x Chromium, often from a handful of tissue contexts. How well a reject class learned that way survives a change of platform is barely tested. The paper’s own evidence is two incidental data points, both in my last post: 0.3% Unassigned on Tabula Sapiens, which mixes droplet and plate chemistries, and 69% on Visium HD bins. One shrug and one refusal, with nothing designed in between. The adjacent evidence is not reassuring. A 2026 benchmark of six ambient-RNA removal tools across droplet and well-plate data found that two of them don’t denoise so much as rebuild the matrix: CellClear replaces over 93% of counts with values derived from matrix factorization, and scAR invents cell types absent from the uncorrected data, including three spurious coarse types in the BD Rhapsody dataset and up to eight in prefrontal cortex. Only two of the six even run on non-droplet platforms without raw count access. Portability across chemistries is clearly not free. So run a reject class somewhere new and if it under-fires you get confident labels on garbage; if it over-fires you silently lose real cells. Neither failure announces itself.

Simulated negatives model soup, not doublets. Generating synthetic negatives by averaging a tissue’s abundant cells gives you a decent model of ambient RNA. It’s a poor model of two real cells stuck together, because two cells stuck together is not the mean of many cells. DoubletFinder builds its artificial doublets by averaging random pairs, which is the right shape. And most of the doublet problem doesn’t need modeling at all: cell hashing tags each sample with its own antibody barcode, so a barcode carrying two tags is empirically a multiplet. Only the cross-sample ones, and the hashing paper itself names same-sample doublets as an open problem for the field. But that’s a real measurement, sitting right there, unused.

And the flattering interpretation is not the only one. In my last post, Azimuth disagreed with human curators most often on cells the humans had called neutrophils, and those cells showed no coherent marker signature at all. The obvious reading is that the model caught human error, and I wrote it up that way. The less flattering reading runs the other way: those were real neutrophils, too RNA-poor for the model to see anything in, and the reject class deleted them.

That isn’t a stretch, because neutrophils are the population this fails on first. The same 2025 comparison that found 10x, Parse, and HIVE all capture them well is explicit about why it had to loosen its own filters: neutrophils “generally have low levels of RNA meaning that they would be filtered out using strict thresholds that have been developed for PBMCs.” A reject class trained on low-count background is that same threshold, learned instead of typed. I don’t know which reading is right, and the frustrating part is that it’s answerable. Protein would settle it: if those barcodes carry CD66b on their surface, they’re real granulocytes no matter how thin the transcriptome looks. The measurement exists. It just wasn’t run, because nothing in the workflow asks for it.

That’s the tell. When the same observation supports “our method is better than humans” and “our method quietly deleted the hardest population in the tissue,” and you can’t distinguish them, you’ve found a place where the field needs a control it doesn’t have.


Why now and not five years ago

The microbiome’s contamination problem didn’t bite when people were asking “is there a gut community and what’s in it.” Dominant taxa are robust; a bit of kit contamination doesn’t change your answer. It bit when the field moved to the rare biosphere — low-abundance taxa, low-biomass sites, subtle differences between conditions. That’s when the noise floor and the signal met.

Single-cell is making the same move right now. “Which cell types are in this tissue” is a robust question, and the field answered it. The questions now are finer: distinguishing states within a type, comparing the same cell type across organs, resolving what’s inside an 8 µm spatial bin. Those are all low-signal questions, and they’re exactly the regime where the noise floor starts producing results.

One example from that paper: it finds a fibroblast state that’s essentially lung-specific, present in over 60% of lung samples and under 0.1% of everything else. The marker genes are G0S2, which is a stress and lipolysis gene, and PPP1R14A, which is smooth-muscle associated. And lung is a tissue where people write whole protocol papers about minimizing dissociation stress. That one benchmarked five protocols and found the choice moves the stress transcripts you induce, the surface markers that survive, and, most relevant here, which fibroblasts you get back at all: a rapid collagenase prep returned proportionally fewer CD90+ fibroblasts than a liberase-elastase prep, which had them at 8.6% of the non-immune population. A stress-flavored, lung-restricted fibroblast state is also precisely what a dissociation artifact would look like. Neither marker is on van den Brink’s dissociation signature, which is the honest caveat. The authors have a real defense — the pattern reproduces in an independent atlas — and I think the result is probably real. But “probably real” is where you end up without a protocol-matched control, and it’s not where anyone wants to be.


The boring fixes

None of this requires a breakthrough. The microbiome’s answer was unglamorous and it worked.

Run blanks. A no-cell library on the same chemistry, same day, same operator, gives you a per-run picture of your ambient and reagent background. Microbiome studies are told to do this per batch, and a 2025 review of 243 insect-microbiota papers found two-thirds of them still didn’t, which is the honest version of the lesson: naming a control is not the same as running it.

Hold out something orthogonal. The cheapest real positive control is a population confirmed at the protein level rather than the transcript level, whether FACS-sorted on surface markers or CITE-seq confirmed, then kept out of training and used to score annotations against truth that didn’t come from a transcriptome reference.

Test abstention adversarially. Feed an annotation model data containing none of the things it claims to find: a pure cell line, a non-human sample, scrambled counts. Check that it declines. A 2025 benchmark ran the tidy version, six out-of-distribution detection methods against held-out cell types and protocol shifts, and found novel types get caught reliably while broader data shifts are harder. The crude version costs an afternoon and would tell you more about a released model’s reject class than any benchmark table.

Report the protocol as a variable. Dissociation method, warm or cold, enzyme, time. If a state only shows up under one protocol, that’s worth knowing before it becomes a cell type.


I want to be careful not to overclaim here, because the honest position is that single-cell is roughly where microbiome was after Salter, not before. The problem is named. The tools exist. What hasn’t happened is the cultural step where controls stop being something a careful lab does and start being something reviewers require.

The microbiome field got as far as it has by losing an argument in public and having to re-read a decade of results. That’s an expensive way to learn it, and the lesson is already written down. Single-cell doesn’t need its own placenta.

What it needs is for the next big atlas paper to include a blank.

Tags:

← Back to Blog