Journal Safari: Malva, An Index Is a Subscription
Published on • 14 min read
You can hand Malva a 24-letter string and get back a list of cells. Not samples, not studies. Cells, each one carrying the tissue it came from, the donor, the disease state. It answers in milliseconds, and what it searches is an index built from 140 terabytes of raw sequencing reads that the rest of the field already threw away.
Ultrafast and reference-free sequence discovery in single-cell data (León-Periñán, Karaiskos & Rajewsky, MDC-BIMSB Berlin, Nature, 2026) is the first paper in this series where the contribution is a running service rather than a result. A method either reproduces or it doesn’t. A service either still exists in 2036 or it doesn’t, and no experiment in this paper can tell you which.
The engineering is excellent and the peer review visibly improved it. My argument is about what happens next, and three of the four referees made it first.
Why This Paper
Here is a question you can’t currently ask: is there Epstein-Barr virus in this nasopharyngeal carcinoma atlas? The count matrix has no EBV genes in it, so you’d have to find the raw FASTQs, build a custom reference, realign and requantify. Days to weeks, for one question. Then somebody asks about cytomegalovirus.
Atlases like CZ CELLxGENE hand you a cell-by-gene matrix. Somebody upstream already picked a reference and an aligner and kept only the counts, so everything off-reference is gone: fusions, viral and bacterial reads, editing sites, non-reference haplotypes, backsplice junctions. Meanwhile BLAST searches databases of genomic sequences, and k-mer indexes like MetaGraph and Fulgor search large sequencing collections but are, in the paper’s words, “bulk, sample-centric.” A hit tells you a database entry or an experiment, not the cell.
So: atlases with no sequence search, sequence search with no single-cell resolution. Malva is the join.
That framing comes straight off the paper’s Figure 1, and the gap is smaller than it looks. sc-SPLASH (Dehghannasiri et al., Nature Biotechnology, 2026) already does statistics-first, reference-free discovery on barcoded scRNA-seq and spatial data, and its preprint went up nine months before Malva was submitted. Malva doesn’t cite it, and none of the four referees raised it. What survives is narrower but real: sc-SPLASH finds things in a dataset you already have, and Malva is an index you query across a corpus you don’t. The paper’s own sentence is hedged more carefully than its artwork: “To our knowledge, Malva is the first platform for large-scale reference-free sequence search and analysis at single-cell resolution.”
What They Actually Did
The coffee-length version: chop every read into 24-letter chunks and remember which cells each chunk appeared in.
The observation everything rests on. Bulk k-mer indexes work over maybe 10² to 10⁴ labels, and many k-mers carry the identical label set, so the winning move is to deduplicate those sets, as MetaGraph’s Multi-BRWT encoding and Fulgor’s unitig monochromaticity both do. Single-cell pushes the label count to 10⁶ and beyond, and the authors claim, in a supplementary note with no distribution behind it, that a typical k-mer appears in a few hundred cells out of millions. Those sets are almost never shared, so the clever compression isn’t as useful.
So the design is a series of refusals. No de Bruijn graph: sample k-mers at non-overlapping positions. One sparse posting list per k-mer, holding only the cells that contain it. Index each sample independently and merge, so adding data costs under 10% of the CPU of a full reindex. The on-disk format is textbook: packed 64-bit k-mers, compressed sparse rows, delta-encoded blocks. The contribution is knowing which pieces to pick for a label space the existing indexes weren’t built for.
Non-overlapping sampling bites users. Indexed k-mers sit at positions 0, 24, 48 and so on. Queries tile the probe into windows of w and call a cell positive if at least τ of a window’s indexed k-mers match. The least intuitive consequence: for a splice junction you make the window smaller. Query a junction-centred 48 bp probe at w = 48, and both indexed k-mers sit inside the flanking exons: you’ve built a detector for “expresses this gene.” Shrink the window to 24 with τ = 1.0 and slide it, and some position is forced across the break. SNVs are worse, because a longer probe picks up k-mers that don’t span the variant and match wild-type happily. As with allele-specific qPCR probes, shorter ones discriminate better.
The corpus is the real artefact. A crawler watches GEO, SRA, ENA, the Human Cell Atlas portal and CNGBdb and pulls FASTQ. An instruction-tuned Llama 3.1 takes a first pass at the deposited metadata, a human curates it, and text2term maps the result into MONDO, UBERON and PATO. As of May 2025 the Index covered around 80% of the Human Cell Atlas and 20% of human single-cell SRA runs. Hold onto that human in the loop: it’s a recurring cost, and it doesn’t obviously parallelise.
The Results That Convinced Me
Two of them, and neither is a benchmark.
The controls are designed properly. PhiX is the perfect positive: present in nearly every Illumina run, and put there by no biology. Mycoplasma is a known positive, since a 2015 survey of the SRA had already found culture contamination; Malva recovers it in 0.5–2% of samples, concentrated at the 23S rRNA locus. Five pathogenic viruses nobody put in these libraries come back clean.
Except for one hit. SARS-CoV-2 turned up in lung samples labelled healthy, and it was a cell line from healthy donors, infected in vitro. The metadata says healthy because the donor was. The sequence says infected because the cells were. That’s the argument for the whole enterprise: annotations are a lossy, out-of-date summary of what was in the tube, and sequences aren’t.
It found bacteria nobody was looking for. They MinHash the 16-mers inside each 24-mer into 100,000 buckets, shrinking an unusable cells-by-billions-of-k-mers matrix into something ordinary clustering can handle. On a head-and-neck tumour section, that reproduces the tumour and stromal compartments and adds two clusters the gene view has no name for.
Assemble the discriminative k-mers with SPAdes and read the contigs back. One cluster is human rRNA, which reference preprocessing routinely strips out. The other, outside the tissue border, is bacterial 23S rRNA from species that occur in the oral mucosa: Kraken2 calls oral commensals including G. morbillorum and C. gingivalis. Oral flora, beside a head-and-neck tumour.
A reference pipeline can find that too, if you suspect it first. INVADEseq (Galeano Niño et al., Nature, 2022) mapped intratumoral bacteria in oral squamous cell carcinoma, and SAHMI exists to separate real microbial signal in single-cell data from contaminants. But those start from a hypothesis, on one study. Here it fell out of clustering, on a corpus, because the non-human reads were never discarded.
What I’d Push On
The number it gives you back is not a count. The index is presence/absence (the rebuttal says each k-mer-and-cell pair is stored “at most once, regardless of read-level frequency”), and what comes back is the number of query windows that matched in a cell. The paper says those pseudocounts “correlate with read counts rather than molecule counts,” and corrects them with one PCR bias factor for the whole library, “the mean R/U ratio across cells,” which by their own description doesn’t address “gene- or sequence-specific amplification biases.” Add no barcode error correction and no doublet removal, and the authors’ own word, semi-quantitative, is right. Trust the cell list, not the abundance.
The microbial calls rest on one of the most conserved sequences in biology. The Mycoplasma and Gemella results both live at 23S rRNA, and the paper’s own wording is careful: contigs “compatible with bacterial 23S rRNA from several species that are known to occur in the oral mucosa.” Cross-species specificity is tested inside one genus, where the negative control is Mycoplasma agassizii, “a respiratory pathogen of turtles.” That control column isn’t empty, either: the largest mark in the kidney row sits under the turtle pathogen, with smaller hits in skin and brain, and the paper doesn’t mention them. Nothing comparable tests the oral-flora call, where the species is the claim. I’d call both results genus-level, and a day’s work would settle it: count how many of the discriminative contigs’ 24-mers are unique to that species across bacterial RefSeq.
Check the benchmark’s scope. The 70 ms for a single k-mer and 400 MB working set come from Extended Data Fig. 2d, one Visium mouse-brain sample; for the full Index the paper claims “milliseconds.” And “comparable query throughput” to Fulgor is qualified as “from disk,” while the numbers live in Supplementary Table 1: 133 seconds per megabase against Fulgor’s 51, with Fulgor holding 326 GB in RAM to Malva’s 0.3 GB. Malva made the right trade. The main text just doesn’t show you the price.
Controlled-access data can never be in it. The clinically interesting data sits in dbGaP, EGA or a hospital, where you need approval to download it, so it can’t be crawled into a public index. Referee #1 raised this in September 2025. The authors’ answer was standalone binaries for building a private index from data you already hold. That keeps your data off their servers, but a local index only searches what you brought, not the corpus.
An Index Is Not a Result
The paper invites the comparison in its second paragraph: “Thirty-five years ago, BLAST pioneered this principle…” So take it seriously.
BLAST didn’t win on optimality. Its 1990 paper sold an approximation on being “an order of magnitude faster than existing sequence comparison tools of comparable sensitivity,” and it kept winning for institutional reasons. NCBI is a standing federal institute, GenBank shipped its 273rd release in August 2026, and the corpus carries a written promise. Under the INSDC policy, shared by GenBank, ENA and DDBJ since 2002, “no use restrictions or licensing requirements will be included in any sequence data records,” and “All database records submitted to the INSD will remain permanently accessible as part of the scientific record.” That permanence is exactly what Malva lacks.
Now the part that cuts against me. AB-BLAST, whose documentation traces its lineage to “the original gapped BLAST (WU-BLAST) in 1996,” says “a Commercial Site license is required for all for-profit entities,” and it had users anyway: its FAQ still fields questions from people holding a “previously licensed copy of WU-BLAST.” So a restrictive licence doesn’t stop adoption, and Malva’s has the same shape: commercial use needs “a separate written license,” by inquiry to the first author. It also excludes “clinical use, diagnostics use,” and names no route for those.
That pushes the weight off the code and onto the corpus. Whichever BLAST you ran, the databases were free to download and search on your own hardware, so NCBI being slow, down or uninterested in your organism was survivable. The algorithm was replaceable. The archive was guaranteed.
Malva has it the other way round. The engineering is better than BLAST’s ever was, and the code is now readable. The institutional picture is thinner: one lab, grant-funded, running the corpus on its own hardware with query limits; a patent application (WO2026/132228, disclosed in the competing-interests statement); an academic-use licence. The 74 million cells are not something you can download.
An index is not a result — it is a subscription.
What Three Referees Bought
Nature published the complete correspondence under CC BY: four referees and five rounds between submission on 17 September 2025 and acceptance on 30 July 2026. The peer review statement names Rob Patro, senior author of Fulgor, the tool Malva is benchmarked against. It’s the best thing attached to this paper.
September 2025. The referees get the API client and nothing else. Referee #1’s first listed concern: “It appears that Malva is intended for commercialization.” Referee #2 notes the core code “seems to be nowhere made available,” asks “Is the intent for that code to remain proprietary?”, and makes the reviewer’s bargain explicit: give me the code, or give me enough detail to reproduce it.
Round one. Standalone binaries, a supplementary chapter of pseudocode, and a cover letter saying the methods were expanded “to make the procedures reproducible, even though the main code remains proprietary.” Referee #4 names the norm: the approach “does not fully align with common practices in the bioinformatics community, where open-source code sharing is strongly encouraged.”
Round two. The response letter: “we have decided to release the full source code under an Academic-use license.”
The process bought more than a licence. The pan-cancer somatic mutation benchmark, the polyASite validation, the sequencing-error robustness analysis, the EGFR lung-cancer cohort and the supplementary chapter on indexing: none of it was in the original submission, and almost all of it arrives in the first rebuttal, the same letter that says the code stays proprietary. The science got better a full round before the licence moved.
So why capitulate? They wanted commercialisation badly enough to file a patent and withhold the core code from referees, then gave up the code and kept the patent. Maybe a Nature paper is worth more than closed source; the authors never say why, anywhere in 22,000 words of correspondence. What you can see is what survived revision. The code moved, and the patent, the licence and the single deployment didn’t. Those were never things a referee could fix, because they aren’t claims. They’re commitments.
And “open source” is doing more work than it can carry. The repository licence grants use “solely for internal academic purposes.” That fails clause 6 of the Open Source Definition, which says a licence “must not restrict anyone from making use of the program in a specific field of endeavor” and gives “genetic research” as its own example. So this is source-available: a real improvement over a binary, and not the same thing. It also means the tool that found bacteria beside a tumour isn’t licensed for diagnostic use.
The Data-Infrastructure Coda
I do data infrastructure for a living and TileDB-SOMA is my day job, so treat this as an interested party’s reading.
Read the Malva Index section of the Methods with the vocabulary swapped and it describes a sparse array store: sorted keys, offsets, fixed-size compressed blocks, memory-mapped range reads, a block cache, and independent write units merged in the background. They reinvented it well, and it’s one of four artefacts they maintain: the k-mer index, a MinHash matrix for the transpose, pre-computed gene matrices for “how much beta-actin is in this cell,” and a relational metadata service, because the index holds integer cell IDs and people ask about tissues. That’s four representations of one corpus, each with its own staleness clock. In the SOMA model the last of those already exists beside the counts, as obs, so a sequence index parked next to it would only need to hand back cell IDs that join. Whether arrays would win on the rest is a benchmark I’d want to run, not a result I can claim.
What I will claim is the version pin. The corpus is crawled continuously, so even the paper’s counts differ: around 74 million cells in the abstract, about 51 million human and 10 million mouse in v2025.04a. The API reports “the current version and statistics,” so you can find out which index you just hit. I can’t find any way to pin a query to a past version, and that’s the difference between an observation and a citation. “I found 4,312 cells expressing X” should still mean 4,312 cells next year.
The speed genuinely impressed me, but what stayed with me is that all of this was already sitting there. The reads were deposited, the barcodes were in them, and the oral bacteria sat outside that tumour section the whole time. They weren’t visible because a reference-based pipeline discards whatever the reference doesn’t recognise, which isn’t a bug in anyone’s code. It’s what the output format is for. Malva didn’t generate a single new base — it declined to discard.
That leaves durability as the question that matters, and it’s the one the paper can’t touch. BLAST is thirty-five years old, and I can still download its databases and search them on my laptop without asking anyone, because an institution outlived every grant that touched it. Somebody has to keep crawling Malva’s corpus too, and a human still curates its metadata. The paper is an excellent argument that it’s worth doing, and it is silent, necessarily, on who pays for it in 2036.
Related Posts
Journal Safari: Seven Million Parameters and a MacBook Air
Journal Safari #3. Pan-human Azimuth puts 27 million cells from 23 tissues on one annotation tree, then classifies them with a model 70× smaller than the foundation model I read in July — and wins. The contribution is not the model. It is the labels.
Journal Safari: Can You Compress a Cell Into a Word?
Kicking off Journal Safari, my biweekly close-read of one recent paper. First up: CellVQ, a 500-million-parameter single-cell foundation model that squashes every cell into a discrete code — and the peer reviews that quietly agreed with my doubts.
Journal Safari: Ask Eleven Models, Keep the Worst Answer
Journal Safari #4. Eleven ESM protein language models were supposed to be near-interchangeable. They have different blind spots that do not overlap, and training each one on the others' best guesses puts sequence-only at the top of the ClinVar board. The gain comes off about 200 proteins, which is the part I keep chewing on.