Journal Safari: Malva, An Index Is a Subscription | Blog Skip to main content
A vast Victorian records hall of card-catalogue drawers receding into cold blue darkness. One man on a plain wooden ladder leaning against the cabinet files a single card into the one drawer under the lamp. Open drawers packed with tabbed cards continue unlit down the aisle beside him, while every other wing of the same catalogue is draped in grey dust sheets and hung with cobwebs.

Journal Safari: Malva, An Index Is a Subscription

Published on • 14 min read

You can hand Malva a 24-letter string and get back a list of cells. Not samples, not studies. Cells, each one carrying the tissue it came from, the donor, the disease state. It answers in milliseconds, and what it searches is an index built from 140 terabytes of raw sequencing reads that the rest of the field already threw away.

Ultrafast and reference-free sequence discovery in single-cell data (León-Periñán, Karaiskos & Rajewsky, MDC-BIMSB Berlin, Nature, 2026) is the first paper in this series where the contribution is a running service rather than a result. A method either reproduces or it doesn’t. A service either still exists in 2036 or it doesn’t, and no experiment in this paper can tell you which.

The engineering is excellent and the peer review visibly improved it. My argument is about what happens next, and three of the four referees made it first.


Why This Paper

Here is a question you can’t currently ask: is there Epstein-Barr virus in this nasopharyngeal carcinoma atlas? The count matrix has no EBV genes in it, so you’d have to find the raw FASTQs, build a custom reference, realign and requantify. Days to weeks, for one question. Then somebody asks about cytomegalovirus.

Atlases like CZ CELLxGENE hand you a cell-by-gene matrix. Somebody upstream already picked a reference and an aligner and kept only the counts, so everything off-reference is gone: fusions, viral and bacterial reads, editing sites, non-reference haplotypes, backsplice junctions. Meanwhile BLAST searches databases of genomic sequences, and k-mer indexes like MetaGraph and Fulgor search large sequencing collections but are, in the paper’s words, “bulk, sample-centric.” A hit tells you a database entry or an experiment, not the cell.

So: atlases with no sequence search, sequence search with no single-cell resolution. Malva is the join.

That framing comes straight off the paper’s Figure 1, and the gap is smaller than it looks. sc-SPLASH (Dehghannasiri et al., Nature Biotechnology, 2026) already does statistics-first, reference-free discovery on barcoded scRNA-seq and spatial data, and its preprint went up nine months before Malva was submitted. Malva doesn’t cite it, and none of the four referees raised it. What survives is narrower but real: sc-SPLASH finds things in a dataset you already have, and Malva is an index you query across a corpus you don’t. The paper’s own sentence is hedged more carefully than its artwork: “To our knowledge, Malva is the first platform for large-scale reference-free sequence search and analysis at single-cell resolution.”


What They Actually Did

The coffee-length version: chop every read into 24-letter chunks and remember which cells each chunk appeared in.

The observation everything rests on. Bulk k-mer indexes work over maybe 10² to 10⁴ labels, and many k-mers carry the identical label set, so the winning move is to deduplicate those sets, as MetaGraph’s Multi-BRWT encoding and Fulgor’s unitig monochromaticity both do. Single-cell pushes the label count to 10⁶ and beyond, and the authors claim, in a supplementary note with no distribution behind it, that a typical k-mer appears in a few hundred cells out of millions. Those sets are almost never shared, so the clever compression isn’t as useful.

So the design is a series of refusals. No de Bruijn graph: sample k-mers at non-overlapping positions. One sparse posting list per k-mer, holding only the cells that contain it. Index each sample independently and merge, so adding data costs under 10% of the CPU of a full reindex. The on-disk format is textbook: packed 64-bit k-mers, compressed sparse rows, delta-encoded blocks. The contribution is knowing which pieces to pick for a label space the existing indexes weren’t built for.

Non-overlapping sampling bites users. Indexed k-mers sit at positions 0, 24, 48 and so on. Queries tile the probe into windows of w and call a cell positive if at least τ of a window’s indexed k-mers match. The least intuitive consequence: for a splice junction you make the window smaller. Query a junction-centred 48 bp probe at w = 48, and both indexed k-mers sit inside the flanking exons: you’ve built a detector for “expresses this gene.” Shrink the window to 24 with τ = 1.0 and slide it, and some position is forced across the break. SNVs are worse, because a longer probe picks up k-mers that don’t span the variant and match wild-type happily. As with allele-specific qPCR probes, shorter ones discriminate better.

The corpus is the real artefact. A crawler watches GEO, SRA, ENA, the Human Cell Atlas portal and CNGBdb and pulls FASTQ. An instruction-tuned Llama 3.1 takes a first pass at the deposited metadata, a human curates it, and text2term maps the result into MONDO, UBERON and PATO. As of May 2025 the Index covered around 80% of the Human Cell Atlas and 20% of human single-cell SRA runs. Hold onto that human in the loop: it’s a recurring cost, and it doesn’t obviously parallelise.


The Results That Convinced Me

Two of them, and neither is a benchmark.

The controls are designed properly. PhiX is the perfect positive: present in nearly every Illumina run, and put there by no biology. Mycoplasma is a known positive, since a 2015 survey of the SRA had already found culture contamination; Malva recovers it in 0.5–2% of samples, concentrated at the 23S rRNA locus. Five pathogenic viruses nobody put in these libraries come back clean.

Dot plot with human tissues as rows and sequence probes as columns. The left block shows HERV-K, PhiX, lentiviral vector, and expected versus negative-control Mycoplasma species, with PhiX dotted densely down nearly every tissue row. The right block shows SARS-CoV-2, hepatitis C, measles, West Nile and Zika, which is almost entirely blank except for a single pair of dots in the Lung row under SARS-CoV-2, both drawn as circles, which the legend marks as healthy samples.
Specificity is the empty half (Fig. 3a, León-Periñán et al. 2026, Nature, CC BY 4.0). Dot size is the fraction of positive cells. The Myco− column holds the negative-control species, and it is not blank.

Except for one hit. SARS-CoV-2 turned up in lung samples labelled healthy, and it was a cell line from healthy donors, infected in vitro. The metadata says healthy because the donor was. The sequence says infected because the cells were. That’s the argument for the whole enterprise: annotations are a lossy, out-of-date summary of what was in the tube, and sequences aren’t.

It found bacteria nobody was looking for. They MinHash the 16-mers inside each 24-mer into 100,000 buckets, shrinking an unusable cells-by-billions-of-k-mers matrix into something ordinary clustering can handle. On a head-and-neck tumour section, that reproduces the tumour and stromal compartments and adds two clusters the gene view has no name for.

Spatial map of a head and neck tumour section, with a dashed line marking the tissue border. Numbered clusters in blue, green, black, pink and purple fill the tissue interior. Two extra clusters are called out separately in the legend: KS1 in red at the upper right, sitting outside the dashed tissue border, and KS2 in yellow running along the lower edge of the section.
The two clusters with no gene-level name (detail from Fig. 5e; León-Periñán et al. 2026, Nature, CC BY 4.0).

Assemble the discriminative k-mers with SPAdes and read the contigs back. One cluster is human rRNA, which reference preprocessing routinely strips out. The other, outside the tissue border, is bacterial 23S rRNA from species that occur in the oral mucosa: Kraken2 calls oral commensals including G. morbillorum and C. gingivalis. Oral flora, beside a head-and-neck tumour.

A reference pipeline can find that too, if you suspect it first. INVADEseq (Galeano Niño et al., Nature, 2022) mapped intratumoral bacteria in oral squamous cell carcinoma, and SAHMI exists to separate real microbial signal in single-cell data from contaminants. But those start from a hypothesis, on one study. Here it fell out of clustering, on a corpus, because the non-human reads were never discarded.


What I’d Push On

The number it gives you back is not a count. The index is presence/absence (the rebuttal says each k-mer-and-cell pair is stored “at most once, regardless of read-level frequency”), and what comes back is the number of query windows that matched in a cell. The paper says those pseudocounts “correlate with read counts rather than molecule counts,” and corrects them with one PCR bias factor for the whole library, “the mean R/U ratio across cells,” which by their own description doesn’t address “gene- or sequence-specific amplification biases.” Add no barcode error correction and no doublet removal, and the authors’ own word, semi-quantitative, is right. Trust the cell list, not the abundance.

The microbial calls rest on one of the most conserved sequences in biology. The Mycoplasma and Gemella results both live at 23S rRNA, and the paper’s own wording is careful: contigs “compatible with bacterial 23S rRNA from several species that are known to occur in the oral mucosa.” Cross-species specificity is tested inside one genus, where the negative control is Mycoplasma agassizii, “a respiratory pathogen of turtles.” That control column isn’t empty, either: the largest mark in the kidney row sits under the turtle pathogen, with smaller hits in skin and brain, and the paper doesn’t mention them. Nothing comparable tests the oral-flora call, where the species is the claim. I’d call both results genus-level, and a day’s work would settle it: count how many of the discriminative contigs’ 24-mers are unique to that species across bacterial RefSeq.

Check the benchmark’s scope. The 70 ms for a single k-mer and 400 MB working set come from Extended Data Fig. 2d, one Visium mouse-brain sample; for the full Index the paper claims “milliseconds.” And “comparable query throughput” to Fulgor is qualified as “from disk,” while the numbers live in Supplementary Table 1: 133 seconds per megabase against Fulgor’s 51, with Fulgor holding 326 GB in RAM to Malva’s 0.3 GB. Malva made the right trade. The main text just doesn’t show you the price.

Log-log scatter of query memory usage in gigabytes against query time in seconds, with two points per method for 1 kbp and 1 Mbp queries. MetaGraph sits near 50 GB, BLAST near 13 GB and slowest, Fulgor and Malva in-memory around 5 to 6 GB, and Malva on disk sits alone near 0.4 GB, roughly an order of magnitude below everything else and correspondingly further right on the time axis.
What Malva optimised (Extended Data Fig. 2d, León-Periñán et al. 2026, Nature, CC BY 4.0). The panel is titled "Query time (Visium mouse brain)": a single sample.

Controlled-access data can never be in it. The clinically interesting data sits in dbGaP, EGA or a hospital, where you need approval to download it, so it can’t be crawled into a public index. Referee #1 raised this in September 2025. The authors’ answer was standalone binaries for building a private index from data you already hold. That keeps your data off their servers, but a local index only searches what you brought, not the corpus.


An Index Is Not a Result

The paper invites the comparison in its second paragraph: “Thirty-five years ago, BLAST pioneered this principle…” So take it seriously.

BLAST didn’t win on optimality. Its 1990 paper sold an approximation on being “an order of magnitude faster than existing sequence comparison tools of comparable sensitivity,” and it kept winning for institutional reasons. NCBI is a standing federal institute, GenBank shipped its 273rd release in August 2026, and the corpus carries a written promise. Under the INSDC policy, shared by GenBank, ENA and DDBJ since 2002, “no use restrictions or licensing requirements will be included in any sequence data records,” and “All database records submitted to the INSD will remain permanently accessible as part of the scientific record.” That permanence is exactly what Malva lacks.

Now the part that cuts against me. AB-BLAST, whose documentation traces its lineage to “the original gapped BLAST (WU-BLAST) in 1996,” says “a Commercial Site license is required for all for-profit entities,” and it had users anyway: its FAQ still fields questions from people holding a “previously licensed copy of WU-BLAST.” So a restrictive licence doesn’t stop adoption, and Malva’s has the same shape: commercial use needs “a separate written license,” by inquiry to the first author. It also excludes “clinical use, diagnostics use,” and names no route for those.

That pushes the weight off the code and onto the corpus. Whichever BLAST you ran, the databases were free to download and search on your own hardware, so NCBI being slow, down or uninterested in your organism was survivable. The algorithm was replaceable. The archive was guaranteed.

Malva has it the other way round. The engineering is better than BLAST’s ever was, and the code is now readable. The institutional picture is thinner: one lab, grant-funded, running the corpus on its own hardware with query limits; a patent application (WO2026/132228, disclosed in the competing-interests statement); an academic-use licence. The 74 million cells are not something you can download.

An index is not a result — it is a subscription.


What Three Referees Bought

Nature published the complete correspondence under CC BY: four referees and five rounds between submission on 17 September 2025 and acceptance on 30 July 2026. The peer review statement names Rob Patro, senior author of Fulgor, the tool Malva is benchmarked against. It’s the best thing attached to this paper.

September 2025. The referees get the API client and nothing else. Referee #1’s first listed concern: “It appears that Malva is intended for commercialization.” Referee #2 notes the core code “seems to be nowhere made available,” asks “Is the intent for that code to remain proprietary?”, and makes the reviewer’s bargain explicit: give me the code, or give me enough detail to reproduce it.

Round one. Standalone binaries, a supplementary chapter of pseudocode, and a cover letter saying the methods were expanded “to make the procedures reproducible, even though the main code remains proprietary.” Referee #4 names the norm: the approach “does not fully align with common practices in the bioinformatics community, where open-source code sharing is strongly encouraged.”

Round two. The response letter: “we have decided to release the full source code under an Academic-use license.”

The process bought more than a licence. The pan-cancer somatic mutation benchmark, the polyASite validation, the sequencing-error robustness analysis, the EGFR lung-cancer cohort and the supplementary chapter on indexing: none of it was in the original submission, and almost all of it arrives in the first rebuttal, the same letter that says the code stays proprietary. The science got better a full round before the licence moved.

So why capitulate? They wanted commercialisation badly enough to file a patent and withhold the core code from referees, then gave up the code and kept the patent. Maybe a Nature paper is worth more than closed source; the authors never say why, anywhere in 22,000 words of correspondence. What you can see is what survived revision. The code moved, and the patent, the licence and the single deployment didn’t. Those were never things a referee could fix, because they aren’t claims. They’re commitments.

And “open source” is doing more work than it can carry. The repository licence grants use “solely for internal academic purposes.” That fails clause 6 of the Open Source Definition, which says a licence “must not restrict anyone from making use of the program in a specific field of endeavor” and gives “genetic research” as its own example. So this is source-available: a real improvement over a binary, and not the same thing. It also means the tool that found bacteria beside a tumour isn’t licensed for diagnostic use.


The Data-Infrastructure Coda

I do data infrastructure for a living and TileDB-SOMA is my day job, so treat this as an interested party’s reading.

Read the Malva Index section of the Methods with the vocabulary swapped and it describes a sparse array store: sorted keys, offsets, fixed-size compressed blocks, memory-mapped range reads, a block cache, and independent write units merged in the background. They reinvented it well, and it’s one of four artefacts they maintain: the k-mer index, a MinHash matrix for the transpose, pre-computed gene matrices for “how much beta-actin is in this cell,” and a relational metadata service, because the index holds integer cell IDs and people ask about tissues. That’s four representations of one corpus, each with its own staleness clock. In the SOMA model the last of those already exists beside the counts, as obs, so a sequence index parked next to it would only need to hand back cell IDs that join. Whether arrays would win on the rest is a benchmark I’d want to run, not a result I can claim.

What I will claim is the version pin. The corpus is crawled continuously, so even the paper’s counts differ: around 74 million cells in the abstract, about 51 million human and 10 million mouse in v2025.04a. The API reports “the current version and statistics,” so you can find out which index you just hit. I can’t find any way to pin a query to a past version, and that’s the difference between an observation and a citation. “I found 4,312 cells expressing X” should still mean 4,312 cells next year.


The speed genuinely impressed me, but what stayed with me is that all of this was already sitting there. The reads were deposited, the barcodes were in them, and the oral bacteria sat outside that tumour section the whole time. They weren’t visible because a reference-based pipeline discards whatever the reference doesn’t recognise, which isn’t a bug in anyone’s code. It’s what the output format is for. Malva didn’t generate a single new base — it declined to discard.

That leaves durability as the question that matters, and it’s the one the paper can’t touch. BLAST is thirty-five years old, and I can still download its databases and search them on my laptop without asking anyone, because an institution outlived every grant that touched it. Somebody has to keep crawling Malva’s corpus too, and a human still curates its metadata. The paper is an excellent argument that it’s worth doing, and it is silent, necessarily, on who pays for it in 2036.

Tags:

← Back to Blog