Journal Safari Hires an Intern | Blog Skip to main content
A cute retro-futuristic desktop robot in cream and beige plastic, its boxy CRT-television head showing glowing amber pixel eyes and a smile under a small antenna, holding a coffee carrier in one clamp hand and a short stack of papers in the other, against a plain warm beige background.

Journal Safari Hires an Intern

Published on • 9 min read

Every other Thursday I read one preprint properly, end to end, largely to hedge against AI’s ability to atrophy my critical-thinking muscles. That part I like. The part nobody writes about is Tuesday, when I open bioRxiv and have to find the one worth opening.

In the three weeks I sampled for this post, the four bioRxiv categories I read took in 826 new papers, plus another three hundred-odd revisions of papers already sitting there. Reading every abstract is not possible. Skimming them is what I actually do, badly, on a phone, usually in a parking lot.

So I built a robot intern to do it. Give a model my reading profile, hand it a title and an abstract, ask for a number from 0 to 1 and a sentence of why. The interesting question was never whether that works. It was which model, and how would I know it worked. Sonnet 5 would obviously do fine, and cost more than the coffee I’d drink while it ran. Something twenty times cheaper would probably be fine too. “Probably fine” is the phrase that has cost me the most rework in my career.


The trick: no labels, just a ceiling

The rigorous way to pick a model is to label a few hundred abstracts yourself and measure everyone against your labels. I’ve done that kind of annotation before, and I know two things about it: it eats a full day, and my own labels drift by lunchtime.

Here’s a faster idea. Score every abstract with the best model I’m willing to pay for (more than a coffee a run, less than a sandwich), freeze those scores, and call them the ceiling: the closest thing to truth I can afford. Then grade every cheap model on how closely it tracks the ceiling. Not “is this model right,” which I can’t answer without reading all the abstracts, but “does it make the same calls as the expensive one,” which I can answer for under $3. The ceiling is anthropic/claude-sonnet-5, everything else is a challenger, and it all runs through OpenRouter, which is the only reason comparing twenty-one configurations is an afternoon instead of a procurement exercise. What a ceiling cannot buy is correctness, and I’ll come back to that.


The setup

Ninety bioRxiv preprints posted or revised in the three weeks ending 11 September 2026 (87 of them revisions): sixty from the four categories I read (bioinformatics, genomics, systems biology, microbiology), thirty from seven I don’t, from neuroscience to paleontology. The thirty give one metric teeth. A judge that likes everything isn’t a judge, and the cheapest way to catch one is to hand it papers it should refuse.

Every model sees every abstract three times, with the same prompt and a request for JSON back. Each is graded on mean absolute error (MAE, its average distance from the ceiling’s score), on overlap with the ceiling’s top ten papers, on how often it calls an off-lane paper a fit, and on whether it returns a score at all.


The table

A selection of single models, best agreement first, best challenger per column in bold. The full table is in the repo.

modelMAE vs ceiling [95% CI]top-10$/call
anthropic/claude-sonnet-5 (ceiling)0.00010/100.00394
minimax/minimax-m3@low0.072 [0.054, 0.092]6/100.00027
tencent/hy30.088 [0.068, 0.109]5/100.00006
minimax/minimax-m30.099 [0.076, 0.123]8/100.00021
anthropic/claude-haiku-4.50.136 [0.117, 0.158]7/100.00161
deepseek/deepseek-v4.1-flash0.140 [0.116, 0.165]8/100.00025
deepseek/deepseek-v4-flash-07310.158 [0.135, 0.182]6/100.00008
openai/gpt-6-luna0.207 [0.168, 0.248]7/100.00012
Scatter plot of mean absolute error against the Sonnet 5 ceiling versus measured dollars per call on a log x-axis, one numbered dot per challenger and ensemble with a key below the plot. Blue is reasoning off or provider default, orange reasoning on, green strict structured output, and hollow dots are ensembles. Each dot carries a vertical bar for its 95% bootstrap interval. Dot 1, MiniMax M3 at low reasoning, has the lowest error; dot 3, Hy3, is close behind near the cheap left end of the axis.
Agreement against price, numbered from closest to the ceiling, ensembles included. Hollow dots are median-of-three ensembles; green dots used strict structured output. The x-axis is measured cost, not the price sheet; the bars are 95% bootstrap intervals over the 86 preprints the ceiling scored.

MiniMax M3 at low reasoning effort tracks the ceiling most closely, at MAE 0.072. Tencent Hy3 is second among single models at 0.088 for $0.00006 per abstract, about 4x cheaper. MiniMax keeps its edge in 96% of two thousand bootstrap resamples, but the interval on the gap still includes zero, and it was the best of twenty on these same papers: a small edge, possibly real, against a fourfold price difference. My intern runs Hy3.

Reasoning helped agreement: medium effort moved DeepSeek V4 Flash from MAE 0.158 to 0.092, at 13.9 seconds a call instead of 4.3. And newer isn’t better: three cheap models released after I set the roster all landed in the bottom third on MAE, two of them further from Sonnet than their predecessors. GPT-6 Luna also called off-lane papers a fit nine times in ninety calls; Hy3, never.


The family discount does not apply

The model I expected to win was Haiku 4.5. Same lab as the ceiling, same house style. If any cheap model were going to think like Sonnet 5, surely it would be the one from the same building.

It came eleventh of twenty on MAE, at 0.136 against Hy3’s 0.088, while costing about 26 times more per call. Yet it ranks papers about as well as Hy3 and agrees with Sonnet more often on which side of 0.5 a paper falls. It just scores papers about 0.13 higher on average, and MAE punishes every bit of that.

The closest match did come from outside the family, though Anthropic has accused MiniMax of mining more than 13 million Claude exchanges to distill it. That campaign predates Sonnet 5 and targeted coding and tool use, so I can’t say it explains things here, but you never know!


The failures were mine

When a model looks broken, suspect your own code first. Most of my early “model failures” were mine:

  • Haiku wrapped all 270 answers in a markdown code fence, even in JSON mode, so my bare json.loads() gave it zero coverage. A parser that reads past the fence scores all 270.
  • One provider put MiniMax’s answers in message.reasoning, a field my unpack() never read, and left message.content empty. Reading both lifted MiniMax M3 with reasoning off from 79.6% coverage to 97.0%.
  • I set a 99% coverage cutoff by eye, and it threw out MiniMax M3 at low effort, whose 10 misses in 270 calls were scattered across ten different papers. The rule now is a score for every paper within three tries, which I also chose after seeing the numbers. Under it, MiniMax M3 at low effort goes from disqualified to the top of the table, and my intern meets it by retrying.

“My code failed” and “the model failed” are different claims, and only the first is cheap to check: read the raw responses.


What the ceiling can’t tell you

The only model that fails the three-tries rule is the ceiling. Sonnet 5 left four abstracts with no score on any attempt: twelve calls came back empty with finish_reason: content_filter, three tries each on influenza antigenic evolution, Cedar virus mRNA editing, tick-borne encephalitis reporter stability, and a Cryptococcus pathogenicity gene. Re-running the Cedar virus call showed that label is OpenRouter’s name for a refusal by Anthropic’s safety classifiers.

The ceiling’s scores are also squashed. It put 12 of its 165 scored in-lane calls at 0.5 or above, which matches my own read: bioRxiv’s microbiology is full of excellent wet-lab virology my profile skips. But it makes MAE too easy. Half the papers the ceiling scored sit at 0.05 or below, and a judge that answers 0.10 to every paper without reading any of them scores MAE 0.098, better than 16 of the 20 challengers. That’s why the top-10 column exists, and it disagrees with MAE. DeepSeek V4.1 Flash matched 8 of the ceiling’s top 10, as many as any single model, while ranking twelfth of twenty on MAE. Hy3 matched 5.

The caveat, stated plainly: I am not calibrating these models against being right. I’m calibrating them against Sonnet 5. The metrics reward a challenger for sharing the ceiling’s blind spots and count it as an error wherever it escapes them. And MAE only compares papers the ceiling managed to score, so those four filtered abstracts sit outside the measurement entirely. A challenger can say anything it likes about the Cedar virus paper and its MAE doesn’t move.

I have made this exact criticism of somebody else’s work. Reading Pan-human Azimuth in August, what bothered me was that every cell’s training label was machine-transferred from older references, so the model mostly compresses what they already believed. They did at least map each tissue reference’s cell type names onto one tree by hand. My judge has no human labels at all, and if it found a paper Sonnet dismissed, my grading would score that as a miss. The stakes differ: their labels go into an atlas other people build on, and mine decide what I read on a Thursday. But I don’t get to publish that critique in August and pretend it doesn’t apply to my own pipeline in October.


What it costs

The whole bake-off is 5,670 calls across twenty-one configurations and cost $2.90. The ceiling was $1.06 of that, over a third of the budget spent on the yardstick rather than the candidates.

Two weeks is about 950 new papers in the categories I read across bioRxiv, medRxiv and arXiv, and one Hy3 pass costs $0.06 to $0.08. Hy3’s scores are coarse (74 of 948 tied at 0.85 in one real two-week batch), so DeepSeek V4.1 re-scores the ties for $0.01 to $0.02 more. A year of it comes to about $2. At the bake-off’s per-call prices, the same year would be about $7 with MiniMax and about $97 with Sonnet.

The code is on GitHub, MIT, and works against your own reading profile if you swap out profile.md and the category lists.

My intern wasn’t running yet when I picked the last Journal Safari paper, so here is a replay. On these ninety, the ceiling’s top pick was a preprint on calibrated variant effect prediction at the residue level, and Hy3’s was the same paper: the one I’d have pulled out myself. Third on the ceiling’s list sat a long-read haplotype study I would have missed entirely. So would my intern, which ranked it 25th. MiniMax M3 at low effort ranked it joint third, for four times the price. That paper is what the cheaper intern costs me, and at least now I know it.

Tags:

← Back to Blog