A benchmark told me my recommendation engine scored 5%. Counting where the questions went, which cost nothing, showed that 46% of them were being scored against a component that returns nothing when run offline.
A metric said precision fell after a fix. It had not. The metric was only defined on 14 of 75 cases before and 70 of 75 after, so the two numbers were averages over different populations.
The most useful control of the day was rereading one answer and not believing it: “of course we have a room for Iggy Azalea.”
The carpentry adage tells us to “measure twice, cut once”. Herein we’ll find that reading is even more important than measuring, in the context of system metrics. This note builds on earlier work where I ran three published benchmarks against a retrieval system I built: a music knowledge graph of 43,567 rooms (one page per artist or style) and Pythia, an oracle that answers questions from them. I wrote predictions before each run so I could not reinterpret the results afterwards. That part worked, and I have written it up here.
This note details some bumps along the road. Every headline number those benchmarks produced was initially misleading, and each of these problems was easily diagnosed with controls, i.e., checks that took minutes and usually cost nothing. Not one of the four conclusions I would have drawn from the raw scores was right.
The gotchas in question can inform projects beyond this one, so here are the four, in the order they happened.
1 · Make sure you’re measuring what you think you are
The recommendation benchmark came back at 5% (the share of requests where one of the crowd’s artists made the top ten), against a prediction of 55% to 75%. With a miss that size, my first concern was that the system was somehow broken when it came to artist recommendations.
Before settling for that answer, I started by looking at the question routing. The system has several channels, or ways of answering, chosen per question, and one of them expands a mood into styles using a language model. Running the evaluation offline, without an API key, that component is replaced by a stub with a small cache of fallback answers. It has no entry for arbitrary text off Reddit, so it returns nothing.
416 of the 900 requests, 46%, were being routed to it. They scored zero because no rooms were obtained at all.
Rescored on the requests that reached a working Pythia channel, the number was 9%. So the measurement was about half broken, and the surviving half was still bad enough to be worth chasing, which is where the real defect turned up. Note however that “broken recommending” and “half the test was unplugged” call for completely different approaches.
The check was looking more closely at routing.
2 · Did the denominator change?
I fixed the defect, re-ran, and one number went the wrong way. Precision fell from 7.9% to 3.9%. Halved. This initially looked like a cost of the fix.
Precision here is the share of the system’s suggestions that the crowd also suggested. It is only defined when the system suggests at least one artist, and before the fix it very rarely did: the metric was computed on 14 of 75 requests. After the fix, on 70 of 75.
So the two numbers were averages over different populations, and the earlier one was an average over the cases where the system had already been confident enough to name somebody. Scored on the same 14 requests both times, precision went 7.9% to 7.1%. Nothing had regressed.
This is pretty unexciting on the surface, but at the same time it’s critical to be aware of the population you are referencing with metrics. A metric that is only defined some of the time will move when the frequency of that condition changes, and it can move in the direction that makes your work look worse or better without anything having happened. When things change, check the denominator before believing the ratio.
3 · Multiple choice questions give away the answer
The lookup benchmark is multiple choice: the right answer is always on the page, as one of four options. I split the results by whether the fact being asked about was in my corpus at all. Where it wasn’t, I expected the system to either decline or guess, and guessing between four options lands near 25%.
Mostly it declined. Of the 413 questions whose answer my data does not contain, it passed on 294. But on the 119 it did answer, it picked the right option 81.5% of the time, four times in five. Those answers could not have come from my data. They had to be coming from the language model’s own memory, and they made up roughly half of all its correct answers.
That’s an uncomfortable finding for a system whose whole pitch is that its answers are grounded in documents you can click through and check. I had a paragraph drafted saying grounding was a tendency rather than a guarantee.
The control was to ask the same questions without the four options, the way a person would actually ask. If the model is answering from memory, it should keep doing so. It didn’t. Asked openly, it got the answer 1.7% of the time, two cases in a hundred and twenty. To make sure the test could detect an answer at all, I ran the same check on facts the corpus does hold: 45%.
One more catch before comparing those numbers. The 81.5% only counts the questions it chose to answer, while the 1.7% counts every question: two averages over different populations, the mistake from section 2 again. Over all 413 questions, declines included, the multiple choice score is 23.5%. That’s the fair comparison: 23.5% with the options in front of it, 1.7% without. Handed four candidates, the model can recognize a fact my rooms don’t hold, but asked openly, it doesn’t volunteer one. The leak was in the exam, not the product.
4 · A metric that was misreading the result
The fourth involved a result that was misleadingly good, and the fact that I’d published it internally before verifying it.
On the counting-and-comparing benchmark, the strongest category was comparisons: 80%, well above what I had predicted. I wrote a paragraph explaining why, something about the system being good at facts it stores directly and weak at facts it has to compute.
Then I read one of the answers it had ‘got right’:
Q. Who is younger, Billie Eilish or Olivia Rodrigo? (answer: Olivia Rodrigo) A. I don’t have information about Olivia Rodrigo’s birth date in my rooms. I know [1] that Billie Eilish was born December 18, 2001, but I cannot answer your question without data on Rodrigo.
The scoring worked by checking whether the correct answer appears in the response. “Olivia Rodrigo” does appear in it, but it’s inside the sentence explaining that the system does not have her.
For a comparison, the right answer is one of the two names the question already contains. Any system that simply echoes the question scores correct in this case. Once I excluded the 25 questions with that property, comparisons fell from 80% to 14.3%, from the best category to nearly the worst, and my explanation of why it was strong evaporated along with the high score.
That episode cost me two wrong explanations before I got to a right one, and I would not have started looking at all if I had only been reading the scores and not the answers as well.
What the controls actually were
Nothing here required a better benchmark or a bigger budget. The complete list:
| control | what it cost | what it settled |
|---|---|---|
| count which component each question went to | nothing | 46% of a test was being scored against a stub |
| rescore both runs on the same subset | nothing | the “regression” was a denominator |
| ask the same questions without the options | about $0.30 | the leak was the exam, not the product |
| exclude questions containing their own answer | nothing | the best category was the second worst |
| run the same model with no retrieval at all | about $0.03 | the graph makes it more careful and more accurate |
| answer every request with the ten most famous artists, ignoring the question | nothing | a floor of 1.3%, so the real gain was real |
Four of the six cost nothing. The entire day, benchmarks and controls together, came to $2.19. Measuring doesn’t have to be expensive.
Two things a score could not have told me
Both came from reading answers, and the first is told in full in the companion piece.
My own test set encoded the wrong shape of answer. Asked what to play when you are feeling sad, my hand-written gold answer was sadcore, slowcore, emo. Styles. Real people asking that question on Reddit wanted artists. The test wasn’t inherently wrong, but it was the wrong shape. Things like that just aren’t visible until you look elsewhere. It’s hard to see your own blind spots.
And after a second fix, the benchmark preferred the broken version. This one changed how the system ranks the artists inside a style once it has decided to name them. Asked for an eighties dance party, a better ranking gives Michael Jackson, David Bowie and Depeche Mode; a worse one gives Taylor Swift. The worse one also scores 24% against the crowd where the better one scores 22.7%, because a crowd naming artists names famous ones. The gap is one request in seventy-five, well inside noise, and it would still have pointed the wrong way if I had been optimizing rather than reading.
The best check is always reading the answers
The first draft of the companion piece quoted this, as an example of the system honestly admitting a gap:
A. My rooms contain Nicki Minaj’s birth date [1], but I have no room for Iggy Azalea. I cannot answer this question.
Rereading it, I stopped there. Of course we have a room for Iggy Azalea.
We did. It ranks first for the question “Who is older, Nicki Minaj or Iggy Azalea?”. The cause was a routing step that had correctly identified both artists and then read only the first of them, and the fix was three lines; comparison questions went from retrieving what they need 35% of the time to 95%. The whole story is in the companion piece.
Nothing in that table found it. What found it was reading one answer and realizing it just didn’t make sense. The output transcripts are crucial data, but it takes work to read them. A lot of mistakes can slip through because a table of scores is so much easier to read than four hundred paragraphs of prose.
What I would take from this
If you are building a system like this, or having one built:
- Treat a benchmark result as a candidate, not a verdict. It’s good at telling you that something is wrong in a particular area, but it’s pretty useless at telling you what, and the obvious interpretation was wrong four times out of four here.
- Verify your results carefully before you spend on another benchmark. A second benchmark gives you another candidate. A verification step resolves the one you already have, and the ones above cost either nothing or the price of a coffee.
- Check the denominator before you believe a ratio, especially on any metric that is only defined some of the time.
- Validate the instrument on a case where you know the answer. The 1.7% scenario (answer neither provided nor does it exist in the room text) means nothing without the 45% comparison (answer does exist in the provided room text).
- Read the output, and get somebody else to read it too. Nothing else on this list would have caught the routing bug.
The uncomfortable version of all of this: I had written a piece about evaluation-first development, about the four separate times my own measurement had turned out to be wrong, and about the fact that every one of those errors flattered the system. I then went on to make four more of the same mistake, and this time the errors ran both ways: two made the system look worse than it was, two made it look better. The only constant was that the first reading was wrong, and one of the four survived until a reread caught it. While it’s a hard lesson to learn in the abstract, I do try to keep in mind a habit of distrusting my own headline number for the twenty minutes it takes to check it.
Pythia is live at ask.orpheus.rocks. Ask it something it does not know.
Siegfried Martens, 26 August 2026. Part of the Daedalus notes: what was built, what it measured, and what the measurement could not see. Questions and corrections to daedalus@s-martens.com.