On 610 questions from a published benchmark that nobody here wrote, the right room arrived in the top ten 97% of the time. On real recommendation requests from Reddit, 5%.
The 5% was half a broken measurement and half a real defect. Finding it cost a control worth $0 and fixed a product that was answering “recommend me artists” with a list of genres.
Pythia said it had no room for Iggy Azalea. It did. A routing step had found both names in the question and then read only the first. Three lines fixed it, and comparison questions went from fetching what they need 35% of the time to 95%.
A few weeks ago I wrote up the evaluation, for Pythia, a retrieval system I’ve built over a music knowledge graph: 43,567 rooms, one per artist or style. Pythia is an ‘oracle’ that answers questions by fetching appropriate rooms and writing prose from them. The headline from the evaluation was recall@10 of 89% across 115 questions in thirteen shapes: in other words, of the rooms that should answer each question, 89% arrived in the top ten.
As I was writing the evaluation, I became gradually more concerned about my lack of comparisons against external standards.
I wrote the 115 questions. I also built the system. So 89% isn’t a score, it’s really a description of my own expectations.
That’s worse than it sounds. My evaluation ‘gold’ set grew by absorbing failures: whenever something got missed, the something became a new test case. This aggregation approach is a good way to avoid regressions, but it doesn’t help you become aware of your own blind spots. If you build your own test set, it converges on the shape of your own imagination.
One fix would have been a held-out set from questions written by others. I initially skipped that in order to get the system off the ground quickly; I had missed that licensed data sets of such questions had already been published.
1 · Three sets, with different intents
I found three relevant data sets, chosen for their different coverage directions.
| Set | What it is | What it tests |
|---|---|---|
| ArtistMus | 1,000 multiple choice questions about 500 artists, sampled to be globally balanced (163 countries, 43% outside the US and Europe) | looking one fact up |
| Mintaka | 20,000 crowdsourced questions typed by complexity, tagged with Wikidata ids; 2,500 are music | counting, comparing, ordering |
| Reddit2Deezer | 186,380 real recommendations from music subreddits, each one resolved to an artist | recommending, graded by the crowd that answered |
The third is the most challenging, but also the most important, for structural reasons. In the first two, the room the system needs for the answer is named in the question: “Where was Thad Jones born?” contains Thad Jones. With the Reddit2Deezer data, the answer isn’t part of the question: they describe a feeling and thirteen strangers reply with artists. That’s a more interesting task.
Matching my system up against these external references wasn’t a problem. ArtistMus matches on artist names: 305 of their 500 artists have a room, yielding 610 usable questions. Mintaka joins on Wikidata ids, which I already use as keys, so all 2,085 questions were keepers. For Reddit2Deezer, I found 23,164 threads where three or more of the recommended artists had rooms in my graph.
2 · Write your predictions down first
Before running an evaluation against these new references, I wrote down what I expected, along with the thresholds that would count as bad news. This was cheap: one paragraph, written first, keeping me from tailoring an explanation to my findings after the fact. My expectations were:
- ArtistMus: hit@10 between 90% and 97%. (Hit@10 is recall@10’s looser cousin: did at least one right room arrive in the top ten.) Every question names its artist, so this is a lookup and it should be easy. Below 85% would mean the name channel has a real gap. The system has several ways of fetching rooms, chosen per question; I call them channels, and the one that matches on names is one of the more simple ones.
- Mintaka: the worst types will be intersection (X and Y) and superlative (most/best). If either falls under 50%, graph traversal is not being reached for questions about membership, which would be a question routing bug.
- Reddit2Deezer: hit@10 between 55% and 75% on “artists like X” requests. The graph already connects related artists, which should reach some of what the crowdsourced answers said.
3 · Two were right
ArtistMus came in at 97%, at the top of the predicted band. Contextual questions scored 97% and factual 98%, a gap of 1% against a predicted 5%. Nothing to report, prediction confirmed.
Mintaka split in a way that made sense. When the question names the subject (“How old was Taylor Swift when she won her first Grammy?”), hit@10 was 94%. Where the answer is itself a room, which the question may not name at all, 67%.
One type fell below the predefined threshold: superlative, 43%, under the 50% line. “Who is the youngest person in Tribe Called Quest?” names the group and wants a member, so nothing in the question points at the room that has to arrive. That is the routing bug the prediction described, called before the number existed.
The prediction also expected intersection as another weak type. Intersection scored 82%, however, so while the threshold was right, the reasoning was wrong: what breaks retrieval isn’t questions about membership, it’s questions that don’t mention the room you need. An intersection names both constraining entities. A superlative names none.
I would not have found that distinction by grading myself, because I would not have written enough superlatives to notice.
4 · The third one found a problem
Reddit2Deezer scored hit@10 of 5% against a predicted 55% to 75%.
Before panicking about some undiscovered system flaw, a miss that big deserves a second look. I decided to look more carefully at question routing. Instrumenting the routing for more insight took a few seconds and cost nothing. It turned out that 416 of the 900 requests, 46%, went to a channel that returns nothing when the system runs without an API key, which is how the offline evaluation runs. They scored zero by construction. Rescored on the requests that reached a working channel, the number was 9%. Much better, but still not great. Run live, with a key, on a sample of 75 requests, it was 13.3%, and that is the starting point for everything below.
So the measurement was half broken. The other half was real, and it was worse than a low score, because reading what the system actually said explained it immediately, and did reveal a real problem. Asked this question:
I’m building a playlist of unique, gritty, sleazy, and funky Americana tracks that feel older or have that vintage vibe.
Pythia answered with seven genres: Outlaw country, roots rock, southern rock, Americana, swamp rock, country soul, Texas blues. A perfectly good gesture in the correct direction. Unfortunately the actual repliers named Sean Hayes, All Them Witches, Radio Moscow, Humble Pie, Creedence Clearwater Revival, Tom Waits and seven others.
Across 25 similar requests, the system cited 170 style rooms and zero artist rooms. It was asked to recommend artists, but instead it kept describing styles.
It’s not that we didn’t have the answers either. Style rooms often carry ranked lists of their artists, and three of the thirteen crowdsourced artists were available from the seven rooms it had just cited: Creedence Clearwater Revival in the top tier of swamp rock, ZZ Top in the second tier of southern rock, Tom Waits in the third tier of Americana. The channel had stopped one hop short: instead of citing the rooms it should have read them.
This small fix, of just reading data already loaded, took hit@10 from 13.3% to 24% on real requests without touching the rest of the evaluation. As an aside, an interesting baseline reference: just answering every request with the ten most followed artists in the corpus and ignoring the question entirely scores 1.3%.
This is an example of an area where my gold set did have a blind spot. For situation questions it expects styles: ask it for music for when you are feeling sad and among the answers I’d been expecting had been sadcore, slowcore and emo. The system scored 96% against that. It was answering the question I had written, correctly, and the question was the wrong shape. Strangers asking the same thing wanted artists, which makes a lot of sense, but just didn’t occur when building the golden set. This is why finding extra reference sets is so important.
5 · A number that isn’t about Pythia at all
Three of these results are impacted by coverage, which is a different system’s problem: the scraper and the graph contents, not the retrieval of the contents.
- 305 of 500 artists on a deliberately globally balanced list have a room. Just 61%.
- Only 76% of the artists Reddit actually recommends have a room. Worse, only 27% of the unique names do, because the long tail is what is missing.
- On ArtistMus the fact needed to answer was found in the room text only 32% of the time.
Michala Petri, Catupecu Machu, Connan Mockasin, Xenia Ghali and 191 others are not in the graph. No amount of retrieval tuning fixes that. This is partially good news: these three question sets handed me a shopping list, written by strangers, for the next six months of scraping.
6 · What Pythia does when it has no idea
The multiple choice set (ArtistMus) turned out to measure something I hadn’t thought of.
Given four options and told to answer with a letter, Pythia declined 59% of the time, in terms that name a reason:
I have no rooms about Thad Jones or his move to Copenhagen in 1979. My rooms contain nothing about him, his decision-making, his work with any big band, or his career trajectory. I cannot answer this question.
Of the 41% it did answer, 81.6% were right. Then I ran the same questions through the same model, but with no retrieval at all, as a control. In other words, I gave the model no access to any room text, forcing it to answer from general knowledge: it answered more often, 58% of the time, and was right 60.9% of the time it did. Use of the graph makes the same model more grounded, careful and accurate, but on a benchmark that rewards guessing, refusing to do so can cost you.
This raised a concern. About half of those correct answers came on questions whose fact isn’t in the room’s text at all, so they couldn’t have been sourced from the rooms. They had to come from the model’s own general knowledge. Checking on this required a control of its own: the same questions without the four options, the way a person asks. On facts the rooms don’t hold, Pythia then produced the right answer just twice in a hundred and twenty, 1.7%, against 23.5% with the options on the page. Handed four candidates, the model recognizes the answer; asked openly, it doesn’t volunteer it. The story of this test, and the control that makes it mean something, is in Read twice, measure once.
Minimizing hallucinations is a property I would demand from a data retrieval system: it would rather say nothing than make something up. This isn’t a guarantee, 1.7% is not zero, and I would not print the word guaranteed. But it is a measurement rather than a promise, and the test that produced it takes an afternoon.
7 · Where Pythia still has challenges
Looking deeper, I found Pythia’s answers for Mintaka, as opposed to its retrieval, to be weak. The rooms hold the answer about three quarters of the time, but Pythia reports it correctly under a third of the time. Comparative questions were the worst of all, and looking at examples makes the problem clear:
Q. Who is older, Nicki Minaj or Iggy Azalea? A. My rooms contain Nicki Minaj’s birth date [1], but I have no room for Iggy Azalea. I cannot answer this question.
I had started to write here about how a comparison needs both sides, and that if each artist independently carries a birth date seven times out of ten, the comparison can only work about half the time. That makes sense, but it’s also just wrong in this case: on re-reading the question, I realized that was nonsense, and said to myself we absolutely have a room for Iggy Azalea.
We do, and it states that she was born on June 7, 1990.
My fallback explanation was that the system had the answer and refused anyway. I’d checked, of course: run that question through the search and Iggy Azalea’s room comes back at rank one, its first line reading “Amethyst Amelia Kelly (born 7 June 1990), known professionally as Iggy Azalea.” Nicki Minaj’s room comes back at rank two, also with her date. Both dates are available to the system, and the answer is 1982 against 1990.
That explanation was also wrong, and the reason why returns to a theme we’ve seen before: I had run the search the way the evaluation runs it, but the evaluation does not run it the way the product does.
Without an API key, the system routes with fixed rules and this question goes to plain name-based search, which returns both artists. With a key, an extra step routes it instead, and that step routed comparisons to the part of the system that walks the graph. That part resolves the first artist named, follows her stylistic edges, and never looks up the second one at all. So in production, this question fetched:
Nicki Minaj, hip-hop, contemporary R&B, pop music, T-Pain
Iggy Azalea is genuinely not there. Across twenty comparison questions, the version I was measuring was fetching everything the question names 95% of the time, while the deployed version was actually managing 35%.
So Pythia was truthful about what it had been handed, but false about what it holds. Those are two different things, and only one of them was said. “I have no room for Iggy Azalea” is a claim about the corpus and it is wrong; “Iggy Azalea is not among the rooms I have here” would have been right, and would have sent the reader to look rather than to conclude the graph was thin. Getting the response language right is important and it shows here.
The uncomfortable part is the other half. The 95% number I had recorded for this shape was measured on the non-production path. The harness even prints a warning that live routing differs, which I had read for weeks as a note about coverage. It is not. It is yet another instance of a mistaken measurement flattering the system, a recurring pattern I’m striving to avoid.
I checked on the ramifications of this evaluation test vs. production code concern. Routed the way production routes, ArtistMus loses 1.7% and Mintaka 6.7% across sixty sampled questions each. So the top-line numbers survive, slightly reduced, and comparisons are the one shape that falls off a cliff. Everything in section 6, the refusal rates and the invention rate, was measured live to begin with and is unaffected.
What it cost to fix
This took three lines to fix. The bug had been in what one step had been handing the next.
The initial question routing step was doing its job perfectly. Given “Who is older, Nicki Minaj or Iggy Azalea?”, it reports back that this is a lookup, that the first entity is Nicki Minaj, and that the second entity is Iggy Azalea. It had found both all along. The next step, the one that turns that report into a fetch, had a bug: it just read the first slot and ignored the second.
In other words, the fix consisted of declining and letting the question fall through to plain search, which correctly fetches both. A guard for this already existed a few lines above for a different two-entity case, and it simply hadn’t been applied to the channels that drop the extra name silently rather than loudly.
With this change, comparison questions now retrieve every entity they name 95% of the time, up from 35%, exactly what the offline path had been getting all along. Overall recall across the whole test set was unchanged, moving from 82.39% to 82.42%. Finally, the question that started this now answers:
Q. Who is older, Nicki Minaj or Iggy Azalea? A. Nicki Minaj is older. She was born December 8, 1982 [2], while Iggy Azalea was born June 7, 1990 [1]. That makes Nicki Minaj nearly eight years her senior.
Two more small things went wrong in the course of this fix. My first attempt detected the second entity using a regular expression looking for capitalized words, and it resolved the “Who” at the start of every question to the band The Who. It was also unnecessary, because the second name was already sitting in the routing report I had not read closely enough. Second, one question shape appeared to get worse by 18.5%, which looked like a real cost until I ran the same measurement twice with the fix turned off, and realized the result was variable: I’d got 39.3% and 22%. The routing step calls a language model, so it is not deterministic, and a three question shape cannot detect anything smaller than the noise. That is the fifth number in this piece that needed a control before it meant anything.
The wording changed as a result of this as well. A refusal now has to be about the rooms in front of the model rather than about the whole graph, because those are different claims and only one of them was ever true.
Superlatives are also a problem, and for a cleaner reason:
Q. Who has won the most Grammy Awards of all time? A. My rooms do not contain information about who has won the most Grammy Awards of all time. They mention specific achievements, Beyoncé winning Album of the Year at the 67th Annual Grammy Awards [3], Adele winning six awards in a single night [5], but they do not identify an overall record holder across Grammy history.
Nothing is wrong with the corpus there. The question wants an operation over all of it, and that operation just hasn’t been added yet.
What I would take from this
The next time I need to build a conversational front-end layered on top of a large confusing mess of heterogeneous data, here are several things I will keep in mind. Note that most of this is not model work, and recall also that we are also attempting to minimize hallucinations by grounding answers in named sources.
- Grade the system on third party questions. A lot of new information about the model emerged from evaluating the three sets, and the two most useful findings came from the one that scored worst. Published benchmarks with a license exist for a lot of domains, and grounding these against my corpus took a name match and a join on ids.
- Write down your predictions first, including the ‘bad news’ thresholds. The superlative bug was called before the number existed, which is the difference between a finding and a rationalization. It also removes the temptation to quietly reframe a ‘bad’ result.
- Measure system refusals too, and check what it says when it does. A system that answers everything isn’t better than one that answers 41% of the time and is right 82% of those; it’s differently wrong, and only one of them tells you when to go and check. But it should say “that isn’t in what I fetched” rather than “I don’t have that”. The first is checkable and usually true. The second is a claim about your whole corpus, and mine made it wrongly about a room it did have.
- Read the transcripts, and better yet get somebody else to do so also. The best findings on this page came from careful reading of the answers, not from reading metrics. A benchmark tells you something is wrong, but it rarely tells you what. The most useful part of this whole exercise was just me saying to myself, rereading one quoted answer: we absolutely have a room for Iggy Azalea. That ended up killing two of my own explanations, and a second reader would have killed them sooner.
Every number here was misleading until a control corrected it, and that pattern turned out to be its own note: Read twice, measure once.
Pythia is live at ask.orpheus.rocks. Ask it something it does not know, and see what it says.
Appendix: the numbers
Measured 2026-08-26 against the shipped code. Retrieval runs cost nothing; the answer runs cost $2.19 in total. The retrieval rows use the evaluation harness’s routing, not the live product’s (section 7): routed the way production routes, ArtistMus loses 1.7 points and Mintaka 6.7, on sixty sampled questions each.
| set | n | what is measured | result |
|---|---|---|---|
| ArtistMus | 610 | right room in the top ten | 97% |
| Mintaka, subject named | 1,368 | right room in the top ten | 94% |
| Mintaka, answer is a room | 717 | at least one needed room in the top ten | 67% |
| Mintaka, superlatives | 164 | as above | 43% |
| Reddit2Deezer | 75 | a crowd’s artist in the top ten | 24%* |
| Reddit2Deezer, ignoring the question | 75 | as above | 1.3% |
| ArtistMus answers | 610 | declined | 59% |
| ArtistMus answers | 250 | correct, of those answered | 81.6% |
| same model, no retrieval | 40 | correct, of those answered | 60.9% |
| ArtistMus, open questions | 120 | states a fact its rooms lack | 1.7% |
* After the fix in section 4. A ranking corrected the same day projects 22.7% and has not been rerun.
Next in the series: Read twice, measure once →
Siegfried Martens, 26 August 2026. Part of the Daedalus notes: what was built, what it measured, and what the measurement could not see. Questions and corrections to daedalus@s-martens.com.