The evaluation was built on day two, before any retrieval channel existed: recall@10 40%. Two days after launch, measured by calling the shipped code: 89%, hit@10 100%, MRR 94%, across 115 questions in thirteen shapes.
Four times the harness measured something other than what ships. Every one of the errors flattered the system.
What to ask anyone building you a retrieval system: the baseline, the blind spots, and whether the eval calls the code you ship or a copy of it.
Most AI demos show you a good answer. That’s not hard to produce: ask enough questions, keep the one that worked, screenshot it. This note describes the ruler instead: the number that says how often the right document showed up at all.
This is a retrieval-augmented system, the shape most companies mean when they say “AI over our own data”: an LLM that isn’t allowed to answer from memory, only from documents fetched out of a corpus I provide. This approach moves the hard problem. The model’s prose is rarely the weak part; the retrieval is. If the wrong documents arrive, a fluent answer is built out of them and sounds exactly as confident as a right one. So the number that matters is not the prose.
Pythia is an oracle over Orpheus, a knowledge graph of musical culture built from Wikipedia, Wikidata, MusicBrainz and Spotify. You ask it a question in plain language and it answers in prose, like any LLM, but in this case its every claim is linked back to the room it came from so you can check it. It went from an empty directory to answering questions on the public internet in three weeks. What follows is how I evaluated its performance, including the four times it was wrong.
Two things happened before any of that was possible, and they exposed a bit of the mess that can hide inside a knowledge graph of forty thousand musicians. Ranking the candidate musicians by how widely they are viewed on Wikipedia put Scarlett Johansson, Keanu Reeves and Pope Francis near the top, and every one of them qualifies on a technicality: Scarlett Johansson has two albums, Keanu Reeves spent a decade in Dogstar, and Francis released a prog-rock record in 2015. And because Wikipedia lists band members by name, and the resolver bound each name to the most famous person who had it, Richard Wagner was recorded as a member of The Impalas and Robert Johnson of KC and the Sunshine Band. The tool that found that class looked for band members who had died before their band existed. Wagner died in 1883. The Impalas formed in 1958.
1. The gold set was built first
Day two, before a single retrieval channel was written, I built a gold set: in other words, an evaluation set associating real questions with the rooms that should answer them, hand-grounded to real entities. Then I ran the only channel (i.e., a way of fetching candidate rooms for an answer) that existed, plain semantic search over the embedded corpus (using a vector index of embeddings of the semantics of Wikipedia sections), and wrote down what it scored.
Recall@10: 40%
Recall@10 is the share of the rooms that should answer a question which actually show up in the top ten results. 100% means all of them did; 40% means less than half did.
That number is the most useful one in the project, because it is bad. Every claim I make below is legible only because that baseline value of 40% existed first. You can’t tell whether you improved something you never measured, and you can’t discover that your ruler is too short if you never had one.
2. Add channels when misses expose gaps
The gold set covered eight question shapes, and the aggregate 40% hid a structure that turned out to be the entire plan. Semantic search did fine on questions about meaning. It scored zero on the ones that are not:
| kind of question | example | semantic search only |
|---|---|---|
| How two things relate | “How does ambient relate to minimalism?” | 87% |
| Where and when | “Where and when did grunge emerge?” | 50% |
| Who or what belongs here | “What are the essential krautrock bands?” | 48% |
| Counting and ranking | “What styles have the most artists?” | 3% |
| Music for a situation | “Music good for getting focus work done” | 5% |
| Puzzles and constraints | “Which artists have exactly three A’s in their name?” | 0% |
| Facts about the names | “Bands with ethereal, ghostlike names” | 0% |
Semantic embeddings can’t answer those four in principle. “What styles have the most artists” is arithmetic. “The opposite of peaceful” requires negation, and a vector index cannot negate: it returned ambient music. “Bands with long, complicated names” is a fact about the strings, and the index returned heavy metal and The Long Ryders.
These results dictated the build order. A bounded compute layer over a precomputed metadata table took counting and puzzles from 3% and 0% to 100%. Inverting the flow for mood questions, letting the model propose styles and then forcing every proposal to resolve against the real corpus, took the situation questions from 5% to 100%, and the negation question started answering noise, grindcore, industrial, powerviolence, harsh noise wall. A lexical channel took the name questions from 0% to 100%, ranking by length times prominence, because pure length surfaces obscure orchestras and pure prominence surfaces the Red Hot Chili Peppers. For bands with long, complicated names it produced The World Is a Beautiful Place & I Am No Longer Afraid to Die, which is a better answer than the one I had written in the answer key. Deterministic graph traversal over the room edges took the relational questions.
The failures on the other side were not subtle either. Asked for the earliest forms of rock and roll, the system answered Philoxenus of Cythera, a Greek dithyrambic poet of the fifth century BC, along with the Minnesänger. It was not malfunctioning. The poets are in the corpus as influence nodes, so asked for the oldest things it knew about, it sorted everything it held by date and answered the question it was asked, which was not the one I’d expected. There is now a guard in the codebase that the project notes refer to as the medieval-poets guard.
Four new retrieval channels in 48 hours took combined recall from 40% to 77%. Once the measurement says which four shapes are broken and why, building is the easy part.
The channels and badges
There are eight ways / channels available to the system to find an answer. A visitor sees a badge that names not the channel but one of four kinds of evidence (the channel itself sits behind a detail toggle), in these words, which are the words it says:
- “followed the edges between rooms” for anything about how things connect, which covers the graph, the city scenes, and the clusters
- “counted across every room” for anything arithmetic, which covers counting and grouping
- “listened for meaning across the rooms” for the ordinary semantic search everybody pictures when they hear AI, plus the mood questions that get turned inside out first
- “weighed the names themselves” for questions about the text of a name rather than the music behind it
It can also decline out of scope / inappropriate requests.
3. First build the channels, then tune the framework
For context, the system consists of a framework: a router uses the question shape to decide which channel(s) should be invoked for the best answer. The channel returns up to 10 rooms, each with the relevant section marked for the LLM to compose a grounded answer.
In the first phase the goal was capability. Every gain came from something that didn’t exist before. Build a channel, watch a shape go from zero to one. Large effects, easy to see.
The second phase expanded reach. No new channels at all. Every gain came from discovering that a channel I had built was not being used for the questions it was intended. For example:
- The router’s phrase list wasn’t tuned to use the new channels. The graph already held the answers to “what are the essential krautrock bands”; the router just never sent it there. Rewording the question to a phrase already in the list took it from 20% to 80%, which proved the retrieval was fine and the door was locked.
- Returned rooms weren’t getting ranked. “What did hip hop spawn” returned the first ten of 131 edges in storage order, which is why it answered Alaskan hip hop and Bangladeshi hip hop.
- The section-grounding pass was discarding the answering section on every question of one shape, i.e., the wrpmg section was returned.
Nothing was missing; things were unreachable. That has a consequence.
4. The plateau was in the measurement
For eight days the number sat at 80% and would not move.
For five of those days I was not working on the oracle at all. I was rebuilding the corpus underneath it: roughly 36,300 artists to over 40,000, and 39,080 rooms to 43,567. There are no retrieval commits at all in that stretch. So the number was flat because nobody was moving it, and the fact worth keeping is the other one: it held at 80% while the ground under it grew by 11.5%. It is the only stretch where the system was asked to survive a moving substrate instead of improve on a fixed one.
The real plateau was the three days after, when work resumed and the number still did not move. The natural reading is that the work was worthless. The real reading was worse: five different ways of choosing what text to index, differing by 72,559 chunks of content between them, had all scored identically, to every decimal the harness printed.
The number that had stopped moving was not the headline one. One of the thirteen shapes in the gold set
is deep_content, and it is scored by section rather than by room: not
whether the right room ranked, but whether the paragraph that answers the question did.
I realized I didn’t have detailed enough coverage here, with just sixteen sections across seven questions.
With a probe at this level of detail, the answer moves in steps
of about six percent, and anything smaller is invisible. In the first phase that was fine,
because the effects are enormous. In the second phase every remaining effect
is smaller than the ruler’s resolution, and a flat line means nothing at all.
Widening that slice to 24 questions and 45 sections immediately adjudicated two changes it had previously scored as noise, and found four live bugs on its own, including two router mis-routes that were returning zero correct rooms to real questions.
A metric that stops moving may be about the metric, rather than about the system.
5. Four times the ruler was wrong
Four times, the test was not testing the thing that ships. A test can be wrong in two ways. It can measure the right system and get the number wrong, which is the failure everyone imagines. Or it can measure something that is not the system at all, and report a perfectly precise number about a thing nobody uses. All four below are the second kind, and the second kind is invisible: the number looks fine, so nothing prompts you to check.
-
The primary harness could not see production at all. This one turned up on the day the system went live, and it turned up by accident: a ranking fix went in, the gold set was re-run, and the numbers came back identical. The reason was that the harness reimplemented the entire router instead of calling it, so the function that had just changed was never executed. The headline figure everyone quotes had been measured against a copy of the system. The same fault had already appeared in miniature elsewhere: another harness needed a selection rule, copied one, got two details wrong, and reported 253 chunks where production sees 120.
-
A metric measuring an algorithm that production does not run. The evaluation scored a global top-k across the whole chunk index (all sections). Production retrieves rooms first, then looks inside each one at the sections. Reported: 11 of 16 sections reaching the composer. Actual: 7. A quarter of the metric was an artifact.
-
An eval scoring an index nobody serves. It measured all 246k chunks while production loaded 173k. Its docstring said it matched production. It did not.
-
A whole shape the harness could not report on. The per-shape block iterated a hardcoded list of twelve shapes that omitted
deep_content: 27 questions, the largest shape in the set, 23% of the gold questions. It counted toward the overall score and was missing from the per-shape table, so its absence read as not yet measured and no amount of re-running would ever have filled it. Like the first, this one surfaced by running the thing rather than by reading it. And when the row finally appeared it read 1.00, which is a ceiling (see the appendix): those questions are answered by a section, and room-level recall structurally cannot see that failure mode.
One more example is worth mentioning as it’s a variation on this theme. A geocoding fix described in the next section was a rule: reject a location that shares no word with the source text the claim came from. It shipped, and it passed the very case it had been written for, because “Vancouver, British Columbia” and “British West Africa” share the word British. The word that caused the bug also cleared the guard against it. Colonial adjectives are now on a stoplist, and the class survives one level in: Cajun music still resolves to Spain, because its source text says New Spain, and unlike British, Spain is a real place name that cannot be banned.
Fixing the first problem exposed a fifth one: the router copy’s top-up logic appended up to ten extra rows, so recall@10 had been scored over as many as twenty candidates. Correcting it moved two shapes down and the headline from 90% to 89%.
Testing should call production code, never a copy of it. A claim about what the system returns is made at the entry point, not at a component underneath it. Where a test can’t do that, it should say so in its own docstring.
Note what all of them, the four above, the fifth hiding inside the first, and the guard, have in common. Every one made the system look better than it was, or made a real change look like no change. Measurement error is not randomly signed. It flatters.
6. The number, and what it still cannot see
Two days after launch, measured by calling the shipped entry point:
Recall@10 89%. Hit@10 100%. MRR 94%. Across 115 questions in thirteen shapes.
In words: of the rooms that should answer each question, 89% arrive in the top ten; every question gets at least one right room in its top ten; and the first right room sits at or near the top almost every time. The per-shape scores are in the appendix. (MRR means ‘mean reciprocal rank’, and measures where in the top-10 rooms the target ones placed)
Three caveats:
- It is not comparable to the earlier ones. The gold set grew from 59 questions to 115, and it grew by absorbing failures. On one representative day the headline read 87% to 91%, but twenty of those questions were new and scored 90% before any fix landed. The same-set gain that day was 82.0% to 84.9%. A real gain, about a third of what the headline implied.
- It does not cover the metered rescue router. That path costs money per call, so the offline eval runs without it, which means roughly a third of the questions route differently in production than in the harness. It is documented. It is not fixed.
- There is a class of defect it scores 1.00 and cannot see. A geocoding bug once had the oracle report highlife, a Ghanaian genre, as originating in Vancouver, because the geocoder matched “British West Africa” to “British Columbia”. The correct room ranked first. Recall was perfect. The answer was wrong. No retrieval metric will ever catch that, which is why the gold set is not the only instrument.
Two more challenges I dealt with:
The Orpheus room for Punchmade Dev, a scam-rap figure from Lexington, Kentucky, carried an Opera tag, fifty opera singers as his related artists, and an origin of Heard Island and McDonald Islands, an uninhabited Antarctic territory. Someone had vandalized his MusicBrainz entry. The pipeline had no way to know that, and he arrived in the Opera room’s top artists next to Callas and Pavarotti. Retrieval was perfect: it was asked for opera and it returned the Opera room.
To stop people using a music oracle as a free general-purpose chatbot, it refuses questions containing words like code, essay and translate. Nobody checked those words against 43,567 rooms named after ordinary English things. So it refused “who are the members of Code Orange?”, and, because her surname contains essay, “tell me about Natalie Dessay.” A stranger asking about a hardcore band was told the oracle only answers questions about music. There is not one refusal in the gold set, so nothing was measuring that in either direction.
Most of the real defects were found by asking the system questions and opening the data on any answer that felt off. The evaluation’s job is to keep the fixes fixed, so the next change does not undo three of them.
What I would take from this
If you are evaluating someone to build this kind of system, the demo tells you almost nothing. Ask for the evaluation set, and ask three things about it:
- What did it score before you started? If there is no baseline, there is no evidence of improvement, only assertion.
- What can it not see? Every measurement has a hole; a good one says where. An evaluation with no stated blind spots has not been examined.
- Does it call the code you ship, or a copy of it? Ask even when you are sure of the answer. The harness in this project announced what it was doing in its own docstring, “mirrors answer.route_and_retrieve”, and that read as diligence rather than as a warning. A mirror is a copy, and a copy drifts. Four times in one month, and every copy flattered the system.
One thing the graph did that nobody asked for:
Asked “how do the Beatles connect to the Rolling Stones”, the graph returned a chain nobody put there: The Beatles → John Lennon → The Dirty Mac → Keith Richards → The Rolling Stones. The Dirty Mac existed for one night in 1968, assembled for the Rock and Roll Circus, and it is a real edge between two rooms because Lennon and Richards really were in a band together, once. Nobody wrote that path down. It was there in the data, and the machinery walked it.
Pythia is live at ask.orpheus.rocks. Ask it something, and click through on anything it tells you.
Appendix: recall@10 per gold set question shape
Each row is a kind of question, and every example below is taken from the gold set. n is how many of that kind the set holds. These are the launch-day figures, measured 2026-08-14, and they are a record of that day rather than a live scoreboard: a re-run on 2026-09-16 put the headline at 90%, with the path between two things up from 82% to 100% after a collaboration channel was added to the graph in September.
| kind of question | example | n | recall@10 |
|---|---|---|---|
| Deep detail about one thing | “What have the critics said about how the Pixies were connected to Nirvana?” | 27 | 100% * |
| What came out of what | “What did punk rock evolve into?” | 15 | 66% |
| Who or what belongs here | “What are the essential krautrock bands?” | 14 | 76% |
| How two things relate | “How does ambient relate to minimalism?” | 12 | 92% |
| Where and when | “Where and when did grunge emerge?” | 10 | 94% |
| Counting and ranking | “What styles have the most artists?” | 9 | 91% |
| Music for a situation | “Music good for getting focus work done” | 7 | 96% |
| The path between two things | “What’s the connection between punk rock and dub music?” | 6 | 82% |
| Facts about an artist | “Who are the members of the Rolling Stones?” | 5 | 100% |
| Facts about the names | “Bands with ethereal, ghostlike names” | 3 | 95% |
| Puzzles and constraints | “Which artists have exactly three A’s in their name?” | 3 | 100% |
| Straight lookup | “Who is Stromae?” | 2 | 100% |
| Category membership | “What are the biggest boy bands?” | 2 | 100% |
Measured 2026-08-14 by calling the shipped entry point, one run without the metered rescue router. Hit@10 is 100% on every shape: every question retrieves at least one correct room in its top ten, and all the variation below is in how many of the expected rooms arrive.
* The 1.00 on the first row is a ceiling, not a score. Those questions are answered by one section inside a room, and this measurement only checks whether the right room arrived. It cannot see the system picking the wrong paragraph out of the right page. Measured at the section level instead, that row is 46 correct out of 47.
Note that the gold set grew by absorbing failures: every play session against the live system pinned the questions it got wrong, so the set is biased toward what the system gets wrong.
Held out (added 2026-08-26). That bias is now measured against three published sets that nobody here wrote, with the predictions filed before the runs. On 610 questions from ArtistMus, hit@10 was 97%; on 2,085 from Mintaka, 85%. The third, 900 real recommendation requests from Reddit, scored 5% on the first run, half of that the harness’s own fault, and working out why turned up a real defect: asked to recommend artists, the system was answering with a list of genres. The numbers, the predictions and the misses are in Someone else’s ruler.
Next in the series: Someone else's ruler →
Siegfried Martens, 14 August 2026. Part of the Daedalus notes: what was built, what it measured, and what the measurement could not see. Questions and corrections to daedalus@s-martens.com.