Notes · Daedalus

The ruler came first

How I built a retrieval system for a 43,000-room music knowledge graph, and what happened when the measurement turned out to be wrong four times.

14 August 2026 · 3,691 words · Siegfried Martens

The evaluation was built on day two, before any retrieval channel existed: recall@10 40%. Two days after launch, measured by calling the shipped code: 89%, hit@10 100%, MRR 94%, across 115 questions in thirteen shapes.

Four times the harness measured something other than what ships. Every one of the errors flattered the system.

What to ask anyone building you a retrieval system: the baseline, the blind spots, and whether the eval calls the code you ship or a copy of it.

Most AI demos show you a good answer. That’s not hard to produce: ask enough questions, keep the one that worked, screenshot it. This note describes the ruler instead: the number that says how often the right document showed up at all.

This is a retrieval-augmented system, the shape most companies mean when they say “AI over our own data”: an LLM that isn’t allowed to answer from memory, only from documents fetched out of a corpus I provide. This approach moves the hard problem. The model’s prose is rarely the weak part; the retrieval is. If the wrong documents arrive, a fluent answer is built out of them and sounds exactly as confident as a right one. So the number that matters is not the prose.

your question search the corpus this is the weak link the right pages the wrong pages the model the model a right answer a wrong answer just as fluent
Why the retrieval is the part that matters. The LLM model at the end is identical on both paths and does its job either way: it writes fluent prose from whatever it gets handed. Everything that decides whether the answer is true happened before the LLM ever ran.

Pythia is an oracle over Orpheus, a knowledge graph of musical culture built from Wikipedia, Wikidata, MusicBrainz and Spotify. You ask it a question in plain language and it answers in prose, like any LLM, but in this case its every claim is linked back to the room it came from so you can check it. It went from an empty directory to answering questions on the public internet in three weeks. What follows is how I evaluated its performance, including the four times it was wrong.

Two things happened before any of that was possible, and they exposed a bit of the mess that can hide inside a knowledge graph of forty thousand musicians. Ranking the candidate musicians by how widely they are viewed on Wikipedia put Scarlett Johansson, Keanu Reeves and Pope Francis near the top, and every one of them qualifies on a technicality: Scarlett Johansson has two albums, Keanu Reeves spent a decade in Dogstar, and Francis released a prog-rock record in 2015. And because Wikipedia lists band members by name, and the resolver bound each name to the most famous person who had it, Richard Wagner was recorded as a member of The Impalas and Robert Johnson of KC and the Sunshine Band. The tool that found that class looked for band members who had died before their band existed. Wagner died in 1883. The Impalas formed in 1958.

1. The gold set was built first

Day two, before a single retrieval channel was written, I built a gold set: in other words, an evaluation set associating real questions with the rooms that should answer them, hand-grounded to real entities. Then I ran the only channel (i.e., a way of fetching candidate rooms for an answer) that existed, plain semantic search over the embedded corpus (using a vector index of embeddings of the semantics of Wikipedia sections), and wrote down what it scored.

Recall@10: 40%

the question “Where and when did grunge emerge?” pages that should answer it three of them what the system returned, in order two of the three arrived, so recall@10 for this question is 67% top ten results 1 2 3 4 5 6 7 8 9 10 filled squares are pages that should have been found the third one never appeared at all and where a right page lands in the ten is a separate number, which is what MRR measures
The whole metric, on one question. Recall@10 asks only one thing: of the pages that ought to have turned up, how many actually did, anywhere in the top ten. It says nothing about whether the page was right at the top, nothing about whether the prose built on it was true, and nothing about the pages it was never told to expect.

Recall@10 is the share of the rooms that should answer a question which actually show up in the top ten results. 100% means all of them did; 40% means less than half did.

That number is the most useful one in the project, because it is bad. Every claim I make below is legible only because that baseline value of 40% existed first. You can’t tell whether you improved something you never measured, and you can’t discover that your ruler is too short if you never had one.

2. Add channels when misses expose gaps

The gold set covered eight question shapes, and the aggregate 40% hid a structure that turned out to be the entire plan. Semantic search did fine on questions about meaning. It scored zero on the ones that are not:

kind of question example semantic search only
How two things relate “How does ambient relate to minimalism?” 87%
Where and when “Where and when did grunge emerge?” 50%
Who or what belongs here “What are the essential krautrock bands?” 48%
Counting and ranking “What styles have the most artists?” 3%
Music for a situation “Music good for getting focus work done” 5%
Puzzles and constraints “Which artists have exactly three A’s in their name?” 0%
Facts about the names “Bands with ethereal, ghostlike names” 0%

Semantic embeddings can’t answer those four in principle. “What styles have the most artists” is arithmetic. “The opposite of peaceful” requires negation, and a vector index cannot negate: it returned ambient music. “Bands with long, complicated names” is a fact about the strings, and the index returned heavy metal and The Long Ryders.

These results dictated the build order. A bounded compute layer over a precomputed metadata table took counting and puzzles from 3% and 0% to 100%. Inverting the flow for mood questions, letting the model propose styles and then forcing every proposal to resolve against the real corpus, took the situation questions from 5% to 100%, and the negation question started answering noise, grindcore, industrial, powerviolence, harsh noise wall. A lexical channel took the name questions from 0% to 100%, ranking by length times prominence, because pure length surfaces obscure orchestras and pure prominence surfaces the Red Hot Chili Peppers. For bands with long, complicated names it produced The World Is a Beautiful Place & I Am No Longer Afraid to Die, which is a better answer than the one I had written in the answer key. Deterministic graph traversal over the room edges took the relational questions.

The failures on the other side were not subtle either. Asked for the earliest forms of rock and roll, the system answered Philoxenus of Cythera, a Greek dithyrambic poet of the fifth century BC, along with the Minnesänger. It was not malfunctioning. The poets are in the corpus as influence nodes, so asked for the oldest things it knew about, it sorted everything it held by date and answered the question it was asked, which was not the one I’d expected. There is now a guard in the codebase that the project notes refer to as the medieval-poets guard.

Four new retrieval channels in 48 hours took combined recall from 40% to 77%. Once the measurement says which four shapes are broken and why, building is the easy part.

The channels and badges

eight ways to find an answer four things a visitor is told graph scenes communities compute groupby vector expand names “followed the edges between rooms” “counted across every room” “listened for meaning across the rooms” “weighed the names themselves” scope guard “declined”
Eight becomes four, on purpose. Which internal function ran is the builder’s business. What kind of evidence sits under an answer is the reader’s, and it is the only part that helps them decide whether to believe it. These are the exact words the thing says, not a paraphrase for this article.

There are eight ways / channels available to the system to find an answer. A visitor sees a badge that names not the channel but one of four kinds of evidence (the channel itself sits behind a detail toggle), in these words, which are the words it says:

It can also decline out of scope / inappropriate requests.

3. First build the channels, then tune the framework

For context, the system consists of a framework: a router uses the question shape to decide which channel(s) should be invoked for the best answer. The channel returns up to 10 rooms, each with the relevant section marked for the LLM to compose a grounded answer.

In the first phase the goal was capability. Every gain came from something that didn’t exist before. Build a channel, watch a shape go from zero to one. Large effects, easy to see.

The second phase expanded reach. No new channels at all. Every gain came from discovering that a channel I had built was not being used for the questions it was intended. For example:

Nothing was missing; things were unreachable. That has a consequence.

4. The plateau was in the measurement

For eight days the number sat at 80% and would not move.

For five of those days I was not working on the oracle at all. I was rebuilding the corpus underneath it: roughly 36,300 artists to over 40,000, and 39,080 rooms to 43,567. There are no retrieval commits at all in that stretch. So the number was flat because nobody was moving it, and the fact worth keeping is the other one: it held at 80% while the ground under it grew by 11.5%. It is the only stretch where the system was asked to survive a moving substrate instead of improve on a fixed one.

40% 60% 80% 24 Jul 40%, before anything was built four channels, 48 hours 31 Jul to 4 Aug corpus grew 11.5% underneath 5 to 7 Aug the ruler ran out of resolution 89% at launch 12 Aug
Three weeks, and the flat part is the part to read. The climb is the easy story: build a thing, watch the number move. The eight flat days in the middle have two different explanations. For the first five nobody was working on retrieval at all, while the corpus underneath grew by 11.5% and the score held. For the next three the work resumed and the number still would not move, and that is the one that mattered: the measurement had gone too coarse to resolve the changes being made. (The last point is not strictly comparable with the first: the test set grew from 59 questions to 115 along the way, which is discussed in section 6.)

The real plateau was the three days after, when work resumed and the number still did not move. The natural reading is that the work was worthless. The real reading was worse: five different ways of choosing what text to index, differing by 72,559 chunks of content between them, had all scored identically, to every decimal the harness printed.

The number that had stopped moving was not the headline one. One of the thirteen shapes in the gold set is deep_content, and it is scored by section rather than by room: not whether the right room ranked, but whether the paragraph that answers the question did. I realized I didn’t have detailed enough coverage here, with just sixteen sections across seven questions. With a probe at this level of detail, the answer moves in steps of about six percent, and anything smaller is invisible. In the first phase that was fine, because the effects are enormous. In the second phase every remaining effect is smaller than the ruler’s resolution, and a flat line means nothing at all.

Widening that slice to 24 questions and 45 sections immediately adjudicated two changes it had previously scored as noise, and found four live bugs on its own, including two router mis-routes that were returning zero correct rooms to real questions.

A metric that stops moving may be about the metric, rather than about the system.

5. Four times the ruler was wrong

Four times, the test was not testing the thing that ships. A test can be wrong in two ways. It can measure the right system and get the number wrong, which is the failure everyone imagines. Or it can measure something that is not the system at all, and report a perfectly precise number about a thing nobody uses. All four below are the second kind, and the second kind is invisible: the number looks fine, so nothing prompts you to check.

  1. The primary harness could not see production at all. This one turned up on the day the system went live, and it turned up by accident: a ranking fix went in, the gold set was re-run, and the numbers came back identical. The reason was that the harness reimplemented the entire router instead of calling it, so the function that had just changed was never executed. The headline figure everyone quotes had been measured against a copy of the system. The same fault had already appeared in miniature elsewhere: another harness needed a selection rule, copied one, got two details wrong, and reported 253 chunks where production sees 120.

  2. A metric measuring an algorithm that production does not run. The evaluation scored a global top-k across the whole chunk index (all sections). Production retrieves rooms first, then looks inside each one at the sections. Reported: 11 of 16 sections reaching the composer. Actual: 7. A quarter of the metric was an artifact.

  3. An eval scoring an index nobody serves. It measured all 246k chunks while production loaded 173k. Its docstring said it matched production. It did not.

  4. A whole shape the harness could not report on. The per-shape block iterated a hardcoded list of twelve shapes that omitted deep_content: 27 questions, the largest shape in the set, 23% of the gold questions. It counted toward the overall score and was missing from the per-shape table, so its absence read as not yet measured and no amount of re-running would ever have filled it. Like the first, this one surfaced by running the thing rather than by reading it. And when the row finally appeared it read 1.00, which is a ceiling (see the appendix): those questions are answered by a section, and room-level recall structurally cannot see that failure mode.

one question what a visitor gets the real router eight channels the answer what the test scored a copy of the router eight channels a score a fix landed here so the score never saw it, and came back identical
The failure that is hardest to notice. Both paths look like the system. Only the top one is. A test that rebuilds the thing it is testing will report a precise, stable, completely sincere number about software nobody runs, and it will go on doing so until something forces the two paths to be compared.

One more example is worth mentioning as it’s a variation on this theme. A geocoding fix described in the next section was a rule: reject a location that shares no word with the source text the claim came from. It shipped, and it passed the very case it had been written for, because “Vancouver, British Columbia” and “British West Africa” share the word British. The word that caused the bug also cleared the guard against it. Colonial adjectives are now on a stoplist, and the class survives one level in: Cajun music still resolves to Spain, because its source text says New Spain, and unlike British, Spain is a real place name that cannot be banned.

Fixing the first problem exposed a fifth one: the router copy’s top-up logic appended up to ten extra rows, so recall@10 had been scored over as many as twenty candidates. Correcting it moved two shapes down and the headline from 90% to 89%.

Testing should call production code, never a copy of it. A claim about what the system returns is made at the entry point, not at a component underneath it. Where a test can’t do that, it should say so in its own docstring.

Note what all of them, the four above, the fifth hiding inside the first, and the guard, have in common. Every one made the system look better than it was, or made a real change look like no change. Measurement error is not randomly signed. It flatters.

6. The number, and what it still cannot see

Two days after launch, measured by calling the shipped entry point:

Recall@10 89%. Hit@10 100%. MRR 94%. Across 115 questions in thirteen shapes.

In words: of the rooms that should answer each question, 89% arrive in the top ten; every question gets at least one right room in its top ten; and the first right room sits at or near the top almost every time. The per-shape scores are in the appendix. (MRR means ‘mean reciprocal rank’, and measures where in the top-10 rooms the target ones placed)

Three caveats:

“Where did highlife come from?” what the measurement looks at rank 1: the highlife room the expected room arrived, first recall@10 = 1.00 a perfect score, correctly awarded what the visitor reads highlife emerged 1870s, Vancouver, British Columbia, Canada the geocoder matched “British West Africa” to “British Columbia” no retrieval score can see this
Where the ruler stops. Recall measures which rooms arrived, and here the right one arrived first. What it cannot measure is whether the room is telling the truth. Two entirely different kinds of correctness, and only one of them has a number.

Two more challenges I dealt with:

The Orpheus room for Punchmade Dev, a scam-rap figure from Lexington, Kentucky, carried an Opera tag, fifty opera singers as his related artists, and an origin of Heard Island and McDonald Islands, an uninhabited Antarctic territory. Someone had vandalized his MusicBrainz entry. The pipeline had no way to know that, and he arrived in the Opera room’s top artists next to Callas and Pavarotti. Retrieval was perfect: it was asked for opera and it returned the Opera room.

To stop people using a music oracle as a free general-purpose chatbot, it refuses questions containing words like code, essay and translate. Nobody checked those words against 43,567 rooms named after ordinary English things. So it refused “who are the members of Code Orange?”, and, because her surname contains essay, “tell me about Natalie Dessay.” A stranger asking about a hardcore band was told the oracle only answers questions about music. There is not one refusal in the gold set, so nothing was measuring that in either direction.

Most of the real defects were found by asking the system questions and opening the data on any answer that felt off. The evaluation’s job is to keep the fixes fixed, so the next change does not undo three of them.

What I would take from this

If you are evaluating someone to build this kind of system, the demo tells you almost nothing. Ask for the evaluation set, and ask three things about it:

  1. What did it score before you started? If there is no baseline, there is no evidence of improvement, only assertion.
  2. What can it not see? Every measurement has a hole; a good one says where. An evaluation with no stated blind spots has not been examined.
  3. Does it call the code you ship, or a copy of it? Ask even when you are sure of the answer. The harness in this project announced what it was doing in its own docstring, “mirrors answer.route_and_retrieve”, and that read as diligence rather than as a warning. A mirror is a copy, and a copy drifts. Four times in one month, and every copy flattered the system.

One thing the graph did that nobody asked for:

Asked “how do the Beatles connect to the Rolling Stones”, the graph returned a chain nobody put there: The Beatles → John Lennon → The Dirty Mac → Keith Richards → The Rolling Stones. The Dirty Mac existed for one night in 1968, assembled for the Rock and Roll Circus, and it is a real edge between two rooms because Lennon and Richards really were in a band together, once. Nobody wrote that path down. It was there in the data, and the machinery walked it.

Pythia is live at ask.orpheus.rocks. Ask it something, and click through on anything it tells you.

Appendix: recall@10 per gold set question shape

Each row is a kind of question, and every example below is taken from the gold set. n is how many of that kind the set holds. These are the launch-day figures, measured 2026-08-14, and they are a record of that day rather than a live scoreboard: a re-run on 2026-09-16 put the headline at 90%, with the path between two things up from 82% to 100% after a collaboration channel was added to the graph in September.

kind of question example n recall@10
Deep detail about one thing “What have the critics said about how the Pixies were connected to Nirvana?” 27 100% *
What came out of what “What did punk rock evolve into?” 15 66%
Who or what belongs here “What are the essential krautrock bands?” 14 76%
How two things relate “How does ambient relate to minimalism?” 12 92%
Where and when “Where and when did grunge emerge?” 10 94%
Counting and ranking “What styles have the most artists?” 9 91%
Music for a situation “Music good for getting focus work done” 7 96%
The path between two things “What’s the connection between punk rock and dub music?” 6 82%
Facts about an artist “Who are the members of the Rolling Stones?” 5 100%
Facts about the names “Bands with ethereal, ghostlike names” 3 95%
Puzzles and constraints “Which artists have exactly three A’s in their name?” 3 100%
Straight lookup “Who is Stromae?” 2 100%
Category membership “What are the biggest boy bands?” 2 100%

Measured 2026-08-14 by calling the shipped entry point, one run without the metered rescue router. Hit@10 is 100% on every shape: every question retrieves at least one correct room in its top ten, and all the variation below is in how many of the expected rooms arrive.

* The 1.00 on the first row is a ceiling, not a score. Those questions are answered by one section inside a room, and this measurement only checks whether the right room arrived. It cannot see the system picking the wrong paragraph out of the right page. Measured at the section level instead, that row is 46 correct out of 47.

Note that the gold set grew by absorbing failures: every play session against the live system pinned the questions it got wrong, so the set is biased toward what the system gets wrong.

Held out (added 2026-08-26). That bias is now measured against three published sets that nobody here wrote, with the predictions filed before the runs. On 610 questions from ArtistMus, hit@10 was 97%; on 2,085 from Mintaka, 85%. The third, 900 real recommendation requests from Reddit, scored 5% on the first run, half of that the harness’s own fault, and working out why turned up a real defect: asked to recommend artists, the system was answering with a list of genres. The numbers, the predictions and the misses are in Someone else’s ruler.

Next in the series: Someone else's ruler →

Siegfried Martens, 14 August 2026. Part of the Daedalus notes: what was built, what it measured, and what the measurement could not see. Questions and corrections to daedalus@s-martens.com.

If your data has a labyrinth in it, let’s talk.

I take a small number of engagements, the interesting kind. The fastest way to find out whether yours is one of them is a conversation.