Every system this studio ships gets written up the same way: the numbers it was measured against, the decisions that produced them, and an honest account of what the measurement is blind to. These are working notes rather than marketing, and each one is about something that is live and clickable while you read.
202626 August
Field notes from how I built and served a music knowledge graph as flat files, including the bumps in the road.
- 43,567 rooms are stored as 111,728 files and 1.4 GB on disk. Every one is written once by an offline Python script and then never computed again.
- The hosting is cheap. Pain points were more subtle, like Cloudflare’s four hour cache that silently hides deploys, and a content delivery network that complicates checking your own work.
Read the note →
202626 August
Wikipedia, Wikidata, MusicBrainz and Spotify each know who an artist is. They do not agree, and the interesting part is what you build when you stop trying to make them.
- Four distinct data sources about music styles. Instead of merging them, I recorded each separately, and evaluated agreement afterwards, as a separate fusion layer that can be recomputed without re-scraping anything.
- I found a style called Corrido that had 64 artists in it. Thirty-eight were Mexican, none were Spanish, and the entry was for the Spanish dance of that name. Nothing about the data looked wrong.
- I built a visualization to dig deeper into the style/artists graph. Putting artists and geography on the same screen is the only reason that kind of error was visible at all.
Read the note →
202626 August
Four numbers, four wrong conclusions, and the cheap boring checks that caught them. A benchmark tells you something is wrong. It almost never tells you what.
- A benchmark told me my recommendation engine scored 5%. Counting where the questions went, which cost nothing, showed that 46% of them were being scored against a component that returns nothing when run offline.
- A metric said precision fell after a fix. It had not. The metric was only defined on 14 of 75 cases before and 70 of 75 after, so the two numbers were averages over different populations.
- The most useful control of the day was rereading one answer and not believing it: “of course we have a room for Iggy Azalea.”
Read the note →
202626 August
I graded my retrieval system using a validation set I’d written myself. Then I found three sets written by strangers. I filed my predictions before running anything, and watched two of them agree while the third revealed problems.
- On 610 questions from a published benchmark that nobody here wrote, the right room arrived in the top ten 97% of the time. On real recommendation requests from Reddit, 5%.
- The 5% was half a broken measurement and half a real defect. Finding it cost a control worth $0 and fixed a product that was answering “recommend me artists” with a list of genres.
- Pythia said it had no room for Iggy Azalea. It did. A routing step had found both names in the question and then read only the first. Three lines fixed it, and comparison questions went from fetching what they need 35% of the time to 95%.
Read the note →
202614 August
How I built a retrieval system for a 43,000-room music knowledge graph, and what happened when the measurement turned out to be wrong four times.
- The evaluation was built on day two, before any retrieval channel existed: recall@10 40%. Two days after launch, measured by calling the shipped code: 89%, hit@10 100%, MRR 94%, across 115 questions in thirteen shapes.
- Four times the harness measured something other than what ships. Every one of the errors flattered the system.
- What to ask anyone building you a retrieval system: the baseline, the blind spots, and whether the eval calls the code you ship or a copy of it.
Read the note →
Already built, not yet written up
Each of these is a system that is running now. Only the note is missing, so the list is
a table of contents. It carries no dates.
- The agent that audits its own corpus. An agent walking 130 rooms,
inventing its own list of what counts as a defect, for $1.81.
- What strangers actually ask an oracle. Once enough of them have.