<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>Notes from Daedalus</title>
  <subtitle>Technical notes from the one-person AI studio of Siegfried Martens, on the systems it ships and how they were measured.</subtitle>
  <link rel="alternate" type="text/html" href="https://s-martens.com/notes/"/>
  <link rel="self" type="application/atom+xml" href="https://s-martens.com/notes/feed.xml"/>
  <id>https://s-martens.com/notes/</id>
  <updated>2026-08-26T00:00:00Z</updated>
  <author><name>Siegfried Martens</name></author>
  <entry>
    <title>Forty-three thousand rooms and no server</title>
    <link rel="alternate" type="text/html" href="https://s-martens.com/notes/forty-three-thousand-rooms/"/>
    <id>https://s-martens.com/notes/forty-three-thousand-rooms/</id>
    <published>2026-08-26T00:00:00Z</published>
    <updated>2026-08-26T00:00:00Z</updated>
    <summary>Orpheus serves a 43,567-room music knowledge graph with no backend, no database and no runtime code. Here is what that buys, what it costs, and the four places the approach actually hurts.</summary>
  </entry>
  <entry>
    <title>Four sources that disagreed</title>
    <link rel="alternate" type="text/html" href="https://s-martens.com/notes/four-sources-that-disagreed/"/>
    <id>https://s-martens.com/notes/four-sources-that-disagreed/</id>
    <published>2026-08-26T00:00:00Z</published>
    <updated>2026-08-26T00:00:00Z</updated>
    <summary>Wikipedia, Wikidata, MusicBrainz and Spotify each have an opinion about who an artist is and what they play. They do not agree. Here is what I did instead of merging them, and the three ways identity broke anyway.</summary>
  </entry>
  <entry>
    <title>Someone else's ruler</title>
    <link rel="alternate" type="text/html" href="https://s-martens.com/notes/someone-elses-ruler/"/>
    <id>https://s-martens.com/notes/someone-elses-ruler/</id>
    <published>2026-08-26T00:00:00Z</published>
    <updated>2026-08-26T00:00:00Z</updated>
    <summary>I graded my own retrieval system on 115 questions I wrote myself. Then I ran it against three public benchmarks written by strangers, with predictions filed before the runs. Two agreed. The third found a real bug.</summary>
  </entry>
  <entry>
    <title>The control did the work</title>
    <link rel="alternate" type="text/html" href="https://s-martens.com/notes/the-control-did-the-work/"/>
    <id>https://s-martens.com/notes/the-control-did-the-work/</id>
    <published>2026-08-26T00:00:00Z</published>
    <updated>2026-08-26T00:00:00Z</updated>
    <summary>I ran three published benchmarks against a retrieval system in one day. Every headline number they produced was misleading until a cheap control corrected it. Four times out of four.</summary>
  </entry>
  <entry>
    <title>The ruler came first</title>
    <link rel="alternate" type="text/html" href="https://s-martens.com/notes/the-ruler-came-first/"/>
    <id>https://s-martens.com/notes/the-ruler-came-first/</id>
    <published>2026-08-14T00:00:00Z</published>
    <updated>2026-08-26T00:00:00Z</updated>
    <summary>How Pythia, an oracle over the Orpheus music knowledge graph, was built evaluation-first: recall@10 0.40 on day two, 0.89 at launch, and six times the measurement itself was wrong.</summary>
  </entry>
</feed>
