Notes · Daedalus

Four sources that disagreed

Wikipedia, Wikidata, MusicBrainz and Spotify each know who an artist is. They do not agree, and the interesting part is what you build when you stop trying to make them.

26 August 2026 · 1,135 words · Siegfried Martens

Four distinct data sources about music styles. Instead of merging them, I recorded each separately, and evaluated agreement afterwards, as a separate fusion layer that can be recomputed without re-scraping anything.

I found a style called Corrido that had 64 artists in it. Thirty-eight were Mexican, none were Spanish, and the entry was for the Spanish dance of that name. Nothing about the data looked wrong.

I built a visualization to dig deeper into the style/artists graph. Putting artists and geography on the same screen is the only reason that kind of error was visible at all.

Every artist and musical style in Orpheus, a music knowledge graph of 43,000 rooms that you can walk around at orpheus.rocks, is assembled from four places that were never designed to work together.

The Wikipedia infobox for the band Pixies. Under Background information it gives Origin: Boston, Massachusetts, U.S.; Genres: alternative rock, indie rock, noise pop, punk rock; Works: discography and songs; Years active: 1986 to 1993 and 2004 to present; eight labels from 4AD to BMG; four spinoff projects including The Breeders; four current members and two past members. The four styles are presented as one flat list, in no particular order, with nothing recording who chose them or how confident anyone was.
A Wikipedia infobox. Note that the styles are provided as a flat list, without ordering or provenance.

From Wikipedia we get prose written by people and an infobox (as shown, for the band Pixies). Wikidata has structured claims, with stable identifiers and a complex ontology. MusicBrainz has vast release data and community-sourced tags. On Spotify we found popularity and follower counts.

If you ask each of these four what an artist plays, you’ll get four answers with different vocabularies, different coverage, and different reasons to be trusted. Wikipedia’s infobox is generous and casual. Wikidata’s style claims are sparser and more careful. MusicBrainz tags are whatever a community member typed.

One obvious move would have been to merge them into one answer at the point of ingestion. This note explains the payoff obtained from avoiding premature data fusion.

Don’t merge. Record data provenance.

Every style claim for an artist is stored as its own row, tagged with the source that made it. From Wikipedia we get two kinds of evidence, from the artist’s infobox and from the style pages they are listed on, so four sources yield five provenance tags: infobox, pages_appears_on, wikidata, mbzngs, spotify. A separate pass computes one more kind of row, consolidated, which is the weighted aggregate the rest of the system reads.

merge at ingest wikipedia infobox wikidata musicbrainz spotify one row who said it: gone to change the rule: re-scrape everything store per source infobox wikidata mbzngs spotify consolidated to change the rule: recompute, from data you already have
Top: merged at ingest, four claims become one row and the provenance is gone, so changing the rule means scraping four APIs again. Bottom: one row per source plus a derived consolidated row, so changing the rule is a script run over data already held.

The consolidation normalizes weights in order to prevent overconfident sources from drowning out more careful ones. More specifically, Wikipedia’s infobox asserts far more styles than Wikidata does, and without normalization it simply wins every disagreement by volume.

The weighting was evaluated several times. Each change cost a script run over already-ingested data. Merging at ingest, on the other hand, would have cost a re-scrape of four APIs each time, which would have made me less willing to re-evaluate: an evaluation you can’t easily redo is an evaluation you stop questioning.

Always build a visual

Before Orpheus became a website, I first built Ariadne, a Tableau dashboard. It was meant as a way to show the graph to other people: a clickable, explorable network diagram of styles.

Building it forced an expansion to the underlying data. A style network alone doesn’t tell you much without any context, so the dashboard grew two more panels: “Who is this style?”, a sortable table of artists, and “Where is this style?”, a map of where those artists are from. Neither was possible without joining artist data onto styles, which the pipeline had not been doing. The visualization is what made that join necessary, which then revealed data issues that had been hidden up to that point.

The problem with Corrido

Everything in the corpus is keyed by Wikidata identifier (QID). While this approach is core to clean data, it rests on an assumption worth saying out loud: that the identifier we stored is the entity we meant.

Things can get a bit weird when that isn’t the case.

The corpus held a style called Corrido, identifier Q15306187, whose Wikidata entry describes “a traditional style of dance and music of Spain”. It had 64 member artists.

the entry says: a traditional style of dance and music of SPAIN where its 64 artists were actually from, August 2026 Mexico by city 46 United States 9 Spain 0 no origin recorded 9
Chalino Sánchez, Valentín Elizalde and Calibre 50, filed under a Spanish folk dance. The 64 artists, as measured in August 2026, by where they were from: 46 Mexican (38 recorded at country level, 8 by city: Mazatlán 5, Guadalajara 3), 9 from the United States, 9 with no origin recorded, none from Spain. The roster is right about the artists; the style it hangs on is a different thing with the same name.

The Mexican corrido is Q869210. We had the Spanish one.

This can happen because each source asserts a style by name. The string "corrido" came from an infobox, got resolved to whichever Wikidata entity carries that label, and a homonym won. Subsequent passes reinforced it, because the name still matched.

It wasn’t alone. We found that Banda music pointed at an entry for the music of the Banda people, a Central African ethnic group, and held 48 artists. Oddly, 31 of them were Mexican. Western music pointed at a French article about European art music, and contained Gene Autry, Roy Rogers and the Sons of the Pioneers.

On the surface, the data looked fine. The rows were populated, the counts healthy, the page rendered, and a style with 64 members and a Spanish-language description is entirely unremarkable, so every automated check was passing.

A good dashboard exposes hidden problems. Building the UI caused me to join the artists’ geography and style, revealing anomalies I hadn’t been aware of.

Two more ways identity breaks

That same assumption about QIDs broke twice more, both in ways that leave healthy-looking data.

The QID moves. Wikidata editors merge duplicate items constantly, and the loser becomes a redirect. Nothing tells you. The corpus held the band Savages under two identifiers with 9 and 8 style memberships respectively, one band split in half, and because a merged identifier looks unfamiliar to the ingest gate, the survivor had been re-imported later as a brand new artist. This kind of failure does not just leave stale rows, it manufactures fresh duplicates on every later pass.

The QID is right and the foreign key is wrong. Chasing a gap of 144 rows after rebuilding the discographies, I found 11 MusicBrainz ids each claimed by two different artists. Of the first five I repaired, four had no discography at all, because the collision had quietly handed every one of their releases to their namesake.

None of these is a scraping bug. In every case the fetch worked, the row is populated, and the result is plausible. That is the whole problem.

Takeaways

If you are pulling one picture of the world out of several sources that disagree:

  1. Never merge at ingest. Store the provenance with the claim. Reconciliation is then a view you can change for the cost of a script run, and re-evaluation, which is important, becomes cheap.
  2. Normalize for source verbosity. The most generous source wins every disagreement by volume, to the detriment of accuracy.
  3. Match on identifiers, and then check that the identifier is the thing you meant. Name matching is how the Spanish corrido got 46 Mexican artists, and identifier matching would have prevented it. But identifiers drift, merge, and get attached to the wrong entity, so being identifier-canonical moves the problem rather than ending it.
  4. Visualize as soon as you can. My early demo instantly started revealing data issues. Counts and null checks cannot see a plausible roster under a wrong heading. A map can.

The uncomfortable summary is that every failure on this page produced data that looked completely healthy, and the thing that caught them was looking at it in a form where the wrongness had somewhere to show.

The graph is at orpheus.rocks, and Ariadne, the dashboard that started it, is on Tableau Public.

Next in the series: Forty-three thousand rooms and no server →

Siegfried Martens, 26 August 2026. Part of the Daedalus notes: what was built, what it measured, and what the measurement could not see. Questions and corrections to daedalus@s-martens.com.

If your data has a labyrinth in it, let’s talk.

I take a small number of engagements, the interesting kind. The fastest way to find out whether yours is one of them is a conversation.