Four distinct data sources about music styles. Instead of merging them, I recorded each separately, and evaluated agreement afterwards, as a separate fusion layer that can be recomputed without re-scraping anything.
I found a style called Corrido that had 64 artists in it. Thirty-eight were Mexican, none were Spanish, and the entry was for the Spanish dance of that name. Nothing about the data looked wrong.
I built a visualization to dig deeper into the style/artists graph. Putting artists and geography on the same screen is the only reason that kind of error was visible at all.
Every artist and musical style in Orpheus, a music knowledge graph of 43,000 rooms that you can walk around at orpheus.rocks, is assembled from four places that were never designed to work together.
From Wikipedia we get prose written by people and an infobox (as shown, for the band Pixies). Wikidata has structured claims, with stable identifiers and a complex ontology. MusicBrainz has vast release data and community-sourced tags. On Spotify we found popularity and follower counts.
If you ask each of these four what an artist plays, you’ll get four answers with different vocabularies, different coverage, and different reasons to be trusted. Wikipedia’s infobox is generous and casual. Wikidata’s style claims are sparser and more careful. MusicBrainz tags are whatever a community member typed.
One obvious move would have been to merge them into one answer at the point of ingestion. This note explains the payoff obtained from avoiding premature data fusion.
Don’t merge. Record data provenance.
Every style claim for an artist is stored as its own row, tagged with the source that made it.
From Wikipedia we get two kinds of evidence, from the artist’s infobox and from the style pages they are
listed on, so four sources yield five provenance tags: infobox, pages_appears_on,
wikidata, mbzngs, spotify. A separate pass computes one more kind of row,
consolidated, which is the weighted aggregate the rest of the system reads.
The consolidation normalizes weights in order to prevent overconfident sources from drowning out more careful ones. More specifically, Wikipedia’s infobox asserts far more styles than Wikidata does, and without normalization it simply wins every disagreement by volume.
The weighting was evaluated several times. Each change cost a script run over already-ingested data. Merging at ingest, on the other hand, would have cost a re-scrape of four APIs each time, which would have made me less willing to re-evaluate: an evaluation you can’t easily redo is an evaluation you stop questioning.
Always build a visual
Before Orpheus became a website, I first built Ariadne, a Tableau dashboard. It was meant as a way to show the graph to other people: a clickable, explorable network diagram of styles.
Building it forced an expansion to the underlying data. A style network alone doesn’t tell you much without any context, so the dashboard grew two more panels: “Who is this style?”, a sortable table of artists, and “Where is this style?”, a map of where those artists are from. Neither was possible without joining artist data onto styles, which the pipeline had not been doing. The visualization is what made that join necessary, which then revealed data issues that had been hidden up to that point.
The problem with Corrido
Everything in the corpus is keyed by Wikidata identifier (QID). While this approach
is core to clean data, it rests on an assumption worth saying out loud:
that the identifier we stored is the entity we meant.
Things can get a bit weird when that isn’t the case.
The corpus held a style called Corrido, identifier Q15306187, whose
Wikidata entry describes “a traditional style of dance and music of Spain”. It
had 64 member artists.
The Mexican corrido is Q869210. We had the Spanish one.
This can happen because each source asserts a style by name. The string
"corrido" came from an infobox, got resolved to whichever Wikidata entity
carries that label, and a homonym won. Subsequent passes reinforced it,
because the name still matched.
It wasn’t alone. We found that Banda music pointed at an entry for the music of the Banda people, a Central African ethnic group, and held 48 artists. Oddly, 31 of them were Mexican. Western music pointed at a French article about European art music, and contained Gene Autry, Roy Rogers and the Sons of the Pioneers.
On the surface, the data looked fine. The rows were populated, the counts healthy, the page rendered, and a style with 64 members and a Spanish-language description is entirely unremarkable, so every automated check was passing.
A good dashboard exposes hidden problems. Building the UI caused me to join the artists’ geography and style, revealing anomalies I hadn’t been aware of.
Two more ways identity breaks
That same assumption about QIDs broke twice more, both in ways that leave healthy-looking data.
The QID moves. Wikidata editors merge duplicate items constantly, and the loser becomes a redirect. Nothing tells you. The corpus held the band Savages under two identifiers with 9 and 8 style memberships respectively, one band split in half, and because a merged identifier looks unfamiliar to the ingest gate, the survivor had been re-imported later as a brand new artist. This kind of failure does not just leave stale rows, it manufactures fresh duplicates on every later pass.
The QID is right and the foreign key is wrong. Chasing a gap of 144 rows after rebuilding the discographies, I found 11 MusicBrainz ids each claimed by two different artists. Of the first five I repaired, four had no discography at all, because the collision had quietly handed every one of their releases to their namesake.
None of these is a scraping bug. In every case the fetch worked, the row is populated, and the result is plausible. That is the whole problem.
Takeaways
If you are pulling one picture of the world out of several sources that disagree:
- Never merge at ingest. Store the provenance with the claim. Reconciliation is then a view you can change for the cost of a script run, and re-evaluation, which is important, becomes cheap.
- Normalize for source verbosity. The most generous source wins every disagreement by volume, to the detriment of accuracy.
- Match on identifiers, and then check that the identifier is the thing you meant. Name matching is how the Spanish corrido got 46 Mexican artists, and identifier matching would have prevented it. But identifiers drift, merge, and get attached to the wrong entity, so being identifier-canonical moves the problem rather than ending it.
- Visualize as soon as you can. My early demo instantly started revealing data issues. Counts and null checks cannot see a plausible roster under a wrong heading. A map can.
The uncomfortable summary is that every failure on this page produced data that looked completely healthy, and the thing that caught them was looking at it in a form where the wrongness had somewhere to show.
The graph is at orpheus.rocks, and Ariadne, the dashboard that started it, is on Tableau Public.
Next in the series: Forty-three thousand rooms and no server →
Siegfried Martens, 26 August 2026. Part of the Daedalus notes: what was built, what it measured, and what the measurement could not see. Questions and corrections to daedalus@s-martens.com.