MS student, Mohamed bin Zayed University of Artificial Intelligence
1 paper at NeurIPS 2025
ViMaR is a two-stage, value-guided inference framework that uses margin-based rewards to produce faster, more accurate, and less hallucinatory captions, enabling scalable and self-improving vision–language models.