PhD student, Information Sciences Institute, University of Southern California
1 paper at NeurIPS 2025
We propose SRPO, a reflection-aware RL method that significantly improves multimodal LLM reasoning by explicitly teaching self-reflection, outperforming state-of-the-art models on multiple benchmarks.