-
Playing with Fire: What Transfers When RL Trains a Language Agent?
Mahesh Ramesh*, Kaousheik Jayakumar*, Hemanth Ram, Pavan Thodima, Ramani Duraiswami, Dinesh Manocha, Aniket Rege, Emmanouil-Vasileios Vlatakis-Gkaragkounis (* equal contribution)
ICML 2026 Workshop · RLxF: Reinforcement Learning from World Feedback
Studies what transfers when models are trained with RL and which factors control that transfer. The paper shows that RL generalization depends on pre-training knowledge, but RL without strong prior knowledge is still useful: hard-task RL-trained models perform better than base models when enough information is provided in context, improve on domains where the model is already competent, and provide a better initialization for further staged RL training.
-
Sparks of Cooperative Reasoning: LLMs as Strategic Hanabi Agents
Mahesh Ramesh, Kaousheik Jayakumar, Aswinkumar Ramkumar, Pavan Thodima, Aniket Rege, Emmanouil-Vasileios Vlatakis-Gkaragkounis
ICML 2026
Introduced a multi-turn benchmark to evaluate state-tracking and cooperation in frontier models. RL-trained a 4B model on curated data, outperforming all non-reasoning baselines and performing comparably to reasoning models such as o4-mini. Released 1500+ (~90K data points) game trajectories for SFT and move-level ratings for RLVR.
-
CuRe: Cultural Gaps in the Long Tail of Text-to-Image Models
Aniket Rege, Zinnia Nie, Mahesh Ramesh, Unmesh Raskar, Zhuoran Yu, Aditya Kusupati, Yong Jae Lee, Ramya Korlakai Vinayak
ICCV 2025 · CVPR DemoDiv Workshop (Oral)
Released a new dataset of 300 cultural artefacts spanning 6 cultural axes and 64 countries to measure long-tail bias in T2I models. Conducted 2,700 artifact-level surveys with culturally aligned annotators. Introduced Marginal Information Attribution, achieving a 2× improvement in Spearman correlation with human preferences over prior baselines.
-
MABViT: Modified Attention Block Enhances Vision Transformers
Mahesh Ramesh, Aswinkumar Ramkumar
Deployable AI Workshop, AAAI 2024 (Oral)
Integrated Gated Linear Units into the attention module for parallel MLP-attention computation.