Home
I am a second-year PhD student at INSAIT and UKP Lab supervised by Prof. Iryna Gurevych. My research interests lie in LLM post-training, specifically leveraging Reinforcement Learning methods align and improve language models. Most recently, I published work in TMLR on Reinforcement Learning with Verifiable Rewards (RLVR) for surrogate code verification, uncovering compute-efficient training methods for verifiers through a controlled evaluation testbed.
I graduated with an Integrated Bachelors and Masters of Technology from IIT Kharagpur. My thesis on Improving legal summarization for low-resource jurisdictions was supervised by Prof. Matthias Grabmair and Prof. Saptarshi Ghosh, and was accepted at the NAACL main conference.
During my undergrad, I also took part in internships at the Australian National University and the University of British Columbia.
Outside of work, I’m usually at the gym, at the tennis court, or at a board game cafe. I also have a major soft spot for the street cats of Sofia (though the feeling is rarely mutual).
News
[Aug ‘26] Our preprint “Parameter Exploration for RLVR via Variational Learning“ is out on arXiv!
[Aug ‘26] Our work “Aletheia: What Makes RLVR For Code Verifiers Tick?“ got accepted at TMLR!
[Nov ‘24] Started my PhD at INSAIT ![]()
[Mar ‘24] Our work “Beyond Borders: Investigating Cross-Jurisdiction Transfer in Legal Case Summarization“ got accepted at NAACL 2024!
[Jan ‘24] Our work “Multi-step Automated Generation of Parameter Docstrings in Python: An Exploratory Study“ got accepted at ICSE Posters 2024!
[May ‘23] Started working as a Research Intern at the Australian National University ![]()
[May ‘22] Started working as a Research Intern at the University of British Columbia ![]()