Home

I am a second-year PhD student at INSAIT and UKP Lab supervised by Prof. Iryna Gurevych. My research interests lie in LLM post-training, specifically leveraging Reinforcement Learning methods align and improve language models. Most recently, I published work in TMLR on Reinforcement Learning with Verifiable Rewards (RLVR) for surrogate code verification, uncovering compute-efficient training methods for verifiers through a controlled evaluation testbed.

I graduated with an Integrated Bachelors and Masters of Technology from IIT Kharagpur. My thesis on Improving legal summarization for low-resource jurisdictions was supervised by Prof. Matthias Grabmair and Prof. Saptarshi Ghosh, and was accepted at the NAACL main conference.

During my undergrad, I also took part in internships at the Australian National University and the University of British Columbia.

Outside of work, I’m usually at the gym, at the tennis court, or at a board game cafe. I also have a major soft spot for the street cats of Sofia (though the feeling is rarely mutual).

News

[Aug ‘26] Our preprint Parameter Exploration for RLVR via Variational Learning is out on arXiv!

[Aug ‘26] Our work Aletheia: What Makes RLVR For Code Verifiers Tick? got accepted at TMLR!

[Nov ‘24] Started my PhD at INSAIT :bulgaria:

[Mar ‘24] Our work Beyond Borders: Investigating Cross-Jurisdiction Transfer in Legal Case Summarization got accepted at NAACL 2024!

[Jan ‘24] Our work Multi-step Automated Generation of Parameter Docstrings in Python: An Exploratory Study got accepted at ICSE Posters 2024!

[May ‘23] Started working as a Research Intern at the Australian National University :australia:

[May ‘22] Started working as a Research Intern at the University of British Columbia :canada: