Research

I build and evaluate AI systems for radiology, and study where they fail once they leave a benchmark and meet real clinical data. Most of the work below sits at the boundary between a working model and a system a clinician can actually trust: labelled data is scarce, deployment conditions shift, and language-model outputs need governance before they reach a report. I do this as Founder and Group Lead of CRASH Lab at Ashoka University.

Fine-grained diagnosis under label scarcity

Rib fractures are undercounted on routine CT because fine-grained, per-fracture annotation is expensive and few centres produce it at scale. AIRib — my MD-thesis work, developed with IIT Jodhpur and AIIMS Delhi under a ₹3.9M TIH iHub Drishti grant — is an end-to-end pipeline for detecting, characterising, and prognosticating traumatic rib fractures from CT. Related work on fine-grained rib-fracture diagnosis, which encodes the fracture hierarchy directly into the model using hyperbolic embeddings and on which I am a co-author, followed at MICCAI 2025.

→ RSNA Trainee Research Prize 2023 (AIRib); MICCAI 2025, Pate et al. (co-author). Publications →

Evaluating frontier AI against clinicians, and where it fails

Radiology's Last Exam (RadLE) is a CRASH Lab benchmarking programme that tests frontier multimodal AI on difficult, clinically grounded radiology cases, measuring not just accuracy but calibrated confidence, safe deferral, and readiness for autonomous use. RadLE 1.0 pitted five frontier models against four board-certified radiologists and four trainees on 50 expert-level cases; board-certified radiologists scored 83% against the best model's 30%. RadLE 2.0 extends this to 200 cases and 16 models. Separately, I work on where LLM-augmented reporting pipelines break under distribution shift, and what governance those failures imply before they reach a clinical report.

→ Radiology's Last Exam (RadLE), arXiv preprint, 2025; RSNA 2023, two Cutting-Edge Oral presentations on LLM-augmented speech recognition for radiology reporting; "From Chatbots to Agentic Workflows: Ensuring Responsible Deployment of Large Language Models in Radiology," Indian Journal of Radiology & Imaging, 2026. Publications →

Multimodal foundation models for medical image interpretation

With the Rajpurkar Lab at Harvard, I work on generalist foundation models that interpret medical images across modalities rather than one narrow task at a time — including MedVersa, developed with Hong-Yu Zhou, Julian Acosta, Subathra Adithan, Eric Topol, and Pranav Rajpurkar, and published in NEJM AI. I led the radiologist evaluation workstream and contributed clinical interpretation of the report-generation results.

→ MedVersa: A Generalist Foundation Model for Diverse Medical Imaging Tasks, NEJM AI, 2026. Publications →

LLM reliability in radiology practice and education

Before a language model is trusted with a clinical decision, its reliability needs to be measured against the specific task, not a general benchmark. I've evaluated LLM accuracy on FRCR examination questions and on imaging-referral decisions for pulmonary embolism, both with Sarangi PK and colleagues.

→ "Evaluating ChatGPT-4's Performance in Identifying Radiological Anatomy in FRCR Part 1 Examination Questions" and "Radiologic Decision-Making for Imaging in Pulmonary Embolism," Indian Journal of Radiology & Imaging, 2024–2025. Publications →