Suvrankar Datta

Founder and Group Lead, CRASH Lab · Simons Ashoka Early Career Fellow

Clinical AI should know when to answer—and when to hand over.

I am a radiologist and physician-scientist evaluating when medical AI can be trusted, when it should defer, and whether it holds up in Indian clinical workflows. My work spans RadLE, MedVersa, equitable radiology models and AI scribes. Explore the research.

Email · LinkedIn · Google Scholar

Suvrankar Datta presenting at RSNA 2025

Selected Work

Radiology's Last Exam (RadLE)

Co-first author and benchmark co-lead

An uncertainty-aware evaluation of frontier multimodal AI against radiologists. RadLE 2.0 evaluates 16 systems across 200 expert-level cases using complementary measures of accuracy, reliability, safety, confidence and handover readiness.

arXiv

PrAImaan

Clinical Research and Technical Teams Supervisor, under PI Prof. Anurag Agrawal

A Gates Foundation-supported programme for blinded, reproducible and locally relevant evaluation of health AI systems in India.

Project page

V2DD

Clinical Evaluation Team Supervisor, under PI Prof. Anurag Agrawal

Evaluation of multilingual AI scribes and voice-to-structured-data systems across transcription fidelity, critical omissions, hallucinations, clinician correction burden, privacy, usability and interoperability.

Grant record

Equitable radiology foundation models

Principal Investigator

A three-year DST–A*STAR India–Singapore programme developing indigenous radiology datasets, subgroup evaluations and culturally congruent vision-language models for South Asian populations.

Announcement

PREDICT-AI

Principal Investigator

An IndiaAI CATCH-supported consortium with Tata Memorial Hospital and Manentia AI developing decision support for post-chemotherapy RP-LND planning in patients with testicular cancer.

Announcement

Additional selected work — AIRib, MICCAI 2025, MedVersa, and LLM-augmented reporting studies — is on the Research page.

The work starts with how care is actually delivered.

Rural primary care — JIPMER

During a two-month residential posting at a primary health centre in Ramanathapuram, I shared alternate emergency coverage, conducted regular NCD clinics and camps, and referred patients to higher centres in a setting affected by medicine shortages and the absence of operative facilities. I worked with patients in Tamil and Hindi.

High-volume radiology — AIIMS New Delhi

I worked in emergency, inpatient and outpatient radiology for patients referred from across India — interpreting more than 200 examinations independently during emergency shifts, mostly CT and ultrasonography, through roughly a year of COVID-related duties and the move from semi-digital to fully digital imaging.

These experiences shape the evaluations I build. A useful system must perform with incomplete information, limited time, variable infrastructure and constrained referral capacity. It must know when not to answer, and it must reduce rather than transfer clinical burden.

What we evaluate for AI systems

Capability
Can the system perform the task under clearly defined conditions? In RadLE 2.0, the RadLE-C confidence-weighted capability index evaluates autonomous diagnostic agents in radiology against human radiologists, making the current capability gap explicit.
Reliability
Is confidence aligned with correctness, and does performance remain stable across repeated runs? RadLE-R measures how often diagnoses delivered with high confidence are correct. We are also developing RADAR and TRUST metrics to assess the reliability of multi-agent systems that produce preliminary clinical records and radiology reports.
Safety
Which errors could cause harm, and can an evaluation distinguish unsafe, confident errors from low-confidence failures? In V2DD, we are developing metrics for ambient AI systems that convert doctor–patient conversations in local languages and dialects into structured clinical records and clinical ontologies, with attention to clinically significant omissions, hallucinations and translation errors. Through PrAImaan, we are extending safety evaluation across diverse Indian clinical contexts within a nationally coordinated framework.
Deployment fit
Does performance hold across dialects, clinicians, sites, interfaces and real clinical workflows? Our multi-institutional studies examine when local, contextual metrics are needed for hard-to-verify or subjective tasks, including differences in physicians’ requirements, preferences and workflows. We presented initial healthcare work using DSPy and GEPA at RSNA 2025; current work benchmarks state-of-the-art small automatic speech recognition models for deployment on edge devices in low-connectivity settings.
Impact
Does the system improve health-worker performance, care quality or patient outcomes? In MedVersa, led by Harvard investigators, generalist-model draft reports reduced radiologists’ reporting time and clinically important discrepancies while supporting accurate reporting.

Benchmarks are increasingly necessary, but current leaderboards are not sufficient. We are extending the evidence pathway from retrospective, leaderboard-based evaluation to workflow testing and prospective validation.

Selected Publications

Full publication list →

Background

I moved from clinical radiology into full-time medical-AI evaluation after training and practice across public-sector, emergency and rural settings in India.

  • 2025 – present

    Simons Ashoka Early Career Fellow

    Koita Centre for Digital Health, Ashoka University

  • 2025 – present

    Founder and Group Lead, CRASH Lab

    Clinician-led evaluation of responsible autonomous systems in healthcare

  • 2023 – 2024

    Former Senior Resident, Radiodiagnosis & IR

    AIIMS New Delhi

  • 2014 – 2023

    MD, AIIMS New Delhi · MBBS, JIPMER

    Radiology training grounded in high-volume public-sector and rural care

Full experience and education →

Contact

Open to research collaborations and speaking engagements in AI and healthcare. Based in New Delhi, India.

ORCID · Google Scholar · ResearchGate