Medical AI Evaluation · Clinician-in-the-Loop Review

A practising surgeon who reviews what medical AI actually says.

I am a laparoscopic surgeon with more than 20 years in clinical practice. I evaluate AI-generated medical answers and illustrations, assess clinical reasoning and differential diagnosis, write difficult cases and scoring rubrics, and design the review steps that keep a qualified clinician accountable for what reaches a patient.

Available for evaluation, rubric design, clinical review and advisory work with healthcare-AI teams.

  • More than 20 years in clinical practice
  • MBBS · MS (Surgery) · FACS · FICS · FIAGES
  • Diploma in Laparoscopic Surgery, Strasbourg, France
Portrait of Dr Rajarshi Mitra
What I Evaluate

The most consequential failures in medical AI are clinical, not merely grammatical

Automated checks can verify structure, schema, links and predefined rules. In this project, those checks did not establish whether a duct had been routed to the wrong anatomical structure, whether a pain marker sat on the wrong side of the body, or whether escalation advice was clinically appropriate. Those questions required clinical judgement before release.

Medical answers and patient-facing health communication

Accuracy, completeness, and what a patient would reasonably do after reading them.

Clinical reasoning and differential diagnosis in model output

Whether the reasoning path and the differential hold up, rather than whether the final answer happens to be right.

Unsafe advice, unsupported claims, and missing escalation

What a clinician would act on that the model did not say, and what it said that no clinician would.

Evidence and guideline adherence

Whether a claim stays inside the evidence available to support it.

Medical illustrations and visual clinical content

Anatomy, laterality, procedural accuracy and unsupported claims introduced by the image rather than the text.

Clinician-in-the-loop review workflows

Where a clinician has to sit in the process, and what they must actually see.

Review vs Testing

A passing technical test suite is not, by itself, a clinical safety signal

What automated checks can establish

  • Schema validity
  • Internal consistency
  • Link and build integrity
  • Structural completeness
  • Regression against known-good output

What requires clinical judgement before release

  • Anatomical and procedural accuracy
  • Correct laterality
  • Whether a claim is supported by the cited evidence
  • Whether escalation advice is present and appropriate
  • Whether a patient would be misled

Both matter. Neither substitutes for the other, and neither alone should authorize publication of medical content.

Ways to Work Together

Where a practising clinician is useful to an AI team

Medical-answer and reasoning evaluation

Scoring model output against clinical rubrics, including difficult and adversarial cases.

Rubric and evaluation-set design

Building difficult cases and scoring criteria that require clinical input to be medically meaningful and consistently assessed.

Clinical review inside a content or product workflow

Acting as the named clinical reviewer with authority to reject and request correction.

Clinician-in-the-loop workflow design

Advising on where review belongs, what reviewers need to see, and what they must be able to stop.

Medical-illustration and patient-communication review

Reviewing generated figures and patient-facing explanations before release.

I contribute clinical judgement and workflow design. I am not a regulatory or compliance adviser, and I do not certify systems as safe or compliant.

Featured Case Study

Medical AI governance · Gallbladder Surgeon Abu Dhabi

From Complex Automation to Clinician-Governed Medical Content

An AI-assisted medical content workflow was simplified from a multi-stage orchestration into a single principal drafting step — without removing the source evidence, structured claims, deterministic checks, clinical review, bounded correction or explicit publication authorization.

  • 16 English articles at the verified 17 August 2026 repository snapshot
  • Three retained correction cycles, rejections included
  • One five-gate approval record
Read the case study
Evidence

What I can show you

Three retained clinical correction cycles

Before-and-after illustration evidence with the rejections preserved, not only the approved output.

A five-gate approval record

One representative article with claims, brief, clinical, editorial and preview gates recorded against a named reviewer and date.

A dated repository snapshot

16 English articles — 11 condition and five surgery — at the verified 17 August 2026 snapshot.

A documented limitation register

The case study states what the evidence does not establish, including where per-run logs were not retained.

Each of these is set out with its source and its limitation in the case study.

Background

Practising surgeon first

More than 20 years in clinical practice as a laparoscopic surgeon in Abu Dhabi. The evaluation work is grounded in that, not adjacent to it — the errors I catch in AI output are the errors I am trained to catch in clinical work.

MBBS · MS (Surgery) · FACS · FICS · FIAGES

Diploma in Laparoscopic Surgery, Strasbourg, France

Get In Touch

Working on medical AI and need a clinician in the loop?

Tell me what you are building and where clinical judgement would help. I read every enquiry personally.