Video summary

AI and Clinical Reasoning: A New Era of Medical Education and Assessment

Main summary

Key takeaways

Educational

Main ideas and lessons

  • Clinical reasoning is fallible and widely affected by diagnostic errors.

    • Even with good training and systems, diagnostic errors still occur and significantly impact patients.
    • The National Academies of Medicine (NAM) report is referenced as a key decade-old source on the pervasive impact of diagnostic errors.
  • Clinical reasoning (as an educational target) is an iterative, multi-step process.

    • It includes:
      • gathering information
      • deciding tests
      • forming a working diagnosis
      • discussing with the patient
      • creating management plans
    • When reasoning goes wrong, it is often multifactorial, but faulty thinking is commonly involved.
    • Mark Graber’s work is referenced as a landmark study showing cognitive faults can contribute.
  • What educators knew before generative AI: proven training strategies

    • Teach clinical reasoning using approaches that promote:
      • Knowledge organization
        • problem representations (abstract representations of a case)
        • illness scripts (how typical disease presentations unfold)
        • management scripts (typical care pathways)
        • diagnostic schemas (structured approaches to problems)
      • Structural reflection
        • metacognition-focused practices that help learners think about their thinking
        • diagnostic timeouts
      • Team-based thinking
        • distributed cognition and diagnostic teams with diverse viewpoints
        • leveraging AI to support team processes (not just individual thinking)
    • The talk notes that AI will alter how these are implemented and emphasized, but many remain relevant.
  • AI’s role in diagnosis: “LLMs outperform humans” is true in constrained settings, but not the end of the story

    • In controlled/simulated cases, LLMs often outperform clinicians when information is curated.
    • Real clinical environments are more complex; therefore, results don’t automatically mean “replace clinicians.”
    • Key point: AI + human does not automatically improve performance unless there is explicit structure.
  • Why “prompting humans to use AI” without structure often fails

    • Studies discussed suggest that simply telling clinicians to use an LLM—or adding the LLM into the workflow without a structured reflection strategy—may not improve performance.
    • The talk links this to structural reflection literature: reflective prompting alone doesn’t reliably improve outcomes.
  • Promising evidence: structured AI-human collaboration can improve diagnostic performance

    • The speaker describes multiple studies (including simulated and real-world clinical settings) showing:
      • AI-first vs human-first ordering matters
        • “human first / then AI second” can outperform cases where humans adopt AI output prematurely
        • AI guidance may be better at improving performance in lower-scoring cases
      • Not all cases improve, but meaningful improvements can occur when the interaction is well-designed.
  • Real clinical workflow examples

    • Urgent care/primary care randomized by clinics (AI + clinician vs standard practice)
      • AI flags potential diagnostic discordance.
      • Example: a clinician is prompted to reconsider an antibiotic choice when a red-flag discordance appears.
      • Independent reviewers found AI + clinician charts scored better across history, testing, diagnosis, and management.
      • Over time, clinicians showed reduced unnecessary AI interventions—suggesting learning from feedback.
    • Patient-interactive AI conversation + clinician in the loop
      • AI gathers patient information via an interactive conversation prior to the visit.
      • Clinicians receive transcript-derived output (not necessarily the AI’s full management plan).
      • Reviewers found similar diagnostic/management quality; clinicians still integrate context the AI lacks.
      • Patients and clinicians reportedly found it helpful (more time for the visit; improved perceived support).
  • Curriculum implication: learners will use AI anyway, so training must guide how

    • The talk emphasizes a “seat at the table” approach: educators must shape policies, workflows, and assessment.
    • Educational tension:
      • Upskilling (human-AI synergy improves performance)
      • vs deskilling/never skilling/mis-skilling (loss of ability or dependence when tools are introduced incorrectly)
  • Ambient documentation: moving from “should we?” to “how should we do it?”

    • AI can reduce documentation burden, affect wellness/burnout, and translate narrative into patient-appropriate literacy.
    • Concerns remain about documentation quality (being studied).
    • Conclusion: it’s already happening—so education must teach safe, critical use and interpretation.
  • Policies and competency development

    • Institutions need:
      • guidelines on what tools can be used for what tasks
      • privacy and data governance (public tools vs HIPAA/FERPA-compliant institutional tools)
      • faculty development so supervision reinforces correct behavior (addressing the “hidden curriculum”)
  • Assessment shift: use AI to enhance feedback and measure “diagnostic performance”

    • The speaker proposes moving assessment from only “diagnostic reasoning” artifacts to also measuring diagnostic performance (e.g., whether a diagnostic delay occurred).
    • Example:
      • Disease-based approach starting with VTE (venous thromboembolism) due to high error rates.
      • Validation uses Safer Dx (human adjudication tool/work by Hardeep Singh) to classify diagnostic delays.
      • A model uses prompts to determine whether residents considered VTE and uses imaging report gold standards to adjudicate outcome/discordance.
      • A dashboard provides iterative feedback with standard-setting and trend views.
  • Ambient assessment and feedback (future direction)

    • Work discussed uses ambient recordings (not only documentation) to give feedback on communication skills (and potentially clinical skills more broadly).
    • Rationale: documentation artifacts may no longer be the only high-quality retrospective signal.

Methodology / instructional framework presented (detailed bullets)

A) Human-AI integration principles (education strategy)

  • Teach structured collaboration, not “AI as an unstructured add-on.”
    • Encourage learners to commit to their own thinking first (to reduce bias from deferring to AI output).
    • Then integrate AI as a second opinion (structured workflow).
  • Use structured reasoning tools derived from clinical reasoning literature
    • Apply a framework conceptually based on structural reflection:
      • generate a differential with supporting vs opposing reasoning
      • estimate probabilities / next steps
  • Require critical appraisal
    • After AI output, teach learners to evaluate evidence quality and safety before acting.

B) “AI-first opinion vs second opinion” teaching model (implementation concept)

  1. Step 1: Human-first differential/plan
    • Learners produce their own differential diagnosis and management considerations.
  2. Step 2: Generate AI-assisted second opinion
    • Learners obtain AI output after committing to their own reasoning.
    • Prompts scaffold the learner through structured reflection.
  3. Step 3: Critically appraise and reconcile
    • Learners compare AI output to their own reasoning:
      • What changed?
      • What was newly considered?
      • What stayed consistent?
  4. Step 4: Reflection and commitment
    • Learners explicitly decide how they will use AI in future cases.
    • They assess whether the AI helped or distracted from correct reasoning.

C) Faculty/supervisor micro-skills framework for ward use (“Depth AI”)

  • The talk references a framework (Depth AI) to guide faculty conversations.
  • Faculty prompts include:
    • What tools did you use?
    • How did you use them? (what prompts did you enter)
    • What evidence did you verify? (e.g., learner answers “I didn’t” → teach evidence checking)
    • Safety / ethics / privacy checks
      • How did you assess accuracy/safety?
    • Feedback and improvement loop
      • How might you change your AI use next time?
    • Tailor teaching to gaps
      • Where the learner is weak determines which general principles to emphasize.

D) Institutional policy approach (privacy/data governance)

  • Do not treat all tools equally
    • Public tools (e.g., “OpenEvidence”) vs evidence-vetted tools (e.g., “UpToDate Expert AI”) vs internal HIPAA/FERPA tools (e.g., “UltraViolet AI”) have different constraints and reliability profiles.
  • Data handling rules
    • Public tools: avoid patient-identifiable or detailed EHR data.
    • Institutional compliant tools: may allow appropriate integration of EHR data (where set up by the institution).
  • Assessment controls
    • For certain assessments, AI usage may be prohibited or restricted.
    • Guardrails must be explicit; otherwise “using AI when not disallowed” becomes a policy gap.

E) Assessment methodology using AI (diagnostic performance tooling)

  • Ground level: define the human behavior / skill being assessed
    • Start by specifying clinical reasoning/performance behaviors (e.g., differential quality, reasoning explanation).
  • Use AI to score artifacts
    • Example: admission note reasoning elements (e.g., did they include a differential; did they explain reasoning).
  • Disease-based performance validation workflow (example VTE)
    • Create prompts to detect whether residents considered the condition (e.g., VTE).
    • Use imaging report narratives and timing rules as gold standards.
    • Use Safer Dx (human-reviewed diagnostic delay adjudication) as the validation benchmark.
    • Provide dashboard feedback on:
      • cases with/without delays
      • trend over time
      • contextual factors and balancing measures (e.g., avoid “everyone ordered imaging” behavior dominating interpretations)
  • Resident advisory panels and quality improvement
    • Residents participate in choosing feedback formats and goals.
    • Implement dashboards with data visibility and reflective exercises.

Speakers / sources featured

People / speakers

  • Bob Wachter (introducer; Chair of the Department of Medicine at Moffitt)
  • Berdychev (primary visiting professor speaker; NYU Grossman School of Medicine; NYU titles described in intro)
  • Michelle Guy (Master Clinician inductee)
  • Raman Chawla (Master Clinician inductee)
  • Vivek Jain (Master Clinician inductee)
  • Scott Steiger (Master Clinician inductee)
  • Sunny Wang (Master Clinician inductee)
  • Gurpreet Dhaliwal (namesake of the master clinician lectureship; mentioned as present)
  • Eric Jameson (audience member; UCSF instructor/faculty; asks questions)
  • Nina (audience member; internal medicine resident at the program; asks questions)
  • Verity (audience member; asks question about hidden curriculum and training residents to supervise students)
  • Dr. Okter (appears to refer to Bob Wachter; introduction contains a brief mis-transcription)
  • Daniel Sartori (implementation lead mentioned)
  • Lauren Hiry (resident advisory panel member mentioned)
  • Bijal Rajbhat (resident advisory panel member mentioned)
  • Dr. Briscoe (referenced as connected to the Depth AI framework; may join virtually)
  • Dr. Abdo (referenced as connected to the Depth AI framework)
  • Mark Triola (NYU colleague referenced as creating the prompt-a-thon/platform; appears in curriculum/tools discussion)
  • Marc Triola (same person as above, referenced with full name)
  • Rodman (referenced as a leader in the field of AI education competencies; name appears in transcript)
  • Hardeep Singh (Safer Dx referenced; leader in diagnostic error field)
  • Yuliya Umcheva (referenced in ambient assessment work)
  • Amy Jessi Burke-Ravelle (referenced in ambient assessment work)

Organizations / reports / tools referenced

  • National Academies of Medicine (NAM) report on diagnostic errors
  • Safer Dx (Hardeep Singh / diagnostic error adjudication tool)
  • UpToDate and UpToDate Expert AI
  • OpenEvidence
  • UltraViolet AI (described as institutional HIPAA/FERPA-compliant)
  • Depth AI framework (for faculty conversations)
  • One-minute preceptor / micro-skills (referenced as an analogy for the feedback framework)
  • OSCEs (assessment context referenced)
  • EHR (Electronic Health Record)
  • HIPAA / FERPA (privacy compliance frameworks referenced)

Original video