Turning a flood of medical records into a ranked, reviewable worklist.
The premise is serious: most people heading toward a critical event are never flagged in time — the warning signs sit buried in lab reports and history no team can read across 200,000 members. We built an AI engine that reads all of it on every new report, ranks each patient into one of five tiers, surfaces the dangerous ones to a clinician, and learns from every verdict the clinician returns.
You can't manually watch 200,000 people for the one signal that means "act now."
Clinical signals arrive constantly — digitised labs, diagnoses, medications, vitals, hospitalisation and family history. Buried in that stream are the rare, time-critical cases: a troponin spike, haemoglobin below 7, a suspected malignancy on imaging. Miss one, and a member's condition can quietly worsen before anyone notices. Review every record by hand, and the team is overwhelmed long before member 200,000.
The design challenge: turn an unbounded, noisy data stream into a ranked, trustworthy worklist — without ever letting an algorithm make a high-stakes clinical call on its own. Speed for the clinician, safety for the patient, accountability for the model.
What "good" had to mean here.
Surface the most urgent, fast
The highest-acuity cases rise to the top instantly, with a clear 0–24h action window — never lost in a list.
Human in the loop, always
The AI proposes; a clinician disposes. Every call is reviewable, overridable, and permanently attributable to a person.
Make the model earn trust
Every verdict feeds live accuracy, false-positive and false-negative tracking — the system gets safer as it runs.
One engine, three very different users — no added friction for any of them.
The clinical reviewer
Triages the queue, validates or overrides the AI, and is accountable for the call. Needs speed, evidence, and a frictionless "the AI got this right / wrong."
The patient
Never sees the dashboard, but everything points at them: the right tier routes them to the right intensity of care, in time.
The care team
Acts on confirmed tiers. A clean, confident hand-off means time spent on care, not on chasing data.
Five tiers, each tied to an action — not just a label.
A tier is only useful if it says what to do and how fast. So each maps to a care intensity and a response window — and never relies on colour alone.
What feeds the call. The engine reads across the member's full record — never a single test in isolation:
A worklist that reads like triage, not a spreadsheet.
The reviewer lands on the cohort at a glance — totals, criticals, pending reviews, and how the model itself is performing. A standing alert flags what changed. Then the table ranks every patient with the risk call, the AI's confidence, and the single signal that triggered it, so a clinician knows what to open first. (See the dashboard up top.)
AI accuracy, false-positive and false-negative sit in the KPI row — the model's vital signs are always in view, not hidden in a report.
A ⚡ bolt means the AI set the tier; a ◓ means a clinician did. Provenance is legible at a glance, on every row.
Under review · Correct · Incorrect · Pending — colour-coded pills, plus an inline Evaluate action when a case still needs a human.
Inside a case: the AI's reasoning, laid bare.
Opening a patient shows the AI's risk call and the evidence behind it — including, crucially, what data it did and didn't have. A "Data Coverage" strip makes the AI's blind spots honest, so a clinician knows how much weight the call deserves.
Showing which of the seven inputs were missing turns the AI from a black box into a colleague that admits what it couldn't see.
Every driver is tied to a real value and severity; the raw vitals and labs sit right below, so the call is checkable in seconds.
Recommendations point back to the exact metrics that triggered them — the bridge from "what's wrong" to "what to do."
The verdict: AI proposes, the clinician disposes.
A reviewer confirms or corrects the AI in a single, structured flow. Mark it Correct and move on; mark it Incorrect and the UI requires a new risk level and a written reason — so no override is silent, and every correction becomes training signal.
Evaluator name: Dr. Shweta
• Initiate antihypertensive therapy review with primary care provider
• Schedule follow-up cholesterol panel & HbA1c in 6 weeks
In medicine, a decision needs a name on it.
Before a clinician can evaluate, they explicitly take ownership — a deliberate friction point. From then on, every action is attributed to them and recorded in an immutable activity log. This is what makes an AI tool safe to use in a clinical setting: not just accuracy, but accountability and audit.
Assign yourself as the evaluator?
By confirming, you'll take over as the evaluator for this AI summary. From this point onward:
- All evaluations on this summary will be attributed to you
- This action is permanent — once assigned, the evaluator cannot be changed
03:30 PM
03:30 PM
03:30 PM
09:15 AM
Risk is a moving target — so the design shows the trajectory, not a snapshot.
New labs re-trigger the engine, so a patient is assessed again and again. The comparison view puts versions side by side — AI vs. the first clinician, then clinician vs. clinician — so a reviewer sees whether someone is improving or sliding. Per-metric trend charts add an AI insight on the direction of travel.
Elevated LDL cholesterol (190 mg/dL) and uncontrolled hypertension (165/95 mmHg) indicate high cardiovascular risk. Multiple risk factors including obesity and pre-diabetic glucose.
- Severely elevated LDL (190 mg/dL) Critical
- Uncontrolled hypertension (165/95) Critical
- Obesity (BMI 31.2) High
Significant cardiovascular risk profile, but potentially manageable through aggressive lifestyle modification and medication. Reclassifying from life-threatening to very high risk pending response to initial interventions.
- Reduce saturated fat; Mediterranean diet + structured exercise
- Initiate antihypertensive therapy review
- Follow-up cholesterol panel & HbA1c in 6 weeks
Value increased by 6.5% over the past 2 months. Consistent upward trend observed, indicating worsening cardiovascular risk profile.
A loop that gets safer every time it runs.
New report
Every digitised lab report triggers a fresh stratification call automatically.
→AI engine
Guided, condition-specific prompts read the full record and assign a tier + confidence.
→Dashboard
The case surfaces in the reviewer's ranked worklist with its evidence and trigger.
→Clinician verdict
Correct, or Incorrect + reason — every override captured and attributed.
→Model hardens
Verdicts feed accuracy / FP / FN tracking and tune prompts — then route to a care program.
A key research finding shaped this: open-ended prompts produced unreliable — sometimes dangerous — calls, especially on radiology. Guided, condition-specific prompts with an explicit "Uncertain — needs clarification" escape hatch performed far better. So the system never forces a verdict on ambiguous data; it routes uncertainty to a human instead of guessing.
Designing for a clinician who is tired, busy, and can't afford a misread.
Scannability first
Risk, confidence, trigger and status live in one row, so a reviewer judges "open now / later" without a click.
Safe by default
Colour never carries meaning alone — every tier and status pairs a hue with a label for colour-blind safety and clinical clarity.
One governed system
KPI cards, the data table, status pills, risk badges, rich-text editors and Excel-style column filters — all from Zyla's governed design system, reused across AI and standard dashboards.
From concept to a working clinician-in-the-loop POC.
Phase 1 shipped the MVP dashboard, the AI-summary review, the evaluation workflow and the comparison/trend views that made the POC real — proving guided prompts could stratify reliably, giving clinicians a fast, accountable way to validate or correct the AI, and turning every verdict into training signal. The accuracy, false-positive and false-negative metrics aren't vanity numbers; they're the instrument panel that tells the team when the model is safe to lean on and where it still needs a human.
Designing the engine into a full care system.
Building on what's shipped, a few directions I'd push to deepen safety, speed and trust.
Live SLA timers for life-threatening cases
A visible 0–24h countdown on every critical case, with auto-escalation to the on-call clinician and care team if it isn't actioned in time.
Evidence drill-down
Tap "triggered by" to jump straight to the exact panic value and the source report it came from — clinicians trust what they can verify in one click.
Confidence-aware routing
Low-confidence and "needs clarification" cases get a distinct state and jump the queue, so ambiguity is never silent — pairing with the AI Confidence already on every row.
Model-health & drift view
Take the KPI accuracy/FP/FN further: trend them over time and per condition, so the data team sees exactly where the model is slipping.
Patient risk timeline
Extend the comparison + trend views into a single timeline of every re-stratification, so a slow slide toward critical is visible at a glance.
Tier → care-program automation
Close the loop the PRD points to: a confirmed tier auto-proposes the matching care-program intensity, hand-off ready for the care team.