Zyla Health · Clinician Console · AI Risk Engine

Catching patients at life-threatening risk — in time to intervene.

An AI risk-stratification engine and clinician review console that reads a member's full medical history on every new lab report, sorts a 200,000+ base into five risk tiers, and routes life-threatening cases to a human for action within 0–24 hours — with every call reviewable, attributable and auditable.

RoleLead Product Designer
SurfaceInternal clinician console
Timeline2026 · Phase 1 POC
TeamClinical · Data Science · Eng
ScopeResearch → IA → UI → Design system
zylaCONSOLEHimanshu Agarwal
← Risk stratification dashboard ⤓ Export
▤ Filter by date
Total Patients
180
Life Threatening Risk
4
Very High Risk
9
Pending Reviews
9
AI Accuracy
87.3%
False Positive
8.2%
False Negative
4.5%
Life threatening risk cases increased by 12% this week
Patient IDAccount codeRisk strat. dateRisk levelAI ConfidenceTriggered byEvaluation statusEvaluator
121334AT1234521 Mar 2026Life threatening94%BP CriticalUnder reviewTrishla
295817ER7612321 Mar 2026◓ Very high risk78%Heart RateCorrectPriyanka
128495LK3210921 Mar 2026High risk85%Lab ValuesCorrectAshish
907362JH2109821 Mar 2026◓ Moderate risk92%Routine CheckCorrectTrishla
286430CV6543221 Mar 2026Low risk88%Sepsis AlertPendingEvaluate
740196SD2345621 Mar 2026Life threatening81%Post-OpIncorrectTrishla
384670AZ8761221 Mar 2026◓ High risk87%Stroke RiskCorrectDr Shweta
Showing 1–20 of 180
The Risk-Stratification dashboard in Zyla's clinician console. ⚡ marks an AI-set tier; ◓ marks one a clinician has set. Recreated from the live design — production swaps in the exported Figma frame.
At a glance

Turning a flood of medical records into a ranked, reviewable worklist.

0
Member base the engine screens
0
Risk tiers, low to life-threatening
0
Action window for life-threatening flags
0
AI accuracy, tracked live vs. clinicians

The premise is serious: most people heading toward a critical event are never flagged in time — the warning signs sit buried in lab reports and history no team can read across 200,000 members. We built an AI engine that reads all of it on every new report, ranks each patient into one of five tiers, surfaces the dangerous ones to a clinician, and learns from every verdict the clinician returns.

The problem

You can't manually watch 200,000 people for the one signal that means "act now."

Clinical signals arrive constantly — digitised labs, diagnoses, medications, vitals, hospitalisation and family history. Buried in that stream are the rare, time-critical cases: a troponin spike, haemoglobin below 7, a suspected malignancy on imaging. Miss one, and a member's condition can quietly worsen before anyone notices. Review every record by hand, and the team is overwhelmed long before member 200,000.

The motto behind the project was blunt: save lives. Predict risk from a member's records and labs, so care reaches them in time — especially when the risk is life-threatening.

The design challenge: turn an unbounded, noisy data stream into a ranked, trustworthy worklist — without ever letting an algorithm make a high-stakes clinical call on its own. Speed for the clinician, safety for the patient, accountability for the model.

Goals & principles

What "good" had to mean here.

Surface the most urgent, fast

The highest-acuity cases rise to the top instantly, with a clear 0–24h action window — never lost in a list.

Human in the loop, always

The AI proposes; a clinician disposes. Every call is reviewable, overridable, and permanently attributable to a person.

Make the model earn trust

Every verdict feeds live accuracy, false-positive and false-negative tracking — the system gets safer as it runs.

Who it serves

One engine, three very different users — no added friction for any of them.

Primary user

The clinical reviewer

Triages the queue, validates or overrides the AI, and is accountable for the call. Needs speed, evidence, and a frictionless "the AI got this right / wrong."

Beneficiary

The patient

Never sees the dashboard, but everything points at them: the right tier routes them to the right intensity of care, in time.

Downstream

The care team

Acts on confirmed tiers. A clean, confident hand-off means time spent on care, not on chasing data.

The core model

Five tiers, each tied to an action — not just a label.

A tier is only useful if it says what to do and how fast. So each maps to a care intensity and a response window — and never relies on colour alone.

Life-threateningHigh probability of rapid, severe deterioration without immediate intervention — panic-range labs, structural hazards on imaging, a sharp decline in trend.Act in 0–24hER / stat consult
Very high riskSevere but potentially manageable with aggressive intervention — e.g. uncontrolled hypertension with compounding factors.DaysIntensive program
High riskClear clinical risk needing proactive management before it escalates.WeeksActive management
Moderate riskBorderline markers worth watching, with guidance and scheduled re-checks.RoutineGuided + re-check
Low riskWithin normal ranges — maintenance, education and periodic screening.MaintenanceLight-touch

What feeds the call. The engine reads across the member's full record — never a single test in isolation:

Digitised lab reportsDiagnosesOngoing medicationChief complaintVitalsHospitalisation historyFamily history
The solution · 01

A worklist that reads like triage, not a spreadsheet.

The reviewer lands on the cohort at a glance — totals, criticals, pending reviews, and how the model itself is performing. A standing alert flags what changed. Then the table ranks every patient with the risk call, the AI's confidence, and the single signal that triggered it, so a clinician knows what to open first. (See the dashboard up top.)

Model on the masthead

AI accuracy, false-positive and false-negative sit in the KPI row — the model's vital signs are always in view, not hidden in a report.

Who set this tier?

A ⚡ bolt means the AI set the tier; a ◓ means a clinician did. Provenance is legible at a glance, on every row.

Status, not guesswork

Under review · Correct · Incorrect · Pending — colour-coded pills, plus an inline Evaluate action when a case still needs a human.

The solution · 02

Inside a case: the AI's reasoning, laid bare.

Opening a patient shows the AI's risk call and the evidence behind it — including, crucially, what data it did and didn't have. A "Data Coverage" strip makes the AI's blind spots honest, so a clinician knows how much weight the call deserves.

zylaCONSOLE
← AI summary Save ChangesClose
PID-121334  ·  Creation Date: 11 Mar 2026  ·  Age: 67  ·  Gender: Male  ·  BMI: 31.2
AI Risk Summary Very high risk Updated on 22 Mar 2026 View previous AI summary
Elevated LDL cholesterol (190 mg/dL) and uncontrolled hypertension (165/95 mmHg) indicate high cardiovascular risk. Patient presents multiple risk factors including obesity (BMI 31.2) and pre-diabetic glucose levels. Immediate intervention recommended with lifestyle modifications and potential medication adjustment.
KEY RISK DRIVERS
Severely elevated LDL cholesterol (190 mg/dL)Critical
Uncontrolled hypertension (165/95 mmHg)Critical
Obesity (BMI 31.2) with metabolic syndromeHigh
Pre-diabetic glucose levels (142 mg/dL)High
Data Coverage
Lab reportsVital signs Diagnosis historyFamily history Chief complaintsMedication recordsHospitalisation history
Clinical Data
VITAL SIGNS
BLOOD PRESSURE
165/95 mmHg Elevated
Recorded on 12 Mar 2026
HEART RATE
88 bpm Elevated
Recorded on 12 Mar 2026
LAB VALUES
CHOLESTEROL
280 Elevated
LDL
190 Elevated
HDL
38 Lower
GLUCOSE
142 Elevated
HBA1C
7.8% Elevated
Recommended Actions
Reduce saturated fat & increase activity to manage cholesterol → Cholesterol, LDLCritical
Monitor blood pressure regularly and reduce sodium intake → Blood PressureCritical
Schedule follow-up glucose monitoring for pre-diabetic range → Glucose, HbA1cHigh Priority
The AI summary detail. Data Coverage (here: only labs + vitals available) tells the clinician exactly how complete the AI's picture was.
Honest about blind spots

Showing which of the seven inputs were missing turns the AI from a black box into a colleague that admits what it couldn't see.

Evidence, not just a verdict

Every driver is tied to a real value and severity; the raw vitals and labs sit right below, so the call is checkable in seconds.

Actions, pre-linked

Recommendations point back to the exact metrics that triggered them — the bridge from "what's wrong" to "what to do."

The solution · 03

The verdict: AI proposes, the clinician disposes.

A reviewer confirms or corrects the AI in a single, structured flow. Mark it Correct and move on; mark it Incorrect and the UI requires a new risk level and a written reason — so no override is silent, and every correction becomes training signal.

Report Evaluation
Current status: Under review
Evaluator name: Dr. Shweta
Is the AI risk assessment correct?
IncorrectCorrect
Update risk level
Very high risk ▾
Reason for incorrect evaluation*
Explain why the AI assessment is incorrect…max. 500 words
Evaluator Summary
BIUS{}🙂
Patient presents a significant cardiovascular risk profile with elevated cholesterol and uncontrolled hypertension. While severe, the risk is potentially manageable through aggressive lifestyle modification — reclassifying pending response to initial interventions.
Evaluator recommendations
BIUS
• Reduce saturated fat; recommend Mediterranean diet with structured exercise
• Initiate antihypertensive therapy review with primary care provider
• Schedule follow-up cholesterol panel & HbA1c in 6 weeks
The Report Evaluation flow. Choosing "Incorrect" unlocks a required risk-level change and a written reason — accountability built into the interaction.
The solution · 04

In medicine, a decision needs a name on it.

Before a clinician can evaluate, they explicitly take ownership — a deliberate friction point. From then on, every action is attributed to them and recorded in an immutable activity log. This is what makes an AI tool safe to use in a clinical setting: not just accuracy, but accountability and audit.

!

Assign yourself as the evaluator?

By confirming, you'll take over as the evaluator for this AI summary. From this point onward:

  • All evaluations on this summary will be attributed to you
  • This action is permanent — once assigned, the evaluator cannot be changed
CancelYes, assign to me
Activity Log
22 Mar 2026
03:30 PM
Evaluation summary
Dr. Shweta
21 Mar 2026
03:30 PM
Evaluator changed
Trishla → Dr. Shweta
20 Mar 2026
03:30 PM
Evaluation summary
Trishla
11 Mar 2026
09:15 AM
Risk assessment generated
AI System
Left: assigning yourself as evaluator is permanent and explicit. Right: the activity log is a full audit trail, from the AI's first call to every human hand-off.
The solution · 05

Risk is a moving target — so the design shows the trajectory, not a snapshot.

New labs re-trigger the engine, so a patient is assessed again and again. The comparison view puts versions side by side — AI vs. the first clinician, then clinician vs. clinician — so a reviewer sees whether someone is improving or sliding. Per-metric trend charts add an AI insight on the direction of travel.

Blood Pressure Trend
1801359045 2023202420252026
Last updated value: 165/95 mmHg
Status: Elevated
↗ AI Insight

Value increased by 6.5% over the past 2 months. Consistent upward trend observed, indicating worsening cardiovascular risk profile.

Left: the evaluation-comparison modal — here the clinician reclassifies the AI's "life-threatening" to "very high" and records why. Right: a per-metric trend with an AI read on the direction of travel.
How it all connects

A loop that gets safer every time it runs.

01

New report

Every digitised lab report triggers a fresh stratification call automatically.

02

AI engine

Guided, condition-specific prompts read the full record and assign a tier + confidence.

03

Dashboard

The case surfaces in the reviewer's ranked worklist with its evidence and trigger.

04

Clinician verdict

Correct, or Incorrect + reason — every override captured and attributed.

05

Model hardens

Verdicts feed accuracy / FP / FN tracking and tune prompts — then route to a care program.

A key research finding shaped this: open-ended prompts produced unreliable — sometimes dangerous — calls, especially on radiology. Guided, condition-specific prompts with an explicit "Uncertain — needs clarification" escape hatch performed far better. So the system never forces a verdict on ambiguous data; it routes uncertainty to a human instead of guessing.

Craft & system

Designing for a clinician who is tired, busy, and can't afford a misread.

Scannability first

Risk, confidence, trigger and status live in one row, so a reviewer judges "open now / later" without a click.

Safe by default

Colour never carries meaning alone — every tier and status pairs a hue with a label for colour-blind safety and clinical clarity.

One governed system

KPI cards, the data table, status pills, risk badges, rich-text editors and Excel-style column filters — all from Zyla's governed design system, reused across AI and standard dashboards.

Impact & status

From concept to a working clinician-in-the-loop POC.

0
Member base the engine is built to screen
0
AI accuracy, measured vs. clinician verdicts
0
Risk tiers mapped to care intensity
0
Of risk calls reviewable & attributable

Phase 1 shipped the MVP dashboard, the AI-summary review, the evaluation workflow and the comparison/trend views that made the POC real — proving guided prompts could stratify reliably, giving clinicians a fast, accountable way to validate or correct the AI, and turning every verdict into training signal. The accuracy, false-positive and false-negative metrics aren't vanity numbers; they're the instrument panel that tells the team when the model is safe to lean on and where it still needs a human.

Where I'd take it next

Designing the engine into a full care system.

Building on what's shipped, a few directions I'd push to deepen safety, speed and trust.

01

Live SLA timers for life-threatening cases

A visible 0–24h countdown on every critical case, with auto-escalation to the on-call clinician and care team if it isn't actioned in time.

02

Evidence drill-down

Tap "triggered by" to jump straight to the exact panic value and the source report it came from — clinicians trust what they can verify in one click.

03

Confidence-aware routing

Low-confidence and "needs clarification" cases get a distinct state and jump the queue, so ambiguity is never silent — pairing with the AI Confidence already on every row.

04

Model-health & drift view

Take the KPI accuracy/FP/FN further: trend them over time and per condition, so the data team sees exactly where the model is slipping.

05

Patient risk timeline

Extend the comparison + trend views into a single timeline of every re-stratification, so a slow slide toward critical is visible at a glance.

06

Tier → care-program automation

Close the loop the PRD points to: a confirmed tier auto-proposes the matching care-program intensity, hand-off ready for the care team.