Validating a clinical documentation assistant for a hospital network
The challenge
The network was piloting an ambient tool that drafts clinical notes from a conversation. The failure that mattered was not a typo. It was a note that added a symptom or a plan the clinician never stated, or dropped one that they did. In a clinical record, that is a safety problem.
Our approach
- Built the risk map with clinical input, so testing focused on the failures that affect patient safety.
- Tested groundedness of every drafted note against the source transcript.
- Ran adversarial and safety probes for fabricated findings and unsafe suggestions.
- Kept expert humans reviewing a sample of outputs, since clinical review cannot be fully automated.
- Produced audit-ready evaluation logs and re-tested after each model update.
Healthcare is where the gap between a good demo and a safe product is widest. A documentation assistant that is right most of the time is not good enough if the times it is wrong add clinical detail that was never said.
What we did
Kaycore has a physician co-founder, so the risk map was built with clinical input rather than guessed at. We measured whether each drafted note was grounded in the actual transcript, ran probes for fabricated findings, and kept expert humans reviewing a sample of the output. That combination of automated scoring and human review is the pattern that holds up in high-stakes settings.
Result
Drafted notes carried fewer unsupported statements, the evaluation produced audit-ready logs for governance review, and clinicians were more willing to rely on the drafts because the failure modes had been measured rather than assumed. Every model update triggers a re-test before it reaches the ward.
Representative engagement. Client details are anonymized and figures are illustrative ranges shown to convey the type and breadth of work, not a specific named result.
