Suki Is Bringing Researchers Together To Define What ‘Good’ AI Means
Healthcare organizations have spent the past several years trying to make artificial intelligence work inside their existing clinical workflows. The harder question that’s emerging is: How do we evaluate technology evolving faster than we can adapt?
That core tension is driving several initiatives to examine evaluation methodologies for healthcare AI. One of the newest comes from Suki, a healthcare AI company launching Science at Suki and the Suki Research Collaborative , a network that includes Regenstrief Institute, the National Center for Human Factors in Healthcare at MedStar Health, the University of Miami Miller School of Medicine’s Office of AI in Medical Education, and Rush University System for Health. The stated goal is to generate research, develop evaluation methodologies, and ultimately collaborate to build more consistent standards for measuring healthcare AI.
The timing reflects the growing importance of clinical artificial intelligence. AI has quickly moved from specific task-based solutions to tools deeply embedded in clinician-technology interactions. And increasingly, AI is shaping the overall experience of care. Ambient scribes in particular have gained significant adoption this decade, using AI to listen to conversations between patients and their clinicians and generate draft documentation for those encounters.
Suki CEO Punit Soni articulated the measurement problem simply: “If I captured everything word by word and transcribed it, that could be garbage from a clinical perspective. What you really want is to understand the insight.”
In other words, listening alone is not enough. A system can show extraordinary performance against a technical metric and still completely fail at what care teams and patients actually need. “As AI becomes more capable, the questions that we have to ask have to get bigger too,” Soni told me.
From Performance To Evidence
As one of the collaborative’s first projects, Suki and Regenstrief are developing what the announcement calls a “gold-standard framework” to measure the performance and impact of ambient clinical intelligence. The expected output is standardized metrics and methods spanning the clinical, operational and financial impact of ambient AI.
“Gold standard” in and of itself, can be a challenging concept to understand. Regenstrief chief research officer Dr. Joshua Vest described something closer to a common language than a certification stamp, building on Regenstrief’s prior work creating LOINC , the widely adopted international standard for identifying medical and laboratory observations. “There are distinctions between adoption, implementation and successful outcomes,” Vest told me. “There’s actually a pretty broad gap. There are a lot of jumps and steps between those aspects of IT usage.”
Those distinctions are ultimately what determine whether new solutions, AI-augmented or not, succeed in clinical practice. Is it useful? Is it usable? And, ultimately, is it used?
Even a seemingly straightforward concept, like determining whether or not there was a hallucination in an AI output, becomes complicated quickly. When I reached out to Abridge, another healthcare AI company that began with ambient documentation, for its perspective on AI measurement, its representatives pointed me to the company’s research on confabulation. The research highlights that not all errors are the same. The framework separates reasonable inferences from questionable ones, unmentioned information and contradictions, then considers the potential clinical severity of an error.
Suki Research Collaborative Chair Dr. Sudha Jayaraman summed up the fragmentation problem succinctly: “Across hospitals, across EHRs, across different ways of measuring, you end up with apples and oranges.” That is where the ambition cannot merely be to prove that AI works but to create enough consistency that we can understand the five W’s and one H: Who? What? Where? When? Why? How?
Soni explained: “Science at Suki’s job is to create general-purpose agreement on what quality should mean.” Representatives from Suki and Regenstrief will join UCSF researcher Dr. Julia Adler-Milstein at Health Datapalooza for a discussion on moving from vendor claims toward trusted real-world evidence for ambient AI. Central to the conversation is a difficult question: How do we match the pace of evidence with the pace of development and deployment? With new updates often coming on a weekly or monthly basis, we simply don’t have time to wait.
This Is Bigger Than Ambient AI
Suki is not alone in recognizing this problem. Abridge, Nabla, Microsoft and others are also building increasingly sophisticated ambient and conversational clinical intelligence tools, while a number of parallel efforts are trying to establish how these systems should be evaluated.
There are now two races happening at once: companies are racing to expand the capabilities of their products while regulators, researchers, industry groups, insurers and health systems are racing to determine how we should assess them.
However, as the technology has grown, even defining the tools with traditional categories is becoming a challenge. “To call it a scribe is somewhat underselling what the power of these technologies are,” Vest told me. “This is really a new and fundamentally different way of interacting with an electronic health record system.”
And that interaction itself is changing. “These ways of interacting are not bound by the keyboard,” Vest said. “It’s opening up to have voice and conversation be the impetus for data and action.”
Determining whether an AI solution can generate an acceptable note is one relatively well-defined problem. But what happens when that same technology begins influencing documentation, clinical reasoning, orders, medications and other workflows? Suki itself describes its platform as spanning multiple domains, including documentation, coding, revenue cycle, clinical reasoning and other workflows. Abridge has similarly moved beyond the traditional ambient-scribe model, describing a platform that includes preparation before the visit, in-visit support, documentation, coding, clinical decision support and other functions.
At this point, these artificial intelligence solutions are no longer simply documentation tools producing notes. They are increasingly becoming an operating layer within healthcare.
The Human-AI Interface Matters
I personally understand the challenges of taking on an effort of this scope. Last year, I helped lead BRIDGE , an open framework developed by Aidoc, NVIDIA and experts across healthcare organizations to address the challenges of designing, deploying, scaling and continuously monitoring clinical AI. During the year-long process it took to put together the final framework, we kept coming back to many of the same five W’s and one H mentioned above. How do you make something broad enough to capture the complexity, narrow enough to be consumable and actionable enough to actually be used?
Because if a framework falls in the cyber forest and no one uses it, does it make an impact?
Other researchers have developed frameworks such as FAIR-AI , while broader coalitions like the Coalition for Health AI have developed more general approaches to responsible AI governance and evaluation.
An AHRQ-funded conference led by Penn State radiologist Dr. Michael Bruno approached the problem through the lens of human factors and patient safety. The resulting research examined gaps across AI development and validation, healthcare-system and workforce redesign, and human-AI team augmentation.
When I spoke with Bruno last week, he distilled the challenge into one sentence:
“AI in healthcare will not improve patient safety or the quality of care by default.”
He described the future as a human-AI dyad. “If we get the human-AI interface right, we will have something that outperforms either the human or AI working alone.”
That same idea is central to the work of Raj Ratwani, PhD, director of the MedStar National Center for Human Factors Engineering in Healthcare, whose tri-level framework separates technical performance from the human-AI interaction and the surrounding healthcare system.
“Let’s think about this as shared responsibility rather than this is an AI developer issue that they have to solve, and this is a healthcare provider organization problem that they have to solve,” Ratwani told me.
That shared responsibility becomes especially important when something goes wrong.
As Jayaraman noted, “Our current root-cause-analysis processes in health systems aren’t set up to look for where the source of this is.”
The challenge we face is not simply determining whether the output of a model was correct. It is understanding how that technology interacted with the clinician, the organization, the workflow and ultimately the patient experiencing the care.
Who Gets To Define “Good”?
The collaborative brings together clinicians, researchers, developers, health systems, academic and research institutions. When asked about the participation of patients and caregivers, Jayaraman was candid in a written response to me. “We’re still in the early stages of the Suki Research Collaborative, and patient and caregiver engagement is absolutely central to where we want to go, but we are not there right now,” she said. “All of us on the team are patients or caregivers ourselves, or often both, and we bring that lived experience to every conversation. That perspective matters.”
But Jayaraman also drew an important distinction: “We know that’s not a substitute for formal partnership with the communities we’re serving.”
She added that the collaborative has already begun engaging with patient safety organizations and “will continue to deepen those relationships as the work evolves. We’re committed to doing this right.”
Maybe we can learn a lesson or two from the ambient scribe approach itself. These tools listen to patient and clinician conversations and turn them into something useful for documentation. Maybe those of us defining what “good” means should do the same.