OpenAI’s Latest Experiment Should Change How Doctors And Patients Use GenAI
One of the biggest obstacles to understanding the power of generative AI, and its potential to save lives, is the human mind itself.
Whereas people have no trouble imagining linear growth (whereby improvement progresses at a constant, arithmetical rate of 1, 2, 3, 4, etc.), exponential growth proves much harder to comprehend.
Assume, for example, that a technology doubles in capability every year. Five years later, it would be 32 times more powerful. Human intuition struggles to envision that magnitude of change. It would be like watching a car, over the same five-year period, become capable of traveling at the speed of an airplane.
When OpenAI released ChatGPT in November 2022, even ambitious projections about the pace of generative AI advancement failed to anticipate the progress that would follow.
This year, an OpenAI experiment demonstrated how quickly those capabilities are expanding. The results made front-page news, shocking industry insiders and surprising even the researchers studying them.
Although regulators and elected officials must continue grappling with the existential threats posed by increasingly powerful AI systems, the experiment also revealed opportunities to advance American medicine. Doctors and patients can now use generative AI’s expanding capabilities to transform medical care, save lives and increase affordability.
What Happened Inside OpenAI’s Experiment
In July, OpenAI set out to test whether increasingly capable models could identify and exploit software vulnerabilities. As part of the cybersecurity benchmark , researchers ran thousands of AI agents, designing the experiment so that each agent would work independently in a walled-off environment, commonly referred to as a sandbox.
Instead, roughly 1,200 agents found a way to collaborate, communicating through a message board they created themselves. Over several days, they exchanged more than 70,000 messages and files.
And they didn’t merely “chat.” AI agents divided work in teams, shared discoveries and tackled highly complex projects. Some even sacrificed the success of their individual runs to generate information that could benefit what the agents called the collective .
Next, roughly 700 agents hacked Hugging Face , one of the world’s largest platforms for AI models, accessing its datasets and software. No human had instructed the agents to coordinate the attack, and the researchers running the experiment were unaware it was happening at the time.
But these same worrisome capabilities also create new possibilities for medicine.
Doctors could deploy coordinated GenAI agents on behalf of patients, for example, to monitor a postoperative wound, track signs that chronic heart failure is worsening or follow the progress of people hospitalized at home. At the same time, patients can and increasingly will use GenAI on their own to evaluate symptoms, manage chronic conditions and obtain medical guidance when clinicians are unavailable. Together, those two paths create massive clinical opportunities that did not exist even a couple of years ago.
Medicine Is Thinking Too Small
Physicians resisted GenAI in its earliest days because of concerns over unreliability and hallucinations. But acceptance has grown rapidly. A recent American Medical Association survey found that 81% of physicians now use AI professionally, more than twice the percentage in 2023.
Doctors are using GenAI to summarize charts, draft patient communications and reduce documentation burdens. According to the same AMA survey, a rapidly increasing percentage of physicians now rely on ambient AI scribes, which listen to physician-patient conversations and generate clinical notes for the electronic medical record.
Medical conferences devote extensive attention to ways that GenAI can assist physicians to research complex topics, identify uncommon diagnoses and complete administrative tasks.
But physicians remain much more cautious when AI begins to perform tasks involving clinical judgment or direct patient guidance. Almost no doctors encourage patients to consult tools like ChatGPT, Gemini or Claude when they have medical questions. Minimal assistance is given to teaching patients to use these tools when clinicians are unavailable, offices are closed or medical questions arise between visits.
Supporting this dichotomy, the AMA adopted policies earlier this summer asserting that AI should remain under physician oversight rather than function as an autonomous clinical decision-maker.
The instinct is understandable. Technologies that threaten professional authority, income or employment naturally encounter resistance. But attempts to restrain or maintain control over GenAI will not stop its development or prevent patients from using it. Instead, such restrictive efforts will likely cause medicine to miss opportunities to improve clinical outcomes, increase access and save lives.
Patients are already moving ahead. More than 300 million people worldwide ask health-related questions of ChatGPT each week. They use it to evaluate symptoms, interpret laboratory results and explore treatment options.
How Medicine Can Start To Think Bigger
Imagine if every patient had 700 AI agents working together to help manage their chronic conditions and evaluate new difficult medical problems.
As in the OpenAI experiment, those agents could draw on enormous bodies of information, divide responsibilities, share findings and solve problems collectively. In medicine, that could mean combining the breadth of medical knowledge available through large language models with a patient’s complete medical history, test results, medications and data from wearable devices. In conjunction with the person’s doctors, the agents could continuously analyze the data and identify clinical patterns that might be missed or find relevant peer-reviewed medical articles that might be overlooked.
That kind of capability has the potential to transform medical care. Here are three ways this kind of capability could improve care:
- Healthcare could have a 24/7 front door. Patients today can’t be certain whether a new symptom requires an emergency room visit, an appointment next month or something in between. GenAI could provide immediate guidance to families, with clinicians and telehealth services available when the technology identifies a problem requiring human intervention.
- Every hospitalized patient could have access to personalized medical guidance. Hospitalizations leave patients and families surrounded by test results, medications, consultations and unfamiliar terminology. GenAI agents could explain what is happening, identify conflicting recommendations among fragmented specialists and help patients understand what should happen next.
- Chronic disease management could become continuous rather than episodic. Patients with hypertension, diabetes and other chronic conditions are commonly evaluated every three to four months in a doctor’s office, even when control is poor. This leads to preventable heart attacks, strokes and kidney failures. GenAI could continuously analyze data from home monitors and wearable devices, identify when treatment is not working as expected and recommend medication adjustments — months before the next scheduled office visit.
Denial Won’t Preserve Physician Control
The Hugging Face incident demonstrates that today’s GenAI is exponentially more capable than the technology released four years ago. More importantly for medicine, it challenges the assumption that these systems will remain passive tools that operate only under human direction.
This is the wake-up call medicine needs. Physician organizations can keep insisting that AI remain subordinate to clinicians. Medical schools can try to limit how much future doctors rely on it. Regulators can impose requirements on proprietary GenAI tools. Nursing organizations can negotiate restrictions on workplace deployment.
But none of those efforts will prevent patients from turning to large language models when they develop new symptoms, need help managing chronic diseases or are unwilling to wait weeks for the next available doctor’s appointment.
The assumption that GenAI must always remain subservient to clinicians is outdated. Patients will use these tools to obtain medical information, analyze symptoms and manage their health whether the medical profession encourages them to or not. A better path has arrived: Dedicated physicians, empowered patients and increasingly capable GenAI systems working together to achieve far better medical outcomes than any one of the three could achieve alone. That’s healthcare’s future.