Skip to Main Content
Quick Read

How Many LLMs Does it Take to Reason Through a Decision?

4 Minute Read

Making clinical decisions is often a team effort. Patients seeking cancer treatment, for instance, may have a surgeon, oncologist, pathologist, and radiologist all involved in their care. These multidisciplinary care teams, in which specialists with diverse perspectives come together, can diagnose a condition or determine the best course of treatment more effectively.

New AI models could dramatically speed up this process.

One new approach comprises multiple specialized large language models (LLMs) that talk to one another and reach a consensus similar to how a committee of doctors would. In 2023, a team of Yale School of Medicine (YSM) researchers led by Mark Gerstein, PhD, Albert L Williams Professor of Biomedical Informatics, introduced the first of this kind of tool, a model they named MedAgents. In MedAgents, LLMs roleplay as different specialists that participate in multiple rounds of discussion to answer various medical questions.

Scientists are also creating other highly trained platforms in which this complex multi-step reasoning occurs within a single LLM. Both platforms have benefits and drawbacks. Now, Gerstein’s team has introduced MedicalAgentsBench, a new tool for comparing different versions of these new AI models. The team published their findings on July 8 in Cell Patterns.

“If we want to build more truly helpful medical agents, we need better ways of evaluation,” says Yanjun "Daniel" Shao, a master’s student at YSM and the study’s first author.

“If we want to build more truly helpful medical agents, we need better ways of evaluation."

Yanjun "Daniel" Shao
Master's Student, Yale School of Medicine

“With MedicalAgentsBench, we’re trying to evaluate how these externalized platforms, which require a lot of interaction between the LLM agents to make a decision, compare to a more complex model that internalizes all the discussion in its training," adds Gerstein, who is the study’s principal investigator. “Is it better to have something that acts like a committee of experts, or is it better to have this incredibly well-trained super oracle that was trained with the knowledge of all of the experts?”

Externalized agents or internalized reasoning?

Researchers in the artificial intelligence space have been debating which type of model is better, Gerstein says. Internalized reasoning models, such as OpenAI's o1 or DeepSeek-R1, are trained with reinforcement learning to "think" through a problem step-by-step inside a single model before answering. Externalized agents, by contrast, achieve similar deliberation by having multiple LLMs talk to each other at inference time.

"It's the difference between training one model to think hard, and orchestrating several models to think together," says Xiangru 'Robert' Tang, who received his PhD in computer science from Yale University in 2026 as a member of Gerstein’s laboratory and is the study’s co-first author. "We didn't know which would be better on hard medical problems—so we built a benchmark to find out."

Existing tools evaluate the performance of these AI models by asking them questions similar to those seen in medical examinations like the MCAT. But newer models are answering these questions with such a high accuracy that it is difficult to distinguish meaningful differences in performance. Another limitation is that they often fail to distinguish reasoning capabilities from simple memorization.

“As the training of models gets more and more complicated and uses more data, we have this problem where the system just memorizes the answers,” Gerstein says. “We wanted to come up with a more sophisticated benchmark that dealt with this ceiling effect.”

To develop MedicalAgentsBench, Gerstein’s team used eight different established medical datasets to create over 800 complex medical questions. The questions challenged the models to use multi-step thinking rather than rely on memorization.

“Is it better to have something that acts like a committee of experts, or is it better to have this incredibly well-trained super oracle that was trained with the knowledge of all of the experts?”

Mark Gerstein, PhD
Albert L Williams Professor of Biomedical Informatics and Professor of Molecular Biophysics & Biochemistry, of Computer Science, and of Statistics & Data Science

Using this framework, the researchers did not find that one type of reasoning was superior to the other, but rather that they can complement one another. For example, adding externalized agents to an internalized reasoning model further improves its performance.

Due to privacy concerns, healthcare systems cannot use publicly available models to aid in clinical decision making. As they develop their own models, MedicalAgentsBench can help institutions evaluate for effectiveness and cost, the team says.

“We’re trying to find a better recipe to help people build their own clinical support system that can answer questions from humans,” Shao says.

Article outro

Author

Isabella Backman
Senior Science Writer/Editor, YSM/YM

The research reported in this news article was supported by the National Institutes of Health (award R01DA063148) and Yale University. The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.

Media Contact

For media inquiries, please contact us.

Explore More

Featured in this article