NewsStocksStudy Finds Similar Omission Errors in OpenEvidence, OpenAI, Anthropic and Doximity Medical AI Tools

Study Finds Similar Omission Errors in OpenEvidence, OpenAI, Anthropic and Doximity Medical AI Tools

Author: Fortune Crypto·

Key Takeaways

  • OpenEvidence’s valuation rose from $1 billion in February to $12 billion in January after a series of funding rounds totaling about $700 million.
  • Doximity’s Ask AI assistant is sold through paid enterprise contracts and includes a human verification layer called PeerCheck.
  • The NOHARM benchmark tested 1,100 real clinical cases across multiple AI systems and ranked Doximity Ask first.
  • Across the systems tested, 76.6% of harmful errors were omissions rather than incorrect statements.
  • The FDA has eased restrictions on some AI clinical decision-support tools, while many states have added new healthcare AI laws requiring human oversight.
Study Finds Similar Omission Errors in OpenEvidence, OpenAI, Anthropic and Doximity Medical AI Tools

Your doctor is probably looking up your symptoms on their phone.

OpenEvidence, founded in 2021, is a free, ad-supported AI search engine for doctors that pulls answers directly from peer-reviewed medical journals at the point of care and labels how strong the evidence is. It has become something of a poster child for the medical AI boom.

The company’s valuation has climbed quickly. In February, a Sequoia-led round valued OpenEvidence at $1 billion. In the months that followed, the figure rose to $3.5 billion by July with GV and Kleiner Perkins co-leading, then to $6 billion in October, and to $12 billion this January in a round co-led by Thrive Capital and DST Global. Altogether, that amounts to roughly $700 million raised in about a year.

Doximity has taken a different route. Best known as a professional networking platform for physicians, the company now sells an AI assistant called Ask that helps doctors summarize patient notes, check drug interactions, and draft documentation. Ask is included in paid enterprise contracts with more than 150 health systems. Each answer passes through a human-review layer called PeerCheck, where physicians verify AI output against the original sources it cites. Doximity reported $145.4 million in quarterly revenue this spring, up 5% year over year.

In mid-July, a new independent benchmark called NOHARM tested both Doximity and OpenEvidence’s AI tools, along with OpenAI’s GPT-5.6 Sol and Anthropic’s Claude Fable 5. Built by researchers at Stanford, Harvard, and the ARISE network, the benchmark ran 1,100 real clinical cases through each model and collected roughly 13,000 physician annotations to measure patient harm. Doximity Ask ranked first. OpenEvidence disputed the accuracy of Doximity’s score.

“We believe that a rigorous study methodology does not allow AIs with perfect memories to ask for ‘re-tests,’” CEO Daniel Nadler told me by email. “We suspect this would not pass peer review. To be clear, the NOHARM study is not even peer-reviewed—the basic table stakes requirement in medicine for even the flimsiest medical conclusions.”

But the broader point was not which system “won.” Across every AI system tested, 76.6% of harmful errors were omissions, meaning the AI left something out rather than stating something factually wrong. That matters in clinical settings because a model can appear helpful while still failing to surface a key detail a clinician would need to check before acting on the answer.

Eric Topol, a cardiologist, Scripps Research scientist, and co-chair of Doximity’s PeerCheck program, has spent much of his career studying diagnostic error. He said that distinction matters. “Errors of omission need to be brought as close to zero as possible,” he said, adding that today’s models still create an “illusion of readiness” that has followed medical AI even as the technology improves. NOHARM did note, however, that doctors equipped with AI provided better care than doctors without it.

The regulatory backdrop makes the timing of NOHARM especially notable. In January, the FDA loosened its stance on AI-powered clinical decision-support tools, giving them more room to operate as long as doctors can independently check the AI’s reasoning. States have moved in the opposite direction, passing more than a dozen new laws in 2026 governing AI use in healthcare, with most requiring a human to sign off before any AI-assisted decision reaches a patient. Malpractice law has not kept pace with either trend, and courts are still working through whether liability falls on the doctor, the hospital, or the AI vendor when a model’s suggestion proves wrong. For hospitals and vendors, that gap helps explain why review layers such as PeerCheck are becoming part of the product story, not just a technical feature.

That legal gray zone is likely to be an early indicator of what regulators, hospital systems, and the next wave of investors will begin asking.

See you tomorrow,

Lily Mae Lazarus X: @LilyMaeLazarus Email: [email protected] Submit a deal for the Term Sheet newsletter here.

Joey Abrams curated the deals section of today’s newsletter. Subscribe here.

This story was originally featured on Fortune.com