Skip to content
For CFO
Executive Brief

Clinical AI Flaw Affects OpenAI, Anthropic, Doximity

NOHARM benchmark reveals dangerous omissions in medical AI-not hallucinations-creating unpriced CFO liability

brown wooden blocks on white surface

The Omission Problem Is a Balance Sheet Problem

A physician receives a clinical AI response. It looks complete. It reads authoritatively. It skips a drug interaction that matters. No error flag fires. No confidence score drops. The output just - doesn't say the thing.

That is the failure mode a new benchmark study surfaced this week, testing clinical AI tools from OpenEvidence, OpenAI, Anthropic, and Doximity. The finding, reported by Fortune on July 29, is that the dominant safety failure across these platforms isn't hallucination - models generating false information - but omission: models generating plausible, incomplete information that clinicians act on as though it were whole.

For healthcare CFOs, the instinct will be to forward this to the CMO. That instinct is expensive.

What Your Vendor Contract Actually Says

Omission failures don't generate audit trails. There is no logged event, no flagged output, no artifact that says "the model skipped this step." When harm results from what the AI didn't say, the liability chain runs through your vendor agreement - specifically through whatever indemnification language your legal team negotiated, almost certainly before omission risk had been benchmarked at scale.

The standard carve-out in clinical AI procurement agreements transfers liability at the moment of clinical use. Language referencing "clinical judgment," "physician discretion," or "final determination" is doing one job: moving the exposure from the vendor's balance sheet to yours. Vendors are also inserting output-related exclusions that disclaim responsibility for outputs generated in response to user prompts - particularly where the health system has customized or fine-tuned the model. Most of these agreements were signed when hallucination was the primary concern. Omission is a structurally different risk, and it is largely unaddressed in the contracts currently sitting in your legal files.

This is not a clinical governance problem that happens to have financial consequences. It is a procurement problem with a clinical origin.

Where the Insurance Gap Lives

The coverage question is more unsettled than most CFOs realize. ISO's generative AI exclusion endorsements - released in early 2026 - give commercial general liability carriers the option to exclude bodily injury and property damage arising from AI use. W.R. Berkley introduced what it describes as an absolute AI exclusion across D&O, E&O, and fiduciary liability products. The exclusion language is broad: "any actual or alleged use, deployment, or development of Artificial Intelligence." That definition doesn't require AI to be the sole cause of harm.

The practical consequence: a CFO who approved clinical AI deployment without notifying carriers may face a coverage dispute on material non-disclosure grounds at the next renewal. That is not a clinical problem. That is a balance sheet problem, and it is one the NOHARM benchmark - by quantifying how often complete-looking outputs are actually incomplete - makes newly legible to underwriters.

Ask your malpractice and D&O carriers one question before renewal: does current coverage explicitly address AI omission events, cases where harm resulted from information the AI failed to surface rather than information it stated incorrectly? If the answer requires a pause, you have a gap that predates this week's study and will outlast it.


The jurisdictional picture adds a layer that a U.S.-centric read tends to flatten. Illinois signed the Artificial Intelligence Safety Measures Act this month, mandating annual independent third-party audits for large AI developers beginning in 2028. Illinois is now the third state - after California and New York - to impose transparency and safety obligations on frontier AI vendors. Because these models don't stop at state lines, the audit disclosures those laws require will have national effects. Health systems operating across multiple states should expect that omission-rate data, once surfaced through mandated audits, will become discoverable in litigation regardless of where the incident occurred. The compliance clock is not uniform, but the liability exposure is.

The Three Moves That Matter Now

This is a tighten-controls-now, renegotiate-on-renewal situation. Clinical AI is embedded deeply enough in large health systems that wholesale removal creates its own operational risk. The immediate decisions are narrower and more tractable.

Pull every active clinical AI vendor agreement and flag indemnification clauses containing "clinical judgment," "physician discretion," or "final determination." Those are your liability transfer points. Then ask vendors for omission-rate data - not accuracy metrics, omission metrics. These are not the same number, and a vendor who responds with only an accuracy scorecard is answering a different question.

Brief your malpractice and D&O carriers before your next renewal on the benchmark findings. Voluntary disclosure is cheaper than a coverage dispute after an incident. Document in writing that governance controls for clinical AI omission risk have been reviewed - not because documentation solves the problem, but because undocumented awareness is the worst of all positions when litigation arrives 18 months from now.

Seventy-five percent of U.S. health systems have deployed at least one AI solution. Eighteen percent have mature governance frameworks. The distance between those two numbers is where this week's study lands.

Source: Fortune, July 29, 2026. Authority: secondary/trade. Claims from supplemental research context are directional and not independently verified.

0
Read0%
Action Plan

1. Pull every active clinical AI vendor agreement. Flag any indemnification clause containing the phrase 'clinical judgment,' 'physician discretion,' or 'final determination.' These are your liability transfer points. 2. Ask your malpractice and D&O carriers one question: 'Does our current coverage explicitly address AI omission events - cases where harm resulted from information the AI failed to surface, not information it stated incorrectly?' If they pause, you have a gap. 3. Require vendors to provide NOHARM benchmark results or equivalent omission-rate data for their specific product version. If they cite only accuracy metrics (what the model gets right), push for omission metrics (what the model fails to include). These are not the same number. 4. Audit your self-insurance reserve assumptions. If your actuarial model for AI-related clinical liability was built on published hallucination rates, it needs to be rerun against omission-frequency data. 5. Before the next board meeting, document in writing that governance controls for clinical AI omission risk have been reviewed. This is your D&O protection - not because it solves the problem, but because undocumented awareness is the worst of all positions.

The CFO who reads this as a clinical story and forwards it to the CMO without a finance action item has made the most expensive mistake. Clinical AI liability that surfaces in litigation 18-36 months from now will be traced back to procurement decisions and governance sign-offs that happened on your watch. The second failure mode: assuming your vendor's indemnification clause covers you. It almost certainly doesn't - 'clinical judgment' carve-outs exist precisely to shift liability back to the institution at the moment of use. The third failure mode is the insurance gap: a CFO who approved clinical AI deployment without notifying carriers may face a coverage dispute on the grounds of material non-disclosure. That's not a clinical problem. That's a balance sheet problem.

Key Takeaways
The article, blog post, or content outline you'd like me to work from?
Any specific themes or topics you want the quotes to emphasize?
Originally Reported ByNaN/10 Minimally Sourced
F
Fortune
fortune.com/2026/07/29/new-medical-ai-study-same-flaw-openevidence-openai-anthropic-doximity
Supporting Sources
C
Censinet / Eliciting Insights 2026 AI Adoption Survey (March 2026)
censinet.com/perspectives/governance-gap-healthcare-ai-leaders-realize
C
Censinet (2026)
censinet.com/perspectives/ai-governance-questions-healthcare-boards-ignore
N
National Law Review
natlawreview.com/article/texas-ags-landmark-ai-settlement-wake-call-health-tech-ai-companies
H
Healthcare Huddle / NOHARM Study Analysis
healthcarehuddle.com/p/clinical-ai-safety-what-the-noharm-study-reveals
G
Greenberg Traurig LLP
gtlaw.com/en/insights/2026/7/illinois-enacts-artificial-intelligence-safety-measures-act
T
ToolixLab / NOHARM Benchmark Summary
toolixlab.com/blog/ai-medical-diagnosis-accuracy-statistics-2026
D
DLA Piper
dlapiper.com/en-us/insights/publications/2026/07/the-artificial-intelligence-safety-measures-act-illinois
C
Contracting with AI Vendors: A Practical Guide for Lawyers (2026)
storage.ghost.io/c/44/95/449506ca-034e-480f-9725-fcde08ef1cc1/content/files/2026/03/Contracting-with-AI-Vendors.pdf
A
ArentFox Schiff / Health Care Counsel Blog
afslaw.com/perspectives/health-care-counsel-blog/ai-service-agreements-health-care-indemnification-clauses
M
Managed Healthcare Executive
managedhealthcareexecutive.com/view/faq-how-ai-is-changing-managed-care-in-2026-and-what-leaders-need-to-watch
B
blueBriX Health
bluebrix.health/articles/ai-reset-a-new-era-for-healthcare-policy
C
Censinet
censinet.com/perspectives/2026-defining-year-ai-governance-healthcare
Affected Workflows
Healthcare AIMedical SafetyAI BenchmarkingClinical RiskVenture FundingFrontier Signal Lane
Research Sources12
  1. As of early 2026, 75% of U.S. health systems have deployed at least one AI solution - up from 59% in 2025 - yet only 18% have mature AI governance frameworks in place, meaning the vast majority of production AI deployments are operating without formal oversight structures. Censinet / Eliciting Insights 2026 AI Adoption Survey (March 2026)
  2. Only 10-15% of large health system boards had formal AI oversight structures in place as of early 2026, and fewer than half of hospitals test deployed AI tools for bias - indicating that board-level disclosure of AI safety gaps, including omission errors, is the exception rather than the rule. Censinet (2026)
  3. The Pieces Technologies settlement requires five years of compliance, though the company may request rescission after one year based on compliance, changes in the regulatory landscape, or changes in generative AI benchmarks and industry standards - making evolving benchmarks like NOHARM a direct contractual trigger for modification. National Law Review
  4. The NOHARM benchmark (Stanford/Harvard, published January 2026) found that the dominant safety failure mode in clinical AI is omission - models are more likely to leave out key management steps than to recommend something overtly dangerous, meaning a plausible-sounding but incomplete answer can slip past clinical review undetected. Healthcare Huddle / NOHARM Study Analysis
  5. On July 6, 2026, Illinois Governor JB Pritzker signed the Artificial Intelligence Safety Measures Act, which mandates annual independent third-party audits for large frontier AI developers (those with annual gross revenues exceeding $500 million), with audit obligations beginning January 1, 2028, and 72-hour incident reporting requirements for AI safety events. Greenberg Traurig LLP
  6. The NOHARM benchmark evaluated 31 leading LLMs - including GPT-5, o3/o4-mini, Gemini, Claude Sonnet 4.5, Llama 4, and two specialized medical platforms - against 100 real primary-care-to-specialist consultation cases, rated by 29 board-certified physicians across 12,747 individual clinical-action ratings, establishing the first standardized safety leaderboard against which deployed clinical AI systems can now be measured. ToolixLab / NOHARM Benchmark Summary
  7. Illinois is now the third state - after California and New York - to impose transparency and safety obligations on large AI developers; because AI models are not geographically confined, the resulting public disclosures have national effects, and legal analysts argue the three states together represent an emerging de facto national standard in the absence of federal legislation. DLA Piper
  8. A negotiation tactic gaining traction in 2026 is tying the vendor's liability cap to its actual insurance coverage amount rather than to subscription fees - for example, if a vendor carries $5 million in E&O coverage, the liability cap for indemnifiable claims should be $5 million, not the $60,000 in annual fees vendors typically propose. Contracting with AI Vendors: A Practical Guide for Lawyers (2026)
  9. Vendors are inserting "output-related exclusions" into clinical AI contracts, disclaiming responsibility for outputs generated in response to user prompts or inputs - particularly in cases where the healthcare provider modifies, customizes, or fine-tunes the model - effectively shifting omission-related liability back to the provider. ArentFox Schiff / Health Care Counsel Blog
  10. Many health system leaders are accepting vendor contract terms as-is rather than actively negotiating them. A common and dangerous misconception is that a vendor's contract transfers clinical liability - when in fact the health system still holds legal responsibility for patient care outcomes. Managed Healthcare Executive
  11. In response to vendor-favorable contract terms, leading health systems are now pushing back by rewriting agreements to include use-case-specific indemnity covering regulatory fines and third-party claims, data breach SLAs with clear notification timelines, and regulatory cooperation clauses obligating vendors to support audits and adapt to new federal and state rules. blueBriX Health
  12. The 2026 HIPAA Security Rule overhaul now explicitly covers AI training data and prediction models, removing previously "addressable" safeguards and creating new mandatory compliance obligations - meaning vendor contracts that fail to address AI-specific data use are now directly out of compliance with federal law. Censinet

Responses

(0)

Responses0



















0