Blog & news

AI Governance in Healthcare: A Physician's Framework for Boards and Executives

Every hospital in America already knows how to govern something that makes clinical decisions, carries real risk, performs differently at different institutions, and occasionally has to be stopped. It is called a doctor.

DR GPT healthcare AI governance article image by Harvey Castro MD

Every hospital in America already knows how to govern something that makes clinical decisions, carries real risk, performs differently at different institutions, and occasionally has to be stopped. It is called a doctor.

You credential them. You grant privileges within a defined scope. You supervise until competence is proven. You review performance on a schedule. When the performance falls apart, you act.

That machinery has existed for a century, and it works. Most AI governance efforts I see fail because someone tried to invent a parallel structure from scratch, staffed it with a committee that meets quarterly, and produced a policy document nobody in the building has read.

Stop inventing. Credential the model like a clinician.

Why boards are running out of time to decide this

The oversight environment shifted twice in the past year.

In September 2025, The Joint Commission and the Coalition for Health AI published joint guidance on the responsible use of AI in healthcare. On June 2, 2026, The Joint Commission launched a voluntary Responsible Use of AI in Healthcare certification structured around five domains: governance, data management, risk and bias reduction, monitoring and validation of safety performance and effectiveness, and transparency with education and training. It certifies organizational practices rather than individual products, and any healthcare organization may apply whether or not it holds Joint Commission accreditation (Fierce Healthcare).

State legislatures moved in the same direction with more teeth. Indiana's HB 1271 took effect July 1, 2026 and bars insurers from using AI as the sole basis for downgrading a claim without professional review. Maryland's HB 1563 took effect June 1, 2026 and requires quarterly reporting on whether AI assisted an adverse decision. Alabama's SB 63 lands October 1, 2026. Utah's SB 319 and Georgia's SB 544 arrive January 1, 2027, both requiring that a licensed professional, not a system, issue an adverse determination. Delaware's HB 191 goes further and bars AI systems from holding professional licensure or using protected titles (Holland & Knight).

Now set that against what hospitals can actually document. Black Book Research surveyed 182 hospital leaders in late 2025 and found 29 percent had implemented and enforced policies covering AI model inventory, lineage, and sign-offs. Twenty-two percent were highly confident they could produce a complete AI audit trail within 30 days (Becker's Hospital Review).

Roughly seven in ten systems cannot show a regulator what they are running.

The five-stage framework

Stage one: credentialing

No model touches a patient until someone can answer four questions in writing.

Where did it come from, and who built it. What population was it trained on. How does it perform on your patients, measured locally before go-live, not on a vendor slide. Who signed off.

That last one is the piece boards skip. Credentialing is an act of accountability with a name attached. If no clinician's name appears next to the approval, no one owns the outcome.

Local validation is the expensive part, and it is the part that pays. The Epic Sepsis Model reached hundreds of US hospitals before independent evaluation. When Michigan Medicine studied it across 38,455 hospitalizations, the model showed an area under the curve of 0.63, missed 67 percent of sepsis cases, and generated alerts on 18 percent of all hospitalized patients. According to PubMed, that is Wong A, Otles E, Donnelly JP, et al., JAMA Internal Medicine, 2021 (DOI).

Nobody credentialed that model. It arrived with the software.

Stage two: privileging

A credentialed clinician does not get to do everything. A cardiologist does not perform craniotomies.

Write the scope. Which patients, which settings, which decisions, and which decisions are explicitly out of bounds. A documentation tool has different privileges than a triage model. A triage model has different privileges than anything touching a treatment recommendation.

Then name the supervising role. Not a department. A role. Someone gets paged when this misbehaves.

Stage three: supervision

Human oversight is meaningless if the human cannot afford to disagree.

Test yours honestly. Can a nurse override the model without filing paperwork. Does overriding it cost her time she does not have. Does the productivity dashboard punish her for it. If a clinician needs courage to say no to your software, you do not have human oversight. You have a rubber stamp with a pulse.

I have watched good clinicians defer to a bad score because the workflow made deference cheaper than judgment. That is a design failure, and it belongs on the board's risk register.

Stage four: ongoing evaluation

Medical staff bylaws already require periodic performance review. Apply the same discipline.

Four metrics, reviewed on a fixed cadence. Discrimination and calibration on current patients. Alert volume and override rate. Performance stratified by race, language, payer, and site, because a model that works on your commercial population and fails on your Medicaid population is a quality problem before it is a legal one. Drift against the baseline you set at credentialing.

Set the threshold that triggers action before you need it. A number written in advance is a governance decision. A number debated after an incident is a negotiation.

Stage five: peer review and revocation

Two things have to exist. A path for a frontline clinician to report that a model harmed or nearly harmed a patient, routed to the same committee that handles other safety events. And a documented authority to shut the model off.

Most AI policies I read have no off switch, no named person who can pull it, and no precedent for having done so. Retire one tool publicly in the first year. The organization learns more from that than from any policy binder.

Nine questions for the board

Directors do not need to understand model architecture. They need to ask questions management cannot answer with a slide.

  1. How many AI tools are running in this organization today, and who maintains that inventory?
  2. Which of them touch a clinical decision?
  3. For each, what was the local validation result before go-live?
  4. Who signed the approval, by name?
  5. What is our alert override rate, and is it rising?
  6. Have we measured performance differences across our patient populations?
  7. What is our documented process for turning a model off, and have we ever used it?
  8. If a state regulator asked for an audit trail tomorrow, how many days would it take?
  9. What percentage of our IT and quality budget funds AI oversight?

On that last one, the median across surveyed hospitals is 4.2 percent, with large systems at 6.8 percent and small hospitals at 2.3 percent. If your number is well under the median, your governance is aspirational.

If you run a small or rural hospital

The framework does not require a data science department. It requires a named physician, a spreadsheet, and a standing agenda item.

Scale matters, and the gap is already visible. Among US hospitals on Epic, 62.6 percent had adopted ambient AI documentation as of June 2025, but adoption ran at 64.7 percent in metropolitan hospitals against 54.3 percent in nonmetropolitan ones, and 70.2 percent at nonprofits against 28.8 percent at for-profit facilities (Yang F, Graetz I, American Journal of Managed Care, 2026). Governance capacity is thinner in exactly the same places. Only 15 percent of small hospitals in the Black Book survey felt confident producing an audit trail within 30 days.

Two practical moves. Join a regional collaborative or a purchasing coalition so validation costs get shared across systems facing the same patient mix. And write the validation requirement into the vendor contract, because a small hospital has the most bargaining power at signature and almost none afterward.

Governance is not the brake

The objection I hear most often is that oversight slows innovation. My experience says the opposite.

Clinicians adopt tools they trust. Trust comes from knowing the thing was checked, that someone competent is watching it, and that saying no carries no penalty. The systems that skipped governance are the ones now spending a year cleaning up alert fatigue, unwinding a vendor relationship, or explaining to a board why nobody can produce a list of what is running.

Credential the model. Privilege it. Supervise it. Review it. Be willing to revoke it.

Hospitals have run that playbook on humans for a hundred years. The machines do not need a new one.

Harvey Castro, MD, MBA is a board-certified emergency physician, 5x TEDx speaker, and author of more than 30 books on AI and healthcare, including AI in Emergency Medicine (Wiley). He serves on Singapore's Ministry of Health Regulatory Advisory Panel and advises the Texas Medical Association's Committee on Health Information Technology.

Related DR GPT™ reading

More from DR GPT™: who DR GPT™ is, healthcare AI keynotes, speaking, TEDx talks, books, and the media kit.

Book Harvey Castro, MD, MBA, known as DR GPT™, for your board, leadership retreat, or healthcare AI event. Start the conversation.