Fred Hutch AI Club - December 2026 Talk
The number of laws, regulations, standards, and frameworks that apply to an organization deploying AI passed 600 in September 2026, up from 250+ in March. More than 200 of them come from US state legislatures. The 2025 and 2026 laws differ from earlier ones: they specify what a system must do in a particular moment, where earlier laws asked only for disclosure AI use. A chatbot that hears suicidal ideation has to interrupt and escalate. A scheduling assistant that hears "chest pain" has to route to emergency care. An assistant asked "are you a doctor?" has to say no. A decision support tool that uses race or sex as an input must be rechecked for discrimination after every vendor update. None of these is a knowledge question, so a model can lead the licensing exam benchmarks while failing all four.
The first half of this talk covers nine questions an institution must answer about its AI testing and monitoring before and after a clinical AI system goes live, including which laws a benchmark covers, whether a given score passes a deployment bar, how a model performs for the patients you serve, and who reviewed the test cases. It shows that the eight medical safety benchmarks published since 2024 answer almost none of them, and introduces PacificMed, a suite of 70+ medical AI benchmarks for bias, fairness, safety, and reliability.
The second half focuses on AI agents, where clinical AI moves from answering to acting: working a prior authorization, running a specialist consultation across dozens of tool calls, or auditing a chart against a standard. Three agentic benchmarks now in the open-source MedHELM project show how large the gap is: agents complete 83% of individual subtasks but 36% of workflows end to end, the best of 13 agents succeeds on 46% of long-horizon clinical tasks, and all 16 systems tested committed to an answer when the patient record didn't support one. Governing an agent means scoring what it did and monitoring it after deployment, and the talk ends with what you can put in place now, including free access to every suite discussed for academic research.
Free