The Biostatistics Program at Fred Hutch hosts a variety of seminar series, symposia, and other events. These events are open to Fred Hutch faculty, staff, and external colleagues. When available, we will post recordings on our YouTube channel. 

To subscribe to the mailing list of these events, please send an empty email to biostatseminars-join@lists.fhcrc.org and follow the subsequent automatic prompts to confirm your address. 


Upcoming Events

There are no events scheduled at this time.

Your search for "" did not match any of our events. Please try another event term or combination of terms and filters.

Past Seminars & Events

2026-09-08
Anru Zhang, Duke University
Making EHR Ready for Analysis and AI: Statistics at the Data Frontier
Electronic health records (EHR) are increasingly central to clinical research and AI, but they are not simply “big data.” They are irregular, heterogeneous, imbalanced, privacy-sensitive, and sometimes internally inconsistent. As a result, reliable statistical learning and inference often requires first solving problems of data readiness. In this talk, we will discuss statistical and computational methods our group has developed to make longitudinal EHR more comparable, learnable, shareable, and trustworthy. Topics include timeline registration for aligning patients along latent disease progression; synthetic data augmentation for learning from imbalanced EHR; privacy-preserving synthetic EHR generation; and large-language-model-assisted detection of clinical documentation inconsistencies. These examples motivate a broader perspective in which preparing EHR for analysis and AI is itself a statistical problem, and a necessary foundation for trustworthy clinical and generative AI.

2026-07-30
Jan Hannig, UNC
Fiducial Generative Models
While generalized fiducial inference (GFI) and its variants have yielded many theoretical and practical results to parametric inference and uncertainty quantification, applying it to generative models remains challenging. We identify three key issues misspecification, metric choices, and over-parameterization hinder the direct application of the GFI to generative models. In this paper, we propose a novel method based on the framework of generalized fiducial inference, designed to construct distributional estimates over the parameter space given observed data, while also enabling uncertainty quantification for generative models. We employ a truncation-based approach and further provide a theoretical analysis of its behavior under varying truncation parameters. Both theoretical results and empirical evidence suggest that, with an appropriately chosen truncation parameter, the truncated distribution derived from generalized fiducial inference achieves valid coverage of the true parameter and leads to improved generalization performance.

2026-07-01
Jiayin Zheng,  U Penn
Statistical Considerations in Developing a Web-Based Personalized Colorectal Cancer Risk Tool
Colorectal cancer (CRC) is one of the most common cancers globally and accurate risk prediction of developing CRC is pivotal in enabling earlier screening and prevention. Polygenic risk scores (PRS) hold promise for identifying high-risk individuals; when combined with lifestyle factors, they substantially improve prediction accuracy compared with models based on lifestyle factors alone. However, few clinical tools currently exist that facilitate this integrated, PRS-enhanced risk assessment. To bridge this gap, we developed MyGeneRisk Colon, a publicly accessible web portal that delivers individualized CRC risk prediction by incorporating genetic, demographic, family history, and lifestyle factors. This talk will highlight key statistical considerations in the development of this risk prediction tool. I will also discuss relevant challenges and opportunities that motivate the development of novel methodologies and could benefit from their application.

2026-06-29
Jean Feng, UCSF
LLMs Judging LLMs and Humans Judging LLMs: Who Can We Trust?
Recording Available
Evaluating the outputs of generative AI systems is extremely time-intensive, and two increasingly common practices have emerged: (i) using LLMs to judge other LLMs, and (ii) having humans adaptively select a small set of outputs to evaluate. The widespread adoption of these approaches suggests they must be practically useful—but what statistical guarantees, if any, do they actually provide? In this talk, we investigate when these practices are valid, when they break down, and how relatively simple modifications can yield rigorous statistical inference.

2026-06-03
Kevin Lin, University of Washington
Disentangling Local Environments and Concurrent Cellular Programs in Complex Diseases
Analyzing high-dimensional single-cell profiles to map cellular dynamics in complex diseases presents two core statistical challenges that universal to many biological systems and pathologies: 1) separating localized microenvironmental/region-specific variations from global, systemic disease progression, and 2) disentangling multi-faceted, overlapping cellular programs to isolate distinct, disease-relevant trajectories. This talk presents two complementary statistical workflows designed to tackle these complexities. Using single-nuclei RNA-sequencing (snRNA-seq) of microglia in Alzheimer's disease (AD) as a concrete focus, we demonstrate how these methodologies can resolve intricate cellular architectures. 

2026-06-02
Luo Xiao, North Carolina State University
Functional Joint Models in Longitudinal Studies
We study functional joint models in longitudinal studies in which longitudinal outcomes are modeled as sparse functional data and are jointly modeled with an event of interest, such as diagnosis of certain disease or mortality. We first provide an overview of functional data methods for modeling longitudinal data. Then we discuss a joint model of survival data and multivariate sparse functional data for measuring progression and diagnosis of Alzheimer’s Disease (AD).  A flexible and computationally feasible model estimation strategy is employed. We shall also discuss the extension of this model to a multi-cohort AD study. Additional related works will also be discussed.

2026-05-27
Yanyao Yi, Eli Lilly & Co
Covariate-Adjusted Log-Rank Test in Randomized Clinical Trials
The log-rank test has long been the standard for assessing treatment effects on time-to-event endpoints in randomized clinical trials. Covariate adjustment for efficiency gain has gained traction following the FDA's 2023 guidance, which emphasizes that, in RCT settings, covariate-adjusted methods should yield valid inference under approximately the same minimal statistical assumptions required for unadjusted estimation—neither altering the estimand nor introducing modeling assumptions. This presentation introduces a covariate-adjusted log-rank test with a simple, closed-form expression and a guaranteed efficiency gain over the unadjusted test. The method is universally appliable across simple randomization and all commonly used covariate-adaptive randomization schemes, including stratified permuted-block randomization. A stratified extension is also available in Ye, Shao and Yi (2024).

2026-05-20
Anil Palepu, Google
Evaluating Conversational AI for Medical Diagnosis and Management
Recording Available
Special Joint Seminar with University of Washington Department of Biostatistics. The medical interview has been termed “the most powerful, sensitive, and versatile instrument available to the physician.” While Large Language Models (LLMs) have achieved expert-level scores on medical board examinations, these static benchmarks fail to capture the essence of clinical practice: the ability to intelligently and compassionately acquire information under conditions of uncertainty. To bridge this gap, we must evaluate AI through frameworks that mirror the complexity of human practice—most notably the Objective Structured Clinical Examination (OSCE), a validated gold standard for assessing clinical competence in medical trainees. 

2025-05-05
Susanne Rafelsi, Allen Institute
Toward a Holistic and Dynamic Stem Cell State Landscape
Part of our AI for Biomedical Data Science Across Scales Statistical AI Symposium. Establishing a conceptual framework for holistic cell states and state transitions is an essential step towards building a multiscale understanding of cell biology. At the Allen Institute for Cell Science, we are working to define and map stem cell states in the context of organizational and transitional dynamics across diverse conditions. Over the past decade, we have developed experimental and analytical pipelines and tools to quantify and computationally represent the intracellular organization of cells from 3D microscopy images and applied these to a variety of biological questions. Building upon this foundation, our new CellScapes initiative aims to understand how cells organize themselves across scales to form complex cell communities and tissues and use this understanding to build predictable and programmable tissue systems. The Institute has applied AI in various ways including the development of 3D bioimage analysis approaches, interpretable representation learning of intracellular structures, and extracting cell state transition dynamics from timelapse movies. At the core of our application of these modern AI approaches is creative framing of the question so it can be answered quantitatively and careful attention to the necessary application-appropriate validation toward biological interpretation of the data. 

2026-05-05
Erick Matsen, Fred Hutch Cancer Center
Thinking Like a Biostatistician, Building Like an ML Researcher: a Case Study with Antibodies
Recording Available
Part of our AI for Biomedical Data Science Across Scales Statistical AI Symposium. Modern AI/ML is a stunningly good toolbox. But the loss functions used often come from other fields, such as language and vision, and don't always fit the biology. Biomedical statisticians, by contrast, obsess over matching the model to the problem. In this talk I'll pitch "thinking like a biostatistician, building like an ML researcher" through a case study on antibodies. Antibodies evolve by mutation and selection. The dominant tool for analyzing them is the protein language model: a transformer trained on hundreds of millions of sequences with a masked language modeling objective. But that objective is a density estimator over observed sequences, and density is not fitness! These models end up learning codon tables, hypermutation rates, and germline identity, none of which are relevant to function.

2026-05-05
William Noble, University of Washington
Foundation Models for Chromatin 3D ARchitecture and Proteomics Mass Spectrometry
Recording Available
Part of our AI for Biomedical Data Science Across Scales Statistical AI Symposium. In machine learning, a foundation model is a large-scale model that is trained in a self-supervised fashion on massive data and that can be easily fine-tuned to solve a variety of downstream tasks. In this talk, I will describe two recent efforts in my lab to develop foundation models. The first is designed to operate on 3D chromatin architecture data, represented as a DNA-DNA contact matrix produced by assays such as Hi-C. We show that the model, trained using techniques borrowed from image analysis, can be used for computing experimental similarity measures, improving effective sequencing depth, detecting chromatin loops, and predicting related, linear epigenomic measurements. The second model was initially designed for de novo sequencing of peptides from proteomics tandem mass spectrometry data. We show that the encoder learned by this model can be re-purposed for a variety of downstream tasks, including prediction of spectrum quality and chimericity, as well as presence of post-translational modifications.

2025-05-05
Lucas Liu, Fred Hutch Cancer Center
Histology to Molecular Insight: Transfer Learning for Molecular Biomarker Prediction in Prostate Cancer
Recording Available
Part of our AI for Biomedical Data Science Across Scales Statistical AI Symposium. Artificial Intelligence (AI) has the great potential to advance pathology and oncology by enabling automated cancer detection, disease grading, and treatment response prediction directly from histopathological slides. However, training an AI model for specific clinical applications such as molecular alteration prediction from scratch is still challenging because labeled training data is often limited. Transfer learning and foundation models are the key to resolve this challenge. Foundation models capture rich and generalizable representations from pretraining on millions of histopathological images, and transfer learning adapts that knowledge to specific tasks.

2025-05-05
Ziqi Rong, University of Washington
Transparent AI Agent for Differential Expression: Marrying Statistical Rigor with Natural Language Interfaces 
Part of our AI for Biomedical Data Science Across Scales Statistical AI Symposium. We present DE-Rigor-Agent, an AI-driven framework that enables differential expression analysis for bulk transcriptomics data through natural language input, while enforcing rigorous, expert-defined statistical assessment and decision-making. DE methods are context-dependent, varying with data properties, and there is no one-size-fits-all approach. Rather than operating as a black box that forces a single method pipeline, the agent systematically evaluates input data against methodological assumptions, screens out inappropriate methods, and executes statistically appropriate methods in parallel. By decoupling user interaction from method selection, our system eliminates subjective cherry-picking, delivering rigorous, reproducible, and fully auditable reports for differential expression analysis.

2025-05-05
Sonali Tamhankar, Fred Hutch Cancer Center
From Models to Decisions: Bridging Statistical AI and Clinical Practice through Interpretable Systems and Workforce Readiness
Recording Available
Part of our AI for Biomedical Data Science Across Scales Statistical AI Symposium. Achieving meaningful impact from healthcare AI requires more than strong models - it demands coordination across methodology, interpretability, deployment, and workforce readiness. In this talk, we articulate a layered framework for translating statistical innovation into clinical decision-making, illustrated through concrete applications and measurable outcomes. At the foundation lies statistics: a rigorous mathematical core that comes into full expression through modern machine learning. We present a Statistical AI project predicting which sickle cell patients are at elevated risk for emergency visits, demonstrating how principled modeling can address clinically relevant and operationally important questions.

2026-04-29
Li Zhang, UCSF
Translating Immune Repertoire Sequencing into Clinical Insight in Cancer through Network Analysis and Deep Learning
Special Joint Seminar with Immunotherapy IRC. T-cell receptor (TCR) recognition of antigenic peptides is central to anti-tumor immunity and immunotherapy response. Advances in high-throughput and single-cell sequencing technologies now generate large-scale immune repertoire and transcriptomic data, creating new opportunities to study immune dynamics in cancer. We developed complementary computational and deep learning frameworks to characterize and predict TCR–antigen interactions. We first introduced NAIR, a network-based approach that constructs TCR similarity networks to analyze clonal architecture, identify disease-associated clusters, and detect shared immune responses across patients and over time. To directly predict antigen specificity, we developed PepTCR-Net, a supervised deep learning framework that integrates sequence and network information to improve TCR–peptide recognition. More recently, we developed ITGAP, a multimodal deep learning framework that integrates TCR sequences with single-cell gene expression profiles to enhance antigen prediction at the cellular level. Applied across multiple cancer cohorts, these frameworks have uncovered clinically relevant immune dynamics—revealing treatment-associated remodeling of the tumor microenvironment in esophageal cancer, identifying circulating tumor-reactive T cell populations in bladder cancer, characterizing immune dysregulation underlying checkpoint inhibitor–associated colitis, and enabling large-scale identification of viral and tumor-associated antigen–reactive TCRs in hepatocellular carcinoma. Together, these studies demonstrate how network modeling and deep learning can translate high-dimensional immune sequencing data into clinically actionable insight, advancing translational cancer immunotherapy research.

2026-04-22
Gary Zhao, Fred Hutch Cancer Center
Mathematical Formulas that Describe the Blood System
Cancer prevention is routinely practiced in dermatology, OB/GYN, and gastroenterology, through population-wide screenings and early surgical interventions. However, leukemia prevention has not been achievable, as no efficient screening method or effective early intervention has been established. Gary’s research aims at solving these bottlenecks by taking a mathematical approach that directly translates to new technologies of leukemia screening and preventative treatments. In this seminar, Gary will provide an overview of his research program, highlighting two projects: (1) an NIH-funded project (PI: Zhao) that uses single-cell behavior-ome analysis in conjuncture with single-cell genomics in deeper characterization of hematopoietic stem cells, and (2) a FH-funded project (co-PI: Zhao) that uses statistical modeling and deep learning in pattern extraction and prediction of gene-gene interactions in blood cells.  Gary wishes to explain how his research is rooted in goal-oriented machine learning, intertwined with biostatistics research, and addressing public health challenges.

2026-04-01
Bo Zhang, Fred Hutch Cancer Center
Two Statistical Problems in Clinical Vaccine Clinical Trials
Recording Available
In one of our Biostat Program Mini Seminars, Bo Zhang presents two recent methodological developments motivated by challenges in clinical trials involving time-to-event outcomes, negative-control endpoints, and biomarkers.

2026-04-01
Kevin Lin, University of Washington
SCOPE: Localizing fate-decision states and their regulatory drivers in single-cell differentiation
One of two Biostat Program Mini Seminars.

2026-03-25
Wanlu Liu, Zhejiang University
Exploring Human T cells and TCRαβ repertoire at the single-cell level
Special Joint Seminar with Computational Biology.  T cell function is defined by both T cell receptors (TCR) and T cell gene expression (GEX). While single-cell technologies now enable the simultaneous capture of both modalities, the field lacks a comprehensive reference atlas and the computational framework necessary to decode fundamental TCR usage rules. In this talk, I will present our work in bridging this gap through the development of a single-cell T cell immune profiling database and novel tools for the integrative analysis of T cell states and TCR sequences. Together, these methods and resources provide a robust foundation for the characterization of disease-associated TCRs and serve as a vital resource for the T cell research community.

2026-01-14
Andrew Portuguese, Fred Hutch Cancer Center
Estimating Treatment Effects Without Randomization: Propensity Scores in Real-World Oncology Studies
Recording Available
Randomized clinical trials remain the gold standard for causal inference, but many clinically important questions in oncology must be addressed using real-world observational data. This seminar will focus on the use of propensity score methods to estimate treatment effects in non-randomized settings, with an emphasis on practical implementation, assumptions, and limitations. Using two real-world oncology case studies, I will illustrate how different propensity score approaches (matching and inverse probability of treatment weighting) align with different causal estimands and clinical questions. The talk will highlight common pitfalls, the importance of thoughtful covariate selection and balance diagnostics, and how observational analyses can inform, but not replace, randomized trials.
 

2025-12-10
Ruishan Liu, USC
AI for Precision Medicine and Clinical Trials
Toward a new era of medicine, our mission is to benefit every patient with individualized medical care. This talk explores how AI can make precision medicine more effective and diverse. I will first discuss Trial Pathfinder, a computational framework to optimize clinical trial designs (Liu et al. Nature 2021). Trial Pathfinder simulates synthetic patient cohorts from medical records, and enables inclusive criteria and data valuation. In the second part, I will discuss how to leverage large real-world data to identify genetic biomarkers for precision oncology (Liu et al. Nature Medicine 2022, Liu et al. Nature Communications 2024), and how to use language models to form individualized treatment plans (Liu et al. Cell Reports Medicine 2024).

2025-11-12
Steve Henikoff, Fred Hutch Cancer Center
Histone Overexpression in Cancer
Genome-wide hypertranscription is common in human cancer and predicts poor prognosis. To understand how hypertranscription might drive cancer, we applied our CUTAC method for mapping RNA polymerase II (RNAPII) genome-wide in formalin-fixed paraffin-embedded (FFPE) sections. RNAPII occupancy at S-phase-dependent histone genes accurately predicted rapid recurrence of meningiomas and corresponded to total whole-arm chromosome losses. Whole-arm losses alone predicted outcome in RNA-sequencing and whole-genome pan-cancer sequencing data. We propose that elevated RNAPII at histone genes both drives hyper-proliferation and displaces the CENP-A histone H3 variant from centromeres, causing centromere breaks and aneuploidies that shape the selective landscape in cancer progenitor cells. Our experimental investigation of the S-phase-dependent histone genes in Drosophila and human cells led to the discovery of histone H4 as a conserved histone gene feedback repressor. Our findings can guide development of anti-cancer therapeutics targeting the ancient histone gene expression machineries.

2025-10-08
Manu Setty, Fred Hutch Cancer Center
Continuous Representation of Cellular Systems
Recording Available
Single-cell technologies generate high-dimensional, noisy, and continuous data that challenge conventional discrete analysis methods. We present a novel probabilistic framework that learns smooth, differentiable density functions in cell-state space, enabling structured representation of cellular populations across development, disease, and perturbation. Our method, Mellon, infers these representations directly from single-cell measurements of any modality, efficiently capturing local geometry and global topology of the phenotypic landscape in high-dimensional feature space. This forms a biologically grounded data representation and foundation for downstream machine learning applications. We illustrate the utility of this representation with Kompot, a complementary framework for continuous differential analysis, which detects subtle, stage-specific gene-regulation changes along aging hematopoiesis, without requiring prior cell-type labeling or clustering. Together, these tools bridge representation learning and statistical inference, offering a foundation for scalable, interpretable modeling in molecular biology.

2025-10-01
Yun Li, UNC Chapel Hill
Investigating spatial omics data from multiple dimensions
Spatial omics technologies revolutionize studies of tissue functions. However, existing methods fail to capture localized, sharp changes characteristic of critical events such as tumor development. I will first present StarTrail, a gradient based method that powerfully defines rapidly changing regions and detects “cliff genes”, genes exhibiting drastic changes at highly localized or disjoint boundaries. StarTrail enables deeper insights into tissue spatial architecture. I will then introduce STimage-1K4M, a comprehensive dataset that provides >4Million paired transcriptomics and sub-tile image data from >1K images. STimage-1K4M offers unprecedented granularity, paving the way for a wide range of advanced research in multi-modal data analysis. Finally, I will introduce our nearest neighbor derivative process (NNGP). By providing a close-form solution, our NNGP reduces computational costs by multiple orders of magnitude, allowing it to be applied to high-resolution spatial omics data with a large number of spatial spots.

2025-09-10
Jingyi Jessica Li, Fred Hutch Cancer Center
Statistical Learning in the Wild: Rethinking Discovery in the AI and Data Era
The data landscape in modern science is shifting rapidly: many new data types are generated not to test pre-specified hypotheses, but to generate new ones. Yet most classical statistical tools were built for a different paradigm—confirming a hypothesis external to the data itself. This mismatch has led to widespread challenges, such as “double dipping,” where exploratory and confirmatory analysis blur, and traditional safeguards like type I error or FDR control may no longer align with the goals of data-driven discovery.

2025-06-11
Will Ma, HopeAI
AI-powered Clinical Development: Integrating Clinical Insights with Statistical Innovations
This study presents novel artificial intelligence methodologies that significantly enhance clinical trial design and execution processes through evidence-based optimization. We demonstrate how integrating comprehensive clinical evidence with statistical innovations transforms traditional drug development workflows. Our suite of AI solutions—PURE Evidence, SynthIPD, and specialized clinical AI assistants—addresses critical challenges in trial design, regulatory requirements, and development timelines. Validation studies conducted in collaboration with Mayo Clinic and Fred Hutch Cancer Center reveal that our models achieve twice the accuracy of conventional large language models in evidence-based medicine assessments. Results demonstrate quantifiable improvements in determining optimal sample sizes, establishing surrogate endpoints for accelerated approval pathways, and strengthening regulatory submissions. This research provides a framework for the evolution toward AI-augmented clinical development under human supervision, with implications for accelerating patient access to novel therapeutics while maintaining scientific rigor.

2025-04-30
Thomas Trikalinos, Brown University
Propagating Ambiguity into Decision Analyses of Test-and-Treat Strategies
I will discuss basics of decision making under ambiguity, also known as ‘deep uncertainty’ or ‘pervasive uncertainty.’ I operationally define ambiguity as uncertainty that the analyst is unwilling or unable to describe with a probability measure model, but is willing to described with alternative uncertainty models, specifically, with uncertainty sets.  For concreteness, I will discuss the decision analysis of whether the US should screen immigrants for latent tuberculosis infection, a problem with deep uncertainty about the performance of screening tests. In the example, I will motivate how ambiguity arises, introduce a natural way to model it, and outline methods to identify which actions are ‘best’ depending on the decision maker's attitude towards ambiguity.

2025-04-24
Xihong Lin, Harvard University
Navigate the Crossroad of Statistics, ML/AI and Genomic and Health Science
Special Joint Seminar with University of Washington Department of Biostatistics. Scalable and robust statistical and ML/AI methods and tools play a pivotal role intrustworthy science by accounting uncertainty, empowering scientific discovery, andimproving interpretability. In this talk, I will discuss the challenges and opportunitiesas we navigate the crossroad of statistics and ML/AI to empower genomic and healthscience. Examples include leveraging the AI/ML-generated synthetic data to empowerstatistical analysis of large biobank data in the presence of missing data, and scalableanalysis of the large whole genome sequencing studies and biobanks by leveragingvariant functional annotation and ensemble methods. We will discuss the analysis ofthe UK biobank of 500,000 subjects in the cloud platform RAP and the All of Us dataof 400,000 subjects in the NIH cloud platform AnVIL. This talk aims to ignite proactiveand thought-provoking discussions, foster cross-disciplinary collaboration, andcultivate open-minded approaches to advance scientific discovery. 

2025-04-09
Yu Shen, University of Texas MD Anderson Cancer Center
Evaluating Cancer Screening Programs: Statistical Methods and Microsimulation Insights
Statistical modeling is an effective tool to estimate medical costs and effectiveness in both cancer screening programs and cancer treatment, which is an important topic in health policy and health economics research. Over the last decades, many cancer screening trials have been conducted for breast, lung, colon, and prostate cancers. These trials generate data that can be used to estimate the preclinical sojourn time distribution, screening sensitivity, and other quantities of interest from the screening cohort. Information on the natural history of cancer is critical in designing optimal screening programs and assessing screening benefits. In this talk, we will review some statistical approaches to estimating these quantities based on screening trials and using them in microsimulations and their implications. We will show examples of using microsimulation modeling to investigate optimal cancer screening strategies by subjects’ risks in terms of cost and benefit at the population level.

2025-04-02
Lucas Liu, Fred Hutch Cancer Center
AI-Driven Digital Pathology: Advancing Precision Oncology
 Artificial Intelligence (AI) on pathology specimens has the potential to accurately identify cancer site, classify cancer subtype and predict treatment effect.  Recent studies suggest that AI methods can also make molecular predictions by learning genotype-phenotype relationships from pathology slides. Since molecular profiling is crucial for treatment and prognostic stratification, using AI-based methods on pathology images could potentially permit early understanding of individual tumor profiles prior to performing genomic tests. As such, AI-driven digital pathology has the potential to provide a cost-effective and time-efficient way to allocate individualized genomic tests and treatments.

2025-03-12
Tim Randolph, Fred Hutch Cancer Center
Simplicial-Complex Structures in Statistical Models for Brain Health
Methods to study co-activation among brain regions typically view the brain as a set of connected nodes. This facilitates the powerful mathematical framework of graph theory, allowing researchers to quantify testable hypotheses about connectivity. Graph theory, however, is limited to pairwise connections among the nodes; it cannot see higher-order co-activation patterns. I describe the larger mathematical framework of simplicial complexes (of which graphs are a special case) allowing us to quantify higher-order interactions: communities of edges, triads, tetrahedra, cliques, and cavities. The methods of Topological Data Analysis and Persistent Homology (cavities) are based on simplicial complex theory, but they fail to quantify much information. Instead, I consider “simplicially aware” statistical models aimed at estimating associations between neurological health and neuroconnectivity patterns that go beyond coarse summaries of homology, node degree, path length, centrality, etc. The methods are aimed at functional MRI data from studies on concussions, radiation therapy, alcohol use, and schizophrenia.

2025-03-05
Bo Zhang, Fred Hutch Cancer Center
Strengthening an Instrumental Variable: Applications to Health Services Research and Multi-center or Multi-stage Clinical Trials
In the first part of the talk,  I will discuss strengthening a continuous instrumental variable (IV) in the design of a matched observational study. We study how strengthening an IV may shorten the partial identification bounds for the sample average treatment effect (SATE) in an IV-based matched cohort study. Unlike the effect ratio estimand (i.e., the Wald estimate), SATE does not depend on who is matched to whom in the design, although a strengthened-IV design has the potential to narrow its partial identification bounds. We applied the method to studying a triage system that directs mothers to hospitals with varying capabilities using excess travel time as an IV.

2025-02-13
Siqi Shen, Fred Hutch Cancer Center
Computational Methods for Single-cell 3D Genomics
Recent advances in single-cell genomics have enabled unprecedented insights into cellular heterogeneity, particularly through high-throughput chromatin conformation capture (scHi-C) technologies that profile long-range genomic interactions. However, technical challenges including noise, sparsity, and missing modalities have limited the full potential of these data. We present a comprehensive computational framework that addresses these challenges through three complementary approaches. First, we introduce BandNorm, an efficient normalization method for scHi-C data that effectively separates cell types. Second, we develop scGAD scores as a dimension reduction tool that enables integration of scHi-C data with other single-cell modalities while accounting for gene-level genomic biases. Finally, we present GLEAM, a graph neural network-based approach that leverages link prediction to address missing modalities across single-cell datasets. Applied to mouse brain data encompassing gene expression, chromatin accessibility, electrophysiology, and chromatin loops, our integrated framework reveals novel connections between 3D chromatin architecture and neuronal electrophysiological properties. Together, these tools provide a robust foundation for multi-modal single-cell analysis, enabling deeper understanding of cellular organization and function.

2025-01-10
Bin Yu, UC Berkeley
Veridical Data Science and Alignment in Medical AI
Alignment and trust are crucial for the successful integration of AI in healthcare including digital twin projects, a field involving diverse stakeholders such as medical personnel, patients, administrators, public health officials, and taxpayers, all of whom influence how these concepts are defined. This talk presents a series of collaborative medical case studies where AI algorithms progressively become, from transparency to more opaque thus with increasing difficulty of alignment assessment. These range from tree-based methods for trauma diagnosis, to prostate cancer detection, to finding genetic drivers of a heart didease, and to interpreting language models for structured data extraction from pathology reports. They are guided by Veridical Data Science (VDS) principles—Predictability, Computability, and Stability (PCS)—for the goal of building trust and interpretability, enabling doctors to assess alignment. The talk concludes with a discussion on applying VDS to medical foundation models and next steps for evaluating AI algorithm alignment in healthcare.