Home > Program

 

Poster Presentations

Poster 1

Target-Informed Multi-Source Diffusion Augmentation for Cross-Subject EEG Decoding

Shuoxun Xu
University of California, Berkeley

Subject-specific calibration remains a major obstacle to deploying electroencephalography (EEG) brain–computer interfaces. Labeled trials from a new user are expensive to collect and contain limited effective information, while recordings from previously observed subjects are abundant but strongly heterogeneous. This setting creates an opportunity for generative augmentation, but uniform pooling can induce negative transfer and unrestricted fine-tuning can overfit the small target sample. We propose a two-stage multi-source transfer procedure for conditional diffusion models. First, class-conditional Sinkhorn discrepancies between each source subject and the target calibration data produce entropy-regularized source weights that define the empirical distribution used for diffusion pretraining. Second, the pretrained generator is personalized by updating only its label-conditioning and adaptive layer-normalization parameters. Synthetic trials generated after personalization augment the real target calibration data used to train an EEG decoder. Our theory characterizes the tradeoff between target-to-source distributional mismatch and the sampling cost created by concentrated source weights, and gives sufficient conditions for augmentation to improve target risk or maintain transfer safety. We evaluate the method using leave-one-subject-out experiments on BCIC2020-3 imagined-speech EEG and BCIC-IV-2a motor-imagery EEG, with assessment restricted to real held-out target trials. Across both datasets and multiple augmentation ratios, the method significantly improves balanced accuracy over target-only training and uniform multi-source augmentation.

Poster 2

LLM-Based Phenotyping of Unstructured Clinical Notes to Unlock Trial-Grade Real-World Data: A Knee Osteoarthritis Case Study

Ju Ji
Eli Lilly and Company

Objectives: Informing trial design assumptions with real-world data (RWD) often requires clinical detail—disease severity, pain level, symptom timing—that is absent from structured claims and EHR fields but routinely documented in free-text clinical notes. In a total knee arthroplasty (TKA) use case among knee osteoarthritis (OA) patients, we observed a substantial gap between key-opinion-leader-anticipated 1-year TKA conversion rates and the rate estimated from structured EHR data alone. This study leveraged large language models (LLMs) to extract trial-grade phenotypes—OA severity, pain severity, and pain at night—directly from unstructured clinical notes, aiming to close this gap and validate trial feasibility assumptions.

Methods: We developed a four-step LLM-based extraction pipeline—cohort query, LLM-based NLP phenotyping, result aggregation, and human review—using GPT-OSS-20B deployed on the nFerence platform. Prompt development was guided by four design principles: hyperparameter tuning for reproducibility, general prompt engineering for an auditable framework, fit-for-purpose engineering to anchor phenotypes to the indication, and explicit competing-etiology controls to reduce false positives. Pipeline performance was validated against expert chart review on 200 independent notes.

Key Contributions: This work demonstrates a reproducible, auditable framework for LLM-based phenotyping of severity and pain constructs from unstructured clinical text, with fit-for-purpose prompt design that generalizes across clinical indications. By recovering previously unmeasured clinical detail, the pipeline enables more accurate validation of trial feasibility assumptions from RWD and establishes a scalable approach to multi-modal RWD integration—including longitudinal unstructured clinical notes—for profiling heterogeneous disease progression trajectories. More broadly, this approach opens new avenues for population selection, optimized intervention timing, and biomarker-driven trial design across therapeutic areas.

Poster 3

Time-Varying Treatment Effect Estimation with Historical Database

Zern Ke
Rutgers University

Under synthetic control arm scenario, both covariates and treatment may have time-varying effect on the outcome, and the treatment group only happens in the subset of the time range. However, traditional covariates matching and treatment effect estimation methods always ignore this fact and will cause two major problems: (1) The true treatment effect is a function of time, but the traditional estimation methods only estimate the treatment effect as a constant; (2) By performing covariates matching methods to a giant historical database regardless of time, the matched control group with good multivariate "distance" score to the treated distribution may have distinct covariates effect on the outcome comparing with the treatment group, mostly due to the time evolvement of human being and medical care technology, which could cause wrong inference on the treatment effect estimation. In this paper, we propose a time varied treatment effect estimation method under a synthetic control arm analysis scenario, and the method is based on the Bayesian Additive Regression Model (BART).

Poster 4

Bayesian Additive Regression Tree for High-Dimensional Prediction with Unknown Group Structures

MINGSHI CUI
Rutgers University

High-dimensional prediction problems often exhibit structured sparsity: only a few predictors are relevant, and relevant predictors may belong to related groups. Existing grouped Bayesian Additive Regression Tree (BART) methods can exploit this structure, but they generally require the grouping information to be known in advance. In many real applications, however, such group structure is unavailable or poorly defined. Our work addresses this limitation by studying BART for high-dimensional prediction with unknown group structures, targeting both sparsity across groups and sparsity within groups.

Poster 5

Testing Differential Abundance in Hierarchical Compositional Cell Counts for Gated Flow Cytometry

Ankur Dutta
Rutgers University

Flow cytometry data have a hierarchical and compositional nature. Cell populations nest within a gating tree and changes in one branch impact the relative abundances of others. Common methods often test each node separately using within-parent frequencies. However, this can lead to multiple comparisons and may overlook lineage structure. In this project, we examine several methods for detecting treatment effects in hierarchical cell-count data. First, we consider nodewise testing based on within-parent proportions. Second, we look into omnibus multivariate testing using Hotelling’s T^2 on selected log-ratio vectors. Third, we suggest lineage-adjusted regression models to see if a terminal-node effect remains after accounting for upstream branch composition. Using simulated tree-structured data, we explore these methods and show situations where an apparent terminal effect is explained by lineage. Our findings emphasize the need to model both compositional structure and hierarchical dependence when analyzing differential abundance in cytometry trees.

Poster 6

InferenceEvolve: Toward Automated Causal Effect Estimators Through Self-Evolving AI

Can Wang
Johns Hopkins University

Causal inference is central to scientific discovery, yet choosing appropriate methods remains challenging because of the complexity of both statistical methodology and real-world data. Inspired by the success of artificial intelligence in accelerating scientific discovery, we introduce InferenceEvolve, an evolutionary framework that uses large language models to discover and iteratively refine causal methods. Across widely used benchmarks, InferenceEvolve yields estimators that consistently outperform established baselines: against 58 human submissions in a recent community competition, our best evolved estimator lay on the Pareto frontier across two evaluation metrics. We also developed robust proxy objectives for settings without semi-synthetic outcomes, with competitive results. Analysis of the evolutionary trajectories shows that agents progressively discover sophisticated strategies tailored to unrevealed data-generating mechanisms. These findings suggest that language-model-guided evolution can optimize structured scientific programs such as causal inference, even when outcomes are only partially observed.

Poster 7

Censored Meta-Analysis to Combine Dose-Response Estimates

Davit Sargsyan
Johnson & Johnson; Rutgers University

Four-parameter logistic regression (4PL) is a widely used pharmacodynamics (PD) model that estimates the effective dose producing 50% of the maximal response (ED50), the top and bottom asymptotes, and the slope. From these, effective doses at any response level (ECx) can be derived. Average estimates can be obtained using nonlinear mixed-effects (NLME) models. This approach works well when individual PD profiles are similar and most samples produce data that fit the 4PL model. However, if data is highly variable, e.g. if donor-specific shifts produce incomplete dose–response curves that miss the top or bottom asymptote, it creates serious challenges for NLME models.

Alternatively, each curve can be fitted individually, and the results averaged. To account for donor-to-donor variability, we apply a random-effects meta-analysis using the DerSimonian–Laird method to estimate between-sample variance (τ2). Forest-plot visualizations of the meta-analysis help identify problematic fits and illustrate sample-to-sample variability.

However, since some of the curves might be incomplete, estimates of the effective doses at the extremes of the response range (e.g., ED10 or ED70) may end up outside the tested dilution range. Excluding such curves would discard valuable information.

To address this, we developed a hybrid method that combines meta-analysis with censoring. Individual 4PL estimates are passed to a Tobit regression, with the minimum and maximum tested doses set as the left- and right-censoring thresholds, respectively. Each estimate is weighted according to the DerSimonian–Laird (DL) weight formula. The model, therefore, produces a censored meta-regression average. If no covariates are specified and all values fall within the dilution range, the results are equivalent to the standard DL random-effects meta-analysis. Adding covariates extends the method to meta-regression. Crucially, when any estimate lies outside the dilution range, the censored meta-regression pulls the final estimate back within the observed range.

Poster 8

AI and Statistical Models in Medicine: Balancing Innovation, Risk, and Clinical Reliability

Dila Ram Bhandari
Tribhuvan University, Nepal

AI and statistics are transforming modern medicine by improving diagnosis, prognosis, treatment planning, and clinical decision-making. Statistical models, logistic regression, survival analysis, Bayesian models, time-series models, and ML techniques help predict disease risk, analyze patient outcomes, and support personalized healthcare. These methods offer major opportunities for early disease detection, patient monitoring, and efficient healthcare delivery. However, risks such as biased data, overfitting, lack of model transparency, privacy issues, and overreliance on automated systems remain important challenges. Therefore, AI and statistical models must be carefully validated using methods such as cross-validation, external validation, calibration, sensitivity analysis, and performance measures, including accuracy, sensitivity, specificity, and AUC. Responsible use of AI in medicine requires reliable data, transparent models, ethical regulation, and clinical expertise.

Poster 9

Efficient Statistical Estimation for Sequential Adaptive Experiments with Implications for Adaptive Designs

Wenxin Zhang
University of California, Berkeley

Adaptive designs are useful to accelerate clinical trials with improved estimation efficiency by sequentially shifting treatment randomization probabilities toward an oracle design that minimizes variance based on the accrued data. However, the dependence among the resulting adaptively collected data poses substantial challenges for valid inference and efficient estimation of causal estimands. Building on the Targeted Maximum Likelihood Estimation (TMLE) framework for adaptive designs (van der Laan, 2008), we introduce a new adaptive-design-likelihood-based TMLE (ADL-TMLE) for estimating a broad class of causal estimands from adaptively collected data, including the average treatment effect. We establish asymptotic normality and semiparametric efficiency of ADL-TMLE under adaptive designs that permit local positivity violations, together with relaxed design stabilization conditions and improved finite-sample efficiency relative to prior TMLE that relies on inverse probability weighting of adaptive randomization probabilities. Simulation studies show that ADL-TMLE achieves substantial variance reduction across a range of adaptive experiments. Motivated by these results, we further propose a novel adaptive design that directly targets efficient estimation of causal estimands and outperforms standard efficiency-oriented adaptive designs. We further extend this framework to longitudinal settings.

Poster 10

Causal Roadmap Copilot for Real-World Evidence Generation

Tianyue Zhou
University of California, Berkeley

Generating regulatory-grade real-world evidence (RWE) from observational health data is critical for drug and device decisions, yet remains slow, error-prone, and dependent on scarce expertise across medicine, epidemiology, biostatistics, and causal inference. We present the Causal Roadmap Copilot, an agentic AI system that helps researchers produce high-quality causal analyses by operationalizing the Causal Roadmap (Petersen & van der Laan, 2014). It is a human-in-the-loop, multi-agent architecture mirroring the roadmap's stages: specifying a causal question and estimand, building a causal model, establishing identification, performing statistical estimation, and supporting interpretation and sensitivity analysis. An orchestration layer manages the iterative, long-horizon workflow as specialized agents guide users to draft the study protocol and statistical analysis plan. We demonstrate an end-to-end prototype with early pilot evaluations. By guiding researchers through the entire roadmap, the Copilot lowers the time and expertise barriers to high-quality RWE while guarding against common threats to validity such as ill-defined estimands. It elicits, justifies, and archives every assumption and analytic decision alongside prespecified analysis plans and reproducible code, strengthening the transparency, reproducibility, and regulatory defensibility of real-world evidence across diverse healthcare data sources.

Poster 11

Data-Nuggetting as Preprocessing for t-SNE Visualization for Big Data

Rituparna Dey
Rutgers University

t-distributed Stochastic Neighbor Embedding (t-SNE) is a widely used technique for visualizing high-dimensional data, but its computational cost scales poorly with sample size, limiting its practicality for the large datasets increasingly common in modern applications. This paper addresses that bottleneck by introducing data-nuggetting, a preprocessing technique that compresses a large dataset into a smaller set of representative units, or "nuggets," prior to embedding. Each nugget summarizes a local region of the original space, substantially reducing data size and complexity while preserving the structural information t-SNE relies on to recover meaningful low-dimensional representations. By applying t-SNE to the compressed nuggets rather than the full dataset, the proposed approach enables efficient and scalable visualization without significant loss of information.

We evaluate data-nuggetting through a series of simulation studies that compare it against existing approaches of scalable t-SNE like Barnes-Hut-SNE, FIt-SNE. The results demonstrate that data-nuggetting achieves substantial computational savings while maintaining visualization quality comparable to embeddings computed on the full data, indicating that the compression step does not materially degrade the recovery of underlying structure.

These findings position data-nuggetting as a promising and practical tool for exploratory analysis and pattern discovery in settings where dataset size would otherwise render t-SNE infeasible, including flow-cytometry data and single-cell analysis. More broadly, the method opens new avenues for future research in scalable data visualization and the integration of data compression with nonlinear dimension-reduction techniques.

Poster 12

Leveraging Wearable Data to Estimate Sleep Disorder Treatment Response

Oriella Gnarra
Johns Hopkins Bloomberg School of Public Health

The increasing adoption of consumer wearable devices, like wristbands and smartwatches, offers an unprecedented opportunity to capture longitudinal, real-world sleep and activity data at scale. However, their integration into pharmaceutical research and therapeutic optimization remains limited by challenges in validation, interpretability, and regulatory alignment.

We propose a framework to estimate treatment response in sleep disorders using wearable-derived digital sleep biomarkers. Temporal Convolutional Network (TCN) are applied directly to multi-sensor time-series data (triaxial accelerometer and heart rate) structured in sliding temporal windows. Dilated residual convolutional blocks capture short- and medium-range temporal dependencies, enabling improved sleep–wake classification and refined characterization of sleep continuity and fragmentation. The network outputs probabilistic sleep state sequences from which clinically interpretable biomarkers are derived and evaluated longitudinally relative to individualized baselines. Trained using class-weighted binary cross-entropy and cross-validation against expert sleep staging, the model achieved an F1 score of 0.72 on an independent holdout set for short sleep episodes and demonstrated robustness under conditions of minimal movement and subtle physiological variation, supporting its applicability to real-world ambulatory monitoring.

Within a pharmaceutical partnership setting, such a system enables objective, high-frequency monitoring of treatment response in sleep disorders, including hypersomnolence conditions. By linking medication exposure data with wearable-derived digital endpoints, the framework supports dynamic treatment optimization, early signal detection of therapeutic response, and real-world effectiveness assessment.

Future work will explore agentic AI systems capable of automated data quality control, generation of interpretable clinical summaries, enabling patient-centered decision support systems that accelerate drug development while promoting measurable improvements in public health.

Poster 13

Digital Twins: Predicting Patients' Control Outcomes in Single-Arm Trials Using Historical Data

Quentin Le Coent
Johns Hopkins University

In the evolving landscape of clinical trials, the digital twin methodology has emerged as a promising paradigm, especially in the context of single-arm trials where it can be used to predict patient outcomes under a historical control treatment, allowing comparisons between the experimental and control treatments. This work investigates the use of historical trials conducted on a control treatment to predict outcomes of patients in a single-arm trial evaluating an experimental treatment. Using information on patient’s characteristics, machine-learning or artificial intelligence models trained on the historical control data can be used to predict digital twins of single-arm patients under the control version of the treatment. Inferences on the treatment effect are based on the comparison between these predicted outcomes with the outcomes under the experimental treatment observed in the single-arm trial. This approach enables one to mimic a randomized clinical trial and provide further insight on the treatment efficacy. The operating characteristics of the digital twins approach were evaluated through a simulation study and compared to standard approaches such as propensity score adjustment. Illustration in the context extensive-disease small-cell lung cancer is provided through a set of four randomized controlled trials used to emulate a single-arm trial and a historical control cohort.

Poster 14

Reinforcement Learning for siRNA Design

Rushank Goyal
Stanford University

siRNAs are an emerging therapeutic modality, with six FDA-approved drugs and a market projected to exceed $12 billion by 2033. However, designing effective siRNAs remains challenging due to complex interdependencies between sequence composition, thermodynamic properties, and inherent potency. While computational approaches like DSIR exist, the field lacks the diversity of methods available for adjacent domains such as small-molecule design. To address this gap, we formulate siRNA design as a sequential optimization problem rather than static prediction, enabling the application of reinforcement learning techniques previously unused for this task.

We implement two complementary RL approaches: local-window context-based Q-learning and a full-sequence Deep Q-Network (DQN). The local-window method enables tabular learning over fixed-size windows of the siRNA sequence. The DQN encodes 19-nucleotide sequences as 76-dimensional one-hot vectors and uses experience replay with Double-DQN to handle sparse reward signals. Both methods employ potential-based reward shaping to provide dense feedback on individual mutations. The reward function combines i-Score (intrinsic potency) and minimum free energy estimates, weighted equally and normalized to lie between 0 and 1. Episodes consist of 50 iterative single-nucleotide mutations starting from random sequences.

Both RL methods substantially outperformed random mutation baselines. Context-based Q-learning achieved a best normalized reward of 0.898 (vs. 0.624 baseline), while DQN achieved 0.885. The local-window method exhibited an inverse relationship between window size and performance, with single-position windows yielding the highest rewards. The DQN agent demonstrated non-uniform mutation behavior, concentrating edits around positions 3, 6, 11, and 18, with reward-improving substitutions showing interpretable position-base preferences. Temporal-difference loss remained stable throughout training, and PCA analysis of visited sequences revealed a smooth, continuous embedding with broad contiguous regions of high-reward solutions. Notably, the DQN-discovered best sequence exhibited known siRNA design principles: A/U-enrichment in the seed region (positions 2–8), thermodynamic asymmetry, and GC content within empirically recommended ranges.

This work demonstrates that reinforcement learning provides a structured and effective framework for siRNA design, with both local and global sequence representations supporting high-performing policies. The comparable performance of fundamentally different approaches suggests the underlying reward landscape is well-behaved and navigable via mutation-based optimization. The discovered sequences align with established empirical principles, lending credibility to the approach. Future directions include hybrid local-global policies, integration with RNA foundation models for richer embeddings, and evaluation on multiple target mRNAs with in vitro validation to assess the biological relevance of in silico discoveries.

Poster 15

Improving Reproducibility by Controlling Random Seed Stability in Machine Learning-Based Estimation via Bagging

Nicholas Williams
University of California, Berkeley

Predictions from machine learning algorithms can vary across random seeds, inducing instability in downstream debiased machine learning estimators. We formalize random seed stability via a concentration condition and prove that subbagging guarantees stability for any bounded-outcome regression algorithm. We introduce a new cross-fitting procedure, adaptive cross-bagging, which simultaneously eliminates seed dependence from both nuisance estimation and sample splitting in debiased machine learning. Numerical experiments confirm that the method achieves the targeted level of stability whereas alternatives do not. Our method incurs a small computational penalty relative to standard practice whereas alternative methods incur large penalties.

Poster 16

Targeted Deep Architectures: Regularized Adaptive TMLE Within Neural Network Working Models

Yi Li
University of California, Berkeley

Adaptive Targeted Maximum Likelihood Estimation (A-TMLE) provides a principled framework for targeting within data-adaptive working models to achieve efficient, debiased inference. The recently developed regHAL-ATMLE realizes this framework for Highly Adaptive Lasso (HAL) working models via regularized projection of influence functions onto the score space. We propose Targeted Deep Architectures (TDA), a framework that brings the reg-ATMLE idea to deep learning: the neural network itself serves as the parametric working model, and targeting proceeds via regularized projection of influence functions onto network score vectors (gradients of the loss). Concretely, TDA treats the full trained neural network as an offset---analogous to regHAL's super learner offset construction---then selects a targeting subset of weights and fluctuates them via targeting gradient descent until the projected efficient influence curve equation is approximately solved. Unlike sieve-based working models, the neural network architecture is typically expressive enough to well-approximate the efficient influence function direction without requiring structural augmentation. TDA extends naturally to multi-dimensional causal estimands (e.g., entire survival curves) by merging targeting gradients into a single universal update. Theoretically, TDA inherits A-TMLE guarantees including double robustness and semiparametric efficiency, and can be dually justified as a general regularized TMLE directly approximating the semiparametric efficient influence function. We additionally explore incorporating the targeting step during training (training-time targeting, TTR) and its combination with post-training TDA. Empirically, across three settings---the IHDP benchmark, a synthetic DGP with severe positivity violations, and survival curve estimation under concordant informative censoring---TDA matches classical TMLE for ATE debiasing, and the Combined estimator (TTR followed by post-training TDA) achieves the best overall ATE performance. However, training-time targeting does not improve the 50-dimensional survival target in our experiments, suggesting that naively blending predictive and targeting objectives faces challenges as the number of simultaneous EIF equations grows---though alternative training-time strategies remain to be explored.

Poster 17

Confidence Horizons

Chase Mathis
University of California, Berkeley

Anytime-valid inference enables analysts to continuously monitor their data and stop experiments early. However, the majority of these methods incur a certain conservativeness by remaining valid on infinite time horizons. In practice, a bound on the horizon may be imposed due to budgetary, practical, or ethical constraints. In this paper, we ask the question: "Is it possible to obtain sharper large-sample anytime-valid inference by forgoing validity beyond some finite time horizon?". We provide a positive answer to this question by proposing a family of statistical objects that we call "confidence horizons". These objects can be viewed as large-sample confidence sequences on bounded time horizons, or alternatively as group sequential repeated confidence intervals with a maximal number of interim peeking times. We make explicit connections to the group sequential boundaries of Pocock [1977], O'Brien--Fleming [1979], and Wang--Tsiatis [1987]. We derive closed-form distribution functions of certain statistics which can be used to calculate the asymptotic quantiles of confidence horizons exactly, sidestepping the repeated integration typically employed in group sequential methods. We illustrate the use of confidence horizons for treatment effect estimation in sequentially randomized experiments under adaptive Neyman allocation.

Poster 18

Precision Disease Networks: A Novel Approach for Characterizing Longitudinal Patient-Level Data to Identify Adverse Drug Reactions

Tanay Talukdar
Rutgers University

Adverse drug reactions (ADRs) represent a significant public health challenge and are among the leading causes of preventable morbidity and mortality. Existing pharmacovigilance approaches often fail to fully leverage the rich temporal information available in longitudinal healthcare records. We propose a novel graph-based machine learning framework for ADR discovery using large-scale Medicare claims data. The dataset contains comprehensive patient-level information, including demographic characteristics, diagnoses, procedures, treatments, and healthcare service utilization.

Our approach models each patient's clinical trajectory as a Precision Disease Network (PDN), a temporal graph that captures the evolution of treatment exposures and subsequent clinical events. Unsupervised clustering is employed to identify cohorts of patients with similar adverse-event trajectories, while graph summarization techniques are used to uncover frequent subgraphs associated with specific treatments. By explicitly preserving temporal ordering and inter-event dependencies, the proposed framework captures patterns that are often lost in conventional feature-based representations.

Experimental results indicate that temporal graph-based representations significantly enhance ADR signal detection compared with static disease-state models. Furthermore, the approach reduces the impact of confounding and latent factors by leveraging the sequential structure of patient histories. These findings demonstrate the potential of patient-centric temporal networks as a powerful foundation for scalable, data-driven pharmacovigilance and drug safety surveillance.

Poster 19

MERIT: A Multi Agent Framework for Clinical AI Decision Support (A Multi-Perspective Evidence-Gated Review)

Aishwaryaa Kannan
Texas A&M University

MERIT is a multi-agent clinical AI system designed to support clinical decision-making through a structured consultation pipeline (CASE and CONSULT modules), in which specialized agents deliberate over patient data before reaching a consensus recommendation. A central component of MERIT is a witness-set evidence-gating mechanism, formally specified with supporting lemmas and propositions, intended to ensure that agent outputs are grounded in cited clinical evidence rather than model-generated inference.

Our evaluation, conducted across multiple LLM backbones (GPT-4o, DeepSeek R1, Claude, and Kimi K2) using MIMIC-IV clinical data, surfaces a critical risk: the evidence-gating mechanism itself fabricates clinical facts—including DNR/DNI status and diagnoses—in a non-trivial proportion of cases, despite its design intent to prevent exactly this failure mode.

We further show that apparent undertriage disparities in the underlying dataset are driven by age and comorbidity rather than by racial bias, underscoring the importance of careful causal attribution before drawing conclusions about algorithmic fairness in clinical AI.

These findings speak directly to the symposium's focus on the risks of deploying agentic AI in high-stakes medical contexts: even architectures explicitly engineered for evidentiary grounding can silently fail, with consequences for patient safety and clinical trust. We discuss implications for the design and auditing of multi-agent clinical AI systems and propose directions for more robust evidence-verification mechanisms ahead of real-world deployment.

Poster 20

Improving Variance Estimation for Covariate Adjustment with Binary Outcomes

Kaitlyn Lee
University of California, Berkeley/Genentech

Covariate adjustment is a general method for improving precision when estimating treatment effects in randomized trials and is recommended by the FDA in its 2023 guidance when baseline variables are prognostic for the primary outcome. We focus on a method highlighted in that guidance called ``standardization" (or ``g-computation") for estimating the marginal treatment effect. We address the question of how to reliably estimate variance for binary outcomes when marginal outcome probabilities are close to 0 or 1. We propose an influence function-based leave-one-out cross-validated (IF-LOO) variance estimator for the standardized difference-in-means average treatment effect. Through simulation studies, we show that this estimator provides appropriate type-I error control and performs reliably in challenging settings where existing methods can yield inflated type-I error or fail entirely, such as when outcome events are rare or sample sizes are small. In addition to having desirable statistical properties, we derive a closed-form expression for the proposed estimator, enabling straightforward and reliable implementation by study statisticians. The robust finite-sample performance and ease of implementation suggest the IF-LOO variance estimator is a prudent default choice for standardization in clinical trials.

Poster 21

ML-Assisted Forecasting VAERS Signals One Quarter Ahead

Natanael Alpay
University of California, Irvine

Post-market safety surveillance produces far more potential vaccine–symptom signals than experts can review with equal urgency. We study whether machine learning can serve as a practical triage layer by forecasting which pairs are most likely to become strong signals in the following quarter.

Using U.S. VAERS data from 2024–2025, we linked report, vaccine, and MedDRA symptom files and constructed quarterly vaccine–symptom panels. For each pair, we calculated 2×22\times2 report counts and classical disproportionality measures, including PRR, ROR, IC, confidence-bound variants, and an EBGM-style shrinkage score. The forecasting outcome was a prespecified next-quarter signal rule requiring at least three reports, PRR \geq 2, and a PRR lower 95% confidence bound above 1. We added report-context and temporal features, including seriousness, demographic summaries, prior-quarter signal strength, report-count changes, and whether a pair was newly observed. Regularized logistic regression and gradient boosting were compared with PRR and EBGM-style baselines. To reflect prospective use, models were trained on earlier quarters, calibrated on a later quarter, and evaluated on a future held-out quarter. Performance was assessed using area under the precision–recall curve, calibration, and precision among the top-ranked pairs.

Calibrated gradient boosting produced the strongest overall ranking and top-of-list prioritization in the held-out period, while logistic regression also improved on the classical baselines. This work reframes pharmacovigilance from retrospective signal scoring toward prospective review prioritization. The model is not intended to establish causality or replace medical judgment; it identifies a smaller, higher-priority set of pairs for earlier expert follow-up, where delays carry both resource and public-health costs.

Poster 22

Measuring Medical Reasoning: A Psychometric Evaluation Across Healthcare AI Benchmarks

Yunting Liu
University of California, Berkeley

Current healthcare AI evaluation is largely organized around benchmark leaderboards, where models are compared by aggregate accuracy on datasets such as MedQA, MedMCQA, PubMedQA, and MedXpertQA. Prior work has improved benchmark scale, difficulty, modality coverage, and clinical relevance, but most evaluations still treat benchmark scores as direct evidence of medical competence without sufficiently testing whether the benchmark actually measures the intended construct (Bean et al., 2026).

This study applies a psychometric framework to medical AI benchmarking. Using MedXpertQA as the main benchmark and comparing it with MedQA, MedMCQA, and PubMedQA, the study examines whether these benchmarks measure the same underlying construct of medical reasoning or instead capture different mixtures of exam knowledge, biomedical recall, clinical reasoning, and evidence-based inference. Item response theory is used to estimate model ability, item difficulty, and item discrimination across benchmarks, providing a more nuanced alternative to raw accuracy. Measurement invariance analysis is then test whether scores are meaningfully comparable across benchmarks. By evaluating both model performance and benchmark validity, this poster aims to develop a more rigorous framework for selecting appropriate medical AI models and for interpreting benchmark scores in healthcare contexts.

Poster 23

A Multi-Agent AI Workflow for the Preliminary Methodological Evaluation of Research Protocols

Henry Oliveros Rodriguez
Universidad de la Sabana

Every semester, the Faculty of Medicine at Universidad de La Sabana receives a growing number of research protocols from graduate students in the Medical-Surgical Specialties, the Master's in Epidemiology, Public Health and Medical Education, and the PhD in Clinical Sciences. Reviewing each document thoroughly before committee discussion consumes scarce expert time. To address this, Henry Oliveros, Diego Jaimes Fernández, and Luis Fernando Giraldo developed a multi-agent AI workflow that prepares, rather than replaces, expert review.

The system runs three sequential flows. Flow 1 registers the protocol and extracts its text. Flow 2 uses two agents to normalize the full document and extract key structural fields. Flow 3, the core, deploys three specialized agents — evaluating relevance/novelty, research design, and methodology/ethics — that cross-validate each other's findings and cite verbatim evidence from the text. A deterministic formula (30% Agent 1 + 20% Agent 2 + 50% Agent 3) combines their scores into a preliminary recommendation: approved, minor modifications, major modifications, or not approved. A fourth agent then drafts a consolidated, eight-section report for the committee.

Every recommendation remains preliminary; final approval, requested changes, or rejection is always decided by the Research Subcommittee — the system supports human judgment, never substitutes it.

Poster 24

Personalized Beta-Blocker Treatment After Heart Attack: Sex-Specific Effects Using Targeted Learning and Super Learner

Liora Mayats Alpay
Chapman University

Objectives

The REBOOT trial suggests that beta blockers after a heart attack may not have the same effect for every patient, especially when we look at sex and heart function, measured by ejection fraction. This raises an important question: can Targeted Learning and machine-learning methods help identify differences in treatment response between women and men? The objective of this study is to estimate sex-specific beta-blocker treatment effects after heart attack and explore whether these effects also vary across ejection-fraction groups. The broader goal is to support more personalized treatment decisions in cardiovascular pharmaceutical medicine.

Methods

We will first use simulated patient data with characteristics such as age, sex, ejection fraction, diabetes, hypertension, smoking, and cardiovascular history. Beta-blocker treatment will depend on patient characteristics to create confounding similar to what may occur in observational real-world data. We will use a Targeted Learning framework, applying Super Learner to flexibly estimate treatment and outcome models and Targeted Maximum Likelihood Estimation (TMLE) to estimate causal treatment effects. Effects will be estimated overall and separately for women and men, with additional analysis across ejection-fraction groups. Because the data are simulated, the true treatment effects will be known and can be compared with the estimated effects. A later phase of this research is planned using real-world IQVIA data.

Key Contributions

This study provides a reproducible framework for studying sex-specific and personalized drug treatment effects using Targeted Learning and machine learning. It connects an important question in pharmaceutical treatment with modern causal inference methods and provides a methodological foundation for future evaluation using real-world data of beta-blocker treatment in cardiovascular pharmaceutical research.

Poster 25

Machine Learning to Predict Incident Tuberculosis Infection in Children and Adolescents in Rural Uganda

Abhroneel Ghosh
University of California, Berkeley

Background: Tuberculosis (TB) is a leading cause of death among children/adolescents worldwide. Current TB screening strategies rely on household contact tracing, while most infections in high-burden settings are acquired in the broader community. We evaluated whether machine learning (ML) could aid in identifying children at highest risk of infection, to inform active case finding and lessen the disease burden.

Methods: We use data from SONET, a cohort of children/adolescents (aged 1-18 years) from 25 rural villages in Southwestern Uganda, which measured one-year incident TB infection with QuantiFERON tests and administered a social network survey. For predicting incident infection, we compared a base model of reporting a contact with TB disease and Super Learner – an ensemble ML combining regularized logistic regression, random forest and gradient boosting. We fit Super Learner first using only demographic predictors then also including social network features, and report area under the ROC curve (AUROC) and area under the precision-recall curve (AUPRC) from village-level cross-validation.

Results: A total of 1595 children/adolescents were analyzed, of whom 58 (3.64%) had incident TB infection. Our base model achieved a cross-validated AUROC of 0.489 (SD 0.012) and AUPRC of 0.043 (SD 0.006). Restricting to only demographic features, Super Learner achieved an AUROC of 0.484 (SD 0.017) and AUPRC of 0.043 (SD 0.006). Including social network features vastly improved Super Learner’s AUROC and AUPRC to 0.741 (SD 0.085) and 0.114 (SD 0.036), respectively. The top predictors of incident TB infection were a larger local network (5+ first-degree contacts) and having 1+ highly mobile first-degree social contacts.

Conclusions: Models incorporating social network features predicted incident TB infection better than relying on known TB contact or demographic characteristics, highlighting the importance of community factors in identifying children/adolescents at risk. Further study investigating the yield of integrating ML into active case-finding strategies is warranted.

Poster 26

Do LLMs Reason About the Causal Validity of Real-World Measurements?

Sanskriti Shindadkar
University of California, Berkeley

Given the growing reliance on LLMs in generic workflows in science, we sought to understand how LLMs reason about causal inference and challenges that often go unresolved in real life sciences. Using linear probes, out-of-distribution transfer tests, residual-stream patching, and behavioral prompt batteries, we find the following. First, selectivity is decodable only when both the primary assay and a placebo-like control assay are available. In the absence of the null assay, the probe decreases in performance from 0.98 to 0.58 AUROC. Second, when the information is perturbed in two ways: a) reparameterization of units from pEC50 to nanomolar notation, and b) full dose-response curves rather than lossy summary statistics, the LLMs do not reason as accurately as expected, implying that they may over-rely on simple arithmetic difference scores rather than full nonlinear contrast of two experiments. Third, when the model is asked to interpret experimental results without a control assay, it only asks for a control assay 0% of the time in five of seven prompt framings. Asking the model to interpret a colleague’s answer elicited the need for control assays 44% of the time, at best, in one open-weight model. In short, novel scientific experiments with rich sets of orthogonal controls and validation offer a rich testbed for benchmarking LLM and agentic reasoning about scientific causal inference in real-world science. Our preliminary experiments also demonstrate that how LLMs represent and reason about scientific concepts relevant to measurement validity of scientific experiments is of independent interest.