Parallel
PS02: Advancing Clinical Evidence Generation: AI, Data Governance, and Informatics in Drug Development
About this session
The rapid evolution of artificial intelligence, data governance, and informatics is transforming real-world evidence (RWE) generation, addressing critical challenges such as bias mitigation, causal inference, and data quality. This session highlights innovative methodologies that integrate large-scale models and advanced natural language processing (NLP) techniques, including large language models (LLMs), to revolutionize healthcare research and applications. Key topics include foundation models designed for real-world data (RWD), such as RWE-GPT, which incorporate negative control outcomes (NCOs) to address unmeasured confounders. These models leverage adversarial objectives, large-scale pretraining, and fine-tuning for robust causal estimation, with applications in drug repositioning and counterfactual modeling. To complement these advancements, the session highlights techniques for improving RWD quality, such as diffusion models and variational autoencoders, which address challenges like missingness and anomalies. Additionally, it explores the transformative role of LLMs in extracting insights from unstructured data sources-including electronic health records (EHRs), social media, and FDA safety reporting systems-for tasks like post-marketing drug safety monitoring. Efforts to represent clinical narratives in frameworks like the OMOP Common Data Model and to standardize workflows further underscore the integration of LLMs into RWE generation. This session demonstrates how advanced statistical methodologies and AI-driven techniques, including LLMs, are reshaping evidence generation and decision-making in healthcare and beyond.
Dr. Yong Ma: Use of NLP in FDA's post-marketing drug safety monitoring NLP has become an essential tool for extracting information from unstructured data, with its applications in FDA's post-marketing drug safety monitoring expanding in recent years. This presentation will highlight several examples of NLP in action, including extracting unstructured data from the FDA's FAERS system, identifying COVID-19 cases and symptoms from Reddit, and leveraging unstructured data from EHRs in Sentinel projects.
Dr. Margaret Gamalo: Transforming Clinical Trial Design: Improving Data Quality in Real-World Data for Informing Clinical Trials Ensuring data quality is essential for using Real-World Data (RWD) effectively, particularly in clinical research involving conditions like asthma and COPD. Challenges include missing data, ambiguity, incorrect data types, and event misordering. Missingness is especially problematic, as it is often non-random, risking selection bias when using complete-case analysis. Traditional imputation methods like mean or multiple imputation fall short due to RWD's high dimensionality. Advanced methods such as Conditional Score-based Diffusion Models for Tabular Data (TabCSDI) and Partial Variational Autoencoder (PartialVAE) show promise. TabCSDI encodes categorical and numerical data for imputation, while PartialVAE uses deep learning to model missing data distributions. For detecting anomalies, encoder-decoder frameworks measure reconstruction errors, while Bayesian networks model conditional dependencies to flag unlikely data points. These advanced methods offer robust solutions, enhancing RWD reliability for clinical applications
Dr. Hua Xu: Unlocking Clinical Textual Data for Real World Evidence Generation in the Era of LLMs Clinical documents within electronic health records (EHRs) are rich sources of patient information, pivotal for real-world studies. Natural language processing (NLP) is key in extracting and leveraging the data contained in these clinical narratives to generate real-world evidence. This presentation will outline the efforts of the OHDSI NLP Workgroup in harnessing textual data for evidence generation. It will cover topics such as representing textual data in the OMOP Common Data Model, applying advanced NLP techniques like large language models (LLMs) for information extraction, and the development of standardized workflows and tools for processing textual data. Additionally, it will also share the challenges encountered and lessons learned, offering valuable insights for researchers looking to implement LLMs in real-world studies.
Dr. Yong Chen: RWE-GPT: Toward a Large-Scale Pretrained Foundation Model with Negative Control Outcomes for Debiased Real-World Evidence Generation Generative pre-trained transformers (GPTs) have transformed natural language processing but lack focus on generating causal evidence from Real-World Data (RWD). We present RWE-GPT, the first GPT model tailored for RWD, incorporating negative control outcomes (NCOs) for debiasing. Pretrained on large datasets and fine-tuned on smaller, relevant cohorts, RWE-GPT uses an adversarial objective to address unmeasured confounders. Evaluated on data from over 20 million patients, it demonstrates robust performance in causal estimation tasks, supporting applications in drug repositioning and counterfactual modeling, showcasing the potential of foundation models in advancing real-world evidence generation.
4 Presentations
1:15 PM - 2:30 PM
1:15 PM - 2:30 PM
1:15 PM - 2:30 PM
1:15 PM - 2:30 PM