Workshop
From Linkage to Inference: Estimating Associations and Causal Quantities with Linked Data (Intermediate; Added Fee)
About this session
Informed decision-making depends on complete and accurate information, yet data relevant to health and policy decisions are often fragmented across health providers, insurance systems, private organizations, and government agencies. Linking records that correspond to the same individual across these sources can substantially improve the ability to estimate important associations and causal relationships, particularly when no single dataset contains all relevant variables. Because privacy regulations frequently restrict access to unique identifiers such as social security numbers, names, or addresses, record linkage methods have become essential tools for integrating data in their absence. However, linked datasets are inherently subject to linkage error because matching decisions must often rely on incomplete, inconsistent, or error-prone identifying information. These challenges are especially pronounced when overlap between data sources is limited or when linkage quality differs across subpopulations, leading to unequal error rates across groups. Even relatively small numbers of incorrect matches can introduce meaningful bias in regression estimates and distort marginal and conditional associations.
This workshop will provide an overview of modern statistical methods for analyzing linked datasets while accounting for possible linkage errors. We will discuss approaches developed for both primary analyses, in which the analyst has access to the original source datasets, and secondary analyses, where only the linked dataset is available, highlighting the assumptions each method requires as well as their practical advantages and limitations. The workshop will cover methods for estimating both causal effects and associations, with particular attention to how linkage uncertainty can affect inference and decision-making. Throughout the session, methods will be illustrated using health-related examples and implemented with available tools in R, allowing participants to gain hands-on experience applying these approaches to realistic data settings. The goal is to develop both a conceptual understanding of statistical inference with linked records and the practical skills needed to use current methods in applied research. Participants should have prior experience using R and be comfortable conducting statistical analyses, as the workshop will involve coding exercises and interpretation of methodological output; prior experience with record linkage is not required.
This workshop is distinct from the previous workshop by concentrating on analysis methods of linked datasets and less on algorithms for implementing linkage.The workshop will be organized into four complementary modules that move from short linkage fundamentals to modern inference methods for linked data.1) Foundations of File Linkage and Analysis of Linked Data
The workshop will begin with a short introduction to file linkage, including why linkage is needed when information is distributed across multiple sources and unique identifiers are unavailable. We will provide brief overview of linkage approaches. This section will also introduce the distinction between primary analysis, where the analyst has access to original source files, and secondary analysis, where only the linked dataset is available, and discuss the assumptions commonly required when drawing inference from linked data.2) Record Linkage as a Missing Data Problem
A central theme of the workshop will be viewing linkage uncertainty as a missing data problem to properly propagate linkage errors into statistical inference. We will present weighting approaches that account for uncertain matches, Bayesian methods that jointly model linkage and analysis, and imputation-based strategies that incorporate multiple plausible linkages. Attention will be given to how these methods address uncertainty differently and when each approach is most appropriate in practice.3) Causal Inference with Linked Datasets
The third module will focus on estimating causal effects when key variables are distributed across linked sources. We will provide a brief introduction to the potential outcomes framework and discuss how linkage errors can affect treatment effect estimation. Methods for causal inference under linkage uncertainty will be introduced, with emphasis on practical considerations and assumptions needed for valid estimation.
4) Practical Implementation and Applied Examples
The final module will translate concepts into practice through implementation in R using available statistical packages. Participants will work through real-data and health-related examples, compare outputs across methods, and learn how to assess the impact of linkage uncertainty on substantive conclusions. By the end of the workshop, participants will have practical experience implementing modern methods for both associational and causal analyses of linked datasets.
1 Instructor
Brown University