Professional Development Course/CE
CE_03C: State-of-the-Art Classification and Regression Trees and Forest (Added Fee)
About this session
This course combines an overview of classification and regression trees and forests with an in-depth presentation of the GUIDE algorithm. Older algorithms (e.g., CART, C4.5, and Random forest) are reviewed, highlighting their strengths, weaknesses, and limitations. The GUIDE material covers key parts of the algorithm designed to overcome those weaknesses and limitations. Practically useful extensions include (i) modeling with missing covariate values without missing-value imputation (dispensing with missingness assumptions on missing covariate values), (ii) identification of subgroups with differential treatment effects for precision medicine, (iii) trees with logistic regression models in the nodes, (iv) multiple and longitudinal responses, (v) circular or periodic predictor variables (e.g., angles and time of day), and (vi) covariate-based clustering of response trajectories (e.g., observational studies on high-school dropouts and Alzheimer's disease). Learning highlights include (1) how GUIDE deals with missing values without requiring imputation, (2) how GUIDE importance scores help with variable selection, and (3) how post-selection inference is performed using a bootstrap calibration technique. To encourage hands-on training, the presentation is interwoven with live demos of GUIDE, RPART, CTREE, and other free software. No commercial software is required. Attendees are expected to be familiar with linear and logistic regression. The target audience is statisticians, data scientists, and researchers in business, government, industry, and academia. The course should be particularly useful for those who need to explore and analyze large and complex datasets with many variables and missing values and who want to learn to use free classification and regression tree software.
Session participants
Wei-Yin Loh
(University of Wisconsin)
Participant