JSM2025
Back to the program
Topic-Contributed Paper Session

Advancing Distributed Learning for Complex and Heterogeneous Data

Mon, Aug 4, 10:30 AM - 12:20 PM Room CC-104D Music City Center
Keren LiOrganizerRan ChenChair
IMS co: Section on Statistical Learning and Data Scienceco: International Chinese Statistical Association Applied

About this session

This session delves into the latest advancements in distributed learning, with a focus on tackling complexities arising from data heterogeneity, privacy constraints, and scalability challenges. As the fields of statistics, data science, and AI rapidly evolve, it is increasingly essential to develop techniques that enable effective learning across distributed systems while addressing communication constraints and data privacy. This session highlights innovative statistical methods and frameworks that address these issues, from collaborative knowledge sharing to high-dimensional data analysis and functional data processing. Key Themes: 1. Distributed Collaborative Learning with Representative Knowledge Sharing: This title emphasizes the collaborative aspect, the use of knowledge sharing, and the idea of distillation on a representative dataset without needing a public dataset. It also subtly hints at addressing heterogeneity through weighted teacher models. 2. Reinforcement Learning in Distributed Systems: This talk discusses online decision-making strategies that exploit both the similarities and differences among nodes in distributed systems. Applications include business and healthcare. 3. High-Dimensional Problems in Distributed Learning and Deep Learning: An examination of the statistical methods used to address the challenges posed by high-dimensional data, particularly in the context of distributed and deep learning environments. 4. Harnessing Deep Learning and Distributed Systems for Next-Generation Functional Data: This presentation explores the innovative use of deep learning and distributed learning techniques to tackle the challenges posed by next-generation functional data. These advanced methodologies unlock new insights and capabilities, enabling more powerful and scalable data analysis solutions in the era of big data and complex functional datasets. 5. Generalized Information Criterion for Ensemble Kernel Learning: This talk introduces a computationally efficient method with using a new information criterion that promotes an ensemble kernel learning for large data analysis. This session is timely, addressing critical challenges in distributed learning, such as managing heterogeneity and enhancing collaborative learning without compromising data privacy. The methods presented offer scalable, innovative solutions in the age of big data, appealing to statisticians, data scientists, and AI researchers alike by providing insights into both theoretical advancements and practical applications. In line with the JSM 2025 theme, "Statistics, Data Science, and AI Enriching Society," this session showcases how statistical innovation in distributed learning contributes to the development of AI systems that are efficient, scalable, and socially beneficial. By emphasizing heterogeneity, privacy-preserving knowledge sharing, and advanced data integration, the session addresses crucial areas for the enrichment of data science and AI applications in society.