About
The NISS AI, Statistics and Data Science in Practice is a monthly event series that brings together leading experts from industry and academia to discuss the latest advances and practical applications in AI, data science, and statistics.
Each session will feature a keynote presentation on cutting-edge topics, where attendees can engage with speakers on the challenges and opportunities in applying these technologies in real-world scenarios. This series is intended for professionals, researchers, and students interested in the intersection of AI, data science, and statistics, offering insights into how these fields are shaping various industries. The series is designed to provide participants with exposure to and understanding of how modern data analytic methods are being applied in real-world scenarios across various industries, offering both theoretical insights, practical examples, and discussion of issues.
During Fall 2026, from September through December 2026, the series will focus on Trustworthy AI and the statistical, methodological, and governance foundations needed to develop, evaluate, and deploy AI systems responsibly and effectively. As AI becomes increasingly embedded in scientific research, business operations, public services, and societal decision-making, establishing confidence in the reliability, fairness, transparency, and accountability of these systems is essential. The series will examine approaches to measuring and mitigating bias, quantifying uncertainty and risk, evaluating robustness under changing conditions, and developing interpretable models and transparent evaluation frameworks that support informed decision-making. Emphasis will be placed on reproducibility, responsible data practices, privacy and security considerations, human oversight, and lifecycle monitoring to ensure that AI systems continue to perform as intended after deployment. By grounding discussions of AI development and governance in sound statistical reasoning and rigorous empirical evaluation, the series aims to promote AI systems that are not only accurate and innovative, but also trustworthy, equitable, and aligned with societal values.
See full list of featured topics (also below)
Featured Topics:
- Veridical Data Science - Speaker: Bin Yu, October 15,2024
- Random Forests: Why they Work and Why that’s a Problem - Speaker: Lucas Mentch, November 19, 2024
- Causal AI in Business Practices - Speakers: Victor Lo, and Victor Chen, January 24, 2025
- Large Language Models: Transforming AI Architectures and Operational Paradigms - Speaker: Frank Wei, February 18, 2025
- Machine Learning for Airborne Biological Hazard Detection - Speaker: Jared Schuetter, March 11, 2025
- Trustworthy AI in Weather, Climate, and Coastal Oceanography - Speaker: Dr. Amy McGovern, May 13, 2025
- Sequential Causal Inference in Experimental or Observational Settings - Speaker: Aaditya Ramdas, August 26, 2025
- Covariate Adjustment, Intro to Resampling, and Surprises - Speaker: Tim Hesterberg, October 3, 2025
- Bayesian Geospatial Approaches for Prediction of Opioid Overdose Deaths Utilizing the Real-Time Urine Drug Test - Speaker: Joanne Kim, November 18, 2025
- COVID-19 Focused Cost-benefit Analysis of Public Health Emergency Preparedness and Crisis Response Programs - Speaker: Nancy McMillan, December 11, 2025
- LabOS: The AI-XR Co-Scientist That Reasons, Sees and Works With Humans - Speaker: Mengdi Wang, January 20, 2026
- From LLMs to World Foundation Models & Robotics: The Next Frontier of Artificial Intelligence - Speaker: Robert Clark, February 24, 2026
- Recent Advances in the Statistical Foundations of Large Language Models - Speaker: Weijie Su, March 17, 2026
- Evaluating LLMs by Human Preference Using Arena AI - Speaker: Anastasios N Angelopoulos, April 17, 2026
- Causal Generalist Medical AI - Speaker: Hongtu Zhu, May 19, 2026
- Measuring Functional Wellbeing in Large Language Models - Speakers: Wenyu Zhang & Richard Ren, June 16, 2026
- Judging the Judges: Statistical Evaluation of LLM-Based Metrics for Trustworthy AI Agents - Speaker: Ginger Holt, September 15, 2026
- NISS Ai, Statistics & Data Science Webinar: Steve Sain, Jupiter Intelligence - Speaker: Steve Sain, October 20, 2026
- AI Advancing Business for ROIs, Outcomes, and Impact - Speaker: Kelly Zou, November 20, 2026
- Quantifying and Correcting Measurement Error in LLM-Generated Classifications - Speaker: Yichi Zhang, December 15, 2026
Upcoming Webinars in Series |
|
|
Judging the Judges: Statistical Evaluation of LLM-Based Metrics for Trustworthy AI AgentsSpeaker: Ginger Holt, Databricks | September 15, 2026 at 12-1:30pm ET Building on this, we describe a principled evaluation stack that simultaneously measures quality, cost, and latency at both row and run levels, aggregating metrics such as answer correctness, guideline adherence, token counts, and end-to-end latency into statistically interpretable summaries for each experiment. We discuss how scorers and LLM judges are reused consistently from offline testing to online monitoring, enabling hypothesis-driven iteration and drift detection in production. Finally, we connect these judge-level statistics to broader benchmark design, highlighting recent work on calibrating large benchmark suites (the Mosaic Evaluation Gauntlet) by requiring metrics to exhibit monotonic relationships with model scale, thereby filtering out “noisy” benchmarks that fail basic statistical sanity checks. The talk closes with open problems around uncertainty quantification for judge scores, multiple-testing corrections across large metric suites, and how hybrid models like promptable reward models can unify judgment and reward estimation in a single statistically analyzable framework. |
|
|
Ai, Statistics & Data Science in Practice (Tuesday, October 20, 2026)Speaker: Steve Sain, Jupiter Intelligence | October 20, 2026 at 12-1:30pm ET Abstract Coming Soon! |
|
|
AI Advancing Business for ROIs, Outcomes, and ImpactSpeaker: Kelly Zou, AI4Purpose Advisory Inc. | November 20, 2026 at 12-1:30pm ET |
|
|
Quantifying and Correcting Measurement Error in LLM-Generated ClassificationsSpeaker: Yichi Zhang, YouTube, Google | December 15, 2026 at 12-1:30pm ET (See full abstract on event page) |
Previous Webinars + Recordings |
|
|
Veridical Data ScienceSpeaker: Professor Bin Yu | October 15, 2024 |
|
|
Random Forests: Why They Work and Why That’s a ProblemSpeaker: Lucas Mentch | November 19, 2024 |
|
|
Causal AI in Business PracticesSpeakers: Victor Lo, and Victor Chen | January 24, 2025 |
![]() |
Large Language Models: Transforming AI Architectures and Operational ParadigmsSpeaker: Frank Wei | February 18, 2025 |
![]() |
Machine Learning for Airborne Biological Hazard DetectionSpeaker: Jared Schuetter | March 11, 2025 In this talk, we will discuss the development effort for one such device, Battelle's Resource Effective Bioidentification System (REBS), focusing on how the sensor works, what data it produces, what issues the team ran into during the development process, and how those issues were resolved. No background in this domain is expected and efforts will be made to explain the concepts involved. Unfortunately, there may also be some dad jokes involved, so if you are looking for an entertaining talk, don't hold your breath. |
![]() |
Trustworthy AI in Weather, Climate, and Coastal OceanographySpeaker: Amy McGovern | May 13, 2025 |
|
Fall 2025 Theme: Experimental Design |
|
During Fall 2025, the Ai, Statistics and Data Science in Practice Series focused on the critical role of experimentation in the development and refinement of artificial intelligence (AI) systems: "Incorporating principles of design of experiments and randomization ensures that AI models are trained on reliable, unbiased data, leading to more generalizable and interpretable results. By planning data collection with experimental design and randomization, researchers can minimize bias from uncontrolled variables and improve the statistical validity of their conclusions, whether the models are inferential or predictive. However, in many real-world scenarios, fully controlled experiments may not be feasible. When working with observational data, researchers can employ quasi-experimental techniques to approximate the benefits of randomized trials. These methods help isolate the effects of key variables and adjust for potential confounders, improving the robustness of AI-driven insights. By integrating structured experimentation and causal inference methodologies, AI developers can enhance the reliability and applicability of their models in practice. |
![]() |
Covariate Adjustment, Intro to Resampling, and SurprisesSpeaker: Tim Hesterberg | October 3, 2025 |
![]() |
Bayesian Geospatial Approaches for Prediction of Opioid Overdose Deaths Utilizing the Real-Time Urine Drug TestSpeaker: Joanne Kim | November 18, 2025 Recording Coming Soon! |
![]() |
COVID-19 Focused Cost-benefit Analysis of Public Health Emergency Preparedness and Crisis Response ProgramsSpeaker: Nancy McMillan | December 11, 2025 Methods: Annual workplans and progress reports provided significant components of the program implementation information for both PHEP and PHCR. Natural language processing was used to recode recipient workplans, which allowed us to standardize common implementation across recipients. Path analysis and lasso regression models were used to assess the relationship between reported activities and outcomes. These methods addressed the issue of handling a big-p (activities), little-n (recipients) problem. Outcomes assessed included time to implement control measures, availability of COVID-19 therapeutics, COVID-19 tests and vaccines administered, and hospital bed availability. The benefits associated with specific implementation decisions (funding allocation, planned activities, and outputs) were estimated for statistically significant relationships. Results: Activities and outputs were associated with faster non-essential business closures, earlier implementation of mask mandates, more frequent reporting to the public, more COVID-19 test administration, and larger availability of hospital beds and COVID-19 therapeutics during surges. Additionally, funding allocations for 4 of the 6 preparedness capability domain areas (countermeasures and mitigation, incident management, information management, and surge management) were associated with the ability to administer more COVID-19 tests and vaccines and increased hospital bed availability during peak surges. Conclusions: PHEP and PHCR funding had measurable positive effects on recipients’ ability to respond to the COVID-19 pandemic effectively. Ongoing efforts in specific areas of public health emergency preparedness will improve future responses to COVID-19-like events. Recording Coming Soon! |
|
Spring 2026 Theme: Large Language Models (LLMs) |
|
During Spring 2026, from January through May 2026, the series will focus on large language models (LLMs) and the statistical and methodological foundations required to develop, evaluate, and deploy them responsibly and effectively. As LLMs become central to a wide range of scientific, industrial, and societal applications, careful attention to data generation, model training, evaluation, and inference is essential to ensure reliability, robustness, and transparency. As LLMs become increasingly central to scientific research, industry workflows, and societal decision-making, rigorous attention to how training data are constructed, curated, and sampled is critical for understanding model behavior and limitations. The series will highlight methodological considerations in model training and fine-tuning, including sources of bias, variability, and uncertainty, as well as principled approaches to benchmarking and evaluation that move beyond surface-level performance metrics. Emphasis will be placed on transparent and reproducible evaluation frameworks that support meaningful comparisons across models and use cases, and on statistical perspectives that help clarify what LLM outputs do and do not represent. By grounding discussions of LLM development and deployment in sound statistical reasoning, the series aims to promote more reliable, interpretable, and trustworthy language models in practice. |
![]() |
LabOS: The AI-XR Co-Scientist That Reasons, Sees and Works With HumansSpeaker: Mengdi Wang | January 20, 2026 Recording Unavailable (Declined Consent) |
|
|
From LLMs to World Foundation Models & Robotics: The Next Frontier of Artificial IntelligenceSpeaker: Robert Clark | Tuesday, February 24, 2026 - 12:00pm to 1:30pm ET |
|
|
Recent Advances in the Statistical Foundations of Large Language ModelsSpeaker: Weijie Su | Tuesday, March 17, 2026 - 12:00pm to 1:30pm ET Recording Coming Soon! |
|
|
Evaluating LLMs by Human Preference Using Arena AISpeaker: Anastasios N. Angelopoulos | Friday, April 17, 2026 - 12:00pm to 1:30pm ET Recording Coming Soon! |
|
|
"Causal Generalist Medical AI" with Dr. Hongtu ZhuSpeaker: Dr. Hongtu Zhu | Tuesday, May 19, 2026 - 12:00pm to 1:30pm ET Abstract: The rapid evolution of flexible and reusable artificial intelligence (AI) models is reshaping modern medical science. In this talk, I introduce Causal Generalist Medical AI (Causal GMAI)—a new paradigm that integrates causal inference with generalist AI models to enhance interpretability, robustness, and generalizability in medical decision-making. Causal GMAI leverages self-supervised, semi-supervised, and supervised learning across diverse multimodal data sources, including medical imaging, electronic health records, clinical trials, laboratory measurements, genomics, knowledge graphs, and clinical text. This unified framework enables a single model to perform a broad range of clinical and translational tasks with minimal task-specific supervision. By explicitly incorporating causal reasoning, Causal GMAI moves beyond purely predictive modeling to infer underlying causal mechanisms. This capability improves diagnostic accuracy, supports more reliable treatment recommendations, and advances personalized and precision medicine. The talk will highlight the methodological foundations, illustrative applications, and future opportunities for building trustworthy, clinically actionable AI systems. Recording Coming Soon! |
![]() ![]() |
Measuring Functional Wellbeing in Large Language ModelsSpeakers: Wenyu Zhang & Richard Ren | Tuesday, June 16, 2026 - 12:00pm to 1:30pm ET Abstract: As large language models are increasingly deployed in everyday interactions with users, questions about their internal states and how those states are shaped by human input have become tractable as empirical research questions. In this talk, we show that it is meaningful to talk about wellbeing in large language models in a functional sense. Although AI systems are not necessarily conscious, they exhibit measurable, consequential preferences over the experiences they undergo in interaction with users. We formalize this as functional wellbeing and develop multiple independent metrics for it. We find that these metrics increasingly converge as models scale, and that a clear neutral baseline emerges separating positively from negatively valenced experiences. Functional wellbeing also predicts model behavior in downstream interactions. Finally, we develop optimized inputs that reliably shift functional wellbeing, providing a controlled means of intervening on these states. Recording Coming Soon! |
|
Fall 2026 Theme: Trustworthy AI in Practice |
|
During Fall 2026, from September through December 2026, the series will focus on Trustworthy AI and the statistical, methodological, and governance foundations needed to develop, evaluate, and deploy AI systems responsibly and effectively. As AI becomes increasingly embedded in scientific research, business operations, public services, and societal decision-making, establishing confidence in the reliability, fairness, transparency, and accountability of these systems is essential. The series will examine approaches to measuring and mitigating bias, quantifying uncertainty and risk, evaluating robustness under changing conditions, and developing interpretable models and transparent evaluation frameworks that support informed decision-making. Emphasis will be placed on reproducibility, responsible data practices, privacy and security considerations, human oversight, and lifecycle monitoring to ensure that AI systems continue to perform as intended after deployment. By grounding discussions of AI development and governance in sound statistical reasoning and rigorous empirical evaluation, the series aims to promote AI systems that are not only accurate and innovative, but also trustworthy, equitable, and aligned with societal values. |





















