NISS Ai, Statistics & Data Science Webinar: Quantifying and Correcting Measurement Error in LLM-Generated Classifications

Tuesday, December 15, 2026 - 12:00pm to 1:30pm

Speaker

Yichi Zhang, Data Scientist at YouTube, Google

Moderator

Grace Deng, Research Data Scientist, Google

Abstract

Large language models are now routinely used to classify text at scale, producing variables that are then analyzed as though they had been observed directly. Because decoding is stochastic, the same input can receive different labels across repeated queries, and common practice is to query several times, aggregate by majority vote, and treat the aggregate as ground truth. Whatever error remains is neither estimated nor propagated, so every quantity computed afterward, including outcome prevalence, per-item classification, and the estimated effect of an intervention, carries uncertainty that goes unreported. We treat the problem as one of measurement. The true label is a latent variable, repeated model judgments are noisy measurements of it, and the classifier’s false acceptance and false rejection rates are estimated jointly with the quantities of interest, so that uncertainty about how often the model errs propagates into every downstream estimate. The framework needs no hand-labeled data when the classifier is accurate enough, and accommodates a small labeled subset otherwise. 
 
To establish what the framework recovers and where it fails, we collected 52,500 judgments from seven open-weight models spanning 1B to 32B parameters, under three prompt conditions with ten independent samples per item, on a public benchmark of reasoning problems carrying human-adjudicated correctness labels. Withholding those labels from the model lets us score every estimate against a known target. Repeated queries turn out to be far from independent. Ten of them carry the information of between 1.06 and 1.81 independent judgments, with a within-item correlation between 0.50 and 0.93 in every one of 21 model-by-prompt cells, and the effect does not depend on model scale or family. Models repeat themselves rather than sampling afresh, so treating ten queries as ten observations understates the standard error on any downstream estimate by a factor near 2.7. A binomial likelihood, which encodes the independence assumption the data violate, is then not merely optimistic but biased. Across 42 cells its stated 95% interval for the false acceptance rate contained the value computed from human labels zero times, and it understated total classification error in all 42. A beta-binomial likelihood carrying a single correlation parameter restores coverage. 
 
Two results govern when the framework can be trusted. Recovery is predicted almost monotonically by Youden’s J, the difference between the classifier’s true positive and false positive rates, which equals one minus its total error rate. Across the 42 cells its rank correlation with per-item classification accuracy is 0.96, and below a J of roughly 0.2 classification falls below chance even though the formal identification condition still holds. Hand labels are worth collecting only where J is small, where five labeled items close 87% of the error in the estimated prevalence, while at wider J the unsupervised estimate is already accurate to within 0.03. Labels improve population quantities at every value of J and per-item classification at none.
 

About the Speaker

Yichi Zhang is currently a data scientist at YouTube, Google. He currently works on developing and applying rigorous experiment designs and statistical methods for measuring impacts and driving decisions of products and partnerships by YouTube. He obtained his PhD degree in Biostatistics from Yale University. His research focused on developing causal inference and machine learning methodology for applications in biomedical and social sciences, aiming to unravel complex data patterns that incur selection and confounding bias to inform responsible, interpretable, and tractable decision-making for effective individual-level interventions or population-level policies.

 

About the Moderator

Grace Deng is currently a research data scientist at Google and formerly interned at Amazon Search and Instagram. She is a member of the NISS Affiliates Leadership Committee, and sits on the Industry Affiliates Subcommittee. Grace completed her PhD in Statistics from Cornell University in 2022 and received her undergraduate degree at UC Berkeley in Statistics and Economics. Her research focus includes generative ML/AI models for synthetic data and Bayesian time series models. She is the recipient of the Cornell Hemmeter Entrepreneurship Award (2020) and JSM Best Student Paper Award at JSM (2021), as well as various hackathon and datathon awards. 


Event Disclaimer

The views and opinions expressed by the speakers during this event are their own and do not necessarily reflect the views, positions, or policies of their employers, affiliated organizations, or any other entity. The speakers are participating in a personal capacity, and their statements should not be attributed to their respective companies or institutions.


 

About AI, StAtIstics and Data Science in Practice

The NISS AI, Statistics and Data Science in Practice is a monthly event series will bring together leading experts from industry and academia to discuss the latest advances and practical applications in AI, data science, and statistics. Each session will feature a keynote presentation on cutting-edge topics, where attendees can engage with speakers on the challenges and opportunities in applying these technologies in real-world scenarios. This series is intended for professionals, researchers, and students interested in the intersection of AI, data science, and statistics, offering insights into how these fields are shaping various industries. The series is designed to provide participants with exposure to and understanding of how modern data analytic methods are being applied in real-world scenarios across various industries, offering both theoretical insights, practical examples, and discussion of issues.

During Fall 2026, from September through December 2026, the series will focus on Trustworthy AI and the statistical, methodological, and governance foundations needed to develop, evaluate, and deploy AI systems responsibly and effectively. As AI becomes increasingly embedded in scientific research, business operations, public services, and societal decision-making, establishing confidence in the reliability, fairness, transparency, and accountability of these systems is essential. The series will examine approaches to measuring and mitigating bias, quantifying uncertainty and risk, evaluating robustness under changing conditions, and developing interpretable models and transparent evaluation frameworks that support informed decision-making. Emphasis will be placed on reproducibility, responsible data practices, privacy and security considerations, human oversight, and lifecycle monitoring to ensure that AI systems continue to perform as intended after deployment. By grounding discussions of AI development and governance in sound statistical reasoning and rigorous empirical evaluation, the series aims to promote AI systems that are not only accurate and innovative, but also trustworthy, equitable, and aligned with societal values.

See full list of featured topics (also below)

Featured Topics:

Location

United States