How to evaluate uncertainty estimates in machine learning for regression?

Neural Netw

Institute for Computing and Information Sciences, Radboud University, Netherlands. Electronic address:

Published: May 2024

As neural networks become more popular, the need for accompanying uncertainty estimates increases. There are currently two main approaches to test the quality of these estimates. Most methods output a density. They can be compared by evaluating their loglikelihood on a test set. Other methods output a prediction interval directly. These methods are often tested by examining the fraction of test points that fall inside the corresponding prediction intervals. Intuitively, both approaches seem logical. However, we demonstrate through both theoretical arguments and simulations that both ways of evaluating the quality of uncertainty estimates have serious flaws. Firstly, both approaches cannot disentangle the separate components that jointly create the predictive uncertainty, making it difficult to evaluate the quality of the estimates of these components. Specifically, the quality of a confidence interval cannot reliably be tested by estimating the performance of a prediction interval. Secondly, the loglikelihood does not allow a comparison between methods that output a prediction interval directly and methods that output a density. A better loglikelihood also does not necessarily guarantee better prediction intervals, which is what the methods are often used for in practice. Moreover, the current approach to test prediction intervals directly has additional flaws. We show why testing a prediction or confidence interval on a single test set is fundamentally flawed. At best, marginal coverage is measured, implicitly averaging out overconfident and underconfident predictions. A much more desirable property is pointwise coverage, requiring the correct coverage for each prediction. We demonstrate through practical examples that these effects can result in favouring a method, based on the predictive uncertainty, that has undesirable behaviour of the confidence or prediction intervals. Finally, we propose a simulation-based testing approach that addresses these problems while still allowing easy comparison between different methods. This approach can be used for the development of new uncertainty quantification methods.

Download full-text PDF

Source
http://dx.doi.org/10.1016/j.neunet.2024.106203DOI Listing

Publication Analysis

Top Keywords

methods output
16
prediction intervals
16
uncertainty estimates
12
prediction interval
12
prediction
9
quality estimates
8
methods
8
output density
8
test set
8
output prediction
8

Similar Publications

Feasibility and preliminary efficacy of a physical activity intervention in adults with lymphoma undergoing treatment.

Pilot Feasibility Stud

January 2025

Department of Internal Medicine - Cardiology, Virginia Commonwealth University, West Hospital 8th Floor, North Wing, Richmond, VA, 23298, USA.

Background: To determine the feasibility, acceptability, and preliminary efficacy of a 6-month tailored non-linear progressive physical activity intervention (PAI) for lymphoma patients undergoing chemotherapy.

Methods: Patients newly diagnosed with lymphoma (non-Hodgkin (NHL) or Hodgkin (HL)) were randomized into the PAI or healthy living intervention (HLI) control (2:1). Feasibility was assessed by examining accrual, adherence, and retention rates.

View Article and Find Full Text PDF

Background: The pathogenesis of non-alcoholic fatty liver disease (NAFLD) with a global prevalence of 30% is multifactorial and the involvement of gut bacteria has been recently proposed. However, finding robust bacterial signatures of NAFLD has been a great challenge, mainly due to its co-occurrence with other metabolic diseases.

Results: Here, we collected public metagenomic data and integrated the taxonomy profiles with in silico generated community metabolic outputs, and detailed clinical data, of 1206 Chinese subjects w/wo metabolic diseases, including NAFLD (obese and lean), obesity, T2D, hypertension, and atherosclerosis.

View Article and Find Full Text PDF

Human interpretable structure-property relationships in chemistry using explainable machine learning and large language models.

Commun Chem

January 2025

Laboratory of Artificial Chemical Intelligence, Institute of Chemical Sciences and Engineering, Ecole Polytechnique Fédérale de Lausanne (EPFL), Lausanne, Switzerland.

Explainable Artificial Intelligence (XAI) is an emerging field in AI that aims to address the opaque nature of machine learning models. Furthermore, it has been shown that XAI can be used to extract input-output relationships, making them a useful tool in chemistry to understand structure-property relationships. However, one of the main limitations of XAI methods is that they are developed for technically oriented users.

View Article and Find Full Text PDF

Algorithmic Audits in Sports Medicine: An Examination of the SpartaScienceTM Force Plate System.

Med Sci Sports Exerc

November 2024

Department of Kinesiology, School of Education and Human Development, University of Virginia, Charlottesville, VA.

Introduction: Force plate systems are increasingly utilized in the armed forces that claim to identify individuals at risk of musculoskeletal injury. However, factors influencing injury risk scores from a force plate system (SpartaScienceTM), and the effects of experimental perturbations on these scores, remain unclear.

Methods: Healthy males (n = 823; 22.

View Article and Find Full Text PDF

Introduction: Angiotensin II may reduce muscle ischemia during intermittent hemodialysis and thereby decrease the incidence and/or intensity of intradialytic muscle cramps. We aimed to test whether angiotensin II infusion during intermittent hemodialysis is safe, feasible, and effective in the attenuation of muscle cramps.

Methods: We performed a pilot, single-blinded, randomized crossover trial of patients receiving intermittent hemodialysis who frequently experience intradialytic muscle cramps.

View Article and Find Full Text PDF

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!