Tree-Values: Selective Inference for Regression Trees.

Anna C Neufeld Lucy L Gao Daniela M Witten

J Mach Learn Res

Departments of Statistics and Biostatistics, University of Washington, Seattle, WA 98195, USA.

Published: January 2022

We consider conducting inference on the output of the Classification and Regression Tree (CART) (Breiman et al., 1984) algorithm. A naive approach to inference that does not account for the fact that the tree was estimated from the data will not achieve standard guarantees, such as Type 1 error rate control and nominal coverage. Thus, we propose a selective inference framework for conducting inference on a fitted CART tree. In a nutshell, we condition on the fact that the tree was estimated from the data. We propose a test for the difference in the mean response between a pair of terminal nodes that controls the selective Type 1 error rate, and a confidence interval for the mean response within a single terminal node that attains the nominal selective coverage. Efficient algorithms for computing the necessary conditioning sets are provided. We apply these methods in simulation and to a dataset involving the association between portion control interventions and caloric intake.

Download full-text PDF	Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC10933572	PMC

Publication Analysis

Top Keywords

selective inference

conducting inference

fact tree

tree estimated

estimated data

type error

error rate

inference

tree-values selective

inference regression

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!