Ensemble Exhaustive-Subsample Partitioning: What Subsample Ensembles Estimate in Clusterwise Least Squares, and When They Are Worth Using
Type: Preprint
DOI: 10.2139/ssrn.7476799
Abstract
Clusterwise least squares partitions regression data into K groups with separate linear fits. We study a subsample ensemble: solve the problem exactly on each of B small random subsamples, extend each solution to the full sample, align the labels, and combine the replicates by a plurality vote, by averaging, or by selection. Each replicate is an empirical K-quantizer on m points, which makes an exact analysis possible. Conditionally on the data the replicates are independent and identically distributed, so the vote converges exponentially fast in B to the plurality partition of the replicate law, which coincides with the criterion minimizer outside a boundary set that does not shrink with B. For the two-group location model we show, under a partial-recovery condition, that m of order 1/π_min suffices for a replicate to be right more often than wrong. Exact minimizers for two groups and one covariate, computed in O(n^3) time, let us check each prediction. On clean data the ensemble is not a better optimizer than multistart alternation. Under gross outliers in the response, an adaptively trimmed variant needs no trimming level and had higher worst-case accuracy than trimmed alternation at any fixed level, with up to 20% contamination.
