This paper introduces Augmented Functional Random Forests, a supervised classification framework for functional data that combines tree-based ensembles with leakage-free functional representation. The method augments smoothed curves with their first and second derivatives, applies blockwise FPCA separately to each functional representation, and tunes the number of retained components and ensemble parameters using training-based out-of-bag criteria. This construction allows level, slope, and curvature information to be incorporated into functional classification while controlling representation complexity. The paper also develops structured conditional permutation diagnostics for augmented FPC scores and combines them with a stability-based ranking criterion to mitigate correlation-induced distortions in feature importance. The resulting scores are interpreted as structured approximations to conditional relevance, rather than as pure conditional importance measures. Simulated scenarios and benchmark datasets show that derivative-based augmentation can improve accuracy when local dynamics are discriminative, while providing stability-aware feature relevance diagnostics.
Augmented Functional Random Forests: From Classifier Construction to Bias Adjusted Importance and Stability Based Feature Selection
Fabrizio Maturo
;Annamaria Porreca
2026-01-01
Abstract
This paper introduces Augmented Functional Random Forests, a supervised classification framework for functional data that combines tree-based ensembles with leakage-free functional representation. The method augments smoothed curves with their first and second derivatives, applies blockwise FPCA separately to each functional representation, and tunes the number of retained components and ensemble parameters using training-based out-of-bag criteria. This construction allows level, slope, and curvature information to be incorporated into functional classification while controlling representation complexity. The paper also develops structured conditional permutation diagnostics for augmented FPC scores and combines them with a stability-based ranking criterion to mitigate correlation-induced distortions in feature importance. The resulting scores are interpreted as structured approximations to conditional relevance, rather than as pure conditional importance measures. Simulated scenarios and benchmark datasets show that derivative-based augmentation can improve accuracy when local dynamics are discriminative, while providing stability-aware feature relevance diagnostics.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

