Title: The Forward Search for Very Large Datasets
Authors: RIANI MarcoPERROTTA DomenicoCERIOLI Andrea
Citation: JOURNAL OF STATISTICAL SOFTWARE vol. 67 no. 1 p. 1-20
Publisher: JOURNAL STATISTICAL SOFTWARE
Publication Year: 2015
JRC N°: JRC83926
ISSN: 1548-7660
URI: http://www.jstatsoft.org/article/view/v067c01
http://publications.jrc.ec.europa.eu/repository/handle/JRC83926
DOI: 10.18637/jss.v067.c01
Type: Articles in periodicals and books
Abstract: The identification of atypical observations and the immunization of data analysis against both outliers and failures of modeling are important aspects of modern statistics. The forward search is a graphics rich approach that leads to the formal detection of outliers and to the detection of model inadequacy combined with suggestions for model enhancement. The key idea is to monitor quantities of interest, such as parameter estimates and test statistics, as the model is fitted to data subsets of increasing size. In this paper we propose some computational improvements of the forward search algorithm and we provide a recursive implementation of the procedure which exploits the information of the previous step. The output is a set of efficient routines for fast updating of the model parameter estimates, which do not require any data sorting, and fast computation of likelihood contributions, which do not require matrix inversion or qr decomposition. It is shown that the new algorithms enable a reduction of the computation time by more than 80%. Furthemore, the running time now increases almost linearly with the sample size. All the routines described in this paper are included in the FSDA toolbox for MATLAB which is freely downloadable from the internet.
JRC Directorate:Space, Security and Migration

Files in This Item:
There are no files associated with this item.


Items in repository are protected by copyright, with all rights reserved, unless otherwise indicated.