4.7 Article

PPCA-Based Missing Data Imputation for Traffic Flow Volume: A Systematical Approach

Journal

Publisher

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC
DOI: 10.1109/TITS.2009.2026312

Keywords

Missing data; probabilistic principal component analysis (PPCA); traffic flow volume

Funding

  1. National Basic Research Program of China (973 Project) [2006CB705506]
  2. Hi-Tech Research and Development Program of China (863 Project) [2007AA11Z222]
  3. National Natural Science Foundation of China [60774034, 50708055]

Ask authors/readers for more resources

The missing data problem greatly affects traffic analysis. In this paper, we put forward a new reliable method called probabilistic principal component analysis (PPCA) to impute the missing flow volume data based on historical data mining. First, we review the current missing data-imputation method and why it may fail to yield acceptable results in many traffic flow applications. Second, we examine the statistical properties of traffic flow volume time series. We show that the fluctuations of traffic flow are Gaussian type and that principal component analysis (PCA) can be used to retrieve the features of traffic flow. Third, we discuss how to use a robust PCA to filter out the abnormal traffic flow data that disturb the imputation process. Finally, we recall the theories of PPCA/Bayesian PCA-based imputation algorithms and compare their performance with some conventional methods, including the nearest/mean historical imputation methods and the local interpolation/regression methods. The experiments prove that the PPCA method provides significantly better performance than the conventional methods, reducing the root-mean-square imputation error by at least 25%.

Authors

I am an author on this paper
Click your name to claim this paper and add it to your profile.

Reviews

Primary Rating

4.7
Not enough ratings

Secondary Ratings

Novelty
-
Significance
-
Scientific rigor
-
Rate this paper

Recommended

No Data Available
No Data Available