4.8 Article

Dominant Data Set Selection Algorithms for Electricity Consumption Time-Series Data Analysis Based on Affine Transformation

Journal

IEEE INTERNET OF THINGS JOURNAL
Volume 7, Issue 5, Pages 4347-4360

Publisher

IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC
DOI: 10.1109/JIOT.2019.2946753

Keywords

Correlation; Data mining; Error analysis; Internet of Things; Power systems; Big Data; Real-time systems; Affine transformation; dominant data set; linear correlation; time-series data (TSD)

Funding

  1. Heilongjiang Provincial Natural Science Foundation of China [F2016035]
  2. Science and Technology Project of State Grid Corporation of China [SGHL0000DKJS1900883]

Ask authors/readers for more resources

In the explosive growth of time-series data (TSD), the scale of TSD suggests that the scale and capability of many Internet of Things (IoT)-based applications has already been exceeded. Moreover, redundancy persists in TSD due to the correlation between information acquired via different sources. In this article, we propose a cohort of dominant data set selection algorithms for electricity consumption TSD with a focus on discriminating the dominant data set that is a small data set but capable of representing the kernel information carried by TSD with an arbitrarily small error rate less than epsilon. Furthermore, we prove that the selection problem of the minimum dominant data set is an NP-complete problem. The affine transformation model is introduced to define the linear correlation relationship between TSD objects. Our proposed framework consists of the scanning selection algorithm with O(n(3)) time complexity and the greedy selection algorithm with O(n(4)) time complexity, which are, respectively, proposed to select the dominant data set based on the linear correlation distance between TSD objects. The proposed algorithms are evaluated on the real electricity consumption data of Harbin city in China. The experimental results show that the proposed algorithms not only reduce the size of the extracted kernel data set but also ensure the TSD integrity in terms of accuracy and efficiency.

Authors

I am an author on this paper
Click your name to claim this paper and add it to your profile.

Reviews

Primary Rating

4.8
Not enough ratings

Secondary Ratings

Novelty
-
Significance
-
Scientific rigor
-
Rate this paper

Recommended

No Data Available
No Data Available