Data Mining of Temporal Data
- Share
- Partager sur Facebook
- Partager sur LinkedIn
Action Team
Scientific description
The massive volume of data, known as the "big data" phenomenon, has revolutionized traditional thinking in the fields of science and computer science, particularly in the area of statistical machine learning.
In many real-world problems—particularly those related to the Internet, though not limited to it—massive streams of data are continuously generated. This is the case, for example, with new types of data describing the dissemination of information on social networks (social dynamics), the organization of blog content (thematic patterns), various human activities in videos (human action recognition), and user preferences (collaborative filtering) available on the Web.
Beyond their sequential nature, data typically exhibit a complex internal structure, such as data describing electricity consumption curves, or data for which the basic assumption of machine learning—that observations are identically and independently distributed according to a fixed probability distribution—no longer holds.
The Khronos Action Team has focused on designing scalable machine learning algorithms capable of solving complex tasks such as large-scale multi-class classification or signal recovery, and of processing data with potentially unknown structures.
Key findings and future work
Large-scale adaptive signal recovery
Dmitrii Ostrovskii’s doctoral dissertation: Distributed estimation of non-stationary signals with an unknown local structure
We are interested in certain high-dimensional statistical problems where the signal has an unknown local structure, such as textured image reconstruction, speech segmentation, and sparse recovery problems in statistical signal processing.
In this case, it is impossible to calculate an effective linear filter a priori. We have proposed using a nonlinear filter that can be estimated from low-complexity data using the discrete Fourier transform operator and the l1 norm. In [2], we provided a simple sufficient condition, called approximate shift invariance, for the effectiveness of this procedure, and we showed that several important statistical models satisfy this condition, notably nonparametric kernel regression and spectral line signal estimation [3].
We have shown that approximate shift invariance guarantees the existence of an oracle filter with good statistical performance, and that the nonlinear filter learned by our procedure performs similarly to the oracle.
Evolutionary learning algorithms for distributed collaborative filtering and large-scale multi-class classification
Bikash Joshi's doctoral dissertation: Learning Algorithms for Big Data: Application to Multi-Class Classification and Asynchronous Distributed Optimization
To avoid undesirable scaling of sample complexity with respect to the number of classes, we designed a new approach based on learning a combination of similarity features between instances and classes.
Similarities are calculated by identifying a class along with its set of representative examples. We studied the consistency of learning with pairs of observations and classes by analyzing the associated dependency graph and showed that reducing the initial multiclass classification of examples to a binary classification of pairs of examples and classes allows for the learning of a single parameter vector whose dimension does not depend on the number of classes.
We have empirically demonstrated that this approach is competitive with state-of-the-art multi-class classification methods, particularly with regard to the F-score, which prioritizes the correct prediction of rare classes over classification accuracy. Furthermore, the number of parameters learned by the algorithm is approximately 107 times lower than that of conventional multi-class classification models, making this approach attractive for large-scale classification.
Coordinators
Massih-Reza Amini (LIG)
Thomas Burger (CEA)
Julie Fontecave (TIMC-IMAG)
Anatoli Juditsky (LJK)
Valuation
Industrial Collaboration
Khronos served as the platform for the Calypso FUI project (2015–2017) with Purch and Kelkoo on online advertising.
Scientific dissemination
- The Conference on Machine Learning (CAP 2017) will be held in Grenoble from June 28 to 30, 2017.
- In partnership with the Persyvact2 action team, which organized an international workshop in Grenoble on statistical tools for data mining (May 22 and 23, 2016), as well as spring and summer schools on high-dimensional statistics and optimization for big data (June 2014) and on large-scale parsimonious learning (April 2015).
Notable publications
[1] Babbar R., Partalas I., Gaussier E., Amini M.-R., Amblard C. Learning Taxonomy Adaptation in Large-scale Classication. Journal of Machine Learning Research (JMLR), 17(98) :1{37, 2016
[2] Ostrovsky, D., Harchaoui, Z., Juditsky, A., & Nemirovski, A. “Structure-blind signal recovery.” 30th Annual Conference on Neural Information Processing Systems (NIPS 29), pp. 4817–4825, 2016.
[3] Harchaoui, Z., Juditsky, A., Nemirovski, A., & Ostrovsky, D. “Adaptive recovery of signals by convex optimization.” Proceedings of the 28th Conference on Learning Theory (COLT), pp. 929–955, 2015.
[4] Joshi B., Amini M.-R., Partalas I., Ralaivola L., Usunier N., Gaussier E. On Binary Reduction of Large-scale Multiclass Classification Problems. 14th International Symposium on Intelligent Data Analysis (IDA), pp. 132–144, 2015
[5] Babbar R., Partalas I., Gaussier E., and Amini M.-R. “On Flat versus Hierarchical Classification in Large-Scale Taxonomies.” 27th Annual Conference on Neural Information Processing Systems (NIPS 2013), pp. 1824–1832, 2013.
- Share
- Partager sur Facebook
- Partager sur LinkedIn