Advanced search
Start date
Betweenand

Efficient, Scalable, and Self-Regulating Methods for Density-Based Machine Learning

Grant number:26/06676-6
Support Opportunities:Regular Research Grants
Start date: October 01, 2026
End date: September 30, 2029
Field of knowledge:Physical Sciences and Mathematics - Computer Science - Computer Systems
Principal Investigator:Murilo Coelho Naldi
Grantee:Murilo Coelho Naldi
Host Institution: Centro de Ciências Exatas e de Tecnologia (CCET). Universidade Federal de São Carlos (UFSCAR). São Carlos , SP, Brazil
City of the host institution:São Carlos
Associated researchers:André Takeshi Endo ; Hermes Senger

Abstract

The significant increase in unlabeled data generated today requires more scalable analysis methods, making density-based machine learning techniques - such as clustering - valuable tools for discovering implicit patterns in an automatic and unsupervised manner. However, traditional algorithms present significant limitations in dealing with this new scenario, as they were created for smaller and static volumes of information, failing to scale adequately and relying on complex parameterizations that require extensive prior knowledge from the user.This research project proposes the development of advanced density-based machine learning methods, focusing on overcoming the scalability limitations and the need for manual parameterization that affect traditional algorithms. The central proposal is based on the use and evolution of the CORE-SG graph, a structure that maps density relationships among the data and allows for performing tasks such as clustering, outlier detection, and semi-supervised learning in a robust manner. To achieve its goals, the project will adopt a multifaceted approach that combines the use of data structures for approximate proximity, aiming to reduce computational complexity from quadratic to sub-quadratic, and the implementation of high-performance algorithms that take advantage of GPU architectures.The strategy to ensure self-regulation is based on the creation of automatic mechanisms for cluster validation and model selection, eliminating the dependence on the user's prior knowledge about the data structure. Furthermore, the project will address the challenge of massive volumes and continuous streams of information through data summarization techniques, such as ``data bubbles'', which allow for representing large sets of points in a statistically compact and efficient manner. The practical feasibility of the proposal will be ensured by integrating the developed tools into open-source ecosystems and by using visualization systems, facilitating technology transfer and the application of the methods to complex real-world problems. (AU)

Articles published in Agência FAPESP Newsletter about the research grant:
More itemsLess items
Articles published in other media outlets ( ):
More itemsLess items
VEICULO: TITULO (DATA)
VEICULO: TITULO (DATA)