| Grant number: | 26/06676-6 |
| Support Opportunities: | Regular Research Grants |
| Start date: | October 01, 2026 |
| End date: | September 30, 2029 |
| Field of knowledge: | Physical Sciences and Mathematics - Computer Science - Computer Systems |
| Principal Investigator: | Murilo Coelho Naldi |
| Grantee: | Murilo Coelho Naldi |
| Host Institution: | Centro de Ciências Exatas e de Tecnologia (CCET). Universidade Federal de São Carlos (UFSCAR). São Carlos , SP, Brazil |
| City of the host institution: | São Carlos |
| Associated researchers: | André Takeshi Endo ; Hermes Senger |
Abstract
The significant increase in unlabeled data generated today requires more scalable analysis methods, making density-based machine learning techniques - such as clustering - valuable tools for discovering implicit patterns in an automatic and unsupervised manner. However, traditional algorithms present significant limitations in dealing with this new scenario, as they were created for smaller and static volumes of information, failing to scale adequately and relying on complex parameterizations that require extensive prior knowledge from the user.This research project proposes the development of advanced density-based machine learning methods, focusing on overcoming the scalability limitations and the need for manual parameterization that affect traditional algorithms. The central proposal is based on the use and evolution of the CORE-SG graph, a structure that maps density relationships among the data and allows for performing tasks such as clustering, outlier detection, and semi-supervised learning in a robust manner. To achieve its goals, the project will adopt a multifaceted approach that combines the use of data structures for approximate proximity, aiming to reduce computational complexity from quadratic to sub-quadratic, and the implementation of high-performance algorithms that take advantage of GPU architectures.The strategy to ensure self-regulation is based on the creation of automatic mechanisms for cluster validation and model selection, eliminating the dependence on the user's prior knowledge about the data structure. Furthermore, the project will address the challenge of massive volumes and continuous streams of information through data summarization techniques, such as ``data bubbles'', which allow for representing large sets of points in a statistically compact and efficient manner. The practical feasibility of the proposal will be ensured by integrating the developed tools into open-source ecosystems and by using visualization systems, facilitating technology transfer and the application of the methods to complex real-world problems. (AU)
| Articles published in Agência FAPESP Newsletter about the research grant: |
| More itemsLess items |
| TITULO |
| Articles published in other media outlets ( ): |
| More itemsLess items |
| VEICULO: TITULO (DATA) |
| VEICULO: TITULO (DATA) |