Busca avançada
Ano de início
Entree
(Referência obtida automaticamente do Web of Science, por meio da informação sobre o financiamento pela FAPESP e o número do processo correspondente, incluída na publicação pelos autores.)

Authorship recognition via fluctuation analysis of network topology and word intermittency

Texto completo
Autor(es):
Amancio, Diego R. [1]
Número total de Autores: 1
Afiliação do(s) autor(es):
[1] Univ Sao Paulo, Inst Math & Comp Sci, Dept Comp Sci, BR-05508070 Sao Paulo, SP - Brazil
Número total de Afiliações: 1
Tipo de documento: Artigo Científico
Fonte: JOURNAL OF STATISTICAL MECHANICS-THEORY AND EXPERIMENT; MAR 2015.
Citações Web of Science: 27
Resumo

Statistical methods have been widely employed in many practical natural language processing applications. More specifically, complex network concepts and methods from dynamical systems theory have been successfully applied to recognize stylistic patterns in written texts. Despite the large number of studies devoted to representing texts with physical models, only a few studies have assessed the relevance of attributes derived from the analysis of stylistic fluctuations. Because fluctuations represent a pivotal factor for characterizing a myriad of real systems, this study focused on the analysis of the properties of stylistic fluctuations in texts via topological analysis of complex networks and intermittency measurements. The results showed that different authors display distinct fluctuation patterns. In particular, it was found that it is possible to identify the authorship of books using the intermittency of specific words. Taken together, the results described here suggest that the patterns found in stylistic fluctuations could be used to analyze other related complex systems. Furthermore, the discovery of novel patterns related to textual stylistic fluctuations indicates that these patterns could be useful to improve the state of the art of many stylistic-based natural language processing tasks. (AU)

Processo FAPESP: 14/20830-0 - Modelagem e reconhecimento de padrões em textos com redes complexas
Beneficiário:Diego Raphael Amancio
Linha de fomento: Auxílio à Pesquisa - Regular