Training state-of-the-art portuguese POS taggers without handcrafted features
Santos, Cicero Nogueira dos; Zadrozny, Bianca
O documento é disponibilizado pela fonte de origem, que mantém a versão integral e as condições de uso.
Resumo
Part-of-speech (POS) tagging for morphologically rich languages normally requires the use of handcrafted features that encapsulate clues about the language's morphology. In this work, we tackle Portuguese POS tagging using a deep neural network that employs a convolutional layer to learn character-level representation of words. We apply the network to three different corpora: the original Mac-Morpho corpus; a revised version of the Mac-Morpho corpus; and the Tycho Brahe corpus. Using the proposed approach, while avoiding the use of any handcrafted feature, we produce state-of-the-art POS taggers for the three corpora: 97.47% accuracy on the Mac-Morpho corpus; 97.31% accuracy on the revised Mac-Morpho corpus; and 97.17% accuracy on the Tycho Brahe corpus. These results represent an error reduction of 12.2%, 23.6% and 15.8%, respectively, on the best previous known result for each corpus.
Ficha do documento
- Tipo
- Artigo científico
- Ano
- 2014
- Instituição
- Springer Int Publishing Ag
- Fonte
- Repositório da FGV
- Idioma
- Inglês
- Acesso
- Acesso restrito
- Identificador
- oai:repositorio.fgv.br:10438/23485
- Temas
- Tecnologia
Conteúdos relacionados
- Artigo científicoA simulation-based approach to analyze the information diffusion in Microblogging Online Social NetworkIEEE · 2013
- OutroRelatório de Inteligência Digital #12 - A disseminação de deepfakes íntimos não consensuais nas redes sociaisFundação Getulio Vargas · 2025
- Artigo científicoA plataformização da influência da publicidade e seu papel no capitalismo de vigilânciaRevista Eptic · 2025
- OutroPodcast Meio Tempo #011 - O fenômeno Taylor SwiftFGV ECMI · 2024