The Corpus-driven Investigation of Personifications in Hungarian
The PerSE corpus
DOI:
https://doi.org/10.31400/dh-hun.2022.6.4409Keywords:
personification, corpus, annotation, analysisAbstract
Despite the recent findings on the conceptual and linguistic organization of personification, we have relatively little knowledge about its lexical patterns and grammatical templates. It is especially true in the case of Hungarian which has remained an understudied language regarding the constructions of figurative meaning generation. The present paper aims to provide a corpus-driven approach to personification analysis in the framework of cognitive linguistics. This approach is based on the building of a semi-automatically processed research corpus (the PerSE corpus) in which personifying linguistic structures are annotated manually. The present test version of the corpus consists of online car reviews written in Hungarian (10468 words altogether): the texts were tokenized, lemmatized, morphologically analyzed, syntactically parsed, and PoS-tagged with the emagyar NLP tool. For the identification of personifications, the adaptation of the MIPVU protocol was used and combined with additional analysis of semantic relations within personifying multi-word expressions. The paper demonstrates the structure of the corpus as well as the levels of the annotation. Furthermore, it gives an overview of possible data types emerging from the analysis: lexical pattern, grammatical characteristics, and the construction-like behaviour of personifications in Hungarian.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2022 the author(s)
This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.