Search CORE

3 research outputs found

Extração de contextos definitórios do Corpus COVID-19 com CQL

Author: Bocorny Ana Eliza Pereira
Kilian Cristiane Krause
Rebechi Rozane
Publication venue: Universidade de São Paulo. Faculdade de Filosofia, Letras e Ciências Humanas
Publication date: 01/01/2022
Field of study

Termos representam os conceitos de um domínio e sua compreensão permite o acesso aos saberes contidos nos textos especializados. Entender o significado dos termos, portanto, é de grande importância não apenas para que pesquisadores possam socializar seus estudos e descobertas, mas também para que profissionais e estudantes de várias áreas possam se valer da informação especializada em contextos de estudo e de trabalho. A evolução rápida do conhecimento muitas vezes não permite que a terminologia criada para designar conceitos seja dicionarizada com a necessária rapidez. Tal fato pode representar um grande desafio para aqueles que necessitam ter acesso ao conhecimento especializado. Tendo em vista o contexto descrito, este estudo parte da revisão de abordagens utilizadas para a extração automática de traços definitórios (TDs) e contextos definitórios (CDs) e propõe a utilização da ferramenta Corpus Query Language (CQL) para a extração de informações que auxiliem no entendimento da terminologia empregada em textos especializados. Em especial, verificamos a utilidade das sintaxes de busca construídas com a CQL para esse propósito, aplicando-as ao Corpus COVID-19. O percurso apresentado neste estudo poderá auxiliar não apenas especialistas da área médica, mas também tradutores, lexicógrafos e professores a processarem, de forma mais rápida e precisa, o conhecimento contido em textos especializados.Terms represent the concepts of a domain and by comprehending them readers have access to the knowledge contained in specialized texts. Therefore, understanding the meaning of terms is of great importance not only for researchers to share the results of their studies, but also for professionals and students from various areas to apply specialized information in their learning and working contexts. The fast-evolving knowledge does not always permit that the terminology created to designate new concepts is quickly inserted in dictionaries, and this may represent a great challenge for those who need access to specialized knowledge. After presenting approaches used in the last twenty years for the automatic extraction of definition traits (DT) and definition contexts (DC), we propose the use of the Corpus Query Language (CQL) tool to retrieve information that helps in understanding the terminology used in specialized texts. In particular, we attested the usefulness of search syntaxes built with CQL for this purpose, applying them to the COVID-19 Corpus. The path presented in this study can help not only specialists in the medical field, but also translators, lexicographers and teachers to process, in a faster and more accurate way, the knowledge contained in specialized texts

Lume 5.8

Cadernos Espinosanos (E-Journal)

Definition context extraction from the COVID-19 corpus with CQL

Author: Bocorny Ana Eliza Pereira
Kilian Cristiane Krause
Rebechi Rozane Rodrigues
Publication venue
Publication date: 01/01/2022
Field of study

Termos representam os conceitos de um domínio e sua compreensão permite o acesso aos saberes contidos nos textos especializados. Entender o significado dos termos, portanto, é de grande importância não apenas para que pesquisadores possam socializar seus estudos e descobertas, mas também para que profissionais e estudantes de várias áreas possam se valer da informação especializada em contextos de estudo e de trabalho. A evolução rápida do conhecimento muitas vezes não permite que a terminologia criada para designar conceitos seja dicionarizada com a necessária rapidez. Tal fato pode representar um grande desafio para aqueles que necessitam ter acesso ao conhecimento especializado. Tendo em vista o contexto descrito, este estudo parte da revisão de abordagens utilizadas para a extração automática de traços definitórios (TDs) e contextos definitórios (CDs) e propõe a utilização da ferramenta Corpus Query Language(CQL) para a extraçãode informações que auxiliem no entendimento da terminologia empregadaem textos especializados. Em especial, verificamos a utilidade das sintaxes de busca construídas com a CQLpara esse propósito, aplicando-as ao Corpus COVID-19. O percurso apresentado neste estudo poderá auxiliar não apenas especialistas da área médica, mas também tradutores, lexicógrafos e professores a processarem, de forma mais rápida e precisa, o conhecimento contido em textos especializados.Terms represent the concepts of a domain and by comprehending them readers have access to the knowledge contained in specialized texts. Therefore, understanding the meaning of terms is of great importance not only for researchers to share the results of their studies, but also for professionals and students from various areas to applyspecialized information in their learning and workingcontexts. The fast-evolving knowledge does not always permit that the terminology created to designate new concepts is quickly inserted in dictionaries, and this may represent a great challenge for those who need access to specialized knowledge. After presenting approaches used in the last twenty years for the automatic extraction of definition traits (DT) and definition contexts (DC), we propose the use of the Corpus Query Language (CQL) tool to retrieveinformation that helps in understanding the terminology used in specialized texts. In particular, we attested the usefulness of search syntaxes built with CQL for this purpose, applying them to the COVID-19 Corpus. The path presented in this study can help not only specialists in the medical field, but also translators, lexicographers and teachers to process, in a faster and more accurate way, the knowledge contained in specialized texts

Lume 5.8

Corpus-based automatic detection of example sentences for dictionaries for Estonian learners

Author: Koppel Kristina
Publication venue
Publication date: 20/02/2020
Field of study

Väitekirja elektrooniline versioon ei sisalda publikatsiooneNäitelause täidab sõnastikus kindlat eesmärki, aidates aru saada sõna tähendusest ja illustreerides sõna erinevaid kasutuskontekste. Näitelausete põhiallikas on mahukas tekstikorpus, kust aga käsitsi on näitelauset leida väga keeruline. Elektroonilise leksikograafia arenguga on Eestisse jõudnud mitmed töövahendid, mis aitavad automaatselt tuvastada eri sõnastike jaoks vajalikku infot, sealhulgas näitelauseid. Väitekirjas uuritakse, missugused parameetrid iseloomustavad Eesti Keele Instituudis koostatud sõnastike "Eesti keele sõnaraamat 2019", "Eesti keele põhisõnavara sõnastik 2014", "Eesti keele naabersõnad 2019" näitelauseid ning "Eesti keele A1−C1 õpikute korpuse 2018" lauseid. Uurimuse eesmärk on välja töötada meetod, mis võimaldab neid parameetreid arvestades korpusest automaatselt tuvastada eesti keele õppijatele sobivaid lauseid. Töö keskmes on reeglipõhine lähenemine, mida rakendatakse korpuspäringusüsteemi Sketch Engine integreeritud tööriista GDEX ehk Good Dictionary Examples näitel. Parameetrite häälestamiseks on osaliselt kasutatud ka masinõppe elemente. Sõnastiku näitelausete ja õpikulausete analüüs näitas, et hea eesti keele näitelause peab olema täislause ja vastama muuhulgas järgmistele parameetritele: on 4–20 sõnet pikk; ei sisalda sõnesid, mis on pikemad kui 20 tähemärki; ei alga teatud sõnaliikidega (nt sidesõnaga) ega tagasi viitavate sõnade (nt sellepärast) või sõnapaaridega (nt sellisel puhul); ei sisalda vulgaarseid ja halvustavaid sõnu, madala sagedusega sõnu jmt. Uurimuse tulemusena on loodud "Eesti keele õppekorpus 2018 (etSkELL)", mis sisaldab ainult välja töötatud parameetritele vastavaid lauseid. Õppekorpus on omakorda aluseks eesti keele õppekeskkonnale Sketch Engine for Estonian Language Learning ehk etSkELL ja veebilausetele Eesti Keele Instituudi keeleportaalis Sõnaveeb.The function of an example sentence in a dictionary is to help the reader understand the meaning of the headword and illustrate its contexts of use. Nowadays, the main source of example sentences is a large text corpus, where suitable sentences are hard to find. Luckily, e-lexicography has generated automatic tools to help detect various information for dictionaries, including example sentences. The dissertation examines certain parameters of the example sentences presented in the Dictionary of Estonian (2019), Basic Estonian Dictionary (2014), Estonian Collocations Dictionary (2019), and Estonian Coursebook Corpus (2018); all four were compiled at the Institute of the Estonian language. The aim of my study is to elaborate an automatic method using parameters which identify sentences suitable for learners of Estonian. To that end, a rule-based approach was applied to the example of Good Dictionary Examples (GDEX) integrated in the Sketch Engine corpus query tool. Machine learning elements were also adopted to fine-tune the parameters. According to the analysis of the example sentences used in the dictionaries and coursebook sentences, a good Estonian example sentence should be a full sentence meeting, inter alia, the following parameters: length 4–20 tokens; no tokens longer than 20 characters; never begins with certain parts of speech (e.g., conjunction) or an anaphoric word (e.g., sellepärast ‘this is why’) or word pair (e.g., sellisel puhul ‘in such a case’); and vulgar or disparaging words, rare words, etc., are excluded. The study resulted in the compilation of the Estonian Corpus for Learners 2018 (etSkELL), which contains no other sentences but those corresponding to the developed parameters. The corpus, in turn, serves as the basis for the corpus-based web tool Sketch Engine for Estonian Language Learning (etSkELL) and the web sentences in the language portal Sõnaveeb of the Institute of the Estonian Language.https://www.ester.ee/record=b530293

DSpace at Tartu University Library