474 research outputs found

    ํ•œ๊ตญ์–ด ์‚ฌ์ „ํ•™์Šต๋ชจ๋ธ ๊ตฌ์ถ•๊ณผ ํ™•์žฅ ์—ฐ๊ตฌ: ๊ฐ์ •๋ถ„์„์„ ์ค‘์‹ฌ์œผ๋กœ

    Get PDF
    ํ•™์œ„๋…ผ๋ฌธ (๋ฐ•์‚ฌ) -- ์„œ์šธ๋Œ€ํ•™๊ต ๋Œ€ํ•™์› : ์ธ๋ฌธ๋Œ€ํ•™ ์–ธ์–ดํ•™๊ณผ, 2021. 2. ์‹ ํšจํ•„.Recently, as interest in the Bidirectional Encoder Representations from Transformers (BERT) model has increased, many studies have also been actively conducted in Natural Language Processing based on the model. Such sentence-level contextualized embedding models are generally known to capture and model lexical, syntactic, and semantic information in sentences during training. Therefore, such models, including ELMo, GPT, and BERT, function as a universal model that can impressively perform a wide range of NLP tasks. This study proposes a monolingual BERT model trained based on Korean texts. The first released BERT model that can handle the Korean language was Google Researchโ€™s multilingual BERT (M-BERT), which was constructed with training data and a vocabulary composed of 104 languages, including Korean and English, and can handle the text of any language contained in the single model. However, despite the advantages of multilingualism, this model does not fully reflect each languageโ€™s characteristics, so that its text processing performance in each language is lower than that of a monolingual model. While mitigating those shortcomings, we built monolingual models using the training data and a vocabulary organized to better capture Korean textsโ€™ linguistic knowledge. Therefore, in this study, a model named KR-BERT was built using training data composed of Korean Wikipedia text and news articles, and was released through GitHub so that it could be used for processing Korean texts. Additionally, we trained a KR-BERT-MEDIUM model based on expanded data by adding comments and legal texts to the training data of KR-BERT. Each model used a list of tokens composed mainly of Hangul characters as its vocabulary, organized using WordPiece algorithms based on the corresponding training data. These models reported competent performances in various Korean NLP tasks such as Named Entity Recognition, Question Answering, Semantic Textual Similarity, and Sentiment Analysis. In addition, we added sentiment features to the BERT model to specialize it to better function in sentiment analysis. We constructed a sentiment-combined model including sentiment features, where the features consist of polarity and intensity values assigned to each token in the training data corresponding to that of Korean Sentiment Analysis Corpus (KOSAC). The sentiment features assigned to each token compose polarity and intensity embeddings and are infused to the basic BERT input embeddings. The sentiment-combined model is constructed by training the BERT model with these embeddings. We trained a model named KR-BERT-KOSAC that contains sentiment features while maintaining the same training data, vocabulary, and model configurations as KR-BERT and distributed it through GitHub. Then we analyzed the effects of using sentiment features in comparison to KR-BERT by observing their performance in language modeling during the training process and sentiment analysis tasks. Additionally, we determined how much each of the polarity and intensity features contributes to improving the model performance by separately organizing a model that utilizes each of the features, respectively. We obtained some increase in language modeling and sentiment analysis performances by using both the sentiment features, compared to other models with different feature composition. Here, we included the problems of binary positivity classification of movie reviews and hate speech detection on offensive comments as the sentiment analysis tasks. On the other hand, training these embedding models requires a lot of training time and hardware resources. Therefore, this study proposes a simple model fusing method that requires relatively little time. We trained a smaller-scaled sentiment-combined model consisting of a smaller number of encoder layers and attention heads and smaller hidden sizes for a few steps, combining it with an existing pre-trained BERT model. Since those pre-trained models are expected to function universally to handle various NLP problems based on good language modeling, this combination will allow two models with different advantages to interact and have better text processing capabilities. In this study, experiments on sentiment analysis problems have confirmed that combining the two models is efficient in training time and usage of hardware resources, while it can produce more accurate predictions than single models that do not include sentiment features.์ตœ๊ทผ ํŠธ๋žœ์Šคํฌ๋จธ ์–‘๋ฐฉํ–ฅ ์ธ์ฝ”๋” ํ‘œํ˜„ (Bidirectional Encoder Representations from Transformers, BERT) ๋ชจ๋ธ์— ๋Œ€ํ•œ ๊ด€์‹ฌ์ด ๋†’์•„์ง€๋ฉด์„œ ์ž์—ฐ์–ด์ฒ˜๋ฆฌ ๋ถ„์•ผ์—์„œ ์ด์— ๊ธฐ๋ฐ˜ํ•œ ์—ฐ๊ตฌ ์—ญ์‹œ ํ™œ๋ฐœํžˆ ์ด๋ฃจ์–ด์ง€๊ณ  ์žˆ๋‹ค. ์ด๋Ÿฌํ•œ ๋ฌธ์žฅ ๋‹จ์œ„์˜ ์ž„๋ฒ ๋”ฉ์„ ์œ„ํ•œ ๋ชจ๋ธ๋“ค์€ ๋ณดํ†ต ํ•™์Šต ๊ณผ์ •์—์„œ ๋ฌธ์žฅ ๋‚ด ์–ดํœ˜, ํ†ต์‚ฌ, ์˜๋ฏธ ์ •๋ณด๋ฅผ ํฌ์ฐฉํ•˜์—ฌ ๋ชจ๋ธ๋งํ•œ๋‹ค๊ณ  ์•Œ๋ ค์ ธ ์žˆ๋‹ค. ๋”ฐ๋ผ์„œ ELMo, GPT, BERT ๋“ฑ์€ ๊ทธ ์ž์ฒด๊ฐ€ ๋‹ค์–‘ํ•œ ์ž์—ฐ์–ด์ฒ˜๋ฆฌ ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•  ์ˆ˜ ์žˆ๋Š” ๋ณดํŽธ์ ์ธ ๋ชจ๋ธ๋กœ์„œ ๊ธฐ๋Šฅํ•œ๋‹ค. ๋ณธ ์—ฐ๊ตฌ๋Š” ํ•œ๊ตญ์–ด ์ž๋ฃŒ๋กœ ํ•™์Šตํ•œ ๋‹จ์ผ ์–ธ์–ด BERT ๋ชจ๋ธ์„ ์ œ์•ˆํ•œ๋‹ค. ๊ฐ€์žฅ ๋จผ์ € ๊ณต๊ฐœ๋œ ํ•œ๊ตญ์–ด๋ฅผ ๋‹ค๋ฃฐ ์ˆ˜ ์žˆ๋Š” BERT ๋ชจ๋ธ์€ Google Research์˜ multilingual BERT (M-BERT)์˜€๋‹ค. ์ด๋Š” ํ•œ๊ตญ์–ด์™€ ์˜์–ด๋ฅผ ํฌํ•จํ•˜์—ฌ 104๊ฐœ ์–ธ์–ด๋กœ ๊ตฌ์„ฑ๋œ ํ•™์Šต ๋ฐ์ดํ„ฐ์™€ ์–ดํœ˜ ๋ชฉ๋ก์„ ๊ฐ€์ง€๊ณ  ํ•™์Šตํ•œ ๋ชจ๋ธ์ด๋ฉฐ, ๋ชจ๋ธ ํ•˜๋‚˜๋กœ ํฌํ•จ๋œ ๋ชจ๋“  ์–ธ์–ด์˜ ํ…์ŠคํŠธ๋ฅผ ์ฒ˜๋ฆฌํ•  ์ˆ˜ ์žˆ๋‹ค. ๊ทธ๋Ÿฌ๋‚˜ ์ด๋Š” ๊ทธ ๋‹ค์ค‘์–ธ์–ด์„ฑ์ด ๊ฐ–๋Š” ์žฅ์ ์—๋„ ๋ถˆ๊ตฌํ•˜๊ณ , ๊ฐ ์–ธ์–ด์˜ ํŠน์„ฑ์„ ์ถฉ๋ถ„ํžˆ ๋ฐ˜์˜ํ•˜์ง€ ๋ชปํ•˜์—ฌ ๋‹จ์ผ ์–ธ์–ด ๋ชจ๋ธ๋ณด๋‹ค ๊ฐ ์–ธ์–ด์˜ ํ…์ŠคํŠธ ์ฒ˜๋ฆฌ ์„ฑ๋Šฅ์ด ๋‚ฎ๋‹ค๋Š” ๋‹จ์ ์„ ๋ณด์ธ๋‹ค. ๋ณธ ์—ฐ๊ตฌ๋Š” ๊ทธ๋Ÿฌํ•œ ๋‹จ์ ๋“ค์„ ์™„ํ™”ํ•˜๋ฉด์„œ ํ…์ŠคํŠธ์— ํฌํ•จ๋˜์–ด ์žˆ๋Š” ์–ธ์–ด ์ •๋ณด๋ฅผ ๋ณด๋‹ค ์ž˜ ํฌ์ฐฉํ•  ์ˆ˜ ์žˆ๋„๋ก ๊ตฌ์„ฑ๋œ ๋ฐ์ดํ„ฐ์™€ ์–ดํœ˜ ๋ชฉ๋ก์„ ์ด์šฉํ•˜์—ฌ ๋ชจ๋ธ์„ ๊ตฌ์ถ•ํ•˜๊ณ ์ž ํ•˜์˜€๋‹ค. ๋”ฐ๋ผ์„œ ๋ณธ ์—ฐ๊ตฌ์—์„œ๋Š” ํ•œ๊ตญ์–ด Wikipedia ํ…์ŠคํŠธ์™€ ๋‰ด์Šค ๊ธฐ์‚ฌ๋กœ ๊ตฌ์„ฑ๋œ ๋ฐ์ดํ„ฐ๋ฅผ ์ด์šฉํ•˜์—ฌ KR-BERT ๋ชจ๋ธ์„ ๊ตฌํ˜„ํ•˜๊ณ , ์ด๋ฅผ GitHub์„ ํ†ตํ•ด ๊ณต๊ฐœํ•˜์—ฌ ํ•œ๊ตญ์–ด ์ •๋ณด์ฒ˜๋ฆฌ๋ฅผ ์œ„ํ•ด ์‚ฌ์šฉ๋  ์ˆ˜ ์žˆ๋„๋ก ํ•˜์˜€๋‹ค. ๋˜ํ•œ ํ•ด๋‹น ํ•™์Šต ๋ฐ์ดํ„ฐ์— ๋Œ“๊ธ€ ๋ฐ์ดํ„ฐ์™€ ๋ฒ•์กฐ๋ฌธ๊ณผ ํŒ๊ฒฐ๋ฌธ์„ ๋ง๋ถ™์—ฌ ํ™•์žฅํ•œ ํ…์ŠคํŠธ์— ๊ธฐ๋ฐ˜ํ•ด์„œ ๋‹ค์‹œ KR-BERT-MEDIUM ๋ชจ๋ธ์„ ํ•™์Šตํ•˜์˜€๋‹ค. ์ด ๋ชจ๋ธ์€ ํ•ด๋‹น ํ•™์Šต ๋ฐ์ดํ„ฐ๋กœ๋ถ€ํ„ฐ WordPiece ์•Œ๊ณ ๋ฆฌ์ฆ˜์„ ์ด์šฉํ•ด ๊ตฌ์„ฑํ•œ ํ•œ๊ธ€ ์ค‘์‹ฌ์˜ ํ† ํฐ ๋ชฉ๋ก์„ ์‚ฌ์ „์œผ๋กœ ์ด์šฉํ•˜์˜€๋‹ค. ์ด๋“ค ๋ชจ๋ธ์€ ๊ฐœ์ฒด๋ช… ์ธ์‹, ์งˆ์˜์‘๋‹ต, ๋ฌธ์žฅ ์œ ์‚ฌ๋„ ํŒ๋‹จ, ๊ฐ์ • ๋ถ„์„ ๋“ฑ์˜ ๋‹ค์–‘ํ•œ ํ•œ๊ตญ์–ด ์ž์—ฐ์–ด์ฒ˜๋ฆฌ ๋ฌธ์ œ์— ์ ์šฉ๋˜์–ด ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด๊ณ ํ–ˆ๋‹ค. ๋˜ํ•œ ๋ณธ ์—ฐ๊ตฌ์—์„œ๋Š” BERT ๋ชจ๋ธ์— ๊ฐ์ • ์ž์งˆ์„ ์ถ”๊ฐ€ํ•˜์—ฌ ๊ทธ๊ฒƒ์ด ๊ฐ์ • ๋ถ„์„์— ํŠนํ™”๋œ ๋ชจ๋ธ๋กœ์„œ ํ™•์žฅ๋œ ๊ธฐ๋Šฅ์„ ํ•˜๋„๋ก ํ•˜์˜€๋‹ค. ๊ฐ์ • ์ž์งˆ์„ ํฌํ•จํ•˜์—ฌ ๋ณ„๋„์˜ ์ž„๋ฒ ๋”ฉ ๋ชจ๋ธ์„ ํ•™์Šต์‹œ์ผฐ๋Š”๋ฐ, ์ด๋•Œ ๊ฐ์ • ์ž์งˆ์€ ๋ฌธ์žฅ ๋‚ด์˜ ๊ฐ ํ† ํฐ์— ํ•œ๊ตญ์–ด ๊ฐ์ • ๋ถ„์„ ์ฝ”ํผ์Šค (KOSAC)์— ๋Œ€์‘ํ•˜๋Š” ๊ฐ์ • ๊ทน์„ฑ(polarity)๊ณผ ๊ฐ•๋„(intensity) ๊ฐ’์„ ๋ถ€์—ฌํ•œ ๊ฒƒ์ด๋‹ค. ๊ฐ ํ† ํฐ์— ๋ถ€์—ฌ๋œ ์ž์งˆ์€ ๊ทธ ์ž์ฒด๋กœ ๊ทน์„ฑ ์ž„๋ฒ ๋”ฉ๊ณผ ๊ฐ•๋„ ์ž„๋ฒ ๋”ฉ์„ ๊ตฌ์„ฑํ•˜๊ณ , BERT๊ฐ€ ๊ธฐ๋ณธ์œผ๋กœ ํ•˜๋Š” ํ† ํฐ ์ž„๋ฒ ๋”ฉ์— ๋”ํ•ด์ง„๋‹ค. ์ด๋ ‡๊ฒŒ ๋งŒ๋“ค์–ด์ง„ ์ž„๋ฒ ๋”ฉ์„ ํ•™์Šตํ•œ ๊ฒƒ์ด ๊ฐ์ • ์ž์งˆ ๋ชจ๋ธ(sentiment-combined model)์ด ๋œ๋‹ค. KR-BERT์™€ ๊ฐ™์€ ํ•™์Šต ๋ฐ์ดํ„ฐ์™€ ๋ชจ๋ธ ๊ตฌ์„ฑ์„ ์œ ์ง€ํ•˜๋ฉด์„œ ๊ฐ์ • ์ž์งˆ์„ ๊ฒฐํ•ฉํ•œ ๋ชจ๋ธ์ธ KR-BERT-KOSAC๋ฅผ ๊ตฌํ˜„ํ•˜๊ณ , ์ด๋ฅผ GitHub์„ ํ†ตํ•ด ๋ฐฐํฌํ•˜์˜€๋‹ค. ๋˜ํ•œ ๊ทธ๋กœ๋ถ€ํ„ฐ ํ•™์Šต ๊ณผ์ • ๋‚ด ์–ธ์–ด ๋ชจ๋ธ๋ง๊ณผ ๊ฐ์ • ๋ถ„์„ ๊ณผ์ œ์—์„œ์˜ ์„ฑ๋Šฅ์„ ์–ป์€ ๋’ค KR-BERT์™€ ๋น„๊ตํ•˜์—ฌ ๊ฐ์ • ์ž์งˆ ์ถ”๊ฐ€์˜ ํšจ๊ณผ๋ฅผ ์‚ดํŽด๋ณด์•˜๋‹ค. ๋˜ํ•œ ๊ฐ์ • ์ž์งˆ ์ค‘ ๊ทน์„ฑ๊ณผ ๊ฐ•๋„ ๊ฐ’์„ ๊ฐ๊ฐ ์ ์šฉํ•œ ๋ชจ๋ธ์„ ๋ณ„๋„ ๊ตฌ์„ฑํ•˜์—ฌ ๊ฐ ์ž์งˆ์ด ๋ชจ๋ธ ์„ฑ๋Šฅ ํ–ฅ์ƒ์— ์–ผ๋งˆ๋‚˜ ๊ธฐ์—ฌํ•˜๋Š”์ง€๋„ ํ™•์ธํ•˜์˜€๋‹ค. ์ด๋ฅผ ํ†ตํ•ด ๋‘ ๊ฐ€์ง€ ๊ฐ์ • ์ž์งˆ์„ ๋ชจ๋‘ ์ถ”๊ฐ€ํ•œ ๊ฒฝ์šฐ์—, ๊ทธ๋ ‡์ง€ ์•Š์€ ๋‹ค๋ฅธ ๋ชจ๋ธ๋“ค์— ๋น„ํ•˜์—ฌ ์–ธ์–ด ๋ชจ๋ธ๋ง์ด๋‚˜ ๊ฐ์ • ๋ถ„์„ ๋ฌธ์ œ์—์„œ ์„ฑ๋Šฅ์ด ์–ด๋Š ์ •๋„ ํ–ฅ์ƒ๋˜๋Š” ๊ฒƒ์„ ๊ด€์ฐฐํ•  ์ˆ˜ ์žˆ์—ˆ๋‹ค. ์ด๋•Œ ๊ฐ์ • ๋ถ„์„ ๋ฌธ์ œ๋กœ๋Š” ์˜ํ™”ํ‰์˜ ๊ธ๋ถ€์ • ์—ฌ๋ถ€ ๋ถ„๋ฅ˜์™€ ๋Œ“๊ธ€์˜ ์•…ํ”Œ ์—ฌ๋ถ€ ๋ถ„๋ฅ˜๋ฅผ ํฌํ•จํ•˜์˜€๋‹ค. ๊ทธ๋Ÿฐ๋ฐ ์œ„์™€ ๊ฐ™์€ ์ž„๋ฒ ๋”ฉ ๋ชจ๋ธ์„ ์‚ฌ์ „ํ•™์Šตํ•˜๋Š” ๊ฒƒ์€ ๋งŽ์€ ์‹œ๊ฐ„๊ณผ ํ•˜๋“œ์›จ์–ด ๋“ฑ์˜ ์ž์›์„ ์š”๊ตฌํ•œ๋‹ค. ๋”ฐ๋ผ์„œ ๋ณธ ์—ฐ๊ตฌ์—์„œ๋Š” ๋น„๊ต์  ์ ์€ ์‹œ๊ฐ„๊ณผ ์ž์›์„ ์‚ฌ์šฉํ•˜๋Š” ๊ฐ„๋‹จํ•œ ๋ชจ๋ธ ๊ฒฐํ•ฉ ๋ฐฉ๋ฒ•์„ ์ œ์‹œํ•œ๋‹ค. ์ ์€ ์ˆ˜์˜ ์ธ์ฝ”๋” ๋ ˆ์ด์–ด, ์–ดํ…์…˜ ํ—ค๋“œ, ์ ์€ ์ž„๋ฒ ๋”ฉ ์ฐจ์› ์ˆ˜๋กœ ๊ตฌ์„ฑํ•œ ๊ฐ์ • ์ž์งˆ ๋ชจ๋ธ์„ ์ ์€ ์Šคํ… ์ˆ˜๊นŒ์ง€๋งŒ ํ•™์Šตํ•˜๊ณ , ์ด๋ฅผ ๊ธฐ์กด์— ํฐ ๊ทœ๋ชจ๋กœ ์‚ฌ์ „ํ•™์Šต๋˜์–ด ์žˆ๋Š” ์ž„๋ฒ ๋”ฉ ๋ชจ๋ธ๊ณผ ๊ฒฐํ•ฉํ•œ๋‹ค. ๊ธฐ์กด์˜ ์‚ฌ์ „ํ•™์Šต๋ชจ๋ธ์—๋Š” ์ถฉ๋ถ„ํ•œ ์–ธ์–ด ๋ชจ๋ธ๋ง์„ ํ†ตํ•ด ๋‹ค์–‘ํ•œ ์–ธ์–ด ์ฒ˜๋ฆฌ ๋ฌธ์ œ๋ฅผ ์ฒ˜๋ฆฌํ•  ์ˆ˜ ์žˆ๋Š” ๋ณดํŽธ์ ์ธ ๊ธฐ๋Šฅ์ด ๊ธฐ๋Œ€๋˜๋ฏ€๋กœ, ์ด๋Ÿฌํ•œ ๊ฒฐํ•ฉ์€ ์„œ๋กœ ๋‹ค๋ฅธ ์žฅ์ ์„ ๊ฐ–๋Š” ๋‘ ๋ชจ๋ธ์ด ์ƒํ˜ธ์ž‘์šฉํ•˜์—ฌ ๋” ์šฐ์ˆ˜ํ•œ ์ž์—ฐ์–ด์ฒ˜๋ฆฌ ๋Šฅ๋ ฅ์„ ๊ฐ–๋„๋ก ํ•  ๊ฒƒ์ด๋‹ค. ๋ณธ ์—ฐ๊ตฌ์—์„œ๋Š” ๊ฐ์ • ๋ถ„์„ ๋ฌธ์ œ๋“ค์— ๋Œ€ํ•œ ์‹คํ—˜์„ ํ†ตํ•ด ๋‘ ๊ฐ€์ง€ ๋ชจ๋ธ์˜ ๊ฒฐํ•ฉ์ด ํ•™์Šต ์‹œ๊ฐ„์— ์žˆ์–ด ํšจ์œจ์ ์ด๋ฉด์„œ๋„, ๊ฐ์ • ์ž์งˆ์„ ๋”ํ•˜์ง€ ์•Š์€ ๋ชจ๋ธ๋ณด๋‹ค ๋” ์ •ํ™•ํ•œ ์˜ˆ์ธก์„ ํ•  ์ˆ˜ ์žˆ๋‹ค๋Š” ๊ฒƒ์„ ํ™•์ธํ•˜์˜€๋‹ค.1 Introduction 1 1.1 Objectives 3 1.2 Contribution 9 1.3 Dissertation Structure 10 2 Related Work 13 2.1 Language Modeling and the Attention Mechanism 13 2.2 BERT-based Models 16 2.2.1 BERT and Variation Models 16 2.2.2 Korean-Specific BERT Models 19 2.2.3 Task-Specific BERT Models 22 2.3 Sentiment Analysis 24 2.4 Chapter Summary 30 3 BERT Architecture and Evaluations 33 3.1 Bidirectional Encoder Representations from Transformers (BERT) 33 3.1.1 Transformers and the Multi-Head Self-Attention Mechanism 34 3.1.2 Tokenization and Embeddings of BERT 39 3.1.3 Training and Fine-Tuning BERT 42 3.2 Evaluation of BERT 47 3.2.1 NLP Tasks 47 3.2.2 Metrics 50 3.3 Chapter Summary 52 4 Pre-Training of Korean BERT-based Model 55 4.1 The Need for a Korean Monolingual Model 55 4.2 Pre-Training Korean-specific BERT Model 58 4.3 Chapter Summary 70 5 Performances of Korean-Specific BERT Models 71 5.1 Task Datasets 71 5.1.1 Named Entity Recognition 71 5.1.2 Question Answering 73 5.1.3 Natural Language Inference 74 5.1.4 Semantic Textual Similarity 78 5.1.5 Sentiment Analysis 80 5.2 Experiments 81 5.2.1 Experiment Details 81 5.2.2 Task Results 83 5.3 Chapter Summary 89 6 An Extended Study to Sentiment Analysis 91 6.1 Sentiment Features 91 6.1.1 Sources of Sentiment Features 91 6.1.2 Assigning Prior Sentiment Values 94 6.2 Composition of Sentiment Embeddings 103 6.3 Training the Sentiment-Combined Model 109 6.4 Effect of Sentiment Features 113 6.5 Chapter Summary 121 7 Combining Two BERT Models 123 7.1 External Fusing Method 123 7.2 Experiments and Results 130 7.3 Chapter Summary 135 8 Conclusion 137 8.1 Summary of Contribution and Results 138 8.1.1 Construction of Korean Pre-trained BERT Models 138 8.1.2 Construction of a Sentiment-Combined Model 138 8.1.3 External Fusing of Two Pre-Trained Models to Gain Performance and Cost Advantages 139 8.2 Future Directions and Open Problems 140 8.2.1 More Training of KR-BERT-MEDIUM for Convergence of Performance 140 8.2.2 Observation of Changes Depending on the Domain of Training Data 141 8.2.3 Overlap of Sentiment Features with Linguistic Knowledge that BERT Learns 142 8.2.4 The Specific Process of Sentiment Features Helping the Language Modeling of BERT is Unknown 143 Bibliography 145 Appendices 157 A. Python Sources 157 A.1 Construction of Polarity and Intensity Embeddings 157 A.2 External Fusing of Different Pre-Trained Models 158 B. Examples of Experiment Outputs 162 C. Model Releases through GitHub 165Docto

    Windows into Sensory Integration and Rates in Language Processing: Insights from Signed and Spoken Languages

    Get PDF
    This dissertation explores the hypothesis that language processing proceeds in "windows" that correspond to representational units, where sensory signals are integrated according to time-scales that correspond to the rate of the input. To investigate universal mechanisms, a comparison of signed and spoken languages is necessary. Underlying the seemingly effortless process of language comprehension is the perceiver's knowledge about the rate at which linguistic form and meaning unfold in time and the ability to adapt to variations in the input. The vast body of work in this area has focused on speech perception, where the goal is to determine how linguistic information is recovered from acoustic signals. Testing some of these theories in the visual processing of American Sign Language (ASL) provides a unique opportunity to better understand how sign languages are processed and which aspects of speech perception models are in fact about language perception across modalities. The first part of the dissertation presents three psychophysical experiments investigating temporal integration windows in sign language perception by testing the intelligibility of locally time-reversed sentences. The findings demonstrate the contribution of modality for the time-scales of these windows, where signing is successively integrated over longer durations (~ 250-300 ms) than in speech (~ 50-60 ms), while also pointing to modality-independent mechanisms, where integration occurs in durations that correspond to the size of linguistic units. The second part of the dissertation focuses on production rates in sentences taken from natural conversations of English, Korean, and ASL. Data from word, sign, morpheme, and syllable rates suggest that while the rate of words and signs can vary from language to language, the relationship between the rate of syllables and morphemes is relatively consistent among these typologically diverse languages. The results from rates in ASL also complement the findings in perception experiments by confirming that time-scales at which phonological units fluctuate in production match the temporal integration windows in perception. These results are consistent with the hypothesis that there are modality-independent time pressures for language processing, and discussions provide a synthesis of converging findings from other domains of research and propose ideas for future investigations

    From Physical Motion to โ€˜Come and Goโ€™: A Spoken Corpus Based Analysis of Kata โ€˜goโ€™-specific Constructions in Korean

    Get PDF
    I analyze one of the motion verbs in Korean, kata โ€˜go,โ€™ and its argument structure constructions. The verb shows an extremely high token frequency and its argument structure constructions have been subject to a great degree of variation in terms of its emergent semantics and syntax. However, there have been recurring issues across the previous studies. First, there is the problem of the so-called โ€œwritten language bias in linguisticsโ€ (Linell, 1982), such that most studies on kata have drawn upon mostly invented sentences or written language data. Secondly, previous studies on kata have focused on the verb itself and have made few efforts on examining the construal of kata as it relates to the argument structure constructions in which the verb appears. Considering what has been pointed out so far, on the basis of contemporary Korean spoken data extracted from Sejong Corpus, the current study aims to establish argument structure constructions focusing on the specification of components, i.e. the subject, the oblique phrase containing the suffix, and kata. Argument structure constructions where kata appears and their components are fully specified are called kata-specific constructions. The objective of this study is to outline the alternations of the argument structure constructions in the physical motion domain, and how and to what extent they are inherited by other semantic domains in accordance with semantic extensions. All the semantic domains are argued to be metaphorically or via constructionalization extended from the physical domain. Further, I aim to examine whether the Principle of Maximized Motivation works or not by virtue of two types of cluster analysis. The first one based on binary coding showed that the metaphorical extension and constructionalization starting from the physical motion domain is not limited to the semantic side, but it also influences how and to what extent the allowed argument structure constructions in the physical motion domain are inherited by other semantic domains. This advocates the Principle of Maximized Motivation. However, the second cluster analysis based on relative frequency showed that abstract motion inherits frequency patterns concerning alternations of argument structure constructions from physical motion to the strongest degree, which weakens the principle

    PARSING AND TAGGING OF BINLINGUAL DICTIONARY

    Get PDF
    Bilingual dictionaries hold great potential as a source of lexical resources for training and testing automated systems for optical character recognition, machine translation, and cross-language information retrieval. In this paper, we describe a system for extracting term lexicons from printed bilingual dictionaries. Our work was divided into three phases - dictionary segmentation, entry tagging, and generation. In segmentation, pages are divided into logical entries based on structural features learned from selected examples. The extracted entries are associated with functional labels and passed to a tagging module which associates linguistic labels with each word or phrase in the entry. The output of the system is a structure that represents the entries from the dictionary. We have used this approach to parse a variety of dictionaries with both Latin and non-Latin alphabets, and demonstrate the results of term lexicon generation for retrieval from a collection of French news stories using English queries. (LAMP-TR-106) (CAR-TR-991) (UMIACS-TR-2003-97

    Natural Language Processing: Emerging Neural Approaches and Applications

    Get PDF
    This Special Issue highlights the most recent research being carried out in the NLP field to discuss relative open issues, with a particular focus on both emerging approaches for language learning, understanding, production, and grounding interactively or autonomously from data in cognitive and neural systems, as well as on their potential or real applications in different domains

    Understanding the structure and meaning of Finnish texts: From corpus creation to deep language modelling

    Get PDF
    Natural Language Processing (NLP) is a cross-disciplinary field combining elements of computer science, artificial intelligence, and linguistics, with the objective of developing means for computational analysis, understanding or generation of human language. The primary aim of this thesis is to advance natural language processing in Finnish by providing more resources and investigating the most effective machine learning based practices for their use. The thesis focuses on NLP topics related to understanding the structure and meaning of written language, mainly concentrating on structural analysis (syntactic parsing) as well as exploring the semantic equivalence of statements that vary in their surface realization (paraphrase modelling). While the new resources presented in the thesis are developed for Finnish, most of the methodological contributions are language-agnostic, and the accompanying papers demonstrate the application and evaluation of these methods across multiple languages. The first set of contributions of this thesis revolve around the development of a state-of-the-art Finnish dependency parsing pipeline. Firstly, the necessary Finnish training data was converted to the Universal Dependencies scheme, integrating Finnish into this important treebank collection and establishing the foundations for Finnish UD parsing. Secondly, a novel word lemmatization method based on deep neural networks is introduced and assessed across a diverse set of over 50 languages. And finally, the overall dependency parsing pipeline is evaluated on a large number of languages, securing top ranks in two competitive shared tasks focused on multilingual dependency parsing. The overall outcome of this line of research is a parsing pipeline reaching state-of-the-art accuracy in Finnish dependency parsing, the parsing numbers obtained with the latest pre-trained language models approaching (at least near) human-level performance. The achievement of large language models in the area of dependency parsingโ€” as well as in many other structured prediction tasksโ€” brings up the hope of the large pre-trained language models genuinely comprehending language, rather than merely relying on simple surface cues. However, datasets designed to measure semantic comprehension in Finnish have been non-existent, or very scarce at the best. To address this limitation, and to reflect the general change of emphasis in the field towards task more semantic in nature, the second part of the thesis shifts its focus to language understanding through an exploration of paraphrase modelling. The second contribution of the thesis is the creation of a novel, large-scale, manually annotated corpus of Finnish paraphrases. A unique aspect of this corpus is that its examples have been manually extracted from two related text documents, with the objective of obtaining non-trivial paraphrase pairs valuable for training and evaluating various language understanding models on paraphrasing. We show that manual paraphrase extraction can yield a corpus featuring pairs that are both notably longer and less lexically overlapping than those produced through automated candidate selection, the current prevailing practice in paraphrase corpus construction. Another distinctive feature in the corpus is that the paraphrases are identified and distributed within their document context, allowing for richer modelling and novel tasks to be defined
    • โ€ฆ
    corecore