Search CORE

4,496 research outputs found

RNN Language Model with Word Clustering and Class-based Output Layer

Author: Johnson Michael T
Liu Jia
Shi Yongzhe
Zhang Wei-Qiang
Publication venue: e-Publications@Marquette
Publication date: 22/07/2013
Field of study

The recurrent neural network language model (RNNLM) has shown significant promise for statistical language modeling. In this work, a new class-based output layer method is introduced to further improve the RNNLM. In this method, word class information is incorporated into the output layer by utilizing the Brown clustering algorithm to estimate a class-based language model. Experimental results show that the new output layer with word clustering not only improves the convergence obviously but also reduces the perplexity and word error rate in large vocabulary continuous speech recognition

epublications@Marquette

Springer - Publisher Connector

Understanding Hidden Memories of Recurrent Neural Networks

Author: Cao Shaozu
Chen Yuanzhe
Li Zhen
Ming Yao
Qu Huamin
Song Yangqiu
Zhang Ruixiang
Publication venue
Publication date: 30/10/2017
Field of study

Recurrent neural networks (RNNs) have been successfully applied to various natural language processing (NLP) tasks and achieved better results than conventional methods. However, the lack of understanding of the mechanisms behind their effectiveness limits further improvements on their architectures. In this paper, we present a visual analytics method for understanding and comparing RNN models for NLP tasks. We propose a technique to explain the function of individual hidden state units based on their expected response to input texts. We then co-cluster hidden state units and words based on the expected response and visualize co-clustering results as memory chips and word clouds to provide more structured knowledge on RNNs' hidden states. We also propose a glyph-based sequence visualization based on aggregate information to analyze the behavior of an RNN's hidden state at the sentence-level. The usability and effectiveness of our method are demonstrated through case studies and reviews from domain experts.Comment: Published at IEEE Conference on Visual Analytics Science and Technology (IEEE VAST 2017

arXiv.org e-Print Archive

Crossref