• 検索結果がありません。

Input Tensor

ドキュメント内 JAIST Repository https://dspace.jaist.ac.jp/ (ページ 52-57)

4.3 Proposed Model

4.3.3 Input Tensor

Input tensor tries to learn the multiple features and converts each flavor-feature to a dense word representation. Dense word representations are concatenated together in order to construct into two Advanced word embeddings:

• Advanced continuous word embedding vi = [ri;ei;li] is constructed by three sub-vectors: the continuous word-level embedding ri ∈Rdword, the charac-ter attention embedding ei ∈Rl, where l is the length of the filter of wide convolutions, the lexicon embedding li ∈Rdscore where dscore is list of senti-ment scores for that word in lexicon datasets. Advanced continuous word embedding contains semantic and is improved information from LexW2Vs and CharAVs.

• Advanced dependency-based word embedding di = [dei, ei;li] is also built by three sub-vectors: the dependency-based word embedding dei ∈Rdword, the character attention embedding ei and the lexicon embedding li. The advanced dependency-based word embedding contains syntactic contexts and is enhanced information from LexW2Vs and CharAVs.

These advanced embeddings deal with three main problems:

• Sentences have any different size.

• Important information of characters that can appear at any position in a word is extracted.

• The interactive knowledge between flavor-features are captured via Bi-CGRNNet in order to produce a Tweet-specific representation.

Different from LexW2Vs and DependencyW2Vs, to learn CharAVs, we use a DeepCNN to capture the morphology of each word via character embeddings as Sub-section 4.3.3. The output of DeepCNN is a single real-valued feature. In the next Sub-sections, we introduce methods to build LexW2Vs, CharAVs and DependencyW2Vs features.

Word/Character Embeddings

The different kinds of embeddings are constructed by using a fixed-sized word vocabulary Vword and a fixed-sized character vocabulary Vchar. Given a word wi is composed from characters{c1, c2, ..., cM}, the character-level embeddings are encoded by column vectorsuiin the embedding matrixWchar ∈Rdchar×|Vchar|, where Vchar is the size of the character vocabulary. For continuous word-level embedding rword, each word wi ∈ Vword, where Vword is a fixed-sized word vocabulary. The word embedding matrix is simply WE ∈ Rd×Vword, where d is the dimension of the word embeddings. Each word embedding wi is mapped by using a pre-trained word embeddings. The character embeddings are constructed by initialization randomly.

Lexicon Embeddings (LexW2Vs)

Semantic embeddings ignore the sentiment polarity of words in the sentence and map words with similar semantic context but opposite sentiment polarity. To inte-grate the sentiment polarity for words in the sentence, LexW2Vs are constructed by taking scores from various lexicon datasets. Sentiment lexicons are valuable resources that can be considered much for building embeddings for deep learn-ing models. The sentiment lexicons present the different states of each word in different contexts which can be trained by using neural networks.

In lexicon datasets, each word contains key-value pairs in which the key is a word, and the value is a list of sentiment scores for that word.

For each word wi ∈ Vword, where Vword is a fixed-sized word vocabulary, a lexicon embedding is constructed by concatenating all of the scores among lexicon datasets with respect to wi. If wi does not exist in a certain dataset, 0 value is substituted. The lexicon embedding is a form of a vector li ∈Rdscore, where dscore is the total number of scores across all lexicon datasets. We use seven lexicon datasets for building LexW2Vs:

• Bing Liu Opinion Lexicon [Hu and Liu, 2004].

• NRC Hashtag Sentiment Lexicon [Mohammad et al., 2013].

• Sentiment140 Lexicon [Go et al., 2009].

• NRC Sentiment140 Lexicon [Go et al., 2009].

• MaxDiff Twitter Sentiment Lexicon [Kiritchenko et al., 2014b].

• National Research Council Canada (NRC) Hash-tag Affirmative and Negated Context Sentiment Lexicon [Kiritchenko et al., 2014b].

• Large-Scale Twitter-specific Sentiment Lexicon [Tang et al., 2014a].

The LexW2Vs capture the different sentiments of words in tweets and is worth incorporating to improve the coverage. Table 4.2 illustrates the type of words for each dataset.

Lexicon dataset The type of words Bing Liu Opinion Lexicon Sentiment adjective words

NRC Hashtag Sentiment Lexicon Hashtag emotion words & Hashtag topic words

Sentiment140 Lexicon Emoticons & Sentiment words

NRC Sentiment140 Lexicon Affirmative context words & Senti-ment140 Negated Context words MaxDiff Twitter Sentiment

Lexi-con

Twitter sentiment words Hashtag Affirmative and Negated

Context Sentiment Lexicon

Hashtag affirmative words & Negated contextual words

Large-Scale Twitter-Specific Sen-timent Lexicon

Colloquial words & Emoticons

Table 4.2: The types of words in lexicon dataset.

Lexicon-based features are considered as lexicon embeddings/ sentiment em-beddings because they are a feature input of deep learning model and describe the properties of words in tweets. Each word in each lexicon datasets has many values that can be built by training a neural network. The deep learning model uses this input for calculating a computational graph (weight matrix) that describe relatedness among words (n-gram order).

Character Attention Vectors (CharAVs)

In this sub-section, we introduce a neural network to develop a Character attention vector (CharAVs feature) which represents the morphology of each word. Figure 4.3 describes DeepCNN with two wide convolutionsto form a Character attention vector. The first convolution produces a fixed-size character feature vectornamed n-gram features by extracting local features around each character window of the given word and using a max pooling over vertical character windows. The second convolution retrieves thefixed-size character feature vector (Character feature ma-trix) and transforms the representation to yield a Character attention vector. This method is most relevant to our work [Nguyen and Nguyen, 2018]. However, we distinguish their work by improving the second convolution. The second convolu-tion transforms the representaconvolu-tion by performing max pooling on each row of the

Figure 4.3: DeepCNN for the sequence of character embeddings of a word. For example with one region size is 2 and Four feature maps in the first convolution and one region size is 3 with three feature maps in the second convolution. The CharAVs is then created by performing max pooling on each row of the attention matrix.

Character feature matrix instead of each column. The purpose of this method is to attend on the highestn-gram feature in order to transform this n-gram features at previous level into representation at a focused abstract level and to produce attention over the best feature vector.

Additionally, DeepCNN is constructed from two wide convolutions which can learn to recognize specific n-grams at every position in a word and allow features to be extracted independently of these positions in the word. These features maintain the order and relative positions of characters and are formed at a higher abstract level. Character attention vector has two advantages: 1) This model could adaptively assign an importance score to each piece of word embeddings concerning its semantic relatedness. 2) Another advantage is that this attention model is differentiated so that it could be easily trained together with other components in an end-to-end fashion. In the next sub-section, we introduce the background structure of Convoluational Neural Network (CNN) with Wide convolution.

The Effect of Convolutional Neural Network: The convolution is an op-eration between a vector of weights m ∈ Rm and a vector of inputs viewed as a sequencew∈Rw. The vectormis the filter of the convolution. Concretely, s as an input word andsi ∈R is a single feature value associated with thei−thcharacter in the word. The convolution has a filter vector m and take the dot product of filterm with each m-grams in the sequence of characterssi ∈R of a word in order to obtain a sequencec:

cj =mTsj−m+1:j (4.1)

Based on Equation 1, we have two types of convolutions that depend on the range of the index j. The narrow type requires that s ≥ m and produce a sequence c ∈Rs−m+1. The wide type does not require on s or m and produce a sequence c ∈Rs+m−1. Out-of-range input values si where i < 1 or i > s are taken to be zero. The result of the narrow convolution is a subsequence of the result of wide convolution.

Wide Convolution: Given a word wi composed ofM characters{c1, c2, ..., cM}, we take a character embedding ui ∈Rdchar for each character ci and construct a character matrix Wchar ∈Rdchar × |Vchar| as following Equation 4.2:

Wchar =

| | | u1 ... uM

| | |

 (4.2)

.

The values of the embeddings ui are parameters that are optimized during training. The trained weights in the filter m correspond to a feature detector which learns to recognize a specific class of n-grams, where n ≤ m, and m is the width of the filter. The use of a wide convolution has more advantages than a narrow convolution because a wide convolution ensures that all weights of the filter m reach the whole characters of a word at the margins. Besides, the wide convolution guarantees that the filterm always produces a valid non-empty result c, independently of the width of m and the sequence length s. Therefore, this is particularly significant when the width of the filter m is set from 7 to 14 to represent a word. The resulting matrix has dimension d×(s+m−1).

Dependency-based Word Vectors (DependencyW2Vs)

To construct context embeddings, we use the idea of [Levy and Goldberg, 2014] to derive syntactic contexts based on the syntactic relations of a word. Most previous works on neural word embeddings take the contexts of a word by computing linear-context words that precede and follow the target word. However, these contexts can be exploited similar by generalizing the SKIP-GRAM model. The model for

learning Dependency-based Word Vectors is improved from SKIP-GRAM model in which the linear bag of words contexts are replaced with arbitrary word contexts from a dependency tree. Syntactic contexts are derived from produced dependency parse-trees. Specifically, the bag-of-words in the SKIP-GRAM model yield broad topical similarities, while the dependency-based contexts yield more functional similarities of a cohyponym nature. In the SKIP-GRAM model, the contexts of a word are the words surrounding it in the text. However, there is a limitation of SKIP-GRAM word embeddings: Contexts no need to correspond to all of the words and the number of context-types maybe larger than the number of word-types. Therefore, dependency-based contexts capture more information than bag-of-words contexts. In Figure 4.4, the contexts are extracted for each word in the sentence, and the contexts of a word are derived from syntactic relations of a word in the sentence. For parsing syntactic dependencies, we use a parser from [Goldberg and Nivre, 2013] for Stanford dependencies and the corpus are tagged with parts-of-speech using Stanford parser 1.

After parsing each sentence, we consider word context as Figure 4.4: For a target wordw with modifiersm1, m2, ..., mnand a head h, we form the contexts as (m1, lbl1), ...,(mn, lbln),(h, lbl−1h ), where lbl is the type of the dependency relation between the head and the modifier, lbl−1 is used to marking the inverse-relation.

The advantages of syntactic dependencies are inclusive and more focused than bag-of-words. Besides, they can capture relations that out-of-reach with small windows and filter out contexts that are not directly related to the target word.

For example, Australian is not used as the context for discovers. We have more focused embeddings that capture more functional and less topical similarity. As such, DependencyW2Vs contains syntactic information captured via dependency trees.

4.3.4 Contextual Gated Recurrent Neural Network

ドキュメント内 JAIST Repository https://dspace.jaist.ac.jp/ (ページ 52-57)

関連したドキュメント