• 検索結果がありません。

Evaluation of invalid input discrimination using BOW for speech-oriented guidance system

N/A
N/A
Protected

Academic year: 2021

シェア "Evaluation of invalid input discrimination using BOW for speech-oriented guidance system"

Copied!
1
0
0

読み込み中.... (全文を見る)

全文

(1)Evaluation of Invalid Input Discrimination Using Bag-of-Words for Speech-Oriented Guidance System. Haruka Majima*, Rafael Torres*, Hiromichi Kawanami*, Sunao Hara**, Tomoko Matsui***, Hiroshi Saruwatari*, Kiyohiro Shikano* *Graduate School of Information Science, Nara Institute of Science and Technology, Japan **Graduate School of Natural Science and Technology, Okayama University, Japan ***Department of Statistical Modeling, The Institute of Statistical Mathematics, Japan Abstract: We investigate a discrimination method for invalid and valid inputs based on machine learning using bag-of-words comprised from automatic speech recognition result as a classification feature. Changing the amount of training data, we elucidate that using 3000 of them (approx. 2 weeks of system inputs) shows enough classification performance.. 5. Features employed for classification 1. BOW (Bag-of-Words) vector consists of frequencies of each word in a vocabulary word list, which is comprised of words from the 10-best ASR candidates of training data. The dimension of BOW vector is determined by the number of words in the word list.. 1. Research background • Automatic speech recognition (ASR) has been widely applied to . dictation, Voice Search, car navigation, etc.. 2. GMM likelihood. Problems. Rejection of invalid inputs is desired.. • Many invalid inputs that the system unnecessarily answers • Invalid inputs decrease the response accuracy. Problems of ASR in real environment. 3. Duration. ASR system Background conversations. is given as the likelihood values of each utterance to six GMMs. The GMMs are trained using six kinds of data, for adults’ valid speech, children’s valid speech, laughter, cough, noise and other invalid speech respectively.. Cough. is the duration of an utterance, determined by voice activity detection of Julius using amplitude and zero crossing.. Cough. Coff coff.. 4. SNR is the signal-to-noise ratio of an utterance.. Noise. 6. Classification methods SVM-based method. 2. Speech-oriented guidance system Takemaru-kun • A real-environment speech-oriented guidance system  placed inside the entrance hall of the Ikoma City North Community Center,  providing guidance to visitors, regarding the center facilities, services, neighboring sightseeing, weather forecast, news, etc.. Processing flow of Takemaru-kun QADB Parallel Decoding. Response Generation. 1 𝑇 min 𝒘 𝒘 + 𝐶 𝒘,𝑏,𝜉 2. background conversations, fuzzy speech, nonsense speech, mistake in VAD (voice activity detection), overflow speech, powerless speech. Speech database. Some tags of invalid inputs are overlapped. Training data are the 15000 data of Nov. 2002 to Dec. 2002 and test data are 14881 data of Aug. 2003.. Detail of system inputs All the inputs to the system. Valid inputs. Valid inputs Inputs that the system should respond Invalid inputs Inputs that the system should not respond Invalid speech Speech, but the contents are of invalid inputs Laughter and cough Non-speech sounds made by human Noise Non-speech sounds that are not made by human. 0. Background conversations Mistake in VAD Cough. Fuzzy Overflow Laughter. Nonsence Powerless Noise. SVM. Output Tool Kernel function. 10-best candidates LIBSVM Radial Basis Function (RBF). ME. Parameter C Tool. 10-2, 10-1,…, 104 Stanford Classifier 2.1.3. • F-measures of SVM are always higher than that of ME. • Both F-measures of SVM and ME are saturated using about 3000 training data. 88% F-measure. Invalid speech 20%. 30%. 40%. 50%. 60%. Detail of invalid inputs. 70%. 80%. 90%. 100%. 4. Proposed method. 86%. 84%. SVM ME. 82% 80% 0. Employ Bag-of-Words as a classification feature Knowledge: ASR results of valid inputs and invalid inputs have different tendency. Examples of valid inputs “Hello.” “Where is the restroom?”. 14881. Results. Noise. 10%. 7607. 15000. Invalid inputs. Laughter and cough. 0%. Test data. 6491 7274. ASR. 50000 Valid inputs. 9509. Total. Engine Julius 4.2 Language model and Acoustic model Takemaru-model. 122939. 106325. Training data. Invalid inputs. Experimental conditions. Detail of invalid inputs 100000. (E, D): training set E: set of class labels D: set of feature represented data points fi(e, d): feature indicator functions. To consider the amount of training data, we experimented the performance of invalid inputs discrimination using Bag-of-words, GMM likelihood, Duration and SNR by SVM or ME.. Unintended inputs with tags of. Inputs. 𝑖=1. 7. Experiments. What is “invalid speech” ?. •. ME (maximum entropy) models – provide a general purpose machine learning technique for classification and prediction. – attempts to maximize the log likelihood. 𝜉𝑖. 𝒘, 𝑏:Parameters of discrimination function 𝐶:Cost parameter of soft margin. 3. Detail of system inputs . 𝑙. subject to 𝑦𝑖 𝒘𝑇 𝒙𝒊 + 𝑏 ≥ 1 − 𝜉𝑖 , 𝜉𝑖 ≥ 0 𝑖 = 1, … , 𝑙. Synthetic Speech Agent Animation. Noise Rejection. •. SVM (support vector machine) – is a supervised learning binary classifier. – estimates a separating hyper-plane with a maximal margin in a higher dimensional space.. Web page URL. Example Selection. ME-based method. Examples of invalid inputs “Aaaaargh.” (noise input) “blah-blah-blah.” (nonsence input). Utilizing linguistic information may improve the discrimination of valid and invalid inputs.. 5000 10000 Amount of training data. 15000. 8. Conclusions • •. We investigated discrimination between invalid and valid spoken inquiries using multiple features, as Bag-of-Words, the likelihood values of GMMs, utterance durations and SNRs. The classification performance was better using larger amount of training data, however it saturated using only 3000 data..

(2)

参照

関連したドキュメント

In particular, the SRS algorithm had a signi fi cantly higher reproducibility and accuracy than the conventional algorithm ( P < 0.01), and a small absolute error and SD of

, Graduate School of Medicine, Kanazawa University of Pathology , Graduate School of Medicine, Kanazawa University Ishikawa Department of Radiology, Graduate School of

*2 Kanazawa University, Institute of Science and Engineering, Faculty of Geosciences and civil Engineering, Associate Professor. *3 Kanazawa University, Graduate School of

Moreover, to obtain the time-decay rate in L q norm of solutions in Theorem 1.1, we first find the Green’s matrix for the linear system using the Fourier transform and then obtain

Nonlinear systems of the form 1.1 arise in many applications such as the discrete models of steady-state equations of reaction–diffusion equations see 1–6, the discrete analogue of

TOPSØE, Some Inequalities for Information Divergence and Related Measures of Discrimination, IEEE Trans. TOUSSAINT, Sharper lower bounds for discrimination information in terms

† Institute of Computer Science, Czech Academy of Sciences, Prague, and School of Business Administration, Anglo-American University, Prague, Czech

Henson, “Global dynamics of some periodically forced, monotone difference equations,” Journal of Di ff erence Equations and Applications, vol. Henson, “A periodically