• 検索結果がありません。

Robust Speech Recognition with Spectral Subtraction in Low SNR

N/A
N/A
Protected

Academic year: 2021

シェア "Robust Speech Recognition with Spectral Subtraction in Low SNR"

Copied!
4
0
0

読み込み中.... (全文を見る)

全文

(1)ROBUSTSPEECH RECOGNITION WITHSPECTRALSU BTRACTION IN LOWSNR Randy Gomez, Akinobu Lee, Hiroshi Saruwatari, Kiyohiro Shikano Graduate School of Information Science Nara Institute of Science and Techonology ,JAPAN E-mail: {randy-g.ri.sawatari.shikano}@is.aist-nara.ac.jp. ABSTRACT Speech recognition in noisy environments is a very diffi­ cuJt task. Jt is is desirabJe to search for parameters that wouJd reJate the speech enhancement technique directJy with the recognizer to optimize the recognition perfor­ mance. Jn this paper, Noise Reduction Rate (NRR) and MeJ Cepstrum Distortion (MeJCD) are investigated when using SpectraJ Subtraction (SS). Under Jow SNR such as OdB,5dB and 10dB, maximizing NRR nor minimizing the MeJCD does not result in a better recognition per­ forrnance. Thus, the conventional SS in which the over­ subtraction parameterαis a function of SNR renders to be ine仔ective in the point-of-view of the recognizer. Our proposed method derives α for SS directly from the training utterances used in creating the Hid­ den Markov ModeJs (HMM) that optimizes the recogni­ tion performance. By superimposing office noise to the SS-denoised noisy speech, we achieved 26.0%and 7.6% for relative increase in word accuracy for the proposed matched and generaJizedαrespectively.. 1. Introduction Spectral Subtraction has been commonJy used in front end of robust speech recognition systems. For an SNR of 25dB, 88.0%accuracy is reported by combining SS with office noise superimposition, and HMM sufficient statis­ tics adaptation using a contantα=2 171. A multispectral approach was also proposed 131 . However,these approaches do not relate SS with the HMMs and it is practical to investigate the e仔ects of the oversubtraction parameter relative to the recognition per­ forrnance of the speech recognizer. The above mentioned SS used in robust speech recogntion addresses the sup­ pression of noise, but on the other hand we do not have any idea on to what extent does noise has to be suppressed and how much distortion is allowed. Meaning to say, SS is implemented soJeJy as a mere speech enhancement technique which is not directly related to the recognition perforrnance. Jn the proposed method, we used NRR and MeJCD in order to study the e仔ects of suppressing noise using SS. and to tailor-fit the noise suppression to the recognizer. We opt to design SS to be related with the model and optimize the recogntion performance because in a speech recognition system we are more interested of recognition rate rather than the quality of the denoised speech. Furtherrnore, we investigated these parameters rela­ tive to the recognition perforrnance during testing and carried out a task-based approach in derivingα.. 2. Background Spectral Subtraction has been applied in robust speech recognition under noisy environments, and is given by Eq. (I ) 1 1 1. IS(kW. =. IY(kWーαID(kW. (1). where IS(kW is the denoised power spectrum. IY(kW and ID(kW are the noisy and noise-only power spec­ trums respectively. α is the oversubtraction parameter that dictates the extent of noise suppression. The oversubtraction parameter can be computed in various ways and one of these is the muJti spectral ap­ proach 13 J . Also, there is a method based on the human auditory system [41. A basic expression of the oversub­ traction parameter is given in Eq.(2) which is a function of SNR [21.. α二αo-;SNR 53NR三25. ω. whereαo 15 a constant. This kind of approach has been effective in the point of view of the human ear. Wìth the same SS used, then it is worth asking whether minimizing noise as perceived by the listener, which is itself a subjective measure, is the same thing as minimizing the mismatch in the point of view of an HMM-based speech recognition system.. 3. Proposed民1ethod Under low SNR,max NRR and min MeJCD do not guar­ antee an optimal recognition perforrnance, thus in our proposed method we resort to a task-based approach in getting the optimal value ofα.. phu 句1・‘ つ'・“.

(2) 同ぺ.. .hu EEZ 85. 6. 8. 10. 12. 14. 16. 18. 4 e。. 6. 8 1. 4. 2 1. 0 1. 8. 6. 4. 2. 四. 4. 10. o 4.. 20. E;. Z3E3 4. 16. 18. and MelCD. Model-based speech recognition is very much dependent on the training data being used. Although SS improves the SNR, distortion is also introduced to the enhanced speech 151. These two parameters give an indirect, yet informative insight of how the recognition performance might be. In this experiment, we investigate these pa­ rameters with the objective of optimizing the recognition performance. NRR given in Eq.(3) is used to characterize the improvement in SNR after SS.. NRR=SNR町山一SNROld. (3). where SNRold and SNRnew are the SNR before and after SS respectively. The higher the NRR means the better the noise is suppressed. In the case of distortion measure (MeICD), lesser MelCD values mean lesser distortion. MelCD is given by Eq.(4).. MelCD. 20 I_� … n 一 (mcie凶)2 1 2 ) :(mc�Tl9) =一一、 lO \1 ム-J ' ln. - t一. 4 ー. 6 -. 8. 10. 12. 14. 16. 18. 田. fj////f//,河け←一a申伊干←?デF戸九-‘--. i笥::ミ;il.r---:--一一一. 20. Figure 1: Max NRR and min MelCD result to max accu­ racy using SNR-based SS (SNR: 25dB). 3.1. NRR. 2 ←. Figure 2: Max NRR and min MelCD do not r巴sult to max accuracy using SNR-based SS (SNR: 5dB). relation between max NRR with max Accuracy at very low SNR (i.e. 10dB, 5dB, OdB). Even min MelCD can­ not be translated into an improvment of the recognition perfo口nance. Short to say, the parameters like NRR due to SS nor MelCD can no longer give as a hint of what the recognition perforrnance might be under low SNR condi­ tions. Figure 2 shows that as the original SNR decreases, the more uncorrelated the three parameters become.. System Implementation. 3.2.. In order to make SS more effective in front end of speech recognition systems, SS should be tailored to optimize the recognition perforrnance. Optimizingαis composed of the following: ・(a) Choosingαin which the word accuracy peaks from the training data. It would make sense if the denoising process is directly derived from the train­ ing data from which the model is created. •. (4). where mc�Tl9 and mci日山 are the Mel Cepstrum c伺ffi­ cients before and after SS respectively. As mentioned earlier, HMM-based speech recogni­ tion system is basically an issue of how acoustically sim ilar the test data are. with the created acoustic model. In other words, the degree of mismatch is very crucial and it can be indirectly attributed with NRR and the degree of distortion(MeICD). At 25dB SNR there exists a cor­ relation between an improved in NRR due to SS and the recognition accuracy,and the same is true with MelCD as shown in Figure 1. This result suggests that the conven­ tional SS in whichαis a function of SNR can be effec­ tively used in a HMM-based speech recognizer as mani­ fested by the correlation of improved NRR and improved word accuracy. On the otherhand, under low SNR, the correlation does not hold true anymore. In fact, there is no more cor-. •. •. (b) From the training data of 20,000 utterances of 260 speakers, 2 utterances per speaker are selected and superimposed with di仔erent types of noise un­ der low SNR conditions OdB, 5dB, 10dB,15dB and 20dB (c) In our experiment we used office, car, booth, and crowd noise and searched for the α-matched condition as shown in Figure 3. (d) We then extended the experiment to incIude several combinations, Iike noise-types are being grouped together as well as groupings of di仔erent SNR conditions to find a generalizedα as shown in Figure 4.. In our experiment,noise-types are defined to be sim­ ilar if their power spectrum are correlated. For example, office noise and car noise are considered similar while office noise and booth noise are dissimilar since the cor relation between offìce and car noise is greater than that of the latter.. - 216-.

(3) ー① -. e Eb 品・. :::21tseJ喧h甘. 口 -34同:! ? - 忌中i. Figure 3: Proposed melhod : α is derived matched to a particular noise-type and a particular SNR condition (matchedα). J�d-I i. za 白書哩EFE E4E智世E邑 oa--aqgas. 届・7 己 wla届 ・. 3z zzsE 4EFZ-E 04・=cgR亘書 • • •. ぷ五lh �. Figure 4・ Proposed method:αis derived by combining noise-types and some SNR conditions (generalizedα). Table 1: Optima1αValues (matchedα/generalizedα) 3.3.. Experimental Condition. The leslset consists of 200 sentences from 46 speakers which are outside from the training data. Initially 4 types of noise are considered: car, office, booth, and crowd noise. The teslsel is then extended to include mall, poster, park, and train noise which are not used in derivingαtn order to test for the degree of robustness to di仔erent types of noise. All in all, the testing procedure consists of 40 testsels. The language model is provided by the IPA dictation loolkil18J. A single PTM 16J speaker-independent acous­ tic HMM is used, which is made by superimposing 25dB office noise to the JNAS database 171. The significance of using a single acoustic model of lhis type instead of matched model is lhe fact that in real application, there exist so many types of noise and it would be impractical to create a matched model for each of these noise. The noise robust speech recognition algorithm with SS and noise superimposition is robusl enough against various noise conditions 171. With this, we will be able to check for the robustness of the proposed melhod under several noise-types using only a single model (JNAS+25dB of­ fice noise ) rather than several matched models. The denoised utterances using the proposed melhod are superimposed with 30dB office noise in order to neu­ lralize the residual noise e仔Ccts after SS 17J. Then these are lesled for recognition performance using JULlUS with 20K-word on Japanese newspaper reading task from JNAS database.. Noise Car Crowd Booth Office. as shown in Figure 4, and “Malched a1pha" refers to that depicted in Figure 3. We use the JNAS+25dB office noise model in all of lhe condilions excepl for the SNR-based SS (matched model ). Tab1e1 1ists the va1ues ofα. The malchedαand gen­ era1ized α are the α va1ues of the proposed method as described in section 3.2. 4.1.. Experiments with other Types ofNoise. Using lhe generalized α, its robustness is further tesled using mall, poster, park, and train noises. These noises are not part of the training data. The resu1t in Figure 7 shows lhat the genera1ized αis robust to a 1imited SNR conditions and various lypes of noise. Tab1e 2 accounts for the relative improvement achieved in using the pro-. 60 •. 90. 60 や 70. 50 マ 40 30. 4. Results and Discussion. 20dB 0.6 /1 .4 1.111.4 2.2 / 1 0.5 1.211. 4. E. .置、. 20. We compare the robustness of the proposed method as oppossed to the SNR-based SS under low SNR. Figures 5 and 6 show the results of lhe basic types of noise. "SNR-based SS" refers to the convenlional SS "SNR­ based SS (matched model )" refers to lhe upper limil of SNR-based SS using malched model, "Generalized al­ pha" refers 10 the robuslness ofαof the proposed melhod. ,. 10 0. Figure 5: Recognition Performance: Car and Crowd Noise. 司J つω.

(4) 9080国?�喝 3。2担4ι 70.. 10-' 0'. ''''''' OdB. 5dB. lOdB H. 15dB 20dB. BQOT ・・SNR-based S� 竃置SNA-based 55 (m副d、edmod叫. C回B. 5dB. 可OdB 15dB OF円CE. 20dB. ・圃Proposed rn割hod (general・zed aI回、剖 Proposed rTM!柑唱d("、atched al回、a). -E・E・-E 叩 叩叩 … 5 E 岨 E 日 t E . r ・E・-E E aE 自 由 r 叩 4園田置量 叩 市 叫旬 四 』量 . 叫 .r 。 E=2. Figure 6: Recognition Petforrnance: 800th and Office Noise. Table 2: Word Accuracy Improvement of the Proposed Method Relative to the Conventional SNR-based SS (generalizedαImatchedα) 5d8 Od8 IOd8 Noise 800th 6.69も1139も 3.5%/3.9% 6.4%17 .1% Car 1.5%13.99も 1.4%/1.8% 1.1 %/1 .6% Train 2.7%/5.5% 0.6%/3.3% 0.7%/1.5% 3.9%/6.1% 0.8%/1.2% 1.1%/2.1% Office Crowd 0.5%10.7% 0.4%10.9% 0.3%/ 1 .0% 1.3%/1.8% 1. 2%/2.49も Park 2.4%14.4% 3.913も113.3% 3.3%/12.3% 1.6%/4.2% Poster 7.6%/20% Mall 4.913も126% 3.3%118. 6%. 60. 6. References [1J. S. 801l,“Suppression of Acoustic Noise in Speech Using Spectral Subtraction", IEEE 7子ω15. Acoustic., Speech, Signal Process., vol 27, pp 113-120, Apr. 1979. [2 J. M. 8erouti, Schwartz and J. Markhoul,“Enhance­ ment of Speech Corruptedby Acoustic Noise", In Proceedings o/IEEE Int. Conf on Acoust., Speech,. Figure 7: Test for robustness ofαusing di仔erent types of nOlse. posed method. Maximizing SNR does not necessarily mean max­ imizing the recognition accuracy using SS under Low SNR conditions (i.e. Od8,5d8, IOd8), then the conven­ tional SS is not very e仔ective in translating the denoising process into an improvement of the recognition petfor­ mance as far as the recognizer is concemed. αreqUlres more than SNR information to e仔ectively translate SS to improving the recognition petforrnanc怠as it shows de­ pendencies with the types of noise. Indeed,it is very dif­ ficult if not impossible to identify the parameters distinc­ tively and formulate an expression to optimize the recog­ nition petformance in an HMM-based system.. 5. CONCLUSION It is shown that the conventional SS is not e仔'ective under low SNR conditions if it is used as a front end speech en hancement for HMM-based speech recognition systems. In this paper, we succesfully tailor fitted SS to optimize the recognition petforrnance by deriving the optimal α from the training data which is directly related with the HMM models. We achieved 26.0% and 7.6% for relative increase in word accuracy for the proposed matched and gener­ aIizedαrespectively as compared with the conventional approaches constantα=2 [7[ and SNR-basedα.. Signal Procs.,. pp.208-21 1, April 1979. 13 J. M. Fujimoto et aI. "Large Vocabulary Speech Recognition Under Real Environments Using Adaptive Sub-band Spectral Subtraction",In Pro­ ceedings o/ICSLP, pp. 1-305-308, 2000.. [4 J. N. V irag, “Speech Enhancement 8ased on Masking Properties of the Human Auditory System" A Mas­ ter's Thesis, Swiss FederaI Institute of Technology 2000. 15 J. R. Gomez, A. Lee, H. Saruwatari, K. Shikano , “Wiener Filtering-based Speech Recognition Under Low SNR" A印刷tical Society 0/ Japan, 2004.. 16J. A. Lee, T. Kawahara, K. Takeda, K. Shikano,“A New Phonetic Tied Mixture Model For Efficient Decoding", ln Proceedings o/ ICASSP, pp. 126911 1272,2000. 17J. Y.. [8J. T. Kawahara et al, “'Free Software Toolkit for Japanese Large Vocabulary Continuous Speech Recognition', /11 Proceedings o/ICSLP, pp. IV-476479, 2000.. Shingo, K. Matsunami, A. 8aba, L. Akinobu, H. Saruwatari,K.Shikano , “Spectral Subtraction In Noisy Environments Applied To Speaker Adapta­ tion 8ased on HMM Sufficient Statistics" In Pro­ ceedings o/ICSLP, pp. 1-1045-1048, 2000.. 00 1ム つω.

(5)

Figure 2:  Max NRR and min MelCD do not r巴sult to max  accuracy using SNR-based SS (SNR:  5dB)
Figure 4・ Proposed method:αis derived by combining  noise-types and some SNR conditions (generalized α)
Table  2:  Word  Accuracy Improvement of  the  Proposed  Method  Relative  to  the  Conventional  SNR-based  SS  (generalized αImatchedα)

参照

関連したドキュメント

In this paper we prove the existence and uniqueness of local and global solutions of a nonlocal Cauchy problem for a class of integrodifferential equation1. The method of semigroups

The main problem upon which most of the geometric topology is based is that of classifying and comparing the various supplementary structures that can be imposed on a

Then it follows immediately from a suitable version of “Hensel’s Lemma” [cf., e.g., the argument of [4], Lemma 2.1] that S may be obtained, as the notation suggests, as the m A

Moreover, in 9, 20, the authors studied the problem of the robust stability of neutral systems with nonlinear parameter perturbations and mixed time-varying neutral and discrete

One important application of the the- orem of Floyd and Oertel is the proof of a theorem of Hatcher [15], which says that incompressible surfaces in an orientable and

The existence of global weak solutions for a class of hemivariational inequalities has been studied by many authors, for example, parabolic type problems in 1–4, and hyperbolic types

Hence, for these classes of orthogonal polynomials analogous results to those reported above hold, namely an additional three-term recursion relation involving shifts in the

Using a clear and straightforward approach, we have obtained and proved inter- esting new binary digit extraction BBP-type formulas for polylogarithm constants.. Some known results