Robust Speech Recognition with Spectral Subtraction in Low SNR
全文
(2) 同ぺ.. .hu EEZ 85. 6. 8. 10. 12. 14. 16. 18. 4 e。. 6. 8 1. 4. 2 1. 0 1. 8. 6. 4. 2. 四. 4. 10. o 4.. 20. E;. Z3E3 4. 16. 18. and MelCD. Model-based speech recognition is very much dependent on the training data being used. Although SS improves the SNR, distortion is also introduced to the enhanced speech 151. These two parameters give an indirect, yet informative insight of how the recognition performance might be. In this experiment, we investigate these pa rameters with the objective of optimizing the recognition performance. NRR given in Eq.(3) is used to characterize the improvement in SNR after SS.. NRR=SNR町山一SNROld. (3). where SNRold and SNRnew are the SNR before and after SS respectively. The higher the NRR means the better the noise is suppressed. In the case of distortion measure (MeICD), lesser MelCD values mean lesser distortion. MelCD is given by Eq.(4).. MelCD. 20 I_� … n 一 (mcie凶)2 1 2 ) :(mc�Tl9) =一一、 lO \1 ム-J ' ln. - t一. 4 ー. 6 -. 8. 10. 12. 14. 16. 18. 田. fj////f//,河け←一a申伊干←?デF戸九-‘--. i笥::ミ;il.r---:--一一一. 20. Figure 1: Max NRR and min MelCD result to max accu racy using SNR-based SS (SNR: 25dB). 3.1. NRR. 2 ←. Figure 2: Max NRR and min MelCD do not r巴sult to max accuracy using SNR-based SS (SNR: 5dB). relation between max NRR with max Accuracy at very low SNR (i.e. 10dB, 5dB, OdB). Even min MelCD can not be translated into an improvment of the recognition perfo口nance. Short to say, the parameters like NRR due to SS nor MelCD can no longer give as a hint of what the recognition perforrnance might be under low SNR condi tions. Figure 2 shows that as the original SNR decreases, the more uncorrelated the three parameters become.. System Implementation. 3.2.. In order to make SS more effective in front end of speech recognition systems, SS should be tailored to optimize the recognition perforrnance. Optimizingαis composed of the following: ・(a) Choosingαin which the word accuracy peaks from the training data. It would make sense if the denoising process is directly derived from the train ing data from which the model is created. •. (4). where mc�Tl9 and mci日山 are the Mel Cepstrum c伺ffi cients before and after SS respectively. As mentioned earlier, HMM-based speech recogni tion system is basically an issue of how acoustically sim ilar the test data are. with the created acoustic model. In other words, the degree of mismatch is very crucial and it can be indirectly attributed with NRR and the degree of distortion(MeICD). At 25dB SNR there exists a cor relation between an improved in NRR due to SS and the recognition accuracy,and the same is true with MelCD as shown in Figure 1. This result suggests that the conven tional SS in whichαis a function of SNR can be effec tively used in a HMM-based speech recognizer as mani fested by the correlation of improved NRR and improved word accuracy. On the otherhand, under low SNR, the correlation does not hold true anymore. In fact, there is no more cor-. •. •. (b) From the training data of 20,000 utterances of 260 speakers, 2 utterances per speaker are selected and superimposed with di仔erent types of noise un der low SNR conditions OdB, 5dB, 10dB,15dB and 20dB (c) In our experiment we used office, car, booth, and crowd noise and searched for the α-matched condition as shown in Figure 3. (d) We then extended the experiment to incIude several combinations, Iike noise-types are being grouped together as well as groupings of di仔erent SNR conditions to find a generalizedα as shown in Figure 4.. In our experiment,noise-types are defined to be sim ilar if their power spectrum are correlated. For example, office noise and car noise are considered similar while office noise and booth noise are dissimilar since the cor relation between offìce and car noise is greater than that of the latter.. - 216-.
(3) ー① -. e Eb 品・. :::21tseJ喧h甘. 口 -34同:! ? - 忌中i. Figure 3: Proposed melhod : α is derived matched to a particular noise-type and a particular SNR condition (matchedα). J�d-I i. za 白書哩EFE E4E智世E邑 oa--aqgas. 届・7 己 wla届 ・. 3z zzsE 4EFZ-E 04・=cgR亘書 • • •. ぷ五lh �. Figure 4・ Proposed method:αis derived by combining noise-types and some SNR conditions (generalizedα). Table 1: Optima1αValues (matchedα/generalizedα) 3.3.. Experimental Condition. The leslset consists of 200 sentences from 46 speakers which are outside from the training data. Initially 4 types of noise are considered: car, office, booth, and crowd noise. The teslsel is then extended to include mall, poster, park, and train noise which are not used in derivingαtn order to test for the degree of robustness to di仔erent types of noise. All in all, the testing procedure consists of 40 testsels. The language model is provided by the IPA dictation loolkil18J. A single PTM 16J speaker-independent acous tic HMM is used, which is made by superimposing 25dB office noise to the JNAS database 171. The significance of using a single acoustic model of lhis type instead of matched model is lhe fact that in real application, there exist so many types of noise and it would be impractical to create a matched model for each of these noise. The noise robust speech recognition algorithm with SS and noise superimposition is robusl enough against various noise conditions 171. With this, we will be able to check for the robustness of the proposed melhod under several noise-types using only a single model (JNAS+25dB of fice noise ) rather than several matched models. The denoised utterances using the proposed melhod are superimposed with 30dB office noise in order to neu lralize the residual noise e仔Ccts after SS 17J. Then these are lesled for recognition performance using JULlUS with 20K-word on Japanese newspaper reading task from JNAS database.. Noise Car Crowd Booth Office. as shown in Figure 4, and “Malched a1pha" refers to that depicted in Figure 3. We use the JNAS+25dB office noise model in all of lhe condilions excepl for the SNR-based SS (matched model ). Tab1e1 1ists the va1ues ofα. The malchedαand gen era1ized α are the α va1ues of the proposed method as described in section 3.2. 4.1.. Experiments with other Types ofNoise. Using lhe generalized α, its robustness is further tesled using mall, poster, park, and train noises. These noises are not part of the training data. The resu1t in Figure 7 shows lhat the genera1ized αis robust to a 1imited SNR conditions and various lypes of noise. Tab1e 2 accounts for the relative improvement achieved in using the pro-. 60 •. 90. 60 や 70. 50 マ 40 30. 4. Results and Discussion. 20dB 0.6 /1 .4 1.111.4 2.2 / 1 0.5 1.211. 4. E. .置、. 20. We compare the robustness of the proposed method as oppossed to the SNR-based SS under low SNR. Figures 5 and 6 show the results of lhe basic types of noise. "SNR-based SS" refers to the convenlional SS "SNR based SS (matched model )" refers to lhe upper limil of SNR-based SS using malched model, "Generalized al pha" refers 10 the robuslness ofαof the proposed melhod. ,. 10 0. Figure 5: Recognition Performance: Car and Crowd Noise. 司J つω.
(4) 9080国?�喝 3。2担4ι 70.. 10-' 0'. ''''''' OdB. 5dB. lOdB H. 15dB 20dB. BQOT ・・SNR-based S� 竃置SNA-based 55 (m副d、edmod叫. C回B. 5dB. 可OdB 15dB OF円CE. 20dB. ・圃Proposed rn割hod (general・zed aI回、剖 Proposed rTM!柑唱d("、atched al回、a). -E・E・-E 叩 叩叩 … 5 E 岨 E 日 t E . r ・E・-E E aE 自 由 r 叩 4園田置量 叩 市 叫旬 四 』量 . 叫 .r 。 E=2. Figure 6: Recognition Petforrnance: 800th and Office Noise. Table 2: Word Accuracy Improvement of the Proposed Method Relative to the Conventional SNR-based SS (generalizedαImatchedα) 5d8 Od8 IOd8 Noise 800th 6.69も1139も 3.5%/3.9% 6.4%17 .1% Car 1.5%13.99も 1.4%/1.8% 1.1 %/1 .6% Train 2.7%/5.5% 0.6%/3.3% 0.7%/1.5% 3.9%/6.1% 0.8%/1.2% 1.1%/2.1% Office Crowd 0.5%10.7% 0.4%10.9% 0.3%/ 1 .0% 1.3%/1.8% 1. 2%/2.49も Park 2.4%14.4% 3.913も113.3% 3.3%/12.3% 1.6%/4.2% Poster 7.6%/20% Mall 4.913も126% 3.3%118. 6%. 60. 6. References [1J. S. 801l,“Suppression of Acoustic Noise in Speech Using Spectral Subtraction", IEEE 7子ω15. Acoustic., Speech, Signal Process., vol 27, pp 113-120, Apr. 1979. [2 J. M. 8erouti, Schwartz and J. Markhoul,“Enhance ment of Speech Corruptedby Acoustic Noise", In Proceedings o/IEEE Int. Conf on Acoust., Speech,. Figure 7: Test for robustness ofαusing di仔erent types of nOlse. posed method. Maximizing SNR does not necessarily mean max imizing the recognition accuracy using SS under Low SNR conditions (i.e. Od8,5d8, IOd8), then the conven tional SS is not very e仔ective in translating the denoising process into an improvement of the recognition petfor mance as far as the recognizer is concemed. αreqUlres more than SNR information to e仔ectively translate SS to improving the recognition petforrnanc怠as it shows de pendencies with the types of noise. Indeed,it is very dif ficult if not impossible to identify the parameters distinc tively and formulate an expression to optimize the recog nition petformance in an HMM-based system.. 5. CONCLUSION It is shown that the conventional SS is not e仔'ective under low SNR conditions if it is used as a front end speech en hancement for HMM-based speech recognition systems. In this paper, we succesfully tailor fitted SS to optimize the recognition petforrnance by deriving the optimal α from the training data which is directly related with the HMM models. We achieved 26.0% and 7.6% for relative increase in word accuracy for the proposed matched and gener aIizedαrespectively as compared with the conventional approaches constantα=2 [7[ and SNR-basedα.. Signal Procs.,. pp.208-21 1, April 1979. 13 J. M. Fujimoto et aI. "Large Vocabulary Speech Recognition Under Real Environments Using Adaptive Sub-band Spectral Subtraction",In Pro ceedings o/ICSLP, pp. 1-305-308, 2000.. [4 J. N. V irag, “Speech Enhancement 8ased on Masking Properties of the Human Auditory System" A Mas ter's Thesis, Swiss FederaI Institute of Technology 2000. 15 J. R. Gomez, A. Lee, H. Saruwatari, K. Shikano , “Wiener Filtering-based Speech Recognition Under Low SNR" A印刷tical Society 0/ Japan, 2004.. 16J. A. Lee, T. Kawahara, K. Takeda, K. Shikano,“A New Phonetic Tied Mixture Model For Efficient Decoding", ln Proceedings o/ ICASSP, pp. 126911 1272,2000. 17J. Y.. [8J. T. Kawahara et al, “'Free Software Toolkit for Japanese Large Vocabulary Continuous Speech Recognition', /11 Proceedings o/ICSLP, pp. IV-476479, 2000.. Shingo, K. Matsunami, A. 8aba, L. Akinobu, H. Saruwatari,K.Shikano , “Spectral Subtraction In Noisy Environments Applied To Speaker Adapta tion 8ased on HMM Sufficient Statistics" In Pro ceedings o/ICSLP, pp. 1-1045-1048, 2000.. 00 1ム つω.
(5)
図
関連したドキュメント
In this paper we prove the existence and uniqueness of local and global solutions of a nonlocal Cauchy problem for a class of integrodifferential equation1. The method of semigroups
The main problem upon which most of the geometric topology is based is that of classifying and comparing the various supplementary structures that can be imposed on a
Then it follows immediately from a suitable version of “Hensel’s Lemma” [cf., e.g., the argument of [4], Lemma 2.1] that S may be obtained, as the notation suggests, as the m A
Moreover, in 9, 20, the authors studied the problem of the robust stability of neutral systems with nonlinear parameter perturbations and mixed time-varying neutral and discrete
One important application of the the- orem of Floyd and Oertel is the proof of a theorem of Hatcher [15], which says that incompressible surfaces in an orientable and
The existence of global weak solutions for a class of hemivariational inequalities has been studied by many authors, for example, parabolic type problems in 1–4, and hyperbolic types
Hence, for these classes of orthogonal polynomials analogous results to those reported above hold, namely an additional three-term recursion relation involving shifts in the
Using a clear and straightforward approach, we have obtained and proved inter- esting new binary digit extraction BBP-type formulas for polylogarithm constants.. Some known results