ORCID
MSU Affiliation
Institute for Signal and Information Processing; Department of Electrical and Computer Engineering; James Worth Bagley College of Engineering
Creation Date
12-6-2002
Creation Date
2026-07-21
Abstract
In this document we describe the Distributed Speech Recognition (DSR) front end large vocabulary continuous speech recognition (LVCSR) evaluations being conducted by the Aurora Working Group of the European Telecommunications Standards Institute (ETSI). The objective of these evaluations is to determine the robustness of different front ends for use in client/server type telecommunications applications. The 5000-word closed-loop vocabulary task based on the DARPA Wall Street Journal (WSJ0) Corpus was chosen for these evaluations. The experiments were designed to test the following focus conditions:
- Additive Noise: six noise conditions collected from street traffic, train stations, cars, babble, restaurants and airports were digitally added to the speech data to simulate degradations in the signal-to-noise ratio of the channel.
- Sample Frequency Reduction: the reduction in accuracy due to decreasing the sample frequency from 16kHz to 8kHz was calibrated.
- Microphone Variation: performance for two microphone conditions (Sennheiser and Second Microphone) contained in the WSJ0 corpus was analyzed.
- Compression: degradations due to data compression of the feature vectors was evaluated.
- Model Mismatch: the degradation due to a mismatch between training and evaluation conditions was calibrated.
- Utterance Detection: performance improvement due to end pointing the WSJ0 corpus.
The first step in this project was to design a baseline system that provides a stable point of comparison with state-of-the-art WSJ0 systems. This baseline system was trained on 7,138 clean utterances from the SI-84 WSJ0 training set. These utterances were parameterized using a standard mel frequency scaled cepstral coefficient (MFCC) front end that uses 12 FFT-derived cepstral coefficients, log energy, and the first and second derivatives of these parameters. From these features, state-tied cross-word triphone acoustic models with 16 Gaussian mixtures per state were generated. The lexicon was extracted from the CMU dictionary (version 0.6) with some local additions to cover the 5000 word vocabulary. Recognition was performed using a single pass dynamic programming-based search guided by a standard backoff bigram language model.
The NIST Nov’92 dev test and evaluation sets were used for our evaluations. Our initial experiments used a 330 utterance subset of the 1206 utterance dev test set, and also used a set of pruning thresholds and scaling parameters based on our Hub-5E conversational speech evaluation system. The baseline system yielded a word error rate (WER) of 10.8% on the dev test set. Tuning various parameters decreased the error rate to 10.1% on the dev test subset and 8.3% on the evaluation set. State-of-the-art systems developed by other sites such as the HTK Group at Cambridge University have achieved a 6.9% WER on the same task. The principal difference between their system and the system presented here seems to be a proprietary lexicon which was developed to improve performance on the WSJ task.
The next step in this project involved the completion of benchmarks using the ETSI standard front end. Since this front end is very similar to the MFCC front end described above, its performance was expected to be consistent with our previous results. We evaluated a total of 154 conditions that involved 7 noise types, 2 compression types, 3 training sets, 2 sample rates, 2 compression types and utterance detection. The results from these experiments are summarized in this report
Publication Date
Fall 12-6-2002
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 International License.
Recommended Citation
Parihar, Naveen and Picone, Joseph, "Aurora Working Group: DSR Front End LVCSR Evaluation AU/384/02" (2002). Publications. 819.
https://scholarsjunction.msstate.edu/works_publications/819