ORCID
MSU Affiliation
Institute for Signal and Information Processing; Department of Electrical and Computer Engineering; James Worth Bagley College of Engineering
Creation Date
7-21-2001
Creation Date
2026-07-21
Abstract
In this document we describe the features of the baseline system to be used in the Distributed Speech Recognition (DSR) front end large vocabulary continuous speech recognition (LVCSR) evaluations being conducted by the Aurora Working Group of the European Telecommunications Standards Institute (ETSI). The objective of these evaluations is to determine the robustness of different front ends for use in client/server type telecommunications applications. As such, our experiments are designed to test the following focus conditions on the DARPA Wall Street Journal (WSJ0) corpus using a 5000-word closed-loop vocabulary and a bigram language model:
- Additive Noise: six noise conditions collected from street traffic, train stations, cars, babble, restaurants and airports will be digitally added to the speech data to simulate degradations in the signal-to-noise ratio of the channel.
- Sample Frequency Reduction: the reduction in accuracy due to decreasing the sample frequency from 16kHz to 8kHz will be calibrated.
- Microphone Variation: performance for both microphone conditions contained in the WSJ0 corpus will be analyzed.
- Compression: degradations due to data compression of the feature vectors will be evaluated.
- Model Mismatch: the degradation due to a mismatch between training an evaluation conditions will be calibrated.
The first step in this project was to design a baseline system that will provide a stable point of comparison with state-of-the-art WSJ0 systems. This baseline system was trained on 7,138 clean utterances from the SI-84 WSJ0 training set. These utterances were parameterized using a standard mel frequency scaled cepstral coefficient (MFCC) front end that uses 12 FFT-derived cepstral coefficients, log energy, and the first and second derivatives of these parameters. From these features, state-tied cross-word triphone acoustic models with 16 Gaussian mixtures per state were generated. The lexicon was extracted from the CMU dictionary (version 0.6) with some local additions to cover the 5000 word vocabulary. Recognition was performed using a single pass dynamic programming-based search guided by a standard backoff bigram language model.
The NIST Nov’92 dev test and evaluation sets were used for our evaluations. Our initial experiments used a 300 utterance subset of the 1206 utterance dev test set, and also used a set of pruning thresholds and scaling parameters based on our Hub-5E conversational speech evaluation system. The baseline system yielded a word error rate (WER) of 10.8% on the dev test set. Tuning various parameters decreased the error rate to 10.1% on the dev test subset and 8.3% on the evaluation set. State-of-the-art systems developed by other sites such as the HTK Group at Cambridge University have achieved a 6.9% WER on the same task. The principal difference between their system and the system presented here seems to be a proprietary lexicon which was developed to improve performance on the WSJ task.
The next step in this project will involve the completion of benchmarks using the ETSI standard front end. Since this front end is very similar to the MFCC front end described above, its performance is expected to be consistent with our previous results. We will evaluate a total of 112 conditions: 6 noise types + clean data x 2 compression types x 2 sampling rates x 2 microphone conditions x 2 training conditions. The system resulting from these experiments will serve as the baseline for all subsequent experiments.
Publication Date
Summer 7-21-2001
Creative Commons License

This work is licensed under a Creative Commons Attribution 4.0 International License.
Recommended Citation
Parihar, Naveen and Picone, Joseph, "Aurora Working Group: DSR Front End LVCSR Evaluation — Baseline Recognition System Description" (2001). Publications. 818.
https://scholarsjunction.msstate.edu/works_publications/818