CA2013263A1 - Rejection method for speech recognition - Google Patents
Rejection method for speech recognitionInfo
- Publication number
- CA2013263A1 CA2013263A1 CA 2013263 CA2013263A CA2013263A1 CA 2013263 A1 CA2013263 A1 CA 2013263A1 CA 2013263 CA2013263 CA 2013263 CA 2013263 A CA2013263 A CA 2013263A CA 2013263 A1 CA2013263 A1 CA 2013263A1
- Authority
- CA
- Canada
- Prior art keywords
- vocabulary
- utterances
- equalized
- templates
- representations
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/04—Segmentation; Word boundary detection
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/10—Speech classification or search using distance or distortion measures between unknown speech and reference templates
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/12—Speech classification or search using dynamic programming techniques, e.g. dynamic time warping [DTW]
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
- G10L25/18—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
- G10L25/24—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being the cepstrum
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Complex Calculations (AREA)
- Machine Translation (AREA)
- Character Discrimination (AREA)
Abstract
A speech recognizer, for recognizing unknown utterances in isolated-word small-vocabulary speech has improved rejection of out of vocabulary utterances. Both a usual spectral representation including a dynamic component and an equalized representation are used to match unknown utterances to templates for in-vocabulary words. In a preferred embodiment, the representations are mel-based cepstral with dynamic components being signed vector differences between pairs of primary cepstra. The equalized representation being the signed difference of each cepstral coefficient less an average value of the coefficients.
Factors are generated from the ordered lists of templates to determine the probability of the top choice being a correct acceptance, with different methods being applied when the usual and equalized representations yield a different match.
For additional enhancement, the rejection method may use templates corresponding to non-vocabulary utterances or decoys. If the top choice corresponds to a decoy, the input is rejected.
Factors are generated from the ordered lists of templates to determine the probability of the top choice being a correct acceptance, with different methods being applied when the usual and equalized representations yield a different match.
For additional enhancement, the rejection method may use templates corresponding to non-vocabulary utterances or decoys. If the top choice corresponds to a decoy, the input is rejected.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CA 2013263 CA2013263C (en) | 1990-03-28 | 1990-03-28 | Rejection method for speech recognition |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CA 2013263 CA2013263C (en) | 1990-03-28 | 1990-03-28 | Rejection method for speech recognition |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| CA2013263A1 true CA2013263A1 (en) | 1991-09-28 |
| CA2013263C CA2013263C (en) | 1995-09-05 |
Family
ID=4144624
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CA 2013263 Expired - Fee Related CA2013263C (en) | 1990-03-28 | 1990-03-28 | Rejection method for speech recognition |
Country Status (1)
| Country | Link |
|---|---|
| CA (1) | CA2013263C (en) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111583907A (en) * | 2020-04-15 | 2020-08-25 | 北京小米松果电子有限公司 | Information processing method, device and storage medium |
-
1990
- 1990-03-28 CA CA 2013263 patent/CA2013263C/en not_active Expired - Fee Related
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111583907A (en) * | 2020-04-15 | 2020-08-25 | 北京小米松果电子有限公司 | Information processing method, device and storage medium |
| CN111583907B (en) * | 2020-04-15 | 2023-08-15 | 北京小米松果电子有限公司 | Information processing method, device and storage medium |
Also Published As
| Publication number | Publication date |
|---|---|
| CA2013263C (en) | 1995-09-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| FI117954B (en) | System for verifying a speaker | |
| US6611801B2 (en) | Gain and noise matching for speech recognition | |
| US5950157A (en) | Method for establishing handset-dependent normalizing models for speaker recognition | |
| US5613037A (en) | Rejection of non-digit strings for connected digit speech recognition | |
| US6058363A (en) | Method and system for speaker-independent recognition of user-defined phrases | |
| US6922668B1 (en) | Speaker recognition | |
| EP0625775A1 (en) | Speech recognition system with improved rejection of words and sounds not contained in the system vocabulary | |
| US6868381B1 (en) | Method and apparatus providing hypothesis driven speech modelling for use in speech recognition | |
| US5758021A (en) | Speech recognition combining dynamic programming and neural network techniques | |
| US5963904A (en) | Phoneme dividing method using multilevel neural network | |
| JPH0876785A (en) | Voice recognition device | |
| KR100698811B1 (en) | Voice recognition rejection method | |
| US4937871A (en) | Speech recognition device | |
| Fukuda et al. | Orthogonalized distinctive phonetic feature extraction for noise-robust automatic speech recognition | |
| EP1005019A3 (en) | Segment-based similarity measurement method for speech recognition | |
| JP2003535366A (en) | Rank-based rejection for pattern classification | |
| AU646060B2 (en) | Adaptation of reference speech patterns in speech recognition | |
| US5425127A (en) | Speech recognition method | |
| JPH07121197A (en) | Learning voice recognition method | |
| Wilpon et al. | Connected digit recognition based on improved acoustic resolution | |
| JPH04332000A (en) | Voice recognition method | |
| JPH07210197A (en) | Method of identifying speaker | |
| Saeta et al. | New speaker-dependent threshold estimation method in speaker verification based on weighting scores | |
| JPS63798B2 (en) | ||
| JPH10307596A (en) | Voice recognition device |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| EEER | Examination request | ||
| MKLA | Lapsed |