KR20140077773A - Apparatus and method for recognizing speech using user location information - Google Patents

Apparatus and method for recognizing speech using user location information Download PDF

Info

Publication number
KR20140077773A
KR20140077773A KR1020120146898A KR20120146898A KR20140077773A KR 20140077773 A KR20140077773 A KR 20140077773A KR 1020120146898 A KR1020120146898 A KR 1020120146898A KR 20120146898 A KR20120146898 A KR 20120146898A KR 20140077773 A KR20140077773 A KR 20140077773A
Authority
KR
South Korea
Prior art keywords
user
location information
language model
model
unit
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
KR1020120146898A
Other languages
Korean (ko)
Inventor
강병옥
이윤근
Original Assignee
한국전자통신연구원
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by 한국전자통신연구원 filed Critical 한국전자통신연구원
Priority to KR1020120146898A priority Critical patent/KR20140077773A/en
Publication of KR20140077773A publication Critical patent/KR20140077773A/en
Withdrawn legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/14Speech classification or search using statistical models, e.g. Hidden Markov Models [HMMs]
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/18Speech classification or search using natural language modelling
    • G10L15/183Speech classification or search using natural language modelling using context dependencies, e.g. language models
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/28Constructional details of speech recognition systems

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Probability & Statistics with Applications (AREA)
  • Artificial Intelligence (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

사용자 위치 정보를 활용하여 맞춤형 음향모델 및 언어모델을 제공함으로써 음성인식 서비스의 성능을 높일 수 있는 사용자 위치 정보를 활용한 음성 인식 기술이 개시된다. 이를 위해, 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 장치는 인식의 대상이 되는 사용자의 음성을 수신하는 음성 수신부; 사용자의 위치 정보를 파악하는 위치 정보 파악부; 사용자의 위치 정보를 이용하여, 사용자가 위치하고 있는 곳의 잡음 환경을 분석하고, 잡음 환경에 대응하는 음향모델을 음향모델 DB(Data Base)에서 추출하는 음향모델 추출부; 사용자의 위치 정보를 이용하여, 사용자가 위치하고 있는 곳에 대응되는 어휘언어모델을 어휘언어모델 DB에서 추출하는 어휘언어모델 추출부; 및 음향모델 및 어휘언어모델을 이용하여 사용자의 음성에 대한 인식을 수행하는 음성 인식부를 포함하는 것을 특징으로 한다. A speech recognition technology using user location information capable of enhancing the performance of a speech recognition service by providing a customized acoustic model and a language model using user location information is disclosed. To this end, the speech recognition apparatus utilizing the user location information according to the present invention includes: a voice receiving unit for receiving a voice of a user to be recognized; A position information acquiring unit for acquiring position information of a user; An acoustic model extracting unit for analyzing a noise environment where a user is located by using the user's location information and extracting an acoustic model corresponding to a noisy environment from an acoustic model DB; A vocabulary language model extraction unit for extracting a vocabulary language model corresponding to a location where the user is located, from the vocabulary language model DB, using the location information of the user; And a speech recognition unit for recognizing the user's speech using the acoustic model and the lexical language model.

Description

사용자 위치정보를 활용한 음성 인식 장치 및 방법{Apparatus and method for recognizing speech using user location information}[0001] The present invention relates to a speech recognition apparatus and method using user location information,

본 발명은 사용자 위치 정보를 활용한 음성 인식 장치 및 방법에 관한 것이다. 더욱 상세하게, 본 발명은 사용자 위치 정보를 활용하여 맞춤형 음향모델 및 언어모델을 제공함으로써 음성인식 서비스의 성능을 높일 수 있는 음성 인식 장치 및 방법에 관한 것이다. The present invention relates to a speech recognition apparatus and method using user location information. More particularly, the present invention relates to a speech recognition apparatus and method capable of enhancing the performance of a speech recognition service by providing a customized acoustic model and a language model using user location information.

휴대폰 등의 통신 단말을 사용하여 음성통화 시에 주변 잡음이 존재하는 경우에는 좋은 통화 품질을 보장하기가 어렵다. 따라서 잡음이 존재하는 환경에서 통화 품질을 높이기 위해서는 주변 잡음 성분을 추정하여 실제 음성 신호만을 추출하는 기술이 필요하다.It is difficult to ensure good call quality when ambient noise exists in a voice call using a communication terminal such as a mobile phone. Therefore, in order to improve the speech quality in the presence of noise, a technique of extracting actual speech signals is necessary.

이와 더불어, 캠코더, 노트북 PC, 네비게이션, 게임기, 휴대폰 등 여러가지 단말기에서 음성을 입력받아 동작하거나 음성 데이터를 저장하는 등 음성 기반의 응용예가 증가하고 있어, 주변 잡음을 감소 또는 제거하여 좋은 품질의 음성을 추출해 내는 기술이 필요하다.In addition, voice-based applications such as a camcorder, a notebook PC, a navigation device, a game device, and a mobile phone operate by receiving voice or storing voice data are increasing in number. Thus, I need a technique to extract it.

종래에도 주변 잡음을 추정하거나 감소시키는 여러가지 방법들이 개시되어 있다. 그러나 시간에 따라 잡음의 통계적 특성이 변화하거나, 잡음의 통계적 특성을 알아내기 위한 초기 단계에서, 예측하지 못한 산발적인(sporadic) 잡음이 발생하는 경우에는 원하는 잡음 감소 또는 제거 성능을 얻지 못한다.Conventionally, various methods for estimating or reducing the ambient noise are disclosed. However, when statistical characteristics of noise change over time, or sporadic noise occurs at an early stage in order to obtain statistical characteristics of noise, desired noise reduction or elimination performance is not obtained.

관련하여, 한국특허출원 제10-2009-0085511호는 "잡음 추정 장치 및 방법과, 이를 이용한 감소 장치"에 관한 기술을 개시하고 있다. Korean Patent Application No. 10-2009-0085511 discloses a technique relating to a noise estimation apparatus and method and a reduction apparatus using the same.

본 발명은 사용자 위치 정보를 활용하여 맞춤형 음향모델 및 언어모델을 제공함으로써 음성인식 서비스의 성능을 높이는 것을 목적으로 한다. 더불어, 본 발명은 누적된 사용자 입력을 바탕으로 갱신 및 관리된 데이터베이스를 이용하여 음성인식 서비스의 성능을 보다 높이는 것을 목적으로 한다. The present invention aims at enhancing the performance of a speech recognition service by providing a customized acoustic model and a language model using user location information. In addition, the present invention aims at enhancing the performance of a speech recognition service using a database updated and managed based on cumulative user input.

그리고, 본 발명은 음향모델 DB 및 어휘언어모델 DB에 미리 저장 및 분류된 음향모델 및 어휘언어모델을 이용하여 음성인식을 수행하여, 음성인식 수행 속도의 저하를 방지하는 것을 목적으로 한다. It is another object of the present invention to prevent voice recognition from being slowed down by performing speech recognition using an acoustic model and a lexical language model stored and classified in advance in an acoustic model DB and a lexical language model DB.

상기한 목적을 달성하기 위한 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 장치는 인식의 대상이 되는 사용자의 음성을 수신하는 음성 수신부; 상기 사용자의 위치 정보를 파악하는 위치 정보 파악부; 상기 사용자의 위치 정보를 이용하여, 상기 사용자가 위치하고 있는 곳의 잡음 환경을 분석하고, 상기 잡음 환경에 맞는 음향모델을 음향모델 DB(Data Base)에서 추출하는 음향모델 추출부; 상기 사용자의 위치 정보를 이용하여, 상기 사용자가 위치하고 있는 곳에 대응되는 어휘언어모델을 어휘언어모델 DB에서 추출하는 어휘언어모델 추출부; 및 상기 음향모델 및 상기 어휘언어모델을 이용하여 상기 사용자의 음성에 대한 인식을 수행하는 음성 인식부를 포함하는 것을 특징으로 한다. According to an aspect of the present invention, there is provided a voice recognition apparatus using user location information, comprising: a voice receiving unit for receiving a voice of a user to be recognized; A position information acquiring unit for acquiring position information of the user; An acoustic model extracting unit for analyzing a noise environment where the user is located using the location information of the user and extracting an acoustic model corresponding to the noise environment from an acoustic model DB; A vocabulary language model extraction unit for extracting a vocabulary language model corresponding to a location of the user from the vocabulary language model DB using the location information of the user; And a voice recognition unit for recognizing the voice of the user using the acoustic model and the lexical language model.

본 발명에 따르면, 사용자 위치 정보를 활용하여 맞춤형 음향모델 및 언어모델을 제공함으로써 음성인식 서비스의 성능을 높일 수 있다. 더불어, 본 발명은 누적된 사용자 입력을 바탕으로 갱신 및 관리된 데이터베이스를 이용하여 음성인식 서비스의 성능을 보다 높일 수 있다. According to the present invention, the performance of the speech recognition service can be improved by providing a customized acoustic model and a language model using the user location information. In addition, the present invention can improve the performance of the speech recognition service using a database updated and managed based on accumulated user input.

그리고, 본 발명은 음향모델 DB 및 어휘언어모델 DB에 미리 저장 및 분류된 음향모델 및 어휘언어모델을 이용하여 음성인식을 수행하여, 음성인식 수행 속도의 저하를 방지할 수 있다. In addition, the present invention can prevent voice recognition speed from being degraded by performing speech recognition using an acoustic model and a lexical language model stored and classified in advance in an acoustic model DB and a lexical language model DB.

도 1은 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 장치의 구성을 나타낸 블록도이다.
도 2는 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 방법을 설명하기 위한 플로우챠트이다.
1 is a block diagram showing a configuration of a speech recognition apparatus using user location information according to the present invention.
FIG. 2 is a flowchart illustrating a speech recognition method using user location information according to the present invention.

본 발명을 첨부된 도면을 참조하여 상세히 설명하면 다음과 같다. 여기서, 반복되는 설명, 본 발명의 요지를 불필요하게 흐릴 수 있는 공지 기능, 및 구성에 대한 상세한 설명은 생략한다. 본 발명의 실시형태는 당 업계에서 평균적인 지식을 가진 자에게 본 발명을 보다 완전하게 설명하기 위해서 제공되는 것이다. 따라서, 도면에서의 요소들의 형상 및 크기 등은 보다 명확한 설명을 위해 과장될 수 있다.
The present invention will now be described in detail with reference to the accompanying drawings. Hereinafter, a repeated description, a known function that may obscure the gist of the present invention, and a detailed description of the configuration will be omitted. Embodiments of the present invention are provided to more fully describe the present invention to those skilled in the art. Accordingly, the shapes and sizes of the elements in the drawings and the like can be exaggerated for clarity.

이하에서는 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 장치의 구성 및 동작에 대하여 설명하도록 한다. Hereinafter, the structure and operation of a speech recognition apparatus using user location information according to the present invention will be described.

도 1은 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 장치의 구성을 나타낸 블록도이다.
1 is a block diagram showing a configuration of a speech recognition apparatus using user location information according to the present invention.

도 1을 참조하면, 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 장치(100)는 음성 수신부(110), 위치 정보 파악부(120), 잡음 환경 판단부(130), 음향모델 추출부(140), 어휘언어모델 추출부(150), 음성 인식부(160)를 포함하여 구성된다. 그리고, 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 장치(100)는 갱신부(170)를 더 포함하여 구성될 수 있다. 1, a voice recognition apparatus 100 using user location information according to the present invention includes a voice receiving unit 110, a location information obtaining unit 120, a noise environment determining unit 130, an acoustic model extracting unit 140, a vocabulary language model extraction unit 150, and a speech recognition unit 160. The voice recognition apparatus 100 using the user location information according to the present invention may further include an update unit 170. [

음성 수신부(110)는 인식의 대상이 되는 사용자의 음성을 수신한다. The voice receiving unit 110 receives the voice of the user to be recognized.

위치 정보 파악부(120)는 음성을 발화하는 사용자의 위치 정보를 파악한다. The location information acquiring unit 120 acquires location information of a user who utteres a voice.

잡음 환경 판단부(130)는 위치 정보 파악부(120)에서 파악된 사용자의 위치 정보를 이용하여 사용자가 위치하고 있는 곳의 잡음 환경을 분석한다. 이러한, 잡음 환경 판단부(130)는 잡음 환경에 대한 모델들이 기 정의 및 저장된 잡음 환경 DB(10)에서 잡음 환경 모델을 추출하여 잡음 환경을 분석한다. 이 때, 잡음 환경 DB(10)는 위치 및 환경에 따라 모델링된 여러 형태의 잡음 환경 모델을 저장하고 있는 데이터베이스이다. 구체적으로, 잡음 환경 DB(10)에는 자동차 내부, 지하철역, 지하철, 길거리, 음식점, 집 내부 등의 잡음 환경 모델에 대한 정의가 기 저장되어 있을 수 있다. 이 때, 잡음 환경 판단부(130)에서 결정된 잡음 환경 모델이 사용자의 입력 음성에서 얻어진 신호 특성과 크게 상이할 경우, 위치에 따른 잡음 환경 모델을 적용하지 않고, 기본적으로 제공되는 범용 잡음 환경 모델이 적용될 수 있다. The noise environment determination unit 130 analyzes the noise environment where the user is located by using the user's location information obtained from the location information determination unit 120. The noise environment determination unit 130 analyzes the noise environment by extracting the noise environment model from the noise environment DB 10 having previously defined and stored models for the noise environment. At this time, the noise environment DB 10 is a database storing various types of noise environment models modeled according to the location and environment. Specifically, the noise environment DB 10 may store definitions of noise environment models in the interior of a car, a subway station, a subway, a street, a restaurant, a house, and the like. In this case, when the noise environment model determined by the noise environment determination unit 130 is significantly different from the signal characteristics obtained from the input speech of the user, the noise environment model based on the position is not applied, Can be applied.

음향모델 추출부(140)는 잡음 환경 판단부(130)에서 판단된 잡음 환경 모델에 대응되는 음향모델을 음향모델 DB(20)에서 추출한다. 이 때, 음향모델의 추출은 사용자가 음성을 입력하기 전에 백그라운드 작업을 통해 미리 수행될 수 있다. 음향모델 DB(20)는 사용자 위치 및 환경에 따른 잡음 환경 모델에 대응하여 다양한 형태로 적응된 음향모델에 대한 데이터베이스이다. 예를 들어, 음향모델 추출부(140)는 잡음 환경 판단부(130)에서 지하철역에 대응되는 잡음환경 모델이 정의된 경우, 지하철 소음 등의 노이즈 신호를 제거할 수 있는 음향모델을 추출한다. The acoustic model extraction unit 140 extracts an acoustic model corresponding to the noise environment model determined by the noise environment determination unit 130 from the acoustic model DB 20. [ At this time, the extraction of the acoustic model can be performed in advance through the background work before the user inputs the voice. The acoustic model DB 20 is a database for acoustic models adapted to various types corresponding to noise environment models according to user's location and environment. For example, when the noise environment model corresponding to the subway station is defined in the noise environment determination unit 130, the acoustic model extraction unit 140 extracts an acoustic model capable of removing noise signals such as subway noise.

어휘언어모델 추출부(150)는 사용자의 위치 정보를 이용하여, 사용자가 위치하고 있는 곳에 대응되는 어휘언어모델을 어휘언어모델 DB(30)에서 추출한다. 이러한, 어휘언어모델 추출부(150)는 사용자의 음성에는 사용자 위치를 중심으로 주변 상호명이나 POI(Point Of Interset) 등이 포함될 가능성이 높으므로, 기존 인식 어휘에 해당 어휘를 추가하여 어휘언어모델을 구성한다. 그리고, 어휘언어모델 추출부(150)는 그 동안 사용자가 발화했던 문장에 대한 언어모델을 추가하여 어휘언어모델을 구성한다. 어휘언어모델 DB(30)는 범용의 인식 어휘 모델과 위치에 따른 인식 어휘를 저장하는 인식 어휘 DB(31), 및 그 동안 사용자가 발화했던 문장에 대한 언어 모델을 분석 및 저장하는 누적 사용자 발화 DB(32)로 구성될 수 있다. 그리고, 이러한 어휘언어모델의 추출은 사용자가 음성을 입력하기 전에 백그라운드 작업을 통해 미리 수행될 수 있다. The vocabulary language model extraction unit 150 extracts, from the vocabulary language model DB 30, a vocabulary language model corresponding to a location where the user is located, using the user's location information. Since the vocabulary language model extraction unit 150 has a high possibility that the user's voice includes the surrounding business name or POI (Point Of Interset) around the user's location, the vocabulary language model is added to the existing recognition vocabulary. . Then, the vocabulary language model extracting unit 150 constructs a vocabulary language model by adding a language model for the sentences that the user has uttered. The vocabulary language model DB 30 includes a recognition vocabulary DB 31 for storing a general recognition vocabulary model and a recognition vocabulary according to a position, and a cumulative user speech database 31 for analyzing and storing a language model for a sentence that the user has uttered (32). The extraction of such a vocabulary language model may be performed in advance through a background operation before the user inputs a voice.

음성 인식부(160)는 음향모델 추출부(140)에서 추출된 음향모델 및 어휘언어모델 추출부(150)에서 추출된 어휘언어모델에 기반하여, 음성 수신부(110)에서 수신한 사용자의 음성을 최종 인식한다. The speech recognition unit 160 recognizes the voice of the user received by the voice receiving unit 110 based on the acoustic model extracted by the acoustic model extraction unit 140 and the lexical language model extracted by the lexical language model extraction unit 150 Finally recognize.

갱신부(170)는 사용자가 음성 인식을 위한 음성을 입력할 때마다 잡음환경 DB(10)의 잡음환경 모델, 음향모델 DB(20)의 음향모델, 어휘언어모델 DB(30)의 어휘언어모델을 해당 인식된 음성을 토대로 갱신한다. 즉, 갱신부(170)는 입력되는 음성 신호를 토대로 잡음환경 DB(10)의 위치에 따른 잡음환경 모델의 정보를 갱신한다. 또한, 갱신부(170)는 입력되는 음성 신호 및 잡음환경 모델을 토대로 음향모델의 정보를 갱신한다. 또한, 갱신부(170)는 사용자가 발화한 문장을 토대로 어휘언어모델 DB(30)의 누적 사용자 발화 DB(32)의 통계값을 갱신한다.
The update unit 170 updates the noise environment model of the noise environment DB 10, the acoustic model of the acoustic model DB 20, the vocabulary language model of the vocabulary language model DB 30, On the basis of the recognized voice. That is, the updating unit 170 updates the information of the noise environment model according to the position of the noise environment DB 10 based on the input voice signal. In addition, the updating unit 170 updates the information of the acoustic model based on the input speech signal and the noise environment model. In addition, the updating unit 170 updates the statistical value of the cumulative user utterance DB 32 of the vocabulary language model DB 30 based on the sentences uttered by the user.

이하에서는 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 방법에 대하여 설명하도록 한다. Hereinafter, a speech recognition method using user location information according to the present invention will be described.

도 2는 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 방법을 설명하기 위한 플로우챠트이다.
FIG. 2 is a flowchart illustrating a speech recognition method using user location information according to the present invention.

도 2를 참조하면, 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 방법은 먼저, 인식의 대상이 되는 사용자의 음성을 수신한다(S10).Referring to FIG. 2, a speech recognition method using user location information according to the present invention receives a voice of a user to be recognized (S10).

그리고, 음성을 발화하는 사용자의 위치 정보를 파악한다(S20).Then, the positional information of the user uttering the voice is grasped (S20).

이 후, S20 단계에서 파악된 사용자의 위치 정보에 기반하여, 잡음환경 DB에서 사용자가 위치한 곳의 잡음 환경 모델을 분석한다(S30). Thereafter, the noise environment model of the place where the user is located in the noise environment DB is analyzed based on the user's location information obtained in step S20 (S30).

그리고, S30 단계에서 분석된 잡음 환경 모델에 대응되는 음향모델을 음향모델 DB에서 추출한다(S40).In step S40, the acoustic model corresponding to the analyzed noise environment model is extracted from the acoustic model DB.

사용자의 위치 정보를 이용하여, 사용자가 위치하고 있는 곳에 대응되는 어휘언어모델을 어휘언어모델 DB에서 추출한다(S50). 이러한, S50 단계에서는 사용자의 음성에는 사용자 위치를 중심으로 주변 상호명이나 POI(Point Of Interset) 등이 포함될 가능성이 높으므로, 기존 인식 어휘에 해당 어휘를 추가하여 어휘언어모델을 구성할 수 있다. 그리고, S50 단계에서는 그 동안 사용자가 발화했던 문장에 대한 언어모델을 추가하여 어휘언어모델을 구성할 수 있다. Using the location information of the user, the lexical language model corresponding to the location where the user is located is extracted from the lexical language model DB (S50). In step S50, the voice of the user is highly likely to include the surrounding business name or POI (Point Of Interset) around the user's location. Thus, the vocabulary language model can be constructed by adding the corresponding vocabulary to the existing recognition vocabulary. In step S50, a lexical language model can be constructed by adding a language model for a sentence that the user has uttered during that time.

S40 단계에서 추출된 음향모델 및 S50 단계에서 추출된 어휘언어모델에 기반하여, S10 단계에서 수신된 사용자의 음성을 최종 인식한다(S60).In step S60, the voice of the user received in step S10 is finally recognized based on the acoustic model extracted in step S40 and the lexical language model extracted in step S50.

그리고, 기존의 잡음환경 DB의 잡음환경 모델, 음향모델 DB의 음향모델, 어휘언어모델 DB의 어휘언어모델을 S60 단계에서 인식된 음성을 토대로 갱신한다(S70).
Then, the noise environment model of the existing noise environment DB, the acoustic model of the acoustic model DB, and the lexical language model of the lexical language model DB are updated based on the voice recognized in step S60 (S70).

이상에서와 같이 본 발명에 따른 사용자 위치 정보를 활용한 음성 인식 장치 및 방법은 상기한 바와 같이 설명된 실시예들의 구성과 방법이 한정되게 적용될 수 있는 것이 아니라, 상기 실시예들은 다양한 변형이 이루어질 수 있도록 각 실시예들의 전부 또는 일부가 선택적으로 조합되어 구성될 수도 있다.As described above, the speech recognition apparatus and method using the user location information according to the present invention are not limited to the configuration and method of the embodiments described above, but the embodiments can be variously modified All or some of the embodiments may be selectively combined.

100; 음성 인식 장치
110; 음성 수신부
120; 위치 정보 파악부
130; 잡음 환경 판단부
140; 음향 모델 추출부
150; 어휘언어모델 추출부
160; 음성 인식부
170; 갱신부
10; 잡음 환경 DB
20; 음향모델 DB
30; 어휘언어모델 DB
31; 인식 어휘 DB
32; 누적 사용자 발화 DB
100; Voice recognition device
110; Voice receiver
120; The location information-
130; The noise environment determination unit
140; The acoustic model extracting unit
150; The lexical language model extraction unit
160; The voice recognition unit
170; Updating unit
10; Noise environment DB
20; Acoustic model DB
30; Lexical language model DB
31; Recognition vocabulary DB
32; Cumulative User Ignition DB

Claims (1)

인식의 대상이 되는 사용자의 음성을 수신하는 음성 수신부;
상기 사용자의 위치 정보를 파악하는 위치 정보 파악부;
상기 사용자의 위치 정보를 이용하여, 상기 사용자가 위치하고 있는 곳의 잡음 환경을 분석하고, 상기 잡음 환경에 대응하는 음향모델을 음향모델 DB(Data Base)에서 추출하는 음향모델 추출부;
상기 사용자의 위치 정보를 이용하여, 상기 사용자가 위치하고 있는 곳에 대응되는 어휘언어모델을 어휘언어모델 DB에서 추출하는 어휘언어모델 추출부; 및
상기 음향모델 및 상기 어휘언어모델을 이용하여 상기 사용자의 음성에 대한 인식을 수행하는 음성 인식부를 포함하는 것을 특징으로 하는 사용자 위치 정보를 활용한 음성 인식 장치.
A voice receiving unit for receiving a voice of a user to be recognized;
A position information acquiring unit for acquiring position information of the user;
An acoustic model extracting unit for analyzing a noise environment where the user is located by using the location information of the user and extracting an acoustic model corresponding to the noise environment from an acoustic model DB;
A vocabulary language model extraction unit for extracting a vocabulary language model corresponding to a location of the user from the vocabulary language model DB using the location information of the user; And
And a speech recognition unit for recognizing the user's speech using the acoustic model and the lexical language model.
KR1020120146898A 2012-12-14 2012-12-14 Apparatus and method for recognizing speech using user location information Withdrawn KR20140077773A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
KR1020120146898A KR20140077773A (en) 2012-12-14 2012-12-14 Apparatus and method for recognizing speech using user location information

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
KR1020120146898A KR20140077773A (en) 2012-12-14 2012-12-14 Apparatus and method for recognizing speech using user location information

Publications (1)

Publication Number Publication Date
KR20140077773A true KR20140077773A (en) 2014-06-24

Family

ID=51129623

Family Applications (1)

Application Number Title Priority Date Filing Date
KR1020120146898A Withdrawn KR20140077773A (en) 2012-12-14 2012-12-14 Apparatus and method for recognizing speech using user location information

Country Status (1)

Country Link
KR (1) KR20140077773A (en)

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2019203580A1 (en) * 2018-04-17 2019-10-24 엘지전자 주식회사 Method and device for providing audio streaming service by using bluetooth low energy technology
CN110634506A (en) * 2019-09-20 2019-12-31 北京小狗智能机器人技术有限公司 Voice data processing method and device
JPWO2022269760A1 (en) * 2021-06-22 2022-12-29
CN116386630A (en) * 2023-03-29 2023-07-04 中国第一汽车股份有限公司 Vehicle control method, device, vehicle and storage medium
WO2024029851A1 (en) * 2022-08-05 2024-02-08 삼성전자주식회사 Electronic device and speech recognition method

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2019203580A1 (en) * 2018-04-17 2019-10-24 엘지전자 주식회사 Method and device for providing audio streaming service by using bluetooth low energy technology
CN110634506A (en) * 2019-09-20 2019-12-31 北京小狗智能机器人技术有限公司 Voice data processing method and device
JPWO2022269760A1 (en) * 2021-06-22 2022-12-29
WO2022269760A1 (en) * 2021-06-22 2022-12-29 ファナック株式会社 Speech recognition device
WO2024029851A1 (en) * 2022-08-05 2024-02-08 삼성전자주식회사 Electronic device and speech recognition method
CN116386630A (en) * 2023-03-29 2023-07-04 中国第一汽车股份有限公司 Vehicle control method, device, vehicle and storage medium

Similar Documents

Publication Publication Date Title
US9711135B2 (en) Electronic devices and methods for compensating for environmental noise in text-to-speech applications
US9711136B2 (en) Speech recognition device and speech recognition method
US8972263B2 (en) System and method for performing dual mode speech recognition
KR102281178B1 (en) Method and apparatus for recognizing multi-level speech
US9177545B2 (en) Recognition dictionary creating device, voice recognition device, and voice synthesizer
US9837068B2 (en) Sound sample verification for generating sound detection model
US9865249B2 (en) Realtime assessment of TTS quality using single ended audio quality measurement
US8438030B2 (en) Automated distortion classification
KR20190100334A (en) Contextual Hotwords
US9881609B2 (en) Gesture-based cues for an automatic speech recognition system
CN104981871B (en) Individualized bandwidth expansion
US20160111090A1 (en) Hybridized automatic speech recognition
CN111354363A (en) Vehicle-mounted voice recognition method and device, readable storage medium and electronic equipment
US9473094B2 (en) Automatically controlling the loudness of voice prompts
US10008205B2 (en) In-vehicle nametag choice using speech recognition
CN112017642B (en) Speech recognition method, device, equipment and computer-readable storage medium
US20160012819A1 (en) Server-Side ASR Adaptation to Speaker, Device and Noise Condition via Non-ASR Audio Transmission
US10229701B2 (en) Server-side ASR adaptation to speaker, device and noise condition via non-ASR audio transmission
CN110826637A (en) Emotion recognition method, system and computer-readable storage medium
US9159315B1 (en) Environmentally aware speech recognition
KR20180012639A (en) Voice recognition method, voice recognition device, apparatus comprising Voice recognition device, storage medium storing a program for performing the Voice recognition method, and method for making transformation model
JP5988077B2 (en) Utterance section detection apparatus and computer program for detecting an utterance section
US11948567B2 (en) Electronic device and control method therefor
CN105047196A (en) Systems and methods for speech artifact compensation in speech recognition systems
US20200321006A1 (en) Agent apparatus, agent apparatus control method, and storage medium

Legal Events

Date Code Title Description
PA0109 Patent application

Patent event code: PA01091R01D

Comment text: Patent Application

Patent event date: 20121214

PG1501 Laying open of application
PC1203 Withdrawal of no request for examination
WITN Application deemed withdrawn, e.g. because no request for examination was filed or no examination fee was paid