JPH11259086A - Voice recognition method and voice recognition device - Google Patents

Voice recognition method and voice recognition device

Info

Publication number
JPH11259086A
JPH11259086A JP10062790A JP6279098A JPH11259086A JP H11259086 A JPH11259086 A JP H11259086A JP 10062790 A JP10062790 A JP 10062790A JP 6279098 A JP6279098 A JP 6279098A JP H11259086 A JPH11259086 A JP H11259086A
Authority
JP
Japan
Prior art keywords
phoneme
standard pattern
word
pattern
similarity
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Granted
Application number
JP10062790A
Other languages
Japanese (ja)
Other versions
JP3289670B2 (en
Inventor
Takeo Oono
剛男 大野
Maki Yamada
麻紀 山田
Masakatsu Hoshimi
昌克 星見
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Panasonic Holdings Corp
Original Assignee
Matsushita Electric Industrial Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Matsushita Electric Industrial Co Ltd filed Critical Matsushita Electric Industrial Co Ltd
Priority to JP06279098A priority Critical patent/JP3289670B2/en
Publication of JPH11259086A publication Critical patent/JPH11259086A/en
Application granted granted Critical
Publication of JP3289670B2 publication Critical patent/JP3289670B2/en
Anticipated expiration legal-status Critical
Expired - Lifetime legal-status Critical Current

Links

Abstract

(57)【要約】 【課題】 少ない話者適応学習用音声データによる音素
標準パターンの話者適応により、高性能な話者適応機能
を有する音声認識装置を提供することを目的とする。 【解決手段】 音響分析部101において、入力音声を
音響パラメータ系列に変換し、不特定話者用音素標準パ
ターン格納部103に格納された、不特定話者用音素標
準パターンと、音素類似度計算部102で照合し、音素
類似度時系列に変換する。パターンマッチング部104
で、音素類似度時系列を特徴量パラメータとして、単語
標準パターンと時間照合し、適応学習パターン抽出部1
08で、単語標準パターン中の高音素類似度フレームに
対応した入力音声フレーム近傍から話者適応学習サンプ
ルパターンを抽出し、音素標準パターン適応部109で
音素標準パターンの話者学習を行い、特定話者用音素標
準パターンとして認識時に用いる。
(57) [Problem] To provide a speech recognition device having a high-performance speaker adaptation function by speaker adaptation of a phoneme standard pattern with a small amount of speaker adaptation learning speech data. SOLUTION: An acoustic analysis unit 101 converts an input speech into an acoustic parameter sequence, and stores a phoneme standard pattern for unspecified speakers stored in a phoneme standard pattern storage unit 103 for unspecified speakers and a phoneme similarity calculation. The collation is performed by the unit 102 and converted into a phoneme similarity time series. Pattern matching unit 104
Then, the phoneme similarity time series is time-matched with a word standard pattern as a feature parameter, and the adaptive learning pattern extraction unit 1
In step 08, a speaker adaptation learning sample pattern is extracted from the vicinity of the input speech frame corresponding to the high phoneme similarity frame in the word standard pattern, and the phoneme standard pattern adaptation unit 109 performs speaker learning of the phoneme standard pattern, and It is used at the time of recognition as a standard phoneme standard pattern.

Description

【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION

【0001】[0001]

【発明の属する技術分野】本発明は、話者適応機能を有
する音声認識方法および音声認識装置に関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a speech recognition method and apparatus having a speaker adaptation function.

【0002】[0002]

【従来の技術】人の発声する音声は、異なる話者が発声
した場合、例え同じ語彙発声しても、喉頭、舌、唇など
で構成される調音器官の特性が個人ごとによって異なる
ため、音声信号を音響分析した結果得られる音響パラメ
ータは、人ごとの調音器官の特性に依存して、微妙な差
が生じる。これを音声信号の個人性と呼ぶ。多くの音声
認識装置では、識別のための特徴量パラメータとして、
調音器官の特性の影響を受けやすい、音響パラメータを
用いている。不特定話者用音声認識装置の標準パターン
は、不特定多数の話者が発声した音声信号から学習され
た音響パラメータの平均値から構成されるため、調音器
官の特性が著しく平均的な特性から異なる話者に対して
は、認識性能が低下してしまうという問題がある。そこ
で、従来、こうした話者の個人性に基づく問題を対処す
るためには、話者の発声した音声信号から、不特定話者
標準パターンをその話者に適した値に適応することで、
音声認識装置に話者適応機能をもたせて、認識性能を維
持する方法がとられてきた。こうした、話者の個人性に
基づく認識率低下を防ぐ話者適応方式としては、特開平
8−22296号公報に記載されたものが知られてい
る。このような、話者の個人性にもとづく認識率の低下
に対して標準パターンを話者適応学習することで対処す
る従来法の一例の構成を、図3に示す。
2. Description of the Related Art When a different speaker utters a voice, even if the same vocabulary is uttered, the characteristics of articulatory organs composed of the larynx, tongue, lips, etc. are different for each individual. Acoustic parameters obtained as a result of acoustic analysis of the signal have subtle differences depending on the characteristics of the articulator for each person. This is called the personality of the audio signal. In many speech recognition devices, as a feature parameter for identification,
Acoustic parameters are used that are susceptible to articulatory characteristics. The standard pattern of the speech recognition device for unspecified speakers is composed of average values of acoustic parameters learned from speech signals uttered by an unspecified number of speakers. For different speakers, there is a problem that the recognition performance deteriorates. Therefore, conventionally, in order to cope with such a problem based on the individuality of a speaker, an unspecified speaker standard pattern is adapted to a value suitable for the speaker from a voice signal uttered by the speaker.
A method has been adopted in which a speech recognition apparatus is provided with a speaker adaptation function to maintain recognition performance. As a speaker adaptation method for preventing a decrease in the recognition rate based on the individuality of the speaker, a method described in Japanese Patent Application Laid-Open No. H8-22296 is known. FIG. 3 shows a configuration of an example of a conventional method for coping with such a decrease in the recognition rate based on the individuality of a speaker by performing speaker adaptive learning of a standard pattern.

【0003】従来法においては、入力音声は、音響分析
部301において、音響パラメータ系列に変換される。
音響パラメータとしては、例えば中川著、「確率モデル
による音声認識」、電子情報通信学会(昭和63年)に
あげられている、LPCケプストラム係数、LPCメル
ケプストラム係数などが用いられる。話者適応時、単語
辞書303は、話者学習用入力音声に対応した単語がい
かなる音声片から構成されるかを表す音声片情報を、不
特定話者用音声片標準パターン格納部304に送る。こ
こで、音声片の単位としては、後続の音素環境を考慮し
た音素バイグラム、あるいは、前後の音素環境を考慮し
た音素トライグラムなどが考えられる。不特定話者用音
声片標準パターン格納部304では、単語辞書303か
ら送られた音声片情報に基づき、音響パラメータから構
成される音声片標準パターンの接続により不特定話者用
単語標準パターンを作成し、パターンマッチング部30
2と音声片標準パターン適応部305に送る。パターン
マッチング部302では、入力音声から得られた音響パ
ラメータ時系列と、不特定話者単語標準パターンが時間
照合され、時間対応結果が算出される。パターンマッチ
ング部302での時間照合の方法としては、例えばDP
マッチング、HMM(Hidden Markov M
odel)などが利用される。話者適応時には、パター
ンマッチング部302から、入力音声の音響パラメータ
時系列と不特定話者用単語標準パターンの時間対応結果
と、入力音声の音響パラメータ時系列が、音声片標準パ
ターン適応部305に送られる。音声片標準パターン適
応部305では、入力音声の音響パラメータ時系列と単
語標準パターンの線形結合などの方法により、単語標準
パターンを特定話者用に適応する。適応後の単語標準パ
ータンは、音声片標準パターン適応部において音声片単
位に分割され、特定話者用音声片標準パターン格納部3
06に格納される。認識時には、特定話者用音声片標準
パターン格納部306に格納された特定話者用音声片標
準パターンを用いて特定話者用単語標準パターンを、単
語辞書303中全ての単語に対して構成し、パターンマ
ッチング部302にて、入力音声の音響パラメータ時系
列と時間照合計算を行い、最も照合スコアの高かった辞
書単語を最終的な認識結果として出力する。
[0003] In the conventional method, an input speech is converted into an acoustic parameter series in an acoustic analysis unit 301.
As the acoustic parameters, for example, LPC cepstrum coefficients, LPC mel cepstrum coefficients, etc., which are described in Nakagawa, "Speech Recognition by Stochastic Model", IEICE (1988), are used. At the time of speaker adaptation, the word dictionary 303 sends, to the unspecified speaker's speech unit standard pattern storage unit 304, speech segment information indicating what speech segment the word corresponding to the speaker learning input speech is composed of. . Here, as a unit of the speech piece, a phoneme bigram considering the following phoneme environment, a phoneme trigram considering the preceding and succeeding phoneme environments, and the like can be considered. In the unspecified speaker's speech unit standard pattern storage unit 304, based on the speech unit information sent from the word dictionary 303, the unspecified speaker's word standard pattern is created by connecting speech unit standard patterns composed of acoustic parameters. And the pattern matching unit 30
2 to the speech unit standard pattern adaptation unit 305. In the pattern matching unit 302, the time series of the acoustic parameter time series obtained from the input voice and the unspecified speaker word standard pattern are time-matched, and the time correspondence result is calculated. As a method of time matching in the pattern matching unit 302, for example, DP
Matching, HMM (Hidden Markov M)
model) is used. At the time of speaker adaptation, the pattern matching unit 302 sends the audio parameter time series of the input speech, the time correspondence result of the word standard pattern for unspecified speakers, and the audio parameter time series of the input speech to the speech piece standard pattern adaptation unit 305. Sent. The speech piece standard pattern adaptation unit 305 adapts the word standard pattern for a specific speaker by a method such as linear combination of the acoustic parameter time series of the input speech and the word standard pattern. The word standard pattern after the adaptation is divided into speech unit units in the speech unit standard pattern adaptation unit, and the speech unit standard pattern storage unit 3 for the specific speaker is used.
06. At the time of recognition, a specific speaker word standard pattern is formed for all the words in the word dictionary 303 using the specific speaker voice segment standard pattern stored in the specific speaker voice segment standard pattern storage unit 306. Then, the pattern matching unit 302 performs a time collation calculation with the acoustic parameter time series of the input speech, and outputs the dictionary word having the highest collation score as the final recognition result.

【0004】[0004]

【発明が解決しようとする課題】しかしながら、音響パ
ラメータで表される音声片標準パターンを、入力音声の
音響パラメータ時系列との時間照合により、特定話者用
に話者適応する従来の方法においては、認識対象となる
単語辞書の語彙が大きく、単語標準パターンを構成する
ために必要な音声片標準パターンの種類が大きい場合に
は、全ての語彙に対して話者適応学習済みの特定話者用
音声片標準パターンを用意するためには、話者適応学習
のために、大量の学習サンプル音声を話者が発声する必
要があり、全ての音声片標準パターンを学習するのに充
分な発声を行うことは話者に多大な負担をかけるという
課題があった。また、話者が膨大な話者適応学習用の発
声を行う場合にも、音声片標準パターンの格納に必要な
メモリー量が膨大な場合には、特定話者用の音声片標準
パターン用にさらに膨大なメモリー量を必要とすること
になり、音声認識装置自体に必要とされるメモリー量
が、大きくなってしまうという課題があった。さらに、
従来の方法において、限られた少量の話者適応学習音声
データから、発声外の音声片標準パターンを適応する場
合には、例えば電子通信情報学会 SP92−16(1
992年)に記載されたVFS(ベクトル場平滑化)な
どの工夫が必要となり、音声認識装置の構成が複雑にな
ってしまうという課題があった。
However, in a conventional method for adapting a speaker for a specific speaker by time collation of a speech unit standard pattern represented by acoustic parameters with a time series of acoustic parameters of input speech. If the vocabulary of the word dictionary to be recognized is large and the type of speech vocal standard pattern necessary to compose the word standard pattern is large, for all the vocabularies, the speaker adaptation learning is performed for the specific speaker. In order to prepare a speech unit standard pattern, the speaker needs to utter a large amount of training sample speech for speaker adaptation learning, and utters enough to learn all the speech unit standard patterns. That puts a heavy burden on the speaker. Also, when a speaker performs vocalization for huge speaker adaptation learning, if a large amount of memory is required for storing a speech unit standard pattern, an additional speech unit standard pattern for a specific speaker is required. A large amount of memory is required, and the amount of memory required for the speech recognition device itself becomes large. further,
In the conventional method, when a speech unit standard pattern other than utterance is adapted from a limited small amount of speaker adaptation learning speech data, for example, the Institute of Electronics, Information and Communication Engineers SP92-16 (1
In such a case, a device such as VFS (Vector Field Smoothing) described in U.S. Pat.

【0005】本発明は、上述の問題を解決するものであ
り、少ないメモリー量によって、少量の話者適応学習デ
ータでも効率的な話者適応を実現できる、高性能な音声
認識装置を提供することを目的とする。
An object of the present invention is to solve the above-mentioned problem, and to provide a high-performance speech recognition apparatus capable of realizing efficient speaker adaptation with a small amount of memory and with a small amount of speaker adaptation training data. With the goal.

【0006】[0006]

【課題を解決するための手段】本発明による音声認識方
法は、入力音声に対して分析時間毎にm個の特徴パラメ
ータ系列を求め、分析時間毎にn種類の音素標準パター
ンとマッチングを行い、分析時間毎に求めたn個の類似
度と、予め作成しておいた音素類似度の時系列で構成さ
れる単語標準パターンとマッチングすることにより単語
を認識する音声認識方法において、前記単語標準パター
ンから音素類似度が高いフレームを抽出しておき、入力
音声と単語標準パターンのマッチングで得られる時間対
応の結果から、単語標準パターンの音素類似度が高いフ
レームに対応した入力音声フレームを求め、このフレー
ムの入力音声の特徴パラメータを音素標準パターン学習
データとして抽出し、前記音素標準パターンと前記音素
標準パターン学習データとを統合し、新たな音素標準パ
ターンを作成するものである。
A speech recognition method according to the present invention obtains m feature parameter sequences for each analysis time for an input speech, performs matching with n types of phoneme standard patterns for each analysis time, In the speech recognition method for recognizing a word by matching n similarities obtained for each analysis time with a word standard pattern composed of a time series of phoneme similarities created in advance, the word standard pattern A frame having a high phoneme similarity is extracted from the input speech, and an input speech frame corresponding to a frame having a high phoneme similarity of the word standard pattern is obtained from the result of the time correspondence obtained by matching the input voice and the word standard pattern. The feature parameters of the input speech of the frame are extracted as phoneme standard pattern training data, and the phoneme standard pattern and the phoneme standard pattern learning are extracted. Integrating the chromatography data is intended to create a new phoneme standard patterns.

【0007】これにより、少ないメモリー量によって、
少数の話者適応学習データでも効率的な話者適応を実現
できる高性能な音声認識装置を提供することができる。
Thus, with a small amount of memory,
It is possible to provide a high-performance speech recognition device that can realize efficient speaker adaptation even with a small number of speaker adaptation learning data.

【0008】[0008]

【発明の実施の形態】本発明の請求項1に記載の発明
は、入力音声と不特定話者用の音素標準パターンとの照
合で得られる音素類似度を認識の特徴量パラメータとす
る音声認識方法において、音素標準パターンに対して話
者適応学習を行うもので、音素標準パターンと音響パラ
メータ時系列との照合によって得られる音素類似度を単
語パターンマッチング時の特徴量パラメータとし、音素
標準パターンの話者適応学習を行うことにより、少量の
話者適応学習データでも効率的な話者適応が可能になる
という作用を有する。
DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention according to claim 1 of the present invention provides a speech recognition system in which a phoneme similarity obtained by collating an input speech with a phoneme standard pattern for an unspecified speaker is used as a feature parameter for recognition. In the method, speaker adaptation learning is performed on the phoneme standard pattern, and the phoneme similarity obtained by matching the phoneme standard pattern with the acoustic parameter time series is used as a feature parameter at the time of word pattern matching. Performing speaker adaptation learning has the effect that efficient speaker adaptation is possible even with a small amount of speaker adaptation learning data.

【0009】請求項2に記載の発明は、入力音声に対し
て分析時間毎にm個の特徴パラメータ系列を求め、分析
時間毎にn種類の音素標準パターンとマッチングを行
い、分析時間毎に求めたn個の類似度と、予め作成して
おいた音素類似度の時系列で構成される単語標準パター
ンとマッチングすることにより単語を認識する音声認識
方法において、前記単語標準パターンから音素類似度が
高いフレームを抽出しておき、入力音声と単語標準パタ
ーンのマッチングで得られる時間対応の結果から、単語
標準パターンの音素類似度が高いフレームに対応した入
力音声フレームを求め、このフレームの入力音声の特徴
パラメータを音素標準パターン学習データとして抽出
し、前記音素標準パターンと前記音素標準パターン学習
データとを統合し、新たな音素標準パターンを作成する
もので、話者適応学習用音声データが少量の場合にも、
あるいは、話者適応機能の実現のために多くのメモリー
量もつことことができない場合にも、音素標準パターン
と音響パラメータ時系列との照合によって得られる音素
類似度を単語パターンマッチング時の特徴量パラメータ
とし、音素標準パターンの話者適応学習を行うことによ
り、少ないメモリー量によって、少量の話者適応学習デ
ータでも効率的な話者適応が可能になるという作用を有
する。
According to a second aspect of the present invention, m feature parameter sequences are obtained for an input voice at each analysis time, and matching is performed with n types of phoneme standard patterns at each analysis time. In the speech recognition method for recognizing a word by matching the n similarities and a previously created word standard pattern composed of a time series of phoneme similarities, the phoneme similarity is calculated from the word standard pattern. A high frame is extracted, and an input voice frame corresponding to a frame having a high phoneme similarity of the word standard pattern is obtained from a time correspondence result obtained by matching the input voice with the word standard pattern. The feature parameters are extracted as phoneme standard pattern learning data, the phoneme standard pattern and the phoneme standard pattern learning data are integrated, and a new Intended to create a phoneme standard pattern, even in the case of small amounts speaker adaptive training for the voice data,
Alternatively, even if a large amount of memory is not available for realizing the speaker adaptation function, the phoneme similarity obtained by matching the phoneme standard pattern with the acoustic parameter time series is used as a feature parameter for word pattern matching. By performing the speaker adaptation learning of the phoneme standard pattern, there is an effect that efficient speaker adaptation can be performed with a small amount of memory and with a small amount of speaker adaptation learning data.

【0010】請求項3に記載の発明は、入力音声に対し
て分析時間毎にm個の特徴パラメータ系列を求め、分析
時間毎にn種類の音素標準パターンとマッチングを行
い、分析時間毎に求めたn個の類似度と、予め作成して
おいた音素類似度の時系列で構成される単語標準パター
ンとマッチングすることにより単語を認識する音声認識
装置において、前記単語標準パターンから音素類似度が
高いフレームを抽出する高音素類似度フレーム情報抽出
手段と、入力音声と単語標準パターンのマッチングで得
られる時間対応の結果から、単語標準パターンの音素類
似度が高いフレームに対応した入力音声フレームを求
め、このフレームの入力音声の特徴パラメータを音素標
準パターン学習データとして抽出する適応学習パターン
抽出手段と、前記音素標準パターンと前記音素標準パタ
ーン学習データとを統合し、新たな音素標準パターンを
作成する音素標準パターン適応手段とを具備するもの
で、話者適応学習用音声データが少量の場合にも、ある
いは、話者適応機能の実現のために多くのメモリー量も
つことことができない場合にも、あるいは、話者適応機
能の実現のために音声認識装置の構成を複雑な構成にで
きない場合にも、音素標準パターンと音響パラメータ時
系列との照合によって得られる音素類似度を単語パター
ンマッチング時の特徴量パラメータとし、音素標準パタ
ーンの話者適応学習を行うことにより、少ないメモリー
量によって、少量の話者適応学習データでも効率的な話
者適応が可能になるという作用を有する。
According to a third aspect of the present invention, m feature parameter sequences are obtained for each analysis time with respect to an input voice, matching is performed with n kinds of phoneme standard patterns at each analysis time, and each feature time is obtained for each analysis time. In the speech recognition device that recognizes a word by matching the n similarities and a previously created word standard pattern composed of a time series of phoneme similarities, the phoneme similarity is calculated from the word standard pattern. A high phoneme similarity frame information extracting means for extracting a high frame, and an input speech frame corresponding to a frame having a high phoneme similarity of a word standard pattern are obtained from a time correspondence result obtained by matching the input speech with the word standard pattern. Adaptive learning pattern extracting means for extracting feature parameters of the input speech of this frame as phoneme standard pattern learning data; It integrates the quasi-pattern and the phoneme standard pattern learning data, and includes phoneme standard pattern adaptation means for creating a new phoneme standard pattern, even when the speaker adaptation learning speech data is small, or The phoneme standard is used even when a large amount of memory is not available to implement the speaker adaptation function, or when the configuration of the speech recognizer cannot be complicated to implement the speaker adaptation function. By using the phoneme similarity obtained by matching the pattern with the acoustic parameter time series as a feature parameter at the time of word pattern matching and performing speaker adaptation learning of phoneme standard patterns, a small amount of memory can be used and a small amount of speaker adaptation learning can be performed. It has the effect that efficient speaker adaptation is possible even with data.

【0011】以下図面を参照しながら本発明の実施の形
態について具体的に説明する。 (実施の形態)図1は、本発明の実施の形態による音声
認識装置のブロック構成図を示す。図1において、10
1は入力音声を分析時間毎に音響パラメータに変換する
音響分析部、102は音響分析部101で得られた音響
パラメータとあらかじめ用意された音素種ごとの不特定
話者用音素標準パターンと照合する音素類似度計算部、
103は得られた音響パラメータ時系列を格納する不特
定話者用音素標準パターン格納部、104は音素類似度
時系列を特徴量パラメータとして、単語標準パターンと
時間照合するパターンマッチング部、105は単語辞
書、106は単語辞書106から送られた音声片情報に
基づき、あらかじめ学習された音素類似度からなる音声
片標準パターンを接続する不特定話者用音声片標準パタ
ーン格納部、107は音素類似度時系列で構成される単
語標準パターンから、いずれかの音素に対する音素類似
度が高いフレームを抽出する高音素類似度フレーム情報
抽出部、108は単語標準パターン中の高音素類似度フ
レームに対応した入力音声フレーム近傍から話者適応学
習サンプルパターンを抽出する適応学習パターン抽出
部、109は音素標準パターンの話者学習を行う音素標
準パターン適応部、110は特定話者用音素標準パター
ンを格納した特定話者用音素標準パターン格納部であ
る。
Hereinafter, embodiments of the present invention will be specifically described with reference to the drawings. (Embodiment) FIG. 1 is a block diagram showing a speech recognition apparatus according to an embodiment of the present invention. In FIG. 1, 10
Reference numeral 1 denotes an acoustic analysis unit that converts an input voice into acoustic parameters for each analysis time, and 102 collates the acoustic parameters obtained by the acoustic analysis unit 101 with a phoneme standard pattern for unspecified speakers prepared for each phoneme type. Phoneme similarity calculator,
103 is a phoneme standard pattern storage for an unspecified speaker that stores the obtained acoustic parameter time series, 104 is a pattern matching unit that performs time matching with a word standard pattern using a phoneme similarity time series as a feature parameter, and 105 is a word A dictionary 106 is a speech unit standard pattern storage unit for an unspecified speaker connecting a speech unit standard pattern composed of phoneme similarities learned in advance based on the speech unit information sent from the word dictionary 106, and 107 is a phoneme similarity A high phoneme similarity frame information extraction unit 108 for extracting a frame having a high phoneme similarity to any one of the phonemes from a word standard pattern composed of a time series, and an input 108 corresponding to the high phoneme similarity frame in the word standard pattern An adaptive learning pattern extraction unit for extracting a speaker adaptive learning sample pattern from the vicinity of a voice frame; Phoneme standard pattern adaptation unit for performing a turn speaker learning, 110 is a speaker-for the phoneme standard pattern storage section for storing phoneme standard patterns for a specific speaker.

【0012】上記のように構成された音声認識装置の動
作を以下に説明する。まず、話者適応時について説明す
る。
The operation of the speech recognition apparatus configured as described above will be described below. First, the case of speaker adaptation will be described.

【0013】音響分析部101は、入力音声を分析時間
(フレームと呼ぶ、本実施例では1フレーム=10ms
ec)毎に、10次元のLPCケプストラム系列、パワ
ーの時間差分、正規化残差の、合計12次元の音響パラ
メータに変換する。
The acoustic analysis unit 101 converts an input voice into an analysis time (called a frame, in this embodiment, one frame = 10 ms).
For each ec), a 10-dimensional LPC cepstrum sequence, a time difference in power, and a normalized residual are converted into a total of 12-dimensional acoustic parameters.

【0014】音素類似度計算部102は、音響分析部1
01で得られた音響パラメータ時系列と、あらかじめ不
特定話者用音素標準パターン格納部103に格納され
た、音素種ごとに用意された不特定話者用音素標準パタ
ーンと照合する。ここで音素は、日本語の一般的な定義
に従い{a、o、u、i、e、z、s、hv、hu、
p、t、k、c、b,d、N、j、w、yv、yu、
m、n、ng、r}の24音素分類を用いる。また、各
音素標準パターンは、音素毎に定義されたの特徴フレー
ム(その音素の特徴をよく表現する時間的位置)を目視
によって正確に検出し、この特徴フレームを中心とした
音響パラメータの時間パターンを使用して作成する。本
実施例では、時間パターンとして特徴フレームの近傍1
0フレーム分の音響パラメータによって計120次元
(音響パラメータ12次元×時間パターン10フレー
ム)のパターンを構成し、不特定多数の話者の発声デー
タから、あらかじめ音素標準パターンを学習しておく。
また、音素類似度は、共分散行列を共通化したマハラノ
ビス距離を用いて求める。
The phoneme similarity calculation unit 102 includes a sound analysis unit 1
01 is compared with the phoneme standard pattern for unspecified speakers prepared for each phoneme type, which is stored in advance in the phoneme standard pattern storage unit 103 for unspecified speakers. Here, phonemes are defined as {a, o, u, i, e, z, s, hv, hu,
p, t, k, c, b, d, N, j, w, yv, yu,
A 24 phoneme classification of m, n, ng, r} is used. In addition, each phoneme standard pattern accurately detects a feature frame defined for each phoneme (temporal position at which the feature of the phoneme is well expressed) by visual observation and obtains a time pattern of acoustic parameters centered on the feature frame. Create using In the present embodiment, as the time pattern, the neighborhood 1 of the feature frame is used.
A total of 120-dimensional (12-dimensional acoustic parameters × 10 temporal patterns) patterns are formed by the acoustic parameters for 0 frames, and a phoneme standard pattern is learned in advance from the utterance data of an unspecified number of speakers.
The phoneme similarity is obtained using a Mahalanobis distance in which a covariance matrix is shared.

【0015】入力音声の音響パラメータ時系列を不特定
話者用音素標準パターンと照合した結果得られた音素類
似度時系列は、音素種の次元(本実施例では24次元)
をもち、単語パターンマッチング時の特徴量パラメータ
として、パターンマッチング部104に送られる。
The phoneme similarity time series obtained as a result of comparing the acoustic parameter time series of the input speech with the phoneme standard pattern for an unspecified speaker is the dimension of the phoneme type (24 dimensions in this embodiment).
And sent to the pattern matching unit 104 as a feature parameter during word pattern matching.

【0016】単語辞書105は、話者学習用入力音声に
対応した単語がいかなる音声片から構成されるかを表す
音声片情報を、音素類似度時系列で構成される不特定話
者用音声片標準パターンを格納する、不特定話者用音声
片標準パターン格納部106に送る。ここで、本実施例
においては、音声片の単位は、前後の音素環境を考慮し
たCV/VC単位で、536種類が存在する。
The word dictionary 105 stores speech segment information indicating what speech segment the word corresponding to the input speech for speaker learning is composed of, based on a phoneme similarity time series, a speech segment for an unspecified speaker. The standard pattern is sent to the unspecified speaker voice segment standard pattern storage unit 106. Here, in the present embodiment, there are 536 types of speech piece units, which are CV / VC units in consideration of the preceding and succeeding phoneme environments.

【0017】不特定話者用音声片標準パターン格納部1
06では、単語辞書105から送られた音声片情報に基
づき、あらかじめ学習された音素類似度からなる音声片
標準パターンを接続し、単語標準パターンをパターンマ
ッチング部104と高音素類似度フレーム情報抽出部1
07に送る。
Speech fragment standard pattern storage unit 1 for unspecified speakers
In step 06, based on the speech piece information sent from the word dictionary 105, a speech piece standard pattern consisting of phoneme similarities learned in advance is connected, and the word standard pattern is combined with the pattern matching section 104 and the high phoneme similarity frame information extraction section. 1
Send to 07.

【0018】パターンマッチング部104では、入力音
声から得られた音素類似度時系列と、単語標準パターン
がDPマッチングにより時間照合され、時間対応結果が
算出される。
In the pattern matching unit 104, the time series of the phoneme similarity obtained from the input speech and the word standard pattern are time-matched by DP matching, and a time correspondence result is calculated.

【0019】高音素類似度フレーム情報抽出部107で
は、音素類似度時系列で構成される単語標準パターンか
ら、いずれかの音素に対する音素類似度が高いフレーム
を抽出し、どの音素がどのフレームで高い音素類似度を
もつかを表す、高音素類似度フレーム情報として、適応
学習パターン抽出部108に送る。本実施例において
は、高音素類似度フレーム情報抽出部107では、24
音素に対するいずれかの音素類似度がしきい値Thより
高いフレームを全て高音素類似度フレームとした。ま
た、しきい値Thは実験により最適値を求めた。
The high phoneme similarity frame information extraction unit 107 extracts a frame having a high phoneme similarity to any one of the phonemes from a word standard pattern composed of a phoneme similarity time series, and determines which phoneme is high in which frame. It is sent to the adaptive learning pattern extraction unit 108 as high phoneme similarity frame information indicating whether or not it has phoneme similarity. In the present embodiment, the high phoneme similarity frame information extraction unit 107
All frames in which one of the phoneme similarities to the phonemes was higher than the threshold value Th were defined as high phoneme similarity frames. The optimum value of the threshold value Th was obtained by an experiment.

【0020】適応学習パターン抽出部108において、
単語標準パターン中の高音素類似度フレーム情報と、入
力音声と単語標準パターンの時間対応結果から、単語標
準パターン中の高音素類似度フレームに対応した入力音
声フレームが算出され、このフレーム近傍の入力音声の
音響パラメータパターンを話者適応学習サンプルパター
ンとして抽出し、音素標準パターン適応部109に送
る。
In the adaptive learning pattern extraction unit 108,
The input speech frame corresponding to the high phoneme similarity frame in the word standard pattern is calculated from the high phoneme similarity frame information in the word standard pattern and the time correspondence result between the input voice and the word standard pattern, and the input near this frame is calculated. The acoustic parameter pattern of the voice is extracted as a speaker adaptation learning sample pattern, and sent to the phoneme standard pattern adaptation unit 109.

【0021】音素標準パターン適応部109では、適応
学習パターン抽出部108からの話者適応学習サンプル
パターンと、不特定話者用音素標準パターン格納部10
3からの不特定話者用音素標準パターンとから、特定話
者用音素標準パターンを算出し、特定話者用音素標準パ
ターン格納部110に格納する。本実施例においては、
特定話者用音素標準パターンの平均値(数1)が、
The phoneme standard pattern adapting unit 109 stores the speaker adaptive learning sample pattern from the adaptive learning pattern extracting unit 108 and the phoneme standard pattern storage unit 10 for unspecified speakers.
The phoneme standard pattern for specific speaker is calculated from the phoneme standard pattern for unspecified speaker from No. 3 and stored in the phoneme standard pattern storage unit for specific speaker 110. In this embodiment,
The average value (Equation 1) of the phoneme standard pattern for a specific speaker is

【0022】[0022]

【数1】 (Equation 1)

【0023】適応前の不特定話者用音素標準パターンの
平均(数2)と、
The average of the phoneme standard pattern for unspecified speakers before adaptation (Equation 2);

【0024】[0024]

【数2】 (Equation 2)

【0025】話者適応学習サンプルパターンの値(数
3)の線形結合として、
As a linear combination of the speaker adaptive learning sample pattern values (Equation 3),

【0026】[0026]

【数3】 (Equation 3)

【0027】(数4)と計算される。(Equation 4) is calculated.

【0028】[0028]

【数4】 (Equation 4)

【0029】ここでαは線形結合の混合比で実験による
最適値を用いる。
Here, α is a mixing ratio of a linear combination, and an optimum value by an experiment is used.

【0030】次に、認識時の処理について説明する。音
響分析部101は、入力音声を分析時間(フレームと呼
ぶ、本実施例では1フレーム=10msec)毎に、1
0次元のLPCケプストラム系列、パワーの時間差分、
正規化残差の、合計12次元の音響パラメータに変換す
る。音素類似度計算部102は、音響分析部101で得
られた音響パラメータ時系列と、特定話者用音素標準パ
ターン格納部110に格納された特定話者用音素標準パ
ターンを用いて音素類似度時系列を算出する。
Next, the processing at the time of recognition will be described. The acoustic analysis unit 101 converts the input voice into one every analysis time (called a frame, one frame = 10 msec in this embodiment).
0-dimensional LPC cepstrum sequence, power time difference,
The normalized residuals are converted into a total of 12-dimensional acoustic parameters. The phoneme similarity calculation unit 102 uses the acoustic parameter time series obtained by the acoustic analysis unit 101 and the phoneme standard pattern for a specific speaker stored in the phoneme standard pattern storage unit for a specific speaker 110 to calculate the phoneme similarity. Calculate the series.

【0031】パターンマッチング部104は、不特定話
者用音声片標準パターン格納部106からの単語標準パ
ターンと照合計算を行い、最も照合スコアの高い辞書単
語を最終的な認識結果として出力する。
The pattern matching unit 104 performs a matching calculation with the word standard pattern from the unspecified speaker voice unit standard pattern storage unit 106, and outputs a dictionary word having the highest matching score as a final recognition result.

【0032】図2は、本実施例における話者適応の効果
を示す、音素類似度時系列の概念図であり、発声『ZA
MA(ざま)』における、音素/a/の音素類似度の時
間変化を示す。上段a)は、不特定話者用音声片標準パ
ターンを接続して作成した単語標準パターンの時間変
化、下段b)の点線は、入力音声と不特定話者用音素標
準パターン/a/とを照合して得られた音素類似度時系
列、下段b)の実線は、入力音声と適応後の特定話者音
素標準パターン/a/とを照合して得られた音素類似度
時系列である。単語標準パターン中の、高音素類似度フ
レームの一つであるフレーム(数5)は高い音素類似度
を持つ。
FIG. 2 is a conceptual diagram of a phoneme similarity time series showing the effect of speaker adaptation in the present embodiment.
MA (Zama) ", the temporal change of the phoneme similarity of the phoneme / a /. The upper part a) shows the time change of the word standard pattern created by connecting the unspecified speaker speech piece standard patterns, and the dotted line in the lower part b) shows the input speech and the phoneme standard pattern / a / for the unspecified speaker. The phoneme similarity time series obtained by collation, the solid line in the lower part b) is a phoneme similarity time series obtained by collating the input voice with the specific speaker phoneme standard pattern / a / after adaptation. The frame (Equation 5) which is one of the high phoneme similarity frames in the word standard pattern has a high phoneme similarity.

【0033】[0033]

【数5】 (Equation 5)

【0034】入力音声と不特定話者用音素標準パターン
/a/との照合で得られる音素類似度時系列で、DPマ
ッチングにより(数5)に対応した入力フレーム(数
6)は、(数5)に比べて小さい類似度である。
The input frame (equation 6) corresponding to (equation 5) by DP matching is a phoneme similarity time series obtained by collating the input speech with the phoneme standard pattern for unspecified speakers / a / The similarity is smaller than 5).

【0035】[0035]

【数6】 (Equation 6)

【0036】一方、本実施例に基づき話者学習された特
定話者用音素標準パターン/a/との照合で得られた音
素類似度系列では、(数5)に対応した入力フレーム
(数7)では、(数5)同様、高い音素類似度を持ち、
本実施例による話者適応効果が確認できる。
On the other hand, in the phoneme similarity sequence obtained by collation with the phoneme standard pattern / a / for a specific speaker, which has undergone speaker learning based on the present embodiment, the input frame (equation 7) corresponding to (equation 5) ) Has a high phoneme similarity as in (Equation 5),
The speaker adaptation effect according to the present embodiment can be confirmed.

【0037】[0037]

【数7】 (Equation 7)

【0038】以上、本実施例の構成を用いて、100単
語を発声した11名の音声データの認識実験を行った。
音素標準パターンの話者適応学習は、評価発声データと
は異なる20単語を用いて行った。11名の平均認識率
が、本実施例に基づく話者適応を行う前は93.2%で
あったのに対し、本実施例に基づく話者適応後は、9
6.1%に認識率が改善され、誤り率が約40%改善さ
れた。
Using the configuration of the present embodiment, an experiment for recognizing voice data of 11 people who uttered 100 words was performed.
The speaker adaptation learning of the phoneme standard pattern was performed using 20 words different from the evaluation utterance data. The average recognition rate of 11 persons was 93.2% before performing the speaker adaptation based on the present embodiment, whereas the average recognition rate after the speaker adaptation based on the present embodiment was 93.2%.
The recognition rate was improved to 6.1%, and the error rate was improved by about 40%.

【0039】なお、本実施例においては、単語マッチン
グ時の特徴量パラメータである音素類似度の計算に、音
響パラメータの時間パタンを用い、距離尺度として共分
散行列を共通化したマハラノビス距離を用いたが、HM
Mから構成される音素モデルから音素類似度を計算する
こともできる。
In the present embodiment, a time pattern of acoustic parameters is used for calculating a phoneme similarity which is a feature parameter at the time of word matching, and a Mahalanobis distance with a common covariance matrix is used as a distance measure. But HM
The phoneme similarity can also be calculated from the phoneme model composed of M.

【0040】また、本実施例においては、認識対象を、
単語としたが、これを連続発声を認識する際に利用する
ことも可能である。
In this embodiment, the recognition target is
Although a word is used, it can be used when recognizing a continuous utterance.

【0041】[0041]

【発明の効果】本発明によれば、音素標準パターンと音
響パラメータ時系列との照合によって得られる音素類似
度を単語パターンマッチング時の特徴量パラメータと
し、音素標準パターンの話者適応学習を行うことによ
り、少ないメモリー量によって、少量の話者適応学習デ
ータでも効率的な話者適応を実現できる、高性能な音声
認識装置を提供できるという効果を得る。
According to the present invention, the speaker adaptation learning of the phoneme standard pattern is performed by using the phoneme similarity obtained by collating the phoneme standard pattern with the acoustic parameter time series as a feature parameter at the time of word pattern matching. Accordingly, an effect is obtained that a high-performance speech recognition device that can realize efficient speaker adaptation with a small amount of memory and a small amount of speaker adaptation learning data can be provided.

【図面の簡単な説明】[Brief description of the drawings]

【図1】本発明の実施の形態における音声認識装置のブ
ロック構成図
FIG. 1 is a block diagram of a speech recognition apparatus according to an embodiment of the present invention.

【図2】本発明の実施例における音素類似度時系列を示
す概念図
FIG. 2 is a conceptual diagram showing a phoneme similarity time series in the embodiment of the present invention.

【図3】従来技術の音声認識装置を示すブロック図FIG. 3 is a block diagram showing a conventional speech recognition apparatus.

【符号の説明】[Explanation of symbols]

101 音響分析部 102 音素類似度計算部 103 不特定話者用音素標準パターン格納部 104 パターンマッチング部 105 単語辞書 106 不特定話者用音声片標準パターン格納部 107 高音素類似度フレーム情報抽出部 108 適応学習パターン抽出部 109 音素標準パターン適応部 110 特定話者用音素標準パターン格納部 Reference Signs List 101 acoustic analysis unit 102 phoneme similarity calculation unit 103 phoneme standard pattern storage unit for unspecified speaker 104 pattern matching unit 105 word dictionary 106 speech unit standard pattern storage unit for unspecified speaker 107 high phoneme similarity frame information extraction unit 108 Adaptive learning pattern extraction unit 109 Phoneme standard pattern adaptation unit 110 Phoneme standard pattern storage unit for specific speaker

Claims (3)

【特許請求の範囲】[Claims] 【請求項1】 入力音声と不特定話者用の音素標準パタ
ーンとの照合で得られる音素類似度を認識の特徴量パラ
メータとする音声認識方法において、音素標準パターン
に対して話者適応学習を行うことを特徴とする音声認識
方法。
1. A speech recognition method in which a phoneme similarity obtained by matching an input speech with a phoneme standard pattern for an unspecified speaker is used as a feature parameter of recognition. A voice recognition method characterized by performing.
【請求項2】 入力音声に対して分析時間毎にm個の特
徴パラメータ系列を求め、分析時間毎にn種類の音素標
準パターンとマッチングを行い、分析時間毎に求めたn
個の類似度と、予め作成しておいた音素類似度の時系列
で構成される単語標準パターンとマッチングすることに
より単語を認識する音声認識方法において、前記単語標
準パターンから音素類似度が高いフレームを抽出してお
き、入力音声と単語標準パターンのマッチングで得られ
る時間対応の結果から、単語標準パターンの音素類似度
が高いフレームに対応した入力音声フレームを求め、こ
のフレームの入力音声の特徴パラメータを音素標準パタ
ーン学習データとして抽出し、前記音素標準パターンと
前記音素標準パターン学習データとを統合し、新たな音
素標準パターンを作成することを特徴とする音声認識方
法。
2. An m-number of feature parameter sequences are obtained for each analysis time with respect to an input voice, matching is performed with n types of phoneme standard patterns at each analysis time, and n is obtained for each analysis time.
In a speech recognition method for recognizing a word by matching a word standard pattern composed of a time series of phoneme similarities prepared in advance and a phoneme similarity, a frame having a high phoneme similarity from the word standard pattern is used. Is extracted, and an input speech frame corresponding to a frame having a high phoneme similarity of the word standard pattern is obtained from the time correspondence result obtained by matching the input speech with the word standard pattern, and the characteristic parameters of the input speech of this frame are obtained. Is extracted as phoneme standard pattern learning data, and the phoneme standard pattern and the phoneme standard pattern learning data are integrated to create a new phoneme standard pattern.
【請求項3】 入力音声に対して分析時間毎にm個の特
徴パラメータ系列を求め、分析時間毎にn種類の音素標
準パターンとマッチングを行い、分析時間毎に求めたn
個の類似度と、予め作成しておいた音素類似度の時系列
で構成される単語標準パターンとマッチングすることに
より単語を認識する音声認識装置において、前記単語標
準パターンから音素類似度が高いフレームを抽出する高
音素類似度フレーム情報抽出手段と、入力音声と単語標
準パターンのマッチングで得られる時間対応の結果か
ら、単語標準パターンの音素類似度が高いフレームに対
応した入力音声フレームを求め、個のフレームの入力音
声の特徴パラメータを音素標準パターン学習データとし
て抽出する適応学習パターン抽出手段と、前記音素標準
パターンと前記音素標準パターン学習データとを統合
し、新たな音素標準パターンを作成する音素標準パター
ン適応手段とを具備することを特徴とする音声認識装
置。
3. An m-number of feature parameter sequences are obtained for each analysis time with respect to an input voice, matching is performed with n kinds of phoneme standard patterns at each analysis time, and n is obtained for each analysis time.
In a speech recognition apparatus for recognizing a word by matching a word standard pattern composed of a time series of phoneme similarities prepared in advance and a phoneme similarity, a frame having a high phoneme similarity from the word standard pattern A high-phoneme-similarity frame information extracting means for extracting the input speech frame corresponding to a frame having a high phoneme-similarity of the word standard pattern from the result of the time correspondence obtained by matching the input speech with the word standard pattern. Adaptive learning pattern extraction means for extracting feature parameters of the input speech of the frame as phoneme standard pattern learning data, and a phoneme standard for integrating the phoneme standard pattern and the phoneme standard pattern learning data to create a new phoneme standard pattern A speech recognition device comprising: a pattern adaptation unit.
JP06279098A 1998-03-13 1998-03-13 Voice recognition method and voice recognition device Expired - Lifetime JP3289670B2 (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP06279098A JP3289670B2 (en) 1998-03-13 1998-03-13 Voice recognition method and voice recognition device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP06279098A JP3289670B2 (en) 1998-03-13 1998-03-13 Voice recognition method and voice recognition device

Publications (2)

Publication Number Publication Date
JPH11259086A true JPH11259086A (en) 1999-09-24
JP3289670B2 JP3289670B2 (en) 2002-06-10

Family

ID=13210505

Family Applications (1)

Application Number Title Priority Date Filing Date
JP06279098A Expired - Lifetime JP3289670B2 (en) 1998-03-13 1998-03-13 Voice recognition method and voice recognition device

Country Status (1)

Country Link
JP (1) JP3289670B2 (en)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2016036163A3 (en) * 2014-09-03 2016-04-21 삼성전자 주식회사 Method and apparatus for learning and recognizing audio signal
CN111276127A (en) * 2020-03-31 2020-06-12 北京字节跳动网络技术有限公司 Voice awakening method and device, storage medium and electronic equipment
CN115966210A (en) * 2022-12-06 2023-04-14 四川启睿克科技有限公司 Phoneme selection method and device for voiceprint recognition

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2016036163A3 (en) * 2014-09-03 2016-04-21 삼성전자 주식회사 Method and apparatus for learning and recognizing audio signal
CN111276127A (en) * 2020-03-31 2020-06-12 北京字节跳动网络技术有限公司 Voice awakening method and device, storage medium and electronic equipment
CN115966210A (en) * 2022-12-06 2023-04-14 四川启睿克科技有限公司 Phoneme selection method and device for voiceprint recognition

Also Published As

Publication number Publication date
JP3289670B2 (en) 2002-06-10

Similar Documents

Publication Publication Date Title
Loizou et al. High-performance alphabet recognition
JP3114468B2 (en) Voice recognition method
EP2048655A1 (en) Context sensitive multi-stage speech recognition
EP1355295A2 (en) Speech recognition apparatus, speech recognition method, and computer-readable recording medium in which speech recognition program is recorded
Vadwala et al. Survey paper on different speech recognition algorithm: challenges and techniques
KR20060066483A (en) Feature Vector Extraction Method for Speech Recognition
JP3798530B2 (en) Speech recognition apparatus and speech recognition method
JP3444108B2 (en) Voice recognition device
JP3289670B2 (en) Voice recognition method and voice recognition device
JP2003177779A (en) Speaker learning method for speech recognition
JP2943445B2 (en) Voice recognition method
JP2943473B2 (en) Voice recognition method
JP3277522B2 (en) Voice recognition method
JPH09114482A (en) Speaker adaptation method for speech recognition
JPH0786758B2 (en) Voice recognizer
JP2692382B2 (en) Speech recognition method
JP3291073B2 (en) Voice recognition method
JP3285047B2 (en) Speech recognition device for unspecified speakers
JP3115016B2 (en) Voice recognition method and apparatus
JP3105708B2 (en) Voice recognition device
JPS6336678B2 (en)
JP3704080B2 (en) Speech recognition method, speech recognition apparatus, and speech recognition program
JP2827590B2 (en) Voice recognition device
JP3357752B2 (en) Pattern matching device
JPH05323990A (en) Speaker recognition method

Legal Events

Date Code Title Description
FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20080322

Year of fee payment: 6

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20090322

Year of fee payment: 7

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20100322

Year of fee payment: 8

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20110322

Year of fee payment: 9

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20110322

Year of fee payment: 9

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20120322

Year of fee payment: 10

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20130322

Year of fee payment: 11

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20130322

Year of fee payment: 11

FPAY Renewal fee payment (event date is renewal date of database)

Free format text: PAYMENT UNTIL: 20140322

Year of fee payment: 12

S111 Request for change of ownership or part of ownership

Free format text: JAPANESE INTERMEDIATE CODE: R313113

S533 Written request for registration of change of name

Free format text: JAPANESE INTERMEDIATE CODE: R313533

R350 Written notification of registration of transfer

Free format text: JAPANESE INTERMEDIATE CODE: R350

EXPY Cancellation because of completion of term