JP2000352989A - ユーザが文字列の発音を設定することを可能にするためにコンピュータ上で実行される方法 - Google Patents
ユーザが文字列の発音を設定することを可能にするためにコンピュータ上で実行される方法Info
- Publication number
- JP2000352989A JP2000352989A JP2000130595A JP2000130595A JP2000352989A JP 2000352989 A JP2000352989 A JP 2000352989A JP 2000130595 A JP2000130595 A JP 2000130595A JP 2000130595 A JP2000130595 A JP 2000130595A JP 2000352989 A JP2000352989 A JP 2000352989A
- Authority
- JP
- Japan
- Prior art keywords
- pronunciation
- user
- characters
- character string
- word
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
- 238000000034 method Methods 0.000 title claims abstract description 58
- 238000006243 chemical reaction Methods 0.000 claims abstract description 8
- 230000008569 process Effects 0.000 claims description 17
- 230000008859 change Effects 0.000 claims description 16
- 238000004519 manufacturing process Methods 0.000 claims 1
- 230000008901 benefit Effects 0.000 abstract description 3
- 238000012360 testing method Methods 0.000 description 9
- 238000010586 diagram Methods 0.000 description 6
- 230000005236 sound signal Effects 0.000 description 2
- 238000013519 translation Methods 0.000 description 2
- 230000014616 translation Effects 0.000 description 2
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/06—Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
- G10L15/063—Training
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L13/00—Speech synthesis; Text to speech systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/04—Segmentation; Word boundary detection
Landscapes
- Engineering & Computer Science (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Artificial Intelligence (AREA)
- Machine Translation (AREA)
- Electrically Operated Instructional Devices (AREA)
- Document Processing Apparatus (AREA)
Abstract
音声認識技術の利益をより多く享受することができるよ
うな簡単な方法で発音の規則および選択肢(オプショ
ン)を表示するユーザインタフェースを実現する。 【解決手段】 ユーザは、発音の設定または変更をした
い単語の入力または選択を行う。単語内の与えられた文
字または文字群の発音を設定するために、ユーザが文字
(群)を選択すると、全部または一部の発音が、選択さ
れた文字の可能な発音とほぼ同一である普通の単語のリ
ストが提示される。好ましくは、サンプルの普通単語の
リストは、通常使用における相関の頻度に基づいて順序
づけられ、最も普通のものがデフォルトのサンプル単語
として指定される。ユーザには、まず、リスト内の単語
のうち、選択される可能性が高い単語からなるサブセッ
トが提示される。
Description
ースに関し、特に、テキスト−音声変換システムおよび
自動音声認識システムとともに用いられるグラフィカル
ユーザインタフェースに関する。
来の入出力装置の代替または補助として重要になってい
る。これは特に、基礎となる使用されるソフトウェア方
法論と、必要な処理および記憶をサポートするハードウ
ェアコンポーネントとにおける改良および進歩が続くと
ともに、正しくなるであろう。これらの技術がますます
一般的に利用可能になり、大衆市場で使用されるように
なるにつれて、音声認識・生成システムを初期化し変更
する際に用いられる技術の改良が必要となる。
み上げられるテキストのファイルをユーザが処理するこ
とを可能にするいくつかの製品が存在する。さらに、音
声言語を入力として処理し、単語やコマンドを識別し、
アクションやイベントをトリガするために用いられるい
くつかのソフトウェア製品がある。従来のいくつかの製
品では、ユーザが、単語を辞書に追加し、辞書中の単語
発音を変更し、また、テキスト−音声エンジンにより生
成される音を変更することが可能である。
ァイルが作成される各言語の文法、発音、および言語規
則についての専門的な情報を理解し使用することが要求
される。さらに、これらの製品の一部では、発音を表現
する手段は、他の分野では一般に使用されていない独特
の発音キーによるマークアップ言語の習得を必要とす
る。
声変換技術および自動音声認識技術を、一般大衆にとっ
て、不自由な(フレキシブルでない)近寄り難いものに
している。これらの製品では、ユーザは、言語規則およ
びプログラミング技術の両方の専門家になることが要求
される。不自由さは、部分的には、これらの製品が、対
象言語の一般的な規則を用いて、コンテクスト(例え
ば、方言の形での地理的コンテクスト)や、名前のよう
な単語の発音に関する個別の好みとは無関係に発音を決
定するために生じる。
の翻訳について得られる結果はあまり満足なものではな
い。これらの製品は、頭字語、固有名、技術用語、商
標、あるいは他言語からとられた単語のような多くのタ
イプの単語に関してはあまりうまく動作しない。また、
これらの製品は、句や文の中での位置に依存する単語の
発音の変化を考慮する際にうまく動作しない(例えば、
addressという単語は、名詞として使用されると
きと動詞として使用されるときでは別様に発音され
る)。
ザがテキスト−音声変換技術や音声認識技術の利益をよ
り多く享受することができるような簡単な方法で発音の
規則および選択肢(オプション)を表示するユーザイン
タフェースの方法およびシステムが必要とされている。
換システムおよび音声認識システムにおける上記の課題
を解決することである。
設定し変更するための単純で直感的なユーザインタフェ
ースを提供することである。
は通常用いられない、あるいは規則に違反する音または
文字の群を、テキスト−音声変換システムおよび音声認
識システムで使用できるようにすることである。
他の目的は、周知の発音とともにオーディオキューおよ
び普通の(ありふれた)単語に基づいて単語および単語
の部分をどのように発音するかについてユーザが決定す
ることを可能にする方法およびユーザインタフェースに
よって、達成される。
変更をしたい単語の入力または選択を行う。単語内の与
えられた文字または文字群の発音を設定するために、ユ
ーザが文字(群)を選択すると、全部または一部の発音
が、選択された文字の可能な発音とほぼ同一である普通
の単語のリストが提示される。好ましくは、サンプルの
普通単語のリストは、通常使用における相関の頻度に基
づいて順序づけられ、最も普通のものがデフォルトのサ
ンプル単語として指定される。ユーザには、まず、リス
ト内の単語のうち、選択される可能性が高い単語からな
るサブセットが提示される。
語についていくつかの異なる発音を辞書に記憶すること
により、コンテクスト上の相違や個別の好みが許容され
る。
いて複数の辞書を記憶するが、ユーザは、特別の単語、
単語の部分、および翻訳を考慮に入れるために、さまざ
まな辞書から発音を選択することができる。その結果、
ユーザは、システムで利用可能な任意の音を有する単語
を、その音が一般的に言語の規則に従う文字群に対応し
ない場合でも、作成し記憶させることが可能である。
て、本発明の実施例によれば、ユーザは、単語を音節も
しくは音節状の文字群または単語サブコンポーネントに
容易に分解し、どの音節にアクセントをおくかを指定す
ることができる。これは、与えられた言語の規則がその
ような文字群(グループ分け)を音節としていない場合
でも可能である。ここで、単語の音節とは、従来の音節
とともに、他のグループ分けも指す。
xt to speech)・自動音声認識(ASR:automated sp
eech recognition)システム10を図1に示す。システ
ム10は、マイクロコントローラまたはマイクロプロセ
ッサ14と、メモリ装置16とを有するコンピュータ化
装置またはシステム(コンピュータ)12を含む。シス
テム10はさらに、ディスプレイ装置18、スピーカ2
0、入力装置22およびマイクロフォン24を有する。
これらのすべてのコンポーネントは通常のものであり、
当業者に周知であって、ここでこれ以上詳細に説明する
必要はない。
タ装置12に組み込まれても、コンピュータ12から離
れて配置されネットワークなどの接続を通じてアクセス
可能としてもよい)は、本発明によるいくつかのプログ
ラムおよびデータファイルを記憶する。発音選択プログ
ラム26は、ここで説明するユーザインタフェースの生
成のためにマイクロコントローラ14上で実行される
と、ユーザの入力の処理、ならびに、データベース28
および30からのデータの取得を行う。辞書データベー
ス28は、システム10によって処理される各言語ごと
に1つずつの、いくつかのデータベースまたはデータフ
ァイルからなり、文字列(STRING)およびそれに
対応する1つ以上の発音(PRON)を記憶する。発音
データベース30は、各言語ごとに1つずつの、いくつ
かのデータベースまたはデータファイルからなる。この
データベースの各レコードは、文字または文字群(CH
AR)と、その文字が発音されるのとほぼ同一の方法で
発音される文字を含む、その文字(群)に対応するいく
つかのサンプル単語(SAMPLE WORD)とを有
する。サンプル単語は、発音データベース30を作成す
る際に、言語の文法的および言語学的規則に基づいて選
択される。好ましくは、各文字または文字群(例えば二
重母音)に対するサンプル単語は、その文字の発音にお
いて、一般に、より普通の使用法からあまり普通でない
使用法へと順序づけられる。
ース30は2つのデータベースとして図示されている
が、これらは、1つのデータファイル、あるいは、ここ
で説明するような発音データの取得を容易にし、与えら
れたアプリケーションまたは使用法の需要を満たすのに
要求されるその他の任意のフォーマットに構造化するこ
とが可能である。
されたTTSモジュール32およびASRモジュール3
4を含む。これらのモジュールは当業者に周知であり、
例えば、IBMから市販されているViaVoice(R)ソフト
ウェアプログラムを含む。これらのモジュール32およ
び34は、ディジタルデータとして記憶されているテキ
ストを、スピーカ20によって出力するためにオーディ
オ信号に変換するとともに、マイクロフォン24を通じ
て入力されるオーディオ信号をディジタルデータに変換
する。これらのモジュールは、辞書データベース28に
記憶されている発音データを取得し利用する。
音データをユーザが容易に変更することを可能にする方
法は、撥音選択プログラム26によって実行される。そ
の概略を図2に示し、詳細には図3および図4に示す。
図2において、ステップ50で、本発明によれば、文字
列(例えば単語、名前など)がディスプレイ装置18に
表示される。ステップ52で、ユーザは、入力装置22
を用いて、文字列から1つ以上の文字を選択する。理解
されるように、発音変化は、母音のような個別の文字に
関連することも、ou、ch、thまたはghのような
文字群に関連することもある。ステップ54で、プログ
ラム26は、発音データベース30に問合せを行い、選
択された文字または文字群に関連づけられたサンプル単
語を検索・取得する。その文字または文字群が発音デー
タベースに存在しない場合、エラーメッセージが送られ
るか、または、それらの文字のうちの1つに対するサン
プル単語が取得される。ステップ56で、サンプル単語
のうちの一部または全部が表示され、ステップ58で、
ユーザはそれらの単語のうちの1つを選択する。次に、
ステップ60で、プログラム26は、選択された文字
(群)の発音を提供するために、そのサンプル単語を用
いて、文字列に対する発音データを生成する。ステップ
62で、文字列および発音データが辞書データベース2
8に記憶される。文字列は、TTSモジュール32の出
力から聞こえるように出力されるか、または、ASRモ
ジュール34の話者照合または発声照合のために使用さ
れる。
スについてさらに詳細に図3および図4に示す。このプ
ロセス中に使用されるユーザインタフェースの実施例を
図6〜図10に示す。図6に示すように、ディスプレイ
装置18に表示されるインタフェース190は、 ・選択された文字の手入力または表示のための入力ボッ
クス200 ・単語が選択されるまで無効であるテスト(TEST)
ボタン202 ・同様に単語が選択されるまで無効である変更(MOD
IFY)ボタン204 ・選択肢SOUND(音)、ACCENT(アクセン
ト)およびSYLLABLE(音節)(またはGROU
PING(グループ分け))からなる選択リスト206 ・ワークスペース208 を含む。
は、異なる言語を表す複数の辞書データベースおよび発
音データベースを含む。図3において、ステップ70
で、ユーザはそれらの言語のうちの1つを選択し、ステ
ップ72で、プログラム26は、選択された言語の辞書
を開く。単語またはその他の文字列を選択するため、ス
テップ74で、ユーザは、選択された辞書をブラウズ
(閲覧)することを選択することができる。この場合、
ユーザは、データベース76から既存の単語を選択す
る。そうでない場合、ステップ78で、ユーザは、入力
ボックス200にタイプ入力することなどによって、単
語を入力する。
ボタン202を選択することによって、単語の発音をテ
ストするかどうかを選択することができる。単語発音の
テストの処理については、図5を参照して後述する。
04を選択することによって、単語の発音を変更するこ
とを選択することができる。選択しない場合、ステップ
84で、ユーザは、ダイアログ190内のOKボタンを
選択することによって、その単語および現在の発音を記
憶させることができる。ステップ86で、単語が辞書デ
ータベース28内の既存の単語でない場合、ステップ8
8で、単語および発音のデータが辞書に記憶される。図
5に関して後述するように、未変更の単語に対する発音
データは、選択された言語の規則に基づくデフォルトの
発音を用いて生成される。単語が既に存在する場合、ス
テップ90で、新しい発音データが辞書内の単語ととも
に記憶される。代わりの発音が、コンテクスト環境から
参照されることも可能である。
ト206内の3つの選択肢が利用可能である。
れている)は、個々の文字に分解され、ワークスペース
208にコピーされる(図7を参照)。ワークスペース
208はさらに、現在の発音に対する音節の分かれ目
(ブレーク、ブレークポイント)(ワークスペース20
8内のダッシュ)およびアクセントマーク(ワークスペ
ース208内のアポストロフィ)を示す。
クを変更することを選択した場合、ブレークポイント記
号210が表示される(図8を参照)。ステップ94
で、記号210は、所望の音節ブレークポイントを指定
するためにユーザが移動させることが可能である。ステ
ップ96で、プログラム26は、既存の音節を、選択さ
れたブレークポイントで2つの音節に分解する。
更することを選択した場合、アクセントタイプ選択アイ
コン群212がインタフェース190に表示される(図
9を参照)。アイコン群212は、3個のアイコン、す
なわち、第1アクセント(主強勢)アイコン212a、
第2アクセント(副強勢)アイコン212b、および無
アクセント(無強勢)212cを含む。ステップ100
で、ユーザは、これらのアイコンのうちの1つをクリッ
クすることによって、アクセントレベルを選択する。次
に、ステップ102で、ユーザは、例えば、ワークスペ
ース208内で、音節の直後のボックスを選択すること
によって、その音節を選択する。ステップ104で、プ
ログラム26は、選択されたアクセントレベルを選択さ
れた音節に指定する。さらに、選択された言語の規則に
従って残りのアクセントを調整することが可能である。
例えば、言語が1つの第1アクセントのある音節を規定
しており、ユーザが第2音節を第1アクセントのために
選択した場合、プログラムは、最初の第1アクセントを
第2アクセントに変更するか、または、残りのすべての
アクセントを完全に削除することが可能である。
ーザがリスト206内の文字音を変更することを選択し
た場合、ステップ108で、ユーザは、ワークスペース
208内の1つ以上の文字を選択する。ステップ110
で、プログラム26は、選択された言語に対する発音デ
ータベース30から、その発音またはその一部の発音が
選択した文字に関連づけられるサンプル単語を検索・取
得する。これらの単語は、単語リスト214に表示され
る(図10を参照)。ステップ112で、選択された文
字に対するデフォルトの発音を表すサンプル単語が強調
表示される。図10では、選択された文字iの発音とし
て、サンプル単語buyが単語リスト214内で強調表
示されている。同じく図10に示されるように、2、3
個のサンプル単語のみが単語リスト214に示され、ユ
ーザが追加の単語を見たり聞いたりするのはオプション
とすることが可能である。
のうちの1つを選択した場合、ステップ116で、選択
された単語の発音データ、またはその一部が、ワークス
ペース208に含まれる選択単語の選択された文字に関
連づけられる。その後、変更された単語は、上記のプロ
セスに従ってさらに変更することも、記憶させることも
可能であり、あるいは、後述のようにテストすることも
可能である。
れるように、英語を含むほとんどの言語は、他の言語か
らとられた単語を含むため、ステップ118で、ユーザ
には、他の言語からの選択された文字に対する発音を選
択するオプション(例えば、単語リスト214におい
て、moreを選択した後)が与えられる。次に、ステ
ップ120で、ユーザが所望の言語を選択すると、ステ
ップ122で、プログラム26は、その選択された言語
に対する発音データベースファイル30から、選択され
た文字に関連づけられたサンプル単語を検索・取得す
る。その後、サンプル単語は、上記のようにユーザの選
択に対して提示される。
とが可能な、簡単でフレキシブルなプロセスが実現され
る。プロセスの容易さおよびフレキシビリティの一例と
して、図10で選択された単語Michaelは、aと
eの間、および、iとchの間に音節ブレークを追加
し、新しい音節chaに第1アクセントを置き、普通の
単語に基づいてi、ch(例えば、ヘブライ語辞書か
ら)、aおよびeに対する適当な発音を選択することに
より、英語発音「MiK’−el」から、ヘブライ名
「Mee−cha’−el」に変更することが可能であ
る。文法的および言語学的な専門知識は不要である。
示す。ステップ140で、単語が既に辞書データベース
28に含まれている場合、ステップ142で、記憶され
ている発音が取得される。単語に対して複数の発音が存
在する場合、ユーザに、1つを選択するか、デフォルト
を使用するかを尋ねることが可能である。単語が存在し
ない場合、各文字または文字群ごとに、ステップ144
で、ユーザがプログラム26を用いて発音を選択する
と、ステップ146で、その発音データが取得され、そ
れ以外の場合、ステップ148で、例えばデフォルトの
発音が選択される。ステップ150で、すべての文字が
調べられた後、ステップ152で、プログラム26は、
取得した文字発音を用いて単語の発音を生成する。最後
に、ステップ154で、TTSモジュールは、取得また
は生成された単語発音の可聴表現を出力する。
する複数の発音が可能であるため、TTSモジュール
は、その単語をどの発音にするかを指定しなければなら
ない。TTSモジュールは、単語が使用されているコン
テクストに基づいて発音を指定することができる。例え
ば、発音は、ネットワーク上のユーザのような目的に関
連づけることが可能であり、その場合、特定のユーザ宛
のメッセージでは、発音が正しく選択されることにな
る。別の例として、TTSモジュールは、名詞対動詞の
ような単語の使用法を指定し、それに従って適当な発音
を選択することも可能である。
専門家ユーザがテキスト−音声変換技術や音声認識技術
の利益をより多く享受することができるような簡単な方
法で発音の規則および選択肢(オプション)を表示する
ユーザインタフェースの方法およびシステムが実現され
る。
ある。
が単語の発音を変更することを可能にするプロセスの概
略を示す流れ図である。
変更することを可能にするプロセスをさらに詳細に示す
流れ図である。
変更することを可能にするプロセスをさらに詳細に示す
流れ図である。
である。
フェースを示すスクリーンディスプレイの図である。
フェースを示すスクリーンディスプレイの図である。
フェースを示すスクリーンディスプレイの図である。
フェースを示すスクリーンディスプレイの図である。
タフェースを示すスクリーンディスプレイの図である。
(ASR)システム 12 コンピュータ 14 マイクロコントローラ 16 メモリ装置 18 ディスプレイ装置 20 スピーカ 22 入力装置 24 マイクロフォン 26 発音選択プログラム 28 辞書データベース 30 発音データベース 32 TTSモジュール 34 ASRモジュール 190 インタフェース 200 入力ボックス 202 テスト(TEST)ボタン 204 変更(MODIFY)ボタン 206 選択リスト 208 ワークスペース 210 ブレークポイント記号 212 アクセントタイプ選択アイコン群 214 単語リスト
Claims (23)
- 【請求項1】 ユーザが文字列の発音を設定することを
可能にするためにコンピュータ上で実行される方法にお
いて、 文字列内の1つ以上の文字をユーザに選択させるステッ
プと、 選択された1つ以上の文字の可能な発音を表す単語また
は単語部分の複数のサンプルを、コンピュータによりア
クセス可能なデータベースから取得し、取得したサンプ
ルを表示する取得表示ステップと、 表示されたサンプルのうちの1つをユーザに選択させる
ステップと、 前記選択された1つ以上の文字に、ユーザによって選択
されたサンプルに対応する発音を割り当てて、前記文字
列を含む第1発音レコードを記憶するステップとを有す
ることを特徴とする、ユーザが文字列の発音を設定する
ことを可能にするためにコンピュータ上で実行される方
法。 - 【請求項2】 前記選択された1つ以上の文字に対する
発音として、ユーザによって選択されたサンプルによっ
て表される発音を用いて前記文字列の発音を生成するス
テップと、 生成された発音を可聴出力するステップとをさらに有す
ることを特徴とする請求項1に記載の方法。 - 【請求項3】 前記生成された発音を可聴出力した後、
もう1つの表示されたサンプルをユーザに選択させるス
テップをさらに有することを特徴とする請求項2に記載
の方法。 - 【請求項4】 第2の表示されたサンプルをユーザに選
択させるステップと、 前記選択された1つ以上の文字に、ユーザによって選択
された第2のサンプルによって表される発音を割り当て
て、前記文字列を含む第2発音レコードを記憶するステ
ップとをさらに有することを特徴とする請求項1に記載
の方法。 - 【請求項5】 前記文字列を含むテキストファイルの可
聴出力を生成するテキスト−音声変換プロセス中に、前
記第1および第2発音レコードのうちの一方を選択する
ステップをさらに有することを特徴とする請求項4に記
載の方法。 - 【請求項6】 前記第1および第2発音レコードを第1
および第2オブジェクトにそれぞれ関連づけるステップ
と、 前記第1および第2オブジェクトのうちの一方を選択す
るステップとをさらに有し、 前記第1および第2発音レコードのうちの一方を選択す
るステップは、選択されたオブジェクトに関連づけられ
た発音レコードを選択することを含むことを特徴とする
請求項5に記載の方法。 - 【請求項7】 音声認識プロセス中に、ユーザによる前
記文字列の発音を認識するステップと、 前記第1および第2発音レコードのうち、認識された発
音に最も良く一致するほうを選択するステップとをさら
に有することを特徴とする請求項4に記載の方法。 - 【請求項8】 前記第1および第2発音レコードを第1
および第2オブジェクトにそれぞれ関連づけるステップ
と、 前記第1および第2オブジェクトのうち、選択された発
音レコードに関連づけられたほうを選択するステップと
をさらに有することを特徴とする請求項7に記載の方
法。 - 【請求項9】 前記文字列の一部を別個の音節としてユ
ーザに指定させるステップをさらに有し、 前記第1発音レコードを記憶するステップは、指定され
た別個の音節を表すデータを記憶することを含むことを
特徴とする請求項1に記載の方法。 - 【請求項10】 前記文字列の一部をアクセントに対応
させるようにユーザに指定させるステップをさらに有
し、 前記第1発音レコードを記憶するステップは、指定され
たアクセントを表すデータを記憶することを含むことを
特徴とする請求項1に記載の方法。 - 【請求項11】 ユーザによる入力として前記文字列を
受け取るステップをさらに有することを特徴とする請求
項1に記載の方法。 - 【請求項12】 コンピュータによりアクセス可能な辞
書データベースから前記文字列をユーザに選択させるス
テップをさらに有することを特徴とする請求項1に記載
の方法。 - 【請求項13】 所望の言語をユーザに選択させるステ
ップをさらに有し、 前記取得表示ステップは、 複数の言語データベースから前記所望の言語のデータベ
ースを選択するステップと、 選択されたデータベースからサンプルを取得するステッ
プとを含むことを特徴とする請求項1に記載の方法。 - 【請求項14】 前記選択された1つ以上の文字に対す
る第2言語をユーザに選択させるステップと、 選択された第2言語に対応する第2データベースから追
加の単語サンプルを取得するステップとをさらに有する
ことを特徴とする請求項1に記載の方法。 - 【請求項15】 実行されると、ユーザが文字列の発音
を設定することを可能にするグラフィカルユーザインタ
フェース方法をコンピュータに実行させるプログラムコ
ードを記憶したコンピュータ可読媒体を含む製品におい
て、前記方法は、 文字列内の1つ以上の文字をユーザに選択させるステッ
プと、 選択された1つ以上の文字の可能な発音を表す単語また
は単語部分の複数のサンプルを、コンピュータによりア
クセス可能なデータベースから取得し、取得したサンプ
ルを表示する取得表示ステップと、 表示されたサンプルのうちの1つをユーザに選択させる
ステップと、 前記選択された1つ以上の文字に、ユーザによって選択
されたサンプルに対応する発音を割り当てて、前記文字
列を含む第1発音レコードを記憶するステップとを有す
ることを特徴とする、コンピュータ可読媒体を含む製
品。 - 【請求項16】 前記方法は、 前記選択された1つ以上の文字に対する発音として、ユ
ーザによって選択されたサンプルによって表される発音
を用いて前記文字列の発音を生成するステップと、 生成された発音を可聴出力するステップとをさらに有す
ることを特徴とする請求項15に記載の製品。 - 【請求項17】 前記方法は、 前記生成された発音を可聴出力した後、もう1つの表示
されたサンプルをユーザに選択させるステップをさらに
有することを特徴とする請求項16に記載の製品。 - 【請求項18】 前記方法は、 第2の表示されたサンプルをユーザに選択させるステッ
プと、 前記選択された1つ以上の文字に、ユーザによって選択
された第2のサンプルによって表される発音を割り当て
て、前記文字列を含む第2発音レコードを記憶するステ
ップとをさらに有することを特徴とする請求項15に記
載の製品。 - 【請求項19】 前記方法は、 前記文字列を含むテキストファイルの可聴出力を生成す
るテキスト−音声変換プロセス中に、前記第1および第
2発音レコードのうちの一方を選択するステップをさら
に有することを特徴とする請求項18に記載の製品。 - 【請求項20】 前記方法は、 前記第1および第2発音レコードを第1および第2オブ
ジェクトにそれぞれ関連づけるステップと、 前記第1および第2オブジェクトのうちの一方を選択す
るステップとをさらに有し、 前記第1および第2発音レコードのうちの一方を選択す
るステップは、選択されたオブジェクトに関連づけられ
た発音レコードを選択することを含むことを特徴とする
請求項19に記載の製品。 - 【請求項21】 前記方法は、 音声認識プロセス中に、ユーザによる前記文字列の発音
を認識するステップと、 前記第1および第2発音レコードのうち、認識された発
音に最も良く一致するほうを選択するステップとをさら
に有することを特徴とする請求項18に記載の製品。 - 【請求項22】 前記方法は、 前記第1および第2発音レコードを第1および第2オブ
ジェクトにそれぞれ関連づけるステップと、 前記第1および第2オブジェクトのうち、選択された発
音レコードに関連づけられたほうを選択するステップと
をさらに有することを特徴とする請求項21に記載の製
品。 - 【請求項23】 ユーザが文字列の発音を変更すること
を可能にするグラフィカルユーザインタフェースにおい
て、 メモリ装置上に記憶され、複数の第1文字列および対応
する発音レコードを含む辞書データベースと、 メモリ装置上に記憶され、それぞれが1つ以上の文字を
含む複数の第2文字列を含み、各第2文字列は複数の単
語に関連づけられ、各単語は関連づけられた第2文字列
が発音されるのとほぼ同様に該単語内で発音される1つ
以上の文字を含む発音データベースと、 ユーザが、前記辞書データベースから第1文字列のうち
の1つを選択し、選択された文字列から1つ以上の文字
を選択し、前記発音データベース内の単語のうちの1つ
を選択することを可能にする入出力システムと、 選択された1つ以上の文字に、ユーザによって選択され
た単語サンプルに対応する発音を割り当てて、選択され
た第1文字列を含む発音レコードを生成するプログラマ
ブルコントローラとを有することを特徴とするグラフィ
カルユーザインタフェース。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US09/303,057 US7292980B1 (en) | 1999-04-30 | 1999-04-30 | Graphical user interface and method for modifying pronunciations in text-to-speech and speech recognition systems |
| US09/303057 | 1999-04-30 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JP2000352989A true JP2000352989A (ja) | 2000-12-19 |
| JP4237915B2 JP4237915B2 (ja) | 2009-03-11 |
Family
ID=23170358
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2000130595A Expired - Fee Related JP4237915B2 (ja) | 1999-04-30 | 2000-04-28 | ユーザが文字列の発音を設定することを可能にするためにコンピュータ上で実行される方法 |
Country Status (6)
| Country | Link |
|---|---|
| US (1) | US7292980B1 (ja) |
| EP (1) | EP1049072B1 (ja) |
| JP (1) | JP4237915B2 (ja) |
| KR (1) | KR100378898B1 (ja) |
| CA (1) | CA2306527A1 (ja) |
| DE (1) | DE60020773T2 (ja) |
Cited By (48)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2008180750A (ja) * | 2007-01-23 | 2008-08-07 | Oki Electric Ind Co Ltd | 音声ラベリング支援システム |
| WO2011089651A1 (ja) * | 2010-01-22 | 2011-07-28 | 三菱電機株式会社 | 認識辞書作成装置、音声認識装置及び音声合成装置 |
| JP2015512062A (ja) * | 2012-03-02 | 2015-04-23 | アップル インコーポレイテッド | 名前発音システム及び方法 |
| JP2015118222A (ja) * | 2013-12-18 | 2015-06-25 | 株式会社日立超エル・エス・アイ・システムズ | 音声合成システム及び音声合成方法 |
| US9668024B2 (en) | 2014-06-30 | 2017-05-30 | Apple Inc. | Intelligent automated assistant for TV user interactions |
| US9865248B2 (en) | 2008-04-05 | 2018-01-09 | Apple Inc. | Intelligent text-to-speech conversion |
| US9934775B2 (en) | 2016-05-26 | 2018-04-03 | Apple Inc. | Unit-selection text-to-speech synthesis based on predicted concatenation parameters |
| US9966060B2 (en) | 2013-06-07 | 2018-05-08 | Apple Inc. | System and method for user-specified pronunciation of words for speech synthesis and recognition |
| US9971774B2 (en) | 2012-09-19 | 2018-05-15 | Apple Inc. | Voice-based media searching |
| US9972304B2 (en) | 2016-06-03 | 2018-05-15 | Apple Inc. | Privacy preserving distributed evaluation framework for embedded personalized systems |
| US9986419B2 (en) | 2014-09-30 | 2018-05-29 | Apple Inc. | Social reminders |
| US10043516B2 (en) | 2016-09-23 | 2018-08-07 | Apple Inc. | Intelligent automated assistant |
| US10049675B2 (en) | 2010-02-25 | 2018-08-14 | Apple Inc. | User profiling for voice input processing |
| US10049663B2 (en) | 2016-06-08 | 2018-08-14 | Apple, Inc. | Intelligent automated assistant for media exploration |
| US10049668B2 (en) | 2015-12-02 | 2018-08-14 | Apple Inc. | Applying neural network language models to weighted finite state transducers for automatic speech recognition |
| US10067938B2 (en) | 2016-06-10 | 2018-09-04 | Apple Inc. | Multilingual word prediction |
| US10079014B2 (en) | 2012-06-08 | 2018-09-18 | Apple Inc. | Name recognition system |
| US10089072B2 (en) | 2016-06-11 | 2018-10-02 | Apple Inc. | Intelligent device arbitration and control |
| US10169329B2 (en) | 2014-05-30 | 2019-01-01 | Apple Inc. | Exemplar-based natural language processing |
| US10192552B2 (en) | 2016-06-10 | 2019-01-29 | Apple Inc. | Digital assistant providing whispered speech |
| US10223066B2 (en) | 2015-12-23 | 2019-03-05 | Apple Inc. | Proactive assistance based on dialog communication between devices |
| US10249300B2 (en) | 2016-06-06 | 2019-04-02 | Apple Inc. | Intelligent list reading |
| US10269345B2 (en) | 2016-06-11 | 2019-04-23 | Apple Inc. | Intelligent task discovery |
| US10297253B2 (en) | 2016-06-11 | 2019-05-21 | Apple Inc. | Application integration with a digital assistant |
| US10318871B2 (en) | 2005-09-08 | 2019-06-11 | Apple Inc. | Method and apparatus for building an intelligent automated assistant |
| US10354011B2 (en) | 2016-06-09 | 2019-07-16 | Apple Inc. | Intelligent automated assistant in a home environment |
| US10356243B2 (en) | 2015-06-05 | 2019-07-16 | Apple Inc. | Virtual assistant aided communication with 3rd party service in a communication session |
| US10366158B2 (en) | 2015-09-29 | 2019-07-30 | Apple Inc. | Efficient word encoding for recurrent neural network language models |
| US10410637B2 (en) | 2017-05-12 | 2019-09-10 | Apple Inc. | User-specific acoustic models |
| US10446143B2 (en) | 2016-03-14 | 2019-10-15 | Apple Inc. | Identification of voice inputs providing credentials |
| US10482874B2 (en) | 2017-05-15 | 2019-11-19 | Apple Inc. | Hierarchical belief states for digital assistants |
| US10490187B2 (en) | 2016-06-10 | 2019-11-26 | Apple Inc. | Digital assistant providing automated status report |
| US10509862B2 (en) | 2016-06-10 | 2019-12-17 | Apple Inc. | Dynamic phrase expansion of language input |
| US10521466B2 (en) | 2016-06-11 | 2019-12-31 | Apple Inc. | Data driven natural language event detection and classification |
| US10567477B2 (en) | 2015-03-08 | 2020-02-18 | Apple Inc. | Virtual assistant continuity |
| US10593346B2 (en) | 2016-12-22 | 2020-03-17 | Apple Inc. | Rank-reduced token representation for automatic speech recognition |
| US10671428B2 (en) | 2015-09-08 | 2020-06-02 | Apple Inc. | Distributed personal assistant |
| US10691473B2 (en) | 2015-11-06 | 2020-06-23 | Apple Inc. | Intelligent automated assistant in a messaging environment |
| US10706841B2 (en) | 2010-01-18 | 2020-07-07 | Apple Inc. | Task flow identification based on user intent |
| US10733993B2 (en) | 2016-06-10 | 2020-08-04 | Apple Inc. | Intelligent digital assistant in a multi-tasking environment |
| US10747498B2 (en) | 2015-09-08 | 2020-08-18 | Apple Inc. | Zero latency digital assistant |
| US10755703B2 (en) | 2017-05-11 | 2020-08-25 | Apple Inc. | Offline personal assistant |
| US10791176B2 (en) | 2017-05-12 | 2020-09-29 | Apple Inc. | Synchronization and task delegation of a digital assistant |
| US10795541B2 (en) | 2009-06-05 | 2020-10-06 | Apple Inc. | Intelligent organization of tasks items |
| US10810274B2 (en) | 2017-05-15 | 2020-10-20 | Apple Inc. | Optimizing dialogue policy decisions for digital assistants using implicit feedback |
| US11010550B2 (en) | 2015-09-29 | 2021-05-18 | Apple Inc. | Unified language modeling framework for word prediction, auto-completion and auto-correction |
| US11080012B2 (en) | 2009-06-05 | 2021-08-03 | Apple Inc. | Interface for a virtual digital assistant |
| US11217255B2 (en) | 2017-05-16 | 2022-01-04 | Apple Inc. | Far-field extension for digital assistant services |
Families Citing this family (54)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7149690B2 (en) | 1999-09-09 | 2006-12-12 | Lucent Technologies Inc. | Method and apparatus for interactive language instruction |
| KR100464019B1 (ko) * | 2000-12-29 | 2004-12-30 | 엘지전자 주식회사 | 음성 인식기의 발음 사전 편집시 발음열 디스플레이 방법 |
| KR100352748B1 (ko) * | 2001-01-05 | 2002-09-16 | (주) 코아보이스 | 온라인 학습형 음성합성 장치 및 그 방법 |
| US6513008B2 (en) * | 2001-03-15 | 2003-01-28 | Matsushita Electric Industrial Co., Ltd. | Method and tool for customization of speech synthesizer databases using hierarchical generalized speech templates |
| GB2388286A (en) * | 2002-05-01 | 2003-11-05 | Seiko Epson Corp | Enhanced speech data for use in a text to speech system |
| GB2393369A (en) * | 2002-09-20 | 2004-03-24 | Seiko Epson Corp | A method of implementing a text to speech (TTS) system and a mobile telephone incorporating such a TTS system |
| US7389228B2 (en) | 2002-12-16 | 2008-06-17 | International Business Machines Corporation | Speaker adaptation of vocabulary for speech recognition |
| CA2501888C (en) | 2003-03-14 | 2014-05-27 | Nippon Telegraph And Telephone Corporation | Optical node device, network control device, maintenance-staff device, optical network, and 3r relay implementation node decision method |
| EP2357623A1 (en) | 2003-04-25 | 2011-08-17 | Apple Inc. | Graphical user interface for browsing, searching and presenting media items |
| US9406068B2 (en) | 2003-04-25 | 2016-08-02 | Apple Inc. | Method and system for submitting media for network-based purchase and distribution |
| US7805307B2 (en) | 2003-09-30 | 2010-09-28 | Sharp Laboratories Of America, Inc. | Text to speech conversion system |
| US7844548B2 (en) | 2003-10-15 | 2010-11-30 | Apple Inc. | Techniques and systems for electronic submission of media for network-based distribution |
| US20060277044A1 (en) * | 2005-06-02 | 2006-12-07 | Mckay Martin | Client-based speech enabled web content |
| US20090291419A1 (en) * | 2005-08-01 | 2009-11-26 | Kazuaki Uekawa | System of sound representaion and pronunciation techniques for english and other european languages |
| US8249873B2 (en) | 2005-08-12 | 2012-08-21 | Avaya Inc. | Tonal correction of speech |
| US20070050188A1 (en) * | 2005-08-26 | 2007-03-01 | Avaya Technology Corp. | Tone contour transformation of speech |
| US8015237B2 (en) | 2006-05-15 | 2011-09-06 | Apple Inc. | Processing of metadata content and media content received by a media distribution system |
| US7827162B2 (en) | 2006-05-15 | 2010-11-02 | Apple Inc. | Media package format for submission to a media distribution system |
| US7962634B2 (en) | 2006-05-15 | 2011-06-14 | Apple Inc. | Submission of metadata content and media content to a media distribution system |
| US7873517B2 (en) * | 2006-11-09 | 2011-01-18 | Volkswagen Of America, Inc. | Motor vehicle with a speech interface |
| US8719027B2 (en) * | 2007-02-28 | 2014-05-06 | Microsoft Corporation | Name synthesis |
| US20090259502A1 (en) * | 2008-04-10 | 2009-10-15 | Daniel David Erlewine | Quality-Based Media Management for Network-Based Media Distribution |
| US9342287B2 (en) | 2008-05-05 | 2016-05-17 | Apple Inc. | Software program ratings |
| US9076176B2 (en) * | 2008-05-05 | 2015-07-07 | Apple Inc. | Electronic submission of application programs for network-based distribution |
| US8990087B1 (en) * | 2008-09-30 | 2015-03-24 | Amazon Technologies, Inc. | Providing text to speech from digital content on an electronic device |
| US8655660B2 (en) * | 2008-12-11 | 2014-02-18 | International Business Machines Corporation | Method for dynamic learning of individual voice patterns |
| US20100153116A1 (en) * | 2008-12-12 | 2010-06-17 | Zsolt Szalai | Method for storing and retrieving voice fonts |
| US8160881B2 (en) * | 2008-12-15 | 2012-04-17 | Microsoft Corporation | Human-assisted pronunciation generation |
| US8775184B2 (en) * | 2009-01-16 | 2014-07-08 | International Business Machines Corporation | Evaluating spoken skills |
| US20100235254A1 (en) * | 2009-03-16 | 2010-09-16 | Payam Mirrashidi | Application Products with In-Application Subsequent Feature Access Using Network-Based Distribution System |
| GB2470606B (en) * | 2009-05-29 | 2011-05-04 | Paul Siani | Electronic reading device |
| US9729609B2 (en) | 2009-08-07 | 2017-08-08 | Apple Inc. | Automatic transport discovery for media submission |
| KR101217653B1 (ko) * | 2009-08-14 | 2013-01-02 | 오주성 | 영어 학습 시스템 |
| US8935217B2 (en) | 2009-09-08 | 2015-01-13 | Apple Inc. | Digital asset validation prior to submission for network-based distribution |
| CN102117614B (zh) * | 2010-01-05 | 2013-01-02 | 索尼爱立信移动通讯有限公司 | 个性化文本语音合成和个性化语音特征提取 |
| US20110184736A1 (en) * | 2010-01-26 | 2011-07-28 | Benjamin Slotznick | Automated method of recognizing inputted information items and selecting information items |
| US9640175B2 (en) * | 2011-10-07 | 2017-05-02 | Microsoft Technology Licensing, Llc | Pronunciation learning from user correction |
| RU2510954C2 (ru) * | 2012-05-18 | 2014-04-10 | Александр Юрьевич Бредихин | Способ переозвучивания аудиоматериалов и устройство для его осуществления |
| US8990188B2 (en) | 2012-11-30 | 2015-03-24 | Apple Inc. | Managed assessment of submitted digital content |
| US9087341B2 (en) | 2013-01-11 | 2015-07-21 | Apple Inc. | Migration of feedback data to equivalent digital assets |
| US10319254B2 (en) * | 2013-03-15 | 2019-06-11 | Joel Lane Mayon | Graphical user interfaces for spanish language teaching |
| KR101487005B1 (ko) * | 2013-11-13 | 2015-01-29 | (주)위버스마인드 | 문장입력을 통해 발음교정을 실시하는 외국어 학습장치 및 그 학습방법 |
| US9633004B2 (en) | 2014-05-30 | 2017-04-25 | Apple Inc. | Better resolution when referencing to concepts |
| US9953631B1 (en) | 2015-05-07 | 2018-04-24 | Google Llc | Automatic speech recognition techniques for multiple languages |
| US11025565B2 (en) | 2015-06-07 | 2021-06-01 | Apple Inc. | Personalized prediction of responses for instant messaging |
| US10102203B2 (en) | 2015-12-21 | 2018-10-16 | Verisign, Inc. | Method for writing a foreign language in a pseudo language phonetically resembling native language of the speaker |
| US9910836B2 (en) | 2015-12-21 | 2018-03-06 | Verisign, Inc. | Construction of phonetic representation of a string of characters |
| US9947311B2 (en) * | 2015-12-21 | 2018-04-17 | Verisign, Inc. | Systems and methods for automatic phonetization of domain names |
| US10102189B2 (en) | 2015-12-21 | 2018-10-16 | Verisign, Inc. | Construction of a phonetic representation of a generated string of characters |
| US11281993B2 (en) | 2016-12-05 | 2022-03-22 | Apple Inc. | Model and ensemble compression for metric learning |
| DK201770383A1 (en) | 2017-05-09 | 2018-12-14 | Apple Inc. | USER INTERFACE FOR CORRECTING RECOGNITION ERRORS |
| DK201770428A1 (en) | 2017-05-12 | 2019-02-18 | Apple Inc. | LOW-LATENCY INTELLIGENT AUTOMATED ASSISTANT |
| CN111968619A (zh) * | 2020-08-26 | 2020-11-20 | 四川长虹电器股份有限公司 | 控制语音合成发音的方法及装置 |
| US12614541B2 (en) * | 2022-11-08 | 2026-04-28 | Jpmorgan Chase Bank, N.A. | Systems and methods for machine-learning based multi-lingual pronunciation generation |
Family Cites Families (30)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| EP0372734B1 (en) | 1988-11-23 | 1994-03-09 | Digital Equipment Corporation | Name pronunciation by synthesizer |
| US5027406A (en) * | 1988-12-06 | 1991-06-25 | Dragon Systems, Inc. | Method for interactive speech recognition and training |
| JPH05204389A (ja) | 1992-01-23 | 1993-08-13 | Matsushita Electric Ind Co Ltd | 音声規則合成用ユーザー辞書登録システム |
| US5393236A (en) * | 1992-09-25 | 1995-02-28 | Northeastern University | Interactive speech pronunciation apparatus and method |
| US5920836A (en) * | 1992-11-13 | 1999-07-06 | Dragon Systems, Inc. | Word recognition system using language context at current cursor position to affect recognition probabilities |
| JPH06176023A (ja) * | 1992-12-08 | 1994-06-24 | Toshiba Corp | 音声合成システム |
| US5799267A (en) * | 1994-07-22 | 1998-08-25 | Siegel; Steven H. | Phonic engine |
| JPH0895587A (ja) | 1994-09-27 | 1996-04-12 | Oki Electric Ind Co Ltd | テキスト音声合成方法 |
| WO1996010795A1 (en) * | 1994-10-03 | 1996-04-11 | Helfgott & Karas, P.C. | A database accessing system |
| US5697789A (en) * | 1994-11-22 | 1997-12-16 | Softrade International, Inc. | Method and system for aiding foreign language instruction |
| US5787231A (en) * | 1995-02-02 | 1998-07-28 | International Business Machines Corporation | Method and system for improving pronunciation in a voice control system |
| JPH08320864A (ja) | 1995-05-26 | 1996-12-03 | Fujitsu Ltd | 音声合成用の辞書登録装置 |
| US5999895A (en) * | 1995-07-24 | 1999-12-07 | Forest; Donald K. | Sound operated menu method and apparatus |
| JP3483230B2 (ja) * | 1995-10-20 | 2004-01-06 | 株式会社リコー | 発声情報作成装置 |
| US5799276A (en) * | 1995-11-07 | 1998-08-25 | Accent Incorporated | Knowledge-based speech recognition system and methods having frame length computed based upon estimated pitch period of vocalic intervals |
| JPH09325787A (ja) * | 1996-05-30 | 1997-12-16 | Internatl Business Mach Corp <Ibm> | 音声合成方法、音声合成装置、文章への音声コマンド組み込み方法、及び装置 |
| US5845238A (en) * | 1996-06-18 | 1998-12-01 | Apple Computer, Inc. | System and method for using a correspondence table to compress a pronunciation guide |
| JP3660432B2 (ja) * | 1996-07-15 | 2005-06-15 | 株式会社東芝 | 辞書登録装置及び辞書登録方法 |
| US5850629A (en) * | 1996-09-09 | 1998-12-15 | Matsushita Electric Industrial Co., Ltd. | User interface controller for text-to-speech synthesizer |
| JPH10153998A (ja) * | 1996-09-24 | 1998-06-09 | Nippon Telegr & Teleph Corp <Ntt> | 補助情報利用型音声合成方法、この方法を実施する手順を記録した記録媒体、およびこの方法を実施する装置 |
| US5950160A (en) | 1996-10-31 | 1999-09-07 | Microsoft Corporation | Method and system for displaying a variable number of alternative words during speech recognition |
| JP3573907B2 (ja) * | 1997-03-10 | 2004-10-06 | 株式会社リコー | 音声合成装置 |
| US5933804A (en) * | 1997-04-10 | 1999-08-03 | Microsoft Corporation | Extensible speech recognition system that provides a user with audio feedback |
| US6226614B1 (en) * | 1997-05-21 | 2001-05-01 | Nippon Telegraph And Telephone Corporation | Method and apparatus for editing/creating synthetic speech message and recording medium with the method recorded thereon |
| JPH10336354A (ja) | 1997-06-04 | 1998-12-18 | Meidensha Corp | マルチメディア公衆電話システム |
| US6016471A (en) * | 1998-04-29 | 2000-01-18 | Matsushita Electric Industrial Co., Ltd. | Method and apparatus using decision trees to generate and score multiple pronunciations for a spelled word |
| US6185535B1 (en) | 1998-10-16 | 2001-02-06 | Telefonaktiebolaget Lm Ericsson (Publ) | Voice control of a user interface to service applications |
| US6275789B1 (en) * | 1998-12-18 | 2001-08-14 | Leo Moser | Method and apparatus for performing full bidirectional translation between a source language and a linked alternative language |
| US6389394B1 (en) * | 2000-02-09 | 2002-05-14 | Speechworks International, Inc. | Method and apparatus for improved speech recognition by modifying a pronunciation dictionary based on pattern definitions of alternate word pronunciations |
| US6865533B2 (en) * | 2000-04-21 | 2005-03-08 | Lessac Technology Inc. | Text to speech |
-
1999
- 1999-04-30 US US09/303,057 patent/US7292980B1/en not_active Expired - Fee Related
-
2000
- 2000-04-20 EP EP00303371A patent/EP1049072B1/en not_active Expired - Lifetime
- 2000-04-20 DE DE60020773T patent/DE60020773T2/de not_active Expired - Lifetime
- 2000-04-25 CA CA002306527A patent/CA2306527A1/en not_active Abandoned
- 2000-04-28 JP JP2000130595A patent/JP4237915B2/ja not_active Expired - Fee Related
- 2000-05-01 KR KR10-2000-0023242A patent/KR100378898B1/ko not_active Expired - Fee Related
Cited By (63)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10318871B2 (en) | 2005-09-08 | 2019-06-11 | Apple Inc. | Method and apparatus for building an intelligent automated assistant |
| JP2008180750A (ja) * | 2007-01-23 | 2008-08-07 | Oki Electric Ind Co Ltd | 音声ラベリング支援システム |
| US9865248B2 (en) | 2008-04-05 | 2018-01-09 | Apple Inc. | Intelligent text-to-speech conversion |
| US11080012B2 (en) | 2009-06-05 | 2021-08-03 | Apple Inc. | Interface for a virtual digital assistant |
| US10795541B2 (en) | 2009-06-05 | 2020-10-06 | Apple Inc. | Intelligent organization of tasks items |
| US10706841B2 (en) | 2010-01-18 | 2020-07-07 | Apple Inc. | Task flow identification based on user intent |
| US11423886B2 (en) | 2010-01-18 | 2022-08-23 | Apple Inc. | Task flow identification based on user intent |
| CN102687197B (zh) * | 2010-01-22 | 2014-07-23 | 三菱电机株式会社 | 声音识别用词典制作装置、声音识别装置及声音合成装置 |
| US9177545B2 (en) | 2010-01-22 | 2015-11-03 | Mitsubishi Electric Corporation | Recognition dictionary creating device, voice recognition device, and voice synthesizer |
| CN102687197A (zh) * | 2010-01-22 | 2012-09-19 | 三菱电机株式会社 | 识别词典制作装置、声音识别装置及声音合成装置 |
| JP4942860B2 (ja) * | 2010-01-22 | 2012-05-30 | 三菱電機株式会社 | 認識辞書作成装置、音声認識装置及び音声合成装置 |
| WO2011089651A1 (ja) * | 2010-01-22 | 2011-07-28 | 三菱電機株式会社 | 認識辞書作成装置、音声認識装置及び音声合成装置 |
| US10049675B2 (en) | 2010-02-25 | 2018-08-14 | Apple Inc. | User profiling for voice input processing |
| US10134385B2 (en) | 2012-03-02 | 2018-11-20 | Apple Inc. | Systems and methods for name pronunciation |
| JP2015512062A (ja) * | 2012-03-02 | 2015-04-23 | アップル インコーポレイテッド | 名前発音システム及び方法 |
| US11069336B2 (en) | 2012-03-02 | 2021-07-20 | Apple Inc. | Systems and methods for name pronunciation |
| US10079014B2 (en) | 2012-06-08 | 2018-09-18 | Apple Inc. | Name recognition system |
| US9971774B2 (en) | 2012-09-19 | 2018-05-15 | Apple Inc. | Voice-based media searching |
| US9966060B2 (en) | 2013-06-07 | 2018-05-08 | Apple Inc. | System and method for user-specified pronunciation of words for speech synthesis and recognition |
| JP2015118222A (ja) * | 2013-12-18 | 2015-06-25 | 株式会社日立超エル・エス・アイ・システムズ | 音声合成システム及び音声合成方法 |
| US10169329B2 (en) | 2014-05-30 | 2019-01-01 | Apple Inc. | Exemplar-based natural language processing |
| US10904611B2 (en) | 2014-06-30 | 2021-01-26 | Apple Inc. | Intelligent automated assistant for TV user interactions |
| US9668024B2 (en) | 2014-06-30 | 2017-05-30 | Apple Inc. | Intelligent automated assistant for TV user interactions |
| US9986419B2 (en) | 2014-09-30 | 2018-05-29 | Apple Inc. | Social reminders |
| US10567477B2 (en) | 2015-03-08 | 2020-02-18 | Apple Inc. | Virtual assistant continuity |
| US10356243B2 (en) | 2015-06-05 | 2019-07-16 | Apple Inc. | Virtual assistant aided communication with 3rd party service in a communication session |
| US10747498B2 (en) | 2015-09-08 | 2020-08-18 | Apple Inc. | Zero latency digital assistant |
| US10671428B2 (en) | 2015-09-08 | 2020-06-02 | Apple Inc. | Distributed personal assistant |
| US11500672B2 (en) | 2015-09-08 | 2022-11-15 | Apple Inc. | Distributed personal assistant |
| US10366158B2 (en) | 2015-09-29 | 2019-07-30 | Apple Inc. | Efficient word encoding for recurrent neural network language models |
| US11010550B2 (en) | 2015-09-29 | 2021-05-18 | Apple Inc. | Unified language modeling framework for word prediction, auto-completion and auto-correction |
| US10691473B2 (en) | 2015-11-06 | 2020-06-23 | Apple Inc. | Intelligent automated assistant in a messaging environment |
| US11526368B2 (en) | 2015-11-06 | 2022-12-13 | Apple Inc. | Intelligent automated assistant in a messaging environment |
| US10049668B2 (en) | 2015-12-02 | 2018-08-14 | Apple Inc. | Applying neural network language models to weighted finite state transducers for automatic speech recognition |
| US10223066B2 (en) | 2015-12-23 | 2019-03-05 | Apple Inc. | Proactive assistance based on dialog communication between devices |
| US10446143B2 (en) | 2016-03-14 | 2019-10-15 | Apple Inc. | Identification of voice inputs providing credentials |
| US9934775B2 (en) | 2016-05-26 | 2018-04-03 | Apple Inc. | Unit-selection text-to-speech synthesis based on predicted concatenation parameters |
| US9972304B2 (en) | 2016-06-03 | 2018-05-15 | Apple Inc. | Privacy preserving distributed evaluation framework for embedded personalized systems |
| US10249300B2 (en) | 2016-06-06 | 2019-04-02 | Apple Inc. | Intelligent list reading |
| US10049663B2 (en) | 2016-06-08 | 2018-08-14 | Apple, Inc. | Intelligent automated assistant for media exploration |
| US11069347B2 (en) | 2016-06-08 | 2021-07-20 | Apple Inc. | Intelligent automated assistant for media exploration |
| US10354011B2 (en) | 2016-06-09 | 2019-07-16 | Apple Inc. | Intelligent automated assistant in a home environment |
| US11037565B2 (en) | 2016-06-10 | 2021-06-15 | Apple Inc. | Intelligent digital assistant in a multi-tasking environment |
| US10509862B2 (en) | 2016-06-10 | 2019-12-17 | Apple Inc. | Dynamic phrase expansion of language input |
| US10067938B2 (en) | 2016-06-10 | 2018-09-04 | Apple Inc. | Multilingual word prediction |
| US10733993B2 (en) | 2016-06-10 | 2020-08-04 | Apple Inc. | Intelligent digital assistant in a multi-tasking environment |
| US10192552B2 (en) | 2016-06-10 | 2019-01-29 | Apple Inc. | Digital assistant providing whispered speech |
| US10490187B2 (en) | 2016-06-10 | 2019-11-26 | Apple Inc. | Digital assistant providing automated status report |
| US10089072B2 (en) | 2016-06-11 | 2018-10-02 | Apple Inc. | Intelligent device arbitration and control |
| US10269345B2 (en) | 2016-06-11 | 2019-04-23 | Apple Inc. | Intelligent task discovery |
| US10521466B2 (en) | 2016-06-11 | 2019-12-31 | Apple Inc. | Data driven natural language event detection and classification |
| US10297253B2 (en) | 2016-06-11 | 2019-05-21 | Apple Inc. | Application integration with a digital assistant |
| US11152002B2 (en) | 2016-06-11 | 2021-10-19 | Apple Inc. | Application integration with a digital assistant |
| US10553215B2 (en) | 2016-09-23 | 2020-02-04 | Apple Inc. | Intelligent automated assistant |
| US10043516B2 (en) | 2016-09-23 | 2018-08-07 | Apple Inc. | Intelligent automated assistant |
| US10593346B2 (en) | 2016-12-22 | 2020-03-17 | Apple Inc. | Rank-reduced token representation for automatic speech recognition |
| US10755703B2 (en) | 2017-05-11 | 2020-08-25 | Apple Inc. | Offline personal assistant |
| US10791176B2 (en) | 2017-05-12 | 2020-09-29 | Apple Inc. | Synchronization and task delegation of a digital assistant |
| US11405466B2 (en) | 2017-05-12 | 2022-08-02 | Apple Inc. | Synchronization and task delegation of a digital assistant |
| US10410637B2 (en) | 2017-05-12 | 2019-09-10 | Apple Inc. | User-specific acoustic models |
| US10482874B2 (en) | 2017-05-15 | 2019-11-19 | Apple Inc. | Hierarchical belief states for digital assistants |
| US10810274B2 (en) | 2017-05-15 | 2020-10-20 | Apple Inc. | Optimizing dialogue policy decisions for digital assistants using implicit feedback |
| US11217255B2 (en) | 2017-05-16 | 2022-01-04 | Apple Inc. | Far-field extension for digital assistant services |
Also Published As
| Publication number | Publication date |
|---|---|
| CA2306527A1 (en) | 2000-10-30 |
| KR20000077120A (ko) | 2000-12-26 |
| KR100378898B1 (ko) | 2003-04-07 |
| EP1049072A2 (en) | 2000-11-02 |
| US7292980B1 (en) | 2007-11-06 |
| JP4237915B2 (ja) | 2009-03-11 |
| EP1049072B1 (en) | 2005-06-15 |
| DE60020773T2 (de) | 2006-05-11 |
| EP1049072A3 (en) | 2003-10-15 |
| DE60020773D1 (de) | 2005-07-21 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP4237915B2 (ja) | ユーザが文字列の発音を設定することを可能にするためにコンピュータ上で実行される方法 | |
| US7881928B2 (en) | Enhanced linguistic transformation | |
| US6363342B2 (en) | System for developing word-pronunciation pairs | |
| CN101223571B (zh) | 音质变化部位确定装置及音质变化部位确定方法 | |
| US7996226B2 (en) | System and method of developing a TTS voice | |
| US7711562B1 (en) | System and method for testing a TTS voice | |
| JP4833313B2 (ja) | 中国語方言判断プログラム | |
| JP2001188777A (ja) | 音声をテキストに関連付ける方法、音声をテキストに関連付けるコンピュータ、コンピュータで文書を生成し読み上げる方法、文書を生成し読み上げるコンピュータ、コンピュータでテキスト文書の音声再生を行う方法、テキスト文書の音声再生を行うコンピュータ、及び、文書内のテキストを編集し評価する方法 | |
| US20070239455A1 (en) | Method and system for managing pronunciation dictionaries in a speech application | |
| JPS6259996A (ja) | 辞書操作装置 | |
| JP2002318595A (ja) | テキスト音声合成システムの韻律テンプレートマッチング | |
| JP5172682B2 (ja) | 音素のnグラムを使用した単語および名前の生成 | |
| US6456973B1 (en) | Task automation user interface with text-to-speech output | |
| US7099828B2 (en) | Method and apparatus for word pronunciation composition | |
| JPH11344990A (ja) | 綴り言葉に対する複数発音を生成し評価する判断ツリ―を利用する方法及び装置 | |
| WO2010136821A1 (en) | Electronic reading device | |
| US20090281808A1 (en) | Voice data creation system, program, semiconductor integrated circuit device, and method for producing semiconductor integrated circuit device | |
| US7742919B1 (en) | System and method for repairing a TTS voice database | |
| JP2010169973A (ja) | 外国語学習支援システム、及びプログラム | |
| JP2005031150A (ja) | 音声処理装置および方法 | |
| JP2006139162A (ja) | 語学学習装置 | |
| WO2022196087A1 (ja) | 情報処理装置、情報処理方法、および情報処理プログラム | |
| JP4026512B2 (ja) | 歌唱合成用データ入力プログラムおよび歌唱合成用データ入力装置 | |
| JP2005241767A (ja) | 音声認識装置 | |
| Hill et al. | Unrestricted text-to-speech revisited: rhythm and intonation. |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| A977 | Report on retrieval |
Free format text: JAPANESE INTERMEDIATE CODE: A971007 Effective date: 20050906 |
|
| A131 | Notification of reasons for refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A131 Effective date: 20051003 |
|
| A601 | Written request for extension of time |
Free format text: JAPANESE INTERMEDIATE CODE: A601 Effective date: 20051227 |
|
| A602 | Written permission of extension of time |
Free format text: JAPANESE INTERMEDIATE CODE: A602 Effective date: 20060105 |
|
| A521 | Request for written amendment filed |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20060113 |
|
| A02 | Decision of refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A02 Effective date: 20061127 |
|
| A601 | Written request for extension of time |
Free format text: JAPANESE INTERMEDIATE CODE: A601 Effective date: 20080826 |
|
| A602 | Written permission of extension of time |
Free format text: JAPANESE INTERMEDIATE CODE: A602 Effective date: 20080829 |
|
| A521 | Request for written amendment filed |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20081028 |
|
| A01 | Written decision to grant a patent or to grant a registration (utility model) |
Free format text: JAPANESE INTERMEDIATE CODE: A01 |
|
| A61 | First payment of annual fees (during grant procedure) |
Free format text: JAPANESE INTERMEDIATE CODE: A61 Effective date: 20081219 |
|
| R150 | Certificate of patent or registration of utility model |
Free format text: JAPANESE INTERMEDIATE CODE: R150 Ref document number: 4237915 Country of ref document: JP Free format text: JAPANESE INTERMEDIATE CODE: R150 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20111226 Year of fee payment: 3 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20121226 Year of fee payment: 4 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| FPAY | Renewal fee payment (event date is renewal date of database) |
Free format text: PAYMENT UNTIL: 20131226 Year of fee payment: 5 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| R250 | Receipt of annual fees |
Free format text: JAPANESE INTERMEDIATE CODE: R250 |
|
| LAPS | Cancellation because of no payment of annual fees |