JP2005106905A - Audio output system and server device - Google Patents

Audio output system and server device Download PDF

Info

Publication number
JP2005106905A
JP2005106905A JP2003336950A JP2003336950A JP2005106905A JP 2005106905 A JP2005106905 A JP 2005106905A JP 2003336950 A JP2003336950 A JP 2003336950A JP 2003336950 A JP2003336950 A JP 2003336950A JP 2005106905 A JP2005106905 A JP 2005106905A
Authority
JP
Japan
Prior art keywords
data
pronunciation
server device
voice
unit
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP2003336950A
Other languages
Japanese (ja)
Inventor
Terushi Kimata
輝志 木全
Kazumasa Takenaka
和正 竹中
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Panasonic Holdings Corp
Original Assignee
Matsushita Electric Industrial Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Matsushita Electric Industrial Co Ltd filed Critical Matsushita Electric Industrial Co Ltd
Priority to JP2003336950A priority Critical patent/JP2005106905A/en
Publication of JP2005106905A publication Critical patent/JP2005106905A/en
Pending legal-status Critical Current

Links

Images

Abstract

【課題】
端末装置の処理の負担を軽減するとともに、通信データ量を減少可能な音声出力システムおよびそのシステムに用いられるサーバ装置を提供すること。
【解決手段】
変換サーバ2は、メールサーバ1やインターネット等の通信網4を介して電子メールやウェブサイト等の情報を受信し、受信した情報に含まれるテキストデータを、記号化および/または数値化された発音データに変換し、基地局等を含む通信網5を介して接続される携帯電話装置3へ情報とともに発音データを送信する。携帯電話装置3は受信した発音データを音声信号に変換して音声出力し、受信した電子メールやウェブサイト等の内容を音声出力する。
【選択図】 図1
【Task】
To provide a voice output system capable of reducing the processing load of a terminal device and reducing the amount of communication data, and a server device used in the system.
[Solution]
The conversion server 2 receives information such as e-mails and websites via the mail server 1 and the communication network 4 such as the Internet, and the text data included in the received information is encoded and / or digitized. The data is converted into data, and the sound generation data is transmitted together with the information to the mobile phone device 3 connected via the communication network 5 including the base station and the like. The cellular phone device 3 converts the received pronunciation data into an audio signal and outputs it as an audio signal, and outputs the content of the received e-mail, website, etc. as an audio message.
[Selection] Figure 1

Description

本発明は、端末装置において文字や文章等を音声で出力音声出力システムおよびサーバ装置に関する。   The present invention relates to a voice output system and a server device that output characters, sentences, and the like by voice in a terminal device.

従来、電子メール等の文章を音声に変換して端末装置等のスピーカから出力する音声出力システムとしては次のものがある。   2. Description of the Related Art Conventionally, there are the following voice output systems that convert sentences such as e-mails into voice and output them from a speaker such as a terminal device.

例えば、特許文献1には、メールサーバから受信した電子メールの文字を解読し、その文字に対応する音声を合成して、その合成された音声を再生して出力する携帯電話機が記載されている。   For example, Patent Document 1 describes a mobile phone that decodes characters of an e-mail received from a mail server, synthesizes speech corresponding to the characters, and reproduces and outputs the synthesized speech. .

次に、例えば特許文献2には、受信した電子メールに対して、内蔵する音声合成装置により音声合成処理を行い、そこに記述されたメッセージに対応する合成音を生成して電話機に送信するサービスプロバイダサーバが記載されている。   Next, for example, Patent Document 2 discloses a service in which a received voice mail is subjected to voice synthesis processing by a built-in voice synthesizer, and a synthesized voice corresponding to a message described therein is generated and transmitted to a telephone. The provider server is described.

また、例えば特許文献3には、画面に表示されているテキストにフォーカスを当て、当てられたフォーカスを当てられたテキストについてのデータを取得し、取得したテキストを音声合成サーバに出力するとともに、音声合成サーバが作成した音声波形データに基づいてスピーカから音声出力する携帯端末を有する携帯端末通信システムが記載されている。   Further, for example, in Patent Document 3, the text displayed on the screen is focused, the data about the focused text is acquired, the acquired text is output to the speech synthesis server, and the voice is also recorded. A mobile terminal communication system having a mobile terminal that outputs voice from a speaker based on voice waveform data created by a synthesis server is described.

しかしながら、例えば特許文献1に記載された携帯電話機のように、端末装置側で文字を音声に変換して出力する音声出力システムにあっては、その変換処理が端末装置へ負担をかけ、特に携帯電話機等の携帯無線装置の処理能力に対しては、その負担が大くなってしまうという事情があった。   However, for example, in a voice output system that converts characters to voice on the terminal device side and outputs them, such as the mobile phone described in Patent Document 1, the conversion processing places a burden on the terminal device, and is particularly portable. There has been a situation in which the processing capacity of portable wireless devices such as telephones is increased.

また、例えば特許文献2や特許文献3に記載されたシステムのように、サーバ装置側で音声データ変換処理を行って端末装置へ音声データを送信する音声出力システムにあっては、音声データの容量が大きくなってしまうという事情があった。
特開2003−150507号公報 特開平09−258764号公報 特開2002−158803号公報
In addition, in a voice output system that performs voice data conversion processing on the server device side and transmits the voice data to the terminal device, such as the systems described in Patent Document 2 and Patent Document 3, the capacity of the voice data There was a situation that became big.
JP 2003-150507 A JP 09-258664 A JP 2002-158803 A

本発明は、上記従来の事情に鑑みてなされたものであって、端末装置の処理の負担を軽減するとともに、通信データ量を減少可能な音声出力システムおよびそのシステムに用いられるサーバ装置を提供することを目的とする。   The present invention has been made in view of the above-described conventional circumstances, and provides an audio output system capable of reducing the processing load of a terminal device and reducing the amount of communication data, and a server device used in the system. For the purpose.

本発明のサーバ装置は、テキストデータを端末装置において音声出力する音声出力システムに用いられるサーバ装置であって、
前記テキストデータを取得する取得手段と、
前記テキストデータを解析して、記号化および/または数値化された発音データを作成する発音データ作成部と、
前記発音データを前記端末装置へ送信する送信部と、
を備える。
The server device of the present invention is a server device used in an audio output system for outputting text data as audio in a terminal device,
Obtaining means for obtaining the text data;
A pronunciation data creation unit that analyzes the text data to create symbolized and / or digitized pronunciation data;
A transmission unit for transmitting the pronunciation data to the terminal device;
Is provided.

この構成により、端末装置の処理の負担を軽減するとともに、サーバ装置と端末装置間の通信データ量を減少させることができる。   With this configuration, it is possible to reduce the processing load on the terminal device and reduce the amount of communication data between the server device and the terminal device.

また、本発明のサーバ装置において、前記発音データ作成部は、文字、音節、および文節のうち少なくとも1種類と、それに対応付けられた発音データを含む発音辞書データを記憶した記憶部を有し、前記発音辞書データを参照して、前記テキストデータから文字、音節、および文節のうち少なくとも1種類を認識するとともに、前記認識された文字、音節、および文節のうち少なくとも1種類に対応する発音を特定して前記発音データを作成するものである。   Further, in the server device of the present invention, the pronunciation data creation unit includes a storage unit that stores pronunciation dictionary data including at least one type of characters, syllables, and phrases, and pronunciation data associated therewith, Referring to the pronunciation dictionary data, recognizes at least one of characters, syllables, and phrases from the text data, and specifies pronunciation corresponding to at least one of the recognized characters, syllables, and phrases Thus, the pronunciation data is created.

この構成により、より正確な発音を再現することが可能となる。   With this configuration, more accurate pronunciation can be reproduced.

また、本発明のサーバ装置において、前記発音辞書データは、漢字、ひらがな、カタカナを含む文字、音節、および文節のうち少なくとも1種類と、それに対応付けられた発音データを含むものである。   In the server device of the present invention, the pronunciation dictionary data includes at least one type of characters, syllables, and phrases including kanji, hiragana, and katakana, and pronunciation data associated with the type.

この構成により、漢字かな混じり文のような複雑な文章の音声出力にも対応することができる。   With this configuration, it is possible to cope with voice output of complicated sentences such as kanji-kana mixed sentences.

また、本発明のサーバ装置において、前記発音データは、発音、強弱、および音程を表す記号および/または数値を含むものである。   In the server device of the present invention, the pronunciation data includes symbols and / or numerical values representing pronunciation, strength, and pitch.

この構成により、より正確な発音を再現することが可能となり、端末装置使用者の音声出力する内容への理解度を向上させることができる。   With this configuration, it is possible to reproduce more accurate pronunciation, and the degree of understanding of the content output by the terminal device user can be improved.

また、本発明のサーバ装置において、前記取得部は、配信された電子メールを受信するとともに、前記電子メールの前記テキストデータを取得し、
前記送信部は前記発音データとともに前記電子メールを前記端末装置へ送信するものである。
Further, in the server device of the present invention, the acquisition unit receives the distributed electronic mail, acquires the text data of the electronic mail,
The transmission unit transmits the electronic mail together with the pronunciation data to the terminal device.

この構成により、端末装置の使用者は、受信した電子メールの文字データとともに、音声で内容を確認することができる。   With this configuration, the user of the terminal device can confirm the contents by voice together with the received e-mail character data.

また、本発明のサーバ装置において、前記取得部は、前記端末装置によって指定されたウェブサイトの情報を受信して前記ウェブサイトの内容を取得するとともに、前記ウェブサイトの内容から前記テキストデータを取得し、
前記送信部は前記発音データとともに前記ウェブサイトの内容を前記端末装置へ送信するものである。
In the server device of the present invention, the acquisition unit receives the information on the website specified by the terminal device to acquire the content of the website, and acquires the text data from the content of the website. And
The transmitting unit transmits the content of the website together with the pronunciation data to the terminal device.

この構成により、端末装置の使用者は、所望のウェブサイトの情報を音声で内容を確認することができる。   With this configuration, the user of the terminal device can confirm the content of desired website information by voice.

本発明の音声出力システムは、
前記サーバ装置と、
前記サーバ装置から通信網を介して接続される通信端末装置と、
を備え、
前記通信端末装置は、
前記サーバ装置の前記通信部から送信された発音データを受信する端末通信部と、
前記端末通信部が受信した発音データから音声信号に変換する音声変換部と、
前記音声変換部で変換した前記音声信号を出力する音声出力部と、
を有する。
The audio output system of the present invention is
The server device;
A communication terminal device connected via a communication network from the server device;
With
The communication terminal device
A terminal communication unit that receives the pronunciation data transmitted from the communication unit of the server device;
A voice conversion unit that converts the sound data received by the terminal communication unit into a voice signal;
An audio output unit for outputting the audio signal converted by the audio conversion unit;
Have

この構成により、端末装置の処理の負担を軽減するとともに、サーバ装置と端末装置間の通信データ量を減少させることができる。   With this configuration, it is possible to reduce the processing load on the terminal device and reduce the amount of communication data between the server device and the terminal device.

また、本発明の音声出力システムにおいて、前記端末装置は、前記サーバ装置の前記送信部から前記テキストデータを受信した場合に、前記テキストデータを表示する表示部を更に備える。   Moreover, the audio | voice output system of this invention WHEREIN: The said terminal device is further provided with the display part which displays the said text data, when the said text data is received from the said transmission part of the said server apparatus.

この構成により、端末装置の使用者は、受信した文字データとともに、音声で内容を確認することができる。   With this configuration, the user of the terminal device can confirm the contents by voice together with the received character data.

また、本発明の音声出力システムにおいて、前記端末装置は、前記発音データに対応した音声情報と、前記発音データと前記音声情報とを対応付けて変換を行う工程を実行可能なソフトウエアとを記憶する端末記憶部を更に備え、
前記音声変換部は前記ソフトウエアに従って動作し、前記発音データに基づいて、前記端末記憶部に記憶された前記音声情報を読出して前記音声信号を生成するものである。
In the voice output system of the present invention, the terminal device stores voice information corresponding to the pronunciation data, and software capable of executing a step of converting the voice generation data and the voice information in association with each other. Further comprising a terminal storage unit,
The voice conversion unit operates according to the software, and reads the voice information stored in the terminal storage unit based on the sound generation data to generate the voice signal.

この構成により、端末装置の処理の負担を軽減するとともに、サーバ装置と端末装置間の通信データ量を減少させることができる。   With this configuration, it is possible to reduce the processing load on the terminal device and reduce the amount of communication data between the server device and the terminal device.

また、本発明の音声出力システムにおいて、前記端末記憶部に記憶された前記ソフトウエアは、前記端末通信部を介して受信する更新データに基づいて更新可能であるものである。   In the audio output system of the present invention, the software stored in the terminal storage unit can be updated based on update data received via the terminal communication unit.

この構成により、音声出力に不具合が生じた場合も、ソフトウエアを更新することで対処することができる。   With this configuration, even if a problem occurs in the audio output, it can be dealt with by updating the software.

また、本発明の音声出力システムにおいて、前記端末装置は携帯無線装置である。   In the audio output system of the present invention, the terminal device is a portable radio device.

この構成により、携帯無線装置の処理の負担を軽減するとともに、音声出力に必要とする記憶容量を減少することにより、装置の大型化を防止して携帯性を損なうことを防ぐことができる。また、音声出力のための通信速度の増大を防ぐことができる。   With this configuration, it is possible to reduce the processing load of the portable wireless device and reduce the storage capacity required for audio output, thereby preventing the device from becoming large and impairing portability. In addition, an increase in communication speed for voice output can be prevented.

本発明によれば、端末装置の処理の負担を軽減するとともに、通信データ量を減少可能な音声出力システムおよびそのシステムに用いられるサーバ装置を提供することができる。   ADVANTAGE OF THE INVENTION According to this invention, while reducing the burden of the process of a terminal device, the audio | voice output system which can reduce communication data amount, and the server apparatus used for the system can be provided.

図1は、本発明の実施形態を説明するための音声出力システムを示す概略構成図である。本実施形態では、音声出力を行う端末装置の一例として携帯電話装置を用いた場合について説明する。図1に示すように、本実施形態の音声出力システムは、メールサーバ1と、変換サーバ2と、携帯電話装置3とを備える。   FIG. 1 is a schematic configuration diagram showing an audio output system for explaining an embodiment of the present invention. In the present embodiment, a case where a mobile phone device is used as an example of a terminal device that performs audio output will be described. As shown in FIG. 1, the audio output system of this embodiment includes a mail server 1, a conversion server 2, and a mobile phone device 3.

メールサーバ1は、インターネット等の通信網4を介して電子メールを受信し、通信回線を介して接続された変換サーバ2へ電子メールを転送する。変換サーバ2は、電子メールを受信すると、電子メールの文字や文章等が記号化および/または数値化されたテキスト形式の発音データに変換し、基地局等を含む通信網5を介して接続される携帯電話装置3へ電子メールとともに発音データを送信する。携帯電話装置3は受信した発音データを音声信号に変換して音声出力し、電子メールの内容を音声出力する。   The mail server 1 receives an electronic mail via a communication network 4 such as the Internet, and transfers the electronic mail to the conversion server 2 connected via a communication line. When the conversion server 2 receives the e-mail, the conversion server 2 converts the character or sentence of the e-mail into pronunciation data in text format that is symbolized and / or digitized, and is connected via a communication network 5 including a base station and the like. The phonetic data is transmitted to the mobile phone device 3 together with the electronic mail. The mobile phone device 3 converts the received pronunciation data into a voice signal and outputs the voice signal, and outputs the contents of the email as voice.

また、携帯電話装置3はウェブサイトを指定して、変換サーバ2へ発音データ要求を行う。変換サーバ2は、指定されたウェブサイトにアクセスし、文字データを取得して発音データに変換し、その発音データを携帯電話装置3へ送信する。携帯電話装置3は受信した発音データを音声信号に変換して音声出力し、指定したウェブサイトの内容を音声出力する。   Further, the mobile phone device 3 designates a website and makes a pronunciation data request to the conversion server 2. The conversion server 2 accesses the designated website, acquires the character data, converts it into pronunciation data, and transmits the pronunciation data to the mobile phone device 3. The cellular phone device 3 converts the received pronunciation data into a voice signal and outputs the voice signal, and outputs the contents of the designated website as a voice.

図2は、本発明の実施形態を説明するための変換サーバの概略構成を示すブロック図である。図2に示すように、本実施形態の変換サーバ2は、第一送受信部21と、サーバ記憶部22と、発音データ変換部23、第二送受信部24と、サーバ制御部25とを備える。   FIG. 2 is a block diagram showing a schematic configuration of a conversion server for explaining the embodiment of the present invention. As shown in FIG. 2, the conversion server 2 of the present embodiment includes a first transmission / reception unit 21, a server storage unit 22, a pronunciation data conversion unit 23, a second transmission / reception unit 24, and a server control unit 25.

第一送受信部21は、メールサーバ1、通信網4に接続されており、メールサーバ1からの電子メールの受信、通信網4を介してのウェブサイトの取得等を行う。記憶部22には、文字、単語、文節、文章等に対応付けられた、発音、音節、アクセント、高低等を表す記号および/または数値を収録した辞書ファイル220が記憶されている。この辞書ファイル220は、内容を書き換えることで更新可能である。この辞書ファイル220は、その収録する容量が大きいほど、正確な発音データ変換を行うことが可能となる反面、携帯電話装置3等の携帯無線装置に記憶させるのが難しくなる。本実施形態では、この辞書ファイル220を変換サーバ2に設けることで、携帯電話装置3に要する音声出力用の記憶容量を減少させることができる。   The first transmission / reception unit 21 is connected to the mail server 1 and the communication network 4, and receives an e-mail from the mail server 1, acquires a website via the communication network 4, and the like. The storage unit 22 stores a dictionary file 220 that records symbols and / or numerical values representing pronunciation, syllables, accents, heights, and the like associated with characters, words, phrases, sentences, and the like. The dictionary file 220 can be updated by rewriting the contents. The dictionary file 220 can be more accurately converted into pronunciation data as its recording capacity increases, but it is difficult to store the dictionary file 220 in a portable wireless device such as the cellular phone device 3. In the present embodiment, by providing the dictionary file 220 in the conversion server 2, it is possible to reduce the storage capacity for voice output required for the mobile phone device 3.

辞書ファイル220に格納される発音データは、日本語の場合は文字(文字列)と音節が対応している場合が殆どなので、例えば、音節とアクセント(強弱)、高低等を示す文字によって構成される。例えば、「秋の七草」は「[ア]↓キノ・ナ[ナ]クサ」、「千載一遇」は「セ[ンザイ]・イ[チグー]」のように、音節はカタカナで、アクセントのある音節には大括弧「[ ]」で囲むことで、高低は矢印「↑」「↓」で、切れ目は中黒「・」で表現する。その他の言語については、それぞれの言語の発音体系に応じた表現方法を用いればよい。   The pronunciation data stored in the dictionary file 220 is mostly composed of characters indicating character (character string) and syllable in the case of Japanese. The For example, “Autumn Herb” is “[A] ↓ Kino Na [Na] Kussa”, “Senri Ichigo” is “Se Nizai Yi [Chigu]”, and the syllables are katakana and accents. By enclosing a certain syllable in square brackets “[]”, the high and low are expressed by arrows “↑” and “↓”, and the break is expressed by middle black “•”. For other languages, an expression method corresponding to the pronunciation system of each language may be used.

発音データ変換部23は、メールサーバ1から受信した電子メールや、アクセスしたウェブサイトの情報に含まれるテキストデータ取得する。そして、記憶部22に記憶された辞書ファイル220を参照してテキストデータを解析し、受信した電子メールやウェブサイトのテキストデータをテキスト形式の発音データに変換する。   The pronunciation data conversion unit 23 acquires text data included in the information of the electronic mail received from the mail server 1 or the accessed website. Then, the text data is analyzed with reference to the dictionary file 220 stored in the storage unit 22 and the received text data of the e-mail or website is converted into text-formatted pronunciation data.

第二送受信部24は、発音データ変換部23によって出力された発音データとともに、電子メールや指定されたウェブサイトの内容を携帯電話装置3に送信する。また、携帯電話装置3からのURL(Uniform Resource Locator)等のウェブサイト指定の指示を受信する。   The second transmission / reception unit 24 transmits the e-mail and the contents of the designated website to the mobile phone device 3 together with the pronunciation data output by the pronunciation data conversion unit 23. In addition, an instruction for specifying a website such as a URL (Uniform Resource Locator) from the mobile phone device 3 is received.

サーバ制御部25は、第一送受信部21、記憶部22、発音データ変換部23、第二送受信部24の動作を制御する。第1送信部21が電子メールを受信した場合、サーバ制御部25は、記憶部22から辞書ファイルを読出して発音データ変換部23に受信した電子メールの発音データ変換指示を行い、第二送受信部24に対しては発音データと受信した電子メールの送信を指示する。   The server control unit 25 controls operations of the first transmission / reception unit 21, the storage unit 22, the pronunciation data conversion unit 23, and the second transmission / reception unit 24. When the first transmission unit 21 receives the e-mail, the server control unit 25 reads the dictionary file from the storage unit 22 and instructs the pronunciation data conversion unit 23 to convert the received pronunciation data to the second transmission / reception unit. 24 is instructed to transmit the pronunciation data and the received electronic mail.

また、サーバ制御部25は、第二送受信部24が携帯電話装置3からウェブサイト指定指示を受信した場合、第一送受信部21に対して指定されたウェブサイトへのアクセスを指示する。第一送受信部21がウェブサイトの内容を取得したら、記憶部22から辞書ファイルを読出して発音データ変換部23で受信したウェブサイトの内容の発音データ変換指示を行い、第二送受信部24に対して発音データと受信したウェブサイトの内容の送信を指示する。   Further, when the second transmission / reception unit 24 receives the website designation instruction from the mobile phone device 3, the server control unit 25 instructs the first transmission / reception unit 21 to access the designated website. When the first transmission / reception unit 21 acquires the contents of the website, the dictionary data is read from the storage unit 22 and the pronunciation data conversion instruction of the website contents received by the pronunciation data conversion unit 23 is given. Instructing the transmission of the pronunciation data and the contents of the received website.

図3は、本発明の実施形態を説明するための携帯電話装置の概略構成を示すブロック図である。ここでは、通常の携帯電話装置としての機能に関してはその説明を省略する。図3に示すように、本実施形態の携帯電話装置3は、端末送受信部31と、端末記憶部32と、音声変換部33と、音声出力部34と、表示部35と、操作部36と、端末制御部37とを備える。   FIG. 3 is a block diagram showing a schematic configuration of a mobile phone device for explaining an embodiment of the present invention. Here, the description of the function as a normal cellular phone device is omitted. As shown in FIG. 3, the cellular phone device 3 of the present embodiment includes a terminal transmission / reception unit 31, a terminal storage unit 32, a voice conversion unit 33, a voice output unit 34, a display unit 35, and an operation unit 36. The terminal control unit 37 is provided.

端末送受信部31は、通信網5を介して、発音データの受信等の変換サーバ2とのデータ通信を行う。端末記憶部32は、発音データの示す発音に対応付けられた音声情報を格納する発音ファイル320と、受信した発音データに対応する音声信号を変換するためのソフトウエアである読上げアプリケーションプログラム(以下、読上げアプリ)321とを記憶する。   The terminal transmission / reception unit 31 performs data communication with the conversion server 2 such as reception of sound generation data via the communication network 5. The terminal storage unit 32 includes a pronunciation file 320 that stores voice information associated with the pronunciation indicated by the pronunciation data, and a reading application program (hereinafter referred to as software) for converting a voice signal corresponding to the received pronunciation data. Reading application) 321 is stored.

この発音ファイル320と読上げアプリ321は、あらかじめ端末記憶部32に記憶させてもよいし、端末送受信部31を介して、図示しないサービスサーバからダウンロードして端末記憶部32に記憶させてもよい。したがって、あらかじめ発音ファイル320と読上げアプリ321を備えない携帯電話装置であっても、プログラムをダウンロード可能であれば、本実施形態の音声出力システムに適用することができる。また、発音ファイル320と読上げアプリ321は、起動時や任意の期間毎、使用者の指示に応じて更新データを取得することで更新可能である。従って、端末側の音声変換に起因する音声出力の不具合が生じた場合も、発音ファイル320または読上げアプリ321を更新することで対処することができる。   The pronunciation file 320 and the reading application 321 may be stored in the terminal storage unit 32 in advance, or may be downloaded from a service server (not shown) via the terminal transmission / reception unit 31 and stored in the terminal storage unit 32. Therefore, even a mobile phone device that does not include the pronunciation file 320 and the reading application 321 in advance can be applied to the audio output system of the present embodiment as long as the program can be downloaded. Further, the pronunciation file 320 and the reading application 321 can be updated by acquiring update data in response to a user instruction at the time of activation or every arbitrary period. Therefore, even when a problem in voice output due to voice conversion on the terminal side occurs, it can be dealt with by updating the pronunciation file 320 or the reading application 321.

発音ファイル320は、上述したように日本語の場合では文字と音節が対応する場合が多いので、例えば音節単位で記憶される。日本語の音節は約114種類あるので、発音ファイル320は、それぞれの音節に対して、アクセント(強弱)2種類、音程(高低)2種類の、合計456種類の音声情報を格納している。   The pronunciation file 320 is stored in units of syllables, for example, since characters and syllables often correspond in the case of Japanese as described above. Since there are about 114 types of Japanese syllables, the pronunciation file 320 stores a total of 456 types of audio information, with two types of accents (strong and weak) and two types of pitches (high and low) for each syllable.

音声変換部33は、読上げアプリ321によって動作し、受信した発音データを、発音ファイル320に格納された音声情報を合成して音声出力用の音声信号に変換する。音声出力部は、音声変換部33によって生成された音声信号を再生出力する。   The voice conversion unit 33 is operated by the reading application 321 and synthesizes the received sound data with the sound information stored in the sound file 320 and converts it into a sound signal for sound output. The audio output unit reproduces and outputs the audio signal generated by the audio conversion unit 33.

表示部35は、受信した電子メールやウェブサイトを表示する。操作部36はテンキーや方向指示キー等を備え、所望のウェブサイト等を指定して、端末送受信部31を介してその指定情報を変換サーバ2へ送信する。ウェブサイトの指定情報は、操作部36においてURL等の情報を入力して変換サーバ2へ送信してもよいし、ウェブサイトの内容が表示部35に表示されているときに、操作部36を用いて音声出力要求を行い、表示されているウェブサイトの情報を変換サーバ2へ送信してもよい。   The display unit 35 displays the received e-mail and website. The operation unit 36 includes a numeric keypad, a direction instruction key, and the like, designates a desired website, and transmits the designation information to the conversion server 2 via the terminal transmission / reception unit 31. The designation information of the website may be transmitted to the conversion server 2 by inputting information such as URL in the operation unit 36, or when the content of the website is displayed on the display unit 35, It is also possible to make a voice output request and transmit the displayed website information to the conversion server 2.

制御部38は、端末送受信部31、端末記憶部32、音声変換部33、音声出力部34、表示部35、および操作部36の動作を制御する。   The control unit 38 controls operations of the terminal transmission / reception unit 31, the terminal storage unit 32, the voice conversion unit 33, the voice output unit 34, the display unit 35, and the operation unit 36.

このような本実施形態の音声出力システムによれば、また、処理量の多いテキストデータの解析を変換サーバで行っているので、携帯電話装置での処理量を減少させることができる。また、変換サーバが容量の大きい辞書ファイルを持っているので、携帯電話装置の記憶領域を増やさずに、より正確な音声出力を実現することができる。   According to such an audio output system of the present embodiment, since the conversion server analyzes text data with a large amount of processing, the amount of processing in the mobile phone device can be reduced. In addition, since the conversion server has a dictionary file with a large capacity, more accurate voice output can be realized without increasing the storage area of the mobile phone device.

更に、テキスト形式の発音データを用いることにより、変換サーバと携帯電話装置との間の通信データ量を減少させることができ、また、SMTPやHTTP等の通常の通信プロトコルを利用できるので、音声出力のための特別なシステムを構築する必要がない。このように、テキストデータを読まずに、アクセント等を含む音声出力を実現することができるので、視覚障害者にも電子メールの送受信等が可能となる。   Further, by using text-formatted pronunciation data, the amount of communication data between the conversion server and the mobile phone device can be reduced, and a normal communication protocol such as SMTP or HTTP can be used. There is no need to build a special system for. In this way, voice output including accents and the like can be realized without reading text data, so that visually impaired persons can also send and receive e-mails.

なお、本実施形態では、携帯電話装置において電子メールやウェブサイトの表示するとともに音声出力を行う場合について説明したが、変換サーバは発音データのみを携帯電話装置に送信して、音声出力のみを行ってもよい。   In the present embodiment, the case where the mobile phone device displays an e-mail or a website and performs voice output has been described. However, the conversion server transmits only the pronunciation data to the mobile phone device and performs only the voice output. May be.

また、受信した電子メールや指定したウェブサイトについて音声出力する場合を例にとって説明したが、使用者が指定した任意の文字列について音声出力を実行してもよい。また、表示部の画面上でハイパーリンク等が張られた部分を指定されたときにその情報を変換サーバへ送信して、その指定された部分を携帯電話機で音声出力してもよい。   Moreover, although the case where the voice is output for the received electronic mail or the specified website has been described as an example, the voice output may be executed for an arbitrary character string specified by the user. In addition, when a part with a hyperlink or the like is designated on the screen of the display unit, the information may be transmitted to the conversion server, and the designated part may be output by voice using a mobile phone.

本発明の音声出力システムおよびサーバ装置は、端末装置の処理の負担を軽減するとともに、通信データ量を減少可能な効果を有し、携帯無線装置等による音声出力システムに有用である。   The voice output system and the server device of the present invention have the effect of reducing the processing load of the terminal device and reducing the amount of communication data, and are useful for a voice output system using a portable wireless device or the like.

本発明の実施形態を説明するための音声出力システムを示す概略構成図The schematic block diagram which shows the audio | voice output system for describing embodiment of this invention 本発明の実施形態を説明するための変換サーバの概略構成を示すブロック図The block diagram which shows schematic structure of the conversion server for describing embodiment of this invention 本発明の実施形態を説明するための携帯電話装置の概略構成を示すブロック図The block diagram which shows schematic structure of the mobile telephone apparatus for describing embodiment of this invention

符号の説明Explanation of symbols

1 メールサーバ
2 変換サーバ
3 携帯電話装置
4、5 通信網
21 第一送受信部
22 サーバ記憶部
23 発音データ変換部
24 第二送受信部
25 サーバ制御部
31 端末送受信部
32 端末記憶部
33 音声変換部
34 音声出力部
35 表示部
36 操作部
37 端末制御部
DESCRIPTION OF SYMBOLS 1 Mail server 2 Conversion server 3 Mobile phone apparatus 4, 5 Communication network 21 1st transmission / reception part 22 Server memory | storage part 23 Pronunciation data conversion part 24 2nd transmission / reception part 25 Server control part 31 Terminal transmission / reception part 32 Terminal storage part 33 Voice conversion part 34 voice output unit 35 display unit 36 operation unit 37 terminal control unit

Claims (11)

テキストデータを端末装置において音声出力する音声出力システムに用いられるサーバ装置であって、
前記テキストデータを取得する取得手段と、
前記テキストデータを解析して、記号化および/または数値化された発音データを作成する発音データ作成部と、
前記発音データを前記端末装置へ送信する送信部と、
を備えるサーバ装置。
A server device used in a voice output system that outputs text data in a terminal device,
Obtaining means for obtaining the text data;
A pronunciation data creation unit that analyzes the text data to create symbolized and / or digitized pronunciation data;
A transmission unit for transmitting the pronunciation data to the terminal device;
A server device comprising:
請求項1記載のサーバ装置であって、
前記発音データ作成部は、文字、音節、および文節のうち少なくとも1種類と、それに対応付けられた発音データを含む発音辞書データを記憶した記憶部を有し、前記発音辞書データを参照して、前記テキストデータから文字、音節、および文節のうち少なくとも1種類を認識するとともに、前記認識された文字、音節、および文節のうち少なくとも1種類に対応する発音を特定して前記発音データを作成するものであるサーバ装置。
The server device according to claim 1,
The pronunciation data creation unit includes a storage unit that stores pronunciation dictionary data including at least one type of characters, syllables, and phrases and pronunciation data associated with the type, and refers to the pronunciation dictionary data, Recognizing at least one of characters, syllables, and phrases from the text data, and generating the pronunciation data by specifying a pronunciation corresponding to at least one of the recognized characters, syllables, and phrases Server device that is.
請求項2記載のサーバ装置であって、
前記発音辞書データは、漢字、ひらがな、カタカナを含む文字、音節、および文節のうち少なくとも1種類と、それに対応付けられた発音データを含むものであるサーバ装置。
The server device according to claim 2,
The server device, wherein the pronunciation dictionary data includes at least one type of characters, syllables, and phrases including kanji, hiragana, and katakana, and pronunciation data associated therewith.
請求項1ないし3のいずれか一項記載のサーバ装置であって、
前記発音データは、発音、強弱、および音程を表す記号および/または数値を含むものであるサーバ装置。
The server device according to any one of claims 1 to 3,
The server device is a server device that includes symbols and / or numerical values representing pronunciation, strength, and pitch.
請求項1ないし4のいずれか一項記載のサーバ装置であって、
前記取得部は、配信された電子メールを受信するとともに、前記電子メールの前記テキストデータを取得し、
前記送信部は前記発音データとともに前記電子メールを前記端末装置へ送信するものであるサーバ装置。
The server device according to any one of claims 1 to 4,
The acquisition unit receives the distributed e-mail, acquires the text data of the e-mail,
The transmission device is a server device that transmits the electronic mail together with the pronunciation data to the terminal device.
請求項1ないし5のいずれか一項記載のサーバ装置であって、
前記取得部は、前記端末装置によって指定されたウェブサイトの情報を受信して前記ウェブサイトの内容を取得するとともに、前記ウェブサイトの内容から前記テキストデータを取得し、
前記送信部は前記発音データとともに前記ウェブサイトの内容を前記端末装置へ送信するものであるサーバ装置。
The server device according to any one of claims 1 to 5,
The acquisition unit receives the information on the website specified by the terminal device to acquire the content of the website, acquires the text data from the content of the website,
The transmission unit is a server device that transmits the content of the website together with the pronunciation data to the terminal device.
請求項1ないし6のいずれか一項記載のサーバ装置と、
前記サーバ装置から通信網を介して接続される通信端末装置と、
を備え、
前記通信端末装置は、
前記サーバ装置の前記通信部から送信された発音データを受信する端末通信部と、
前記端末通信部が受信した発音データから音声信号に変換する音声変換部と、
前記音声変換部で変換した前記音声信号を出力する音声出力部と、
を有する音声出力システム。
A server device according to any one of claims 1 to 6;
A communication terminal device connected via a communication network from the server device;
With
The communication terminal device
A terminal communication unit that receives the pronunciation data transmitted from the communication unit of the server device;
A voice conversion unit that converts the sound data received by the terminal communication unit into a voice signal;
An audio output unit for outputting the audio signal converted by the audio conversion unit;
An audio output system.
請求項7記載の音声出力システムであって、
前記端末装置は、前記サーバ装置の前記送信部から前記テキストデータを受信した場合に、前記テキストデータを表示する表示部を更に備える音声出力システム。
The voice output system according to claim 7,
The said terminal device is an audio | voice output system further provided with the display part which displays the said text data, when the said text data is received from the said transmission part of the said server apparatus.
請求項7または8記載の音声出力システムであって、
前記端末装置は、前記発音データに対応した音声情報と、前記発音データと前記音声情報とを対応付けて変換を行う工程を実行可能なソフトウエアとを記憶する端末記憶部を更に備え、
前記音声変換部は前記ソフトウエアに従って動作し、前記発音データに基づいて、前記端末記憶部に記憶された前記音声情報を読出して前記音声信号を生成するものである音声出力システム。
The voice output system according to claim 7 or 8,
The terminal device further includes a terminal storage unit that stores voice information corresponding to the pronunciation data, and software capable of executing a process of converting the pronunciation data and the voice information in association with each other.
The voice output system, wherein the voice conversion unit operates according to the software and reads the voice information stored in the terminal storage unit based on the sound generation data to generate the voice signal.
請求項7ないし9のいずれか一項記載の音声出力システムであって、
前記端末記憶部に記憶された前記ソフトウエアは、前記端末通信部を介して受信する更新データに基づいて更新可能であるものである音声出力システム。
The voice output system according to any one of claims 7 to 9,
The voice output system in which the software stored in the terminal storage unit can be updated based on update data received via the terminal communication unit.
請求項7ないし10のいずれか一項記載の音声出力システムであって、
前記端末装置は携帯無線装置である、音声出力システム。
The audio output system according to any one of claims 7 to 10,
The voice output system, wherein the terminal device is a portable radio device.
JP2003336950A 2003-09-29 2003-09-29 Audio output system and server device Pending JP2005106905A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2003336950A JP2005106905A (en) 2003-09-29 2003-09-29 Audio output system and server device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP2003336950A JP2005106905A (en) 2003-09-29 2003-09-29 Audio output system and server device

Publications (1)

Publication Number Publication Date
JP2005106905A true JP2005106905A (en) 2005-04-21

Family

ID=34532913

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2003336950A Pending JP2005106905A (en) 2003-09-29 2003-09-29 Audio output system and server device

Country Status (1)

Country Link
JP (1) JP2005106905A (en)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP2112650A1 (en) 2008-04-23 2009-10-28 Sony Ericsson Mobile Communications Japan, Inc. Speech synthesis apparatus, speech synthesis method, speech synthesis program, portable information terminal, and speech synthesis system
KR100923942B1 (en) 2007-12-04 2009-10-29 엔에이치엔(주) Method, system and computer readable recording medium for extracting text from a web page and converting it into a voice data file

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR100923942B1 (en) 2007-12-04 2009-10-29 엔에이치엔(주) Method, system and computer readable recording medium for extracting text from a web page and converting it into a voice data file
EP2112650A1 (en) 2008-04-23 2009-10-28 Sony Ericsson Mobile Communications Japan, Inc. Speech synthesis apparatus, speech synthesis method, speech synthesis program, portable information terminal, and speech synthesis system
EP3086318A1 (en) 2008-04-23 2016-10-26 Sony Mobile Communications Japan, Inc. Speech synthesis apparatus, speech synthesis method, speech synthesis program, and portable information terminal
US9812120B2 (en) 2008-04-23 2017-11-07 Sony Mobile Communications Inc. Speech synthesis apparatus, speech synthesis method, speech synthesis program, portable information terminal, and speech synthesis system
US10720145B2 (en) 2008-04-23 2020-07-21 Sony Corporation Speech synthesis apparatus, speech synthesis method, speech synthesis program, portable information terminal, and speech synthesis system

Similar Documents

Publication Publication Date Title
US8705705B2 (en) Voice rendering of E-mail with tags for improved user experience
JP3884851B2 (en) COMMUNICATION SYSTEM AND RADIO COMMUNICATION TERMINAL DEVICE USED FOR THE SAME
JP2009265279A (en) Voice synthesizer, voice synthetic method, voice synthetic program, personal digital assistant, and voice synthetic system
JP2005346252A (en) Information transmission system and information transmission method
US20110022948A1 (en) Method and system for processing a message in a mobile computer device
CN1329739A (en) Voice control of a user interface to service applications
US20100268525A1 (en) Real time translation system and method for mobile phone contents
JP2005310129A (en) How to display messages on the terminal
JP3714159B2 (en) Browser-equipped device
JPH11308270A (en) Communication system and terminal equipment used for the same
JPH10149361A (en) Information processing method and apparatus and storage medium
KR100826778B1 (en) Browser-based wireless terminal for multi-modal, Browser-based multi-modal server and system for wireless terminal and its operation method
JPH0766830A (en) Mail system
JP4403284B2 (en) E-mail processing apparatus and e-mail processing program
US20060129402A1 (en) Method for reading input character data to output a voice sound in real time in a portable terminal
JPH09258785A (en) Information processing method and information processing apparatus
JP2006171498A (en) Speech synthesis system, speech synthesis method, and speech synthesis server
JP2005266009A (en) Data conversion program and data conversion apparatus
JPH0764583A (en) Text reading method and device
JP4082249B2 (en) Content distribution system
JPH10143352A (en) Document information conversion device and conversion method
CN101014996A (en) Speech synthesis
US20090100150A1 (en) Screen reader remote access system
JP2002140086A (en) Short message-to-voice output converter for mobile phones
JP2005107320A (en) Data generator for voice reproduction

Legal Events

Date Code Title Description
RD04 Notification of resignation of power of attorney

Free format text: JAPANESE INTERMEDIATE CODE: A7424

Effective date: 20060325