JPH10282991A - Speech rate converting device - Google Patents

Speech rate converting device

Info

Publication number
JPH10282991A
JPH10282991A JP9083640A JP8364097A JPH10282991A JP H10282991 A JPH10282991 A JP H10282991A JP 9083640 A JP9083640 A JP 9083640A JP 8364097 A JP8364097 A JP 8364097A JP H10282991 A JPH10282991 A JP H10282991A
Authority
JP
Japan
Prior art keywords
segment
block
memory
pitch
output
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP9083640A
Other languages
Japanese (ja)
Inventor
Takao Katayama
貴夫 片山
Motoyasu Ono
元康 大野
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Panasonic System Solutions Japan Co Ltd
Original Assignee
Matsushita Graphic Communication Systems Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Matsushita Graphic Communication Systems Inc filed Critical Matsushita Graphic Communication Systems Inc
Priority to JP9083640A priority Critical patent/JPH10282991A/en
Publication of JPH10282991A publication Critical patent/JPH10282991A/en
Pending legal-status Critical Current

Links

Landscapes

  • Compression, Expansion, Code Conversion, And Decoders (AREA)

Abstract

PROBLEM TO BE SOLVED: To facilitate calculation of a correlative function by presuming pitch lengths of a head segment and the next segment the same ratio, overlapping both segments, shifting so that waveforms are matched with each other and evaluating the correlation function in the shifted state. SOLUTION: A first correlater 41 calculates the correlative function between the data in a first memory 38 and a second memory 39, and obtains a first correlative delay amount obtained by shifting relative positions so that the large positions of the correlative function of two segments coincide with each other. On the other hand, a waveform deviation value calculator 42 obtains pitch information from a pitch buffer 34, and presumes that the waveforms of the head segment and the next segment have the same pitch, and only different phases, and obtains a deviation value shifting so as to be matched with each other by shifting. A second correlater 43 calculates the correlative function based on the data of the first memory 38 and the second memory 39 and the deviation value from the waveform deviation value calculator 42, and obtains a second correlative delay amount obtained by shifting the relative positions so that the large positions of the correlative function of two segments coincide with each other.

Description

【発明の詳細な説明】DETAILED DESCRIPTION OF THE INVENTION

【0001】[0001]

【発明の属する技術分野】本発明は、符号化した音声の
音の高さを変えず速度を変化させる音声速度変換装置に
関する。
BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a voice speed conversion device for changing the speed of a coded voice without changing its pitch.

【0002】[0002]

【従来の技術】留守番電話機などの伝言メッセージやテ
レビの映像を録画したビデオなどでは、聞きたいとか見
たいところを普通の速度で再生し、あまり重要でないと
ころや興味の無いところは高速で再生したい。また打ち
合わせ記録などを再生する場合、重要なところは遅い速
度で再生したい。このような目的のための音声を音の高
さ(周波数)を変えることなく速度のみを変化させる技
術が開発されている。
2. Description of the Related Art In a message message of an answering machine or a video recording a television picture, a user wants to play at a normal speed what he wants to hear or watch, and to play it at a high speed in places that are not important or in which he is not interested. . When playing back meeting recordings, it is important to play back important parts at a slow speed. Techniques have been developed to change only the speed of voice for such purpose without changing the pitch (frequency) of the sound.

【0003】特開平4−104200号にはこのような
音声速度変換装置が示されている。これは音声をセグメ
ント単位で表し、複数のセグメントで1ブロックを構成
し、このブロックのセグメント数を増減して速度を低速
または高速にする。速度変化によって音声が不自然とな
るのをできるだけ少なくするため、速くする場合は1番
目と2番目のセグメントを重み付けして加算して1セグ
メントとし、これに3番目以降のセグメントを連続する
ことにより1ブロックのセグメントを1個減少して高速
にする。また低速にする場合は、1番目と2番目のセグ
メントを重み付けして加算して1セグメントとし、これ
に2番目以降のセグメントを連続し、さらにブロックの
最終セグメントの後にこれに続くセグメントを1個追加
することにより1ブロックのセグメント数を1個増加し
て低速にする。
[0003] Japanese Patent Application Laid-Open No. 4-104200 discloses such an audio speed converter. In this method, voice is expressed in units of segments, one block is composed of a plurality of segments, and the number of segments in this block is increased or decreased to make the speed low or high. In order to minimize the unnaturalness of the sound due to the speed change, when the speed is increased, the first and second segments are weighted and added to form one segment, and the third and subsequent segments are successively added thereto. One block segment is reduced by one to increase the speed. When the speed is to be reduced, the first and second segments are weighted and added to form one segment, the second and subsequent segments are continued, and one segment following this is added after the last segment of the block. By adding, the number of segments in one block is increased by one to reduce the speed.

【0004】図17は特開平4−104200号の音声
速度変換装置を示すブロック図である。音声速度変換は
速度比αに基づいて行われる。 α=(速度変換後の再生時間長)/(速度再生前の再生時間長)……(1) 入力信号A/D変換器11でデジタル信号に変換され、
データバッファ12に格納される。音声データは図2の
(A)に示すようにピッチLの波形が連続する場合が多
い。このため(B)に示すようにピッチL内のデータが
入るようなセグメントを設定し、このセグメントを複数
個まとめて、1ブロックとし、このブロックのセグメン
ト数を増減して音声の速度を増減する。
FIG. 17 is a block diagram showing a voice speed converter disclosed in Japanese Patent Laid-Open No. 4-104200. The voice speed conversion is performed based on the speed ratio α. α = (reproduction time length after speed conversion) / (reproduction time length before speed reproduction) (1) The input signal is converted into a digital signal by the A / D converter 11, and
The data is stored in the data buffer 12. As shown in FIG. 2A, audio data often has a continuous waveform having a pitch L. For this reason, as shown in (B), a segment in which data within the pitch L is set is set, and a plurality of these segments are grouped into one block, and the number of segments in this block is increased or decreased to increase or decrease the voice speed. .

【0005】速度制御回路13は分割器14を制御して
ブロックの最初のセグメントのデータを第1メモリ15
に格納し、次のセグメントのデータ第2メモリ16に格
納する。相関器17は第1メモリ15と第2メモリ16
のデータの相関関数の演算を行い相関遅延量(rk)を
算出する。相関遅延量rkは図3(B)または(C)に
示すrk+またはrk−である。(A)は各セグメント
内の音声波形を示す。bの波形を基準とし、aまたはc
の波形がbの波形よりずれる量をrkとし、bに対して
右側(時間的に後の方)にずらした場合を(+)とし、
左側(時間的に前の方)にずらした場合を(−)とす
る。
The speed control circuit 13 controls the divider 14 to store the data of the first segment of the block in the first memory 15.
And the data of the next segment is stored in the second memory 16. The correlator 17 includes a first memory 15 and a second memory 16
The calculation of the correlation function of the data is performed to calculate the correlation delay amount (rk). The correlation delay amount rk is rk + or rk- shown in FIG. 3B or 3C. (A) shows a speech waveform in each segment. a or c based on the waveform of b
Rk is the amount by which the waveform of b is deviated from the waveform of b, and (+) is the case where the waveform is shifted to the right (later in time) with respect to b.
A case shifted to the left side (time forward) is defined as (-).

【0006】窓関数発生器20は第1メモリ15と第2
メモリ16のデータの重み係数を発生する。重み係数は
速度比αや相関器17の発生する相関関数に基づいて
増、減する窓関数を発生し、第1メモリ15と第2メモ
リ16の対応する各データに対して、それぞれの重み係
数の和が1.0になるように設定されている。
The window function generator 20 includes a first memory 15 and a second memory 15.
A weight coefficient for the data in the memory 16 is generated. The weight coefficient is based on the speed ratio α and the correlation function generated by the correlator 17.
A window function to increase or decrease is generated, and the sum of the respective weighting factors is set to 1.0 for each corresponding data in the first memory 15 and the second memory 16.

【0007】第1乗算器18で第1メモリ15のデータ
と重み係数を乗算し、第2乗算器19で第2メモリ16
のデータと重み係数を乗算し、結果を加算器21に出力
する。加算器21は相関器17の相関遅延量(rk)に
基づいて、図3の(B)または(C)に示すように相関
の高い位置に一方をrkずらして加算して結合器22に
出力する。速度制御回路13は分割器14から、低速化
または高速化に応じて第2セグメント以降、または第3
セグメント以降のセグメントを出力させ、結合器22は
この値を加算値の後に続けて1ブロックのデータとして
D/A変換器23に出力する。D/A変換器23ではデ
ジタル信号をアナログ信号に変換して出力する。
The first multiplier 18 multiplies the data in the first memory 15 by the weighting coefficient, and the second multiplier 19 multiplies the data in the second memory 16.
Is multiplied by the weight coefficient, and the result is output to the adder 21. The adder 21 adds one to the position having a high correlation by shifting it by rk based on the correlation delay amount (rk) of the correlator 17 as shown in FIG. I do. The speed control circuit 13 outputs the signal from the divider 14 to the second and subsequent segments or the third
The segments 22 and subsequent segments are output, and the combiner 22 outputs this value to the D / A converter 23 as one block of data following the added value. The D / A converter 23 converts the digital signal into an analog signal and outputs the analog signal.

【0008】図18は速度比αが3/4と3/2の場合
の音声速度変換結果を示す。(a)〜(d)は速度比α
が3/4の場合を示し、(e)〜(h)はα=3/2の
場合を示す。(a)は1ブロックに4セグメントが入っ
ている元の音声データを示す。(b)は相関遅延量rk
が0の場合、つまりセグメント1とセグメント2の波形
は同一の場合を示す。このとき第1セグメントと第2セ
グメントの加算値は元のセグメントの長さと同じとな
り、1ブロックの長さは3/4になる。なお先頭のセグ
メントの斜線は重み係数を示し、元の第1セグメントへ
の重み係数は1より0まで漸減し、元の第2セグメント
への重み係数は0より1まで漸増する。セグメントの長
さは、加算して生成したセグメントの長さは増減するが
その他はtsで固定である。
FIG. 18 shows the result of voice speed conversion when the speed ratio α is 3/4 and 3/2. (A) to (d) show the speed ratio α
Shows the case where 3/4, and (e) to (h) show the case where α = 3/2. (A) shows original audio data in which one block contains four segments. (B) is the correlation delay amount rk
Is 0, that is, the waveforms of segment 1 and segment 2 are the same. At this time, the added value of the first segment and the second segment is the same as the length of the original segment, and the length of one block is /. The oblique line of the first segment indicates a weighting factor, and the weighting factor for the original first segment gradually decreases from 1 to 0, and the weighting factor for the original second segment gradually increases from 0 to 1. As for the length of the segment, the length of the segment generated by the addition is increased or decreased, but the others are fixed at ts.

【0009】(c)は相関遅延量rkが(+)の場合
で、図3の(B)の場合である。この場合、第1セグメ
ントと第2セグメントとを加算したセグメントはrkだ
け元のセグメントの長さtsより長くなる。また(d)
は相関遅延量rkが(−)の場合で、図3の(C)に示
すように第1セグメントと第2セグメントとを加算した
セグメントはrkだけ元のセグメントの長さtsより短
くなる。このため(c),(d)の場合は速度比αは3
/4にはならない。データポインタP1,P2は第1メ
モリ15、第2メモリ16に入力するデータの先頭番地
を示し、P1は(a)に示すように3セグメントの長さ
(3ts)の次の位置を示し、これは次のブロックの最
初のセグメントの先頭位置を示し、(b)〜(d)とも
同じ位置である。P2は次のブロックの第2セグメント
の先頭位置を示し、この場合セグメント6の先頭アドレ
スを示す。
FIG. 3C shows the case where the correlation delay amount rk is (+), which is the case of FIG. 3B. In this case, the segment obtained by adding the first segment and the second segment is longer than the original segment length ts by rk. (D)
Is the case where the correlation delay amount rk is (-), and the segment obtained by adding the first segment and the second segment is shorter than the original segment length ts by rk as shown in FIG. 3C. Therefore, in the cases (c) and (d), the speed ratio α is 3
It does not become / 4. The data pointers P1 and P2 indicate the start addresses of the data input to the first memory 15 and the second memory 16, and P1 indicates the next position of the length of 3 segments (3 ts) as shown in FIG. Indicates the start position of the first segment of the next block, and is the same position in (b) to (d). P2 indicates the start position of the second segment of the next block, and in this case, indicates the start address of segment 6.

【0010】(e)〜(h)は速度比αが3/2の場合
を示し、(e)は1ブロックが2セグメントよりなり、
これにセグメント3を追加し速度比αを3/2にするこ
とが示されている。(f)は相関遅延量rkが0の場合
で、第1セグメントと第2セグメントを加算して元のセ
グメントの長さtsと同じ長さにし、第2セグメントと
次のブロックのセグメント3を加算して1ブロックの長
さを3セグメントの長さ(3ts)にしている。(g)
はrkが(+)の場合で加算したセグメントの長さは1
セグメントの長さtsより大きくなる。(h)は、rk
が(−)の場合で、加算したセグメントの長さは1セグ
メントの長さtsより短くなる。これにより(g)では
1ブロックが3tsより長くなり、(h)では短くな
る。ポインタP1はセグメント2の後にあり、ポインタ
P2は(f)のセグメント3の後で次のブロックの先頭
セグメントの先頭位置を示す。
(E) to (h) show the case where the speed ratio α is 3/2, and (e) shows that one block consists of two segments,
It is shown that a segment 3 is added to this to make the speed ratio α 3/2. (F) is a case where the correlation delay amount rk is 0, the first segment and the second segment are added to have the same length as the original segment length ts, and the second segment and the segment 3 of the next block are added. Thus, the length of one block is set to the length of three segments (3 ts). (G)
Is the case where rk is (+) and the length of the added segment is 1
It is larger than the segment length ts. (H) is rk
Is (-), the length of the added segment is shorter than the length ts of one segment. As a result, one block becomes longer than 3 ts in (g), and becomes shorter in (h). The pointer P1 is after the segment 2, and the pointer P2 indicates the start position of the start segment of the next block after the segment 3 in (f).

【0011】[0011]

【発明が解決しようとする課題】相関関数を演算する相
関器の演算量が多大で、大容量の演算器が必要になる。
また図18の(c),(g)に示すように相関遅延量r
k+が大きいとポインタP1とP2とが離れてしまい、
第1メモリと第2メモリのデータの類似性が損なわれる
ことにより、相関器の出力する相関遅延量の正確性が失
われ、再生音声の品質が劣化する。また加算して生成さ
れるセグメント以外はセグメント長さtsが固定である
ため、音声データのピッチがセグメント長tsより長い
場合、ブロックの最初のセグメントと次のセグメントの
類似性の比較が精度よくできないという問題もある。
The amount of calculation of the correlator for calculating the correlation function is enormous, and a large-capacity calculator is required.
Also, as shown in FIGS. 18C and 18G, the correlation delay amount r
If k + is large, the pointers P1 and P2 are separated,
When the similarity between the data in the first memory and the data in the second memory is impaired, the accuracy of the correlation delay amount output from the correlator is lost, and the quality of the reproduced sound is degraded. In addition, since the segment length ts is fixed except for the segments generated by addition, if the pitch of the audio data is longer than the segment length ts, the similarity between the first segment and the next segment of the block cannot be compared with high accuracy. There is also a problem.

【0012】本発明は上述の問題点に鑑みてなされたも
ので、相関器の相関関数の演算を容易にすることを目的
とする。また、ポインタの位置を次のブロックの第1セ
グメントと第2セグメントの先頭位置を示すようにし、
2つのポインタの位置が離れ過ぎないようにする。また
音声データのピッチ長さがセグメントより長い場合、セ
グメントの長さを可変とすることを目的とする。
The present invention has been made in view of the above problems, and has as its object to facilitate calculation of a correlation function of a correlator. Further, the position of the pointer is set to indicate the start position of the first segment and the second segment of the next block,
Make sure that the two pointers are not too far apart. It is another object of the present invention to make the length of the segment variable when the pitch length of the audio data is longer than the segment.

【0013】[0013]

【課題を解決するための手段】本発明の音声速度変換装
置においては、ブロックの先頭セグメントと次のセグメ
ントのピッチ長を同一比と仮定し両セグメントを重ねて
波形が一致するようにずらし、そのずらし量を求め、次
に相関器によりずらした状態で両セグメントの相関関数
を算出する。
In the voice speed converter of the present invention, the pitch length of the first segment of a block and the next segment are assumed to have the same ratio, and both segments are overlapped and shifted so that their waveforms match. The shift amount is obtained, and then the correlation function of both segments is calculated in a state shifted by the correlator.

【0014】本発明によれば2つの隣接するセグメント
の相関関数の算出が容易になる。
According to the present invention, the calculation of the correlation function between two adjacent segments is facilitated.

【0015】[0015]

【発明の実施の形態】請求項1の発明では、符号化した
音声データを復号化する復号器と、この復号器より得ら
れる音声データのピッチ情報と1つのピッチ内のデータ
が隣接する時間的に後のピッチのデータに類似するか前
のピッチのデータに類似するかを示すピッチ極性情報と
を格納するピッチ格納部と、前記復号器より得られる音
声データのピッチが明確か不明確かを示したフラグを格
納するフラグ格納部と、前記復号器からのデータを所定
長さのセグメントに分割し連続する所定数のセグメント
を1つのブロックとし、この各ブロックのセグメントを
入力した順に分けて出力する分割器と、この分割器が出
力する各ブロックの最初のセグメントを格納する第1メ
モリと、前記分割器が出力する各ブロックの2番目のセ
グメントを格納する第2メモリと、前記第1メモリの内
容と前記第2メモリの内容との相関関数を算出し、第1
相関遅延量を演算する第1相関器と、前記ピッチ格納部
のデータより前記最初のセグメントの波形と前記次のセ
グメントの波形とがほぼ一致するようにずらして重ねそ
のずれ量を算出する波形ずれ量算出器と、前記第1メモ
リの内容と前記第2メモリの内容との相関関数を前記波
形ずれ量算出器の出力に基づいて算出し第2相関遅延量
を演算する第2相関器と、セグメントのデータに対する
重み係数を発生する窓関数発生器と、前記第1メモリの
内容に前記重み係数を乗算する第1乗算器と、前記第2
メモリの内容に前記重み係数を乗算する第2乗算器と、
前記第1相関器または前記第2相関器の出力に基づき前
記第1乗算器の出力と前記第2乗算器の出力とを相関関
数の値が大きい位置で加算する加算器と、前記第1相関
器と前記第2相関器との出力を前記フラグ格納部のフラ
グに応じて切り替えて出力する切替えスイッチと、前記
加算器の出力とこれに続けて、前記分割器より入力する
低速化のときは第2セグメント以降、高速化の時は第3
セグメント以降のセグメントで構成されるスルー区間セ
グメントを出力する結合器と、音声の速度の変換比を表
す速度比に基づき前記分割器と前記結合器を制御する速
度制御部と、を備える。
According to the first aspect of the present invention, there is provided a decoder for decoding coded audio data, and a temporal information in which pitch information of audio data obtained from the decoder and data within one pitch are adjacent to each other. A pitch storage unit for storing pitch polarity information indicating whether the data is similar to the data of the subsequent pitch or the data of the previous pitch, and indicates whether the pitch of the audio data obtained from the decoder is clear or unknown. And a flag storage unit for storing the generated flag, and divides the data from the decoder into segments of a predetermined length, and sets a predetermined number of continuous segments into one block, and divides and outputs the segments of each block in the order of input. A divider, a first memory for storing a first segment of each block output from the divider, and a second memory for storing a second segment of each block output from the divider. Calculated a second memory, the correlation function between the contents of the content and the second memory of the first memory, the first
A first correlator for calculating a correlation delay amount, and a waveform shift for calculating a shift amount by shifting the waveform of the first segment from the data of the pitch storage unit so that the waveform of the next segment substantially coincides with the waveform of the next segment. An amount calculator, a second correlator that calculates a correlation function between the contents of the first memory and the contents of the second memory based on the output of the waveform shift amount calculator and calculates a second correlation delay amount; A window function generator for generating a weighting factor for the data of the segment; a first multiplier for multiplying the content of the first memory by the weighting factor;
A second multiplier for multiplying the content of the memory by the weighting factor;
An adder for adding the output of the first multiplier and the output of the second multiplier at a position where the value of the correlation function is large based on the output of the first correlator or the output of the second correlator; A switch for switching and outputting the outputs of the adder and the second correlator in accordance with the flag in the flag storage unit, and the output of the adder and, subsequently, when the speed is reduced from the divider, From the second segment onwards, when speeding up, the third
A combiner that outputs a through section segment composed of segments subsequent to the segment, and a speed control unit that controls the divider and the combiner based on a speed ratio that represents a conversion ratio of the speed of audio.

【0016】符号化したデータを復号化するがその際、
音声データのピッチ情報、ピッチ極性情報およびピッチ
が明確に表れているか否かを表すフラグが得られる。音
声データをセグメントに分割し、所定数のセグメントを
1ブロックとし、このブロック単位で音声速度変換を行
う。速度変換は変換後の再生時間長を変換前の再生時間
長で割った速度比αに基づいて行われ、1ブロックのセ
グメント数はこの速度比αに応じて予め定められる。速
度制御部は分割器を制御してブロックの最初のセグメン
トを第1メモリへ出力し、2番目のセグメントを第2メ
モリへ出力し、低速化または高速化に応じて第2セグメ
ント以降または第3セグメント以降を結合器へ出力す
る。第1相関器は最初のセグメントと2番目のセグメン
トとの相関関数を算出し、両セグメントの相関の高い位
置間の距離を示す第1相関遅延量を演算する。
[0016] Decoding the encoded data,
The pitch information, the pitch polarity information, and the flag indicating whether the pitch is clearly shown are obtained. The audio data is divided into segments, and a predetermined number of segments are made into one block, and audio speed conversion is performed in units of this block. The speed conversion is performed based on the speed ratio α obtained by dividing the converted playback time length by the playback time length before conversion, and the number of segments in one block is predetermined according to the speed ratio α. The speed control unit controls the divider to output the first segment of the block to the first memory, outputs the second segment to the second memory, and outputs the second and subsequent segments or the third segment according to the speed reduction or the speed increase. The segment and subsequent segments are output to the combiner. The first correlator calculates a correlation function between the first segment and the second segment, and calculates a first correlation delay amount indicating a distance between highly correlated positions of both segments.

【0017】波形ずれ量算出器では、ピッチ情報に基づ
いて最初のセグメントの波形と2番目のセグメントの波
形がほぼ一致するようにずらして重ねそのずれ量を算出
する。両セグメントの波形と位相が同じであれば、ずれ
量は0となる。このように両セグメントの波形をほぼ一
致するようにずらした状態で、第2相関器は第1メモリ
の最初のセグメントと第2メモリの2番目のセグメント
の相関関数を計算することにより相関の高い位置を求め
るための探索範囲が大幅に狭くなり、計算量が大幅に減
少する。これにより第2相関遅延量の演算を容易に行う
ことができる。
The waveform shift amount calculator calculates the shift amount by shifting the waveform of the first segment and the waveform of the second segment so as to substantially coincide with each other based on the pitch information. If the waveform and the phase of both segments are the same, the shift amount is zero. The second correlator calculates the correlation function between the first segment of the first memory and the second segment of the second memory in such a state that the waveforms of both segments are shifted so as to substantially coincide with each other, so that the correlation is high. The search range for finding the position is greatly narrowed, and the amount of calculation is greatly reduced. This makes it possible to easily calculate the second correlation delay amount.

【0018】窓関数発生器は最初のセグメントと2番目
のセグメントのデータに対してそれぞれ重み係数を予め
定められた窓関数により発生する。重み係数は、両セグ
メントを加算して新たな加算セグメントとした場合、こ
の加算セグメントのデータの大きさが元のセグメントの
データと同じ大きさになるようにし、加算によって音声
が不自然にならないようにするため設けられている。こ
のため両セグメントの加算するデータの一方の重み係数
と他方の重み係数との和は1.0となる。また一方のセ
グメントの重み係数はセグメントの先頭から後端にゆく
に従い漸増し、他方のセグメントの重み係数はセグメン
トの先頭から後端にゆくに従い漸減するようにしてい
る。
The window function generator generates a weight coefficient for each of the data of the first segment and the data of the second segment by a predetermined window function. The weighting factor is such that, when both segments are added to form a new added segment, the data size of the added segment is the same as the data of the original segment, so that the addition does not make the sound unnatural. It is provided in order to. Therefore, the sum of one weight coefficient and the other weight coefficient of the data to be added in both segments is 1.0. The weight coefficient of one segment increases gradually from the head of the segment to the rear end, and the weight coefficient of the other segment decreases gradually from the head of the segment to the rear end.

【0019】第1乗算器は第1メモリのデータと重み係
数を乗算し、第2乗算器は第2メモリのデータと重み係
数を乗算する。切替スイッチはフラグ格納部のフラグに
応じて第1相関遅延量か第2相関遅延量かを出力する。
加算器は第1および第2乗算器の出力を第1または第2
相関遅延量に基づき、相関関数の大きな位置に一方をず
らして加算を行う。結合器はこの加算によって得られた
加算セグメントの後に分割器から入力したスルー区間セ
グメントを続けて1ブロックのデータとして出力する。
これにより第2相関器を用いる場合は相関遅延量の算出
を迅速に行うことができる。
The first multiplier multiplies the data in the first memory by the weighting factor, and the second multiplier multiplies the data in the second memory by the weighting factor. The changeover switch outputs the first correlation delay amount or the second correlation delay amount according to the flag in the flag storage unit.
The adder outputs the output of the first and second multipliers to the first or second output.
Based on the correlation delay amount, one is shifted to a position where the correlation function is large, and the addition is performed. The combiner continuously outputs the through segment input from the divider after the addition segment obtained by the addition as one block of data.
Thus, when the second correlator is used, the calculation of the correlation delay amount can be performed quickly.

【0020】請求項2の発明では、前記速度制御部は、
前記ピッチ情報に基づいて前記セグメントの長さを変更
する。
According to the second aspect of the present invention, the speed control section includes:
The length of the segment is changed based on the pitch information.

【0021】音声データは女性の声と男性の声ではピッ
チが異なり、セグメントの長さを固定にしておくと、セ
グメントの長さを超えたピッチのデータの場合セグメン
ト間の相関値が劣化する。このためピッチに応じてセグ
メントの長さを可変にすることにより、相関値が良好に
なる。つまり音質が向上する。
The pitch of voice data differs between a female voice and a male voice. If the length of a segment is fixed, the correlation value between segments deteriorates in the case of data having a pitch exceeding the length of the segment. Therefore, by making the length of the segment variable according to the pitch, the correlation value is improved. That is, the sound quality is improved.

【0022】請求項3の発明では、前記セグメントの長
さを女性の声を基準にした初期セグメント長さに設定
し、前記ピッチ情報のピッチがこの初期セグメント長さ
を超える場合、セグメントの長さを初期セグメントの2
倍とする。
According to the third aspect of the present invention, the length of the segment is set to an initial segment length based on a female voice, and when the pitch of the pitch information exceeds the initial segment length, the length of the segment To the initial segment 2
Double it.

【0023】初期(ディフォルト)セグメントの長さを
女性の声を基準にした長さに設定しておき、ピッチがこ
の初期セグメントを越える場合は初期セグメントの2倍
の長さのセグメントに設定し、これに応じてブロックの
長さも2倍となる。1ブロックのセグメント数は変わら
ない。これにより1セグメント内には必ず1ピッチ以上
のデータが存在し、2つのセグメントの相関関数の算出
が可能となる。
The length of the initial (default) segment is set to a length based on the female voice, and if the pitch exceeds this initial segment, the length is set to a segment twice as long as the initial segment. Accordingly, the length of the block is also doubled. The number of segments in one block does not change. As a result, data of one pitch or more always exists in one segment, and the correlation function of two segments can be calculated.

【0024】請求項4の発明では、前記セグメントの長
さを切り替えたブロックにおいては、前記切替スイッチ
は、前記第1相関器の出力を前記加算器に出力する。
According to a fourth aspect of the present invention, in the block in which the length of the segment is switched, the switch outputs the output of the first correlator to the adder.

【0025】セグメントの長さを切り替えたブロックで
はピッチがやや不安定な部分であるため、波形ずれ量算
出器のずれ量が不正確になるため、第1相関器の相関遅
延量を採用し、音声の品質低下を防止する。
In the block in which the length of the segment is switched, the pitch is a part that is slightly unstable, so that the waveform shift amount calculator becomes inaccurate. Therefore, the correlation delay amount of the first correlator is used. Prevent audio quality degradation.

【0026】請求項5の発明では、前記切替スイッチ
は、前記フラグ格納部の音声データのピッチが明確な場
合、前記第2相関器の出力を前記加算器に出力する。
In the invention according to claim 5, the changeover switch outputs the output of the second correlator to the adder when the pitch of the audio data in the flag storage section is clear.

【0027】ピッチが不明確な場合、波形ずれ量算出器
ではずれ量の算出が不正確になるので、第1相関器の出
力を用いて加算する。ピッチが明確な場合は、ずれ量を
かなり正確に算出できるので、第2相関器の出力を用い
て第2相関遅延量の算出を迅速に行う。
When the pitch is unclear, the calculation of the shift amount becomes inaccurate in the waveform shift amount calculator, so that the addition is performed using the output of the first correlator. When the pitch is clear, the amount of deviation can be calculated quite accurately, so that the second correlation delay amount is quickly calculated using the output of the second correlator.

【0028】請求項6の発明では、前記ピッチ極性情報
に基づき前記ブロックの最初のセグメントと次のセグメ
ントのデータが類似しているとき前記切替スイッチは前
記第2相関器の出力を前記加算器に出力する。
According to the invention of claim 6, when the data of the first segment and the next segment of the block are similar based on the pitch polarity information, the changeover switch outputs the output of the second correlator to the adder. Output.

【0029】ブロックの最初のセグメントと次のセグメ
ントのデータが類似している時は、波形ずれ量算出器に
より両セグメントの波形が一致するようにずらすことが
正確にできるので、正確なずれ量が求まる。これにより
第2相関器の第2相関遅延量が正確かつ迅速に求まるの
で、第2相関遅延量により加算器は加算処理をする。
When the data of the first segment of the block and the data of the next segment are similar, the waveform shift amount calculator can accurately shift the waveforms of both segments so that they coincide with each other. I get it. As a result, the second correlation delay amount of the second correlator can be accurately and quickly obtained, and the adder performs an addition process based on the second correlation delay amount.

【0030】請求項7の発明では、前記初期セグメント
の2倍を新セグメントとし、最初の新セグメントはデー
タ伝送順に第1初期セグメントと第2初期セグメントか
らなり、次の新セグメントはデータ伝送順に第3初期セ
グメントと第4初期セグメントからなり、第1初期セグ
メントが第2初期セグメントに類似し、第2初期セグメ
ントと第3初期セグメントが類似し、第4初期セグメン
トが第3初期セグメントに類似している場合、前記切替
スイッチは前記第2相関器の出力を前記加算器に出力す
る。
In the present invention, twice the initial segment is set as a new segment, the first new segment is composed of a first initial segment and a second initial segment in the order of data transmission, and the next new segment is a second segment in the order of data transmission. The first initial segment is similar to the second initial segment, the second initial segment is similar to the third initial segment, and the fourth initial segment is similar to the third initial segment. If so, the changeover switch outputs the output of the second correlator to the adder.

【0031】第1初期セグメントと第2初期セグメント
が類似し、第2初期セグメントと第3初期セグメントが
類似し、第3初期セグメントと第4初期セグメントが類
似することにより、最初の新セグメントと次の新セグメ
ントは類似することになる。これにより波形ずれ量算出
器により両新セグメントの波形が一致するように正確に
ずらすことができ、正確なずれ量が求まる。このずれ量
を用いて第2相関器で第2相関遅延量を正確かつ迅速に
求めることができるので、第2相関遅延量により加算器
は加算する。
The first initial segment is similar to the second initial segment, the second initial segment is similar to the third initial segment, and the third initial segment is similar to the fourth initial segment. Will be similar. As a result, the waveforms of the two new segments can be accurately shifted by the waveform shift amount calculator so that the waveforms of the two new segments match, and an accurate shift amount is obtained. Since the second correlation delay amount can be accurately and quickly obtained by the second correlator using the shift amount, the adder adds the second correlation delay amount.

【0032】請求項8の発明では、前記第1メモリおよ
び前記第2メモリに入力するデータの先頭は前記セグメ
ントの先頭番地を示す。
In the present invention, the head of data input to the first memory and the second memory indicates a head address of the segment.

【0033】従来では図18で説明したように速度比α
が1以下ではポインタP1はセグメント長tsの整数倍
の位置にあり、相関遅延量rkが正の場合、セグメント
の中間にくる。一方ポインタP2は次のブロックの第2
セグメントの先頭位置にくるため、(c)に示すように
P1とP2の距離が離れると互いのセグメント間の類似
性が少なくなり、P1で始まるデータとP2で始まるデ
ータの加算値と、次に続くデータの連続性が悪くなって
音声として不自然な再生音となる。P1を次のブロック
の先頭セグメントの先頭番地、P2を次のセグメントの
先頭番地とすることにより、P1とP2間は常に先頭セ
グメントの長さとなり、離れすぎて類似性が低下するこ
とを防止できる。
Conventionally, as described with reference to FIG.
Is less than or equal to 1, the pointer P1 is located at a position that is an integral multiple of the segment length ts. On the other hand, the pointer P2 is the second block of the next block.
As shown in (c), when the distance between P1 and P2 increases, the similarity between the segments decreases, and the sum of data starting with P1 and data starting with P2, and The continuity of subsequent data deteriorates, resulting in an unnatural reproduction sound as voice. By setting P1 as the head address of the head segment of the next block and P2 as the head address of the next segment, the length of the head segment is always between P1 and P2, and it is possible to prevent the similarity from being reduced due to being too far apart. .

【0034】請求項9の発明では、前記第1相関遅延量
と前記第2相関遅延量とを積算する相関遅延量積算手段
と、この積算値がセグメント長にほぼ達したとき遅延フ
ラグを発生する遅延フラグ発生器とをさらに備え、前記
速度制御部は、この遅延フラグに基づきブロックのスル
ー区間のセグメントを追加または削除する。
According to a ninth aspect of the present invention, a correlation delay amount integrating means for integrating the first correlation delay amount and the second correlation delay amount, and a delay flag is generated when the integrated value substantially reaches the segment length. A delay flag generator, wherein the speed control unit adds or deletes a segment of a through section of the block based on the delay flag.

【0035】相関遅延量rkが正の場合、1ブロックの
長さは、rkが0で速度比αが予め定められた値となる
ブロック長さよりもrk分長くなる。(図18の(c)
の場合)各ブロックのrkを積算し、+1ディフォルト
セグメント長以上に達したとき、スルー区間のセグメン
トを1つ削除する。またrkが負の場合、速度比αが予
め定めた値となるブロックの長さよりもrk分短くな
る。(図18の(d)の場合)各ブロックのrkを積算
し、−1ディフォルトセグメント長以下になった時スル
ー区間のセグメントを1つ追加する。これにより1つ1
つのブロックでは速度比αを維持できなくても、ある範
囲のブロック数を考えれば速度比αを維持できる。
When the correlation delay amount rk is positive, the length of one block is longer by rk than the block length at which rk is 0 and the speed ratio α has a predetermined value. ((C) of FIG. 18)
In the case of (1), the rk of each block is integrated, and when the rk reaches the +1 default segment length or more, one segment in the through section is deleted. When rk is negative, the speed ratio α is shorter by rk than the length of the block where the speed ratio α has a predetermined value. (In the case of (d) in FIG. 18) The rk of each block is integrated, and when the length becomes equal to or smaller than the −1 default segment length, one segment of the through section is added. This is one by one
Even if one block cannot maintain the speed ratio α, the speed ratio α can be maintained considering the number of blocks in a certain range.

【0036】請求項10の発明では、前記速度制御部
は、前記ピッチ極性情報に基づき、現在のブロックの先
頭セグメントが次のセグメントと類似し、現在のブロッ
クの最後のセグメントが次のブロックの先頭セグメント
と類似している場合、現在のブロックの最後のセグメン
トを除去して現在のブロックのスルー区間を短縮する。
According to the tenth aspect of the present invention, based on the pitch polarity information, the speed control unit may determine that the first segment of the current block is similar to the next segment, and that the last segment of the current block is the first segment of the next block. If it is similar to the segment, the last segment of the current block is removed to shorten the through section of the current block.

【0037】現在のブロックの先頭セグメントと次のセ
グメントが類似しており、最後のセグメントが次のブロ
ックの先頭セグメントと類似している場合、最後のセグ
メントは次のブロックの先頭セグメントと加算した方が
自然な再生音声が得られる。このため最後のセグメント
は現在のブロックから除去する。これによりセグメント
は1つ減少し、ブロックは短くなる。
If the first segment of the current block is similar to the next segment and the last segment is similar to the first segment of the next block, the last segment is calculated by adding the first segment to the next block. However, a natural reproduced sound can be obtained. Therefore, the last segment is removed from the current block. This reduces the segment by one and shortens the block.

【0038】請求項11の発明では、前記速度制御部
は、前記ピッチ極性情報に基づき、現在のブロックの先
頭セグメントが1つ前のブロックの最後のセグメントと
類似しており、次のブロックの先頭セグメントがそのブ
ロックの次のセグメントを類似している場合、前のブロ
ックの最後のセグメントを現在のブロックに取り込んで
先頭のセグメントとし、現在のブロックの旧先頭セグメ
ントを第2セグメントとして加算し、現在のブロックの
スルー区間を1セグメント増加する。
According to the eleventh aspect of the present invention, based on the pitch polarity information, the speed control section determines that the first segment of the current block is similar to the last segment of the immediately preceding block, If the segment is similar to the next segment of the block, the last segment of the previous block is taken into the current block as the first segment, the old first segment of the current block is added as the second segment, and the current segment is added. Is increased by one segment.

【0039】現在のブロックの先頭セグメントが1つ前
のブロックの最後のセグメントと類似しており、次のブ
ロックの先頭セグメントが次のセグメントと類似してい
る場合は、1つ前のブロックの最後のセグメントを取り
込んで現在のブロックの先頭セグメントと加算すること
により、自然な再生音声が得られる。これにより現在の
ブロックのセグメントは1つ増加しブロックは長くな
る。
If the first segment of the current block is similar to the last segment of the previous block, and the first segment of the next block is similar to the next segment, the last segment of the previous block is Is obtained and added to the head segment of the current block, a natural reproduced sound can be obtained. This increases the segment of the current block by one and makes the block longer.

【0040】請求項12の発明では、前記速度制御部
は、前記ピッチ極性情報に基づき、現在のブロックの先
頭セグメントが1つ前のブロックの最後のセグメントと
類似し、現在のブロックの最後のセグメントが次のブロ
ックの先頭セグメントと類似する場合には、現在のブロ
ックは1つ前のブロックのセグメントを先頭セグメント
とし、現在のブロックの旧先頭セグメントを第2セグメ
ントとして加算し、最後のセグメントを次のブロックの
先頭セグメントとして加算する。
In the twelfth aspect of the present invention, the speed control unit may determine that the first segment of the current block is similar to the last segment of the immediately preceding block based on the pitch polarity information, and the last segment of the current block. Is similar to the first segment of the next block, the current block is set to the segment of the previous block as the first segment, the old first segment of the current block is added as the second segment, and the last segment is set to the next segment. Is added as the first segment of the block.

【0041】現在のブロックが1つ前のブロックの最後
のセグメントに類似し、現在のブロックの最後のセグメ
ントが次のブロックの先頭セグメントと類似する場合
は、1つ前のブロックの最後のセグメントと現在のブロ
ックの先頭セグメントで加算し、現在のブロックの最後
のセグメントは次のブロックの先頭セグメントと加算す
る。これにより自然な再生音が得られる。
If the current block is similar to the last segment of the previous block and the last segment of the current block is similar to the first segment of the next block, the last segment of the previous block is It adds at the first segment of the current block, and adds the last segment of the current block to the first segment of the next block. As a result, a natural reproduced sound can be obtained.

【0042】請求項13の発明では、符号化した音声デ
ータを復号化する復号器と、この復号器より得られる音
声データのピッチ情報と1つのピッチ内のデータが隣接
する時間的に後のピッチのデータに類似するか前のピッ
チのデータに類似するかを示すピッチ極性情報とを格納
するピッチ格納部と、前記復号器からのデータを所定長
さのセグメントに分割し連続する所定数のセグメントを
1つのブロックとし、この各ブロックのセグメントを入
力した順に分けて出力する分割器と、この分割器が出力
する各ブロックの最初のセグメントを格納する第1メモ
リと、前記分割器が出力する各ブロックの2番目のセグ
メントを格納する第2メモリと、前記ピッチ格納部のデ
ータより前記最初のセグメントの波形と前記次のセグメ
ントの波形とがほぼ一致するようにずらして重ねそのず
れ量を算出する波形ずれ量算出器と、前記第1メモリの
内容と前記第2メモリの内容との相関関数を前記波形ず
れ量算出器の出力に基づいて算出し相関遅延量を演算す
る相関器と、セグメントのデータに対する重み係数を発
生する窓関数発生器と、前記第1メモリの内容に前記重
み係数を乗算する第1乗算器と、前記第2メモリの内容
に前記重み係数を乗算する第2乗算器と、前記相関器の
出力に基づき前記第1乗算器の出力と前記第2乗算器の
出力とを相関関数の値が大きい位置で加算する加算器
と、前記加算器の出力とこれに続けて、前記分割器より
入力する低速化のときは第2セグメント以降、高速化の
時は第3セグメント以降のセグメントで構成されるスル
ー区間セグメントを出力する結合器と、音声の速度の変
換比を表す速度比に基づき前記分割器と前記結合器を制
御する速度制御部と、を備える。
According to the thirteenth aspect of the present invention, a decoder for decoding coded audio data, a pitch information of the audio data obtained from the decoder, and a pitch within a pitch which is adjacent to the temporally succeeding pitch information. A pitch storage unit for storing pitch polarity information indicating whether the data is similar to the data of the previous pitch or similar to the data of the previous pitch; and a predetermined number of continuous segments obtained by dividing the data from the decoder into segments of a predetermined length. As one block, a divider for dividing and outputting the segments of each block in the order of input, a first memory for storing the first segment of each block output by the divider, and a The second memory for storing the second segment of the block, and the waveform of the first segment and the waveform of the next segment are approximately determined from the data of the pitch storage unit. A waveform shift amount calculator for calculating the shift amount by shifting so as to coincide with each other, and calculating a correlation function between the content of the first memory and the content of the second memory based on the output of the waveform shift amount calculator. A correlator for calculating the amount of correlation delay, a window function generator for generating a weighting factor for the data of the segment, a first multiplier for multiplying the content of the first memory by the weighting factor, A second multiplier for multiplying the content by the weighting coefficient, and an adder for adding the output of the first multiplier and the output of the second multiplier at a position where the value of the correlation function is large based on the output of the correlator Then, an output of the adder and, subsequently, a through section segment composed of the second and subsequent segments when the speed is reduced, and the third and subsequent segments when the speed is increased, input from the divider are output. With the coupler Comprising a speed control unit for controlling the divider and the combiner based on the speed ratio representing the conversion ratio of the sound speed, the.

【0043】符号化したデータを復号化するがその際、
音声データのピッチ情報、ピッチ極性情報が得られる。
音声データをセグメントに分割し、所定数のセグメント
を1ブロックとし、このブロック単位で音声速度変換を
行う。速度変換は変換後の再生時間長を変換前の再生時
間長で割った速度比αに基づいて行われ、1ブロックの
セグメント数はこの速度比αに応じて予め定められる。
速度制御部は分割器を制御してブロックの最初のセグメ
ントを第1メモリへ出力し、2番目のセグメントを第2
メモリへ出力し、低速化または高速化に応じて第2セグ
メント以降または第3セグメント以降を結合器へ出力す
る。
When the encoded data is decoded,
Pitch information and pitch polarity information of audio data can be obtained.
The audio data is divided into segments, and a predetermined number of segments are made into one block, and audio speed conversion is performed in units of this block. The speed conversion is performed based on the speed ratio α obtained by dividing the converted playback time length by the playback time length before conversion, and the number of segments in one block is predetermined according to the speed ratio α.
The speed controller controls the divider to output the first segment of the block to the first memory, and the second segment to the second memory.
The data is output to the memory, and the second and subsequent segments or the third and subsequent segments are output to the coupler according to the speed reduction or the speed increase.

【0044】波形ずれ量算出器では、ピッチ情報に基づ
いて最初のセグメントの波形と2番目のセグメントの波
形がほぼ一致するようにずらして重ねそのずれ量を算出
する。両セグメントの波形と位相が同じであれば、ずれ
量は0となる。このように両セグメントの波形をほぼ一
致するようにずらした状態で、相関器は第1メモリの最
初のセグメントと第2メモリの2番目のセグメントの相
関関数を計算することにより相関の高い位置を求めるた
めの探索範囲が大幅に狭くなり、計算量が大幅に減少す
る。これにより相関遅延量の演算を容易に行うことがで
きる。
The waveform shift amount calculator calculates the shift amount by shifting the waveform of the first segment and the waveform of the second segment so as to substantially coincide with each other based on the pitch information. If the waveform and the phase of both segments are the same, the shift amount is zero. With the waveforms of both segments shifted so as to substantially coincide with each other, the correlator calculates the correlation function between the first segment of the first memory and the second segment of the second memory to find a position having a high correlation. The search range for finding is greatly narrowed, and the amount of calculation is greatly reduced. Thereby, the calculation of the correlation delay amount can be easily performed.

【0045】窓関数発生器は最初のセグメントと2番目
のセグメントのデータに対してそれぞれ重み係数を予め
定められたパターンにより発生する。重み係数は、両セ
グメントを加算して新たな加算セグメントとした場合、
この加算セグメントのデータの大きさが元のセグメント
のデータと同じ大きさになるようにし、加算によって音
声が不自然にならないようにするため設けられている。
このため両セグメントの加算するデータの一方の重み係
数と他方の重み係数との和は1.0となる。また一方の
セグメントの重み係数はセグメントの先頭から後端にゆ
くに従い漸増し、他方のセグメントの重み係数はセグメ
ントの先頭から後端にゆくに従い漸減するようにしてい
る。
The window function generator generates a weighting coefficient for the data of the first segment and the data of the second segment according to a predetermined pattern. The weighting factor is calculated by adding both segments to create a new added segment.
It is provided so that the size of the data of the added segment is the same as that of the data of the original segment, and the addition does not make the sound unnatural.
Therefore, the sum of one weight coefficient and the other weight coefficient of the data to be added in both segments is 1.0. The weight coefficient of one segment increases gradually from the head of the segment to the rear end, and the weight coefficient of the other segment decreases gradually from the head of the segment to the rear end.

【0046】第1乗算器は第1メモリのデータと重み係
数を乗算し、第2乗算器は第2メモリのデータと重み係
数を乗算する。加算器は第1および第2乗算器の出力を
相関遅延量に基づき、相関関数の大きな位置に一方をず
らして加算を行う。結合器はこの加算によって得られた
加算セグメントの後に分割器から入力したスルー区間セ
グメントを続けて1ブロックのデータとして出力する。
The first multiplier multiplies the data in the first memory by the weighting factor, and the second multiplier multiplies the data in the second memory by the weighting factor. The adder adds the outputs of the first and second multipliers by shifting one of them to a position having a large correlation function based on the correlation delay amount. The combiner continuously outputs the through segment input from the divider after the addition segment obtained by the addition as one block of data.

【0047】請求項14の発明では、音声データを入力
して所定長さのセグメントに分割し連続する所定数のセ
グメントを1つのブロックとし、この各ブロックのセグ
メントを入力した順に分けて出力する分割器と、この分
割器が出力する各ブロックの最初のセグメントを格納す
る第1メモリと、前記分割器が出力する各ブロックの2
番目のセグメントを格納する第2メモリと、前記第1メ
モリの内容と前記第2メモリの内容との相関関数を算出
し相関遅延量を演算する相関器と、セグメントのデータ
に対する重み係数を発生する窓関数発生器と、前記第1
メモリの内容に前記重み係数を乗算する第1乗算器と、
前記第2メモリの内容に前記重み係数を乗算する第2乗
算器と、前記相関器の出力に基づき前記第1乗算器の出
力と前記第2乗算器の出力とを相関関数の値が大きい位
置で加算する加算器と、前記加算器の出力とこれに続け
て、前記分割器より入力する低速化のときは第2セグメ
ント以降、高速化の時は第3セグメント以降のセグメン
トで構成されるスルー区間セグメントを出力する結合器
と、音声の速度の変換比を表す速度比に基づき前記分割
器と前記結合器を制御する速度制御部と、を備え、前記
速度制御部は前記分割部を制御して前記セグメントの長
さを可変にする。
According to a fourteenth aspect of the present invention, the audio data is input, divided into segments of a predetermined length, and a predetermined number of continuous segments are formed into one block, and the segments of each block are divided and output in the order of input. A first memory for storing the first segment of each block output by the divider; and a second memory for each block output by the divider.
A second memory for storing a third segment; a correlator for calculating a correlation function between the contents of the first memory and the contents of the second memory to calculate a correlation delay amount; and generating a weight coefficient for the data of the segment. A window function generator;
A first multiplier for multiplying the content of the memory by the weight coefficient;
A second multiplier for multiplying the content of the second memory by the weighting coefficient, and a position where the value of the correlation function is large between the output of the first multiplier and the output of the second multiplier based on the output of the correlator. , And an output of the adder followed by a through segment composed of the second and subsequent segments input from the divider when the speed is low, and the third and subsequent segments when the speed is high. A combiner that outputs a section segment, and a speed control unit that controls the divider and the combiner based on a speed ratio that represents a conversion ratio of a voice speed, and the speed control unit controls the divider. To make the length of the segment variable.

【0048】入力した音声データをセグメントに分割
し、所定数のセグメントを1ブロックとし、このブロッ
ク単位で音声速度変換を行う。速度変換は変換後の再生
時間長を変換前の再生時間長で割った速度比αに基づい
て行われ、1ブロックのセグメント数はこの速度比αに
応じて予め定められる。速度制御部は分割器を制御して
ブロックの最初のセグメントを第1メモリへ出力し、2
番目のセグメントを第2メモリへ出力し、低速化または
高速化に応じて第2セグメント以降または第3セグメン
ト以降を結合器へ出力する。相関器は最初のセグメント
と2番目のセグメントとの相関関数を算出し、両セグメ
ントの相関の高い位置間の距離を示す相関遅延量を演算
する。
The input audio data is divided into segments, and a predetermined number of segments are made into one block, and audio speed conversion is performed for each block. The speed conversion is performed based on the speed ratio α obtained by dividing the converted playback time length by the playback time length before conversion, and the number of segments in one block is predetermined according to the speed ratio α. The speed controller controls the divider to output the first segment of the block to the first memory,
The second segment is output to the second memory, and the second and subsequent segments or the third and subsequent segments are output to the coupler according to the speed reduction or the speed increase. The correlator calculates a correlation function between the first segment and the second segment, and calculates a correlation delay amount indicating a distance between highly correlated positions of both segments.

【0049】窓関数発生器は最初のセグメントと2番目
のセグメントのデータに対してそれぞれ重み係数を予め
定められたパターンにより発生する。重み係数は、両セ
グメントを加算して新たな加算セグメントとした場合、
この加算セグメントのデータの大きさが元のセグメント
のデータと同じ大きさになるようにし、加算によって音
声が不自然にならないようにするため設けられている。
このため両セグメントの加算するデータの一方の重み係
数と他方の重み係数との和は1.0となる。また一方の
セグメントの重み係数はセグメントの先頭から後端にゆ
くに従い漸増し、他方のセグメントの重み係数はセグメ
ントの先頭から後端にゆくに従い漸減するようにしてい
る。
The window function generator generates a weight coefficient for each of the data of the first segment and the data of the second segment according to a predetermined pattern. The weighting factor is calculated by adding both segments to create a new added segment.
It is provided so that the size of the data of the added segment is the same as that of the data of the original segment, and the addition does not make the sound unnatural.
Therefore, the sum of one weight coefficient and the other weight coefficient of the data to be added in both segments is 1.0. The weight coefficient of one segment increases gradually from the head of the segment to the rear end, and the weight coefficient of the other segment decreases gradually from the head of the segment to the rear end.

【0050】第1乗算器は第1メモリのデータと重み係
数を乗算し、第2乗算器は第2メモリのデータと重み係
数を乗算する。加算器は第1および第2乗算器の出力を
相関遅延量に基づき、相関関数の大きな位置に一方をず
らして加算を行う。結合器はこの加算によって得られた
加算セグメントの後に分割器から入力したスルー区間セ
グメントを続けて1ブロックのデータとして出力する。
The first multiplier multiplies the data of the first memory by the weighting factor, and the second multiplier multiplies the data of the second memory by the weighting factor. The adder adds the outputs of the first and second multipliers by shifting one of them to a position having a large correlation function based on the correlation delay amount. The combiner continuously outputs the through segment input from the divider after the addition segment obtained by the addition as one block of data.

【0051】速度制御部は分割器を制御してセグメント
の長さを可変とする。これにより男性や女性のように音
声の異なるデータを適切に再生することができる。
The speed controller controls the divider to change the length of the segment. This makes it possible to appropriately reproduce data with different voices, such as a man or a woman.

【0052】以下、本発明の実施の形態について図面を
参照して説明する。図1は本発明の各実施形態を実現す
る音声速度変換装置のブロック図である。A/D変換器
31は入力信号(音声データ)をアナログ信号からデジ
タル信号に変換する。音声符号化/復号器32は、本実
施の形態の場合、符号化された入力信号を復号化する。
音声符号化/復号器32により得られた波形はバッファ
33に書き込まれ、同時に音声符号化/復号器32によ
り得られるピッチ情報及びピッチの極性情報はピッチバ
ッファ(ピッチ格納部)34に、音声データのピッチが
明確に表れている場合有声、不明確な場合無声とし、こ
れをフラグで表示する有声/無声判別フラグ情報はフラ
グバッファ(フラグ格納部)35にそれぞれ書き込まれ
る。
Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 is a block diagram of an audio speed conversion device that realizes each embodiment of the present invention. The A / D converter 31 converts an input signal (audio data) from an analog signal to a digital signal. In the case of the present embodiment, the audio encoder / decoder 32 decodes the encoded input signal.
The waveform obtained by the audio encoder / decoder 32 is written into the buffer 33, and the pitch information and the pitch polarity information obtained by the audio encoder / decoder 32 are simultaneously stored in a pitch buffer (pitch storage unit) 34. Is voiced when the pitch is clearly expressed, and unvoiced when the pitch is unclear, and voiced / unvoiced discrimination flag information, which is indicated by a flag, is written to the flag buffer (flag storage unit) 35.

【0053】速度制御回路36には予め速度比αが設定
されており、また初期セグメント長さ、およびこの2倍
の2倍セグメント長が設定されている。速度変換を表す
図18で説明したように、速度比αに応じて1ブロック
のセグメント数を決め、先頭セグメントと次のセグメン
トに重み係数を乗じて加算して1つのセグメントとし、
高速にする場合は、この加算セグメントに第3セグメン
ト以降のセグメント(これをスルー区間セグメントと称
する)を加えて、1セグメント少ないブロックとする。
低速にする場合は加算セグメントの後に第2セグメント
以降のセグメントを加えて1セグメント多いブロックと
する。1ブロックのセグメント数を変更することにより
任意の速度比αにすることができる。
In the speed control circuit 36, a speed ratio α is set in advance, and an initial segment length and a double segment length twice as long as the initial segment length are set. As described with reference to FIG. 18 showing the speed conversion, the number of segments of one block is determined according to the speed ratio α, and the first segment and the next segment are multiplied by a weighting factor and added to form one segment.
To increase the speed, the third segment and subsequent segments (hereinafter referred to as “through-segment segments”) are added to this added segment to make the block one segment less.
When the speed is to be reduced, the second and subsequent segments are added after the addition segment to make the block one segment more. An arbitrary speed ratio α can be obtained by changing the number of segments in one block.

【0054】速度制御回路36はピッチバッファ34か
らピッチ情報およびピッチ極性情報を入力し、分割器3
7を制御し、また後述する結合器49を制御する。分割
器37はデータバッファ33から音声データを読み出し
初期セグメントの大きさに分割し、1ブロック内のセグ
メント数を定め、先頭セグメントを第1メモリ38に出
力し、次のセグメントを第2メモリ39に出力し、低速
化のときは第2セグメント以降のセグメントと次のブロ
ックの先頭セグメント、高速化の時は第3セグメント以
降のブロック内セグメント(これらのセグメントをスル
ー区間セグメントと称する)を結合器49に出力する。
The speed control circuit 36 receives pitch information and pitch polarity information from the pitch buffer 34 and
7 and a coupler 49 described later. The divider 37 reads out the audio data from the data buffer 33, divides it into the size of the initial segment, determines the number of segments in one block, outputs the first segment to the first memory 38, and stores the next segment in the second memory 39. When the speed is reduced, the segments after the second segment and the first segment of the next block are combined, and when the speed is increased, the segments in the block after the third segment (these segments are referred to as through-segment segments) are combined. Output to

【0055】第1相関器41は第1メモリ38と第2メ
モリ39のデータの相関関数を算出し、2つのセグメン
トの相関関数の大きな位置が一致するように相対位置を
ずらして得られる第1相関遅延量rkを求める。一方波
形ずれ量(rk)算出器42はピッチバッファ34より
ピッチ情報を得て、先頭セグメントと次のセグメントの
波形が同じピッチを有し、位相のみ異なると仮定して、
重ねて一致するようにずらしたずらし量(つまり位相
差)を求める。このずらし量は上記のような仮定により
容易に求められる。
The first correlator 41 calculates the correlation function between the data in the first memory 38 and the data in the second memory 39, and shifts the relative position so that the large positions of the correlation functions of the two segments coincide with each other. The correlation delay rk is obtained. On the other hand, the waveform shift amount (rk) calculator 42 obtains the pitch information from the pitch buffer 34, and assumes that the waveforms of the first segment and the next segment have the same pitch and differ only in the phase.
A shift amount (that is, a phase difference) shifted so as to coincide with each other is obtained. This shift amount can be easily obtained based on the above assumption.

【0056】第2相関器43は第1メモリ38と第2メ
モリ39のデータと、波形ずれ量(rk)算出器42か
らのずれ量に基づいて相関関数を算出し、2つのセグメ
ントの相関関数の大きな位置が一致するように相対位置
をずらして得られる第2相関遅延量rkを求める。ずれ
量がわかっているので第2相関器43が相関関数を計算
するときのデータの探索範囲が第1相関器41の場合に
比べて大幅に少なくなり短時間に正確な相関関数が得ら
れる。
The second correlator 43 calculates a correlation function based on the data in the first memory 38 and the second memory 39 and the shift amount from the waveform shift (rk) calculator 42, and calculates the correlation function of the two segments. The second correlation delay amount rk obtained by shifting the relative position so that the position having the larger value of. Since the shift amount is known, the data search range when the second correlator 43 calculates the correlation function is significantly smaller than that in the case of the first correlator 41, and an accurate correlation function can be obtained in a short time.

【0057】窓関数発生器45は第1メモリ38と第2
メモリ39のデータに対する重み係数をそれぞれ別々に
出力する。この重み係数は速度比αや第1相関器41、
第2相関器43の発生する相関関数に基づいてセグメン
トの先頭から後端のデータにゆくに従い漸増または漸減
する窓関数を発生する。この窓関数は、第1メモリ38
と第2メモリ39の対応する各データに対してそれぞれ
の重み係数の和が1.0になるように設定され、加算さ
れた先頭および次のセグメントの加算値による音声の高
さが加算前と変わらないようにしている。
The window function generator 45 includes a first memory 38 and a second
The weight coefficients for the data in the memory 39 are output separately. This weighting factor is determined by the speed ratio α, the first correlator 41,
On the basis of the correlation function generated by the second correlator 43, a window function that gradually increases or decreases as the data goes from the head to the tail of the segment is generated. This window function is stored in the first memory 38.
And the sum of the respective weighting factors for the corresponding data in the second memory 39 is set so as to be 1.0. I try not to change.

【0058】第1乗算器47は第1メモリ38のデータ
と窓関数発生器45からの重み係数を乗算し、第2乗算
器48は第2メモリ39と窓関数発生器45からの重み
係数を乗算し、加算器44に出力する。切替スイッチ4
0は、ピッチバッファ34のピッチ情報、ピッチ極性情
報およびフラグバッファ35のフラグにより第1相関器
41の第1相関遅延量rkまたは第2相関器43の第2
相関遅延量rkを加算器44に出力する。
The first multiplier 47 multiplies the data in the first memory 38 by the weight coefficient from the window function generator 45, and the second multiplier 48 calculates the weight coefficient from the second memory 39 and the window function generator 45. The result is multiplied and output to the adder 44. Changeover switch 4
0 is the first correlation delay rk of the first correlator 41 or the second correlation delay rk of the second correlator 43 according to the pitch information of the pitch buffer 34, the pitch polarity information and the flag of the flag buffer 35.
The correlation delay amount rk is output to the adder 44.

【0059】加算器44は、第1乗算器47の出力と第
2乗算器48の出力を、第1相関遅延量rkまたは第2
相関遅延量rkに基づいて相関関数の大きな値に一方を
ずらして加算して加算セグメントを生成し、結合器49
に出力する。結合器49は速度制御回路36からの制御
により、加算器44からの加算セグメントの後に分割器
37から送られてくる、低速化時には第2セグメント以
降のセグメントと、次のブロックの先頭セグメントを1
ブロックのデータとして出力し、高速化時には第3セグ
メント以降のブロック内のセグメントを1ブロックとし
て出力する。D/A変換器50は結合器49からのデー
タをデジタル信号からアナログ信号に変換して出力す
る。これにより音声再生が可能となる。
The adder 44 compares the output of the first multiplier 47 and the output of the second multiplier 48 with the first correlation delay rk or the second
Based on the correlation delay amount rk, one of the values is shifted and added to a large value of the correlation function to generate an added segment.
Output to Under the control of the speed control circuit 36, the combiner 49 sends the segment after the second segment and the first segment of the next block sent from the divider 37 after the addition segment from the adder 44 at the time of speed reduction.
The data is output as block data, and at the time of high-speed operation, the segments in the block after the third segment are output as one block. The D / A converter 50 converts the data from the combiner 49 from a digital signal to an analog signal and outputs it. This enables sound reproduction.

【0060】加算器46は第1相関器41および第2相
関器43が発生する第1相関遅延量rkおよび第2相関
遅延量rkを発生順に積算し、rkバッファ51に格納
する。rkバッファ51の積算値が+1ディフォルトセ
グメントの長さに達すると、又は−1ディフォルトセグ
メントの長さに満たなくなったとき、遅延フラグ発生器
52からフラグが立てられ、このフラグ情報は速度制御
回路36に出力される。
The adder 46 integrates the first correlation delay rk and the second correlation delay rk generated by the first correlator 41 and the second correlator 43 in the order of generation, and stores the result in the rk buffer 51. When the integrated value of the rk buffer 51 reaches the length of the +1 default segment or becomes shorter than the length of the -1 default segment, a flag is set from the delay flag generator 52, and this flag information is stored in the speed control circuit 36. Is output to

【0061】速度制御回路36はピッチ情報によりピッ
チが初期セグメント長より大きくなった場合、初期セグ
メントの2倍の2倍長セグメントに切り替え、1ブロッ
クの長さも2倍にする。セグメント長シフトレジスタ5
3はセグメント長が切り替わったことを記憶し、この切
り替え情報により切替スイッチ40は第1相関器41と
第2相関器43の出力を切り替える。
When the pitch becomes larger than the initial segment length due to the pitch information, the speed control circuit 36 switches to a double-length segment which is twice the initial segment, and doubles the length of one block. Segment length shift register 5
Reference numeral 3 indicates that the segment length has been switched, and the switch 40 switches the output of the first correlator 41 and the output of the second correlator 43 based on this switching information.

【0062】図2は音声データとセグメントを説明する
図である。音声データは(A)に示すように1つの大き
な振幅と小さな複数の振幅のデータが1周期となり、こ
れが繰り返されるものが多い。大きい振幅間の長さをピ
ッチLとする。このピッチLは例えば10ms程度であ
る。女性の声の統計的な近似値としてL=7.5msを
ディフォルト長として設定すれば各セグメントに1つの
大きな振幅が入るようになる。(B)はこのようにして
設定したセグメントを示す。このセグメントを初期(デ
フォルト)セグメントと称する。
FIG. 2 is a diagram for explaining audio data and segments. As shown in (A), the audio data has one cycle of data having one large amplitude and a plurality of small amplitudes, and is often repeated. The length between large amplitudes is referred to as a pitch L. This pitch L is, for example, about 10 ms. If L = 7.5 ms is set as the default length as a statistical approximation of the female voice, one large amplitude will enter each segment. (B) shows the segment set in this way. This segment is called an initial (default) segment.

【0063】これに対して男性の声などはピッチが長
く、L=7.5msを越える場合が多い。(C)はこの
ような音声データを示す。ピッチL1は(B)のセグメ
ントより長いので、1つのセグメント内に大きな振幅が
現れないものも生じる。そこで(D)に示すようにセグ
メント長の2倍の2倍長セグメントを採用する。1ブロ
ック内のセグメント数を一定とすると1ブロック長も2
倍となる。
On the other hand, the pitch of a male voice or the like is long and often exceeds L = 7.5 ms. (C) shows such audio data. Since the pitch L1 is longer than the segment of (B), there may be a case where a large amplitude does not appear in one segment. Therefore, as shown in (D), a double length segment which is twice the segment length is adopted. If the number of segments in one block is fixed, the length of one block is also 2
Double.

【0064】(E)はピッチ極性情報を説明する図であ
る。隣接する2つのセグメントのデータが類似する場
合、右側(時間の進行方向)のセグメントに類似する
か、左側のセグメントに類似するかを表したものをピッ
チ極性情報と言う。図では第2セグメントと第3セグメ
ントが類似している場合を示す。第2セグメントを中心
にし、右側のセグメントに類似している場合をL(2)
>0で表す。また第3セグメントを中心にし、左側のセ
グメントに類似する場合はL(3)<0と表す。なお、
(A)に示すように両側のセグメントに類似している場
合は、L(2)>0,L(2)<0で表す。
(E) is a diagram for explaining pitch polarity information. When the data of two adjacent segments are similar, whether the data is similar to the segment on the right side (time progress direction) or similar to the segment on the left side is called pitch polarity information. The figure shows a case where the second segment and the third segment are similar. L (2) when the second segment is centered and similar to the right segment
> 0. L (3) <0 when the third segment is the center and is similar to the left segment. In addition,
If the segments are similar on both sides as shown in (A), they are represented by L (2)> 0 and L (2) <0.

【0065】図3は相関遅延量rkを説明する図であ
る。(A)は3つのセグメントのそれぞれの波形a,
b,cを示す。tsは初期セグメント長を示す。(B)
は波形bに対して波形aを最大振幅が一致するように重
ねた状態を示す。この場合波形aを右側にrkずらすと
両セグメントの波形a,bはほぼ一致する。このときの
ずらし量rkをrk+とする。(C)は波形bに対して
波形cを左側にrkずらして最大振幅が一致するように
重ねた状態を示す。このときずらし量をrk−とする。
このrk+,rk−を相関遅延量rkとする。
FIG. 3 is a diagram for explaining the correlation delay amount rk. (A) shows the waveforms a,
b and c are shown. ts indicates the initial segment length. (B)
Indicates a state in which the waveform a is superimposed on the waveform b so that the maximum amplitude coincides. In this case, when the waveform a is shifted by rk to the right, the waveforms a and b of both segments substantially coincide. The shift amount rk at this time is rk +. (C) shows a state where the waveform c is shifted rk to the left with respect to the waveform b and overlapped so that the maximum amplitudes match. At this time, the shift amount is rk-.
These rk +, rk- are set as the correlation delay amount rk.

【0066】次に第1実施形態を説明する。図4は速度
比α=3/4の場合のセグメント可変処理を示す。本処
理においては相関遅延量rkを簡単のため0とした標準
形で示す。(a)はデータバッファ33の音声データを
セグメントに分割し、セグメントxに同期したピッチ長
さ情報L(x)をピッチバッファ34から得る。1ブロ
ックは4セグメントとし、ピッチ情報L(x)は、各ブ
ロックの先頭セグメントのものを示している。
Next, a first embodiment will be described. FIG. 4 shows the segment variable processing when the speed ratio α = 3/4. In this processing, the correlation delay amount rk is shown in a standard form with 0 for simplicity. 3A, the audio data in the data buffer 33 is divided into segments, and pitch length information L (x) synchronized with the segment x is obtained from the pitch buffer. One block is composed of four segments, and the pitch information L (x) indicates the first segment of each block.

【0067】(b)は(a)に示したデータに基づい
て、L(1)<デフォルトセグメント(初期セグメン
ト)が成立するときの分割器37の制御手順を示してい
る。この場合セグメントの長さは初期セグメント長と
し、(a)に示したポインタP1(a)を先頭アドレス
として、セグメントの長さ分をそれぞれ第1メモリ3
8、第2メモリ39に入力し、それに続く第3セグメン
ト、第4セグメントをそのまま結合器49に出力する。
第1メモリ38、第2メモリ39のデータは窓関数によ
る重み付け加算され、結合器49に出力される。結合器
49の出力は図のようになる。
(B) shows a control procedure of the divider 37 when L (1) <default segment (initial segment) is established based on the data shown in (a). In this case, the length of the segment is the initial segment length, and the pointer P1 (a) shown in FIG.
8. The data is input to the second memory 39, and the subsequent third and fourth segments are output to the combiner 49 as they are.
The data in the first memory 38 and the data in the second memory 39 are weighted and added by the window function and output to the combiner 49. The output of the combiner 49 is as shown in the figure.

【0068】(c)はピッチ長さL(5)が初期セグメ
ント長より大きい場合で、セグメント長を初期セグメン
ト長×2とした場合の分割器37の制御手段を示してい
る。(b)に示したポインタP1(b),P2(b)を
先頭アドレスとしてセグメント長分をそれぞれ第1メモ
リ38と第2メモリ39に出力する。これに続くデータ
は、元のセグメント8と9とが第3の2倍長セグメント
に入り、元のセグメント10,11が第4の2倍長セグ
メントに入り、結合器49に出力される。第1メモリ3
8、第2メモリ39のデータは窓関数による重み付け加
算され、結合器49に出力され、結合器49からは図に
示すようなデータ列が出力される。
(C) shows the control means of the divider 37 when the pitch length L (5) is larger than the initial segment length, and when the segment length is equal to the initial segment length × 2. With the pointers P1 (b) and P2 (b) shown in (b) as the start address, the segment length is output to the first memory 38 and the second memory 39, respectively. In the subsequent data, the original segments 8 and 9 enter the third double-length segment, and the original segments 10 and 11 enter the fourth double-length segment and are output to the combiner 49. First memory 3
8. The data in the second memory 39 is weighted and added by the window function and output to the combiner 49, which outputs a data string as shown in the figure.

【0069】(d)は(c)の場合と同じで、ピッチL
(13)が初期セグメント長より大きいため2倍長セグ
メントが用いられている。(e)はピッチL(13)が
初期セグメント長より短く、(b)と同じ場合を示す。
このように2倍長セグメントから初期セグメントに切り
替える場合は、ポインタP2(c)は2倍長セグメント
が続くとして定められているので、加算するセグメント
は12と14になり、セグメント13が欠落するので再
生音が多少悪くなる。
(D) is the same as (c), and the pitch L
Since (13) is larger than the initial segment length, a double-length segment is used. (E) shows a case where the pitch L (13) is shorter than the initial segment length and is the same as (b).
When switching from the double-length segment to the initial segment in this way, since the pointer P2 (c) is determined to be followed by the double-length segment, the segments to be added are 12 and 14, and the segment 13 is missing. Playback sound is slightly worse.

【0070】図5は速度比α=3/2のときのセグメン
ト長を可変とした処理を示す。相関遅延量rkは簡単の
ため0とした場合を示す。(a)〜(e)共図4の
(a)〜(e)とほぼ同様である。相違点は窓関数の重
み付けを示す加算セグメントの斜め線が図4の高速化の
場合と逆になっている。これは加算セグメントの後に第
2セグメントが続くが、加算セグメントの後端側が第1
セグメントに近いと、次の第2セグメントによく接続
し、再生音が自然になるためである。以上のようにセグ
メント長をピッチLに合わせて2段に切り替えることに
より、女性、男性音を正確に再生でき、汎用性のある処
理を行うことができる。
FIG. 5 shows a process in which the segment length is variable when the speed ratio α = 3/2. The case where the correlation delay amount rk is set to 0 for simplicity. (A) to (e) are almost the same as (a) to (e) in FIG. The difference is that the diagonal line of the addition segment indicating the weight of the window function is opposite to that in the case of speeding up in FIG. This is because the addition segment is followed by the second segment, but the end of the addition segment is the first
This is because, when the segment is close to the segment, the segment is well connected to the next second segment, and the reproduced sound becomes natural. By switching the segment length to two steps according to the pitch L as described above, female and male sounds can be accurately reproduced, and versatile processing can be performed.

【0071】図6は波形ずれ量(rk)算出器42と第
2相関器43の詳細処理を示す。(A)は第1セグメン
トと第2セグメント内の音声データを示し、セグメント
内の縦線は図2で説明した大きな振幅のデータを示し、
Lはこの振幅のピッチを示す。aは第1セグメントの先
頭から最初の大きな振幅までの長さを示し、bは第2セ
グメントの先頭から最初の大きな振幅の長さを示す。第
2相関器43を用いる場合は、ピッチLが第1および第
2セグメント共ほぼ同じ程度と仮定している。
FIG. 6 shows detailed processing of the waveform shift amount (rk) calculator 42 and the second correlator 43. (A) shows the audio data in the first segment and the second segment, and the vertical line in the segment shows the large amplitude data described in FIG.
L indicates the pitch of this amplitude. a indicates the length from the beginning of the first segment to the first large amplitude, and b indicates the length of the first large amplitude from the beginning of the second segment. When the second correlator 43 is used, it is assumed that the pitch L is substantially the same for both the first and second segments.

【0072】(B)は第1セグメントと第2セグメント
とを各セグメントの先頭を合わせて重ねた状態である。
rk−は第2セグメントを左側(時間軸の進行と反対方
向)に移動して大きな振幅の位置を一致させた場合の移
動量を示し、rk+は右側(時間軸の進行方向)に移動
して大きな振幅の位置を一致させた場合の移動量を示
す。(C)の左側の図は第1セグメントに対して第2セ
グメントをrk−移動させて波形をほぼ一致させた状態
で第2相関器43による相関関数の計算を行うことを示
す。(A)で仮定したようにピッチLが両セグメントで
全て同一であれば、rk−又はrk+が相関遅延量rk
と一致するが実際は両セグメント内でピッチLは異なる
場合が多い。しかし、ピッチLの差はあまり大きくない
ので、両セグメントをrk−ずらした状態で相関関数を
計算する場合、探索範囲を少なくしてよい。例えば、初
期セグメント長が7.5msに設定された場合、8kH
zサンプリング離散データのとき、ディフォルトセグメ
ントデータが60サンプルあり、通常、相関関数を求め
るには40サンプルぐらいのデータについて乗算と加算
を行うが、rk−ずらした後では5〜6サンプルのデー
タについて行えば相関関数が求められる。(C)の右側
の図は第1セグメントに対して第2セグメントをrk+
移動させて波形をほぼ一致させた状態で第2相関器43
で相関関数の計算を行う場合を示す。左右両方の計算か
らrkを決定することにより、精度のよい相関遅延量r
kが得られる。これにより第1相関器41で計算する場
合に比べ演算量を大幅に削減できる。
(B) shows a state in which the first segment and the second segment are overlapped with the beginning of each segment.
rk- indicates the amount of movement when the second segment is moved to the left (in the direction opposite to the progress of the time axis) to match the position of large amplitude, and rk + is to the right (in the direction of progress of the time axis). This shows the amount of movement when the positions of large amplitudes are matched. The diagram on the left side of (C) shows that the correlation function is calculated by the second correlator 43 in a state where the second segment is moved rk- with respect to the first segment and the waveforms are almost matched. If the pitch L is the same in both segments as assumed in (A), rk− or rk + is the correlation delay rk
However, the pitch L is often different between the two segments. However, since the difference between the pitches L is not so large, when calculating the correlation function in a state where both segments are shifted by rk-, the search range may be reduced. For example, if the initial segment length is set to 7.5 ms, 8 kHz
In the case of z sampling discrete data, there are 60 samples of default segment data. Usually, multiplication and addition are performed on data of about 40 samples in order to obtain a correlation function. For example, a correlation function is obtained. The figure on the right side of (C) shows that the second segment is rk +
The second correlator 43 is moved in a state where the waveforms are substantially matched.
Shows the calculation of the correlation function. By determining rk from both left and right calculations, an accurate correlation delay amount r
k is obtained. Thus, the amount of calculation can be significantly reduced as compared with the case where the calculation is performed by the first correlator 41.

【0073】次に第2実施形態を説明する。本実施の形
態は切替スイッチ40により、第1相関器41と第2相
関器43の切替えに関するものである。フラグバッファ
35の有声/無声判別フラグにより、ピッチの明確な有
声データの場合は第2相関器43の出力を用い、ピッチ
の不明確な無声データのときは、第1相関器41の出力
により加算器44は加算を行う。さらにピッチ極性情報
により第2相関器43の選択を行う。
Next, a second embodiment will be described. This embodiment relates to switching between a first correlator 41 and a second correlator 43 by a changeover switch 40. According to the voiced / unvoiced discrimination flag of the flag buffer 35, the output of the second correlator 43 is used in the case of voiced data having a clear pitch, and is added by the output of the first correlator 41 in the case of unvoiced data having an unclear pitch. The unit 44 performs addition. Further, the second correlator 43 is selected based on the pitch polarity information.

【0074】図7は速度比α=3/4のとき相関遅延量
rkを導出する際の切替スイッチの切り替えを示す。
(a)は原音声の入った初期セグメント列を示す。
(b)はセグメントが初期セグメントでかつフラグバッ
ファ35のフラグが有声の場合である。なお本図におけ
るL(x)>0又はL(x)<0は図2の(E)で説明
したピッチ極性情報を示す。すなわちL(1)>0は第
1セグメントが右側の第2セグメントと類似し、L
(2)<0は第2セグメントが左側の第1セグメントに
類似していることを示す。の場合は、L(1)>0で
かつL(2)<0のときは第1セグメントと第2セグメ
ントは互いに類似しているので、第2相関器43の第2
相関遅延量rkは十分な精度が得られる。これにより第
2相関器43を選択し、その他は第1相関器41を選択
する。の場合も同様で、セグメント列は(b)に示
すようになる。
FIG. 7 shows the switching of the changeover switch when deriving the correlation delay amount rk when the speed ratio α = 3/4.
(A) shows an initial segment sequence containing the original voice.
(B) is a case where the segment is the initial segment and the flag of the flag buffer 35 is voiced. Note that L (x)> 0 or L (x) <0 in the drawing indicates the pitch polarity information described with reference to FIG. That is, if L (1)> 0, the first segment is similar to the second segment on the right, and L (1)> 0
(2) <0 indicates that the second segment is similar to the first segment on the left. In the case of (1), when L (1)> 0 and L (2) <0, the first segment and the second segment are similar to each other.
The correlation delay amount rk has sufficient accuracy. Thereby, the second correlator 43 is selected, and the others select the first correlator 41. The same applies to the case of (1), and the segment sequence is as shown in (b).

【0075】(c)はセグメント長が初期セグメント長
の2倍の2倍長セグメントの場合が続き、かつフラグが
有声の場合である。はL(1)>0,L(2)>0,
L(3)<0,L(4)<0の場合のみ第2相関器43
の出力に切り替える。この場合第1セグメントは第2セ
グメントに類似し、第2セグメントは第3セグメントに
類似する。第3セグメントは第2セグメントに類似し、
第4セグメントは第3セグメントに類似する。これは第
1と第2セグメントが類似し、第3と第4セグメントが
類似し、かつ第2と第3セグメントが類似している場合
である。の場合も同様である。
(C) is a case where the segment length is twice as long as the initial segment length, and the flag is voiced. Are L (1)> 0, L (2)> 0,
Second correlator 43 only when L (3) <0 and L (4) <0
Switch to output. In this case, the first segment is similar to the second segment, and the second segment is similar to the third segment. The third segment is similar to the second segment,
The fourth segment is similar to the third segment. This is the case when the first and second segments are similar, the third and fourth segments are similar, and the second and third segments are similar. The same applies to the case of.

【0076】図8は、速度比α=3/2のときの相関遅
延量rkを導出する際の切替スイッチの切り替えを示
す。(a)は原音声を格納したセグメント列を示し、
(b)は初期セグメント長を有し、かつフラグバッファ
35のフラグが有声を示してる場合である。はL
(1)>0でかつL(2)<0の場合のみ第2相関器4
3の出力を用い、その他は第1相関器41の出力を用い
る。,の場合も同様である。(c)はセグメントに
2倍長セグメントを使用し、かつ有声の場合である。
のときはL(1)>0,L(2)>0,L(3)<0,
L(4)<0の場合のみ第2相関器43の出力を選択
し、その他の場合は第1相関器41の出力を選択する。
の場合も同様である。図7、図8に示す処理により、
波形を窓関数により重み付けした後、加算する際に、よ
り類似した波形どうしの場合のみ、第2相関器43の出
力を選択することにより、相関遅延量rkの信頼性を向
上させることができる。また第2相関器43の出力を選
択することにより演算量を大幅に減少することができ
る。
FIG. 8 shows switching of the changeover switch when deriving the correlation delay amount rk when the speed ratio α = 3/2. (A) shows a segment string storing the original audio,
(B) shows a case where the flag has an initial segment length and the flag of the flag buffer 35 indicates voiced. Is L
Second correlator 4 only when (1)> 0 and L (2) <0
3 and the output of the first correlator 41 is used for the others. , Is the same. (C) shows a case where a double-length segment is used as a segment and the segment is voiced.
When L (1)> 0, L (2)> 0, L (3) <0,
The output of the second correlator 43 is selected only when L (4) <0, and the output of the first correlator 41 is selected otherwise.
The same applies to the case of. By the processing shown in FIGS. 7 and 8,
When the waveforms are weighted by the window function and then added, the reliability of the correlation delay amount rk can be improved by selecting the output of the second correlator 43 only when the waveforms are more similar. Further, by selecting the output of the second correlator 43, the amount of calculation can be significantly reduced.

【0077】次に第3実施形態を説明する。本実施の形
態では加算する第1セグメントと第2セグメントの先頭
アドレスを示すポインタP1とP2の位置を示すもの
で、従来例と本発明の場合を比較して示している。図9
は速度比α=3/4のときのポインタP1とP2を従来
方法と本発明と比較して示している。(a)は、相関遅
延量rk>0の場合の処理、(b)はrk<0のときの
処理である。(a)の場合、従来の場合よりポインタP
1をrk分強制的に右側に(時間軸の進行方向)に進行
させ次のブロックの先頭セグメントの先頭位置に移動さ
せ、次のブロックの処理を開始する。(b)のときは従
来のときよりもポインタP1をrk分強制的に左側に
(時間軸の後退方向)に後退させて次のブロックの処理
を開始する。なお、P2は従来も次のブロックの第2セ
グメントの先頭位置を示しているので、変更する必要は
ない。なお、速度変換前の1ブロックは4セグメントよ
りなる。
Next, a third embodiment will be described. In the present embodiment, the positions of pointers P1 and P2 indicating the start addresses of the first segment and the second segment to be added are shown, and a comparison is made between the conventional example and the present invention. FIG.
Shows pointers P1 and P2 when the speed ratio α = 3/4 in comparison with the conventional method and the present invention. (A) is a process when the correlation delay amount rk> 0, and (b) is a process when rk <0. In the case of (a), the pointer P
1 is forcibly advanced to the right (in the direction of progress of the time axis) by rk, moved to the head position of the head segment of the next block, and processing of the next block is started. In the case of (b), the pointer P1 is forcibly retracted to the left (reverse direction on the time axis) by rk as compared with the conventional case, and the processing of the next block is started. It should be noted that P2 has conventionally shown the head position of the second segment of the next block, and therefore does not need to be changed. One block before the speed conversion is composed of four segments.

【0078】図10は、速度比α=3/2のときのポイ
ンタP1,P2の位置を従来の方法と本発明の場合とで
比較して示す。(a)は相関遅延量rk<0の場合で、
従来よりもポインタP2をrk分強制的に右側に進行さ
せて次のブロックの処理を開始する。(b)のときは従
来よりもポインタP2をrk分強制的に左側に後退させ
て次のブロックの処理を開始する。
FIG. 10 shows the positions of the pointers P1 and P2 when the speed ratio α = 3/2, in comparison between the conventional method and the present invention. (A) is a case where the correlation delay amount rk <0,
The pointer P2 is forcibly advanced to the right by rk as compared with the related art, and the processing of the next block is started. In the case of (b), the pointer P2 is forcibly retracted to the left by rk as compared with the conventional case, and the processing of the next block is started.

【0079】図9、図10に示した処理により、図9の
(a)、図10の(b)に示すようにP1とP2が離れ
過ぎて、相関遅延量rkの信頼性が低下するのを防止す
ることができる。これにより自然な音声を再生すること
ができる。
The processes shown in FIGS. 9 and 10 cause P1 and P2 to be too far apart as shown in FIGS. 9 (a) and 10 (b), thereby reducing the reliability of the correlation delay rk. Can be prevented. Thereby, natural sound can be reproduced.

【0080】次に第4実施形態を説明する。本実施の形
態では相関遅延量rkが蓄積されると速度比αを維持で
きなくなるため、蓄積量が一定量、例えばセグメント長
になる度にその時の処理ブロックに1セグメント追加
(rk−の場合)、または1セグメント削除(rk+の
場合)する。これにより速度比αを維持できる。
Next, a fourth embodiment will be described. In the present embodiment, if the correlation delay amount rk is accumulated, the speed ratio α cannot be maintained. Therefore, each time the accumulated amount becomes a fixed amount, for example, a segment length, one segment is added to the processing block at that time (in the case of rk−). , Or one segment is deleted (in the case of rk +). Thus, the speed ratio α can be maintained.

【0081】図11は速度比α=3/4のときrkの蓄
積値が1個の初期セグメント長以上となったとき、スル
ー区間のセグメント(加算セグメントに続くセグメン
ト)を1個削除または追加することを示す。(a)はr
k=0の場合で標準形を示す。(b)はrk+の場合で
rk+の蓄積値が初期セグメント長tsを越えた時、そ
の時の処理中のブロックから初期セグメントを1個削除
する。(c)はrk−の積算値が初期セグメントの1個
分を越えた時、その時の処理中のブロックのスルー区間
に初期セグメントを1個追加する。これにより速度比α
=3/4を維持することができる。つまり(a)の標準
形と全体として同じ状態にすることができる。
FIG. 11 shows that when the accumulated value of rk is equal to or longer than the length of one initial segment when the speed ratio α = 3/4, one segment in the through section (the segment following the addition segment) is deleted or added. Indicates that (A) is r
The standard form is shown when k = 0. (B) In the case of rk +, when the accumulated value of rk + exceeds the initial segment length ts, one initial segment is deleted from the block being processed at that time. In (c), when the integrated value of rk- exceeds one initial segment, one initial segment is added to the through section of the block being processed at that time. This gives the speed ratio α
= 3/4 can be maintained. That is, it can be in the same state as the standard form of (a) as a whole.

【0082】図12は、速度比α=3/2のときrkの
絶対値の蓄積量が1初期セグメント長を越えた時の処理
を示す。(a)はrk=0の場合で標準形を示し、
(b)はrk−の場合で、rkの値は(−)の値を持っ
ている。rkの絶対値の積算値が初期セグメント長ts
を越えた時、その時処理中のブロックのスルー区間から
1個のセグメントを除去する。(c)はrk+の場合で
rkは正の値を有する。rkの積算値が初期セグメント
長tsを越えた時、処理中のブロックのスルー区間に1
個のセグメントを追加する。これにより速度比α=3/
2を維持できる。つまり(a)の標準形と全体として同
じ状態にすることができる。
FIG. 12 shows processing when the accumulated amount of the absolute value of rk exceeds one initial segment length when the speed ratio α = 3/2. (A) shows the standard form when rk = 0,
(B) is the case of rk-, where the value of rk has the value of (-). The integrated value of the absolute value of rk is the initial segment length ts
Is exceeded, one segment is removed from the through section of the block currently being processed. (C) is the case of rk +, where rk has a positive value. When the integrated value of rk exceeds the initial segment length ts, 1 is added to the through section of the block being processed.
Add segments. As a result, the speed ratio α = 3 /
2 can be maintained. That is, it can be in the same state as the standard form of (a) as a whole.

【0083】次に第5実施形態を説明する。本実施の形
態ではピッチ極性情報L(x)>0またはL(x)<0
に基づいて加算するセグメントを選択し、これにより1
ブロック内のセグメントの変化が生じる場合があるので
ブロック内のセグメント数の調整も行うようにしたもの
である。
Next, a fifth embodiment will be described. In the present embodiment, pitch polarity information L (x)> 0 or L (x) <0
Select the segments to be added based on
Since the segments in the block may change, the number of segments in the block is also adjusted.

【0084】図13は、速度比α=3/4の場合で、
(a)はピッチ極性情報によりスルー区間のセグメント
の増減を行わない標準形を示す。(b)は次のブロック
の処理に用いるピッチ極性情報を先読みして、現在の処
理ブロックで用いるピッチ極性情報と比較する。はL
(1)>0でかつL(5)<0のときで、隣接するブロ
ックのピッチ極性情報が互いに向き合っているときであ
る。このときは第1ブロックの第1セグメントと第2セ
グメントが類似し、第1ブロックの第4セグメントと第
2ブロックの第1セグメント(セグメント5)が類似し
ている場合である。このときは、第1ブロックではセグ
メント1とセグメント2を加算セグメントとし、第2ブ
ロックでは第1ブロックのセグメント4を取り込みセグ
メント5と加算セグメントを生成する。これにより互い
に類似したセグメントどうしを加算セグメントとするこ
とにより自然な再生音が得られる。しかし、このため第
1ブロックは初期セグメント1個分短くなる。
FIG. 13 shows a case where the speed ratio α = 3/4.
(A) shows a standard form in which the number of segments in a through section is not increased or decreased according to pitch polarity information. In (b), the pitch polarity information used in the processing of the next block is read in advance and compared with the pitch polarity information used in the current processing block. Is L
(1)> 0 and L (5) <0 when the pitch polarity information of adjacent blocks face each other. In this case, the first segment and the second segment of the first block are similar, and the fourth segment of the first block and the first segment (segment 5) of the second block are similar. At this time, in the first block, the segments 1 and 2 are set as addition segments, and in the second block, the segment 4 of the first block is fetched, and the segment 5 and the addition segment are generated. Thus, a natural reproduced sound can be obtained by using segments similar to each other as addition segments. However, the first block is therefore shortened by one initial segment.

【0085】はL(5)<0でかつL(9)<0の場
合である。この場合第2ブロックは第1ブロックからセ
グメント4を取り込みセグメント5と加算セグメントを
生成する。また第3ブロックにセグメント8を渡してし
まう。これにより第2ブロックは第1ブロックからセグ
メント4を取り込み、第3ブロックへセグメント8を渡
してしまうので、ブロック内のセグメント数は変わらな
い。ただし加算セグメントは互いに類似するセグメント
どうしなので自然な再生音が得られる。もの場合と
同じである。
Is the case where L (5) <0 and L (9) <0. In this case, the second block takes in the segment 4 from the first block and generates the segment 5 and the addition segment. In addition, the segment 8 is passed to the third block. As a result, the second block takes in the segment 4 from the first block and passes the segment 8 to the third block, so that the number of segments in the block does not change. However, since the added segments are similar to each other, a natural reproduced sound can be obtained. It is the same as the case.

【0086】はL(13)<0でかつL(17)>0
の場合で、第4ブロックの両側のピッチ極性情報が互い
に反発している場合である。つまり先頭セグメント13
は第3ブロックのセグメント12と類似している。第5
ブロックでは先頭セグメント17と次のセグメント18
が互いに類似していて、第4ブロックには関係ない。第
4ブロックは第3ブロックのセグメント12を取り込
み、先頭セグメント13と加算セグメントを生成する。
これにより第4ブロックは第3ブロックより初期セグメ
ントを1個取り込んだことにより、セグメント数は1個
増加する。
Is that L (13) <0 and L (17)> 0
In this case, the pitch polarity information on both sides of the fourth block repel each other. That is, the first segment 13
Is similar to segment 12 of the third block. Fifth
In the block, the first segment 17 and the next segment 18
Are similar to each other and are not relevant to the fourth block. The fourth block takes in the segment 12 of the third block, and generates the leading segment 13 and the addition segment.
As a result, the fourth block fetches one initial segment from the third block, thereby increasing the number of segments by one.

【0087】図14は、速度比α=3/2の場合で、
(a)はピッチ極性情報によりブロック内のセグメント
を増減しない標準形を示す。(b)は図13の(b)と
同様で、次のブロックで用いるピッチ極性情報を先読み
して、現在の処理ブロックで用いるピッチ極性情報と比
較する。のように互いに向き合っているときはスルー
区間を1初期セグメント長分短くし、のように互いに
反発している場合はスルー区間に1初期セグメント長分
長くして標準形の時間軸の時刻にずれが生じないように
する。この処理により、再生音声を指定した速度比αに
維持とすることができ、かつ窓関数による重み付け加算
の際で、ピッチ極性情報が向き合う場合をより多く取り
易くすることができ、相関遅延量rkの信頼性を維持お
よび向上させることにより自然な音声を再生することが
できる。
FIG. 14 shows a case where the speed ratio α = 3/2.
(A) shows a standard form in which the number of segments in a block is not increased or decreased according to pitch polarity information. (B) is the same as (b) of FIG. 13, prefetches the pitch polarity information used in the next block and compares it with the pitch polarity information used in the current processing block. If they are facing each other, the through section is shortened by one initial segment length, and if they repel each other, as in, the through section is extended by one initial segment length and shifted to the time on the standard time axis. Should not occur. By this processing, the reproduced voice can be maintained at the specified speed ratio α, and in the case of weighting addition by the window function, it is easier to take more cases where the pitch polarity information faces each other, and the correlation delay amount rk A natural sound can be reproduced by maintaining and improving the reliability of.

【0088】図15は第6実施形態を示す。本実施形態
は図1に示す第1実施形態の簡易化を計ったもので、図
1の第1相関器41を廃止し第2相関器43のみとしこ
れを相関器43aとし、さらに、フラグバッファ35、
スイッチ40、加算器46、rkバッファ51、遅延フ
ラグ発生器52、セグメント長シフトレジスタ53を廃
止し、残りの機器のこれら廃止した機器に関連する機能
も廃止したものである。
FIG. 15 shows a sixth embodiment. This embodiment is a simplification of the first embodiment shown in FIG. 1. The first correlator 41 in FIG. 1 is abolished and only the second correlator 43 is used, which is referred to as a correlator 43a. 35,
The switch 40, the adder 46, the rk buffer 51, the delay flag generator 52, and the segment length shift register 53 are abolished, and the functions of the remaining devices related to the abolished devices are also abolished.

【0089】図16は第7実施形態を示す。本実施形態
は図17に示す従来例を改良したもので、セグメントの
長さを可変とし、女性音声データや男性音声データを適
切に処理できるようにしたものである。
FIG. 16 shows a seventh embodiment. This embodiment is an improvement of the conventional example shown in FIG. 17, in which the length of a segment is variable so that female voice data and male voice data can be appropriately processed.

【0090】[0090]

【発明の効果】以上のように本発明は次の効果を奏す
る。 ピッチ情報、ピッチ極性情報、有声/無声判別フラグ
情報を用い、従来相関器のみで長時間かけて導出してい
た相関遅延量rkを波形ずれ量(rk)算出器と相関器
を用いることにより短時間に導出することができる。 ピッチ情報の大小により分割器を制御してセグメント
長を2段に切り替えることにより、男声、女声を含めた
汎用性のある音声再生処理が可能となる。 第1メモリと第2メモリのデータに、窓関数による重
み付け加算する際、音声データの波形の前後の相関関係
に依存したピッチ極性情報を用い、より相関の高い波形
どうしのみに、ピッチ情報からrk算出器、第2相関器
で第2相関遅延量rkを求めることにより、相関遅延量
rkの信頼性を向上させると共に演算量を削減すること
ができ、再生音の品質を向上できる。
As described above, the present invention has the following effects. Using the pitch information, the pitch polarity information, and the voiced / unvoiced discrimination flag information, the correlation delay amount rk, which was conventionally derived over a long period of time using only a correlator, can be reduced by using a waveform shift amount (rk) calculator and a correlator. Can be derived in time. By controlling the divider according to the magnitude of the pitch information and switching the segment length to two stages, versatile audio reproduction processing including male and female voices can be performed. When weighting and adding by the window function to the data in the first memory and the second memory, pitch polarity information depending on the correlation before and after the waveform of the audio data is used, and only the waveforms having higher correlation are used to calculate rk from the pitch information. By calculating the second correlation delay amount rk by the calculator and the second correlator, the reliability of the correlation delay amount rk can be improved, the amount of calculation can be reduced, and the quality of reproduced sound can be improved.

【0091】第1メモリに入力するデータの先頭アド
レスを示すP1,及び第2メモリに入力するデータの先
頭アドレスを示すP2の距離が、相関遅延量rkの大小
によって離れ過ぎてしまい、その結果、相関遅延量rk
の信頼性が失われるのを防ぐため、P1とP2の間隔を
一定にして相関遅延量rkの信頼性を向上させることに
より、再生音声の品質を向上できる。 第1相関器、第2相関器によって出力する相関遅延量
rkを積算し、1セグメント長以上になったとき1セグ
メント追加または削除することにより、再生音声を指定
した速度比αに維持することができる。 重み付け加算及びスルー区間制御を含めた一連の処理
で、次の処理に用いるピッチ極性情報を先読みし、その
結果を用いてスルー区間の長さを制御することにより、
指定した速度比αを維持しつつ、ピッチ極性情報の向か
い合う場合での重み付け加算処理をできるだけ取り易く
することによって、相関遅延量rkの信頼性を向上さ
せ、自然な再生音声を得ることができる。
The distance between P1, which indicates the head address of the data input to the first memory, and P2, which indicates the head address of the data input to the second memory, is too far depending on the magnitude of the correlation delay rk. Correlation delay rk
In order to prevent loss of reliability, the interval between P1 and P2 is kept constant to improve the reliability of the correlation delay amount rk, thereby improving the quality of reproduced sound. By accumulating the correlation delay amount rk output by the first correlator and the second correlator, and adding or deleting one segment when the length becomes equal to or more than one segment length, the reproduced voice can be maintained at the specified speed ratio α. it can. In a series of processing including weighting addition and through section control, by pre-reading the pitch polarity information used in the next processing, and controlling the length of the through section using the result,
By maintaining the specified speed ratio α and making the weighted addition process possible when pitch polarity information is opposed as much as possible, the reliability of the correlation delay amount rk can be improved, and a natural reproduced voice can be obtained.

【0092】相関器を波形ずれ量から相関値を算出す
るもののみとすることにより、装置を簡易化できる。 従来の装置にセグメント可変機能を付加することによ
り音声データを適切に処理できる。
By using only a correlator that calculates a correlation value from the amount of waveform deviation, the apparatus can be simplified. By adding a segment variable function to a conventional device, audio data can be appropriately processed.

【図面の簡単な説明】[Brief description of the drawings]

【図1】本発明の実施形態の音声速度変換装置のブロッ
ク図
FIG. 1 is a block diagram of an audio speed conversion device according to an embodiment of the present invention.

【図2】音声データとセグメントを説明する図FIG. 2 is a diagram illustrating audio data and segments.

【図3】相関遅延量rkを説明する図FIG. 3 is a diagram for explaining a correlation delay amount rk.

【図4】速度比α=3/4の場合のセグメント可変処理
説明図
FIG. 4 is an explanatory diagram of a segment variable process when the speed ratio α = 3/4.

【図5】速度比α=3/2の場合のセグメント可変処理
説明図
FIG. 5 is an explanatory diagram of segment variable processing when the speed ratio α = 3/2.

【図6】rk算出器及び第2相関器の処理説明図FIG. 6 is an explanatory diagram of the processing of the rk calculator and the second correlator

【図7】速度比α=3/4における相関遅延量導出の際
の切替スイッチ処理説明図
FIG. 7 is an explanatory diagram of a changeover switch process when deriving a correlation delay amount at a speed ratio α = 3/4.

【図8】速度比α=3/2における相関遅延量導出の際
の切替スイッチ処理説明図
FIG. 8 is an explanatory diagram of a changeover switch process when deriving a correlation delay amount at a speed ratio α = 3/2.

【図9】速度比α=3/4におけるポインタP1,P2
説明図
FIG. 9 shows pointers P1 and P2 at a speed ratio α = 3/4.
Illustration

【図10】速度比α=3/2におけるポインタP1,P
2説明図
FIG. 10 shows pointers P1 and P at a speed ratio α = 3/2.
2 explanatory diagram

【図11】速度比α=3/4におけるrk逐次加算処理
に伴う補正処理を説明する図
FIG. 11 is a view for explaining correction processing accompanying rk successive addition processing at a speed ratio α = 3/4.

【図12】速度比α=3/2におけるrk逐次加算処理
に伴う補正処理を説明する図
FIG. 12 is a view for explaining correction processing accompanying rk successive addition processing at a speed ratio α = 3/2.

【図13】速度比α=3/4におけるピッチ極性情報に
基づくスルー区間調整処理説明図
FIG. 13 is an explanatory diagram of through section adjustment processing based on pitch polarity information at a speed ratio α = 3/4.

【図14】速度比α=3/2におけるピッチ極性情報に
基づくスルー区間調整処理説明図
FIG. 14 is an explanatory diagram of through section adjustment processing based on pitch polarity information at a speed ratio α = 3/2.

【図15】本発明の第6実施形態のブロック図FIG. 15 is a block diagram of a sixth embodiment of the present invention.

【図16】本発明の第7実施形態のブロック図FIG. 16 is a block diagram of a seventh embodiment of the present invention.

【図17】従来の音声速度変換装置のブロック図FIG. 17 is a block diagram of a conventional voice speed converter.

【図18】従来の速度変換模式図FIG. 18 is a schematic diagram of a conventional speed conversion.

【符号の説明】[Explanation of symbols]

31 A/D変換器 32 音声符号化/復号器 33 データバッファ 34 ピッチバッファ 35 フラグバッファ 36 速度制御回路 37 分割器 38 第1メモリ 39 第2メモリ 40 切替スイッチ 41 第1相関器 42 波形ずれ量(rk)算出器 43 第2相関器 43a 相関器 44 加算器 45 窓関数発生器 46 加算器 47,48 乗算器 49 結合器 50 D/A変換器 51 rkバッファ 52 遅延フラグ発生器 53 セグメント長シフトレジスタ Reference Signs List 31 A / D converter 32 Audio encoder / decoder 33 Data buffer 34 Pitch buffer 35 Flag buffer 36 Speed control circuit 37 Divider 38 First memory 39 Second memory 40 Changeover switch 41 First correlator 42 Waveform shift amount ( rk) calculator 43 second correlator 43a correlator 44 adder 45 window function generator 46 adder 47,48 multiplier 49 combiner 50 D / A converter 51 rk buffer 52 delay flag generator 53 segment length shift register

Claims (14)

【特許請求の範囲】[Claims] 【請求項1】 符号化した音声データを復号化する復号
器と、この復号器より得られる音声データのピッチ情報
と1つのピッチ内のデータが隣接する時間的に後のピッ
チのデータに類似するか前のピッチのデータに類似する
かを示すピッチ極性情報とを格納するピッチ格納部と、
前記復号器より得られる音声データのピッチが明確か不
明確かを示したフラグを格納するフラグ格納部と、前記
復号器からのデータを所定長さのセグメントに分割し連
続する所定数のセグメントを1つのブロックとし、この
各ブロックのセグメントを入力した順に分けて出力する
分割器と、この分割器が出力する各ブロックの最初のセ
グメントを格納する第1メモリと、前記分割器が出力す
る各ブロックの2番目のセグメントを格納する第2メモ
リと、前記第1メモリの内容と前記第2メモリの内容と
の相関関数を算出し、第1相関遅延量を演算する第1相
関器と、前記ピッチ格納部のデータより前記最初のセグ
メントの波形と前記次のセグメントの波形とがほぼ一致
するようにずらして重ねそのずれ量を算出する波形ずれ
量算出器と、前記第1メモリの内容と前記第2メモリの
内容との相関関数を前記波形ずれ量算出器の出力に基づ
いて算出し第2相関遅延量を演算する第2相関器と、セ
グメントのデータに対する重み係数を発生する窓関数発
生器と、前記第1メモリの内容に前記重み係数を乗算す
る第1乗算器と、前記第2メモリの内容に前記重み係数
を乗算する第2乗算器と、前記第1相関器または前記第
2相関器の出力に基づき前記第1乗算器の出力と前記第
2乗算器の出力とを相関関数の値が大きい位置で加算す
る加算器と、前記第1相関器と前記第2相関器との出力
を前記フラグ格納部のフラグに応じて切り替えて出力す
る切替えスイッチと、前記加算器の出力とこれに続け
て、前記分割器より入力する低速化のときは第2セグメ
ント以降、高速化の時は第3セグメント以降のセグメン
トで構成されるスルー区間セグメントを出力する結合器
と、音声の速度の変換比を表す速度比に基づき前記分割
器と前記結合器を制御する速度制御部と、を備えた音声
速度変換装置。
1. A decoder for decoding encoded audio data, and pitch information of audio data obtained from the decoder and data in one pitch are similar to data of an adjacent temporally later pitch. A pitch storage unit for storing pitch polarity information indicating whether the data is similar to the data of the previous pitch,
A flag storage unit for storing a flag indicating whether the pitch of the audio data obtained from the decoder is clear or unknown, and dividing the data from the decoder into segments of a predetermined length and storing a predetermined number of continuous segments as 1 A divider that divides and outputs the segments of each block in the order of input, a first memory that stores the first segment of each block that is output by the divider, and a first memory that stores the first segment of each block that is output by the divider. A second memory for storing a second segment, a first correlator for calculating a correlation function between the contents of the first memory and the contents of the second memory, and calculating a first correlation delay amount; The waveform of the first segment and the waveform of the next segment are superimposed so that the waveforms of the first segment and the next segment are substantially coincident with each other, and a waveform deviation amount calculator that calculates the deviation amount; A second correlator that calculates a correlation function between the contents of one memory and the contents of the second memory based on the output of the waveform shift amount calculator to calculate a second correlation delay amount; A generated window function generator, a first multiplier for multiplying the content of the first memory by the weighting factor, a second multiplier for multiplying the content of the second memory by the weighting factor, An adder for adding the output of the first multiplier and the output of the second multiplier at a position where the value of the correlation function is large based on the output of the second correlator or the first correlator; (2) a changeover switch for switching the output from the correlator according to the flag in the flag storage unit and outputting the output; and the output from the adder and, subsequently, the second segment and subsequent segments when the speed is input from the divider. 3rd segment for high speed A speech speed converter comprising: a combiner that outputs a through section segment composed of subsequent segments; and a speed control unit that controls the divider and the combiner based on a speed ratio that represents a conversion ratio of the speed of sound. apparatus.
【請求項2】 前記速度制御部は、前記ピッチ情報に基
づいて前記セグメントの長さを変更することを特徴とす
る請求項1記載の音声速度変換装置。
2. The audio speed conversion device according to claim 1, wherein the speed control unit changes the length of the segment based on the pitch information.
【請求項3】 前記セグメントの長さを女性の声を基準
にした初期セグメント長さに設定し、前記ピッチ情報の
ピッチがこの初期セグメント長さを超える場合、セグメ
ントの長さを初期セグメントの2倍とし、ブロックの長
さも2倍とすることを特徴とする請求項2記載の音声速
度変換装置。
3. The length of the segment is set to an initial segment length based on a female voice, and when the pitch of the pitch information exceeds the initial segment length, the length of the segment is set to 2 of the initial segment. 3. The audio speed conversion device according to claim 2, wherein the length is doubled and the length of the block is doubled.
【請求項4】 前記セグメントの長さを切り替えたブロ
ックにおいては、前記切替スイッチは、前記第1相関器
の出力を前記加算器に出力することを特徴とする請求項
3記載の音声速度変換装置。
4. The audio speed conversion device according to claim 3, wherein in the block in which the length of the segment is switched, the switch outputs the output of the first correlator to the adder. .
【請求項5】 前記切替スイッチは、前記フラグ格納部
の音声データのピッチが明確な場合、前記第2相関器の
出力を前記加算器に出力し、ピッチが不明確な場合、前
記第1相関器の出力を前記加算器に出力することを特徴
とする請求項1記載の音声速度変換装置。
5. The switch outputs the output of the second correlator to the adder when the pitch of the audio data in the flag storage unit is clear, and outputs the first correlation when the pitch is unclear. 2. The audio speed conversion device according to claim 1, wherein the output of the audio device is output to the adder.
【請求項6】 前記ピッチ極性情報に基づき前記ブロッ
クの最初のセグメントと次のセグメントのデータが類似
しているとき前記切替スイッチは前記第2相関器の出力
を前記加算器に出力することを特徴とする請求項1記載
の音声速度変換装置。
6. The changeover switch outputs an output of the second correlator to the adder when data of a first segment and a next segment of the block are similar based on the pitch polarity information. The audio speed conversion device according to claim 1, wherein
【請求項7】 前記初期セグメントの2倍を新セグメン
トとし、最初の新セグメントはデータ伝送順に第1初期
セグメントと第2初期セグメントからなり、次の新セグ
メントはデータ伝送順に第3初期セグメントと第4初期
セグメントからなり、第1初期セグメントが第2初期セ
グメントに類似し、第2初期セグメントと第3初期セグ
メントが類似し、第4初期セグメントが第3初期セグメ
ントに類似している場合、前記切替スイッチは前記第2
相関器の出力を前記加算器に出力することを特徴とする
請求項3記載の音声速度変換装置。
7. Double the initial segment as a new segment, wherein the first new segment comprises a first initial segment and a second initial segment in data transmission order, and the next new segment comprises a third initial segment and a second initial segment in data transmission order. If the first initial segment is similar to the second initial segment, the second initial segment is similar to the third initial segment, and the fourth initial segment is similar to the third initial segment, The switch is the second
The audio speed conversion device according to claim 3, wherein an output of a correlator is output to the adder.
【請求項8】 前記第1メモリおよび前記第2メモリに
入力するデータの先頭は前記セグメントの先頭番地を示
すことを特徴とする請求項1記載の音声速度変換装置。
8. The audio speed conversion device according to claim 1, wherein a head of data input to the first memory and the second memory indicates a head address of the segment.
【請求項9】 前記第1相関遅延量と前記第2相関遅延
量とを積算する相関遅延量積算手段と、この積算値がセ
グメント長にほぼ達したとき遅延フラグを発生する遅延
フラグ発生器とをさらに備え、前記速度制御部は、この
遅延フラグに基づきブロックのスルー区間のセグメント
を追加または削除することを特徴とする請求項1記載の
音声速度変換装置。
9. A correlation delay amount integrating means for integrating the first correlation delay amount and the second correlation delay amount, and a delay flag generator for generating a delay flag when the integrated value substantially reaches a segment length. The audio speed conversion apparatus according to claim 1, further comprising: a speed control unit that adds or deletes a segment of a through section of the block based on the delay flag.
【請求項10】 前記速度制御部は、前記ピッチ極性情
報に基づき、現在のブロックの先頭セグメントが次のセ
グメントと類似し、現在のブロックの最後のセグメント
が次のブロックの先頭セグメントと類似している場合、
現在のブロックの最後のセグメントを除去して現在のブ
ロックのスルー区間を短縮することを特徴とする請求項
1記載の音声速度変換装置。
10. The speed controller, based on the pitch polarity information, wherein the first segment of the current block is similar to the next segment, and the last segment of the current block is similar to the first segment of the next block. If you have
2. The apparatus according to claim 1, wherein a last segment of the current block is removed to shorten a through section of the current block.
【請求項11】 前記速度制御部は、前記ピッチ極性情
報に基づき、現在のブロックの先頭セグメントが1つ前
のブロックの最後のセグメントと類似しており、次のブ
ロックの先頭セグメントがそのブロックの次のセグメン
トを類似している場合、前のブロックの最後のセグメン
トを現在のブロックに取り込んで先頭のセグメントと
し、現在のブロックの旧先頭セグメントを第2セグメン
トとして加算し、現在のブロックのスルー区間を1セグ
メント増加することを特徴とする請求項1記載の音声速
度変換装置。
11. The speed controller according to claim 1, wherein the first segment of the current block is similar to the last segment of the immediately preceding block, and the first segment of the next block is determined based on the pitch polarity information. If the next segment is similar, the last segment of the previous block is taken into the current block as the first segment, the old first segment of the current block is added as the second segment, and the through section of the current block is added. 2. The audio speed conversion device according to claim 1, wherein the number is increased by one segment.
【請求項12】 前記速度制御部は、前記ピッチ極性情
報に基づき、現在のブロックの先頭セグメントが1つ前
のブロックの最後のセグメントと類似し、現在のブロッ
クの最後のセグメントが次のブロックの先頭セグメント
と類似する場合には、現在のブロックは1つ前のブロッ
クのセグメントを先頭セグメントとし、現在のブロック
の旧先頭セグメントを第2セグメントとして加算し、最
後のセグメントを次のブロックの先頭セグメントとして
加算することを特徴とする請求項1記載の音声速度変換
装置。
12. The speed controller according to claim 1, wherein the first segment of the current block is similar to the last segment of the immediately preceding block, and the last segment of the current block is determined based on the pitch polarity information. If it is similar to the first segment, the current block is set to the segment of the previous block as the first segment, the old first segment of the current block is added as the second segment, and the last segment is set to the first segment of the next block. The audio speed conversion device according to claim 1, wherein the addition is performed as:
【請求項13】 符号化した音声データを復号化する復
号器と、この復号器より得られる音声データのピッチ情
報と1つのピッチ内のデータが隣接する時間的に後のピ
ッチのデータに類似するか前のピッチのデータに類似す
るかを示すピッチ極性情報とを格納するピッチ格納部
と、前記復号器からのデータを所定長さのセグメントに
分割し連続する所定数のセグメントを1つのブロックと
し、この各ブロックのセグメントを入力した順に分けて
出力する分割器と、この分割器が出力する各ブロックの
最初のセグメントを格納する第1メモリと、前記分割器
が出力する各ブロックの2番目のセグメントを格納する
第2メモリと、前記ピッチ格納部のデータより前記最初
のセグメントの波形と前記次のセグメントの波形とがほ
ぼ一致するようにずらして重ねそのずれ量を算出する波
形ずれ量算出器と、前記第1メモリの内容と前記第2メ
モリの内容との相関関数を前記波形ずれ量算出器の出力
に基づいて算出し相関遅延量を演算する相関器と、セグ
メントのデータに対する重み係数を発生する窓関数発生
器と、前記第1メモリの内容に前記重み係数を乗算する
第1乗算器と、前記第2メモリの内容に前記重み係数を
乗算する第2乗算器と、前記相関器の出力に基づき前記
第1乗算器の出力と前記第2乗算器の出力とを相関関数
の値が大きい位置で加算する加算器と、前記加算器の出
力とこれに続けて、前記分割器より入力する低速化のと
きは第2セグメント以降、高速化の時は第3セグメント
以降のセグメントで構成されるスルー区間セグメントを
出力する結合器と、音声の速度の変換比を表す速度比に
基づき前記分割器と前記結合器を制御する速度制御部
と、を備えた音声速度変換装置。
13. A decoder for decoding encoded audio data, and pitch information of audio data obtained from the decoder and data in one pitch are similar to data of an adjacent temporally later pitch. Or a pitch storage unit for storing pitch polarity information indicating whether the data is similar to the data of the previous pitch, and dividing the data from the decoder into segments of a predetermined length and dividing a predetermined number of continuous segments into one block. A first memory for storing the first segment of each block output by the divider; a second memory for storing the first segment of each block output by the divider; and a second memory of each block output by the divider. A second memory for storing a segment, and a shift of the waveform of the first segment and the waveform of the next segment based on the data of the pitch storage unit so that the waveform substantially matches the waveform of the next segment. And a waveform shift amount calculator for calculating the shift amount of the overlap, and calculating a correlation function between the content of the first memory and the content of the second memory based on the output of the waveform shift amount calculator. , A window function generator for generating a weight coefficient for the data of the segment, a first multiplier for multiplying the content of the first memory by the weight coefficient, and a weight for the content of the second memory. A second multiplier for multiplying a coefficient, an adder for adding an output of the first multiplier and an output of the second multiplier at a position where a value of a correlation function is large, based on an output of the correlator; A combiner for outputting a through section segment composed of the second and subsequent segments at the time of low-speed input and the third and subsequent segments at the time of high-speed input from the divider; Audio speed conversion ratio A speed control unit for controlling the coupler and the splitter based on a speed ratio representative, voice speed converting apparatus having a.
【請求項14】 音声データを入力して所定長さのセグ
メントに分割し連続する所定数のセグメントを1つのブ
ロックとし、この各ブロックのセグメントを入力した順
に分けて出力する分割器と、この分割器が出力する各ブ
ロックの最初のセグメントを格納する第1メモリと、前
記分割器が出力する各ブロックの2番目のセグメントを
格納する第2メモリと、前記第1メモリの内容と前記第
2メモリの内容との相関関数を算出し相関遅延量を演算
する相関器と、セグメントのデータに対する重み係数を
発生する窓関数発生器と、前記第1メモリの内容に前記
重み係数を乗算する第1乗算器と、前記第2メモリの内
容に前記重み係数を乗算する第2乗算器と、前記相関器
の出力に基づき前記第1乗算器の出力と前記第2乗算器
の出力とを相関関数の値が大きい位置で加算する加算器
と、前記加算器の出力とこれに続けて、前記分割器より
入力する低速化のときは第2セグメント以降、高速化の
時は第3セグメント以降のセグメントで構成されるスル
ー区間セグメントを出力する結合器と、音声の速度の変
換比を表す速度比に基づき前記分割器と前記結合器を制
御する速度制御部と、を備え、前記速度制御部は前記分
割部を制御して前記セグメントの長さを可変にすること
を特徴とする音声速度変換装置。
14. A divider that inputs audio data, divides it into segments of a predetermined length, divides a predetermined number of continuous segments into one block, and divides and outputs the segments of each block in the order of input. A first memory for storing the first segment of each block output by the divider, a second memory for storing a second segment of each block output by the divider, the contents of the first memory and the second memory And a window function generator for generating a weighting factor for the data of the segment, and a first multiplication for multiplying the content of the first memory by the weighting factor. A multiplier, a second multiplier for multiplying the content of the second memory by the weighting coefficient, and a correlation function between an output of the first multiplier and an output of the second multiplier based on an output of the correlator. An adder for adding at a position where the value of is larger, an output of the adder, and a subsequent segment input from the divider when the speed is low, and after the third segment when the speed is high. And a speed control unit that controls the splitter and the combiner based on a speed ratio representing a conversion ratio of the speed of voice, and the speed control unit includes the speed control unit. An audio speed conversion device characterized in that the length of the segment is made variable by controlling a dividing unit.
JP9083640A 1997-04-02 1997-04-02 Speech rate converting device Pending JPH10282991A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP9083640A JPH10282991A (en) 1997-04-02 1997-04-02 Speech rate converting device

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP9083640A JPH10282991A (en) 1997-04-02 1997-04-02 Speech rate converting device

Publications (1)

Publication Number Publication Date
JPH10282991A true JPH10282991A (en) 1998-10-23

Family

ID=13808058

Family Applications (1)

Application Number Title Priority Date Filing Date
JP9083640A Pending JPH10282991A (en) 1997-04-02 1997-04-02 Speech rate converting device

Country Status (1)

Country Link
JP (1) JPH10282991A (en)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2005266571A (en) * 2004-03-19 2005-09-29 Sony Corp Variable speed reproduction method and apparatus, and program
JP2007511162A (en) * 2003-11-11 2007-04-26 コスモタン インク. Digital audio signal and audio / video signal shift processing method and digital broadcast signal shift reproduction method using the same
JP2009053618A (en) * 2007-08-29 2009-03-12 Yamaha Corp Speech processing device and program

Cited By (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2007511162A (en) * 2003-11-11 2007-04-26 コスモタン インク. Digital audio signal and audio / video signal shift processing method and digital broadcast signal shift reproduction method using the same
JP2005266571A (en) * 2004-03-19 2005-09-29 Sony Corp Variable speed reproduction method and apparatus, and program
JP2009053618A (en) * 2007-08-29 2009-03-12 Yamaha Corp Speech processing device and program
US8214211B2 (en) 2007-08-29 2012-07-03 Yamaha Corporation Voice processing device and program

Similar Documents

Publication Publication Date Title
JP3451900B2 (en) Pitch / tempo conversion method and device
US5842172A (en) Method and apparatus for modifying the play time of digital audio tracks
KR100303913B1 (en) Sound processing method, sound processor, and recording/reproduction device
JP2004505304A (en) Digital audio signal continuously variable time scale change
JPH10282991A (en) Speech rate converting device
JP2005512134A (en) Digital audio with parameters for real-time time scaling
JPS6054580A (en) Reproducing device of video signal
US5679912A (en) Music production and control apparatus with pitch/tempo control
JP3162945B2 (en) Video tape recorder
JPS642960B2 (en)
JP3156020B2 (en) Audio speed conversion method
JP4442239B2 (en) Voice speed conversion device and voice speed conversion method
JPH1078791A (en) Pitch converter
KR100359988B1 (en) real-time speaking rate conversion system
JP2812379B2 (en) Sound source device
JP2890530B2 (en) Audio speed converter
US5959561A (en) Digital analog converter with means to overcome effects due to loss of phase information
JP2624538B2 (en) Audio synchronization method for television format conversion
JP2000181458A (en) Time stretching device
JP2669088B2 (en) Audio speed converter
JP2709198B2 (en) Voice synthesis method
JPH0477320B2 (en)
KR100659883B1 (en) How to play back videos in sync with audio
JPH10187180A (en) Tone generator
JPH01267700A (en) Speech processor