JPH064689A - Linear array parallel DSP processor - Google Patents
Linear array parallel DSP processorInfo
- Publication number
- JPH064689A JPH064689A JP4158264A JP15826492A JPH064689A JP H064689 A JPH064689 A JP H064689A JP 4158264 A JP4158264 A JP 4158264A JP 15826492 A JP15826492 A JP 15826492A JP H064689 A JPH064689 A JP H064689A
- Authority
- JP
- Japan
- Prior art keywords
- memory
- data
- section
- cell
- processor
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Landscapes
- Image Processing (AREA)
- Memory System (AREA)
- Dram (AREA)
Abstract
(57)【要約】
【目的】 同一命令サイクル中に書き込みと次の処理の
ための読み出しを同時に実行する。
【構成】 メモリAセル3及びメモリBセル5に、3ト
ランジスタのメモリセルを使用し、かつその書き込みビ
ット線と読み出しビット線を専用化し、個別に持つよう
にした。すなわち、入力SAM部からデータメモリA部
を経由してALUアレイ部に至る各プロセッサエレメン
トごとに1本の縦のビット線は、従来例で入力SAM部
からデータメモリA部を経由してALUアレイ部に至る
読み出しビット線と、ALUアレイ部からデータメモリ
A部に至る書き込みビット線の、各プロセッサエレメン
トごとに2本の縦のビット線になっている。この時、メ
モリAセル3は、その書き込みのためのひとつのトラン
ジスタと、読み込みのための2つのトランジスタは、そ
れぞれ別の書き込みビット線と読み出しビット線に接続
されている。
(57) [Summary] [Purpose] Write and read for the next process are executed simultaneously during the same instruction cycle. [Structure] A three-transistor memory cell is used for the memory A cell 3 and the memory B cell 5, and the write bit line and the read bit line thereof are dedicated and individually provided. That is, one vertical bit line for each processor element from the input SAM section to the ALU array section via the data memory A section has a conventional ALU array from the input SAM section via the data memory A section. There are two vertical bit lines for each processor element, that is, a read bit line reaching the area and a write bit line extending from the ALU array section to the data memory A section. At this time, in the memory A cell 3, one transistor for writing and two transistors for reading are connected to different write bit lines and read bit lines, respectively.
Description
【0001】[0001]
【産業上の利用分野】本発明は、テレビなどの映像信号
の処理をプログラマブルに実現するリニアアレイ型の並
列DSPプロセッサに関するもので、特にその構成要素
のメモリセルの、読み出しと書き込みのビット線を別々
にして、メモリアクセスのリードとライトを別々とし
て、パイプライン動作を可能とし、そのプロセッサの処
理性能を向上させたものに関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a linear array type parallel DSP processor capable of programmatically processing a video signal of a television or the like, and particularly to a read / write bit line of a memory cell of its constituent elements. The present invention relates to a memory access and a memory access read and write separately, enabling pipeline operation and improving the processing performance of the processor.
【0002】[0002]
【従来の技術】本発明は、テレビなどの映像信号をソフ
トウェアプログラムでディジタル信号処理するための、
SIMD制御のプログラマブルなリニアアレイ型の並列
DSPプロセッサの構成に関するもので、特にその処理
能力を向上させるために、各プロセッサエレメントにお
けるALUとデータメモリの間の書き込み及び読み出し
をパイプライン化した構成とすることを特徴とするもの
である (SIMD制御:Single Instruction stream Mu
lti Data stream/全てのプロセッサエレメントが一つ
のプログラムにより連動して動作する)。BACKGROUND OF THE INVENTION The present invention relates to a digital signal processing of a video signal of a television or the like by a software program,
The present invention relates to the configuration of a SIMD-controlled programmable linear array type parallel DSP processor, and in particular, in order to improve its processing capability, the configuration is such that the writing and reading between the ALU and the data memory in each processor element are pipelined. (SIMD control: Single Instruction stream Mu
lti Data stream / All processor elements work together by one program).
【0003】従来技術として、テレビなどの映像信号の
ディジタル信号処理をプログラマブルに実現するプロセ
ッサの構成として、SIMD制御のリニアアレイ型プロ
セッサがある(例えば、JIM CHILDERS,et al "SVP:SERI
AL VIDEO PROCESSSOR" IEEE1990 CUSTOM INTEGRATED CI
RCUITS CONFERENCE 17.3 )。As a conventional technique, there is a SIMD-controlled linear array type processor (for example, JIM CHILDERS, et al "SVP: SERI) as a configuration of a processor for programmable digital signal processing of a video signal of a television or the like.
AL VIDEO PROCESSSOR "IEEE1990 CUSTOM INTEGRATED CI
RCUITS CONFERENCE 17.3).
【0004】このプロセッサは、図2のように1ビット
ALUによる演算アレイをVRAMに組み込んだ形にな
っている。以下この図2から説明する。As shown in FIG. 2, this processor has a form in which a 1-bit ALU arithmetic array is incorporated in a VRAM. Hereinafter, description will be made from FIG.
【0005】このプロセッサは、大きくは、入力SAM
(シリアルアクセスメモリ)部、データメモリA部、A
LUアレイ部、データメモリB部、出力SAM部、プロ
グラム制御部などに分けられる。This processor is mainly composed of an input SAM.
(Serial access memory) section, Data memory A section, A
It is divided into an LU array section, a data memory B section, an output SAM section, a program control section, and the like.
【0006】プログラム制御部にはプログラムメモリと
そのシーケンス制御回路がある。また入力SAM部、デ
ータメモリA部、ALUアレイ部、データメモリB部、
出力SAM部は、全体でリニアアレイ(直線配列)型の
多数並列化したプロセッサエレメント群を構成してお
り、プログラム制御部内にある、共通の一つのプログラ
ム制御部により連動してSIMD制御される。The program control unit includes a program memory and its sequence control circuit. Also, an input SAM unit, a data memory A unit, an ALU array unit, a data memory B unit,
The output SAM unit constitutes a large number of linear array type linearly arranged processor element groups as a whole, and SIMD control is performed in conjunction with one common program control unit in the program control unit.
【0007】なお入力SAM部、データメモリA部、デ
ータメモリB部、出力SAM部は、基本的にメモリであ
り、それらのメモリのためのROWアドレスデコード
は、詳細に説明しないが、図2においては、このプログ
ラム制御部に含まれているものとして以下説明する。The input SAM unit, the data memory A unit, the data memory B unit, and the output SAM unit are basically memories, and the ROW address decoding for these memories will not be described in detail, but in FIG. Will be described below as being included in this program control unit.
【0008】多数並列化されたプロセッサエレメントの
単一エレメントは、図2で斜線で示したような縦の細長
い範囲であり、これが図で横方向に直線配列で並んでい
る。即ち、ひとつのプロセッサエレメントを構成するの
に必要な、ごく一般的な図3のようなプロセッサの構成
を、図2の斜線で示した縦の細長いプロセッサエレメン
トがそれぞれ実現している。A single element of a number of paralleled processor elements is a vertical elongated area, as shown by the diagonal lines in FIG. 2, which are arranged in a horizontal linear arrangement in the figure. That is, the vertically long and slender processor elements shown by the slanted lines in FIG. 2 respectively realize the configuration of the processor as shown in FIG. 3 that is generally necessary for constructing one processor element.
【0009】入力バッファメモリ(IQ)に相当するの
が入力SAM部である。出力バッファメモリ(OQ)に
相当するのが出力SAM部である。第1のデータメモリ
(RFB)に相当するのがデータメモリB部である。第
2のデータメモリ(RFA)に相当するのがデータメモ
リA部である。第1のデータメモリと第2のデータメモ
リのデータを、必要に応じて選んで演算するためのセレ
クタ(SEL)及びALUに相当するのが、ALUアレ
イ部である。The input SAM unit corresponds to the input buffer memory (IQ). The output SAM unit corresponds to the output buffer memory (OQ). The data memory B section corresponds to the first data memory (RFB). The data memory A section corresponds to the second data memory (RFA). The ALU array unit corresponds to the selector (SEL) and ALU for selecting and operating the data of the first data memory and the data of the second data memory as needed.
【0010】このプロセッサエレメントの普通のプロセ
ッサとの違いは、普通のプロセッサではそのハードウェ
アはワードプロセッサであり、ワードを単位として処理
するが、このプロセッサの場合はそのハードウェアはビ
ットプロセッサであり、ビットを単位として処理する点
である。ビットプロセッサはハードウェアが小さく、普
通には実現できない程多数の並列数を実現できる。なお
プロセッサエレメントの直線配列の並列数は、映像信号
の一水平走査期間の画素数(H)に一致させる。The difference between this processor element and an ordinary processor is that in an ordinary processor, the hardware is a word processor and the word is processed as a unit. In the case of this processor, the hardware is a bit processor, and Is a point to be processed. Bit processors have small hardware and can realize a large number of parallel numbers that cannot be realized normally. The parallel number of the linear array of processor elements is set to match the number of pixels (H) in one horizontal scanning period of the video signal.
【0011】更にこのプロセッサエレメントの構造は、
図4のように概略書くことができる。入力SAM部の一
つのプロセッサエレメント分は、図4における一つの入
力ポインタセル1と縦に並んだ複数の入力SAMセル2
からなる。入力SAMセル2は図2の入力ビット数分
(ISB)縦に並べて用意されるのだが、図4ではそれ
を省略して一つだけ代表して表示している。なお入力ポ
インタセル1は、シフトレジスタを構成するためのフリ
ップフロップ(FF)である。Further, the structure of this processor element is
It can be roughly written as shown in FIG. One processor element of the input SAM unit is composed of one input pointer cell 1 in FIG. 4 and a plurality of input SAM cells 2 arranged vertically.
Consists of. The input SAM cells 2 are prepared by vertically arranging them by the number of input bits (ISB) in FIG. 2, but in FIG. 4, they are omitted and only one is shown as a representative. The input pointer cell 1 is a flip-flop (FF) for forming a shift register.
【0012】データメモリA部の一つのプロセッサエレ
メント分は、図4におけるメモリAセル3を、図2のM
ABのビット数分用意されて縦にならんでいるのだが、
図4ではそれを省略して一つだけ代表して表示してい
る。For one processor element of the data memory A section, the memory A cell 3 in FIG.
The number of bits of AB is prepared and arranged vertically,
In FIG. 4, it is omitted and only one is shown as a representative.
【0013】ALUアレイ部の一つのプロセッサエレメ
ント分は、図4におけるALUセル4である。ここでA
LUセル4中の本当のALU部分は1ビットALUであ
り、全加算器(フルアダー)程度のものである。またA
LUセル4中のセレクタ(SEL)は、1ビットALU
の入力選択のためのものであり、図中の複数のX印で示
すバスとの交点のうちの一つのバスからのデータを選択
する。なおFFはフリップフロップ(1ビットレジス
タ)である。One processor element of the ALU array section is the ALU cell 4 in FIG. Where A
The real ALU portion in the LU cell 4 is a 1-bit ALU, which is about the full adder. Also A
The selector (SEL) in the LU cell 4 is a 1-bit ALU
Input selection, and selects data from one of the intersections with the buses indicated by a plurality of X's in the figure. FF is a flip-flop (1-bit register).
【0014】データメモリB部の一つのプロセッサエレ
メント分は、図4におけるメモリBセル5を、図2のM
BBのビット数分用意されて縦に並んでいるのだが、図
4ではそれを省略して一つだけ代表して表示している。
なおメモリAセル3とメモリBセル5は同じもので良
い。For one processor element of the data memory B section, the memory B cell 5 in FIG.
Although the number of bits of BB is prepared and arranged vertically, it is omitted in FIG. 4 and only one is displayed as a representative.
The memory A cell 3 and the memory B cell 5 may be the same.
【0015】出力SAM部の一つのプロセッサエレメン
ト分は、図4における一つの出力ポインタセル7と縦に
並んだ複数の出力SAMセル6からなる。出力SAMセ
ル6は図2の出力ビット数分(OSB)縦に並べて用意
されるのだが、図4ではそれを省略して一つだけ代表し
て表示している。なお出力ポインタセル7は、シフトレ
ジスタを構成するためのフリップフロップ(FF)であ
る。なお出力ポインタセル7と出力SAMセル6は、そ
れぞれ入力ポインタセル1と入力SAMセル2と同様の
もので良い。One processor element of the output SAM section is composed of one output pointer cell 7 in FIG. 4 and a plurality of output SAM cells 6 arranged vertically. The output SAM cells 6 are prepared by arranging them vertically for the number of output bits (OSB) in FIG. 2, but in FIG. 4, they are omitted and only one is displayed. The output pointer cell 7 is a flip-flop (FF) for forming a shift register. The output pointer cell 7 and the output SAM cell 6 may be the same as the input pointer cell 1 and the input SAM cell 2, respectively.
【0016】入力SAM読みだし信号(IR)、メモリ
アクセス信号(AA及びAB)、出力SAM書き込み信
号(OW)などは、メモリセルのワード線であり、アド
レスデコードがされているものとする。またリードモデ
ィファイライトのために、読み出しのための信号はサイ
クルの前半、書き込みのための信号はサイクルの後半の
タイミングで発生される。The input SAM read signal (IR), the memory access signal (AA and AB), the output SAM write signal (OW), etc. are the word lines of the memory cell and are assumed to be address-decoded. For read-modify-write, a signal for reading is generated in the first half of the cycle and a signal for writing is generated in the latter half of the cycle.
【0017】なお図4において、セルを縦に通過する接
続線(ビット線)は、縦に並ぶ回路要素を同様に接続し
ながら通過するものとする。また横方向の接続線のう
ち、メモリのワード線および入力データバスは、横に並
ぶ回路要素を同様に接続しながら通過し、ポインタの線
は隣に並ぶ対応する回路要素同志手をつなぐように接続
される。In FIG. 4, connection lines (bit lines) that vertically pass through the cells are assumed to pass while vertically connecting circuit elements. Of the horizontal connection lines, the word line of the memory and the input data bus pass while connecting the circuit elements arranged side by side in the same manner, and the pointer line connects the corresponding circuit elements arranged next to each other. Connected.
【0018】次にこのプロセッサの動作を、図2、図4
を使って説明する。入力信号は入力SAM部に導かれ
る。入力ポインタセル1を構成するフリップフロップは
シフトレジスタを構成しており、このシフトレジスタが
リレーする論理“H”を1つ立てた1ビット信号即ち入
力ポインタ信号(IP)が作られ、その論理“H”で指
定されたプロセッサエレメントの入力SAMセル2に入
力データ(IN)が書き込まれる。入力データバス及び
入力SAMセル2はそれぞれISBビットだけあるが、
図4では1ビット分だけを示している。Next, the operation of this processor will be described with reference to FIGS.
Use to explain. The input signal is guided to the input SAM section. The flip-flops that compose the input pointer cell 1 compose a shift register, and a 1-bit signal, that is, an input pointer signal (IP), which is one logic "H" relayed by this shift register, is generated and the logic "H" is generated. Input data (IN) is written to the input SAM cell 2 of the processor element designated by H ″. The input data bus and the input SAM cell 2 each have only ISB bits,
In FIG. 4, only one bit is shown.
【0019】入力データは、映像信号の一水平走査期間
ごとに、入力ポインタにより入力SAM部の左端のプロ
セッサエレメントのSAMから順に右方向のプロセッサ
エレメントのSAMに記憶していくことが出来、並んだ
プロセッサエレメント数が映像信号の一水平走査期間の
画素数分(H)であるので、入力映像信号のデータレー
トに合わせたクロックで、一水平走査期間右方向へSA
M書き込みを続け、一水平走査期間分の入力データを入
力SAM部に蓄積できる。このような入力動作は、水平
走査期間毎に繰り返される。The input data can be stored in the SAM of the processor element in the right direction in order from the SAM of the leftmost processor element of the input SAM section by the input pointer every horizontal scanning period of the video signal. Since the number of processor elements is equal to the number of pixels (H) in one horizontal scanning period of the video signal, the clock SA corresponding to the data rate of the input video signal causes SA to the right in the one horizontal scanning period.
M writing can be continued and input data for one horizontal scanning period can be stored in the input SAM section. Such an input operation is repeated every horizontal scanning period.
【0020】プログラム制御部は、入力SAM部、デー
タメモリA部、ALUアレイ部、データメモリB部、出
力SAM部を以下のようにSIMD制御して、処理を実
行する。The program control unit controls the input SAM unit, the data memory A unit, the ALU array unit, the data memory B unit, and the output SAM unit by SIMD as follows, and executes the processing.
【0021】一水平走査期間分の入力SAM部に蓄積さ
れた入力データは、次の一水平走査期間において、必要
に応じてプログラム制御部の制御のもとに入力SAM部
からデータメモリB部へ移され、演算処理に使われる。
この動作はプログラムで入力SAM部の必要なビットの
記憶内容を入力SAM読みだし信号(IR)によりアク
セスしては、転送先のデータメモリB部の所定のメモリ
セルへメモリアクセス信号(AB)を出して書き込んで
いくことにより実現する。The input data accumulated in the input SAM section for one horizontal scanning period is transferred from the input SAM section to the data memory B section under the control of the program control section as needed in the next one horizontal scanning period. Moved and used for arithmetic processing.
In this operation, the stored contents of the necessary bits of the input SAM section are accessed by the program by the input SAM read signal (IR), and then the memory access signal (AB) is sent to the predetermined memory cell of the data memory B section of the transfer destination. It is realized by taking out and writing.
【0022】ここで入力SAM読みだし信号(IR)と
メモリアクセス信号(AB)はワード線であり、それぞ
れ複数あるが、これはアドレスデコーダでデコードされ
ている。またこれらワード線はリードモディファイライ
トのために、読み出しのための信号はサイクルの前半、
書き込みのための信号はサイクルの後半のタイミングで
発生される。このデータ転送は縦方向のビット線を経由
して1ビット1ビット行われる。なおこのデータ移動に
際してALUで処理することは何もないが、ALUセル
4を通るようになっており、その際ALU出力制御信号
(BB)が所定のタイミングで発生されている。The input SAM read signal (IR) and the memory access signal (AB) are word lines, and there are a plurality of word lines, which are decoded by the address decoder. In addition, these word lines are for read-modify-write, so the signals for reading are in the first half of the cycle.
The signal for writing is generated at the timing of the latter half of the cycle. This data transfer is performed bit by bit through the vertical bit lines. There is nothing to be processed by the ALU at the time of this data movement, but it is designed to pass through the ALU cell 4, and at that time, the ALU output control signal (BB) is generated at a predetermined timing.
【0023】入力SAM部の各入力SAMセル2からの
読みだし信号(IR)とデータメモリA部の各メモリセ
ルへメモリアクセス信号(AA)は同じアドレス空間に
あり、メモリの同じROWデコーダでデコードされて、
ワード線として与えられるものである。The read signal (IR) from each input SAM cell 2 of the input SAM section and the memory access signal (AA) to each memory cell of the data memory A section have the same address space and are decoded by the same ROW decoder of the memory. Has been
It is given as a word line.
【0024】データの演算処理にあたり、必要に応じて
データメモリA部とデータメモリB部の間では、所定の
メモリセルへメモリアクセス信号(AA、AB)を出し
て読み出し或いは書き込みを行い、データを移動でき
る。これも入力SAM部からデータメモリB部へのデー
タ転送同様リードモディファイライトで、縦方向のビッ
ト線を経由して1ビット1ビット行われる。またこの時
もデータ移動に際してALUで処理することは何もない
が、ALUセル4を通るようになっており、その際AL
U出力制御信号(BA或いはBB)が所定のタイミング
で発生される。In the data arithmetic processing, if necessary, between the data memory A section and the data memory B section, a memory access signal (AA, AB) is issued to a predetermined memory cell to read or write the data, and the data is stored. You can move. This is also read-modify-write as in the data transfer from the input SAM section to the data memory B section, and one bit per bit is performed via the vertical bit line. Also at this time, there is nothing to be processed by the ALU at the time of data movement, but it is configured to pass through the ALU cell 4, and at that time
The U output control signal (BA or BB) is generated at a predetermined timing.
【0025】よって、データメモリA部とデータメモリ
B部には、過去に上述のようにして書き込まれた入力デ
ータや演算途中のデータが記憶されている。それらのデ
ータ或いはALUセル4中の1ビットレジスタ(FF)
に記憶したデータを用いて、それぞれのサイクルにおい
て必要なALUでのビット演算処理ができる。Therefore, in the data memory A section and the data memory B section, the input data written in the past as described above and the data in the middle of calculation are stored. 1-bit register (FF) in those data or ALU cell 4
Using the data stored in, the bit arithmetic processing in the ALU required in each cycle can be performed.
【0026】このALUセル4での演算動作は、ALU
制御信号(ALU−CONT)によりプログラムから指
定される。ALUセル4で演算した結果は、ALUセル
4中の1ビットレジスタに記憶されるか、必要に応じて
再びデータメモリA部或いはデータメモリB部に書き込
まれる。The arithmetic operation in this ALU cell 4 is
It is specified by the program by the control signal (ALU-CONT). The result calculated by the ALU cell 4 is stored in the 1-bit register in the ALU cell 4, or is written again in the data memory A section or the data memory B section as needed.
【0027】例えばデータメモリA部のあるビットのメ
モリAセル3のデータとデータメモリB部のあるビット
のメモリBセル5のデータを加算してデータメモリB部
の今読み出したビットのメモリBセル5に加算結果を書
き込む場合は以下のようになる。データメモリA部の所
定のビットのメモリAセル3へ読み出し信号(AA)、
またデータメモリB部の所定のビットのメモリBセル5
へは読み出し信号(AB)をサイクルの前半に出す。デ
ータメモリA部から読み出されたデータとデータメモリ
B部から読み出されたデータは、ALUアレイ部のAL
Uで演算処理される。ALUからの出力は、データメモ
リB部の先のビットのメモリBセル5へ書き込み信号
(AB)をサイクルの後半に出してそこに書き込む。そ
の際にALU出力制御信号(BB)が所定のタイミング
で与えられる。For example, the data of the memory A cell 3 of a certain bit in the data memory A section and the data of the memory B cell 5 of a certain bit of the data memory B section are added to each other, and the memory B cell of the bit just read in the data memory B section. The case of writing the addition result to 5 is as follows. A read signal (AA) to the memory A cell 3 of a predetermined bit in the data memory A section,
In addition, the memory B cell 5 of a predetermined bit in the data memory B section
A read signal (AB) is issued in the first half of the cycle. The data read from the data memory A section and the data read from the data memory B section are AL of the ALU array section.
U is calculated. The output from the ALU outputs a write signal (AB) to the memory B cell 5 of the previous bit of the data memory B section in the latter half of the cycle and writes it there. At that time, the ALU output control signal (BB) is given at a predetermined timing.
【0028】このようにして上下に存在するデータメモ
リA部とデータメモリB部から、プログラムに応じてデ
ータを読み出しては、ALUアレイ部で必要な算術演算
或いは論理演算を施し、再びデータメモリA部或いはデ
ータメモリB部の所定のアドレスに書き込むことが出来
る。この演算処理は全てビット処理であり、1ビット1
ビット処理を進める。In this way, the data is read from the data memory A and the data memory B located above and below according to the program, and the necessary arithmetic operation or logical operation is performed in the ALU array section, and the data memory A is read again. Part or data memory B part can be written to a predetermined address. This arithmetic processing is all bit processing, 1 bit 1
Advance bit processing.
【0029】一水平走査期間分の演算処理が終わると、
その水平走査期間のうちに、プログラムの最後の部分で
その水平走査期間分の出力データを出力SAM部に移す
必要がある。今出力すべきデータがデータメモリA部に
あるとする時、所定のメモリセルへメモリアクセス信号
(AA)をサイクルの前半に出して読みだしを行い、ま
た出力SAM部の所定のビットの出力SAMセル6にデ
ータ転送されるように、その出力SAMセル6にサイク
ルの後半に書き込み信号(OW)が発生される。When the calculation processing for one horizontal scanning period is completed,
During the horizontal scanning period, it is necessary to transfer the output data for the horizontal scanning period to the output SAM section at the last part of the program. If the data to be output is present in the data memory A section, the memory access signal (AA) is output to a predetermined memory cell in the first half of the cycle to read the data, and the output SAM of a predetermined bit of the output SAM section is also used. A write signal (OW) is generated in the output SAM cell 6 in the latter half of the cycle so that the data is transferred to the cell 6.
【0030】データは縦方向のビット線を経して1ビッ
ト1ビットデータ転送される。またこの時もデータ移動
に際してALUで処理することは何もないが、ALUセ
ル4を通るようになっており、その際ALU出力制御信
号(BB)が所定のタイミングで発生される。Data is transferred one bit at a time through one bit line in the vertical direction. Also at this time, there is nothing to be processed by the ALU at the time of data movement, but the data is passed through the ALU cell 4, and at that time, the ALU output control signal (BB) is generated at a predetermined timing.
【0031】出力SAM部の各出力SAMセル6への書
き込み信号(OW)とデータメモリB部の各メモリセル
へメモリアクセス信号(AB)は同じアドレス空間にあ
り、メモリの同じROWデコーダでデコードされて、ワ
ード線として与えられるものである。The write signal (OW) to each output SAM cell 6 of the output SAM section and the memory access signal (AB) to each memory cell of the data memory B section have the same address space and are decoded by the same ROW decoder of the memory. Given as a word line.
【0032】以上のように、一水平走査期間の時間のう
ちに、入力SAM部に蓄積された入力データの読み出
し、必要な演算処理、必要なデータ移動、そして出力S
AM部への出力データの書き込みまでが、ビットを単位
とするSIMD制御プログラムで制御される。このプロ
グラム処理は水平走査期間を単位として繰り返される。As described above, during the time of one horizontal scanning period, reading of the input data accumulated in the input SAM section, necessary arithmetic processing, necessary data movement, and output S
The writing of output data to the AM section is controlled by the SIMD control program in units of bits. This program process is repeated in units of horizontal scanning period.
【0033】このプログラム処理が終わって出力SAM
部に移された出力データは、更に次の水平走査期間に、
以下のように出力SAM部から出力される。Output SAM after this program processing is completed
The output data transferred to the section is further in the next horizontal scanning period,
It is output from the output SAM unit as follows.
【0034】出力信号は出力SAM部から出力データバ
スへ導かれ、このプロセッサの外へ出力される。出力ポ
インタセル7を構成するフリップフロップはシフトレジ
スタを構成し、このシフトレジスタでリレーする論理
“H”を1つ立てた1ビット信号即ち出力ポインタ信号
(OP)が作られ、その論理“H”で指定されたプロセ
ッサエレメントの出力SAMセル6から出力データが出
力データバスに読み出され、出力データ(OUT)とな
る。出力データバス及び出力SAMセル6は、それぞれ
OSBビットだけあるが、図4では1ビット分だけを示
している。The output signal is led from the output SAM section to the output data bus and output to the outside of this processor. The flip-flops that form the output pointer cell 7 form a shift register, and a 1-bit signal, ie, an output pointer signal (OP), which is one logic "H" relayed by this shift register, is generated, and the logic "H" is generated. The output data is read from the output SAM cell 6 of the processor element designated by the above to the output data bus and becomes the output data (OUT). The output data bus and the output SAM cell 6 have only OSB bits, but only one bit is shown in FIG.
【0035】出力データは映像信号の一水平走査期間ご
とに、出力ポインタにより、出力SAM部の左端のプロ
セッサエレメントのSAMから順に右方向のプロセッサ
エレメントのSAMへ場所を移しながら読み出していく
ことが出来、並んだプロセッサエレメント数が映像信号
の一水平走査期間の画素数分(H)であるので、出力映
像信号のデータレートに合わせたクロックで、一水平走
査期間分の出力データを出力SAM部から出力できる。
このような出力動作は水平走査期間毎に繰り返される。The output data can be read every horizontal scanning period of the video signal by the output pointer while sequentially shifting the location from the SAM of the leftmost processor element of the output SAM section to the SAM of the right processor element. Since the number of arranged processor elements is equal to the number of pixels (H) in one horizontal scanning period of the video signal, the output data for one horizontal scanning period is output from the SAM unit with a clock adapted to the data rate of the output video signal. Can be output.
Such an output operation is repeated every horizontal scanning period.
【0036】1.入力データの入力SAM部への書き込
みによる入力動作。 2.プログラム制御部のSIMD制御による、入力SA
M部からの入力データ読み出し、データメモリA部、A
LUアレイ部、データメモリB部による演算処理の実
行、データ転送、そして出力SAM部への出力データ書
き込み。 3.出力データの出力SAM部からの読み出しによる出
力動作。 の3つの動作は、映像信号の一水平走査期間を単位とす
るパイプライン動作になっており、ひとつの水平走査期
間の入力データに注目すれば、それぞれの動作は一水平
走査期間の時間づつずれた形で実行されるが、3つの動
作は連続して同時に並行して進行できる。1. Input operation by writing input data to the input SAM section. 2. Input SA by SIMD control of program control unit
Input data read from M section, data memory A section, A
Execution of arithmetic processing by the LU array unit and the data memory B unit, data transfer, and writing of output data to the output SAM unit. 3. Output operation by reading output data from the output SAM section. The above three operations are pipeline operations in units of one horizontal scanning period of the video signal, and if attention is paid to the input data of one horizontal scanning period, the respective operations are shifted by the time of one horizontal scanning period. However, the three operations can proceed sequentially and concurrently.
【0037】[0037]
【発明が解決しようとする課題】従来例の構成では、プ
ログラム制御部のSIMD制御による入力SAM部、デ
ータメモリA部、ALUアレイ部、データメモリB部、
出力SAM部による演算処理の実行において、ALU入
力のデータソースは、入力SAMセル2、データメモリ
A部、データメモリB部、ALUセル4の中の1ビット
レジスタのうちの2つである。またALU出力のデータ
ディスティネーションは、出力SAMセル6、データメ
モリA部、データメモリB部、ALUセル4の中の1ビ
ットレジスタのうちの1つである。In the configuration of the conventional example, the input SAM section by the SIMD control of the program control section, the data memory A section, the ALU array section, the data memory B section,
In the execution of the arithmetic processing by the output SAM unit, two data sources of the ALU input are the input SAM cell 2, the data memory A unit, the data memory B unit, and the 1-bit register in the ALU cell 4. The data destination of the ALU output is one of the output SAM cell 6, the data memory A section, the data memory B section, and the 1-bit register in the ALU cell 4.
【0038】このプロセッサにおける命令サイクルは、
2つのデータソースからデータを読み出してALUで演
算処理し、データディスティネーションに書き込むまで
である。すなわち、全ての動作ではデータは基本的にA
LUを通ることになっており、また全ての処理サイクル
はリードモディファイライト動作になっている。The instruction cycle in this processor is
That is, data is read from two data sources, arithmetic processing is performed by the ALU, and writing is performed in the data destination. That is, data is basically A in all operations.
It is supposed to pass through the LU, and all the processing cycles are read-modify-write operations.
【0039】一般にこのプロセッサの入力SAMセル
2、メモリAセル3、メモリBセル5、出力SAMセル
6は、DRAM構造で作るのでただでさえアクセスタイ
ムが遅い。その上リードとライトの両動作を1サイクル
中に実行するリードモディファイライト動作をさせてい
るので、処理サイクルは遅い動作になっている。Generally, the input SAM cell 2, the memory A cell 3, the memory B cell 5 and the output SAM cell 6 of this processor are made of a DRAM structure, so that the access time is slow even by themselves. In addition, since the read-modify-write operation is performed in which both read and write operations are executed in one cycle, the processing cycle is slow.
【0040】この図2のアーキテクチャのプロセッサ
は、一水平走査期間を単位としてプログラム処理するの
で、データレートの高い映像信号に対しても、リアルタ
イム処理で高いプログラマビリティを実現している。そ
れを実現したのは、処理単位をビット単位として一つの
プロセッサエレメントを極端に小型化し、それによって
非常に多数のプロセッサエレメントの並列化をしたから
である。Since the processor having the architecture shown in FIG. 2 performs program processing in units of one horizontal scanning period, it realizes high programmability by real-time processing even for a video signal having a high data rate. This is realized because one processor element is extremely miniaturized with a processing unit being a bit unit, and thus a large number of processor elements are parallelized.
【0041】このアーキテクチャのプロセッサの処理性
能は、例えば図2の縦方向の長さを長く、即ち、入力S
AM部、データメモリA部、データメモリB部、出力S
AM部のメモリサイズを大きくしても、それぞれのデー
タメモリのアドレス空間が広がるだけで、ワーキングメ
モリが増えるだけである。また図2の横方向の長さを長
く、即ちプロセッサエレメントの並列数を増やしても、
このプロセッサエレメントの並列数は適用する映像信号
の一水平走査期間の画素数に対応させて使うので意味が
ないThe processing performance of the processor of this architecture is such that the length in the vertical direction of FIG.
AM section, data memory A section, data memory B section, output S
Even if the memory size of the AM section is increased, the address space of each data memory is expanded and the working memory is increased. In addition, even if the horizontal length of FIG. 2 is increased, that is, the number of parallel processor elements is increased,
This parallel number of processor elements is meaningless because it is used in correspondence with the number of pixels in one horizontal scanning period of the applied video signal.
【0042】このアーキテクチャのプロセッサの処理性
能は、命令サイクルの高速化、或いはALUの並列化、
或いはプロセッサ全体の並列化によるしかない。図2の
プロセッサの命令サイクルの高速化のために、図4の1
トランジスタによるメモリAセル3及びメモリBセル5
(これらは同じ)の代わりに、3トランジスタによる図
5のような出力の大きなメモリセルを使い、これによっ
てセンスアンプにかかっていた負担を減らすことにより
高速化を計るというようなことはあった。しかしなおリ
ードモディファイライト動作であった。図5において、
Wは書き込みアクセス信号線(ライトワード線)、また
Rは読み出しアクセス信号線(リードワード線)であ
る。The processing performance of the processor of this architecture is such that the instruction cycle is accelerated or the ALU is parallelized.
Or it could only be done by parallelizing the entire processor. In order to speed up the instruction cycle of the processor of FIG.
Memory A cell 3 and memory B cell 5 by transistors
Instead of (these are the same), a memory cell having a large output as shown in FIG. 5 by using three transistors was used, and thereby the load on the sense amplifier was reduced to speed up the operation. However, it was still a read modify write operation. In FIG.
W is a write access signal line (write word line), and R is a read access signal line (read word line).
【0043】本件は、このプロセッサにおける命令サイ
クルが遅いという従来の欠点を解決しようとするもので
ある。本件では、メモリの読み出し、書き込みを、別サ
イクルとして命令サイクルの高速化を計っている。This case is intended to solve the conventional drawback that the instruction cycle in this processor is slow. In this case, the memory read and write are set as separate cycles to speed up the instruction cycle.
【0044】[0044]
【課題を解決するための手段】本発明による第1の手段
は、リニアアレイ型の並列DSPプロセッサにおいて、
データメモリ部のメモリセルに接続される読み出しビッ
ト線及び書き込みビット線とを有することを特徴とする
リニアアレイ型の並列DSPプロセッサである。A first means according to the present invention is a linear array type parallel DSP processor,
A linear array parallel DSP processor having a read bit line and a write bit line connected to a memory cell of a data memory unit.
【0045】本発明による第2の手段は、第1の手段記
載のリニアアレイ型の並列DSPプロセッサにおいて、
少なくともSAM(シリアルアクセスメモリ)部、デー
タメモリ部、ALUアレイ部よりなり、上記データメモ
リ部とALUアレイ部の間のデータのやり取りであるメ
モリ読み出し又は書き込みを独立して行うことを特徴と
するリニアアレイ型の並列DSPプロセッサである。A second means according to the present invention is the linear array parallel DSP processor according to the first means,
At least a SAM (serial access memory) unit, a data memory unit, and an ALU array unit, and a memory read or write that is data exchange between the data memory unit and the ALU array unit are independently performed. It is an array type parallel DSP processor.
【0046】本発明による第3の手段は、第2の手段記
載のリニアアレイ型の並列DSPプロセッサにおいて、
データメモリ部のメモリセルの読み出し線をALU部の
入力に接続するとともに、上記データメモリ部のメモリ
セルの書き込み線をALU部の出力に接続することを特
徴とするリニアアレイ型の並列DSPプロセッサであ
る。A third means according to the present invention is the linear array parallel DSP processor according to the second means.
A linear array parallel DSP processor characterized in that a read line of a memory cell of a data memory unit is connected to an input of an ALU unit and a write line of a memory cell of the data memory unit is connected to an output of an ALU unit. is there.
【0047】本発明による第4の手段は、第3の手段記
載のリニアアレイ型の並列DSPプロセッサにおいて、
上記プロセツサは全てのプロセツサが1つのプログラム
により連動して動作をするSIMD制御であることを特
徴とするリニアアレイ型の並列DSPプロセッサであ
る。A fourth means according to the present invention is the linear array parallel DSP processor according to the third means,
The processor is a linear array type parallel DSP processor characterized in that all the processors are SIMD control in which one processor operates in synchronization with one program.
【0048】本発明による第5の手段は、第4の手段記
載のリニアアレイ型の並列DSPプロセッサにおいて、
上記プロセツサは映像信号の水平解像度に相当する個数
のプロセツサよりなり、全てのプロセツサが1期間にお
いて1水平走査線の情報を処理することを特徴とするリ
ニアアレイ型の並列DSPプロセッサである。A fifth means according to the present invention is the linear array parallel DSP processor according to the fourth means,
The processor is a linear array type parallel DSP processor characterized by comprising a number of processors corresponding to the horizontal resolution of a video signal, and all the processors process information of one horizontal scanning line in one period.
【0049】[0049]
【作用】これによれば、テレビなどの映像信号をソフト
ウェアプログラムでディジタル信号処理するための、S
IMD制御のプログラマブルなリニアアレイ型プロセッ
サの構成および制御に関するもので、特に、SIMD制
御のリニアアレイ型プロセッサであって、その構成要素
である各プロセッサエレメントのデータメモリのビット
線を、書き込みと読み出しの各々に専用に持つようにし
て、各プロセッサエレメントのALUとデータメモリの
間の書き込み及び読み出しをパイプライン化している。
これにより命令サイクルが高速化し、また同一命令サイ
クル中に書き込みと次の処理のための読み出しを同時に
実行できる。According to this, the S signal for digital signal processing of the video signal of the television etc. by the software program is provided.
The present invention relates to a configuration and control of a programmable linear array type processor of IMD control, and particularly to a linear array type processor of SIMD control, in which a bit line of a data memory of each processor element, which is a component thereof, is written and read. Write and read between the ALU and the data memory of each processor element are pipelined so that each processor element has a dedicated one.
This speeds up the instruction cycle, and during the same instruction cycle, writing and reading for the next processing can be executed simultaneously.
【0050】[0050]
【実施例】図1が実施例の一つのプロセッサエレメント
のモデルである。従来例が図2を基本にして、図4のモ
デルにより実現していたのに対して、実施例では図2を
基本にして、図1のモデルにより実現する。FIG. 1 is a model of one processor element of the embodiment. While the conventional example is realized by the model of FIG. 4 based on FIG. 2, the embodiment is realized by the model of FIG. 1 based on FIG.
【0051】図4のモデルと図1のモデルの相違は、メ
モリAセル3及びメモリBセル5(これらは同じ)に、
図5と同様な3トランジスタのメモリセルを使用し、か
つその書き込みビット線と読み出しビット線を専用化
し、個別に持つようにしたことである。The difference between the model of FIG. 4 and the model of FIG. 1 is that the memory A cell 3 and the memory B cell 5 (these are the same) are
The memory cell of 3 transistors similar to that of FIG. 5 is used, and the write bit line and the read bit line thereof are dedicated and provided individually.
【0052】また実施例においては、もはやリードモデ
ィファイライト動作はさせない。即ち一つのビット線の
上では、各サイクルでは読み出し或いは書き込みしかし
ない。しかし、従来に比べてビット線を増やしているの
で、同じだけのデータ転送が可能になっている。ひとつ
のデータ転送にかかわるリード動作とライト動作は、A
LUセル4中の1ビットレジスタ(FF)によってパイ
プライン動作になっていて、そのふたつの動作は連続す
るふたつのサイクルに分かれる。In the embodiment, the read modify write operation is no longer performed. That is, on one bit line, only reading or writing is performed in each cycle. However, since the number of bit lines is increased compared to the conventional one, the same data transfer is possible. Read operation and write operation related to one data transfer are
The 1-bit register (FF) in the LU cell 4 performs a pipeline operation, and the two operations are divided into two continuous cycles.
【0053】すなわち、図4で入力SAM部からデータ
メモリA部を経由してALUアレイ部に至る各プロセッ
サエレメントごとに1本の縦のビット線は、図1におい
て入力SAM部からデータメモリA部を経由してALU
アレイ部に至る読み出しビット線と、ALUアレイ部か
らデータメモリA部に至る書き込みビット線の、各プロ
セッサエレメントごとに2本の縦のビット線になってい
る。この時、メモリAセル3は、その書き込みのための
ひとつのトランジスタと、読み込みのための2つのトラ
ンジスタは、図5と違って、それぞれ別の書き込みビッ
ト線と読み出しビット線に接続されている。That is, one vertical bit line for each processor element from the input SAM section to the ALU array section via the data memory A section in FIG. Via ALU
There are two vertical bit lines for each processor element, a read bit line reaching the array portion and a write bit line extending from the ALU array portion to the data memory A portion. At this time, in the memory A cell 3, one transistor for writing and two transistors for reading are connected to different write bit lines and read bit lines, respectively, unlike FIG.
【0054】また図4でALUアレイ部からデータメモ
リB部を経由して出力SAM部へ至る各プロセッサエレ
メントごとに1本の縦のビット線は、図1においてAL
Uアレイ部からデータメモリB部を経由して出力SAM
部へ至る書き込みビット線と、データメモリB部からA
LUアレイ部に至る読み出しビット線の、各プロセッサ
エレメントごとに2本の縦のビット線になっている。こ
の時、メモリBセル5(メモリAセル3と同じ)は、そ
の書き込みのためのひとつのトランジスタと、読み込み
のための2つのトランジスタは、図5と違って、それぞ
れ別の書き込みビット線と読み出しビット線に接続され
ている。Further, one vertical bit line for each processor element from the ALU array section to the output SAM section via the data memory B section in FIG.
Output SAM from U array section via data memory B section
To the write bit line and the data memory from the B section to the A section
There are two vertical bit lines for each processor element of the read bit lines reaching the LU array section. At this time, the memory B cell 5 (same as the memory A cell 3) has one transistor for writing and two transistors for reading, which are different from those in FIG. It is connected to the bit line.
【0055】従って図2のプロセッサでの 1.入力データの入力SAM部への書き込みによる入力
動作。 2.プログラム制御部のSIMD制御による入力SAM
部、データメモリA部、ALUアレイ部、データメモリ
B部、出力SAM部による演算処理の実行。 3.出力データの出力SAM部からの読み出しによる出
力動作。 の3つの動作のうち、1.及び3.は変更ない。以下
2.の動作を説明する。Therefore, in the processor of FIG. Input operation by writing input data to the input SAM section. 2. Input SAM by SIMD control of program control unit
Of arithmetic processing by the unit, the data memory A unit, the ALU array unit, the data memory B unit, and the output SAM unit. 3. Output operation by reading output data from the output SAM section. Of the three operations of 1. And 3. Does not change. Below 2. The operation of will be described.
【0056】一水平走査期間分の入力SAM部に蓄積さ
れた入力データは、次の一水平走査期間において、必要
に応じてプログラム制御部の制御のもとに入力SAM部
からデータメモリB部へ移される。この動作は、プログ
ラムで入力SAM部の必要なビットを入力SAM読みだ
し信号(IR)によりアクセスしては、そのビットの入
力SAMセル2からこのプロセッサの上半分の読み出し
ビット線を通ってALUセル4の1ビットレジスタに一
旦記憶し、次のサイクルでALUを経由して、このプロ
セッサの下半分の書き込みビット線を通ってデータメモ
リB部の所定のビットメモリBセル5に届け、そのメモ
リBセル5へメモリ書き込み信号(WB)を出して書き
込むことにより実現する。The input data accumulated in the input SAM section for one horizontal scanning period is transferred from the input SAM section to the data memory B section under the control of the program control section as needed in the next one horizontal scanning period. Be transferred. This operation is performed by accessing a required bit of the input SAM section by an input SAM read signal (IR) by a program, and from the input SAM cell 2 of the bit through the read bit line of the upper half of the processor to the ALU cell. It is temporarily stored in the 1-bit register 4 and is delivered to the predetermined bit memory B cell 5 of the data memory B section through the write bit line in the lower half of this processor via the ALU in the next cycle. It is realized by issuing a memory write signal (WB) to the cell 5 and writing.
【0057】このデータ転送は1ビット1ビット行われ
る。なおこのデータ移動に際してALUで処理すること
は何もないが、ALUセル4を通るようになっており、
その際ALU出力制御信号(BB)が必要なタイミング
で発生される。ひとつのデータ転送が2サイクルに見え
るが、パイプラインなので多数のビットを連続して転送
するような場合はサイクルは増えない。This data transfer is carried out bit by bit. Although there is nothing to be processed by the ALU at the time of this data movement, it is designed to pass through the ALU cell 4,
At that time, the ALU output control signal (BB) is generated at a necessary timing. Although one data transfer looks like two cycles, the cycle does not increase when a large number of bits are transferred continuously because it is a pipeline.
【0058】また、必要に応じてデータメモリA部とデ
ータメモリB部の間で、プログラムで所定のメモリセル
へ書き込み信号(RA或いはRB)と読み出し信号(W
A或いはWB)を出してデータを移動できる。If necessary, a write signal (RA or RB) and a read signal (W) are written between the data memory A section and the data memory B section to a predetermined memory cell by a program.
You can move the data by issuing A or WB).
【0059】例えばデータメモリA部からデータメモリ
B部へのデータ転送では、まず移すべきデータメモリA
部の所定のビットのメモリAセル3へ読み出し信号(R
A)を出し、そのビットのメモリAセル3からこのプロ
セッサの上半分の読み出しビット線を通ってALUセル
4の1ビットレジスタに一旦記憶し、次のサイクルでA
LUを経由して、このプロセッサの下半分の書き込みビ
ット線を通ってデータメモリB部の所定のビットメモリ
Bセル5に届け、そのビットのメモリBセル5へ書き込
み信号(WB)を出して書き込む。For example, in the data transfer from the data memory A section to the data memory B section, first, the data memory A to be moved.
A read signal (R
A) from the memory A cell 3 of the bit, and once stored in the 1-bit register of the ALU cell 4 through the read bit line in the upper half of this processor, and then in the next cycle
The data is sent to a predetermined bit memory B cell 5 of the data memory B section through the write bit line in the lower half of this processor via the LU, and a write signal (WB) is issued to the memory B cell 5 of the bit to write. .
【0060】これも1ビット1ビット行う。なおこのデ
ータ移動に際してALUで処理することは何もないが、
ALUセル4を通るようになっており、その際ALU出
力制御信号(BB)が必要なタイミングで発生される。
ひとつのデータ転送が2サイクルに見えるが、パイプラ
インなので多数のビットを連続して転送するような場合
はサイクルは増えない。This is also performed for each one bit. Although there is nothing to process with the ALU when moving this data,
It passes through the ALU cell 4, and at that time, the ALU output control signal (BB) is generated at a necessary timing.
Although one data transfer looks like two cycles, the cycle does not increase when a large number of bits are transferred continuously because it is a pipeline.
【0061】よって、データメモリA部とデータメモリ
B部には、過去に上述のようにして書き込まれた入力デ
ータや演算途中のデータが記憶されている。それらのデ
ータ或いはALUセル4中の1ビットレジスタ(FF)
のデータを用いて、必要なALUでのビット演算処理が
できる。このALUセル4での演算動作は、ALU制御
信号(ALU−CONT)により指定される。ALUセ
ル4で演算した結果は、ALUセル4中の1ビットレジ
スタに記憶するか、必要に応じてデータメモリA部或い
はデータメモリB部に書き込まれる。Therefore, in the data memory A section and the data memory B section, the input data written in the past as described above and the data in the middle of calculation are stored. 1-bit register (FF) in those data or ALU cell 4
It is possible to perform the necessary bit arithmetic processing in the ALU using the data of. The arithmetic operation in this ALU cell 4 is designated by the ALU control signal (ALU-CONT). The result calculated by the ALU cell 4 is stored in the 1-bit register in the ALU cell 4 or written in the data memory A section or the data memory B section as needed.
【0062】例えばデータメモリA部のあるビットのメ
モリAセル3のデータとデータメモリB部のあるビット
のメモリBセル5のデータを加算してデータメモリB部
の今読み出したビットのメモリBセル5に加算結果を書
き込む場合は以下のようになる。まず最初のサイクルで
データメモリA部の所定のビットのメモリAセル3へ読
み出し信号(RA)を出し、そのビットのメモリAセル
3からこのプロセッサの上半分の読み出しビット線を通
ってALUセル4のALU入力の一つの1ビットレジス
タに一旦記憶する。また同時に同様にデータメモリB部
の所定のビットのメモリBセル5へは読み出し信号(R
B)を出し、そのビットのメモリBセル5からこのプロ
セッサの下半分の読み出しビット線を通ってALUセル
4のALU入力のもう一つの1ビットレジスタに一旦記
憶する。For example, the data of the memory A cell 3 of a certain bit in the data memory A section and the data of the memory B cell 5 of a certain bit of the data memory B section are added to each other, and the memory B cell of the currently read bit in the data memory B section is added. The case of writing the addition result to 5 is as follows. First, in the first cycle, a read signal (RA) is issued to the memory A cell 3 of a predetermined bit of the data memory A section, and the ALU cell 4 is output from the memory A cell 3 of that bit through the read bit line of the upper half of this processor. It is temporarily stored in one 1-bit register of the ALU input. At the same time, a read signal (R) is similarly sent to the memory B cell 5 of a predetermined bit in the data memory B section.
B) from the memory B cell 5 of that bit, and through the read bit line of the lower half of this processor, and temporarily stores it in another 1-bit register of the ALU input of the ALU cell 4.
【0063】そしてその次のサイクルで加算のためのA
LU制御信号(ALU−CONT)を出してALUで加
算をし、このプロセッサの下半分の書き込みビット線を
通ってデータメモリB部の所定のビットメモリBセル5
に届け、そのビットのメモリBセル5へ書き込み信号
(WB)を出して書き込む。これも1ビット1ビット行
う。ALU出力制御信号(BA及びBB)は必要なタイ
ミングで発生される。ひとつの演算が2サイクルに見え
るが、パイプラインであり、各データ経路は一つの目的
にのみ使っているので、多数のビットを連続してサイク
ル毎に演算することが可能である。なお図でALUは3
入力2出力であるが、説明を省いた1入力1出力はキャ
リー(桁上げ)用で、第3のALU入力の前の1ビット
レジスタに記憶して、下位ビットからの演算結果を次の
サイクルでひとつ上位のキャリー入力とする。Then, in the next cycle, A for addition
A LU control signal (ALU-CONT) is issued, addition is performed by the ALU, and a predetermined bit memory B cell 5 of the data memory B section is passed through the write bit line in the lower half of this processor.
To write to the memory B cell 5 of that bit by issuing a write signal (WB). This is also performed by 1 bit 1 bit. The ALU output control signals (BA and BB) are generated at the required timing. Although one operation looks like two cycles, it is a pipeline and since each data path is used for only one purpose, it is possible to operate a large number of bits consecutively in each cycle. ALU is 3 in the figure.
Input 2 output, 1 input 1 output (explained) is for carry (carry), stored in the 1 bit register before the 3rd ALU input, and the operation result from the lower bit is stored in the next cycle. To make it one higher carry input.
【0064】ALUアレイ部では、このようにして上下
に存在するデータメモリA部とデータメモリB部から、
プログラムに応じてデータを読み出しては、必要な算術
演算或いは論理演算を施し、再びデータメモリA部或い
はデータメモリB部の所定のアドレスに書き込む。この
演算処理は全てビット処理であり、1ビット1ビット処
理を進める。In the ALU array section, from the data memory A section and the data memory B section which are present above and below,
The data is read according to the program, the necessary arithmetic operation or logical operation is performed, and the data is again written to a predetermined address of the data memory A section or the data memory B section. This arithmetic processing is all bit processing, and 1 bit 1 bit processing is advanced.
【0065】一水平走査期間分の演算処理が終わると、
その一水平走査期間分の時間のうちに、プログラムの最
後の部分でその一水平走査期間分の出力データを出力S
AM部に移す必要がある。今出力すべきデータがデータ
メモリA部にあるとする時は、まず最初のサイクルで所
定のメモリAセル3へ読み出し信号(RA)を出し、そ
のビットのメモリAセル3からこのプロセッサの上半分
の読み出しビット線を通ってALUセル4の1ビットレ
ジスタに一旦記憶し、次のサイクルでALUを経由し
て、このプロセッサの下半分の書き込みビット線を通っ
て出力SAM部の所定のビットの出力SAMセル6に届
け、その出力SAMセル6に書き込み信号(OW)を出
して書き込む。When the calculation processing for one horizontal scanning period is completed,
The output data for the one horizontal scanning period is output S at the last part of the program within the time for the one horizontal scanning period.
It is necessary to move to the AM section. When it is assumed that the data to be output is in the data memory A section, a read signal (RA) is first issued to a predetermined memory A cell 3 in the first cycle, and the memory A cell 3 of that bit outputs the upper half of this processor. It is temporarily stored in the 1-bit register of the ALU cell 4 through the read bit line of, and is output through the write bit line of the lower half of this processor via the ALU in the next cycle. The write signal (OW) is delivered to the output SAM cell 6 and written to the output SAM cell 6.
【0066】これも縦方向のビット線を経由して1ビッ
ト1ビットデータ転送される。またこの時もデータ移動
に際してALUで処理することは何もないが、ALUセ
ル4を通るようになっており、その際ALU出力制御信
号(BB)が必要なタイミングで発生される。ひとつの
データ転送が2サイクルに見えるが、パイプラインなの
で多数のビットを連続して転送するような場合はサイク
ルは増えない。In this case as well, 1-bit 1-bit data is transferred via the vertical bit lines. Also at this time, there is nothing to be processed by the ALU at the time of data movement, but the data is passed through the ALU cell 4, and at that time, the ALU output control signal (BB) is generated at a necessary timing. Although one data transfer looks like two cycles, the cycle does not increase when a large number of bits are transferred continuously because it is a pipeline.
【0067】以上のように、一水平走査期間の時間のう
ちに、入力SAM部に蓄積された入力データの読み出
し、必要な演算処理、必要なデータ移動、そして出力S
AM部への出力データの書き込みまでが、ビットを単位
とするSIMD制御プログラムで制御される。このプロ
グラム処理は水平走査期間を単位として繰り返される。
ほかの、従来例と同じことは説明を省略する。As described above, during the time of one horizontal scanning period, reading of the input data accumulated in the input SAM section, necessary arithmetic processing, necessary data movement, and output S
The writing of output data to the AM section is controlled by the SIMD control program in units of bits. This program process is repeated in units of horizontal scanning period.
Descriptions of other points that are the same as those of the conventional example are omitted.
【0068】多少、バリエーションなどについて補足す
る。上の説明では、入力SAM部のデータをALUでの
演算入力にしたり、ALU出力を出力SAM部へ書き込
んだりするケースがないが、これらもこの実施例の図の
構成からは可能である。Some variations will be supplemented. In the above description, there is no case where the data of the input SAM section is used as the arithmetic input in the ALU and the ALU output is written to the output SAM section, but these are also possible from the configuration of the drawing of this embodiment.
【0069】上の説明では、ALUでの演算に必要な2
つの演算入力と1つの演算出力のうち、1つの演算入力
と1つの演算出力は、メモリセル上で同じアドレスに一
致させて説明していたが、これら2つの演算入力と1つ
の演算出力の合計3つのメモリセル上のアドレスは、全
て別々にすることが、実施例の図の構成からは可能であ
る。In the above explanation, the two required for the ALU operation are
It was explained that one operation input and one operation output among one operation input and one operation output are matched with the same address on the memory cell, but the sum of these two operation inputs and one operation output is described. It is possible that the addresses on the three memory cells are all different from the configuration shown in the drawing of the embodiment.
【0070】[0070]
【発明の効果】従来例の構成では、入力SAM部、デー
タメモリA部、データメモリB部、出力SAM部の各メ
モリセルの動作がリードモディファイライト動作であっ
たために1サイクルの周期が長かったが、本件の実施例
によれば、リードとライトは別サイクルであり、1サイ
クルの周期は短くなる。In the configuration of the conventional example, one cycle is long because the operation of each memory cell of the input SAM section, the data memory A section, the data memory B section, and the output SAM section is the read modify write operation. However, according to the embodiment of the present application, read and write are different cycles, and the cycle of one cycle is shortened.
【0071】半分以下にするのは困難だが、60〜70
%ほどにできて1.5倍以上の高速化が期待できる。こ
れはそのままプロセッサの処理性能を向上させる。この
時、1ビットのデータ転送や、1ビットの演算を見ると
見かけ上2サイクルかかるように見えるが、実際にはパ
イプライン動作ができるので、処理サイクルが増えてし
まうことはない。It is difficult to reduce it to less than half, but 60 to 70
%, Which is expected to be 1.5 times faster. This directly improves the processing performance of the processor. At this time, 1-bit data transfer and 1-bit operation seem to take 2 cycles, but since the pipeline operation is actually possible, the processing cycle does not increase.
【0072】実施例は、従来例(図4)においてデータ
メモリA部とデータメモリB部のメモリセルを図5とす
るものと比べる時、トランジスタ数では増えていない。
実施例と、従来例(図4)では、各プイロセッサエレメ
ントにおける、縦のビット線が1本増えたことが、ハー
ドウェアの増加であり、これは軽微である。In the embodiment, the number of transistors does not increase when comparing the memory cells of the data memory A portion and the data memory B portion in FIG. 5 in the conventional example (FIG. 4).
In the embodiment and the conventional example (FIG. 4), the increase in the vertical bit line in each processor element is an increase in hardware, and this is minor.
【図1】本発明によるリニアアレイ型の並列DSPプロ
セッサのプロセッサエレメントの一例のモデル図であ
る。FIG. 1 is a model diagram of an example of a processor element of a linear array parallel DSP processor according to the present invention.
【図2】リニアアレイ型プロセッサの構成図である。FIG. 2 is a configuration diagram of a linear array type processor.
【図3】一般的なプロセッサの構成を説明するための図
である。FIG. 3 is a diagram illustrating a configuration of a general processor.
【図4】従来のプロセッサエレメントのモデル図であ
る。FIG. 4 is a model diagram of a conventional processor element.
【図5】別のメモリセル構成図である。FIG. 5 is another memory cell configuration diagram.
1 入力ポインタセル 2 入力SAMセル 3 メモリAセル 4 ALUセル 5 メモリBセル 6 出力SAMセル 7 出力ポインタセル 1 Input Pointer Cell 2 Input SAM Cell 3 Memory A Cell 4 ALU Cell 5 Memory B Cell 6 Output SAM Cell 7 Output Pointer Cell
Claims (5)
において、 データメモリ部のメモリセルに接続される読み出しビッ
ト線及び書き込みビット線とを有することを特徴とする
リニアアレイ型の並列DSPプロセッサ。1. A linear array type parallel DSP processor having a read bit line and a write bit line connected to a memory cell of a data memory section in the linear array type parallel DSP processor.
SPプロセッサにおいて、 少なくともSAM(シリアルアクセスメモリ)部、デー
タメモリ部、ALUアレイ部よりなり、 上記データメモリ部とALUアレイ部の間のデータのや
り取りであるメモリ読み出し又は書き込みを独立して行
うことを特徴とするリニアアレイ型の並列DSPプロセ
ッサ。2. The linear array type parallel D according to claim 1.
In the SP processor, at least a SAM (serial access memory) unit, a data memory unit, and an ALU array unit are provided, and memory reading or writing, which is data exchange between the data memory unit and the ALU array unit, can be independently performed. Characteristic linear array type parallel DSP processor.
SPプロセッサにおいて、 データメモリ部のメモリセルの読み出し線をALU部の
入力に接続するとともに、上記データメモリ部のメモリ
セルの書き込み線をALU部の出力に接続することを特
徴とするリニアアレイ型の並列DSPプロセッサ。3. The linear array type parallel D according to claim 2.
In the SP processor, the read line of the memory cell of the data memory unit is connected to the input of the ALU unit, and the write line of the memory cell of the data memory unit is connected to the output of the ALU unit. Parallel DSP processor.
SPプロセッサにおいて、 上記プロセツサは全てのプロセツサが1つのプログラム
により連動して動作をするSIMD制御であることを特
徴とするリニアアレイ型の並列DSPプロセッサ。4. The linear array type parallel D according to claim 3.
In the SP processor, the processor is SIMD control in which all the processors operate in synchronization with one program, which is a linear array parallel DSP processor.
SPプロセッサにおいて、 上記プロセツサは映像信号の水平解像度に相当する個数
のプロセツサよりなり、全てのプロセツサが1期間にお
いて1水平走査線の情報を処理することを特徴とするリ
ニアアレイ型の並列DSPプロセッサ。5. The linear array type parallel D according to claim 4.
In the SP processor, the processor is composed of a number of processors corresponding to the horizontal resolution of a video signal, and all the processors process information of one horizontal scanning line in one period, which is a linear array parallel DSP processor.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP4158264A JPH064689A (en) | 1992-06-17 | 1992-06-17 | Linear array parallel DSP processor |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP4158264A JPH064689A (en) | 1992-06-17 | 1992-06-17 | Linear array parallel DSP processor |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH064689A true JPH064689A (en) | 1994-01-14 |
Family
ID=15667820
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP4158264A Pending JPH064689A (en) | 1992-06-17 | 1992-06-17 | Linear array parallel DSP processor |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH064689A (en) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR960029967A (en) * | 1995-01-19 | 1996-08-17 | 윌리엄 이. 힐러 | Digital signal processing method and apparatus, and memory cell readout method |
| CN102337896A (en) * | 2011-09-08 | 2012-02-01 | 上海盾构设计试验研究中心有限公司 | Roller type cutter rectangular pipe jacking machine |
-
1992
- 1992-06-17 JP JP4158264A patent/JPH064689A/en active Pending
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR960029967A (en) * | 1995-01-19 | 1996-08-17 | 윌리엄 이. 힐러 | Digital signal processing method and apparatus, and memory cell readout method |
| CN102337896A (en) * | 2011-09-08 | 2012-02-01 | 上海盾构设计试验研究中心有限公司 | Roller type cutter rectangular pipe jacking machine |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US6195106B1 (en) | Graphics system with multiported pixel buffers for accelerated pixel processing | |
| EP0539595A1 (en) | Data processor and data processing method | |
| JPH077260B2 (en) | Image data rotation processing apparatus and method thereof | |
| JP2947664B2 (en) | Image-dedicated semiconductor storage device | |
| US5926583A (en) | Signal processing apparatus | |
| EP1314099B1 (en) | Method and apparatus for connecting a massively parallel processor array to a memory array in a bit serial manner | |
| JP3086769B2 (en) | Multi-port field memory | |
| JP2001084229A (en) | SIMD type processor | |
| US7593016B2 (en) | Method and apparatus for high density storage and handling of bit-plane data | |
| JPH06274528A (en) | Vector processing unit | |
| US7284113B2 (en) | Synchronous periodical orthogonal data converter | |
| EP0456394B1 (en) | Video memory array having random and serial ports | |
| JPH0969061A (en) | Video signal processor | |
| JPH04295953A (en) | Parallel data processor with built-in two-dimensional array of element processor and sub-array unit of element processor | |
| JPH064689A (en) | Linear array parallel DSP processor | |
| JPH0256760B2 (en) | ||
| JPH01283676A (en) | Read-out processing system for window image data | |
| JP2002269067A (en) | Matrix arithmetic unit | |
| JPH0522629A (en) | Video signal processor | |
| JPH07210545A (en) | Parallel processor | |
| JPH064284A (en) | Linear array parallel DSP processor | |
| JPH064690A (en) | Linear array parallel DSP processor | |
| JP2855899B2 (en) | Function memory | |
| JPH08297652A (en) | Array processor | |
| JP2664420B2 (en) | Image processing apparatus and buffering apparatus for image processing apparatus |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| A02 | Decision of refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A02 Effective date: 20040330 |