JPH012176A - Product-sum calculation method - Google Patents
Product-sum calculation methodInfo
- Publication number
- JPH012176A JPH012176A JP62-156427A JP15642787A JPH012176A JP H012176 A JPH012176 A JP H012176A JP 15642787 A JP15642787 A JP 15642787A JP H012176 A JPH012176 A JP H012176A
- Authority
- JP
- Japan
- Prior art keywords
- data
- output
- processor
- adder
- intermediate storage
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Abstract
(57)【要約】本公報は電子出願前の出願データであるた
め要約のデータは記録されません。(57) [Summary] This bulletin contains application data before electronic filing, so abstract data is not recorded.
Description
【発明の詳細な説明】
産業上の利用分野
本発明は、画像データ等の時系列及び空間配列のデータ
に対し、ある分布曲線に対応する係数を乗じ、その後加
算する積和演算方式に関する。DETAILED DESCRIPTION OF THE INVENTION Field of the Invention The present invention relates to a product-sum operation method in which time-series and spatially arranged data such as image data is multiplied by a coefficient corresponding to a certain distribution curve and then added.
従来の技術
画像データの鮮鋭化、空間フィルタリング等に使用され
るたたみ込み積分器のような、時系列及び空間配列のデ
ータから順次n個のデータを取込み、このn個のデータ
に、ある分布曲線に対応する係数を各々乗じ、その後金
係数が乗じられたn個のデータを加算して出力する積和
演算方式においては、nlJのデータと各々の係数との
積を同時に求め、これらを加算して出力し、次にデータ
を1つシフトさせ、シフトされたn個のデータと各々の
係数の積を同時に求め、これらを加算して出力するとい
う動作を順次行っている。この場合、n個のデータに対
し、各々の係数を同時に乗じる必要があることから、乗
算するプロセッサをn個並列に設けた積和P#口器のシ
ストリック・アレイで処理を行っている。Conventional technology A convolution integrator used for image data sharpening, spatial filtering, etc. sequentially captures n pieces of data from time-series and spatial array data, and applies a certain distribution curve to these n pieces of data. In the product-sum operation method, which multiplies each coefficient by a coefficient corresponding to , and then adds and outputs n data multiplied by a gold coefficient, the product of nlJ data and each coefficient is simultaneously calculated, and these are added. Then, the data is shifted by one, the product of the n shifted data and each coefficient is simultaneously determined, and these are added and output. In this case, since it is necessary to simultaneously multiply n pieces of data by respective coefficients, processing is performed using a systolic array of product-sum P# mouthpieces in which n multiplication processors are provided in parallel.
即ち、データ列がal、a2.a3・・・・・・am
。That is, the data strings are al, a2 . a3...am
.
・・・・・・と連続しているものとし、この内n個のデ
ータを順次取込み係数に1−koを各々同時に乗じた後
、加t)シて出力Yを出力するものとすると、積和演算
器の出力Yは次のようになる。. . . , and if n data are sequentially taken in and multiplied by 1-ko for each at the same time, then added t) and output Y, the product is The output Y of the summation unit is as follows.
第1回[1の出力Y1は、
Yl−=k a 十k a +に3a3・・・
・・・ka ・・・・・・(1)n
2回目の出力Y2は、
Y 2 −− k 1 a 2 +
k 2 a 3 −) k 3 a
a・・・・・・+knan+1 ・・・・・・
(2)1回目の出力Yiは、
Y i =に1at +に2a、+1+に3a、。The output Y1 of 1st [1 is Yl-=k a + k a + 3a3...
・・・ka ・・・・・・(1)n The second output Y2 is Y 2 −− k 1 a 2 +
k 2 a 3 -) k 3 a
a...+knan+1...
(2) The first output Yi is 1at for Y i = 2a for +, 3a for +1+.
・・・・・・+ka−・・・・・・(3)nn÷巨1
上述のように、計数とデータのn個の積は同時に求めね
ばならないことから、乗算器としてのプロセッサをn個
並列に設けたシストリック・アレイを用いている。・・・・・・+ka−・・・・・・(3) nn ÷ huge 1 As mentioned above, n products of count and data must be calculated simultaneously, so n processors are used as multipliers. It uses systolic arrays arranged in parallel.
発明が解決しようとする問題点
しかしながら、乗算器として用いられるプ【〕セセラは
高速動作が要求されCいるため高価格である。イのため
、乗算する係数の数が多くなればなるほど積和演算器の
価格は高価となるという欠点がある。Problems to be Solved by the Invention However, the processor used as a multiplier is required to operate at high speed and is therefore expensive. Therefore, the disadvantage is that the more the number of coefficients to be multiplied, the more expensive the product-sum calculator becomes.
そこで、本発明の目的は、乗算器としてのプロセッサの
数を減少させ、低価格で積和演算器を構成できる積和演
算方式を提供することばある。SUMMARY OF THE INVENTION Therefore, an object of the present invention is to provide a product-sum calculation method that can reduce the number of processors as multipliers and configure a product-sum calculation unit at a low cost.
問題点を解決するための手段
n個の係数を対応するデータに乗じ、多積を加nする積
和演算方式において、データ列の各)−一タがシフトさ
れながら入力され、入力されたデータに係数を乗じるプ
ロセッサをn / d個有し、各ブロセッ奢すの出力を
加口して出力するブ[1セツ1y・アレイと、プL:l
Izツサ・アレイの出力と各シフトに対応する中間ス
トレージのアドレスの記憶値を加尊する加n器と、各シ
フト毎の加n器の出力を対応するアドレスに格納する中
間ストレージを有し、同一データ列を上記プロセッサ・
アレイに6回入力し、データ列を入力する各回毎に各プ
ロセッサに設定する係数値を決められた値に設定変更し
、各回のデータ列の入力毎に上記中間ストレージからの
データの読出し及び[記加n鼎からの出力の山込みをn
/d周tI′iiらせ、6回目の上記加算器の出力をn
個の係数とデータの積の加算結果として出力するように
構成づることにより、データに係数を乗する積n器とし
てのプロセッサの数を少なくして上記問題点を解決した
。Means for Solving the Problem In the product-sum operation method of multiplying corresponding data by n coefficients and adding n multiproducts, each () - ta of the data string is input while being shifted, and the input data is It has n/d processors that multiply the output of each processor by a coefficient, and has a block array that multiplies the output of each processor and outputs the output.
It has an adder that processes the output of the Iztusa array and the stored value of the address of the intermediate storage corresponding to each shift, and an intermediate storage that stores the output of the adder for each shift in the corresponding address, The same data string is processed by the above processor.
The data is input to the array six times, and each time the data string is input, the coefficient value set in each processor is changed to a predetermined value, and each time the data string is input, the data is read from the intermediate storage and [ The pile of output from the record is n
/d cycles tI'ii, and the output of the above adder for the sixth time is n
By configuring the present invention to output the result of adding the products of coefficients and data, the number of processors serving as multipliers for multiplying data by coefficients is reduced, thereby solving the above problem.
作 用
データ列をa、a、a3・・・・・・、プロセッサ・ア
レイのプat?ツサの数n / d == rrlとす
るとし、まず始めは、各プロセッサに各々に、に2.・
・・・・・k、が設定され、データ列が入力されると(
以下データ列a 、a2 、a3・・・・・・をプロ
セッサ・アレイにすべて入力することをパスという)。The operation data string is a, a, a3..., the processor array's output? Assume that the number of tufts is n/d == rrl, and at first, each processor has 2.・
...When k is set and a data string is input (
(Hereinafter, inputting all data strings a, a2, a3, . . . to the processor array is called a pass).
上記加算器の最初の周期の出力P1−1は次のようにな
る(なお、中間ストレージの各アドレスに格納された値
は初期化でrOJにクリアされている)。The output P1-1 of the first cycle of the adder is as follows (note that the values stored at each address of the intermediate storage are cleared to rOJ during initialization).
1)1−1=に1a1+に2a2+川…+kIllaI
Ilそして、この出力P1−1は中間ストレージの対応
するアドレスへ1に格納される。そして、入力データが
シフトされた次のシフト周期では上記加算器の出力P1
〜2は
Pl−2−に1a2+に2a3+
°°°゛°°klall。1
となり、中間ストレージの対応するアドレスへ2に格納
される。以下同様に処理が行われ、1回目の周期におい
ては加算器の出力P1−iは次のようになる。1) 1-1=to 1a1+to 2a2+river...+kIllaI
This output P1-1 is then stored as 1 in the corresponding address of the intermediate storage. Then, in the next shift period after the input data is shifted, the output P1 of the adder is
~2 is Pl-2- to 1a2+ to 2a3+ °°°゛°°klall. 1 and is stored at the corresponding address in the intermediate storage as 2. Processing is performed in the same manner thereafter, and in the first cycle, the output P1-i of the adder is as follows.
Pl−i−に1a、 十に、 ai+1+”” ””
km a i+a−f
そして、この値’1−iは中間ストレージのアドレスA
iに格納される。1a to Pl-i-, ai+1+”” ”” to Pl-i-
km a i+a-f And this value '1-i is the address A of the intermediate storage
stored in i.
かくして、積和演篩されるべきデータ列a1゜a2・・
・・・・の寸べてのシフトを終えると、次にブ[1セツ
サ・アレイの各プロセッサの係数をkIIl+1〜に2
Ilk:変更し、再びデータ列a 1. a 2・・・
・・・をプロセッサ・アレイに入力し前述同様の処理を
行う。Thus, the data strings a1, a2, etc. to be subjected to the product-sum operation
After completing the shift of ..., the coefficients of each processor in the block [1 setter array are kIIl+1~2]
Ilk: Change data string a again.1. a2...
... is input to the processor array and the same processing as described above is performed.
即ち、この2回目のパスにおいては、中間ストレ−ジか
らのデータの読出し、書込みをm周期遅らせ、m+i回
目のシフ1〜時から順に中間ストレージのアドレスA1
よりデータを読出し、このデータにm+1回[1のシフ
ト時のプロセッサ・アレイの出力値を加算し、その和を
アドレスA1から順に格納する。That is, in this second pass, the reading and writing of data from the intermediate storage is delayed by m cycles, and the address A1 of the intermediate storage is sequentially changed from shift 1 to the m+ith shift.
The data is read out, the output value of the processor array at the time of shift of 1 is added to this data m+1 times, and the sum is stored in order starting from address A1.
即ち、2回目のデータ入力におけるm→1回目のシフト
時において加n器の出力P2−1は次のようになる。That is, at the time of the first shift from m in the second data input, the output P2-1 of the adder becomes as follows.
P2−1−’−[)1−1 トki+1 am+1
十k11)2 al11+2−1°” ”” 2n+
82m
= k a 十k a +・・・−・・+ k
2.a、、。P2-1-'-[)1-1 ki+1 am+1
10k11) 2 al11+2-1°""" 2n+
82m = ka 10ka +...-...+k
2. a...
そして、次のシフト周期では
P −P +k a 十に、ll+2am、3
2−2 1−2 m÷1111÷2” ”’ ”
・k2m a2m+ 1
= k 1 a 2 + k 2 a 3+”” ”’
十に2i a2m+ 1
また、i回目のシフト時の周期では
P2−i ” Pl−i 十k111)1 a
II++i+ k a ・°= =゛+ k 2
.a 2Il+H−11+2 l+1+1
−− k 1 a 1+ k 2 a i+1””””
k2ma2a+i−1
、そして、これらの加暮器の出力P2−2・・・・・・
P2−+・・・は各々中間ストレージのアドレスA2・
・・・・・Δi・・・・・・に格納される。Then, in the next shift period, P −P +k a 10, ll+2am, 3
2-2 1-2 m÷1111÷2” ”’ ”
・k2m a2m+ 1 = k 1 a 2 + k 2 a 3+”” ”'
10 to 2i a2m+ 1 Also, in the period of the i-th shift, P2-i ” Pl-i 10k111) 1 a
II++i+ k a ・°= =゛+ k 2
.. a 2Il+H-11+2 l+1+1 -- k 1 a 1+ k 2 a i+1""""
k2ma2a+i-1, and the outputs of these Kaaku devices P2-2...
P2-+... are intermediate storage addresses A2 and 2, respectively.
. . . Stored in Δi . . .
以下同様にして、処理が行われ、最後のデータ人力5即
ち、d回目のパス時においては、ブ[1セツサ・アレイ
の各プロセッサに設定する係数をk 、k
・・・・・・ko (dm=nであるたn−11+I
n−1÷2
め、1パスでm個の係数が設定されd回目では、プロセ
ッサ・アレイの最後のプロレッナにはkdl、1=kn
が設定されることとなる)に変更し、中間ス1−レージ
からのデータの読出し、書込みを(d−1)m周期遅ら
せて(d−1)m+1回目のシフト時から順に中間スト
レージのアドレスA1よリデータを読出し、プロセッサ
・アレイの出力と加算し、これを出力として出力する。Processing is performed in the same manner thereafter, and in the final data input 5, that is, at the d-th pass, the coefficients to be set for each processor of the setter array are set to k, k.
・・・・・・ko (dm=n so n-11+I
n-1 ÷ 2, m coefficients are set in one pass, and at the dth time, the last prorena of the processor array has kdl, 1=kn
is set), reading and writing data from the intermediate storage 1-storage is delayed by (d-1)m cycles, and the addresses of the intermediate storage are set sequentially from the (d-1)m+1st shift. Read data from A1, add it to the output of the processor array, and output it as an output.
即ち、1回目の出力Y1は次のようになる。That is, the first output Y1 is as follows.
Yl” Pd−1”” p(d−1) −1+kn−m
+1” a (d−1)l+1+ kn−142a
(d−1)+2”’ ”” + kn a (d−1
)e+m= pl al + k2 a2 =−°”
+ k(d−1)m+k a
” a(d−1)s n−a+ (d−
1)sul”” ”” 十kn a(d−1)its”
−k 1 a 1 −ト k 2 a 2 +
°°゛ °゛° k n−I a n−鋼・・・・・
・ka
” kn−m+I n−m+I n 。Yl"Pd-1"" p(d-1) -1+kn-m
+1” a (d-1)l+1+ kn-142a
(d-1)+2”' ””+kna (d-1
)e+m= pl al + k2 a2 =-°”
+ k(d-1)m+k a ” a(d-1)s n-a+ (d-
1) sul”” ”” 10kna a(d-1)its”
-k 1 a 1 -k 2 a 2 +
°°゛ °゛° k n-I a n-steel...
・ka” kn-m+I n-m+I n.
=k a 十k a ト・・・・・・−トk
n aol 1 2 2
・・・・・・(4)
同様に、i回目のシフト時の周期での出力Yiは次の第
(5)式のようになる。= k a 10k a t・・・・−tk
n aol 1 2 2 (4) Similarly, the output Yi in the period of the i-th shift is expressed by the following equation (5).
Yi”” ’−P(d−1)−i+kn−t)1d−+
” ” (d−t)m+i+ kn−m+2 a(
d−1)m+国+”’ ”” n E″(d−1)a+
i+ 1−1−に1at 十に2 ” tit ”””
kn an+i−1・・・・・・(5)
上記第(4)式、第(5)式は各々第(1)式、第(3
)式と同一であり、データ列から順次11個のデータを
取込みシフl−L、ながら各データに係数に1〜knを
、各々乗じてその積を加算した出力を得る積和演算が行
われたことを意味する。Yi""'-P(d-1)-i+kn-t)1d-+"" (d-t)m+i+ kn-m+2 a(
d-1) m+country+”' ”” n E”(d-1)a+
i+ 1-1- 1at 10-2 ” tit ”””
kn an+i-1...(5) The above equations (4) and (5) are the equations (1) and (3), respectively.
) is the same as the formula, and a sum-of-products operation is performed to obtain the output by sequentially taking in 11 data from the data string, shifting l-L, multiplying each data by a coefficient of 1 to kn, and adding the products. It means something.
実施例
図は本発明の一実施例のブロック図で、1は画像データ
等の時系列又は空間配列の被処理データが格納されたダ
イナミックRAM等で構成されたメモリで、例えば画像
データであると、1画素分の8ビツトのデータを記憶す
るメモリセルがCCDカメラ等により対象物を走査した
ー走査分として256個設l5れ、25611の走査分
として、合語256×256個のメモリヒルで構成され
ている。2は被処理データ1の各データに分布曲線に対
応する係数を乗じる積算器としてのプロセッサが多数配
設されたプロセッサ・アレイで、少なくと6被処理デー
タ1の各データを順次シフトするシフトレジスタ、係数
を記憶する係数記憶部、及び係数を切換えるための係数
切換f段、さらに各プロセッサの出力、即ちデータと係
数の積を足し合わせる加q器を右し【いる。3は、プロ
セッサ・アレイ2の出力、即ち各データと係数との積の
和と中間ストレージ4に記憶させたデータを加粋づる加
算器、4は積和演算の出力Yとして出力するまでの中間
処理の段階のデータを一時記憶するための中間ストレー
ジ、5,6.7はゲート回路である。そして、被処理デ
ータ1からプロセッサ・アレイ2へのデータの入力、プ
ロはツサ・アレイ2中のデータのシフト、乗韓、加算処
理、係数の切換、加算器の加算処理、ゲート5.6.7
の制御、中間ストレージ4へのデータの書込み。The embodiment diagram is a block diagram of one embodiment of the present invention, and 1 is a memory configured with a dynamic RAM etc. in which time-series or spatially arranged data to be processed such as image data is stored. , there are 256 memory cells that store 8-bit data for one pixel when the object is scanned by a CCD camera, etc., and 256 x 256 memory cells are provided for 25611 scans. has been done. 2 is a processor array in which a large number of processors are arranged as multipliers that multiply each piece of data 1 to be processed by a coefficient corresponding to a distribution curve, and at least 6 shift registers that sequentially shift each piece of data 1 to be processed. , a coefficient storage section for storing coefficients, a coefficient switching f stage for switching coefficients, and a q adder for adding up the output of each processor, that is, the product of data and coefficients. 3 is an adder that adds the output of the processor array 2, that is, the sum of the products of each data and coefficient, and the data stored in the intermediate storage 4; 4 is an intermediate until output as the output Y of the product-sum operation; Intermediate storage 5, 6.7 is a gate circuit for temporarily storing data at a processing stage. Then, the input of data from the processed data 1 to the processor array 2, the processing of data in the TSAR array 2, the multiplication, the addition processing, the switching of coefficients, the addition processing of the adder, the gate 5.6. 7
control, and write data to intermediate storage 4.
読出しの制御等は図示しないIII御装置によってマイ
クロプログラム制御が行われている。Control of reading and the like is performed by microprogram control by a III control device (not shown).
積和演算を行うべき一群の被処理データ列、例えば上記
画像データにおいては一走査分の256個の画素分のデ
ータ列に対し、該データ列a1゜a2・・−・・・を順
にシフトしながらn個のデータに対し分布曲線に対応づ
る係数に1〜knを乗じ、その積を加粋し′(その出力
Yを第(3)式、第(5)式に示すように
Yi−k a、+k a、 十−koao、。For a group of data strings to be processed on which a product-sum operation is to be performed, for example, in the above image data, data strings for 256 pixels for one scan are sequentially shifted. For n pieces of data, multiply the coefficient corresponding to the distribution curve by 1 to kn, add the product, and calculate the output Y as shown in equations (3) and (5). a, +k a, 10-koao,.
1 1 2 1+1
(i=1.2.3・・・・・・)
として出力するものとする場合において、本発明におい
ては、積算するプロセッサの数mをn/d(dはnの約
数)として、プロセッサ・アレイ2に入力するデータ列
を6回パス(入力)させるようにしている。そして、本
実施例においては、n=4.d=2.m=4÷2=2と
してプ11セツυ・アレイ2のプロセッサの数を2とし
、4個のa・ 、a H+ 3 (i−1。1 1 2 1+1 (i=1.2.3...) In the present invention, the number m of processors to be integrated is n/d (d is a divisor of n). ), the data string input to the processor array 2 is passed (input) six times. In this embodiment, n=4. d=2. Assuming that m=4÷2=2, the number of processors in the array 2 is 2, and there are 4 a, aH+3 (i-1.
データai ・ai+1 ・++2
2.3・・・・・・)に各々係数k 、k 、に3
.に4を乗じた各種を加算し次の第6式に示す積和演n
を行う場合について説明する。data ai ・ai+1 ・++2 2.3...), coefficient k, k, 3, respectively
.. By adding various products multiplied by 4, we obtain the product-sum operation n shown in the following equation 6.
The case where this is done will be explained below.
Y + =に1 a H+ k2 a 国+ k3a
++2→−に4ai+3 ・・・
・・・(6)まず、プロセッサ・アレイ2の2つのプロ
セッサに係数に、に2を各々設定し、被処理データ1か
ら積和演n′TJる一群のデータa−、a2・・・・・
・をブロアレイに順次シフトしながら入力する。例えば
、256個の画素データa ”a256をプロレツサ・
アレイに入力する場合、まず始めにプ[1セツサΦアレ
イ2の2つのプロセッサには各々a1.a2が入力され
、各プロセラυに設定された係r!ik、k とこの
データa 1. a 2 (7)栢k 1a 1. k
2 a 2が求められ加算され出力される。この出力
に1a1→・k2a2は加算器3に入りされて中間スト
レージ4からのデータと加粋されるが、第1回目のデー
タのパス時においてはゲート回路5.7は口1じられゲ
ート回路6のみがマイクロプログラム制御方式の制御l
装置より信号S1が入力され間どされ、その結果、加算
器3はブ[:I t:: ッ4j・7レイ2の出力に1
a1+に2a2を出力P どして出力する。Y + = 1 a H+ k2 a country + k3a
++2→-4ai+3...
(6) First, set the coefficients to 2 for each of the two processors of the processor array 2, and calculate a group of data a-, a2, etc. by product-sum operation n'TJ from the data to be processed 1.・
・Input while shifting sequentially to the blow array. For example, 256 pixel data a "a256"
When inputting to the array, first, the two processors of the processor Φ array 2 each have a1. a2 is input and the relation r! set for each processor υ! ik, k and this data a 1. a 2 (7) 梢k 1a 1. k
2 a 2 is calculated, added, and output. This output 1a1→・k2a2 is input to the adder 3 and added to the data from the intermediate storage 4, but during the first data pass, the gate circuit 5.7 is closed and the gate circuit Only 6 is controlled by microprogram control method.
The signal S1 is input from the device and is delayed, and as a result, the adder 3 outputs 1 to the output of the 4j.7 ray 2.
Output 2a2 to a1+ through output P.
Pl−1−= kl a1+ k2a2 ””
”(7)この出力P1−1はゲート回路6を介し【、中
間ストレージ4のアドレスA1に書込まれる。Pl-1-= kl a1+ k2a2 ””
(7) This output P1-1 is written to the address A1 of the intermediate storage 4 via the gate circuit 6.
イして、データが1つシフトされて、プロセラlす・ア
レイ2の各プロセッサにはa 2 、 a 3のデータ
が入力され、このデータに係数に1.に2が各々乗じら
れ、加算されグロッセザ・アレイ2から出力される。イ
の結果中間ストレージ4からのデータはないから加算器
3の出力P1−2は次の第(8)式のようになり、中間
ストレージ4のアドレス△2に記憶される。Then, the data is shifted by one, and the data of a 2 and a 3 is input to each processor of the processor array 2, and this data is given a coefficient of 1. are respectively multiplied by 2, added, and output from the grosser array 2. As a result of step A, there is no data from the intermediate storage 4, so the output P1-2 of the adder 3 becomes as shown in the following equation (8) and is stored at address Δ2 of the intermediate storage 4.
P 1−2 =k 1a 2 + k2a 3””・・
(8)このように、データa、a2・・・・・・が、順
次シフトされ各々係数に、に2が乗じられ、2つの積の
和が中間ストレージ4の各アドレスに格納されることと
なり、i番目の周期においては、加算器の出力P1−i
は次の第(9)式のようになり中間ストレージ4のアド
レスAiに格納される。P 1-2 = k 1a 2 + k2a 3""...
(8) In this way, data a, a2, etc. are sequentially shifted, each coefficient is multiplied by 2, and the sum of the two products is stored in each address of the intermediate storage 4. , in the i-th period, the adder output P1-i
is expressed as the following equation (9) and is stored at address Ai of intermediate storage 4.
P 、 =k a、+に2ai+t −・・
・・・<9>1−tl+
かりシて、積和演算を行うすべてのデータをプロセッサ
・アレイに1回バスさゼると、例えば上記画像データの
例で、256個の画素データをすべてパスさせると、中
間ストレージ4のA1−A256には1回目のパスにお
ける積和演算の算出ll1p −p が各々
格納されたこととなる。P, = k a, + to 2ai + t -...
...<9>1-tl+ If all the data to be subjected to the multiply-accumulate operation are passed to the processor array once, for example, in the image data example above, all 256 pixel data will be passed. Then, the calculations ll1p -p of the sum of products operation in the first pass are stored in A1 to A256 of the intermediate storage 4, respectively.
(なお、算出値P は意味をもたないので実際は格
納しないようにする。)
次に、プロセッサ・アレイの各プロセッサに設定する係
数をに、に4に設置変更し、再び同じデータ列をブ臼口
ッナ・アレイ2にバスさせる。(Note that the calculated value P has no meaning, so it is not actually stored.) Next, change the coefficients set for each processor in the processor array to 4, and block the same data string again. Let Morguchina Array 2 bus.
このとさ、第1回目の周期においてはプロUツリー・ア
レイ2の出力はに3a1−トk 、s a 2となり、
加算器3に入力されることとなるが、このときはすべて
のグーl−回路を閉じておき、加算器3の出力P
=、k a −+−に4a2は中間スt−L/ −
シ4に入力されず、また、出力Yとしても出力されない
。そして、データが1つシフ]〜され、この周期におけ
る加算器の出力P ・−k 3a 2 +k 4a
3も中間ストレージ4に入力されず、出力Yどしても出
力されない。そして、2周II遅れた3周期目から、即
ち、プロセッサ・アレイ2のブ1コセッナの数(m=
2 )だけ遅れた次の周期から、制御装置よりS 、S
3を送出しゲート回路5゜6を開とし、3周期から加算
器3に中間ストレージ4のアドレスA1に格納されてい
た第(7)式に示づ゛データP =k a +に2a
2を入力ず1−11す
る。In this case, in the first cycle, the output of the pro U-tree array 2 is 3a1-k, s a 2,
It will be input to the adder 3, but at this time, all the circuits are closed and the output P of the adder 3 is
=, k a −+−, 4a2 is the intermediate st−L/−
It is not input to Y4, nor is it output as output Y. Then, the data is shifted by one] and the output of the adder in this period is P ・−k 3a 2 +k 4a
3 is not input to the intermediate storage 4, and the output Y is not outputted either. Then, from the third cycle delayed by two cycles II, that is, the number of blocks in the processor array 2 (m=
From the next cycle delayed by 2), S, S
3, the gate circuit 5.6 is opened, and from the 3rd period, the adder 3 receives data P = k a + 2a as shown in equation (7) stored at the address A1 of the intermediate storage 4.
Run 1-11 without entering 2.
その結果、加算器3ではプロセッサ・アレイ2から出力
された積の和k a +に4a4と1回目のバスの
1回目の周期の積和演口の値P1−1が加t1され、こ
れがゲート回路7を介して出力Y1として出力される。As a result, in the adder 3, 4a4 and the value P1-1 of the product-sum operator of the first cycle of the first bus are added to the sum of products k a + output from the processor array 2, and this is added to the sum k a + of the products output from the processor array 2. It is output via circuit 7 as output Y1.
即ち、
Yl =P2.==P1.+に3a3+に4a4=に1
a1+に2a2+に3a3
+に4a4 ・−・・−(10)次の周期では
プロセッサ・アレイ2の出力に3a4+に4a5と中間
ストレージ4のアドレスA2に配憶されたデータP1−
2が加算され、次の出力Y2が出力される。That is, Yl = P2. ==P1. + to 3a3+ to 4a4=to 1
a1+ to 2a2+ 3a3 + to 4a4 (10) In the next cycle, the output of processor array 2 is 3a4+ to 4a5 and the data P1- stored at address A2 of intermediate storage 4
2 is added and the next output Y2 is output.
Y2==P =P +k
a +k 4 as=−k 1a 2 +
k 2 a 3+ k 3a a + k 4a s
・・・・・・(11)
同様に、i番目の周期においては、プロセッサ・アレイ
2の出力k a、 +に4a、、と中間3 1+
2
ストレージ4のアドレスAiに記憶するデータP1−i
が加算され次の出力Y1が出力される。Y2==P=P+k
a +k 4 as=-k 1a 2 +
k 2 a 3+ k 3a a + k 4a s (11) Similarly, in the i-th period, the output k a of processor array 2 is 4a, and the intermediate 3 1+
2 Data P1-i stored at address Ai of storage 4
are added and the next output Y1 is output.
Yi==P −−=P ・+に3a1+22−+
1−+
+ k 4 a 1 +3
=k a、+に2a、+1
十kaai+2十に4a1+3
・・・・・・ (11)
(i=1.2・・・・・・)
即ら、第(11)式と第(6)式は同一であり、4a・
、ai+3 (i六
個のデータai 、ai+1 、++21.2・・・・
・・で上記256の画素データであるとY254’〜Y
256は意味をもたないためi =1〜253となる。Yi==P --=P ・+3a1+22-+
1-+ + k 4 a 1 +3 = k a, + 2a, +1 10 kaai + 20 4a1 + 3 ...... (11) (i=1.2...) That is, the (th) Equation 11) and Equation (6) are the same, and 4a.
, ai+3 (i six data ai , ai+1 , ++21.2...
...and the above 256 pixel data is Y254'~Y
Since 256 has no meaning, i=1 to 253.
)に各々係数に1.に2.に3゜k4が乗じられ、その
積の和が各周期毎に出力されることとなる。) to each coefficient of 1. 2. is multiplied by 3°k4, and the sum of the products is output for each cycle.
上記実施例は積和演算を行うデータの数nを4゜プロセ
ッサの数mを2.データ列をバスさせる回数dを2とし
たが、データの数nが6.プロセッサの数mを2.デー
タ列をバスさせる回数dを3どすると、1回目のバス時
においては、プロセッサ・アレイ2及び加算器3で積和
された値を中間ストレージのアドレスA 1 J:り順
に格納し、2回口のバス時には、プロセッサの数2 (
=m)だけ遅らせて3回目の周期よりゲート回路5.6
を開ぎ、加算器3で3周期目の各プロセッサの出力とア
ドレスA1のデータを加算し、アドレスA1に格納し、
以下類に各周期毎中間ストレージのアドレスA2より順
に読出し、該アドレスの記憶データとプロセッサ・アレ
イの各プロセッサからの出力を加算し、読出したアドレ
スに加算器3の出力を書込み、3回目のバス時において
は、さらに2周期、即ち、mX2=2X2=4周期遅ら
せ、5周期目からグーl−回路5.7を開きアドレスΔ
1より記憶データを順に読出し、加算器3で各シフト時
のプロセッサ・アレイからの出力と該読出した記憶デー
タを加算して、順に出力Yを出力させるようにすればよ
い。In the above embodiment, the number n of data to be subjected to the product-sum operation is 4 degrees, and the number m of processors is 2 degrees. The number of times d of busing the data string was set to 2, but the number of data n was set to 6. Let the number m of processors be 2. If the number of times d of busing the data string is 3, then during the first bus, the values summed by the products in the processor array 2 and the adder 3 are stored in the order of addresses A 1 J: of the intermediate storage, and the values are transferred twice. When the bus is opened, the number of processors is 2 (
Gate circuit 5.6 from the third period after delaying by = m)
, the adder 3 adds the output of each processor in the third cycle and the data at address A1, and stores it at address A1.
The following groups are sequentially read from address A2 of the intermediate storage every cycle, the data stored at the address and the output from each processor in the processor array are added, the output of adder 3 is written to the read address, and the third bus In some cases, the circuit is delayed by two more cycles, that is, mX2=2X2=4 cycles, and from the fifth cycle, the circuit 5.7 is opened and the address Δ
1, the stored data may be sequentially read out, and the adder 3 may add the output from the processor array at each shift to the read stored data, and the output Y may be output in sequence.
即ち、積和演算を行うデータの数nをプロセッサ・アレ
イ2に設けたプロセッサの数mで除し、n/m=6回だ
けデータ列をバスすると共に、1回目のバスではプロセ
ッサ・アレイ2の各プロセッサからの各データと係数と
の積の和を、中間ストレージ4の各アドレスに格納し、
2回目のバスではプロセッサの数mだけ遅れて中間スト
レージ4のアドレスA1からの読出し、及び書込みを行
わせ、プロセッサ・アレーfの出力と中間ストレージか
ら読出したを加Qし中間ス1〜レージΔ1から順に格納
し、以下、各パス毎、中間ストレージ4のアドレスΔ1
から順に読出し、内込むタイミングをプロセッサ−mの
数だけ順次遅ら往、最後のバス時には加p器3の出力を
グー1−回路7を介して出力するようにりればよい。That is, the number n of data to be subjected to the product-sum operation is divided by the number m of processors provided in the processor array 2, the data string is bused only n/m = 6 times, and in the first bus, the data string is bused to the processor array 2. The sum of the products of each data and coefficient from each processor is stored in each address of the intermediate storage 4,
On the second bus, reading and writing from address A1 of intermediate storage 4 are performed with a delay of several m of processors, and the output of processor array f and the read from intermediate storage are added to Q, and the output from intermediate storage 1 to storage Δ1 is From then on, address Δ1 of intermediate storage 4 is stored for each path.
It is only necessary to sequentially read out the signals from 1 and 2, delay the input timing sequentially by the number of processors m, and output the output of the adder 3 through the circuit 7 at the time of the last bus.
その結果、データの数nと等しい数の1【=1セツ1す
がある積和病い方式と比べ、処理時間はd倍になるが、
プロセッサの数はn/dとなり高価なプロセッサの数を
減らすことができる。As a result, compared to the sum-of-products method, which has a number of 1s equal to the number of data n, the processing time will be d times as much.
The number of processors is n/d, which allows the number of expensive processors to be reduced.
発明の効果
以上述べたように、本発明は、積和演尊を行うデータの
数よりも少ない数(データの数の約数)のプロセッサで
積和油筒が行えるため、高価な乗客1を行うプロセッサ
を減らすことができ、積和演棹を安価む装置で行うこと
ができる。Effects of the Invention As described above, the present invention can perform multiplication processing using a smaller number of processors (a divisor of the number of data) than the number of data that performs processing. The number of processors required can be reduced, and the product-sum calculation can be performed using inexpensive equipment.
図は本発明の一実施例のブロック図である。
1・・・被処理データ、2・・・プロセッサ・アレイ、
3・・・加口器、4・・・中間ストレージ、5,6.7
・・・ゲート回路。The figure is a block diagram of one embodiment of the present invention. 1... Processed data, 2... Processor array,
3...Kakuchi device, 4...Intermediate storage, 5,6.7
...Gate circuit.
Claims (1)
和演算方式において、データ列の各データがシフトされ
ながら入力され、入力されたデータに係数を乗じるプロ
セッサをn/d個有し、各プロセッサの出力を加算して
出力するプロセッサ・アレイと、プロセッサ・アレイの
出力と各シフトに対応する中間ストレージのアドレスの
記憶値を加算する加算器と、各シフト毎の加算器の出力
を対応するアドレスに格納する中間ストレージとを有し
、同一データ列を上記プロセッサ・アレイにd回入力し
、データ列を入力する各回毎に各プロセッサに設定する
係数値を決められた値に設定変更し、各回のデータ列の
入力毎に上記中間ストレージからのデータの読出し及び
上記加算器からの出力の書込みをn/d周期遅らせ、加
算器の出力をn個の係数とデータの積の加算結果として
出力するようにした積和演算方式。In the product-sum calculation method in which n coefficients are multiplied by corresponding data and the respective products are added, each data in a data string is input while being shifted, and there are n/d processors that multiply the input data by a coefficient. , a processor array that adds and outputs the output of each processor, an adder that adds the output of the processor array and the stored value of the intermediate storage address corresponding to each shift, and the output of the adder for each shift. and intermediate storage for storing data at corresponding addresses, input the same data string to the processor array d times, and change the coefficient value set in each processor to a predetermined value each time the data string is input. Then, for each data string input, reading data from the intermediate storage and writing the output from the adder is delayed by n/d cycles, and the output of the adder is calculated as the result of adding the product of n coefficients and data. A product-sum calculation method that outputs as .
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP62156427A JPS642176A (en) | 1987-06-25 | 1987-06-25 | Sum of products operating system |
| PCT/JP1988/000636 WO1988010473A1 (en) | 1987-06-25 | 1988-06-25 | System for calculating sum of products |
| EP19880906038 EP0321584A4 (en) | 1987-06-25 | 1988-06-25 | System for calculating sum of products |
| US07/314,055 US4987557A (en) | 1987-06-25 | 1988-06-25 | System for calculation of sum of products by repetitive input of data |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP62156427A JPS642176A (en) | 1987-06-25 | 1987-06-25 | Sum of products operating system |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPH012176A true JPH012176A (en) | 1989-01-06 |
| JPS642176A JPS642176A (en) | 1989-01-06 |
Family
ID=15627510
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP62156427A Pending JPS642176A (en) | 1987-06-25 | 1987-06-25 | Sum of products operating system |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US4987557A (en) |
| EP (1) | EP0321584A4 (en) |
| JP (1) | JPS642176A (en) |
| WO (1) | WO1988010473A1 (en) |
Families Citing this family (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH03192477A (en) * | 1989-12-22 | 1991-08-22 | Nec Corp | Two-dimensional fir digital filtering system and circuit |
| EP0891903B1 (en) * | 1997-07-17 | 2009-02-11 | Volkswagen Aktiengesellschaft | Automatic emergency brake function |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US3961167A (en) * | 1974-07-22 | 1976-06-01 | Gte Automatic Electric Laboratories Incorporated | PCM tone receiver using optimum statistical technique |
| JPS6053349B2 (en) * | 1981-06-19 | 1985-11-25 | 株式会社日立製作所 | image processing processor |
| US4489393A (en) * | 1981-12-02 | 1984-12-18 | Trw Inc. | Monolithic discrete-time digital convolution circuit |
| JPS58181171A (en) * | 1982-04-16 | 1983-10-22 | Hitachi Ltd | Parallel picture processing processor |
| JPS581275A (en) * | 1982-06-14 | 1983-01-06 | Hitachi Ltd | Shading circuit |
| JPS62105287A (en) * | 1985-11-01 | 1987-05-15 | Fanuc Ltd | Signal processor |
-
1987
- 1987-06-25 JP JP62156427A patent/JPS642176A/en active Pending
-
1988
- 1988-06-25 US US07/314,055 patent/US4987557A/en not_active Expired - Fee Related
- 1988-06-25 WO PCT/JP1988/000636 patent/WO1988010473A1/en not_active Ceased
- 1988-06-25 EP EP19880906038 patent/EP0321584A4/en not_active Withdrawn
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US4635292A (en) | Image processor | |
| US4601006A (en) | Architecture for two dimensional fast fourier transform | |
| US4791598A (en) | Two-dimensional discrete cosine transform processor | |
| US4644488A (en) | Pipeline active filter utilizing a booth type multiplier | |
| US4945506A (en) | High speed digital signal processor capable of achieving realtime operation | |
| JPS6053349B2 (en) | image processing processor | |
| JPS6247786A (en) | Exclusive memory for adjacent image processing | |
| JP7435602B2 (en) | Computing equipment and computing systems | |
| US4853887A (en) | Binary adder having a fixed operand and parallel-serial binary multiplier incorporating such an adder | |
| JPS63167967A (en) | Digital signal processing integrated circuit | |
| US5805476A (en) | Very large scale integrated circuit for performing bit-serial matrix transposition operation | |
| JPH012176A (en) | Product-sum calculation method | |
| JPH07152730A (en) | Discrete cosine transform device | |
| JP3523315B2 (en) | Digital data multiplication processing circuit | |
| EP0321584A1 (en) | System for calculating sum of products | |
| JPH06223166A (en) | General processor for image processing | |
| RU2709160C1 (en) | Triangular matrix handling device | |
| JP3875183B2 (en) | Arithmetic unit | |
| JP2953918B2 (en) | Arithmetic unit | |
| JP2568179B2 (en) | Interpolation enlargement calculation circuit | |
| JPS63164640A (en) | Cosine transformation device | |
| KR940007569B1 (en) | Matrix multiplication circuit | |
| JPS58163061A (en) | Parallel image processing processor and device | |
| JPH0298777A (en) | Parallel sum of product arithmetic circuit and vector matrix product arithmetic method | |
| JPS6280775A (en) | Image processor |