JPH11110215A - Information processor using multi thread program - Google Patents
Information processor using multi thread programInfo
- Publication number
- JPH11110215A JPH11110215A JP28766297A JP28766297A JPH11110215A JP H11110215 A JPH11110215 A JP H11110215A JP 28766297 A JP28766297 A JP 28766297A JP 28766297 A JP28766297 A JP 28766297A JP H11110215 A JPH11110215 A JP H11110215A
- Authority
- JP
- Japan
- Prior art keywords
- thread
- instruction
- information processing
- arithmetic
- storage means
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
- 238000012545 processing Methods 0.000 claims abstract description 67
- 238000003860 storage Methods 0.000 claims abstract description 52
- 230000010365 information processing Effects 0.000 claims abstract description 44
- 238000012546 transfer Methods 0.000 claims description 22
- 238000004891 communication Methods 0.000 claims description 14
- 238000004364 calculation method Methods 0.000 claims description 11
- 230000005540 biological transmission Effects 0.000 claims description 10
- 230000004044 response Effects 0.000 claims description 8
- 230000015654 memory Effects 0.000 abstract description 271
- 238000000034 method Methods 0.000 description 72
- 238000012937 correction Methods 0.000 description 30
- 238000007667 floating Methods 0.000 description 27
- 238000010586 diagram Methods 0.000 description 25
- 230000008569 process Effects 0.000 description 25
- 230000007246 mechanism Effects 0.000 description 22
- 230000006870 function Effects 0.000 description 19
- 230000002829 reductive effect Effects 0.000 description 15
- 230000008859 change Effects 0.000 description 12
- 238000009826 distribution Methods 0.000 description 9
- 238000013500 data storage Methods 0.000 description 8
- 230000000694 effects Effects 0.000 description 8
- 238000005516 engineering process Methods 0.000 description 8
- 230000008901 benefit Effects 0.000 description 7
- 238000002360 preparation method Methods 0.000 description 7
- 230000006872 improvement Effects 0.000 description 5
- 230000009467 reduction Effects 0.000 description 3
- 230000002457 bidirectional effect Effects 0.000 description 2
- 238000013461 design Methods 0.000 description 2
- 238000007726 management method Methods 0.000 description 2
- 238000009738 saturating Methods 0.000 description 2
- 239000004065 semiconductor Substances 0.000 description 2
- 238000013459 approach Methods 0.000 description 1
- 230000006399 behavior Effects 0.000 description 1
- 238000007796 conventional method Methods 0.000 description 1
- 238000012217 deletion Methods 0.000 description 1
- 230000037430 deletion Effects 0.000 description 1
- 238000000605 extraction Methods 0.000 description 1
- 230000002401 inhibitory effect Effects 0.000 description 1
- 238000009434 installation Methods 0.000 description 1
- 230000000670 limiting effect Effects 0.000 description 1
- 238000005457 optimization Methods 0.000 description 1
- 230000002093 peripheral effect Effects 0.000 description 1
- 230000001737 promoting effect Effects 0.000 description 1
- 230000000717 retained effect Effects 0.000 description 1
- 229920006395 saturated elastomer Polymers 0.000 description 1
- 230000003068 static effect Effects 0.000 description 1
- 239000000725 suspension Substances 0.000 description 1
- 230000001360 synchronised effect Effects 0.000 description 1
Landscapes
- Advance Control (AREA)
Abstract
Description
【0001】[0001]
【産業上の利用分野】本発明は、デジタル回路を使用し
た情報処理装置に関し、特にマイクロプロセッサおよび
それを利用した情報処理システムに関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to an information processing apparatus using a digital circuit, and more particularly to a microprocessor and an information processing system using the same.
【0002】[0002]
【従来の技術】汎用マイクロプロセッサの性能を向上さ
せるためには、主に2つの方法がある。動作周波数の向
上、および並列化である。動作周波数の向上には、半導
体プロセス技術の改善と、パイプライン等の回路構成の
変更が考えられる。前者の進化は今後も続くものと考え
られるが、後者はメモリの速度等がネックになって向上
はほとんど見込めない。よって今後は、プロセス技術の
改善と並列化によって性能を向上させることになる。2. Description of the Related Art There are mainly two methods for improving the performance of general-purpose microprocessors. Improvement of operating frequency and parallelization. The operating frequency can be improved by improving semiconductor process technology and changing the circuit configuration such as a pipeline. The former is expected to continue to evolve in the future, but the latter is hardly expected to improve because of the speed of the memory. Therefore, in the future, performance will be improved by improving process technology and parallelizing.
【0003】ところが並列化には、さまざまな制限があ
る。プログラムの互換性を保ったまま並列化する手段
は、スーパスカラ方式として良く知られているが、スー
パスカラ方式は、命令の並列配分のための回路が本体の
演算装置より大規模になり、消費電力の点で優れている
とは言えない。少なくとも、消費電力の削減という観点
では、プログラムの互換性という制限がなければ避ける
べきアプローチといえる。そして、スーパースカラ方式
は、プログラムの局所的な並列抽出である。それは、プ
ログラムの中でデータ依存関係の一番多い部分をあえて
並列化する方式と言える。そのため、並列化を行うハー
ドウェアの増加に対して、それによる性能向上率は飽和
しつつある。However, parallelization has various limitations. Means for performing parallelization while maintaining program compatibility is well known as the superscalar method.However, the superscalar method requires a larger circuit for parallel allocation of instructions than the main processing unit, and consumes less power. It is not good in terms of point. At least from the viewpoint of reducing power consumption, it can be said that this approach should be avoided unless there are restrictions on program compatibility. The superscalar method is local parallel extraction of a program. It can be said that it is a method to intentionally parallelize the part with the most data dependencies in the program. Therefore, the performance improvement rate due to the increase in hardware for parallelization is saturating.
【0004】そこで、プログラムの互換性をあきらめて
並列化を進める方式の代表的なものが次に示す、従来例
1のVLIW方式と、従来例2の共有メモリマルチプロ
セッサ方式である。[0004] Therefore, typical systems for giving up compatibility of programs and proceeding with parallelization are the following VLIW system of the first conventional example and the shared memory multiprocessor system of the second conventional example.
【0005】<従来例1>図2を参照して、本発明の元
となった従来のVLIW型マイクロプロセッサの模式図
を示す。以下この例を従来例1とする。<Conventional Example 1> Referring to FIG. 2, a schematic diagram of a conventional VLIW type microprocessor on which the present invention is based is shown. Hereinafter, this example is referred to as Conventional Example 1.
【0006】VLIW型マイクロプロセッサは、命令ア
ドレス生成ユニット201、そして命令キャッシュ20
2、命令デコード203、データクロスバスイッチ20
4、分岐ユニット205、ロードストアユニット20
6、2つの演算器207、多ポートレジスタファイル2
08、データキャッシュ209、外部バスインターフェ
ース210で構成される。The VLIW type microprocessor includes an instruction address generation unit 201 and an instruction cache 20.
2. Instruction decode 203, data crossbar switch 20
4, branch unit 205, load store unit 20
6, two operation units 207, multi-port register file 2
08, a data cache 209, and an external bus interface 210.
【0007】VLIW型マイクロプロセッサは、命令キ
ャッシュ202から発行される命令が同時に分岐ユニッ
ト205、ロードストアユニット206、2つの演算機
207を制御する。そのため、命令キャッシュ202に
は複数の命令が1つのラインに混在して格納される。In the VLIW microprocessor, instructions issued from the instruction cache 202 simultaneously control the branch unit 205, the load store unit 206, and the two arithmetic units 207. Therefore, the instruction cache 202 stores a plurality of instructions in a single line.
【0008】そして、分岐ユニット205、ロードスト
アユニット206、2つの演算器207は全て1つのレ
ジスタファイル208を共有し、データクロスバスイッ
チ204によって配分される。分岐ユニット205、ロ
ードストアユニット207、2つの演算器207は同時
に動作できるため、すべての入力データ、出力データは
レジスタファイル208に接続される。The branch unit 205, the load store unit 206, and the two arithmetic units 207 all share one register file 208, and are distributed by the data crossbar switch 204. Since the branch unit 205, the load store unit 207, and the two arithmetic units 207 can operate simultaneously, all input data and output data are connected to the register file 208.
【0009】VLIW型マイクロプロセッサの特徴は、
演算機207やロードストアユニット206等を複数同
時に制御することを可能にしつつ、命令デコード203
の複雑化を抑制できることにある。ところが、さらに演
算機207を増加させると、演算器間の転送バスの数、
レジスタファイル208のポート数が極度に増大すると
いう問題が生じる。これが回路規模の増加やクロック速
度の低下をもたらすことになる。この欠点を解消するた
めにはレジスタファイル208やデータクロスバスイッ
チ204を分割するしかない。しかし、分割のために
は、同時に実行される命令列の中に、互いにデータ依存
関係のない独立した処理が共存していなければならな
い。VLIW型マイクロプロセッサでは、あらかじめ独
立した処理を、コンパイラあるいはプログラマが事前に
1つの命令ラインに混在させて命令キャッシュ202に
格納しておく必要がある。これは超並列における動作の
自由度を失わせる結果になる。要約すると、今後の並列
処理性能の向上には、通常の依存関係のある複数のプロ
グラムを、動的に組み合わせて動作できる構造が必要に
なる。The features of the VLIW type microprocessor are as follows.
The instruction decoder 203 is capable of controlling a plurality of arithmetic units 207 and load store units 206 at the same time.
Is to suppress the complexity of However, when the number of arithmetic units 207 is further increased, the number of transfer buses between the arithmetic units,
There is a problem that the number of ports in the register file 208 is extremely increased. This leads to an increase in circuit scale and a decrease in clock speed. The only solution to this disadvantage is to divide the register file 208 and the data crossbar switch 204. However, for the purpose of division, independent processes having no data dependency must coexist in an instruction sequence executed at the same time. In the VLIW type microprocessor, it is necessary for a compiler or a programmer to mix independent processing in one instruction line beforehand and store it in the instruction cache 202 in advance. This results in a loss of freedom of operation in massively parallel. In summary, in order to improve parallel processing performance in the future, it is necessary to have a structure capable of operating by dynamically combining a plurality of programs having ordinary dependencies.
【0010】また、VLIW型マイクロプロセッサは、
特定アプリケーションに特化した回路を演算機と並列に
組み込まれることが多い。これは、メディアプロセッサ
と呼ばれるマルチメディア専用のマイクロプロセッサで
多くみられる方式である。メディアプロセッサの目的
は、専用回路の高速性と汎用マイクロプロセッサが持つ
自由度の双方を兼ね備えることある。しかし、今後の性
能向上のために、専用演算機を大量に搭載することにな
ると、それらの専用演算機とのデータのやりとりをする
前後の処理が問題となる。専用回路の組み合わせの自由
度と性能を維持するためには、専用回路間の仲立ちをす
る整数演算機207やロードストアユニット206を増
加させる必要がある。しかし、それには先に述べたVL
IW型に依存する並列化の限界がある。だが、これらを
利用せず、専用回路同士を直接接続すると、組み合わせ
の自由度がなくなり汎用性を失うことになる。今後のメ
ディアプロセッサには、専用回路間の接続を動的に変更
でき、かつ高いバンド幅のデータを処理できる手段が求
められている。The VLIW type microprocessor is
A circuit specialized for a specific application is often incorporated in parallel with the arithmetic unit. This is a method often used in a microprocessor dedicated to multimedia called a media processor. The purpose of a media processor is to provide both the high speed of a dedicated circuit and the flexibility of a general-purpose microprocessor. However, if a large number of special-purpose processing units are to be mounted in order to improve performance in the future, processing before and after data exchange with these special-purpose processing units becomes a problem. In order to maintain the degree of freedom and the performance of the combination of the dedicated circuits, it is necessary to increase the number of the integer arithmetic units 207 and the load store units 206 which mediate between the dedicated circuits. However, it is based on the VL
There is a limit of parallelization depending on the IW type. However, if these circuits are not used and the dedicated circuits are directly connected to each other, the degree of freedom of combination is lost and the versatility is lost. Future media processors will require means that can dynamically change the connections between dedicated circuits and process high bandwidth data.
【0011】<従来例2>図3を参照して、本発明の元
となった従来のマイクロプロセッサを複数使用した構成
の模式図を示す。以下、この例を従来例2とする。<Conventional Example 2> Referring to FIG. 3, a schematic diagram of a configuration using a plurality of conventional microprocessors based on which the present invention is based is shown. Hereinafter, this example is referred to as Conventional Example 2.
【0012】図3は、共有メモリ型マルチプロセッサと
呼ばれる構成である。データ共有のための機構を備えた
マイクロプロセッサ301を複数搭載したシステムであ
る。図では3つ搭載され、共有バス304に接続されて
いる。メモリバンク305、306は共に共有バス30
4に接続され、同時にアクセスすることはできない。ま
た、使用頻度の低い演算を共有するため、共有機能ユニ
ット307は共有バス304に接続される。I/Oペリ
フェラル308も共有バス304に接続される。FIG. 3 shows a configuration called a shared memory type multiprocessor. This is a system equipped with a plurality of microprocessors 301 having a mechanism for sharing data. In the figure, three are mounted and connected to the shared bus 304. The memory banks 305 and 306 are both shared bus 30
4 and cannot be accessed simultaneously. The shared function unit 307 is connected to the shared bus 304 in order to share the infrequently used operations. The I / O peripheral 308 is also connected to the shared bus 304.
【0013】単体のマイクロプロセッサ301は、命令
キャッシュ311、命令発行ユニット312、整数演算
器313、浮動小数点演算器314、ロードストアユニ
ット315、データキャッシュ316、キャッシュコヒ
ーレンシ制御機構317を全て内蔵する。キャッシュコ
ヒーレンシ制御機構317は、共有バス304の内容を
常にスヌープし、自身のデータキャッシュ316が持つ
コピーの共有状態を制御する。A single microprocessor 301 includes an instruction cache 311, an instruction issuing unit 312, an integer arithmetic unit 313, a floating point arithmetic unit 314, a load store unit 315, a data cache 316, and a cache coherency control mechanism 317. The cache coherency control mechanism 317 always snoops the contents of the shared bus 304, and controls the copy sharing state of its own data cache 316.
【0014】この共有メモリマルチプロセッサ方式は、
従来例1のVLIW方式と違い、マイクロプロセッサは
それぞれ全く独立した処理を自由な組み合わせで処理で
きる。そして、この共有メモリマルチプロセッサ方式
は、他の方式のマルチプロセッサに対して次のような長
所を持つ。単体のマイクロプロセッサの為に作成された
ソフトウェアがある程度使用できる、動的なプロセッサ
間通信のプログラムを作成しやすい、マイクロプロセッ
サごとのローカルキャッシュによって実際の通信のバン
ド幅を縮小できる、等である。This shared memory multiprocessor system is
Unlike the VLIW method of the conventional example 1, the microprocessors can perform completely independent processing in any combination. The shared memory multiprocessor system has the following advantages over multiprocessors of other systems. Software created for a single microprocessor can be used to some extent, dynamic inter-processor communication programs can be easily created, and the actual communication bandwidth can be reduced by a local cache for each microprocessor.
【0015】この共有メモリマルチプロセッサ方式の欠
点としては、性能を向上させるのに必要なハードウェア
が最も複雑になることが挙げられる。A disadvantage of the shared memory multiprocessor system is that the hardware required to improve the performance becomes the most complicated.
【0016】まず、すべてのマイクロプロセッサ301
が命令キャッシュ311を持つことが挙げられる。ルー
プなどの中粒度の並列をマルチプロセッサで実行する場
合、同じプログラムのコピーをそれぞれの命令キャッシ
ュ311に持つことになる。そして、命令キャッシュリ
プレースのバンド幅もプロセッサの数だけ必要になる。
ループは同じ命令を再利用することであるため、命令メ
モリのマイクロプロセッサ301間の共有が望ましい。First, all the microprocessors 301
Has an instruction cache 311. When a medium-level parallel processing such as a loop is executed by a multiprocessor, each instruction cache 311 has a copy of the same program. Also, the bandwidth of the instruction cache replacement is required by the number of processors.
Since the loop is to reuse the same instruction, sharing of instruction memory between microprocessors 301 is desirable.
【0017】次に、データキャッシュ316について述
べる。バスを共有するので、大規模なローカルキャッシ
ュ、およびキャッシュコヒーレンシ調停機構を搭載する
必要がある。スレッドはプロセッサ301〜303の間
で非決定的に配分され、データキャッシュコヒーレンシ
制御317、共有バス304を使用してデータを転送す
る必要がある。マイクロプロセッサ301の並列度を増
やすと、共有バスの転送能力の限界によって性能が制限
される。ところが、バスを分割してバンド幅を稼ぐよう
にすると、双方に対してキャッシュの制御を行うため、
ハードウェアの増加は甚大なものになる。メモリバンク
305を自然にマイクロプロセッサ301に直結できる
構造が望ましい。Next, the data cache 316 will be described. Since the bus is shared, it is necessary to mount a large-scale local cache and a cache coherency arbitration mechanism. The threads are non-deterministically distributed among the processors 301 to 303, and need to transfer data using the data cache coherency control 317 and the shared bus 304. When the parallelism of the microprocessor 301 is increased, the performance is limited by the limit of the transfer capability of the shared bus. However, if the bus is divided to gain bandwidth, cache control is performed for both buses,
The increase in hardware will be enormous. It is desirable that the memory bank 305 can be directly connected to the microprocessor 301 naturally.
【0018】そして、演算機の共有が難しいため、すべ
てのマイクロプロセッサが浮動小数点演算ユニット31
4を有する。ところが、浮動小数点演算ユニット314
の命令の発生頻度や稼働率は、特殊なアプリケーション
を除けば整数演算器313より低い。そのために、プロ
セッサの数だけ専用演算機を搭載するのは冗長と言え
る。しかし、機能ユニット307のように、共有バス3
04に接続すると、マイクロプロセッサ301との転送
バンド幅が不足することになる。そのため、専用演算器
の共有とマイクロプロセッサとの転送バンド幅の両立が
求められる。Since it is difficult to share the arithmetic units, all of the microprocessors
4 However, the floating point arithmetic unit 314
The occurrence frequency and operation rate of the instruction are lower than that of the integer arithmetic unit 313 except for special applications. For this reason, it can be said that mounting dedicated arithmetic units as many as the number of processors is redundant. However, like the functional unit 307, the shared bus 3
04, the transfer bandwidth with the microprocessor 301 becomes insufficient. For this reason, it is required to achieve both the sharing of the dedicated arithmetic unit and the transfer bandwidth with the microprocessor.
【0019】[0019]
【発明が解決しようとする課題】プログラムの互換性を
保ったまま並列化を進める方式はスーパースカラ方式と
して知られるが、この方式は回路規模に対する性能比が
極端に低くなる。そこで、プログラムの互換性をあきら
めて性能を出す方式の代表的なものが次に示す、従来例
1のVLIW方式と、従来例2の共有メモリマルチプロ
セッサ方式である。A method of promoting parallelization while maintaining program compatibility is known as a superscalar method, but this method has an extremely low performance ratio to the circuit scale. Therefore, the typical systems that give up the performance by giving up the compatibility of the program are the VLIW system of the conventional example 1 and the shared memory multiprocessor system of the conventional example 2 shown below.
【0020】しかし、従来例1のVLIW方式は並列数
に上限があり、かつ並列プログラミングにおける汎用性
が低く、垂直性がない。そして、従来例2の共有メモリ
マルチプロセッサはは従来のマイクロプロセッサの回路
を全て内蔵するため、冗長な回路が多く、消費電力や回
路規模は最大になる。また、バスを共有するため、バス
のバンド幅に制限されて、並列度に対する性能が飽和す
ることが欠点といえる。However, the VLIW system of the prior art 1 has an upper limit on the number of parallel operations, is low in versatility in parallel programming, and lacks verticality. Since the shared memory multiprocessor of Conventional Example 2 incorporates all the circuits of the conventional microprocessor, there are many redundant circuits, and the power consumption and the circuit scale are maximized. In addition, since the bus is shared, the drawback is that the performance with respect to the parallelism is saturated due to the limitation of the bus bandwidth.
【0021】よって、並列度に上限がなく、プログラム
構造が従来のマイクロプロセッサに近く、かつ同じプロ
グラムで並列性能を生かすことができ、構造が単純で消
費電力の少ないマイクロプロセッサが求められている。Therefore, there is a demand for a microprocessor having a simple structure and low power consumption, which has no upper limit on the degree of parallelism, has a program structure close to that of a conventional microprocessor, can utilize the parallel performance with the same program, and has a simple structure.
【0022】<自由度を維持しつつ演算能力の高い汎用
マイクロプロセッサの要求>専用に設計された回路は汎
用プロセッサと比較して常に高速である。しかし、それ
は単機能を実現するという前提に基づく。しかし、今後
要求されるアプリケーションは、多くの機能を実現する
ことが求められる。そのためには、専用回路を機能の数
だけ用意することになる。だが、複雑化の一途をたどる
アプリケーションの機能の全てをまかなうには、専用回
路を全て提供することは不可能である。そのため、それ
ほど性能を必要としない機能については、順次汎用マイ
クロプロセッサによって取って換わるようになった。専
用回路を使用する必要があるものにおいても、マイクロ
プロセッサが複数の専用回路の仲立ちをすることで自由
度を確保している。<Requirement of General-purpose Microprocessor with High Computing Ability while Maintaining Degree of Freedom> A circuit designed specifically for use is always faster than a general-purpose processor. However, it is based on the premise that a single function is realized. However, applications required in the future are required to realize many functions. For that purpose, dedicated circuits are prepared by the number of functions. However, it is impossible to provide all dedicated circuits to cover all of the increasingly complex application functions. Therefore, functions that do not require much performance have been successively replaced by general-purpose microprocessors. Even in the case where a dedicated circuit needs to be used, the degree of freedom is ensured by the microprocessor mediating a plurality of dedicated circuits.
【0023】専用回路を複数使用する機能は、専用回路
同士の接続の自由度とバンド幅を同時に確保する必要が
ある。しかし、バス転送速度や仲立ちをするマイクロプ
ロセッサの性能向上は飽和しつつある。よって、バスと
マイクロプロセッサの分割が必要になる。ところが、バ
スやマイクロプロセッサの分割は、従来例2に見られる
とおり、負荷分散、同期機構が非常に複雑になる。バス
やマイクロプロセッサを分割して局所的な性能を維持し
つつ、複数の専用回路をすべてのプロセッサで共有でき
るマイクロプロセッサが理想的である。The function of using a plurality of dedicated circuits needs to simultaneously secure the degree of freedom of connection between the dedicated circuits and the bandwidth. However, the improvement in bus transfer speed and the performance of a middleman microprocessor is saturating. Therefore, it is necessary to divide the bus and the microprocessor. However, the division of the bus and the microprocessor makes the load distribution and the synchronization mechanism very complicated as seen in the second conventional example. A microprocessor that can share a plurality of dedicated circuits with all processors while maintaining local performance by dividing a bus or a microprocessor is ideal.
【0024】<消費電力に対して高性能なハードウェア
の要求>コストに対して最高性能のマイクロプロセッサ
は、単一プロセッサと互換性の制約が無ければ、消費電
力、チップコストに対して最高の性能を出すことができ
るプロセッサであると言える。単体のコストが低けれ
ば、そのプロセッサを大量に使用すれば最高性能のシス
テムを構築できる。<Requirement of High-Performance Hardware for Power Consumption> A microprocessor having the highest performance for cost has the highest power consumption and chip cost if there is no compatibility constraint with a single processor. It can be said that it is a processor that can provide high performance. If the cost of a single unit is low, the highest performance system can be constructed by using a large amount of the processor.
【0025】つまり、マイクロプロセッサを突き詰めれ
ば低消費電力になる。低消費電力のための手段はアーキ
テクチャレベル、回路レベル、プロセスレベルにそれぞ
れ存在する。その中でも、アーキテクチャレベルの手段
は、直接処理性能に関係ない回路や配線を極力無くすこ
とに尽きる。In other words, the power consumption can be reduced if the microprocessor is squeezed. Means for low power consumption exist at the architecture level, the circuit level, and the process level, respectively. Among them, the means at the architectural level is simply to eliminate circuits and wirings not directly related to the processing performance as much as possible.
【0026】既存のスーパースカラ型マイクロプロセッ
サは、既存のプログラムを単一のマイクロプロセッサで
動作させることを至上命題としているため、既存のプロ
グラムを内部の命令レベル並列マイクロプロセッサに翻
訳する回路が存在する。そして、近年はこの回路が巨大
化する一方である。これでは低消費電力と性能は両立で
きない。Since the existing superscalar-type microprocessor has a paradigm of operating an existing program on a single microprocessor, there is a circuit for translating the existing program into an internal instruction-level parallel microprocessor. . In recent years, this circuit has been increasing in size. This makes low power consumption and performance incompatible.
【0027】マイクロプロセッサの本質は、演算器とメ
モリに尽きる。それ以外は基本的に冗長な回路である。
その冗長な回路を自由度を維持しつつ最小限にすること
が、アーキテクチャレベルの低消費電力につながる。The essence of a microprocessor is an arithmetic unit and a memory. The rest is basically a redundant circuit.
Minimizing the redundant circuit while maintaining the degree of freedom leads to low power consumption at the architecture level.
【0028】<要求されるソフトウェアの変化>並列プ
ログラムを記述する方法としては、古くはベクトルプロ
セッサ方式スーパーコンピュータに使用されたループの
コンパイラによるベクトル展開、最近ではコンパイラに
よるSIMD命令やVLIWへの展開などの技術があ
る。この方式は、ループのような特別な処理のみが並列
に動作するので、ループの前後の処理が性能の足かせに
なる。このことはAmdahlの法則として広く知られ
る。<Changes in Required Software> As a method of describing a parallel program, vector expansion by a compiler using a loop used in a vector processor type supercomputer in the past, expansion into SIMD instructions and VLIW by a compiler recently, etc. Technology. In this method, only special processing such as a loop operates in parallel, and processing before and after the loop hinders performance. This is widely known as Amdahl's law.
【0029】しかし、現在は、並列処理を行うプログラ
ムの汎用性、処理の切り替えの高速化、処理の間の高速
通信を両立できる技術である、マルチスレッドと呼ばれ
る手法が主流になりつつある。このマルチスレッドは、
従来例2で示した共有メモリマルチプロセッサに最適な
ように作成されている。マルチスレッドはOSに登録し
ておけば、自動的に空いたマイクロプロセッサに負荷分
散される。マルチスレッドプログラムは、スレッド間の
同期は明示的に記述する必要があり、それ以外は独立し
て動作することを保証するため、ループの並列展開と比
較してAmdahlの法則の影響を減らすことができ
る。このマルチスレッドプログラミングモデルを使用す
れば、十分な汎用性と超並列への負荷分散を両立するこ
とができる。更に、マルチスレッドはSIMD命令など
との両立も可能である。However, at present, a technique called multithreading, which is a technique capable of achieving both versatility of a program for performing parallel processing, high-speed switching of processing, and high-speed communication during processing, is becoming mainstream. This multi-thread,
It is created so as to be optimal for the shared memory multiprocessor shown in the conventional example 2. If the multithread is registered in the OS, the load is automatically distributed to the free microprocessor. In a multi-threaded program, the effect of Amdahl's law can be reduced compared to parallel unrolling of loops to ensure that synchronization between threads must be explicitly described and otherwise operate independently. it can. By using this multi-thread programming model, sufficient versatility and load balancing to massively parallel can be achieved at the same time. Further, multi-threading can be compatible with SIMD instructions and the like.
【0030】データと処理を対にするのは、オブジェク
ト指向と呼ばれる技術の一部であり、大規模化、分散化
される今後のアプリケーションを支える技術である。こ
のオブジェクト指向に適合するように、マイクロプロセ
ッサとメモリで構成されるプロセッサ単位を作成する。
そして、処理能力の必要なオブジェクトをそのプロセッ
サ単位の1つ、あるいは複数に割り当て、処理能力の必
要ないオブジェクトは、共有のプロセッサ単位に割り当
てる。この方式は、オブジェクト指向のソフトウェアに
とって一番自然な負荷分散の形態の1つと言える。The pairing of data and processing is a part of a technology called object-oriented technology, which is a technology that supports future applications that will be scaled up and distributed. A processor unit composed of a microprocessor and a memory is created so as to conform to the object orientation.
Then, an object requiring a processing capability is assigned to one or more of the processor units, and an object requiring no processing capability is assigned to a shared processor unit. This method is one of the most natural forms of load distribution for object-oriented software.
【0031】近年は、マイクロプロセッサの主な用途
は、科学技術計算からマルチメディアアプリケーション
に移行しつつある。このマルチメディアアプリケーショ
ンは、定型処理が多く、大量のデータ転送能力を要求す
る傾向がある。それでいて、大量の処理の間でデータ依
存関係が比較的少ないという特徴があり、データと処理
の分散化は容易と言える。そして、複数の処理が同時に
動作し、その組み合わせは非決定的である。よって、非
決定的な処理の組み合わせを効率的に行うためには、専
用回路の導入とともに、その仲立ちをする汎用マイクロ
プロセッサの性能向上が不可欠である。In recent years, the main use of microprocessors has been shifting from scientific computing to multimedia applications. This multimedia application tends to require a large amount of data transfer capability with a lot of routine processing. Nevertheless, there is a characteristic that data dependency is relatively small between a large amount of processing, and it can be said that the distribution of data and processing is easy. Then, a plurality of processes operate simultaneously, and the combination is non-deterministic. Therefore, in order to efficiently perform the non-deterministic combination of processes, it is essential to improve the performance of a general-purpose microprocessor that mediates with the introduction of a dedicated circuit.
【0032】そして、マルチメディアアプリケーション
は高速リアルタイム性能が求められる。現在のところ、
複雑なプライオリティーを使用したプロセススケジュー
リングを採用しているが、今後は、プロセスに対して複
数のマイクロプロセッサを一定量だけ静的に割り当て
て、マイクロプロセッサの割り当て数で負荷分散とする
ような単純な構造の方が望ましい。プライオリティーを
設定するのは、一定の演算資源の割り当てを必要とする
からである。[0032] Multimedia applications require high-speed real-time performance. at present,
Although process scheduling using complex priorities is adopted, in the future, multiple microprocessors will be statically allocated to a process by a fixed amount and the load will be distributed by the number of allocated microprocessors. A simple structure is more desirable. The reason for setting the priority is that it is necessary to allocate a certain calculation resource.
【0033】結論として、マルチスレッドプログラミン
グによる汎用性と、オブジェクト指向による明示的な負
荷分散を可能にするマイクロプロセッサが望ましいと考
えられる。In conclusion, it would be desirable to have a microprocessor that allows for versatility with multi-thread programming and explicit load balancing in an object-oriented manner.
【0034】[0034]
【発明の実施の形態】次に、本発明について図面を参照
して説明する。Next, the present invention will be described with reference to the drawings.
【0035】<実施例1>図1を参照する。本発明のマ
イクロプロセッサ1は、命令発行制御手段2を1つ以上
持つ。そして、プログラムカウンタ記憶手段3、命令格
納手段4、演算手段5、データ格納手段7を1セットに
して複数内蔵する。そして、演算手段5とデータ格納手
段7と、外部インターフェース手段8は、1つの動的信
号接続手段6に接続される。<Embodiment 1> Referring to FIG. The microprocessor 1 of the present invention has one or more instruction issuing control means 2. Then, a plurality of the program counter storage means 3, the instruction storage means 4, the arithmetic means 5, and the data storage means 7 are incorporated as a set. Then, the arithmetic means 5, the data storage means 7, and the external interface means 8 are connected to one dynamic signal connection means 6.
【0036】命令発行制御手段2は、外部インターフェ
ース8からの要求、あるいは演算ユニット5の要求に応
じて動作する。命令発行制御手段2は、命令格納手段4
のプログラムカウンタを最初のプログラムカウンタ記憶
手段3に転送して演算処理を開始する。The instruction issuance control means 2 operates in response to a request from the external interface 8 or a request from the arithmetic unit 5. The instruction issuance control means 2 includes an instruction storage means 4
Is transferred to the first program counter storage means 3 to start the arithmetic processing.
【0037】プログラムカウンタ記憶手段3は、直列に
接続され、個々に命令格納手段4に接続される。演算ユ
ニットの演算の終了と同時に右隣のプログラムカウンタ
記憶手段2にプログラムカウンタを送信し、左隣のプロ
グラムカウンタ記憶手段3から、更新されたプログラム
カウンタを受信する。右端のプログラムカウンタ記憶手
段3は、命令発行制御2に接続され、一連の演算の終了
を通知する。The program counter storage means 3 is connected in series and individually connected to the instruction storage means 4. At the same time as the completion of the operation of the arithmetic unit, the program counter is transmitted to the program counter storage means 2 on the right side, and the updated program counter is received from the program counter storage means 3 on the left side. The rightmost program counter storage means 3 is connected to the instruction issuance control 2 and notifies the end of a series of operations.
【0038】命令格納手段4は、全て演算手段5に接続
され、それぞれ演算手段5の機能を選択する命令コード
を伝達する。読み出される命令コードは、プログラムカ
ウンタ記憶手段3からの値によって選択される。The instruction storage means 4 are all connected to the operation means 5 and each transmit an instruction code for selecting a function of the operation means 5. The instruction code to be read is selected according to the value from the program counter storage means 3.
【0039】演算手段5は、命令格納手段4の命令を受
理して演算処理を行い、演算に使用する値を動的信号接
続手段6から読み出す。演算結果は、右に隣接する演算
手段5、あるいは動的信号接続手段6に伝達する。The operation means 5 receives the instruction from the instruction storage means 4 and performs an operation process, and reads out a value used for the operation from the dynamic signal connection means 6. The calculation result is transmitted to the calculation means 5 adjacent to the right or the dynamic signal connection means 6.
【0040】動的信号接続手段6は、すべての演算手段
5と、データ格納手段7と、外部インターフェース手段
7同士で同時に通信を行うことができる。動的信号接続
手段6は、演算手段5からの要求に応じて、すべてのデ
ータ格納手段7と外部インターフェース手段7の中から
1つを選択して、データ値を演算手段5に転送する。ま
たは、逆に演算結果を1つのデータ格納手段7か外部イ
ンターフェース手段7を選択して転送する。The dynamic signal connection means 6 allows all the arithmetic means 5, the data storage means 7, and the external interface means 7 to simultaneously communicate with one another. The dynamic signal connection means 6 selects one of all the data storage means 7 and the external interface means 7 in response to a request from the calculation means 5 and transfers the data value to the calculation means 5. Alternatively, the calculation result is transferred by selecting one of the data storage means 7 or the external interface means 7.
【0041】データ格納手段7は、同時刻には1つの読
み出し、あるいは1つの書き込みだけを処理できれば十
分である。すべての演算手段5は、同時に同じデータ格
納手段7、あるいは外部インターフェース手段7と通信
することがないように命令発行制御2によって調停でき
る。It is sufficient that the data storage means 7 can process only one read or one write at the same time. All the arithmetic means 5 can be arbitrated by the instruction issuing control 2 so as not to communicate with the same data storage means 7 or the external interface means 7 at the same time.
【0042】外部インターフェース手段8は、演算手段
5の要求に応じてマイクロプロセッサ1の演算に必要な
データを、マイクロプロセッサ1の外部から読み出す。
あるいは、演算手段5の演算結果をマイクロプロセッサ
1の外部へ出力する。The external interface means 8 reads data necessary for the operation of the microprocessor 1 from outside the microprocessor 1 in response to a request from the operation means 5.
Alternatively, the operation result of the operation means 5 is output to the outside of the microprocessor 1.
【0043】<実施例2> <マイクロプロセッサ401の内部構造>図4を参照し
て本発明の第二の実施例のマイクロプロセッサ401に
ついて概説する。<Second Embodiment><Internal Structure of Microprocessor 401> Referring to FIG. 4, a microprocessor 401 according to a second embodiment of the present invention will be outlined.
【0044】4つの整数演算ユニット群405と、4つ
のスレッド制御ユニット403は、分岐クロスバスイッ
チ402によって接続される。4つのスレッド制御ユニ
ット403には、それぞれ命令メモリ404が接続され
る。命令メモリ404には、それぞれに整数演算ユニッ
ト群405が接続される。整数演算ユニット群405
は、4つの整数演算ユニットを内蔵する。よって、マイ
クロプロセッサ401は最大16の整数演算を同時実行
できる。The four integer operation unit groups 405 and the four thread control units 403 are connected by a branch crossbar switch 402. An instruction memory 404 is connected to each of the four thread control units 403. The instruction memory 404 is connected to an integer operation unit group 405, respectively. Integer operation unit group 405
Incorporates four integer arithmetic units. Therefore, the microprocessor 401 can simultaneously execute a maximum of 16 integer operations.
【0045】すべての整数演算ユニットは1つのローカ
ルキャッシュクロスバスイッチ406に接続される。1
Kバイトのローカルデータキャッシュ407は整数演算
ユニットの数と等しい数だけ実装され、全て1つのロー
カルメモリクロスバスイッチ406に接続される。ロー
カルメモリクロスバスイッチ406はグローバルアクセ
ス制御408と、共有浮動小数点演算ユニット409に
も接続される。グローバルアクセス制御408は、共有
浮動小数点演算ユニット409にも接続され、その制御
を行う。同様に分岐クロスバスイッチ402と、ローカ
ルメモリクロスバスイッチ406の制御も行う。All integer arithmetic units are connected to one local cache crossbar switch 406. 1
The K-byte local data caches 407 are mounted in a number equal to the number of integer operation units, and are all connected to one local memory crossbar switch 406. The local memory crossbar switch 406 is also connected to the global access control 408 and the shared floating point operation unit 409. The global access control 408 is also connected to the shared floating point arithmetic unit 409 and controls the same. Similarly, the control of the branch crossbar switch 402 and the local memory crossbar switch 406 is performed.
【0046】分岐ユニット402クロスバスイッチは、
4つの整数演算ユニット群405からのスレッド状態信
号419と、とアクセス制御ユニット408からのスレ
ッド再開要求信号430を受理して、適切なスレッド発
行ユニット403にスレッド状態信号411を伝達する
役割を果たす。同時に4つの分岐を処理可能である。左
端のスレッド制御ユニット403は、その右に配置され
た別のスレッド制御ユニット403からスレッド制御信
号413を受理し、プログラムカウンタ信号414を伝
達する。また、スレッド制御ユニット403は、右端の
スレッド制御ユニット403からスレッド制御信号41
3を伝達し、プログラムカウンタ信号414を受理す
る。右端のスレッド制御ユニット403は、プログラム
カウンタ信号413をグローバルアクセス制御ユニット
408に接続し、インクリメントされたプログラムカウ
ンタ信号412を受理する。The branch unit 402 crossbar switch
It receives the thread status signal 419 from the four integer operation unit groups 405 and the thread resumption request signal 430 from the access control unit 408, and transmits the thread status signal 411 to the appropriate thread issuing unit 403. Four branches can be processed at the same time. The leftmost thread control unit 403 receives a thread control signal 413 from another thread control unit 403 disposed on the right side, and transmits a program counter signal 414. Also, the thread control unit 403 receives a thread control signal 41 from the rightmost thread control unit 403.
3 is received and the program counter signal 414 is received. The rightmost thread control unit 403 connects the program counter signal 413 to the global access control unit 408 and receives the incremented program counter signal 412.
【0047】そして、命令メモリ404は、内蔵する命
令メモリの内容に基づいて命令デコード信号415を隣
接する整数演算ユニット群405に伝達する。同時に、
その右に配置された別の命令メモリ404にデコード信
号415の一部である命令デコード信号431を伝達す
る。通常のパイプラインでは次のパイプラインステージ
に命令デコード結果を送るが、本発明のプロセッサでは
次のパイプラインステージは常に右に隣接する演算ユニ
ットで実行するためである。右端の命令メモリ404
は、左端の命令メモリに命令デコード信号431を伝達
する。The instruction memory 404 transmits the instruction decode signal 415 to the adjacent integer operation unit group 405 based on the contents of the built-in instruction memory. at the same time,
An instruction decode signal 431 which is a part of the decode signal 415 is transmitted to another instruction memory 404 arranged on the right side. In a normal pipeline, an instruction decode result is sent to the next pipeline stage. However, in the processor of the present invention, the next pipeline stage is always executed by an arithmetic unit adjacent to the right. Rightmost instruction memory 404
Transmits an instruction decode signal 431 to the leftmost instruction memory.
【0048】整数演算ユニット群405は、4つの整数
演算ユニットと4つの分岐ユニットと1つの分岐調停ユ
ニットで構成される。整数演算ユニット群405は、隣
接した命令メモリ404から受け取った命令デコード信
号415を使用して機能を自在に変更できる。この点で
は一般的なマイクロプロセッサと同一である。そして、
演算の入力データとして、ローカルメモリクロスバスイ
ッチ406とのアドレス、データのやり取りを行う信号
417が4つと、グローバルメモリアクセスバス42
7、428が4本接続される。この4つという数は整数
演算ユニット群405が内蔵する整数演算ユニットの数
に等しい。さらに本発明のプロセッサでは、同時に左に
隣接する整数演算ユニット405からもデータをパイプ
ラインデータバス信号416を介して受理することにな
る。演算ユニット405が左端の場合は、右端の演算ユ
ニット405から受理する。ローカルメモリクロスバス
イッチへの接続は、データバスとアドレスバスと制御バ
スから構成される。データバスは32ビットだが、アド
レスバスは内部表記のため1Kバイト分の8ビットのみ
が使用されている。なお、グローバルメモリアクセスア
ドレスバス信号417は32ビット長である。The integer operation unit group 405 includes four integer operation units, four branch units, and one branch arbitration unit. The function of the integer operation unit group 405 can be freely changed using the instruction decode signal 415 received from the adjacent instruction memory 404. In this respect, it is the same as a general microprocessor. And
As input data for the operation, four signals 417 for exchanging addresses and data with the local memory crossbar switch 406 and the global memory access bus 42
7, 428 are connected. The number of four is equal to the number of integer operation units included in the integer operation unit group 405. Further, in the processor of the present invention, data is also received from the integer arithmetic unit 405 adjacent to the left via the pipeline data bus signal 416 at the same time. When the arithmetic unit 405 is at the left end, the data is received from the rightmost arithmetic unit 405. The connection to the local memory crossbar switch includes a data bus, an address bus, and a control bus. Although the data bus has 32 bits, the address bus uses only 8 bits of 1 Kbyte for internal notation. The global memory access address bus signal 417 is 32 bits long.
【0049】さらに、整数演算ユニット群405は、全
て分岐クロスバスイッチ402に向けてスレッド状態信
号419を伝送する。スレッド状態信号409はスレッ
ド状態信号419のうちのプログラムカウンタ信号の内
容によって4つのスレッド制御ユニット403のうちの
1つに分配される。Further, all the integer operation unit groups 405 transmit the thread state signal 419 to the branch crossbar switch 402. The thread status signal 409 is distributed to one of the four thread control units 403 according to the contents of the program counter signal in the thread status signal 419.
【0050】ローカルデータキャッシュ407は1Kバ
イトの容量を持つ、1ポートのキャッシュメモリであ
る。1つのマイクロプロセッサ401上に全部で16個
配置される。それらは全てローカルメモリクロスバスイ
ッチ406と接続される。ローカルメモリアドレスバス
421は26ビットの信号である。ローカルメモリデー
タバス422は32ビットの双方向信号である。ローカ
ルデータキャッシュ407を16個用意することによ
り、マイクロプロセッサ401が内蔵する整数演算ユニ
ット全てが同時にローカルデータキャッシュ407をア
クセスできることになる。The local data cache 407 is a one-port cache memory having a capacity of 1 Kbyte. A total of 16 microprocessors are arranged on one microprocessor 401. They are all connected to the local memory crossbar switch 406. The local memory address bus 421 is a 26-bit signal. The local memory data bus 422 is a 32-bit bidirectional signal. By preparing 16 local data caches 407, all the integer operation units incorporated in the microprocessor 401 can access the local data cache 407 at the same time.
【0051】ローカルメモリクロスバスイッチ406
は、16本のローカルメモリアクセス信号418の内の
アドレス信号を受理して、任意のローカルメモリアドレ
スバス421にそのまま転送する。ローカルメモリアク
セス信号418のうちのデータバス信号は、ローカルメ
モリデータバス422に、それぞれクロスバスイッチ制
御信号426の内容に応じて接続される。ローカルメモ
リアクセス信号418のうちのデータバス信号と、ロー
カルメモリデータバス422の信号の方向は全て双方向
である。Local memory crossbar switch 406
Receives an address signal among the 16 local memory access signals 418 and transfers it to an arbitrary local memory address bus 421 as it is. The data bus signal of the local memory access signal 418 is connected to the local memory data bus 422 according to the content of the crossbar switch control signal 426, respectively. The data bus signal of the local memory access signal 418 and the signal of the local memory data bus 422 are all bidirectional.
【0052】グローバルアクセス制御ユニット408
は、整数演算ユニット群405からの分岐要求信号43
3、434、435、436を受理して、分岐クロスバ
スイッチ402を制御する。同時に、ローカルメモリク
ロスバスイッチ406を、制御信号426を使用して制
御する。また、グローバルアクセス制御ユニット408
にはグローバルメモリアクセスアドレス信号427とグ
ローバルメモリアクセスデータ信号428が接続され
る。命令キャッシュリプレース信号432による命令キ
ャッシュのリプレース、およびデータキャッシュリプレ
ース信号425によるデータキャッシュのリプレース行
う。マイクロプロセッサ401の外部へのアクセスであ
る場合は、外部アドレスバス452、外部制御バス45
1を使用してチップ外部メモリをアクセスする。外部デ
ータバス450は、外部書き込みのときは出力、外部読
み込みの時は入力となる。外部からのデータロードやキ
ャッシュのリプレースが終わった時点で、ロード待ちで
休眠しているスレッドを生起するために、分岐クロスバ
スイッチ制御信号430を通してスレッド制御ユニット
403にスレッド再開要求を行う。さらに、外部割りこ
み信号453を受理して、それをスレッド発行に変換し
てスレッド状態信号437を伝達し、スレッド制御ユニ
ット403に割り込み受理スレッドを生起させる。Global access control unit 408
Is a branch request signal 43 from the integer operation unit group 405.
3, 434, 435 and 436 are received, and the branch crossbar switch 402 is controlled. At the same time, the local memory crossbar switch 406 is controlled using the control signal 426. Also, the global access control unit 408
Are connected to a global memory access address signal 427 and a global memory access data signal 428. The instruction cache is replaced by the instruction cache replace signal 432 and the data cache is replaced by the data cache replace signal 425. When the access is to the outside of the microprocessor 401, the external address bus 452, the external control bus 45
1 is used to access the external memory of the chip. The external data bus 450 is an output during external writing and an input during external reading. When data loading from the outside or replacement of the cache is completed, a thread resumption request is issued to the thread control unit 403 through the branch crossbar switch control signal 430 in order to generate a sleeping thread waiting for loading. Further, it receives the external interrupt signal 453, converts it into a thread issue, transmits the thread status signal 437, and causes the thread control unit 403 to generate an interrupt accepting thread.
【0053】共有浮動小数点演算ユニット409は、ロ
ーカルキャッシュクロスバスイッチ406にデータバス
424、アドレスバス423を通じて接続される。そし
て、共有浮動小数点ユニット409は、ローカルキャッ
シュクロスバスイッチ406を介して整数演算ユニット
群405からデータを受理し、整数演算ユニット群40
5に直接データを伝送する。共有浮動小数点演算ユニッ
ト409は、演算を要求した1つの整数演算ユニット群
405にとっては、ローカルクロスバスイッチ406に
よって、ローカルキャッシュメモリ407の代わりに接
続される形になる。The shared floating point arithmetic unit 409 is connected to the local cache crossbar switch 406 via a data bus 424 and an address bus 423. The shared floating-point unit 409 receives data from the integer operation unit group 405 via the local cache crossbar switch 406, and
5 directly. The shared floating-point operation unit 409 is connected to one integer operation unit group 405 that has requested the operation, instead of the local cache memory 407 by the local crossbar switch 406.
【0054】<整数演算ユニット群405>次に、図5
を参照して、整数演算ユニット群405の内部構造のブ
ロック図について概説する。<Integer operation unit group 405> Next, FIG.
, A block diagram of the internal structure of the integer operation unit group 405 will be outlined.
【0055】整数演算ユニット群405は、4つの整数
演算ユニット501と、4つの分岐ユニット502と、
分岐ユニット502の調停を行う分岐アービター503
から構成される。The integer operation unit group 405 includes four integer operation units 501, four branch units 502,
Branch arbiter 503 for arbitrating branch unit 502
Consists of
【0056】整数演算ユニット401は、主に整数演算
機とテンポラリデータレジスタから構成される。前述の
命令メモリ404から、演算機の実行制御の為の命令デ
コード信号511を受理する。前段の整数演算ユニット
401からテンポラリデータレジスタを含むすべての内
部状態を受理して、1つの演算を行い。次の段の整数演
算ユニット401に演算結果を含むすべての状態を伝達
する。1クロックごとにすべての状態を右方向に伝達す
る点で他のプロセッサと異なっている。演算の結果発生
したメモリへのロードストア、分岐を実行するために、
分岐ユニット502に対して制御信号を送り、メモリか
ら読み込まれたデータを分岐ユニット502から受理す
る。The integer operation unit 401 mainly comprises an integer operation unit and a temporary data register. An instruction decode signal 511 for controlling execution of the arithmetic unit is received from the instruction memory 404 described above. One internal operation is performed by receiving all internal states including the temporary data register from the integer operation unit 401 in the preceding stage. All states including the operation result are transmitted to the integer operation unit 401 in the next stage. It differs from other processors in that all states are transmitted to the right every clock. To execute load store / branch to memory generated as a result of operation,
A control signal is sent to the branch unit 502, and data read from the memory is received from the branch unit 502.
【0057】分岐ユニット502は、分岐命令の実行
と、ロードストアの実行を行う。ローカルメモリのロー
ドストアはすべての分岐ユニット502が同時に実行で
きる。しかし、分岐の実行は4つの分岐ユニット502
で共有され、1つだけが同時に実行できる。さらに、グ
ローバルメモリのロードストアは、マイクロプロセッサ
401全体で同時に1つとなっている。本発明のプロセ
ッサでは、スレッド切り替えだけでなく、分岐と不定レ
イテンシ時間のロード、これらが全てスレッドの退避と
コンテキストスイッチを行う必要がある。そのため、分
岐の度に分岐要求信号群によってスレッドに必要な状態
を分岐ユニット403に伝達する必要がある。本発明の
マイクロプロセッサではプログラムカウンタ、スタック
ポインタ、スレッドIDの3つの信号を最小限必要なス
レッドの状態として使用している。これらは分岐の実行
に応じて分岐要求信号群525を通して分岐アービター
503に伝達される。分岐ユニット状態信号群516
は、分岐に必要な内部状態である。The branch unit 502 executes a branch instruction and executes load store. A local memory load store can be performed by all branch units 502 simultaneously. However, execution of the branch is performed by four branch units 502.
And only one can run at the same time. Further, the load store of the global memory is one at a time for the entire microprocessor 401. In the processor of the present invention, it is necessary to perform not only thread switching, but also branching and loading of an indefinite latency time, all of which need to perform thread saving and context switching. Therefore, it is necessary to transmit a state necessary for a thread to the branch unit 403 by a branch request signal group each time a branch is performed. In the microprocessor of the present invention, three signals of a program counter, a stack pointer, and a thread ID are used as a minimum necessary thread state. These are transmitted to the branch arbiter 503 through the branch request signal group 525 according to the execution of the branch. Branch unit status signal group 516
Is an internal state required for branching.
【0058】分岐ユニット502のローカルメモリロー
ドストアは、単にローカルキャッシュクロスバスイッチ
406に向かって無条件でローカルアドレス521と、
ローカルデータ523をストアすれば良い。ロードの場
合は1クロック遅れて右のローカルデータ523に読み
込まれる。スケジューリングはグローバルアクセス制御
ユニット408が、ローカルキャッシュクロスバスイッ
チ406を制御している。スレッドに対してデータキャ
ッシュメモリ407を1つ静的に割り当てるため、ロー
カルメモリクロスバスイッチ406の制御は基本的に静
的である。分岐ユニット502が割り当てられていない
データキャッシュメモリ407については、分岐ユニッ
ト502がスレッドの右への伝達と無関係にその場で保
持し、分岐ユニット502に対象となるデータキャッシ
ュメモリバンク407が割り当てられるのを待つ。割り
当てられていないデータキャッシュメモリバンク407
へのアクセスレイテンシは最小16クロック必要とす
る。その代わり、16のスレッドの同一バンクへのバン
ク外アクセスの調停も可能になる。The local memory load store of the branch unit 502 simply stores the local address 521 unconditionally toward the local cache crossbar switch 406;
The local data 523 may be stored. In the case of loading, the data is read into the right local data 523 with a delay of one clock. In the scheduling, the global access control unit 408 controls the local cache crossbar switch 406. Since one data cache memory 407 is statically allocated to a thread, the control of the local memory crossbar switch 406 is basically static. For the data cache memory 407 to which the branch unit 502 is not assigned, the branch unit 502 holds the data on the spot regardless of the transmission of the thread to the right, and the target data cache memory bank 407 is assigned to the branch unit 502. Wait for. Unassigned data cache memory bank 407
Access latency requires a minimum of 16 clocks. Instead, arbitration of out-of-bank access to the same bank by 16 threads becomes possible.
【0059】分岐アービター503は、4つの分岐ユニ
ット502からの分岐要求信号群525〜528を全て
受理して、1つを選択し、分岐要求信号419を分岐ク
ロスバスイッチ402に伝達する。選択されなかった分
岐要求信号群525〜528は、パイプラインストール
となり、次のクロック以降に発行される。分岐の頻度
が、基本的に命令の4分の1以下であることを想定した
インプリメントである。The branch arbiter 503 receives all the branch request signal groups 525 to 528 from the four branch units 502, selects one, and transmits the branch request signal 419 to the branch crossbar switch 402. The branch request signal groups 525 to 528 that are not selected become pipeline stall and are issued after the next clock. This implementation is based on the assumption that the frequency of branching is basically one-fourth or less of the instruction.
【0060】<整数演算ユニット501>次に、図6を
参照する。整数演算ユニット501の内部構造のブロッ
ク図である。<Integer operation unit 501> Next, reference will be made to FIG. FIG. 3 is a block diagram of an internal structure of an integer operation unit 501.
【0061】601から619は全てエッジトリガフリ
ップフロップである。プログラムカウンタ601、スタ
ックポインタ602、スレッド情報603は、分岐、ス
レッドの新規発行の時に分岐ユニット403から伝達さ
れる信号である。デコード済み命令コードラッチ604
は、常に命令メモリ404から伝達される信号である。Reference numerals 601 to 619 denote edge trigger flip-flops. The program counter 601, the stack pointer 602, and the thread information 603 are signals transmitted from the branch unit 403 when a new branch or thread is issued. Decoded instruction code latch 604
Is a signal always transmitted from the instruction memory 404.
【0062】プログラムカウンタ605は、26ビット
の信号であり、そのパイプラインステージのスレッドの
命令アドレスを示すものであり、分岐先の計算に使用さ
れる。条件実行命令制御ラッチ606は、整数演算ユニ
ット501の演算結果の書き戻しや分岐命令実行の中止
を実行するための制御信号である。フラグレジスタラッ
チ607は、フラグを更新しない場合のための状態保持
レジスタである。フラグレジスタ更新ラッチ608は、
第1オペランドバス627に接続されており、レジスタ
の値をフラグに直接代入するのに使用される。ALU演
算結果フラグラッチ609は、左に配置される1クロッ
ク前の整数演算ユニット501の演算結果のフラグを保
持する。一般的にフラグの算出はクリティカルパスにな
り易いので、フラグレジスタ607への代入は演算の次
のステージで行うことになる。ALU演算結果ラッチ6
10は、ALU630の演算結果を保持するラッチであ
る。バレルシフタ演算結果フラグラッチ611は、バレ
ルシフタ631で発生したフラグ変更を保持するラッチ
である。バレルシフタ演算結果ラッチ612は、バレル
シフタの演算結果を保持するラッチである。ストアデー
タラッチ613は、演算の次のクロックで実行されるス
トアまでストアデータを保持するラッチである。The program counter 605 is a 26-bit signal that indicates the instruction address of the thread in the pipeline stage, and is used for calculating the branch destination. The condition execution instruction control latch 606 is a control signal for executing write back of the operation result of the integer operation unit 501 or suspension of execution of a branch instruction. The flag register latch 607 is a state holding register for a case where the flag is not updated. The flag register update latch 608 is
It is connected to the first operand bus 627 and is used to directly substitute the value of a register into a flag. The ALU operation result flag latch 609 holds the flag of the operation result of the integer operation unit 501 one clock before, which is arranged on the left. In general, the calculation of a flag tends to be a critical path, and therefore, the assignment to the flag register 607 is performed in the next stage of the operation. ALU operation result latch 6
Reference numeral 10 denotes a latch for holding the operation result of the ALU 630. The barrel shifter operation result flag latch 611 is a latch that holds a flag change generated in the barrel shifter 631. The barrel shifter operation result latch 612 is a latch that holds the operation result of the barrel shifter. The store data latch 613 is a latch that holds store data until the store is executed at the next clock of the operation.
【0063】レジスタライトバックラッチ614は、第
1汎用レジスタラッチ615〜第4汎用レジスタラッチ
618に結果を書き戻すタイミングを遅らせるために用
いられる。レジスタの書き戻しは、ローカルメモリから
の読み込みデータ書き込みと同じタイミングで行う必要
があるため、時間的に1クロック遅らせる必要がある。The register write-back latch 614 is used to delay the timing of writing the result back to the first to fourth general-purpose register latches 615 to 618. Since the write-back of the register needs to be performed at the same timing as the writing of the read data from the local memory, it is necessary to temporally delay by one clock.
【0064】第1汎用レジスタラッチ615から第4汎
用レジスタラッチ618は、演算に使用するオペランド
レジスタに使用する。スタックポインタラッチ619
は、演算結果の書き戻し、状態の一時退避などに使用す
るスタックポインタアドレスを保持する。The first to fourth general-purpose register latches 615 to 618 are used for operand registers used for operations. Stack pointer latch 619
Holds a stack pointer address used for writing back the operation result, temporarily saving the state, and the like.
【0065】プログラムカウンタ更新セレクタ621
は、次のプログラムカウンタ601に伝達すべき値を前
のプログラムカウンタ605と、更新するプログラムカ
ウンタ601から選択する。分岐や不定レイテンシロー
ドの終了によるスレッドの再開時には更新するプログラ
ムカウンタ601を選択する。それ以外は前のプログラ
ムカウンタ605をそのまま使用する。一般的なマイク
ロプロセッサと異なり、プログラムカウンタ加算器は必
要ない。Program counter update selector 621
Selects the value to be transmitted to the next program counter 601 from the previous program counter 605 and the program counter 601 to be updated. When the thread is restarted due to the termination of the branch or the undefined latency load, the program counter 601 to be updated is selected. Otherwise, the previous program counter 605 is used as it is. Unlike a general microprocessor, no program counter adder is required.
【0066】プログラム定数生成回路622は、プログ
ラムカウンタ605の中に含まれない下位のPCを生成
する。16の整数演算機501に対して0から15まで
の数値が割り当てられて値を生成する。The program constant generation circuit 622 generates a lower-order PC that is not included in the program counter 605. Numerical values from 0 to 15 are assigned to the 16 integer arithmetic units 501 to generate values.
【0067】フラグレジスタ更新セレクタ623は、デ
コードされた命令コード604に応じて、前のフラグレ
ジスタラッチ607、前段のフラグレジスタ更新レジス
タの608、前段のALU演算結果フラグラッチ60
9、前段のバレルシフタ演算結果ラッチ610の中から
次段のフラグレジスタラッチ607へ伝達する信号を選
択する。フラグレジスタへの代入命令では608、AL
U命令では609、シフト命令では611、それ以外で
は607を選択することになる。In response to the decoded instruction code 604, the flag register update selector 623 provides a previous flag register latch 607, a previous flag register update register 608, and a previous ALU operation result flag latch 60.
9. A signal to be transmitted to the flag register latch 607 of the next stage is selected from the barrel shifter operation result latches 610 of the previous stage. 608, AL in the assignment instruction to the flag register
609 is selected for the U instruction, 611 for the shift instruction, and 607 for the other instructions.
【0068】条件実行命令制御回路624は、フラグレ
ジスタ608と、デコード済み命令コード604の内容
に応じて、分岐やデータのフォワーディングやライトバ
ックの制御を行う。演算を即座に停めるわけではない。
分岐の制御には、分岐ユニット502へ条件実行制御信
号646を伝達する。結果データ642の汎用レジスタ
ラッチ615〜619へのライトバック制御には、汎用
レジスタ更新レジスタ634〜638を制御して、条件
が成立した場合のみデータライトバックバス642の値
を代入ようにする。データフォワーディングに関して
は、オペランドセレクタ627〜629によるフォワー
ディングで、演算結果バス641やライトバックバス6
42のデータを代入するかどうかを選択する。実際の命
令にレジスタデータ依存関係があっても、条件が成立し
た場合のみフォワーディングを行うことになる。The conditional execution instruction control circuit 624 controls branching, data forwarding and write-back according to the contents of the flag register 608 and the decoded instruction code 604. It does not stop the operation immediately.
For branch control, a condition execution control signal 646 is transmitted to the branch unit 502. For the write-back control of the result data 642 to the general-purpose register latches 615 to 619, the general-purpose register update registers 634 to 638 are controlled to substitute the value of the data write-back bus 642 only when the condition is satisfied. Regarding the data forwarding, the operation result bus 641 and the write-back bus 6 are forwarded by the operand selectors 627 to 629.
Select whether to substitute the data of 42. Even if the actual instruction has a register data dependency, forwarding is performed only when the condition is satisfied.
【0069】定数生成回路626は、デコードされた命
令コード604の内容に応じて、演算に使用する定数を
生成し、第2オペランドセレクタ628に代入する。The constant generation circuit 626 generates a constant to be used for the operation in accordance with the content of the decoded instruction code 604 and substitutes it for the second operand selector 628.
【0070】第1オペランドセレクタ627は、第1オ
ペランドバス643と、結果伝送バス641と、データ
ライトバックバス642と、プログラムカウンタバス6
25から1つを選択してALU回路630、 バレルシ
フタ631、フラグレジスタ更新ラッチ608に代入す
る。The first operand selector 627 includes a first operand bus 643, a result transmission bus 641, a data write back bus 642, and a program counter bus 6.
One is selected from 25 and substituted into the ALU circuit 630, the barrel shifter 631, and the flag register update latch 608.
【0071】第2オペランドセレクタ628は、第2オ
ペランドバス644と、結果伝送バス641と、データ
ライトバックバス642と、定数生成回路626から1
つを選択してALU回路630、バレルシフタ631の
シフト数制御入力に代入する。The second operand selector 628 includes a second operand bus 644, a result transmission bus 641, a data write back bus 642, and a constant generation circuit 626.
One is selected and substituted into the shift number control input of the ALU circuit 630 and the barrel shifter 631.
【0072】ストアデータセレクタ629は、第3オペ
ランドバス645と、結果伝送バス641と、データラ
イトバックバス642と、プログラムカウンタバス62
5と、フラグレジスタ更新セレクタ623の出力から選
択して、ストアアライナ632へ代入する。The store data selector 629 includes a third operand bus 645, a result transmission bus 641, a data write back bus 642, and a program counter bus 62.
5 and the output of the flag register update selector 623, and assigns it to the store aligner 632.
【0073】ストアアライナ632は、バイト、16ビ
ットワード単位のストアを、ストアアドレスに応じて、
32ビットのデータバスに適切に配置するユニットであ
る。The store aligner 632 stores a byte or a 16-bit word unit according to a store address.
This unit is appropriately arranged on a 32-bit data bus.
【0074】スレッド間レジスタフォワードバス633
は、整数演算ユニット501の演算結果バス641の値
を、後続の別のスレッドがオペランドとして使用するた
めのバスである。これにより、マルチスレッド間の通信
を高速に行うことができる。Thread register forward bus 633
Is a bus for another subsequent thread to use the value of the operation result bus 641 of the integer operation unit 501 as an operand. Thereby, communication between multi-threads can be performed at high speed.
【0075】第1汎用レジスタ更新セレクタ634は、
第1汎用レジスタラッチ615とデータライトバックバ
ス642のどちらかを選択して次の段の第1汎用レジス
タラッチ615に代入する。第2汎用レジスタ更新セレ
クタ635から第4汎用レジスタ更新セレクタ637も
同様の動作を行う。The first general register update selector 634 is
One of the first general-purpose register latch 615 and the data write-back bus 642 is selected and assigned to the first general-purpose register latch 615 of the next stage. The second general register update selector 635 to the fourth general register update selector 637 perform the same operation.
【0076】スタックポインタ更新セレクタ638は、
スタックポインタラッチ619とデータライトバックバ
ス642と、さらにスタックポインタ更新ラッチ602
の出力を選択して、次の段のスタックポインタレジスタ
638に代入する。The stack pointer update selector 638
Stack pointer latch 619, data write back bus 642, and stack pointer update latch 602
Is selected and assigned to the stack pointer register 638 of the next stage.
【0077】メモリストア用データバス640は、ロー
カルメモリ407、外部メモリに格納する32ビットデ
ータである。The memory store data bus 640 is 32-bit data stored in the local memory 407 and the external memory.
【0078】演算結果バス641は、演算結果をローカ
ルメモリ407に書き戻すために用いられる。ローカル
メモリ407、外部メモリのアクセスの為のアドレスと
して使用される。The operation result bus 641 is used to write the operation result back to the local memory 407. The local memory 407 is used as an address for accessing an external memory.
【0079】データライトバックバス642は、演算結
果、あるいはロードデータをを汎用レジスタ615〜6
19に書き戻すためのバスである。The data write back bus 642 transfers the operation result or load data to the general purpose registers 615 to
This is a bus for writing back to 19.
【0080】第1オペランドバス643は、汎用レジス
タ616〜618およびスタックポインタ619を読み
出すのに用いられる。第2オペランドバス644と第3
オペランドバス645も同様である。The first operand bus 643 is used for reading the general-purpose registers 616 to 618 and the stack pointer 619. The second operand bus 644 and the third operand bus 644
The same applies to the operand bus 645.
【0081】<分岐・ロードストアユニット502>次
に、図7を参照して、分岐ユニット502の内部構造を
概説する。<Branch / Load Store Unit 502> Next, the internal structure of the branch unit 502 will be outlined with reference to FIG.
【0082】分岐ユニット502は、スレッド状態ラッ
チ701と、ロードデータラッチ702と、ロードデー
タ保持バッファ703、ローカルメモリライトアドレス
ラッチ704、ローカルメモリライトデータラッチ70
5、アドレスセレクタ706、アドレスバッファ70
7、ストアデータセレクタ708、ストアデータバッフ
ァ709、スレッド状態更新セレクタ710、分岐制御
ユニット711、ロードアライナ712、ローカルメモ
リアクセス保護検査713で構成される。The branch unit 502 includes a thread state latch 701, a load data latch 702, a load data holding buffer 703, a local memory write address latch 704, and a local memory write data latch 70.
5, address selector 706, address buffer 70
7, a store data selector 708, a store data buffer 709, a thread state update selector 710, a branch control unit 711, a load aligner 712, and a local memory access protection check 713.
【0083】スレッド状態ラッチ701は、スレッドI
Dなど、分岐に必要なスレッド状態を保持している。ス
レッド状態信号647の値をスレッド状態更新セレクタ
710によって選択することで更新できる。The thread status latch 701 is provided for the thread I
The thread state required for branching, such as D, is held. It can be updated by selecting the value of the thread status signal 647 by the thread status update selector 710.
【0084】ロードデータラッチ702は、32ビット
のラッチで、データバスから読み込んだ値を次のクロッ
クまで保持するのに使用される。ロードデータ保持バッ
ファ703は、所有しないローカルメモリからのリード
データを保持しておくのに使用される。The load data latch 702 is a 32-bit latch used to hold the value read from the data bus until the next clock. The load data holding buffer 703 is used to hold read data from a non-owned local memory.
【0085】ローカルメモリライトデータラッチ70
4、およびローカルメモリライトアドレスラッチ705
は、32ビットのラッチで、ローカルメモリへの書き込
みを遅らせるのに使用される。Local memory write data latch 70
4, and local memory write address latch 705
Is a 32-bit latch used to delay writing to local memory.
【0086】アドレスセレクタ706は、ローカルデー
タキャッシュメモリ407、あるいはグローバルバス4
27に出力するアドレス値を、演算結果バス641と、
前クロックからのローカルメモリライトアドレスラッチ
705と、アドレスバッファ707の3つの中から選択
する。The address selector 706 is connected to the local data cache memory 407 or the global bus 4
27, the address value to be output to the operation result bus 641,
A local memory write address latch 705 from the previous clock and an address buffer 707 are selected from the three.
【0087】アドレスバッファ707は、32ビットの
バッファで、バンク外へのローカルキャッシュメモリ4
07のアクセス、グローバルバス427への出力のアー
ビトレーションを待ち合わせるために、2つまでロード
ストアアドレスを蓄積できる。The address buffer 707 is a 32-bit buffer and stores the local cache memory 4 outside the bank.
Up to two load store addresses can be stored in order to wait for the access of 07 and the arbitration of the output to the global bus 427.
【0088】ストアデータセレクタ708は、ローカル
メモリ407、あるいはグローバルバス427に出力す
るデータ値を、メモリストア用データバス640か、ロ
ーカルメモリライトデータラッチ704、あるいはデー
タバスバッファ709の3つの中から選択する。The store data selector 708 selects a data value to be output to the local memory 407 or the global bus 427 from the memory store data bus 640, the local memory write data latch 704, or the data bus buffer 709. I do.
【0089】データバスバッファ710は、32ビット
のバッファで、バンク外へのローカルメモリのアクセ
ス、グローバルバス427への出力のアービトレーショ
ンを待ち合わせるために、2つまでストアデータを蓄積
できる。The data bus buffer 710 is a 32-bit buffer and can store up to two store data in order to wait for local memory access outside the bank and arbitration of output to the global bus 427.
【0090】分岐制御ユニット711は、バンク外ロー
カルメモリアクセス、グローバルロードストア命令、分
岐命令の実行、および分岐アービター503との調停を
行う。分岐の場合は、更新するプログラムカウンタは演
算結果バス641から受理する。更新するスタックポイ
ンタは、メモリストア用データバス640から受理す
る。スレッドIDなどのその他のスレッド状態はスレッ
ド状態ラッチ701から受理する。これらをまとめて分
岐要求信号群525から分岐調停ユニット503に向け
て出力する。ここで、条件付き実行命令のため、分岐を
取り消す場合は、分岐取り消し信号646を受理して分
岐動作を停める。The branch control unit 711 performs out-of-bank local memory access, execution of a global load / store instruction, execution of a branch instruction, and arbitration with the branch arbiter 503. In the case of a branch, the program counter to be updated is received from the operation result bus 641. The stack pointer to be updated is received from the memory store data bus 640. Other thread states, such as the thread ID, are received from the thread state latch 701. These are collectively output from the branch request signal group 525 to the branch arbitration unit 503. Here, when the branch is canceled due to the conditional execution instruction, the branch cancel operation is accepted and the branch operation is stopped.
【0091】ロードアライナ712は、バイト単位でデ
ータを読み出す場合に、データのシフト、32ビットへ
の符号の拡張等を行う。When reading data in byte units, the load aligner 712 shifts the data, extends the code to 32 bits, and the like.
【0092】ローカルメモリアクセス検査713は、ロ
ーカルメモリに発行されるアドレス信号が固有のローカ
ルメモリへのアクセスかどうか検査を行う。固有のロー
カルメモリへのライトアクセスの場合は、データラッチ
704、アドレスラッチ705を使用して次のクロック
でローカルメモリへの書き込みを行う。固有のローカル
メモリ以外の場合は、アドレスバッファ708、データ
バッファ710を使用して保持し、ローカルメモリ、あ
るいはグローバルメモリバス427に出力る。The local memory access check 713 checks whether an address signal issued to the local memory is an access to a unique local memory. In the case of write access to a unique local memory, writing to the local memory is performed at the next clock using the data latch 704 and the address latch 705. In the case other than the unique local memory, the data is held using the address buffer 708 and the data buffer 710, and is output to the local memory or the global memory bus 427.
【0093】分岐調停ユニット525から、分岐受理信
号714を受理できない場合は、分岐のアービトレーシ
ョンに失敗して分岐が実行できない場合である。その場
合は、すべての分岐情報を、データバス群515を使用
して右に隣接する演算ユニット501に伝達する。本来
のスレッドはパイプラインストールとなり、データの退
避を待つ。次のクロックで、右に隣接する分岐ユニット
502が再度同じ分岐を実行する。If the branch arbitration unit 525 cannot receive the branch acknowledgment signal 714, the arbitration of the branch fails and the branch cannot be executed. In that case, all the branch information is transmitted to the right adjacent arithmetic unit 501 using the data bus group 515. The original thread becomes a pipeline stall and waits for data to be saved. At the next clock, the right adjacent branch unit 502 executes the same branch again.
【0094】<命令発行ユニット404>図8を参照し
て、命令発行ユニット404の内部構造について概説す
る。命令発行ユニット404は、分岐ユニット403か
ら与えられたプログラムカウンタを元に、整数演算機群
405に必要な命令コードを出力するユニットである。<Instruction Issuing Unit 404> The internal structure of the instruction issuing unit 404 will be outlined with reference to FIG. The instruction issuing unit 404 is a unit that outputs an instruction code required for the integer operation unit group 405 based on the program counter provided from the branch unit 403.
【0095】命令発行ユニット404は、Xデコーダ8
01、128ビット幅のSRAMセル802、センスア
ンプおよびYセレクタ803、命令メモリアクセス制御
ユニット804、3つの命令コードセレクタ805、6
つの32ビット命令コードラッチ806、命令デコード
ユニット807で構成される。The instruction issuing unit 404 includes the X decoder 8
01, 128-bit SRAM cell 802, sense amplifier and Y selector 803, instruction memory access control unit 804, three instruction code selectors 805, 6
It comprises two 32-bit instruction code latches 806 and an instruction decode unit 807.
【0096】プログラムカウンタ821の信号を元に、
SRAMセル802から命令コードを読み出す。SRA
Mセンスアンプ803からは同時に4つ分の命令コード
が読み出される。Based on the signal of the program counter 821,
The instruction code is read from the SRAM cell 802. SRA
Four instruction codes are simultaneously read from the M sense amplifier 803.
【0097】128ビット幅のリプレース信号851
は、直接センスアンプ803に接続される。命令のリプ
レースは1クロックで実行される。Replacement signal 851 having a width of 128 bits
Are directly connected to the sense amplifier 803. The instruction replacement is executed in one clock.
【0098】命令アドレスの最も若い命令コード830
は即座に命令デコードユニット807に送られる。しか
し、次の命令コード831は命令コードラッチ806に
格納されて1クロック遅れて発行される。最後の命令コ
ードは3クロック遅れて発行される。同時刻には、4つ
の独立したスレッドの命令コードが4つの命令デコード
ユニット807に送信される。The instruction code 830 with the youngest instruction address
Is immediately sent to the instruction decode unit 807. However, the next instruction code 831 is stored in the instruction code latch 806 and issued one clock later. The last instruction code is issued three clocks later. At the same time, instruction codes of four independent threads are transmitted to four instruction decode units 807.
【0099】スレッド状態信号821は、分岐やスレッ
ド新規生成の時にのみ左端の整数演算ユニット501に
送られる。分岐は常に4命令単位で行う。分岐先アドレ
スが4の倍数の先頭でないときは、分岐先までの整数演
算ユニット501は使用できないため性能が低下する。
しかし、ソフトウェアで分岐先のアドレスを常に4の倍
数に配置すれば性能の低下はない。The thread status signal 821 is sent to the leftmost integer operation unit 501 only at the time of branching or new thread generation. Branching is always performed in units of four instructions. When the branch destination address is not the beginning of a multiple of 4, the integer operation unit 501 up to the branch destination cannot be used, so that the performance is reduced.
However, if the branch destination address is always arranged in a multiple of 4 by software, the performance does not decrease.
【0100】演算ユニット制御信号835は、複数のス
レッドの命令コードが混在した信号である。演算ユニッ
ト制御信号835は、1つの整数演算ユニット501と
1つの分岐ユニット502で使用される制御信号の全て
が含まれる。そのクロックで使用される制御信号だけで
構成される。次以降のクロックで実行される制御信号は
デコード済み命令836によって次のクロックの命令デ
コードユニット807に渡されて、次のクロックで演算
ユニット制御信号837として出力される。The arithmetic unit control signal 835 is a signal in which instruction codes of a plurality of threads are mixed. The operation unit control signal 835 includes all the control signals used in one integer operation unit 501 and one branch unit 502. It consists only of control signals used in the clock. The control signal executed at the next clock is passed to the instruction decode unit 807 of the next clock by the decoded instruction 836, and is output as the operation unit control signal 837 at the next clock.
【0101】<スレッド発行ユニット403>図9を参
照して、スレッド発行ユニット403の内部構造につい
て概説する。スレッド発行ユニット403は、分岐命令
やスレッド発行命令の実行を行うユニットである。ま
た、動作していないスレッドや、外部のメモリアクセス
などで休止しているスレッド情報を格納するユニットで
もある。<Thread Issue Unit 403> The internal structure of the thread issue unit 403 will be outlined with reference to FIG. The thread issuing unit 403 is a unit that executes a branch instruction and a thread issuing instruction. It is also a unit that stores information on threads that are not operating or threads that are suspended due to external memory access or the like.
【0102】スレッド発行ユニット403は、スレッド
発行アービトレーションユニット901、スタックポイ
ンタ連想メモリ902、スレッド状態メモリ903、ス
レッド開始準備フラグ904、命令キャッシュアドレス
タグメモリ905、4つのシフトレジスタ構成を取るロ
ーカルキャッシュバンク番号タグ906、命令キャッシ
ュアドレス比較器907、スレッド発行ユニット90
8、プログラムカウンタシフトレジスタ909で構成さ
れる。The thread issuing unit 403 includes a thread issuing arbitration unit 901, a stack pointer associative memory 902, a thread state memory 903, a thread start preparation flag 904, an instruction cache address tag memory 905, and a local cache bank number having four shift register configurations. Tag 906, instruction cache address comparator 907, thread issuing unit 90
8, a program counter shift register 909.
【0103】まず、分岐クロスバスイッチ402の機能
について概説する。分岐ユニット502から発行された
スレッド状態信号419は、プログラムカウンタの下位
2ビットの内容に応じて4つのスレッド制御ユニット4
03の中から選択してスレッド状態信号411を伝送す
る。この伝送を行うのが分岐クロスバスイッチ402で
ある。同時に同じスレッド制御ユニット403への分岐
が発生した場合は、1つだけが伝送され、残りは整数演
算ユニット群405の中で次のクロック以降の分岐受理
を待つことになる。First, the function of the branch crossbar switch 402 will be outlined. The thread status signal 419 issued from the branch unit 502 is divided into four thread control units 4 according to the contents of the lower two bits of the program counter.
The thread status signal 411 is transmitted by selecting the thread status signal 411 from among the three types. The branch crossbar switch 402 performs this transmission. When a branch to the same thread control unit 403 occurs at the same time, only one is transmitted, and the rest waits for branch acceptance after the next clock in the integer operation unit group 405.
【0104】スレッド発行アービトレーション901
は、分岐制御信号430の要求に応じて、グローバルロ
ード等を待つスレッドを再開させる制御を行うユニット
である。スレッド発行ユニット403のスレッドが連続
して発行できる場合に、隣接するスレッド発行ユニット
403の待ちスレッドがいつまでも発行できないことを
防ぐためのラウンドロビンスケジューラである。Thread arbitration 901
Is a unit that performs control to restart a thread waiting for a global load or the like in response to a request from the branch control signal 430. This is a round-robin scheduler for preventing that the waiting thread of the adjacent thread issuing unit 403 cannot issue forever if the threads of the thread issuing unit 403 can issue continuously.
【0105】分岐クロスバスイッチ402からスレッド
に必要な最小限の情報であるスレッド状態信号411が
伝送されると、スレッド状態信号411は基本的には到
着順にスレッド状態メモリ903に格納される。同時
に、同じアドレスのスタックポインタ連想メモリ902
にスタックポインタの上位4ビットが格納される。スレ
ッド生成や、分岐命令である場合は、スレッド開始準備
フラグ904は最初から1に設定される。ロード命令
や、共有浮動小数点演算ユニット904待ちのスレッド
の場合は、スレッド開始準備フラグ904は0に設定さ
れる。When the thread state signal 411, which is the minimum information necessary for a thread, is transmitted from the branch crossbar switch 402, the thread state signal 411 is basically stored in the thread state memory 903 in the order of arrival. At the same time, the stack pointer associative memory 902 at the same address
Stores the upper 4 bits of the stack pointer. In the case of a thread generation or a branch instruction, the thread start preparation flag 904 is set to 1 from the beginning. In the case of a load instruction or a thread waiting for the shared floating-point operation unit 904, the thread start preparation flag 904 is set to 0.
【0106】ローカルキャッシュバンク番号906は、
4ビットレジスタ4段で構成されるシフトレジスタであ
る。先頭の値が、ローカルメモリバンク信号930を介
してスタックポインタ連想メモリ902に入力される。
ローカルメモリバンクが一致するスレッドの中で最初に
登録された1つが選択され、該当するスレッド状態メモ
リ903からスレッド状態信号936が出力される。さ
らに、スレッド開始準備フラグ904も参照され、スレ
ッド発行ユニット908に伝達される。The local cache bank number 906 is
This is a shift register composed of four 4-bit registers. The leading value is input to the stack pointer associative memory 902 via the local memory bank signal 930.
The first registered one of the threads whose local memory banks match is selected, and a thread state signal 936 is output from the corresponding thread state memory 903. Further, the thread start preparation flag 904 is also referred to and transmitted to the thread issuing unit 908.
【0107】プログラムカウンタシフトレジスタ909
は、32ビットのレジスタ4段で構成されるシフトレジ
スタである。分岐やスレッド発行が無い場合は、先頭の
値が命令キャッシュのタグの比較に使用される。Program counter shift register 909
Is a shift register composed of four stages of 32-bit registers. If there is no branch or thread issuance, the first value is used for comparing the instruction cache tag.
【0108】次に、命令キャッシュ判定機能について説
明する。命令メモリ404は、スレッド発行メモリ40
3のアクセスの度に1度、つまり1つのスレッドから見
れば4命令に1回だけキャッシュのアクセスを行う。し
かし、4つのスレッドが同時に存在することにより、ス
レッドの発行如何に関わらず命令キャッシュのチェック
は毎クロック行うことになる。プログラムカウンタシフ
トレジスタ909の先頭か、あるいはスレッド状態メモ
リ903から、プログラムカウンタがスレッド状態信号
935に出力される。このプログラムカウンタの下位の
ビットを使用してプログラムカウンタインデックス信号
937とし、命令キャッシュタグメモリ905の参照を
行う。プログラムカウンタ429の上位アドレス20ビ
ットは、プログラムカウンタアドレスタグ938として
出力される。そして、命令キャッシュアドレスタグ93
9と20ビット比較器907で一致するかどうかを確認
する。一致しない場合はキャッシュミスとなる。この場
合、今後発行される命令は取り消され、実行されている
スレッドは強制的にその命令メモリ404の左端アドレ
スへの分岐となる。Next, the instruction cache determination function will be described. The instruction memory 404 is the thread issue memory 40
The cache is accessed once every three accesses, that is, once per four instructions from one thread. However, since four threads exist at the same time, the instruction cache is checked every clock regardless of whether the thread is issued. The program counter is output to the thread status signal 935 from the top of the program counter shift register 909 or from the thread status memory 903. The lower bits of the program counter are used as a program counter index signal 937 to refer to the instruction cache tag memory 905. The upper 20 bits of the program counter 429 are output as a program counter address tag 938. Then, the instruction cache address tag 93
The 9-bit and 20-bit comparators 907 confirm whether they match. If they do not match, a cache miss occurs. In this case, the instruction issued in the future is canceled, and the thread being executed is forcibly branched to the leftmost address of the instruction memory 404.
【0109】最終的に、スレッドは、 1.スレッド発行アービトレーション901によりスレ
ッド発行の権限を取得 2.スタックポインタがローカルキャッシュバンク90
6の左端に一致する。 3.スレッド開始準備フラグ904が1である。 4.命令キャッシュアドレスが一致する。 これらの条件を全て満たしたときに初めてスレッドが発
行される。Finally, the thread: 1. Acquisition authority of thread issuance by thread issuance arbitration 901 The stack pointer is in the local cache bank 90
6 matches the left end. 3. The thread start preparation flag 904 is 1. 4. Instruction cache addresses match. A thread is issued only when all of these conditions are satisfied.
【0110】命令キャッシュリプレース要求、ローカル
メモリリプレース要求が発生した場合は、とりあえずス
レッド状態信号414は発行されるが、命令は実行され
ず、分岐ユニット502がグローバルアクセス制御40
8に要求するだけである。When an instruction cache replacement request or a local memory replacement request occurs, a thread status signal 414 is issued for the time being, but no instruction is executed, and the branch unit 502
8 only.
【0111】<グローバルアクセス制御ユニット408
>図10を参照して、グローバルアクセス制御ユニット
408の内部構造について概説する。ローカルメモリク
ロスバスイッチ406、分岐クロスバスイッチ402の
制御、およびそれぞれのスレッドの固有のローカルメモ
リ407以外のメモリアクセス調停、そしてマイクロプ
ロセッサ401の外のメモリアクセスの動作、命令、デ
ータメモリリプレース動作、割り込みの受理を行う。<Global access control unit 408>
> The internal structure of the global access control unit 408 will be outlined with reference to FIG. Control of the local memory crossbar switch 406 and branch crossbar switch 402, arbitration of memory access other than the local memory 407 unique to each thread, and operation of memory access outside the microprocessor 401, instruction, data memory replacement operation, and interruption of memory Perform acceptance.
【0112】グローバルアクセス制御ユニット408
は、分岐クロスバスイッチ制御1001、割り込みベク
タ生成ユニット1002、割り込み入力ユニット100
3、グローバルデータキャッシュタグメモリ1004、
グローバルデータキャッシュ比較器1005、グローバ
ルデータキャッシュメモリ1006、外部バスインター
フェースユニット1007、ローカルメモリクロスバス
イッチ制御1008、内部ロードストアインターフェー
ス1009、分岐受理ユニット1010、プログラムカ
ウンタインクリメンタ1011、命令キャッシュリプレ
ースバッファ1012、データキャッシュリプレースバ
ッファ1013で構成される。Global access control unit 408
Are the branch crossbar switch control 1001, the interrupt vector generation unit 1002, the interrupt input unit 100
3, global data cache tag memory 1004,
Global data cache comparator 1005, global data cache memory 1006, external bus interface unit 1007, local memory crossbar switch control 1008, internal load store interface 1009, branch accepting unit 1010, program counter incrementer 1011, instruction cache replacement buffer 1012, data It is composed of a cache replacement buffer 1013.
【0113】分岐クロスバスイッチ制御1001は、分
岐受理ユニット1010からの要求を元に、分岐クロス
バスイッチ402を制御する。The branch crossbar switch control 1001 controls the branch crossbar switch 402 based on a request from the branch receiving unit 1010.
【0114】割り込みベクタ生成ユニット1002は、
割り込み制御ユニット1003への割り込み入力453
に対して、特別なスレッドを生成する。このマイクロプ
ロセッサでは、割り込みは最高優先スレッド生成と等価
である。スレッドの情報は分岐クロスバスイッチ制御ユ
ニット1001によって分岐ユニットに伝達される。The interrupt vector generation unit 1002
Interrupt input 453 to the interrupt control unit 1003
Create a special thread for In this microprocessor, an interrupt is equivalent to creating a highest priority thread. The thread information is transmitted to the branch unit by the branch crossbar switch control unit 1001.
【0115】グローバルアクセスキャッシュメモリ10
04と、グローバルアドレスタグメモリ1005は、1
28ビット幅のキャッシュメモリである。キャッシュの
リプレースおよびローカルメモリ以外のアクセスを高速
化する目的で使用される。内部ロードストアインターフ
ェース1009からの要求で、グローバルアドレス信号
1021の示す内容を参照する。そして、グローバルア
クセスタグメモリ1005の内容と、グローバルアドレ
ス信号1021の上位ビットが一致すれば、グローバル
データキャッシュメモリ1004とデータバス信号10
22のやり取りを行う。Global access cache memory 10
04 and the global address tag memory 1005 are 1
This is a 28-bit width cache memory. Used to speed up cache replacement and non-local memory access. A request from the internal load store interface 1009 refers to the content indicated by the global address signal 1021. If the contents of the global access tag memory 1005 match the upper bits of the global address signal 1021, the global data cache memory 1004 and the data bus signal 1010
22 are exchanged.
【0116】外部バスインターフェースユニット100
7は、外部バス制御信号451と、外部アドレス信号4
52と、外部データバス信号453を出力し、外部から
のデータを外部データバス信号453から受理する。External bus interface unit 100
7 is an external bus control signal 451 and an external address signal 4
52, and an external data bus signal 453, and receives external data from the external data bus signal 453.
【0117】メモリクロスバスイッチ制御1008は、
毎クロック毎に接続が変更されるローカルメモリクロス
バスイッチ406の制御を行う。ただし、接続の変更は
演算器群405とデータキャッシュメモリ407とのロ
ーテート動作のみであり、自由に接続を変更できるわけ
ではない。The memory crossbar switch control 1008
The local memory crossbar switch 406 whose connection is changed every clock is controlled. However, the connection is changed only by the rotation operation between the arithmetic unit group 405 and the data cache memory 407, and the connection cannot be changed freely.
【0118】ロードストアインターフェース1009
は、グローバルアクセスアドレスバス427と、グロー
バルメモリアクセスデータバス428と接続される。ま
ず、スレッドからのロードストア要求であるグローバル
アクセスを1つ受理する。そして、グローバルアドレス
バス1021に出力する。外部メモリへのストアアクセ
スなら、即座にストアデータをグローバルデータバス1
022にアラインして出力する。外部から読み込まれた
データは、データキャッシュリプレースでなければグロ
ーバルメモリアクセスデータバス428に書き戻す。Load store interface 1009
Are connected to a global access address bus 427 and a global memory access data bus 428. First, one global access as a load store request from a thread is received. Then, the data is output to the global address bus 1021. For store access to external memory, store data is immediately transferred to global data bus 1
022 and output. Data read from outside is written back to the global memory access data bus 428 unless the data cache is replaced.
【0119】ロードストア分岐受理ユニット1010
は、4つの分岐ユニット502からの分岐要求信号43
3〜436を受理して、その要求に応じて分岐クロスバ
スイッチ1001に制御を要求する。Load store branch accepting unit 1010
Is the branch request signal 43 from the four branch units 502
3 to 436, and requests the branch crossbar switch 1001 for control in response to the request.
【0120】プログラムカウンタ制御ユニット1011
は、マイクロプロセッサ401の起動時にブート時のプ
ログラムカウンタを供給する。動作中は、スレッド制御
ユニット403からのプログラムカウンタ信号429を
入力して、インクリメントして次のプログラムカウンタ
信号412をスレッド制御ユニット403に出力する。Program counter control unit 1011
Supplies a boot-time program counter when the microprocessor 401 is started. During operation, a program counter signal 429 from the thread control unit 403 is input, incremented, and the next program counter signal 412 is output to the thread control unit 403.
【0121】命令キャッシュリプレースバッファ101
2は、命令キャッシュのリプレースをタイミングをスレ
ッドの再開に合わせるための同期を行うためのラッチで
ある。Instruction cache replacement buffer 101
Reference numeral 2 denotes a latch for synchronizing the replacement of the instruction cache with the resumption of the thread.
【0122】データキャッシュリプレースバッファ10
13は、4つのローカルデータキャッシュ407のリプ
レースを同時に行う。128ビット幅のリプレースデー
タを内部バス1022から受理し、4クロックに分けて
32ビットごとに1つのローカルデータキャッシュ40
7に転送する。 <データキャッシュメモリ407>図11を参照して、
データキャッシュ407の内部構造について概説する。
データキャッシュ407は、1つのポートのみを持つS
RAMであり、それに加えて他のスレッドからのアクセ
スを禁止するロック機構などを有している。Data cache replacement buffer 10
13 simultaneously replaces the four local data caches 407. Replacement data having a width of 128 bits is received from the internal bus 1022, and divided into four clocks, and one local data cache 40 is provided for every 32 bits.
Transfer to 7. <Data cache memory 407> Referring to FIG.
The internal structure of the data cache 407 will be outlined.
The data cache 407 is an S with only one port
It is a RAM, and further has a lock mechanism and the like for inhibiting access from other threads.
【0123】データキャッシュメモリは、1ポートのデ
ータキャッシュRAM1101、データタグRAM11
02、アドレスタグ比較器1103、そしてロック機構
制御機能1104で構成される。The data cache memory includes a one-port data cache RAM 1101, a data tag RAM 11
02, an address tag comparator 1103, and a lock mechanism control function 1104.
【0124】バスロック機能は、特定のスレッド以外の
アクセスを禁止する機構である。バスロック状態になる
と、その他のスレッドからのアクセスは、制御回路11
04が判定して、アクセス違反通知信号1114を返
す。The bus lock function is a mechanism for prohibiting access other than a specific thread. When the bus lock state is reached, access from other threads is restricted by the control circuit 11
04 returns an access violation notification signal 1114.
【0125】<共有浮動小数点ユニット409>図12
を参照して、共有浮動小数点ユニット409の内部構造
について概説する。共有浮動小数点ユニットは同時に1
つのスレッドだけが使用できる。浮動小数点演算が終了
した時に休止しているスレッドの生起要求を行う。<Shared floating point unit 409> FIG.
, The internal structure of the shared floating-point unit 409 will be outlined. Shared floating point units are 1 at the same time
Only one thread can be used. When the floating-point operation is completed, a request is made to generate a sleeping thread.
【0126】共有浮動小数点ユニット409は、ローカ
ル浮動小数点レジスタ1201、浮動小数点データパス
1202、アドレスデコードユニット1203、命令デ
コード1204で構成される。The shared floating point unit 409 comprises a local floating point register 1201, a floating point data path 1202, an address decode unit 1203, and an instruction decode 1204.
【0127】浮動小数点データバス1202は、データ
バス423から命令を受理するのと同時に、データバス
423からレジスタからのデータも受理してローカルレ
ジスタ1201に格納する。演算の終了とともに、グロ
ーバルアクセス制御ユニット408にスレッドの再開を
要求する。再開されたスレッドは、結果データをローカ
ルレジスタ1201から汎用レジスタに転送する。The floating-point data bus 1202 receives an instruction from the data bus 423 and, at the same time, receives data from a register from the data bus 423 and stores it in the local register 1201. At the end of the operation, the global access control unit 408 is requested to restart the thread. The resumed thread transfers the result data from the local register 1201 to the general-purpose register.
【0128】<マイクロプロセッサ401のプログラム
モデル>マイクロプロセッサ401の命令セットは、ス
レッド生成の高速化、コンテキストスイッチの高速化、
従来のマイクロプロセッサより大きい分岐レイテンシの
隠蔽を目的として作成されている。それ以外は従来のフ
ォンノイマン型マイクロプロセッサの命令セットに可能
な限り近づけてある。フォワードできないレジスタ依存
関係の発生によりパイプラインを自動的に停める、パイ
プラインストールの機能は実施例2ではインプリメント
されていないが可能である。<Program Model of Microprocessor 401> The instruction set of the microprocessor 401 is designed to speed up thread generation, speed up context switch,
It is created for the purpose of concealing branch latency larger than conventional microprocessors. Others are as close as possible to the instruction set of a conventional von Neumann microprocessor. The pipeline stall function for automatically stopping the pipeline due to the occurrence of a register dependency that cannot be forwarded is not implemented in the second embodiment, but is possible.
【0129】<命令セット>図13を参照して、本発明
のマイクロプロセッサ401の基本的命令セットを示
す。<Instruction Set> Referring to FIG. 13, a basic instruction set of microprocessor 401 of the present invention will be described.
【0130】基本的には一般的なRISCマイクロプロ
セッサと同じく、汎用レジスタ間だけで演算を行い、明
示的なロードストア命令でメモリにアクセスする命令セ
ットを持つ。ただし、汎用レジスタは、基本的に分岐の
前後で保持されない。キャッシュミス等ではレジスタの
退避は自動的に行われる。Basically, similarly to a general RISC microprocessor, there is an instruction set for performing operations only between general-purpose registers and accessing a memory with an explicit load / store instruction. However, general-purpose registers are not basically retained before and after branching. When a cache miss or the like occurs, the register is automatically saved.
【0131】フラグレジスタは、整数演算の結果によっ
て変化する。すべての命令はコンディションフィールド
を持ち、フラグレジスタの内容によって実行を制御する
ことができる。たとえば、ゼロフラグが1の場合に動作
する、キャリーフラグが0の場合以外で動作するという
ように使用する。よって特別な条件分岐命令はない。既
存のパイプラインプロセッサよりも分岐レイテンシが大
きいために、分岐を使用するのを極力避ける必要がある
ためである。The flag register changes according to the result of the integer operation. All instructions have a condition field, whose execution can be controlled by the contents of the flag register. For example, the operation is performed when the zero flag is 1, and the operation is performed when the carry flag is other than 0. Therefore, there is no special conditional branch instruction. This is because it is necessary to avoid using a branch as much as possible because the branch latency is higher than that of the existing pipeline processor.
【0132】演算命令は、整数演算に関するレジスタ間
の演算を行う。乗算、浮動小数点演算等は、演算ユニッ
トをスレッド間で共有するため、コプロセッサ命令とな
る。The operation instruction performs an operation between registers regarding an integer operation. Multiplication, floating-point operation, and the like are coprocessor instructions because the operation unit is shared between threads.
【0133】分岐命令は、指定されたレジスタとイミデ
ィエイト値を加算したアドレスに分岐する。分岐の前後
では汎用レジスタは保持されないため、分岐前には必要
なレジスタをスタックに退避し、分岐後に必要なレジス
タをスタックから再度読み出す必要がある。条件実行は
分岐命令に対して当然適用できるため、専用の条件分岐
命令は存在しない。A branch instruction branches to an address obtained by adding a specified register and an immediate value. Since general-purpose registers are not held before and after the branch, necessary registers must be saved on the stack before the branch, and the necessary registers must be read from the stack again after the branch. Since conditional execution is naturally applicable to branch instructions, there is no dedicated conditional branch instruction.
【0134】スレッド管理命令は、スレッドの生成、終
了に使用される。FORK命令は現在のスレッドと別
に、汎用レジスタで指定したアドレス、スタックポイン
タ、スレッド情報のスレッドを新規に生成できる。CH
GTH命令は従来のスレッドを終了して指定したスレッ
ドを新規に生成する。EXIT命令は従来のスレッドを
終了する。The thread management instruction is used for creating and terminating a thread. The FORM instruction can newly generate a thread of the address, the stack pointer, and the thread information specified by the general-purpose register separately from the current thread. CH
The GTH instruction terminates a conventional thread and newly generates a specified thread. The EXIT instruction terminates the conventional thread.
【0135】同期命令は、マルチスレッドに不可欠なア
トミックロードストアをサポートする。LLOCK命令
は、データを読み出すと同時に、アクセスしたローカル
メモリバンク407をロックする。他のスレッドがロッ
ク状態のメモリロードを行うと強制的にパイプラインス
トールとなりスレッドが休眠状態となる。STULK命
令はデータをストアすると同時にロック状態のメモリを
アンロックする。Synchronous instructions support atomic load stores that are essential for multithreading. The LLOCK instruction locks the accessed local memory bank 407 at the same time as reading data. When another thread loads the memory in the locked state, the pipeline is forcibly installed, and the thread enters a sleep state. The STULK instruction unlocks the locked memory at the same time as storing the data.
【0136】システム割り込み命令は、レジスタの値が
示す値とスレッドIDの値が一致する休眠状態のスレッ
ドを再開させる。外部からの割り込みの挙動と全く同一
である。The system interrupt instruction restarts the sleeping thread whose thread ID value matches the value indicated by the register value. The behavior is exactly the same as that of an external interrupt.
【0137】コプロセッサ命令は、スレッド間で共有す
る演算ユニットを使用する命令である。コプロセッサ命
令は、整数演算ユニットの汎用レジスタを直接オペラン
ドとして転送して使用ことになる。同時に使用できるス
レッドは常に1つという制限がある。A coprocessor instruction is an instruction that uses an operation unit shared between threads. The coprocessor instruction uses the general-purpose register of the integer operation unit by directly transferring it as an operand. There is always a limit of one thread that can be used simultaneously.
【0138】<マルチスレッドプログラミングモデル>
図14を参照して、マルチスレッドプログラムがどのよ
うに実行されるかを示す。<Multithread programming model>
Referring to FIG. 14, how the multi-thread program is executed will be described.
【0139】本発明のマイクロプロセッサ401は、同
時に動作出来るスレッドは16個であるが、常に演算機
の数以上の数のスレッドを同時処理できる。実行待ち状
態のスレッド1401は、使用するスタックポインタが
示すローカルメモリバンクで分割される。スレッドは各
メモリバンクごとに格納され、実行状態スレッド140
2の空きを待つ。実行状態のスレッド1402は、1つ
のローカルメモリ407に対して1つだけのスレッドが
割り当てられる。ローカルメモリバンクの異なるスレッ
ドはスレッドが空いていても実行状態に送ることはでき
ない。The microprocessor 401 of the present invention can simultaneously operate 16 threads, but can always simultaneously process more threads than the number of arithmetic units. The thread 1401 waiting to be executed is divided by the local memory bank indicated by the stack pointer to be used. Threads are stored for each memory bank, and the execution state thread 140
Wait for an empty 2. Only one thread is allocated to one local memory 407 for the thread 1402 in the execution state. Threads in different local memory banks cannot be sent to the running state, even if the thread is free.
【0140】実行状態の全てのスレッドは、FORK命
令によって別のスレッドを生成することができる。生成
されたスレッドは待ち状態スレッド1401にスタック
ポインタのローカルメモリアドレスに応じた番号に格納
される。スレッドはEXIT命令で自分自身を停止する
ことも可能である。All the threads in the execution state can generate another thread by the FORK instruction. The generated thread is stored in the waiting state thread 1401 at a number corresponding to the local memory address of the stack pointer. A thread can stop itself with an EXIT instruction.
【0141】本発明のマイクロプロセッサでは、分岐命
令はスレッドの切り替えと全く等価である。分岐命令は
スレッドは待ち状態スレッド1401に入れられ、別の
待ち状態のスレッドが1つ実行状態スレッド1402に
送られ、実行される。単体のスレッドのレイテンシに関
しては従来のマイクロプロセッサに対してかなり劣る
が、別のスレッドの動作によってレイテンシは隠蔽され
る。単体のスレッドの処理レイテンシよりも、全体の演
算ユニットの使用効率を最大にすることを優先させる。In the microprocessor of the present invention, the branch instruction is completely equivalent to the thread switching. In the branch instruction, the thread is put into the waiting thread 1401, and another waiting thread is sent to the execution thread 1402 to be executed. The latency of a single thread is considerably inferior to that of a conventional microprocessor, but the operation of another thread hides the latency. Priority is given to maximizing the use efficiency of the entire arithmetic unit over the processing latency of a single thread.
【0142】命令メモリ1403は、4つの命令メモリ
404から構成される。すべてのスレッドは1つの命令
メモリ1403をアクセスできる。スレッドはループの
個別のイタレーションを発行することが多く、同じコー
ドを使用するため、命令メモリ1403の共有は有用で
ある。The instruction memory 1403 is composed of four instruction memories 404. All threads can access one instruction memory 1403. Because threads often issue separate iterations of the loop and use the same code, sharing instruction memory 1403 is useful.
【0143】また、スレッドの発行がローカルメモリの
バンクによって制限される構成にした理由は、16個に
も上るローカルメモリアクセスの調停を省略するためで
ある。16の入力に対してその都度アービトレーション
を行うと、メモリアクセスのレイテンシの低下は避けら
れない。The reason why the thread issuance is limited by the bank of the local memory is to arbitrate as many as 16 local memory accesses. When arbitration is performed for each of the 16 inputs, a reduction in the latency of memory access is inevitable.
【0144】<レジスタセット>図15を参照して、マ
イクロプロセッサ401のレジスタセットについて説明
する。<Register Set> The register set of the microprocessor 401 will be described with reference to FIG.
【0145】R0は常に0が読み出され、書き込んだ値
が保持されないレジスタである。R1,R2、R3,R
4は汎用レジスタであり、自由に読み書きができる。R0 is a register from which 0 is always read and the written value is not held. R1, R2, R3, R
Reference numeral 4 denotes a general-purpose register which can be freely read and written.
【0146】SPはスタックポインタであり、常にロー
カルメモリのアドレスを示し、1クロックでアクセスで
きる。コンテキストスイッチの時は、レジスタが自動的
に格納され、SPだけが退避される。また、スレッドの
時はSPのアドレスが示すメモリバンクをローカルメモ
リとする。SP is a stack pointer, which always indicates the address of the local memory, and can be accessed in one clock. At the time of a context switch, the register is automatically stored, and only the SP is saved. In the case of a thread, the memory bank indicated by the address of the SP is set as a local memory.
【0147】TIはスレッド状態を示す。TIDは現在
実行しているスレッドを番号で示すスレッドIDであ
り、すべてのスレッドに対して別にに与える必要があ
る。FLGは4ビットのフラグレジスタの実体である。[0147] TI indicates a thread state. The TID is a thread ID indicating the currently executing thread by a number, and needs to be separately given to all threads. FLG is an entity of a 4-bit flag register.
【0148】PCは現在の命令アドレスに対して相対的
な分岐を実行するために使用する読み出し専用レジスタ
である。PC is a read-only register used to execute a branch relative to the current instruction address.
【0149】これらのレジスタの中で、スレッドの再開
に必要な情報はPC、SP、TIだけである。その他の
レジスタはコンテキストスイッチの際に全てローカルメ
モリのSPの示すアドレスに退避する。Among these registers, the information necessary for resuming the thread is only PC, SP, and TI. All other registers are saved to the address indicated by the SP of the local memory at the time of the context switch.
【0150】<メモリマップと負荷分散>図16を参照
して、ローカルメモリの配置およびスレッドの負荷分散
方法を説明する。<Memory Map and Load Balancing> Referring to FIG. 16, a local memory arrangement and a thread load balancing method will be described.
【0151】本発明のマイクロプロセッサはマルチスレ
ッドを前提とする。そのため同時に存在するスレッドは
全て固有のスタックを所持することになる。スレッドへ
のローカルメモリバンクの割り当ては、スタックポイン
タの示すメモリバンクと同義である。The microprocessor of the present invention is premised on multi-threading. Therefore, all the threads that exist at the same time have their own stack. The assignment of a local memory bank to a thread is synonymous with the memory bank indicated by the stack pointer.
【0152】本発明のマイクロプロセッサはメモリは1
6個に分散させる。スレッドはローカルヒープメモリへ
のアクセスを極力使用するようにすれば最高性能を出せ
る。逆にいえば、本発明のマイクロプロセッサは、OS
による負荷分散を考慮したヒープメモリの配置を前提と
している。In the microprocessor of the present invention, the memory is 1
Disperse into 6 pieces. Threads perform best if they make the best use of local heap memory access. Conversely, the microprocessor according to the present invention uses the OS
It is assumed that heap memory is arranged in consideration of load distribution by
【0153】メモリ空間は、全てローカルキャッシュメ
モリバンクに割り付けられる。1つのローカルメモリバ
ンクは1Mバイトの連続したメモリ空間をアクセス対象
とする。ローカルメモリのバンクは、アドレスバスの2
0から23ビット目によって一意に決定する。All memory spaces are allocated to local cache memory banks. One local memory bank accesses a continuous memory space of 1 Mbyte. The local memory bank is the address bus 2
It is uniquely determined by the 0th to 23rd bits.
【0154】16のパイプラインに適切にプログラムの
負荷分散を行うということは、プロセスに対して個別の
ローカルメモリバンクを割り当てるということと等価で
ある。頻度の低いプロセスは1つのローカルメモリバン
クを共用し、性能が必要であるか、あるいはリアルタイ
ムレイテンシの必要なプロセスは1つ以上のローカルメ
モリバンクを占有すれば良い。それだけでプライオリテ
ィー制御などのOSの制御を必要としない負荷分散が可
能になる。Properly distributing the program load to the 16 pipelines is equivalent to allocating individual local memory banks to processes. Infrequent processes may share one local memory bank and require performance, or processes requiring real-time latency may occupy one or more local memory banks. This alone enables load distribution that does not require OS control such as priority control.
【0155】スレッドに割り当てられたローカルキャッ
シュメモリへのアクセスは1クロックで終了する。しか
し、それ以外に割り当てられたローカルメモリキャッシ
ュへのアクセスは16クロック以上要する。キャッシュ
がミスした場合は更に外部へのアクセスになり、レイテ
ンシは不定である。しかし、ローカルキャッシュ以外の
アクセスレイテンシは、マルチスレッド機構によりある
ていど隠蔽可能である。Access to the local cache memory allocated to the thread is completed in one clock. However, access to the other allocated local memory cache requires 16 clocks or more. If the cache misses, access to the outside is further performed, and the latency is undefined. However, access latencies other than the local cache can be concealed by the multi-thread mechanism.
【0156】演算ユニットがアクセスできるメモリバン
クは、起動時に順に割り当てられている。分岐には、分
岐先のアドレスの演算ユニットが空いているだけではだ
めで、その時点でのメモリバンクがスタックポインタと
一致する必要がある。そのため、最悪の場合、分岐命令
発行から16クロック後にスレッドを再開するケースが
発生する。その場合は、同じメモリバンクを使用し、命
令アドレスの異なる別のスレッドを用意して先に実行で
きる方からスレッドを発行する。当然、単体のスレッド
のレイテンシは更に低下するが、全体の性能低下を防止
でき、単体のスレッドのレイテンシを隠蔽できる。The memory banks that can be accessed by the arithmetic unit are sequentially allocated at the time of startup. For branching, it is not enough that the operation unit at the address of the branch destination is empty, and the memory bank at that time must match the stack pointer. Therefore, in the worst case, the thread may be restarted 16 clocks after the issuance of the branch instruction. In this case, the same memory bank is used, another thread having a different instruction address is prepared, and a thread is issued from a person who can execute the thread first. Naturally, the latency of a single thread is further reduced, but the overall performance can be prevented from being reduced, and the latency of a single thread can be hidden.
【0157】<マイクロプロセッサ401の内部動作>
次に、マイクロプロセッサ401の動作について概説す
る。<Internal Operation of Microprocessor 401>
Next, the operation of the microprocessor 401 will be outlined.
【0158】図17を参照して、基本的な命令の動作を
中心に説明する。マイクロプロセッサ401の命令ごと
の基本動作を示す為に、縦軸を時間軸にとってブロック
をそれぞれの動作順序に従って配置している。単体の命
令の実行は、初段のスレッド発行、汎用レジスタを除け
ば、従来のRISCパイプラインプロセッサに近い構成
である。標準的なRISCパイプラインについては、コ
ンピューターの構造と設計ハードウェアとソフトウェア
のインターフェース(著者 David A.Patt
erson/John L.Hennessy 出版社
日経BP社)の記述を参照のこと。Referring to FIG. 17, the operation of the basic instruction will be mainly described. In order to show the basic operation of each instruction of the microprocessor 401, the blocks are arranged in the order of operation with the vertical axis as the time axis. The execution of a single instruction has a configuration similar to that of a conventional RISC pipeline processor, except for the first-stage thread issue and the general-purpose register. For a standard RISC pipeline, see Computer Structure and Design Hardware and Software Interfaces (Author David A. Patt
erson / John L. Hennessy publisher Nikkei BP).
【0159】最初のThreadIssueステージで
は、スレッドの発行を行う。今後は略称としてTIステ
ージと呼ぶ。In the first ThreadIssue stage, a thread is issued. In the future, it will be referred to as the TI stage as an abbreviation.
【0160】このTIステージに限り、すべての命令に
対して必要であるわけではなく、分岐やスレッド生成の
直後、あるいは4命令おきに実行される。This TI stage is not necessary for all instructions, and is executed immediately after branching or thread generation, or every four instructions.
【0161】命令読み出しは4クロック単位でまとめて
行われるため、命令キャッシュのTAGのチェックも4
命令置きになる。分岐やパイプラインストールの再開な
どで、命令実行プログラムカウンタが41でアラインさ
れていない場合は、その命令まで演算器が動作しないこ
とになるが、分岐は可能である。Since instruction reading is performed in units of four clocks, the TAG check of the instruction cache is also performed in four clocks.
It becomes a command place. If the instruction execution program counter is not aligned at 41 due to branching or restart of pipeline stall, the arithmetic unit will not operate until that instruction, but branching is possible.
【0162】命令キャッシュ、データキャッシュのリプ
レースは、そのメモリバンク407を使用する、リプレ
ースの必要のないスレッドが一切なくなるまで実行され
ない。The replacement of the instruction cache and the data cache is not executed until there are no threads that use the memory bank 407 and do not need to be replaced.
【0163】次のInstructionFetchス
テージでは、命令の読み出し、デコードを行う。略称と
してIFステージと呼ぶ。In the next InstructionFetch stage, instructions are read and decoded. It is called an IF stage as an abbreviation.
【0164】1つのプログラムカウンタを受理し、4つ
の命令を一度に発行する。最初の1つの命令だけが即座
に命令デコード814に入力される。続く命令は1クロ
ック分ラッチ806で保存されて次のクロックでデコー
ドされる。3クロック後の命令は3クロック分のラッチ
806で維持することになる。命令デコード814は1
命令単位で命令デコードを行う。パイプラインEXステ
ージの制御信号は即座に発行され、後続のパイプライン
DFステージ,WBステージの右に隣接する命令デコー
ドに渡されて次以降のステージの制御を行う。逆に、左
に隣接する命令のDFステージ制御、WBステージ制御
信号と共に演算ユニット501に出力される。One program counter is accepted and four instructions are issued at a time. Only the first one instruction is immediately input to instruction decode 814. The following instruction is stored in the latch 806 for one clock, and is decoded at the next clock. The instruction after three clocks is maintained in the latch 806 for three clocks. Instruction decode 814 is 1
Instruction decoding is performed in instruction units. The control signal of the pipeline EX stage is immediately issued, and is passed to the instruction decode adjacent to the right of the subsequent pipeline DF stage and WB stage to control the next and subsequent stages. Conversely, it is output to the arithmetic unit 501 together with the DF stage control and WB stage control signals of the instruction adjacent to the left.
【0165】通常のマイクロプロセッサでは、レジスタ
は、ラッチより比較的低速なレジスタファイルに格納さ
れる。そして、デコードとレジスタのアクセスを対にし
てインストラクションデコードステージ(略称IDステ
ージ)と呼ばれる1つのパイプラインステージを形成す
る。しかし、実施例2のマイクロプロセッサ401は、
レジスタ数が少なく、毎クロックごとに右にシフトする
構成を取るため、レジスタファイルを持たず、ラッチか
ら直接ドライブする。しかし、命令キャッシュ読み出
し、命令デコードの速度によっては独立したIDステー
ジが必要になる。In a typical microprocessor, registers are stored in a register file that is relatively slower than a latch. Then, one pipeline stage called an instruction decode stage (abbreviated ID stage) is formed by pairing the decode and the access of the register. However, the microprocessor 401 of the second embodiment includes:
Since the number of registers is small and the clock is shifted to the right at every clock, the drive is performed directly from the latch without the register file. However, an independent ID stage is required depending on the speed of instruction cache reading and instruction decoding.
【0166】次のExecutionステージでは、演
算の実行を行う。略称としてEXステージと呼ぶ。In the next Execution stage, an operation is performed. It is called EX stage as an abbreviation.
【0167】デコードされた命令内容に応じて、レジス
タや既に実行された命令のDFステージ、WBステージ
からのデータから、演算に使用するオペランドから選択
する。オペランド値は3つ用意され、整数演算ALU6
30、バレルシフタ631、ストアアライナ632など
に入力され、結果をラッチ609などにとりあえず格納
する。In accordance with the decoded instruction content, a register or data from the DF stage or WB stage of an already executed instruction is selected from the operands used in the operation. Three operand values are prepared and the integer operation ALU6
30, the barrel shifter 631, the store aligner 632, etc., and store the result in the latch 609 or the like.
【0168】また、命令キャッシュのアクセスを行った
場合、命令TAGのチェックを比較器909によって行
う。命令キャッシュにミスが生じた場合は、その命令の
後続のDF、WBのステージは無効になる。When the instruction cache is accessed, the instruction TAG is checked by the comparator 909. When a miss occurs in the instruction cache, the DF and WB stages following the instruction become invalid.
【0169】次のDataFetchステージでは、デ
ータメモリのアクセスを行う。略称としてDFステージ
と呼ぶ。In the next DataFetch stage, the data memory is accessed. It is called a DF stage as an abbreviation.
【0170】ロードストア命令では、ALU630を使
用してアドレス計算を行う。算出したアドレスは即座に
クロスバスイッチ406を介してローカルキャッシュメ
モリ407に送られる。ロード命令の場合は即座にデー
タを読み出し、クロスバスイッチ406、ロードアライ
ナ710を介して元のプロセッサに送られる。In the load store instruction, the ALU 630 is used to calculate an address. The calculated address is immediately sent to the local cache memory 407 via the crossbar switch 406. In the case of a load instruction, the data is immediately read out and sent to the original processor via the crossbar switch 406 and the load aligner 710.
【0171】ストア命令の場合は、ローカルメモリアド
レス、キャッシュミスの判定が必要なため、DFステー
ジで即座に書き込むことはしない。DFステージで判断
できた後に次の命令のDFステージに相当するタイミン
グでデータを書き込む。この場合、次の命令のDFステ
ージのリードとかち合うことになるため、データキャッ
シュ407は、独立したリードライトを同時に処理でき
る構成が望ましい。In the case of a store instruction, since it is necessary to determine a local memory address and a cache miss, it is not immediately written in the DF stage. After the determination in the DF stage, data is written at a timing corresponding to the DF stage of the next instruction. In this case, the data cache 407 is desirably configured to be able to process independent read / write simultaneously, because the read operation in the DF stage of the next instruction is mixed.
【0172】グローバルメモリへのアクセスかどうかは
ローカルメモリへのアクセスと並行してDFステージで
判断され、グローバルリードアクセスであった場合は分
岐ユニット709によってスレッド待機状態に移行す
る。The access to the global memory is determined in the DF stage in parallel with the access to the local memory. If the access is a global read access, the branch unit 709 shifts to a thread waiting state.
【0173】分岐命令の場合は、分岐先プログラムカウ
ンタはALU630で算出され、分岐ユニット709に
よってSPやTIレジスタと共にスレッド発行ユニット
403に送られる。In the case of a branch instruction, the branch destination program counter is calculated by the ALU 630 and sent to the thread issuing unit 403 by the branch unit 709 together with the SP and the TI register.
【0174】グローバルデータリードなどのパイプライ
ンストールの場合は、現在のプログラムカウンタがその
まま分岐ユニット709に送られ、分岐の処理が行われ
る。分岐命令が成立した場合は、内蔵シーケンサによ
り、後続のクロックのDFステージではその他の汎用レ
ジスタのスタックへの書き戻しが自動的に行われる。In the case of pipeline stall such as global data read, the current program counter is sent to the branch unit 709 as it is, and branch processing is performed. When the branch instruction is taken, the built-in sequencer automatically rewrites the other general-purpose registers to the stack in the DF stage of the subsequent clock.
【0175】最後のWriteBackステージでは、
レジスタへの書き戻しを行う。略称としてWBステージ
と呼ぶ。In the last WriteBack stage,
Write back to the register. Abbreviated name is called WB stage.
【0176】従来のマイクロプロセッサでは比較的低速
なレジスタファイルのために設けられるステージであ
る。ローカルメモリへのアクセスはクリティカルパスに
なりやすいため、ローカルキャッシュメモリ407から
のロードデータがラッチ702が間に合わない場合は、
WBステージにもロードアライナ710などの処理が入
る。その場合、WBステージからEXステージのALU
へのデータフォワードが不可能になり、データロードレ
イテンシが低下する。In a conventional microprocessor, this is a stage provided for a relatively slow register file. Since access to the local memory is likely to be a critical path, if the load data from the local cache memory 407 cannot keep up with the latch 702,
Processing such as the load aligner 710 also enters the WB stage. In that case, ALU from WB stage to EX stage
Data forward is impossible, and the data load latency is reduced.
【0177】<パイプラインの全体図>図18を参照し
て、本発明のマイクロプロセッサ401におけるパイプ
ラインの共存について説明する。<Overall Diagram of Pipeline> The coexistence of pipelines in the microprocessor 401 of the present invention will be described with reference to FIG.
【0178】本発明のマイクロプロセッサは、16のパ
イプラインマイクロプロセッサが互いに結合して内蔵さ
れている。よって、4+64=68のパイプラインステ
ージに匹敵する回路が1つのチップ上に同時に存在す
る。The microprocessor of the present invention has 16 pipeline microprocessors connected to each other. Therefore, circuits equivalent to 4 + 64 = 68 pipeline stages are simultaneously present on one chip.
【0179】通常のマルチプロセッサでは、複数のパイ
プラインステージは全く独立しており、1つのスレッド
は1つのプロセッサに静的に割付られる。本発明のマイ
クロプロセッサ401では、1つのスレッドは、これら
の68のパイプラインステージを全て使用することがで
きるが、同時刻に使用できるのは5つだけである。だ
が、1つのスレッドに着目すれば、1801のように通
常の5段パイプラインプロセッサの動作と同等の表記が
できる。ただし、同じ回路を連続して使用することはな
く、常に右に隣接するパイプラインステージを使用して
いる点が異なる。本発明のマイクロプロセッサは、Th
read2からThread8までの16のスレッドが
同時に動作できる。同時刻に同じパイプラインステージ
を使用することはない。たとえば、Thread2がI
F2を使用している時刻では、Thread14はIF
3を使用している。スレッド2がDF2を使用している
時刻ではThread14はDF3を使用している。In a normal multiprocessor, a plurality of pipeline stages are completely independent, and one thread is statically allocated to one processor. In the microprocessor 401 of the present invention, one thread can use all of these 68 pipeline stages, but only five can be used at the same time. However, if attention is paid to one thread, a notation equivalent to the operation of a normal five-stage pipeline processor such as 1801 can be obtained. However, the difference is that the same circuit is not used successively and the pipeline stage adjacent to the right is always used. The microprocessor of the present invention has a Th
16 threads from read2 to Thread8 can operate simultaneously. No two pipeline stages are used at the same time. For example, if Thread2 is I
At the time when F2 is used, Thread 14 is set to IF
3 is used. At the time when the thread 2 uses DF2, the Thread 14 uses DF3.
【0180】<分岐命令におけるパイプライン動作>図
19を参照して、分岐命令の動作について説明する。<Pipeline Operation in Branch Instruction> The operation of a branch instruction will be described with reference to FIG.
【0181】分岐命令は、その前後でレジスタの保存が
行われない。そのため、明示的にレジスタの退避を行う
必要がある。分岐命令はDFステージで実行され、スレ
ッド1を待ち状態とする。空き状態となったバンクは、
次のTIステージで待ち状態のスレッド4が発行され
る。The branch instruction does not save the register before and after it. Therefore, it is necessary to explicitly save the register. The branch instruction is executed in the DF stage, and places thread 1 in a wait state. The vacant bank is
In the next TI stage, the thread 4 in the waiting state is issued.
【0182】待ち状態のスレッド1は、分岐先の命令ア
ドレスの下位4ビットのアドレスが一致する整数演算ユ
ニットが未使用であれば、分岐を受理してスレッド1を
再開する。スレッドを発行しても、パイプラインが命令
アドレスの下位2ビットも一致する演算ユニットまで到
達して初めて動作を開始する。If the integer arithmetic unit whose lower four bits of the instruction address of the branch destination match is not used, the thread 1 in the waiting state accepts the branch and restarts the thread 1. Even if a thread is issued, the operation starts only when the pipeline reaches an arithmetic unit that also matches the lower two bits of the instruction address.
【0183】<パイプラインストールにおけるパイプラ
インの動作>図20を参照して、キャッシュミス時の動
作について説明する。<Operation of Pipeline in Pipeline Install> The operation at the time of a cache miss will be described with reference to FIG.
【0184】キャッシュミスなどの突発事項は、命令の
処理から隠蔽される必要があるため、レジスタの退避は
自動的に行う必要がある。EXステージで発覚した命令
キャッシュのミスは、その直後以降の命令のEXステー
ジをスタックポインタアドレス計算にまわす。4つのレ
ジスタ状態を退避した時点で別のスレッド4を開始でき
るようなら、スレッドを再開する。Since unexpected matters such as cache misses need to be hidden from the processing of instructions, it is necessary to save registers automatically. An instruction cache miss found in the EX stage causes the EX stage of the immediately following instruction to be passed to the stack pointer address calculation. If another thread 4 can be started when the four register states have been saved, the thread is restarted.
【0185】スレッド1の再開は、4命令分前の命令か
ら行われ、4つのレジスタの再度格納が行われる。Thread 1 is restarted from the instruction four instructions before, and the four registers are stored again.
【0186】<マイクロプロセッサの内部動作の詳細>
ここからは、これまで説明した命令の内部動作がどのよ
うに実現されているかを詳細に示す。<Details of Internal Operation of Microprocessor>
The following describes in detail how the internal operations of the instructions described so far are implemented.
【0187】最初に、分岐、スレッド生成動作について
説明する。本発明のマイクロプロセッサでは、分岐とス
レッドの生成は、どちらもスレッド発行ユニット403
へのスレッド構造体の伝達を行う。分岐とスレッド生成
の違いは、前者は元のパイプラインを開放して別のスレ
ッドを動作させるが、後者は元のパイプラインをそのま
ま実行を続けることになる。First, the branching and thread generation operations will be described. In the microprocessor of the present invention, both the branch and the thread generation are performed by the thread issuing unit 403.
Of the thread structure to the The difference between branching and thread creation is that the former releases the original pipeline and runs another thread, while the latter continues executing the original pipeline.
【0188】次に、パイプラインストールの動作につい
て説明する。本発明のマイクロプロセッサでは、パイプ
ラインストールは自動的なレジスタの退避を伴う自身ア
ドレスへの分岐となる。分岐は、キャッシュリプレース
や例外処理などの動作を終了させた後、最短で16クロ
ック後に実行される。例外事象でパイプラインを停める
場合は、常にパイプラインストールとなる。Next, the operation of the pipeline stall will be described. In the microprocessor of the present invention, the pipeline stall is a branch to its own address accompanied by automatic register saving. The branch is executed at least 16 clocks after ending operations such as cache replacement and exception processing. When stopping the pipeline due to an exception event, the pipeline is always installed.
【0189】次に、命令キャッシュミスについて説明す
る。命令キャッシュの判定は常にスレッド発行ユニット
403で行われる。命令キャッシュのリプレースが必要
な場合は、即座にパイプラインストールとなる。同時
に、命令キャッシュリプレースの要求をグローバルアク
セスユニット408とスレッド発行ユニット403に要
求する。命令キャッシュリプレースサイクルは、命令の
リプレースとスレッドの再開は同時に実行される。他の
スレッドの干渉を防ぐために128ビット幅の命令リプ
レースバス432を使用し、再開されたスレッドの読み
出しの最初の1クロックで行われる。Next, an instruction cache miss will be described. The instruction cache is always determined by the thread issuing unit 403. If the instruction cache needs to be replaced, it immediately becomes pipeline stall. At the same time, a request for instruction cache replacement is made to the global access unit 408 and the thread issuing unit 403. In the instruction cache replacement cycle, instruction replacement and thread resumption are performed simultaneously. The instruction replacement bus 432 having a width of 128 bits is used to prevent the interference of other threads, and is performed in the first clock of the read of the resumed thread.
【0190】次に、ローカルキャッシュメモリ407へ
のメモリロードストアについて説明する。スレッドが固
有に所持する1つのローカルキャッシュメモリ407は
1クロックで無条件にアクセスができる。データストア
は、ローカルバンクの判定、キャッシュのタグの判定を
待たなければ書き込みができないため、次の命令のDF
ステージで書き込みを行うことになる。次の命令のDF
ステージがロードである場合は、データライトは次のク
ロックに持ち越される。しかし、キャッシュミスのばあ
いを除いて、ローカルキャッシュへのリードとライトの
組み合わせでパイプラインが止まることはない。Next, a memory load store to the local cache memory 407 will be described. One local cache memory 407 uniquely owned by a thread can be accessed unconditionally in one clock. The data store cannot be written without waiting for the judgment of the local bank and the judgment of the tag of the cache.
Writing will be performed on the stage. DF of next instruction
If the stage is a load, the data write is carried over to the next clock. However, except in the case of a cache miss, the pipeline does not stop by the combination of reading and writing to the local cache.
【0191】次に、ローカルメモリ以外へのメモリスト
ア動作について説明する。スレッドが所持していないロ
ーカルキャッシュメモリ407は、そのローカルメモリ
を所持するスレッドがメモリをアクセスしないタイミン
グを狙う必要がある。アドレスバッファ707、データ
バッファ709に保持され、アクセス対象のキャッシュ
バンクを所持する他のスレッドが到達した時点でローカ
ルアドレス信号521、ローカルデータバス信号524
に出力する。ただし、本来のスレッドがロードストアを
行う場合は1回目はそちらが優先である。16クロック
待ち、2回目は本来のスレッドをパイプラインストール
させてローカルキャッシュへの出力を行うことができ
る。本来のスレッドのロードストアは、アドレスラッチ
704、データラッチ705に格納され、命令と関係な
く次のクロックで実行される。Next, the operation of storing data in a memory other than the local memory will be described. For the local cache memory 407 not owned by the thread, it is necessary to aim for a timing at which the thread possessing the local memory does not access the memory. The local address signal 521 and the local data bus signal 524 are held when another thread held in the address buffer 707 and the data buffer 709 and holding the cache bank to be accessed arrives.
Output to However, when the original thread performs load store, that thread has priority in the first time. After waiting for 16 clocks, the second time, the original thread can be pipeline-installed and output to the local cache. The original thread load store is stored in the address latch 704 and the data latch 705 and executed at the next clock regardless of the instruction.
【0192】次に、データキャッシュミスについて説明
する。データキャッシュメモリバンクをSPのメモリバ
ンクとして所有するスレッドは、データキャッシュミス
が発生した時点で、パイプラインストールを発生させ、
レジスタの退避を行うことができる。スタックポインタ
のキャッシュミスの場合は、ラッチを保持し、スタック
領域のキャッシュのリプレースを待ってレジスタの退避
を行う。Next, a data cache miss will be described. The thread that owns the data cache memory bank as the SP memory bank generates a pipeline stall when a data cache miss occurs,
The register can be saved. In the case of a cache miss of the stack pointer, the latch is held, and the register is saved after the cache in the stack area is replaced.
【0193】データキャッシュメモリバンクをSPのメ
モリバンクとして所持しないスレッドは、データキャッ
シュミスが発生したら即座にグローバルメモリアクセス
に切り替え、外部から値を読み込む。キャッシュリプレ
ース動作は行わない。A thread that does not possess the data cache memory bank as the SP memory bank immediately switches to global memory access when a data cache miss occurs, and reads a value from outside. No cache replacement operation is performed.
【0194】スタックポインタのキャッシュミスは、リ
プレース中の別のスレッドの動作を禁止し、状態をパイ
プライン上でシフトして保持する。The cache miss of the stack pointer inhibits the operation of another thread being replaced, and shifts and holds the state on the pipeline.
【0195】次に、グローバルメモリアクセスについて
説明する。ローカルキャッシュ407のキャッシュミ
ス、キャッシュにアサインされていない外部のメモリへ
のアクセスで発生する。Next, global memory access will be described. A cache miss of the local cache 407 or access to an external memory not assigned to the cache occurs.
【0196】グローバルメモリアクセスが発生した場
合、分岐制御711が分岐・ロードストアユニットにグ
ローバルアクセスバス427、428のアービトレーシ
ョンをリクエストする。バスが取得できるまではアドレ
スバッファ707、データバッファ708に格納され
る。When a global memory access occurs, the branch control 711 requests the arbitration of the global access buses 427 and 428 to the branch / load store unit. Until the bus can be acquired, the data is stored in the address buffer 707 and the data buffer 708.
【0197】グローバルメモリは、一連のデータのリー
ドライト、キャッシュのリプレースが終了するまで他の
スレッドの入力を受け付けない。リードライト動作の終
了と同時に分岐クロスバ制御ユニット1001からスレ
ッド発行ユニット503に伝達される。The global memory does not accept input from another thread until a series of data read / write and cache replacement are completed. The data is transmitted from the branch crossbar control unit 1001 to the thread issuing unit 503 simultaneously with the end of the read / write operation.
【0198】読み込まれた値はデータキャッシュリプレ
ースバス425に入力され、4クロック要してローカル
キャッシュに伝達される。同時に4つのデータキャッシ
ュリプレースを可能にする。The read value is input to the data cache replacement bus 425, and transmitted to the local cache after four clocks. This allows four data cache replacements at the same time.
【0199】次に、外部メモリアクセスについて説明す
る。外部メモリアクセスは、グローバルメモリのデータ
キャッシュミスの時に発生する。アドレスバス1022
の内容はそのまま外部アドレスバス452に出力され
る。データは32ビットごとに順次データバス1023
からグローバルバス1022に読み込まれる。Next, an external memory access will be described. External memory access occurs at the time of a global memory data cache miss. Address bus 1022
Is output to the external address bus 452 as it is. Data is sequentially transferred to the data bus 1023 every 32 bits.
To the global bus 1022.
【0200】最後に、割り込みについて説明する。本発
明のマイクロプロセッサ401では、高速コンテキスト
スイッチ機構の採用により、割り込み応答はスレッド生
成と等価としている。つまり、スレッド生成ユニット4
03に用意されているスレッド状態に対して、スレッド
を生起するだけで良い。割り込みユニット1005は、
割り込みベクタ生成ユニット1002に割り込みベクタ
番号を伝達する。割り込みベクタ生成ユニット1002
は、割り込みの種類に応じたプログラムカウンタ、スタ
ックポインタ、スレッド情報を生成し、スレッド制御ユ
ニット403に伝達する。Finally, the interruption will be described. In the microprocessor 401 of the present invention, an interrupt response is equivalent to thread generation by employing a high-speed context switch mechanism. That is, the thread generation unit 4
It is only necessary to create a thread for the thread state prepared in 03. The interrupt unit 1005
The interrupt vector number is transmitted to the interrupt vector generation unit 1002. Interrupt vector generation unit 1002
Generates a program counter, a stack pointer, and thread information according to the type of interrupt, and transmits the generated information to the thread control unit 403.
【0201】<マルチスレッドプログラミングの例>図
21を参照して、実際のスレッドの使用方法について説
明する。スレッドAは、スレッドBを生成して並列処理
を行い。双方の終了を確認して元のスレッドAを再開す
る。<Example of Multithread Programming> Referring to FIG. 21, an actual method of using threads will be described. Thread A generates thread B and performs parallel processing. After confirming both ends, the original thread A is resumed.
【0202】スレッドAは、共有変数に2を格納し、ス
レッドBの状態を作成するためにOSから未使用のSP
およびTIの取得を行う。The thread A stores 2 in the shared variable, and uses the unused SP from the OS to create the state of the thread B.
And TI are obtained.
【0203】双方のスレッドは、同じLabelCに分
岐し、共有変数Counterにアトミックにアクセス
する。2回目に到達した方のスレッドはCounter
を0にするため、その時点で続きのコードのLabel
Dを再開できる。そして、元のスレッドAのSP、TI
状態を受け継ぐ。同じ要領で並列動作させるスレッドは
いくらでも増加させることができる。Both threads branch to the same LabelC, and access the shared variable Counter atomically. The thread that reached the second time is Counter
To 0, at that point the Label of the following code
D can be restarted. And the SP and TI of the original thread A
Inherit the state. Any number of threads can be run in parallel in the same way.
【0204】<負荷分散の方法>図22を参照して、負
荷分散につて説明する。<Load Balancing Method> The load balancing will be described with reference to FIG.
【0205】すべてのプロセス、スレッドはOSからス
タック、ヒープ領域を要求する。処理能力の必要なプロ
セスは、1つ、あるいは複数のメモリバンクに渡るスタ
ック、ヒープ領域をOSから独占して取得することがで
きる。All processes and threads request a stack and a heap area from the OS. A process that requires processing power can exclusively acquire a stack or heap area over one or more memory banks from the OS.
【0206】この図では、プロセスAは3つ分のプロセ
ッサを、プロセッサBは2つ分のプロセッサを所持し、
他のプロセスに妨害されることはほとんどない。このよ
うにして、リアルタイム処理能力の必要なプロセスは、
負荷分散の保証を得ることができる。In this figure, process A has three processors, processor B has two processors,
There is little interruption by other processes. In this way, processes that require real-time processing power
A guarantee of load distribution can be obtained.
【0207】<実施例3>図23を参照して、本発明の
第3の実施例のマイクロプロセッサ2301の内部構造
について解説する。<Embodiment 3> The internal structure of a microprocessor 2301 according to a third embodiment of the present invention will be described with reference to FIG.
【0208】マイクロプロセッサ2301は、実施例2
のマイクロプロセッサ401に加えて、乗算器や浮動小
数点演算器を整数パイプライン上に搭載し、マルチプロ
セッサ構造を導入している。マルチプロセッサには、メ
モリ共有機構だけでなく、スレッド管理をマルチプロセ
ッサ間で行うための機構を導入している。メモリはあえ
てマイクロプロセッサに直接接続させて分散し、メモリ
との通信バンド幅を確保している。The microprocessor 2301 is used in the second embodiment.
In addition to the microprocessor 401, a multiplier and a floating-point arithmetic unit are mounted on an integer pipeline, and a multiprocessor structure is introduced. In the multiprocessor, not only a memory sharing mechanism but also a mechanism for performing thread management between the multiprocessors is introduced. The memory is deliberately connected directly to the microprocessor and distributed to ensure a communication bandwidth with the memory.
【0209】実施例3のマイクロプロセッサ2301と
実施例2のマイクロプロセッサ401との違いは、演算
ユニット群2305と、グローバルアクセス制御ユニッ
ト2308だけである。The difference between the microprocessor 2301 of the third embodiment and the microprocessor 401 of the second embodiment is only the operation unit group 2305 and the global access control unit 2308.
【0210】<演算ユニット群2305>図24を参照
して、演算ユニット群2305の内部構造について詳し
く説明する。整数演算ユニット群405に加え、乗算、
および浮動小数点機能を内蔵する。<Operation Unit Group 2305> The internal structure of operation unit group 2305 will be described in detail with reference to FIG. In addition to the integer operation unit group 405, multiplication,
And built-in floating point function.
【0211】4つの演算ユニット501と同時に、2つ
の乗算・SIMDユニット2401と、1つの浮動小数
点演算ユニット2402を内蔵する。これらは浮動小数
点レジスタ2403を共用する。そして、オペランドデ
ータクロスバスイッチ2404が命令実行タイミングを
調節する。At the same time as the four operation units 501, two multiplication / SIMD units 2401 and one floating point operation unit 2402 are incorporated. They share the floating point register 2403. Then, the operand data crossbar switch 2404 adjusts the instruction execution timing.
【0212】浮動小数点レジスタ2403は、整数演算
ユニット内の汎用レジスタと同じシフトレジスタであ
り、4クロックで伝送される。The floating-point register 2403 is the same shift register as the general-purpose register in the integer operation unit, and is transmitted at four clocks.
【0213】オペランドデータクロスバスイッチ240
4は、すべての整数演算ユニット501からの命令発行
を、浮動小数点演算ユニット2402の左端のオペラン
ド入力に接続する。そして、結果データを浮動小数点レ
ジスタ2403の4つのタイミング全てに対して転送す
る。Operand data crossbar switch 240
4 connects instruction issuance from all integer arithmetic units 501 to the leftmost operand input of floating point arithmetic unit 2402. Then, the result data is transferred to all four timings of the floating-point register 2403.
【0214】これらの乗算ユニット2401は命令とし
てインプリメントされ、コプロセッサではない。しか
も、分岐やグローバルデータロードのようにスレッドを
待ち状態にして同期を取る必要はない。よって、1つの
演算ユニット群2305の中で、4つのプロセッサが2
つの乗算機、1つの浮動小数点演算機を共有するのと同
じことになる。[0214] These multiplication units 2401 are implemented as instructions and are not coprocessors. In addition, there is no need to wait for a thread to synchronize as in the case of a branch or global data load. Therefore, in one arithmetic unit group 2305, four processors are 2
It is the same as sharing one multiplier and one floating point arithmetic unit.
【0215】この浮動小数点ユニット2402は、1つ
の演算ユニット群2305につき1つの浮動小数点命令
を実行できる。命令のアドレスは任意で良い。ただし、
4命令中に2命令以上の浮動小数点演算が現れた場合
は、パイプラインストールとなり、16クロック後に再
開される。乗算命令の場合は2命令までである。よっ
て、プログラムモデルでは浮動小数点ユニットの配置の
制限はない。単に制限を超えると実行速度が低下するだ
けである。This floating point unit 2402 can execute one floating point instruction per one operation unit group 2305. The address of the instruction may be arbitrary. However,
If a floating-point operation of two or more instructions appears in four instructions, the pipeline is stuck and restarted after 16 clocks. In the case of a multiplication instruction, the number is up to two instructions. Therefore, there is no restriction on the arrangement of floating point units in the program model. Simply exceeding the limit only slows down execution.
【0216】浮動小数点ユニット2402は、除算命令
などの可変レイテンシ命令にも対応できる。ただし、命
令を実行したスレッドは除算命令の終了まで待ち状態と
なる。グローバルメモリへのロード命令と同じ動作であ
る。[0216] The floating point unit 2402 can cope with variable latency instructions such as a division instruction. However, the thread that has executed the instruction is in a waiting state until the end of the division instruction. This is the same operation as the load instruction to the global memory.
【0217】<グローバルアクセスユニット2308>
図25を参照して、グローバルアクセスユニット230
8の内部構造について説明する。実施例2のグローバル
アクセスユニット408との違いは、外部ローカルバス
と外部共有バスを分離したこと、外部共有バスからのス
レッド生成を可能にすることが挙げられる。よって、違
いは共有バスインターフェース2501と、ローカルバ
スインターフェース2502だけである。<Global access unit 2308>
Referring to FIG. 25, global access unit 230
8 will be described. The difference from the global access unit 408 of the second embodiment is that the external local bus and the external shared bus are separated, and that a thread can be generated from the external shared bus. Therefore, the only difference is the shared bus interface 2501 and the local bus interface 2502.
【0218】外部共有バスインターフェース2501
は、外部のデバイスからデータを読み込むと同時に、外
部のマイクロプロセッサ2301から自身のローカルバ
スをアクセスさせる機能を持つ。External shared bus interface 2501
Has a function of reading data from an external device and allowing the external microprocessor 2301 to access its own local bus.
【0219】外部からのメモリアクセスの場合、ローカ
ルメモリアクセス制御ユニット2503は、内部アドレ
スバス1022、内部データバス1023のアービトレ
ーションを取得し、グローバルキャッシュメモリ100
8、ローカルバスインターフェース2502を制御す
る。In the case of external memory access, the local memory access control unit 2503 obtains the arbitration of the internal address bus 1022 and the internal data bus 1023, and
8. Control the local bus interface 2502.
【0220】本発明のマイクロプロセッサ2301は、
マルチプロセッサ間のキャッシュのコヒーレンシ維持の
方法も多少異なる。ローカルメモリキャッシュ407同
士は同じアドレスのデータを所持することはありえない
ので調停の必要はない。しかし、グローバルキャッシュ
メモリは、チップ外部からのアクセスを高速化する目的
のため、他のプロセッサのローカルキャッシュメモリ4
07上に実体が存在しうる。この場合、グローバルキャ
ッシュ1007の間でキャッシュコヒーレンシ制御を行
う必要が発生する。[0220] The microprocessor 2301 of the present invention comprises:
The method of maintaining cache coherency between multiprocessors is somewhat different. Since the local memory caches 407 cannot have data of the same address, there is no need for arbitration. However, for the purpose of accelerating the access from outside the chip, the global cache memory has a local cache memory 4 of another processor.
07 may exist. In this case, it becomes necessary to perform cache coherency control between the global caches 1007.
【0221】<マルチプロセッサシステム>図26を参
照して、本発明の第3の実施例のマイクロプロセッサ2
301を使用したシステムの例を挙げる。<Multiprocessor System> Referring to FIG. 26, a microprocessor 2 according to the third embodiment of the present invention will be described.
An example of a system using 301 will be described.
【0222】メモリ2611は、ローカルバス2612
を通して、マイクロプロセッサ2301に直接接続され
る。共有バス2613は、マイクロプロセッサ2301
同士の通信に使用され、共有I/Oデバイス2614を
接続する。The memory 2611 has a local bus 2612
Through to the microprocessor 2301 directly. The shared bus 2613 is connected to the microprocessor 2301
It is used for communication between each other and connects a shared I / O device 2614.
【0223】ローカルI/Oデバイス2615は、グラ
フィックアクセラレータ等のメモリバンド幅を要求する
デバイスである。マイクロプロセッサ2301のローカ
ルバス2612に直接接続されて、マイクロプロセッサ
2301との通信バンド幅を確保している。The local I / O device 2615 is a device requiring a memory bandwidth such as a graphic accelerator. It is directly connected to the local bus 2612 of the microprocessor 2301 to secure a communication bandwidth with the microprocessor 2301.
【0224】本発明のマイクロプロセッサのスレッド通
信機構のハードウェア化により、メモリを分散させるこ
とができる。自身のマイクロプロセッサが持たないメモ
リ領域へのアクセスは、実際にマイクロプロセッサ間で
転送要求を出す方法もある。しかし、直接データを転送
しなくても、そのメモリを持つマイクロプロセッサ23
01にデータ処理を行うスレッドを生成させることがで
きる。この場合、共有バス2613のバンド幅を最小に
できる。The memory can be distributed by hardware of the thread communication mechanism of the microprocessor of the present invention. For access to a memory area not owned by the microprocessor itself, there is a method of actually issuing a transfer request between the microprocessors. However, even if the data is not directly transferred, the microprocessor 23 having the memory
01 can generate a thread for performing data processing. In this case, the bandwidth of the shared bus 2613 can be minimized.
【0225】<マルチプロセッサ版メモリマップ>図2
7を参照して、本発明のマイクロプロセッサ2301を
使用したマルチプロセッサシステムのメモリ空間につい
て説明する。<Multiprocessor Memory Map> FIG.
7, the memory space of a multiprocessor system using the microprocessor 2301 of the present invention will be described.
【0226】各マイクロプロセッサ2301のメモリ空
間は、システムのメモリ空間2702に静的に割付られ
る。従来例2のマイクロプロセッサ301と異なり、共
有バスに出力したトランザクションの受理先のマイクロ
プロセッサはアドレスに対して一意に決定する。The memory space of each microprocessor 2301 is statically allocated to the memory space 2702 of the system. Unlike the microprocessor 301 of the conventional example 2, the microprocessor which receives the transaction output to the shared bus is uniquely determined with respect to the address.
【0227】通信帯域幅が必要なオブジェクト2711
と、オブジェクト2712は、同一のマイクロプロセッ
サ2703に配置される。オブジェクト2713はオブ
ジェクト2711との通信帯域幅を必要としないため、
マイクロプロセッサ2704に配置できる。このよう
に、メモリの割付によってオブジェクト間の通信バンド
幅を最適化できる。Object 2711 Requiring Communication Bandwidth
And the object 2712 are arranged in the same microprocessor 2703. Since the object 2713 does not require a communication bandwidth with the object 2711,
It can be located in the microprocessor 2704. Thus, the communication bandwidth between objects can be optimized by allocating the memory.
【0228】<マルチプロセッサ共有バストランザクシ
ョン>図28を参照して、マルチプロセッサ間のバスト
ランザクションの内容を示す。<Multiprocessor Shared Bus Transaction> Referring to FIG. 28, the contents of a bus transaction between multiprocessors will be described.
【0229】通常のマイクロプロセッサが持つシングル
データアクセス、バーストデータアクセスに加え、共有
メモリに必要なロック付きシングルアクセス、そしてマ
イクロプロセッサ間スレッド生成要求コマンドが存在す
る。逆に、従来例2のマルチプロセッサに存在するキャ
ッシュコヒーレンシ制御コマンドは必要ない。In addition to a single data access and a burst data access of a normal microprocessor, there are a single access with a lock required for a shared memory and a thread generation request command between microprocessors. Conversely, the cache coherency control command existing in the multiprocessor of the second conventional example is not required.
【0230】<ループアンローリングの効果>図29を
参照して、ループアンローリングの効果について説明す
る。<Effect of Loop Unrolling> The effect of loop unrolling will be described with reference to FIG.
【0231】本発明のマイクロプロセッサ401は、命
令の下位4ビットに対応した演算ユニットで常に実行さ
れる。ということは、利用率の高い命令の下位4ビット
が偏っていては、使用されない演算ユニットが存在する
ことになる。この問題は、一般的な高速化テクニックで
あるループアンローリングによって解消できる。The microprocessor 401 of the present invention is always executed by an arithmetic unit corresponding to the lower 4 bits of an instruction. That is, if the lower 4 bits of the instruction having a high utilization rate are biased, there will be an unused operation unit. This problem can be solved by a general speed-up technique, loop unrolling.
【0232】2901は、1つのスレッドのプログラム
である。2902は16の演算機を示す。2901のコ
ードは、6命令のみを使用しており、最初の6つの演算
機だけを使用することになる。つまり、残りのパイプラ
イン2904はそっくり他のスレッドに渡すことにな
る。ところが、それをうめるだけの他のスレッドの要求
がない場合は性能を発揮できないことになる。Reference numeral 2901 denotes a program of one thread. Reference numeral 2902 denotes 16 arithmetic units. The code of 2901 uses only six instructions, and uses only the first six arithmetic units. In other words, the remaining pipeline 2904 is completely passed to another thread. However, if there is no request from other threads to make up for it, performance cannot be exhibited.
【0233】そのため、演算機を全て使用するように3
つのスレッドを1つにまとめ、1つのスレッドを14命
令に拡張する。2911がインライン化された1つのス
レッドのプログラムである。これによって、演算機29
12は14個使用されることになる。[0233] Therefore, 3
Combine one thread into one and extend one thread to 14 instructions. Reference numeral 2911 denotes an inline one thread program. Thereby, the arithmetic unit 29
12 will be used 14 pieces.
【0234】また、一般的な長いスレッドについても、
16の倍数の長さの命令長に調節すれば、最も性能が発
揮できる。Also, for a general long thread,
Adjusting to an instruction length that is a multiple of 16 gives the best performance.
【0235】[0235]
【発明の効果】本発明のマイクロプロセッサの効果を、
クロック当たりの論理性能、周波数性能、プログラミン
グモデル、チップ面積、低消費電力の項目に分けて説明
する。The effect of the microprocessor of the present invention is as follows.
The logic performance per clock, frequency performance, programming model, chip area, and low power consumption will be described separately.
【0236】<クロック当たりの論理性能><Logical performance per clock>
【0237】本発明のマイクロプロセッサは、演算性能
において以下の長所を持つ。 1.マルチスレッド動作による並列処理が可能である。 2.マルチスレッドによって外部メモリへのアクセスレ
イテンシの隠蔽が可能である。 3.スレッドの命令およびデータの共有が容易に実現で
き、小規模イタレーション型の並列にも有効である。The microprocessor of the present invention has the following advantages in the operation performance. 1. Parallel processing by multi-thread operation is possible. 2. It is possible to hide the access latency to the external memory by multithreading. 3. Thread instructions and data can be easily shared, and it is also effective for small-scale iteration-type parallelism.
【0238】本発明のマイクロプロセッサは、単体のス
レッドの速度では分岐のレイテンシが巨大であるため従
来のパイプラインマイクロプロセッサより劣る。しか
し、それはマルチスレッドと並列処理によって補って余
りある。The microprocessor of the present invention is inferior to the conventional pipeline microprocessor because the branch latency is large at the speed of a single thread. However, it is more than compensated for by multithreading and parallel processing.
【0239】以下、汎用アプリケーションにおける性能
の定量的な予測を示す。Computer Archi
techture: A Quantitive Ap
proach Second Edition(著者
John H.Hennessy AND David
A.Patterson 出版社 MorganKa
ufmann Pubishers,Inc.)に記載
されている、統計データを使用する。以下この文献を参
考文献1とする。Hereinafter, quantitative prediction of performance in general-purpose applications will be described. Computer Archi
technology: A Quantitative Ap
Provision Second Edition (Author)
John H. Hennessy AND David
A. Patterson Publisher MorganKa
ufmann Pubishers, Inc. Use the statistical data described in). Hereinafter, this document is referred to as Reference Document 1.
【0240】単体のスレッドからみれば、本発明のマイ
クロプロセッサは、以下の条件のマイクロプロセッサと
ほぼ等価である。 ・汎用レジスタ5 ロードストアレジスタモデル ・16Kバイトダイレクトマップ命令キャッシュ ・1Kバイトダイレクトマップデータキャッシュ ・分岐レイテンシが不定 簡単のため、単体のスレッドはローカルメモリバンクの
みを使用するとする。From the standpoint of a single thread, the microprocessor of the present invention is substantially equivalent to a microprocessor under the following conditions. -General-purpose register 5 load store register model-16K byte direct map instruction cache-1K byte direct map data cache-Branch latency is undefined For simplicity, it is assumed that a single thread uses only a local memory bank.
【0241】本発明のマイクロプロセッサは、十分なス
レッドが共有されるという前提であれば、全体の性能
は、ロードストア、分岐、キャッシュミスに起因するコ
ンテキストスイッチのオーバーヘッドを差し引いたもの
になる。レジスタ依存関係によるインタロックはコンパ
イラのスケジューリングで除去されているものとする。
参考文献1のp105、p384の統計データによる
と、 ・ロードストア命令の発生頻度:全命令中35% ・分岐命令の発生頻度:全命令中20% ・ロードストアのうちデータキャッシュのミスの確率:
24.61% ・16Kバイト命令キャッシュのキャッシュミスの確
率:0.64%In the microprocessor of the present invention, assuming that sufficient threads are shared, the overall performance is obtained by subtracting the overhead of the context switch caused by load store, branch, and cache miss. It is assumed that the interlock due to the register dependency has been removed by the scheduling of the compiler.
According to the statistical data of p105 and p384 in Reference 1, the frequency of load store instructions: 35% of all instructions, the frequency of branch instructions: 20% of all instructions, and the probability of a data cache miss in load stores:
24.61% ・ Probability of cache miss of 16K byte instruction cache: 0.64%
【0242】以下は、コンテキストスイッチに要するク
ロック数の予測値である。コンテキストスイッチ先のス
レッドの命令アドレスは均等に配分されているものと仮
定する。そして、データキャッシュリプレースのペナル
ティーは別のスレッドの動作で隠蔽できるものと仮定す
る。 ・分岐命令:平均7.5クロック、20%の頻度で発生 ・パイプラインストール:平均9.5クロック、9.2
5%の頻度で発生The following is a predicted value of the number of clocks required for a context switch. It is assumed that the instruction addresses of the thread at the context switch destination are evenly distributed. It is assumed that the penalty for data cache replacement can be hidden by the operation of another thread. Branch instructions: 7.5 clocks on average, occurring at a frequency of 20% Pipelining: 9.5 clocks on average, 9.2
Occurs at a frequency of 5%
【0243】以上の条件で、1つのパイプラインのIP
C性能(クロック単位の命令実行数)は約0.3とな
る。本発明のマイクロプロセッサ全体では、16のパイ
プラインがほとんど互いに干渉しないので、、この数値
を16倍した4.8弱が全体のIPC性能になる。Under the above conditions, one pipeline IP
The C performance (the number of executed instructions in clock units) is about 0.3. In the overall microprocessor of the present invention, since the 16 pipelines hardly interfere with each other, a value less than 4.8, which is 16 times this value, is the overall IPC performance.
【0244】このように、本発明のマイクロプロセッサ
では、分岐命令のペナルティーが非常に大きく性能の妨
げとなる。ループアンローリングの手法で分岐を削減す
れば、更に大きく性能を向上させることができる。As described above, in the microprocessor of the present invention, the penalty of the branch instruction is very large, which hinders the performance. If the number of branches is reduced by the loop unrolling technique, the performance can be further improved.
【0245】以下、ループアンローリングによって分岐
命令の頻度を64命令に1回に削減したと仮定した場合
の性能を予測する。この場合、IPCは約0.5とな
る。全体のIPCは8近くに向上する。In the following, the performance will be predicted on the assumption that the frequency of branch instructions is reduced to once every 64 instructions by loop unrolling. In this case, the IPC is about 0.5. The overall IPC improves to near 8.
【0246】さらに、演算ユニットを増加した場合も、
データキャッシュのリプレースの転送能力を増強すれ
ば、単体のパイプラインの性能に干渉することはほとん
どない。よって、投入したハードウェア資源に対する性
能の線形増加を実現できる。Further, when the number of arithmetic units is increased,
Increasing the transfer capacity of the replacement of the data cache hardly interferes with the performance of a single pipeline. Therefore, a linear increase in performance with respect to the input hardware resources can be realized.
【0247】<周波数性能>本発明のマイクロプロセッ
サは、演算器を大量に集積しているが、命令単位のパイ
プラインの構造は通常のマイクロプロセッサとほぼ同等
である。並列数の増加によるセレクタ入力やレジスタポ
ートの増大はない。<Frequency Performance> Although the microprocessor of the present invention has a large number of arithmetic units integrated therein, the structure of a pipeline for each instruction is almost the same as that of a normal microprocessor. There is no increase in selector inputs or register ports due to an increase in the number of parallel operations.
【0248】VLIW型マイクロプロセッサは、演算器
の増加に対して、互いの結果データ、オペランド、レジ
スタファイル間の転送が複雑化して速度を抑制する。そ
れに対し本発明のマイクロプロセッサでは、演算ユニッ
ト間のデータ送信を単一方向に制限することにより、演
算器間のオペランドの転送も単純な回路で実現でき、回
路段数が削減できる。In the VLIW type microprocessor, the transfer between the result data, operands, and the register file becomes complicated as the number of arithmetic units increases, and the speed is suppressed. On the other hand, in the microprocessor of the present invention, by limiting data transmission between the operation units in a single direction, the transfer of operands between the operation units can be realized by a simple circuit, and the number of circuit stages can be reduced.
【0249】また、演算ユニットが隣接していることに
より、演算器間の配線が最短距離である。演算器とロー
カルデータキャッシュ間の転送が唯一の長距離配線とな
る。Further, since the arithmetic units are adjacent to each other, the wiring between the arithmetic units is the shortest distance. The transfer between the arithmetic unit and the local data cache is the only long-distance wiring.
【0250】以上の効果により、従来のパイプライン方
式、VLIW方式等のマイクロプロセッサと同等、ある
いはそれ以上の周波数性能を出すことができる。特に、
並列性能の向上と、周波数性能の両立が容易な点で優れ
ている。With the above effects, frequency performance equal to or higher than that of a conventional pipeline system, VLIW system, or other microprocessor can be obtained. Especially,
It is excellent in that it is easy to achieve both improved parallel performance and frequency performance.
【0251】<プログラミングモデル>本発明のマイク
ロプロセッサは、命令セットの用法は一般的なRISC
マイクロプロセッサとほとんど同じであり、さらに一般
的なマルチスレッドの概念でプログラムを作成できる。<Programming Model> The microprocessor of the present invention employs a general instruction set
It is almost the same as a microprocessor, and can create programs using the general concept of multithreading.
【0252】本発明の既存のマイクロプロセッサに対す
るプログラミングモデルにおける長所は、主に、並列度
向上に伴う垂直性、マルチスレッドサポート、単純な負
荷分散、マルチプロセッサ垂直性である。それぞれにつ
いて詳細に説明する。The advantages of the programming model for the existing microprocessor of the present invention are mainly verticality with increased parallelism, multi-thread support, simple load balancing, and multi-processor verticality. Each will be described in detail.
【0253】まず、並列度向上に伴う垂直性について述
べる。従来例1のVLIWマイクロプロセッサでは、同
時に実行する処理をその都度コンパイラなどが適切に配
置する必要があるが、本発明のマイクロプロセッサは、
プログラムを改変することなく全く独立した処理を同時
に大量に実行できる。更に、演算器405の数を更に増
加した場合も、最初に作成したマルチスレッドプログラ
ムをそのまま使用して性能を向上することができる。First, the verticality associated with the improvement in the degree of parallelism will be described. In the VLIW microprocessor of the first conventional example, it is necessary for a compiler or the like to appropriately arrange processes to be executed simultaneously each time.
A large number of completely independent processes can be executed simultaneously without modifying the program. Further, even when the number of the arithmetic units 405 is further increased, the performance can be improved by using the multi-thread program created first as it is.
【0254】次に、マルチスレッドサポートについて述
べる。既存のマイクロプロセッサでのマルチスレッドプ
ログラムの同期には、スピンロックによるスレッド同期
待ちか、明示的なOSのスケジューラの呼び出しのどち
らかが必要になる。本発明のマイクロプロセッサではそ
のどちらも必要なく、スレッド生成、消滅命令や自動コ
ンテキストスイッチ機構により、OSを介在しない高速
なコンテキストスイッチを可能にしている。これによ
り、記述が自然で、かつ高速なスレッド間同期を実現で
きる。Next, multithread support will be described. Synchronization of a multi-thread program in an existing microprocessor requires either thread synchronization waiting by spin lock or explicit call of a scheduler of the OS. The microprocessor of the present invention does not require either of them, and enables high-speed context switching without the intervention of an OS by using a thread generation / deletion instruction and an automatic context switching mechanism. This makes it possible to realize high-speed synchronization between threads with a natural description.
【0255】次に、負荷分散について述べる。オブジェ
クトのスタックやヒープメモリ配置がそのまま負荷分散
となる機構である。そのため、複雑なプロセス間プライ
オリティー制御などのOS機能を使用する必要がなく、
マイクロプロセッサの演算資源をスレッドに対して一定
に分配できる。Next, load distribution will be described. This is a mechanism in which the stack of objects and the memory arrangement of the heap distribute the load as it is. Therefore, there is no need to use complicated OS functions such as priority control between processes.
The computing resources of the microprocessor can be uniformly distributed to the threads.
【0256】最後に、マルチプロセッサ負荷分散につい
て述べる。本発明のマイクロプロセッサを複数使用する
場合は、単一の本発明のマイクロプロセッサ上で開発し
たマルチスレッドプログラムがそのまま使用できる。O
Sのヒープメモリ割り当て機構だけがマルチプロセッサ
の情報を管理して負荷分散すれば良い。Finally, multiprocessor load distribution will be described. When a plurality of microprocessors of the present invention are used, a multi-thread program developed on a single microprocessor of the present invention can be used as it is. O
Only the heap memory allocation mechanism of S needs to manage the information of the multiprocessor and distribute the load.
【0257】<チップ面積>本発明のマイクロプロセッ
サのチップ面積縮小の効果を示すため、Compute
r Architechture Pipelined
And Parallel Processor D
esign(著者 Michael J.Flynn)
のp96の記載に基づいたモデルで示す。<Chip Area> To show the effect of reducing the chip area of the microprocessor of the present invention, Compute
r Architecture Pipelined
And Parallel Processor D
design (author Michael J. Flynn)
The model is based on the description of p96.
【0258】本発明の第2の実施例において、各ユニッ
トの面積予測値を以下の通りとする。単位Aは1ミクロ
ンプロセスにおける1mm平方に相当する。 ・整数演算ユニット405 7A×16 ・データキャッシュ407 4A×16 ・命令メモリ404 13.4A×4 ・スレッド発行ユニット402 約4A×4 ・グローバルデータキャッシュ8K 26.6A ・共有浮動小数点ユニット409 37.8A この条件において、整数演算性能は最大16並列とな
る。In the second embodiment of the present invention, the predicted area value of each unit is as follows. Unit A corresponds to 1 mm square in a 1 micron process. -Integer operation unit 405 7A x 16-Data cache 407 4A x 16-Instruction memory 404 13.4A x 4-Thread issuing unit 402 Approx. 4A x 4-Global data cache 8K 26.6A-Shared floating point unit 409 37.8A Under this condition, the integer operation performance is 16 parallel at maximum.
【0259】メモリに必要な制御回路を30%と仮定
し、配線に伴うオーバーヘッドを、ロジック部で50
%、メモリ部で15%とする。この場合、ロジック部は
248.7A、メモリ部は216Aになり、最終的なプ
ロセッサ領域は464.7Aとなる。Assuming that the control circuit required for the memory is 30%, the overhead associated with wiring is reduced by 50% in the logic section.
% And 15% in the memory section. In this case, the logic section becomes 248.7A, the memory section becomes 216A, and the final processor area becomes 464.7A.
【0260】実際の半導体のダイに実装するオーバーヘ
ッドを20%としてチップ面積は557Aとなり、0.
25ミクロンにおけるチップ面積は34mm角となる。Assuming that the overhead to be mounted on an actual semiconductor die is 20%, the chip area is 557A.
The chip area at 25 microns is 34 mm square.
【0261】本発明の第3の実施例の条件を示す。以下
のユニットが第2の実施例に付加される。演算性能は、
整数演算性能が8ビット単位で16並列となり、浮動小
数点が4並列となる。 ・16ビットSIMD整数乗算器2401 20A×8 ・倍精度浮動小数点加算乗算器2402 37.8A×4The conditions of the third embodiment of the present invention will be described. The following units are added to the second embodiment. The calculation performance is
The integer operation performance becomes 16 parallel in 8-bit units, and the floating point becomes 4 parallel. 16-bit SIMD integer multiplier 2401 20A × 8 Double-precision floating-point addition multiplier 2402 37.8A × 4
【0262】同じ前提条件において、ロジック部は65
8Aとなり、最終的なプロセッサ領域は878Aとな
る。チップ面積は1050Aで、0.25ミクロンにお
けるチップ面積は65mm角となる。Under the same preconditions, the logic section is 65
8A, and the final processor area is 878A. The chip area is 1050A, and the chip area at 0.25 micron is 65 mm square.
【0263】既存のマイクロプロセッサと比較のため
に、米Intel社のMMXPentiumマイクロプ
ロセッサを引用する。このマイクロプロセッサの0.2
5μm版のチップサイズが95mm角である。このマイ
クロプロセッサの整数並列度は最大で2、浮動小数点が
1、バイト単位のSIMD命令を使用しても最大16で
ある。For comparison with the existing microprocessor, the MMXPentium microprocessor from Intel Corporation of the United States is cited. 0.2 of this microprocessor
The chip size of the 5 μm plate is 95 mm square. The microprocessor has a maximum degree of integer parallelism of 2, a floating point of 1, and a maximum of 16 even when using SIMD instructions in byte units.
【0264】さらに、本発明のマイクロプロセッサは、
演算機などの並列数の自然な拡張が可能である。それは
ハードウェアの構成と、ソフトウェアの互換性の双方の
理由である。Furthermore, the microprocessor of the present invention
Natural expansion of the number of parallel units such as arithmetic units is possible. That is the reason for both the hardware configuration and the software compatibility.
【0265】まず、演算ユニットのn倍の増加に対し
て、データキャッシュメモリはO(n)、そして命令メ
モリの増加はO(n)以下である。First, for an n-fold increase in the number of arithmetic units, the data cache memory is O (n), and the increase in the instruction memory is O (n) or less.
【0266】そして、演算ユニット間のバス配線の増加
をO(n)に押さえることができる。既存のスーパース
カラ方式、VLIW方式はすべての演算器の間の自由な
転送を保証するため、O(n×n)であるのと対照的で
ある。The increase in the number of bus lines between the operation units can be reduced to O (n). The existing super scalar method, VLIW method, is in contrast to O (n × n) in order to guarantee free transfer between all the arithmetic units.
【0267】本発明のマイクロプロセッサで線形増加以
上の回路規模増加になるのは、分岐クロスバスイッチ4
02とデータクロスバスイッチ406である。そのまま
ではO(n×n)のオーダーの増加になることは避けら
れない。だが、データクロスバスイッチ406について
は、転送先のメモリバンクの配置順序が固定であるた
め、ローカルメモリアクセスレイテンシの低下を容認す
るならバレルシフタと同じ方式を採用できる。この場
合、回路増加のオーダーはO(n×logn)にでき
る。In the microprocessor according to the present invention, the increase in the circuit size beyond the linear increase is caused by the branch crossbar switch 4.
02 and the data crossbar switch 406. It is unavoidable that the increase will be on the order of O (n × n) as it is. However, regarding the data crossbar switch 406, since the arrangement order of the memory banks at the transfer destination is fixed, the same system as the barrel shifter can be adopted if the reduction in local memory access latency is tolerated. In this case, the order of circuit increase can be O (n × logn).
【0268】<低消費電力>低消費電力のための技術は
大量に存在するが、アーキテクチャレベルの低消費電力
化の手段は、性能に対する回路や配線を最小限にするこ
とで実現できる。本発明のマイクロプロセッサは、性能
に対する回路および配線を最小限にできる。具体的に
は、従来のマイクロプロセッサと比較して以下の長所が
ある。<Low Power Consumption> There are a large number of technologies for low power consumption, but means for reducing power consumption at the architectural level can be realized by minimizing circuits and wiring for performance. The microprocessor of the present invention minimizes circuitry and wiring for performance. Specifically, there are the following advantages as compared with the conventional microprocessor.
【0269】スーパースカラプロセッサと比較して、複
雑な結果の転送を必要とする命令レベル並列を行わない
ことにより、データの転送の自由度を最小限にしたこと
による、回路や配線の削減を可能にした。Compared with a superscalar processor, by eliminating instruction-level parallelism that requires the transfer of complicated results, the circuit and wiring can be reduced by minimizing the degree of freedom in data transfer. I made it.
【0270】マルチプロセッサでありながら、演算の結
果の伝送の距離が最小限であることが言える。クリティ
カルになる長距離のバス配線はローカルメモリキャッシ
ュへの信号のみである。よって、チップ全体の配線容量
を最小限にできる。It can be said that the distance of the transmission of the result of the operation is minimum even though it is a multiprocessor. The only long-distance bus lines that become critical are signals to the local memory cache. Therefore, the wiring capacity of the entire chip can be minimized.
【0271】演算器間で命令キャッシュを共有すること
による命令メモリ容量、リプレース頻度の削減が可能で
ある。命令メモリは定型処理を行う上で同じコピーを持
つ可能性が多いので、命令キャッシュの共有により命令
メモリの容量を削減できる。同じ理由により命令キャッ
シュのリプレース頻度を下げることができる。It is possible to reduce the instruction memory capacity and the replacement frequency by sharing the instruction cache between the arithmetic units. Since the instruction memory is likely to have the same copy in performing the routine processing, the capacity of the instruction memory can be reduced by sharing the instruction cache. For the same reason, the replacement frequency of the instruction cache can be reduced.
【0272】演算器間でデータメモリを共有することに
よるデータメモリの省略も挙げられる。通常のマルチプ
ロセッサでは、データの共有を行う場合でも、それぞれ
がデータキャッシュを所有する必要があるが、本発明の
マイクロプロセッサではデータは全て1つのデータキャ
ッシュに格納され、複数のデータキャッシュが同じ内容
を格納することはない。It is also possible to omit the data memory by sharing the data memory between the arithmetic units. In a general multiprocessor, even when sharing data, each of them must own a data cache. However, in the microprocessor of the present invention, all data is stored in one data cache, and a plurality of data caches have the same contents. Is not stored.
【0273】これらの効果により、性能に対する消費電
力を最小限にできる。With these effects, power consumption for performance can be minimized.
【0274】<犠牲にしたもの>本発明のマイクロプロ
セッサは、ほとんどの面で従来のマイクロプロセッサを
凌駕する性能を持つが、同時に従来のマイクロプロセッ
サにはなかった短所もいくつか存在する。<Sacrifices> Although the microprocessor of the present invention has performance that surpasses conventional microprocessors in most respects, it also has some disadvantages not found in conventional microprocessors.
【0275】まず、互換性がないことが挙げられる。マ
ルチスレッドプログラムを前提としているため、当然で
ある。しかし、オブジェクト指向への最適化とマルチス
レッド機能を除けば、プログラミングモデルは従来から
あるRISC命令に近く、ソフトウェアの移植は容易で
ある。First, there is no compatibility. Naturally, a multi-thread program is assumed. However, except for the object-oriented optimization and the multi-thread function, the programming model is close to the conventional RISC instruction, and software porting is easy.
【0276】次に、単体のスレッドの性能は従来のマイ
クロプロセッサより低いことが挙げられる。理由は、分
岐レイテンシが巨大であること、バンク外のメモリアク
セスのレイテンシも比較的大きい為である。しかし、ど
ちらのレイテンシ時間でも別のスレッドが動作できるた
め、全体性能としてはある程度隠蔽可能である。Second, the performance of a single thread is lower than that of a conventional microprocessor. The reason is that the branch latency is huge and the latency of memory access outside the bank is relatively large. However, since another thread can operate at either latency time, the overall performance can be concealed to some extent.
【0277】次に、データアクセスのレイテンシが大き
いことが挙げられる。オブジェクト指向プログラミング
モデルなどで、データアクセスの範囲、配置を常にロー
カルメモリに配置する努力が必要になる。しかし、オブ
ジェクト指向プログラミングモデルは現在も広く使用さ
れるソフトウェア作成手法であり、それに適合する形で
明示的に指定して性能を向上できるのは有用である。Second, the latency of data access is large. In an object-oriented programming model or the like, it is necessary to always allocate data access ranges and locations to local memory. However, the object-oriented programming model is still a widely used technique for creating software, and it is useful to be able to improve performance by explicitly specifying it in a form that is compatible with it.
【0278】最後に、従来のスーパースカラ方式等と比
較して演算器の使用効率が低いことが挙げられる。本発
明のマイクロプロセッサは、メモリバンクと命令のアド
レス配置が一致するまでスレッドが再開できないため、
適切なスレッドが開始できずに全く演算器が動作できな
い状況が発生する。しかし、分岐とメモリバンク外のメ
モリアクセスを行わなければその問題は発生しない。よ
って、性能を出すべきアプリケーションでは、ループア
ンローリングなどの手段でチューニングを行えば良い。
全体の並列度を圧倒的に高くできるため、チューニング
による性能向上も大きい。Lastly, the operation efficiency of the arithmetic unit is lower than that of the conventional super scalar system or the like. In the microprocessor of the present invention, since the thread cannot be restarted until the address arrangement of the memory bank and the instruction match,
A situation occurs in which the arithmetic unit cannot operate at all because the appropriate thread cannot be started. However, the problem does not occur unless the branch and the memory access outside the memory bank are performed. Therefore, in an application that should provide high performance, tuning may be performed by means such as loop unrolling.
Since the overall degree of parallelism can be greatly increased, the performance improvement by tuning is large.
【0279】<さらに性能を向上させるために>汎用レ
ジスタの増加はローカルメモリへのバンド幅の削減に貢
献する。しかし、コンテキストスイッチの退避状態の増
加をもたらすため、トレードオフで決定する必要があ
る。レジスタの退避を行わず、レジスタウィンドウによ
って切り替えるなどの手段も考えられるが、当然規模の
増大をもたらす。<To further improve the performance> The increase in the number of general-purpose registers contributes to the reduction of the bandwidth to the local memory. However, it is necessary to make a trade-off decision to increase the evacuation state of the context switch. Means such as switching by a register window without saving the register can also be considered, but this naturally increases the scale.
【0280】データキャッシュのリプレースのバンド幅
向上の効果は大きい。現在のところ、データキャッシュ
のリプレースには、1つの32ビットグローバルデータ
バスを使用しているため、リプレースのペナルティーが
大きいことは自明である。リプレース用に128ビット
バスなどを採用する、複数のデータキャッシュを並列に
リプレースを行う、などの処置が望ましい。The effect of improving the bandwidth of the replacement of the data cache is great. At present, since one 32-bit global data bus is used to replace the data cache, it is obvious that the replacement penalty is large. It is desirable to adopt a 128-bit bus or the like for replacement, or to replace a plurality of data caches in parallel.
【0281】分岐予測機構はコンテキストスイッチの間
のパイプラインの空きの削減に役立つ。しかし、レイテ
ンシ隠蔽機能を持つ本発明のマイクロプロセッサでは、
分岐予測機構の効力はそれほど絶対的ではない。トラン
ジスタを演算機の並列数の増大に使用するか、分岐予測
機構を強化するかどうかは、要求性能に対するトレード
オフで決定することになる。The branch prediction mechanism helps reduce pipeline vacancies during context switches. However, in the microprocessor of the present invention having the latency hiding function,
The effectiveness of the branch prediction mechanism is not so absolute. Whether to use transistors for increasing the number of parallel processing units or to enhance the branch prediction mechanism is determined by a trade-off with respect to required performance.
【0282】<従来方式との比較のまとめ>最後に、本
発明のマイクロプロセッサと、従来のマイクロプロセッ
サとの比較結果をまとめる。<Summary of Comparison with Conventional Method> Finally, the results of comparison between the microprocessor of the present invention and the conventional microprocessor will be summarized.
【0283】スーパースカラ方式に対して低消費電力、
並列性の限界がないという長所がある。逆に、互換性が
ないという短所がある。Low power consumption compared to superscalar system
There is an advantage that there is no limit on parallelism. On the contrary, there is a disadvantage that they are not compatible.
【0284】従来例1のVLIW方式に対して、並列性
の限界がない、演算機の増加に対して、あるいはマルチ
プロセッサ構成でも全て同じプログラムで性能を出すこ
とができる。という長所がある。短所としては、スレッ
ド発行機構が複雑であるということが言える。Compared to the VLIW system of the first conventional example, the same program can be used to achieve the performance with no parallelism limit, an increase in the number of arithmetic units, or even in a multiprocessor configuration. There is an advantage. The disadvantage is that the thread issuing mechanism is complicated.
【0285】従来例2の共有メモリマルチプロセッサに
対して、性能に対する回路規模が小さい、メモリの分割
による性能の向上が容易であるという長所がある。逆に
短所としては、既存のマルチプロセッサ対応OSが使用
できないことが言える。Compared to the shared memory multiprocessor of Conventional Example 2, there are advantages that the circuit scale for the performance is small and that the performance can be easily improved by dividing the memory. On the contrary, it can be said that the existing multiprocessor compatible OS cannot be used.
図1 本発明の第1の実施例の図 図2 従来例1のVLIWプロセッサの図 図3 従来例2のマルチプロセッサシステムの図 図4 本発明の第2の実施例の図 図5 図4の整数演算ユニット群に関する図 図6 図5の1つの整数演算ユニットの詳細な構成図 図7 図5の1つの分岐ユニットの詳細な構成図 図8 図4の命令発行機構の構成図 図9 図4のスレッド発行ユニットの構成図 図10 図4のグローバルアクセスユニットの構成図 図11 図4のデータキャッシュメモリの構成図 図12 図4の共有浮動小数点ユニットの構成図 図13 命令セット表 図14 スレッド生成の図 図15 レジスタセット表 図16 メモリ構成図 図17 単体の命令のパイプライン実行の図 図18 チップ全体のパイプライン動作の概念図 図19 通常分岐におけるスレッド切り替え動作の図 図20 パイプラインストールの動作の図 図21 スレッドを使用したプログラム例 図22 負荷分散の方法 図23 第3の実施例の全体図 図24 図23の演算ユニット群の詳細な図 図25 図23のグローバルアクセスユニットの図 図26 第3の実施例のマイクロプロセッサを使用した
システム構成例 図27 マルチプロセッサメモリ配置図 図28 共有バスのバストランザクション図 図29 ループインライン展開の効用の図1 is a diagram of a first embodiment of the present invention. FIG. 2 is a diagram of a VLIW processor of a first conventional example. FIG. 3 is a diagram of a multiprocessor system of a second conventional example. FIG. 4 is a diagram of a second embodiment of the present invention. Fig. 6 is a detailed configuration diagram of one integer operation unit in Fig. 5 Fig. 7 is a detailed configuration diagram of one branch unit in Fig. 5 Fig. 8 is a configuration diagram of an instruction issuing mechanism of Fig. 4 FIG. 10 is a block diagram of the global access unit of FIG. 4. FIG. 11 is a block diagram of the data cache memory of FIG. 4. FIG. 12 is a block diagram of the shared floating-point unit of FIG. 4. FIG. FIG. 15 Register set table FIG. 16 Memory configuration diagram FIG. 17 Diagram of pipeline execution of a single instruction FIG. 18 Conceptual diagram of pipeline operation of entire chip FIG. Red switching operation diagram FIG. 20 Pipeline installation operation diagram FIG. 21 Program example using threads FIG. 22 Load distribution method FIG. 23 Overall diagram of third embodiment FIG. 24 Detailed diagram of arithmetic unit group in FIG. 25 Diagram of global access unit in FIG. 23 FIG. 26 Example of system configuration using microprocessor of third embodiment FIG. 27 Multiprocessor memory layout diagram FIG. 28 Bus transaction diagram of shared bus FIG. 29 Utility of loop inline expansion
1 マイクロプロセッサ 2 命令発行制御 3 プログラムカウンタ記憶手段 4 命令格納手段 5 演算手段 6 動的信号接続手段 7 データ格納手段 8 外部インターフェース 201 命令アドレス生成 202 命令キャッシュ 203 命令デコード 204 データクロスバスイッチ 205 分岐ユニット 206 ロードストアユニット 207 演算ユニット 208 レジスタファイル 209 データキャッシュ 210 外部バスインターフェース 301 マイクロプロセッサ 302 マイクロプロセッサ 303 マイクロプロセッサ 304 共有バスクロスバスイッチ 305 データメモリ 306 データメモリ 307 専用演算ユニット 308 I/Oデバイス 311 命令キャッシュ 312 命令発行制御 313 整数演算機 314 浮動小数点演算機 315 ロードストアユニット 316 データキャッシュ 317 キャッシュコヒーレンシ制御ユニット 401 マイクロプロセッサ 402 分岐クロスバスイッチ 403 スレッド制御ユニット 404 命令メモリ 405 整数演算ユニット群 406 ローカルキャッシュアクセスクロスバスイッチ 407 ローカルキャッシュバンク 408 グローバルアクセス制御 409 共有特殊演算ユニット 411 スレッド状態信号 412 プログラムカウンタ信号 413 スレッド制御ユニット間制御信号 414 スレッド状態信号 415 命令コード信号 416 整数演算ユニット間パイプラインデータバス 417 グローバルメモリアクセスアドレスバス 418 グローバルメモリアクセスデータバス 419 スレッド状態信号 421 ローカルキャッシュアドレスバス 422 ローカルキャッシュデータバス 423 共有演算機アドレスバス 424 共有演算機データバス 425 データキャッシュリプレースバス 426 ローカルキャッシュクロスバスイッチ406制
御信号 427 グローバルアクセスアドレスバス 428 グローバルアクセスデータバス 429 プログラムカウンタ信号 430 分岐クロスバスイッチ402制御信号 431 命令コード信号 432 命令リプレースデータバス 433、434、435、436 グローバルアクセス
要求信号 437 スレッド状態信号 450 プロセッサ外部データバス 451 プロセッサ外部制御バス 452 プロセッサ外部アドレスバス 453 プロセッサ外部割り込み要求信号 501 整数演算ユニット 502 分岐ユニット 503 分岐アービター 511、512、513、514 命令デコード信号 515 データバス群 516 分岐ユニット状態信号群 521 ローカルメモリアクセスアドレスバス 522 グローバルメモリアクセスアドレスバス 523 ローカルメモリアクセスデータバス 524 グローバルメモリアクセスデータバス 525、526、527、528 分岐要求信号群 601 プログラムカウンタラッチ 602 スタックポインタラッチ 603 スレッドIDラッチ 604 デコード済み命令コードラッチ 605 プログラムカウンタラッチ 606 条件実行命令制御信号ラッチ 607 フラグレジスタラッチ 608 フラグレジスタ更新ラッチ 609 ALU演算結果フラグラッチ 610 ALU演算結果ラッチ 611 バレルシフタ演算結果フラグラッチ 612 バレルシフタ演算結果ラッチ 613 ストアデータラッチ 614 レジスタライトバックデータ保持ラッチ 615 第1汎用レジスタラッチ 616 第2汎用レジスタラッチ 617 第3汎用レジスタラッチ 618 第4汎用レジスタラッチ 619 スタックポインタラッチ 621 プログラムカウンタ更新セレクタ 622 プログラムカウンタ定数生成回路 623 フラグレジスタ更新セレクタ 624 条件実行命令制御回路 625 プログラムカウンタバス 626 定数生成回路 627 第1オペランドセレクタ 628 第2オペランドセレクタ 629 ストアデータセレクタ 630 ALU回路 631 バレルシフタ 632 ストアアライン回路 633 スレッド間レジスタフォワードバス 634 第1汎用レジスタ更新セレクタ 635 第2汎用レジスタ更新セレクタ 636 第3汎用レジスタ更新セレクタ 637 第4汎用レジスタ更新セレクタ 638 スタックポインタ更新セレクタ 639 スタックポインタ信号 640 メモリストア用データバス 641 演算結果バス 642 データライトバックバス 643 第1オペランドバス 644 第2オペランドバス 645 第3オペランドバス 646 条件実行制御信号 647 スレッド状態信号 701 スレッド状態ラッチ 702 ロードデータラッチ 703 ロードデータ保持バッファ 704 ローカルメモリライトデータラッチ 705 ローカルメモリライトアドレスラッチ 706 アドレスセレクタ 707 アドレスバッファ 708 ストアデータセレクタ 709 ストアデータバッファ 711 スレッド状態更新セレクタ 711 分岐制御ユニット 712 ロードアライナ 713 ローカルメモリアクセス検査 714 分岐受理信号 801 Xデコーダ 802 RAMセル 803 センスアンプおよびライトバッファ 804 命令メモリアクセス制御 805 命令セレクタ 806 命令コードラッチ 807 命令デコードユニット 821 スレッド状態信号 822 プログラムカウンタ 830、831、832、833 命令コード 834 デコード済み命令 835 演算ユニット制御信号 836 デコード済み命令 837 演算ユニット制御信号 838 デコード済み命令 839 演算ユニット制御信号 840 デコード済み命令 841 演算ユニット制御信号 842 デコード済み命令 851 命令リプレースバス 901 スレッド発行アービトレーション 902 スタックポインタ連想メモリ 903 スレッド状態メモリ 904 スレッド開始準備フラグ 905 命令キャッシュアドレスタグ 906 ローカルキャッシュバンク番号 907 命令キャッシュアドレス比較器 908 スレッド発行ユニット 909 プログラムカウンタシフトレジスタ 921 スレッド発行アービトレーション要求 922 スレッド発行アービトレーション要求 923、924 ローカルキャッシュバンク信号 925、926 プログラムカウンタ信号 931、932、933 スレッド状態メモリワードラ
イン信号 934 ローカルキャッシュバンク信号 935 スレッド開始準備フラグ信号 936 スレッド状態信号 937 プログラムカウンタインデックス 938 プログラムカウンタアドレスタグ 939 命令キャッシュアドレスタグ 1001 分岐クロスバスイッチ402制御 1002 割り込みベクタ生成ユニット 1003 割り込み入力ユニット 1004 グローバルアドレスタグメモリ 1005 グローバルアドレスタグ比較器 1006 グローバルデータキャッシュメモリ 1007 外部バスインターフェースユニット 1008 ローカルメモリクロスバスイッチ406制御 1009 内部ロードストアインターフェース 1010 ロードストア分岐受理ユニット 1011 プログラムカウンタインクリメンタ 1012 命令キャッシュリプレースバッファ 1013 データキャッシュリプレースバッファ 1021 アドレスバス 1022 データバス 1101 データキャッシュRAM 1102 データキャッシュタグRAM 1103 アドレスタグ比較器 1104 保護機構チェック 1111 アドレスバス 1112 データバス 1113 リードライト制御信号 1114 アクセス違反通知信号 1201 浮動小数点レジスタ 1202 浮動小数点データパス 1203 アドレスデコードユニット 1204 命令デコードユニット 1401 待ち状態のスレッド 1402 実行中のスレッド 1403 命令メモリ 1801、1802、1803 スレッド 2301 マイクロプロセッサ 2305 演算ユニット群 2308 グローバルアクセス制御ユニット 2350 グローバルアドレス・データバス 2351 グローバル制御バス 2352 割り込み信号 2353 ローカル制御バス 2354 ローカルアドレスバス 2355 ローカルデータバス 2401 整数乗算ユニット 2402 浮動小数点演算ユニット 2403 浮動小数点レジスタユニット 2404 結果出力バス 2501 共有バスインターフェース 2502 ローカルバスインターフェース 2611 ローカルメモリ 2612 ローカルバス 2613 共有バス 2614 共有I/Oデバイス 2615 ローカルI/Oデバイス 2616 ローカルI/Oデバイス 2701 OS 2702 メモリ空間 2703、2704、2705 マイクロプロセッサロ
ーカルメモリ空間 2711、2712、2713 オブジェクト 2901 1スレッド分のプログラム 2902 整数演算機アレイ 2903 使用状態の整数演算機 2904 休止状態の整数演算機 2911 インライン化スレッドのプログラム 2911 整数演算機アレイ 2913 使用状態の整数演算機DESCRIPTION OF SYMBOLS 1 Microprocessor 2 Instruction issue control 3 Program counter storage means 4 Instruction storage means 5 Operation means 6 Dynamic signal connection means 7 Data storage means 8 External interface 201 Instruction address generation 202 Instruction cache 203 Instruction decode 204 Data crossbar switch 205 Branch unit 206 Load store unit 207 Operation unit 208 Register file 209 Data cache 210 External bus interface 301 Microprocessor 302 Microprocessor 303 Microprocessor 304 Shared bus crossbar switch 305 Data memory 306 Data memory 307 Dedicated operation unit 308 I / O device 311 Instruction cache 312 Instruction Issue control 313 Integer arithmetic unit 314 Floating point arithmetic unit 31 Load store unit 316 Data cache 317 Cache coherency control unit 401 Microprocessor 402 Branch crossbar switch 403 Thread control unit 404 Instruction memory 405 Integer operation unit group 406 Local cache access crossbar switch 407 Local cache bank 408 Global access control 409 Shared special operation unit 411 Thread status signal 412 Program counter signal 413 Control signal between thread control units 414 Thread status signal 415 Instruction code signal 416 Pipeline data bus between integer operation units 417 Global memory access address bus 418 Global memory access data bus 419 Thread status signal 421 Local cache Ad Bus 422 local cache data bus 423 shared processor address bus 424 shared processor data bus 425 data cache replacement bus 426 local cache crossbar switch 406 control signal 427 global access address bus 428 global access data bus 429 program counter signal 430 branch crossbar switch 402 Control signal 431 Instruction code signal 432 Instruction replacement data bus 433, 434, 435, 436 Global access request signal 437 Thread status signal 450 Processor external data bus 451 Processor external control bus 452 Processor external address bus 453 Processor external interrupt request signal 501 Integer operation Unit 502 Branch unit 503 Branch arbiter 5 11, 512, 513, 514 Instruction decode signal 515 Data bus group 516 Branch unit status signal group 521 Local memory access address bus 522 Global memory access address bus 523 Local memory access data bus 524 Global memory access data bus 525, 526, 527, 528 Branch request signal group 601 Program counter latch 602 Stack pointer latch 603 Thread ID latch 604 Decoded instruction code latch 605 Program counter latch 606 Condition execution instruction control signal latch 607 Flag register latch 608 Flag register update latch 609 ALU operation result flag latch 610 ALU Operation result latch 611 Barrel shifter Operation result flag latch 612 Barrel shift Operation result latch 613 Store data latch 614 Register write-back data holding latch 615 First general register latch 616 Second general register latch 617 Third general register latch 618 Fourth general register latch 619 Stack pointer latch 621 Program counter update selector 622 Program counter Constant generation circuit 623 Flag register update selector 624 Condition execution instruction control circuit 625 Program counter bus 626 Constant generation circuit 627 First operand selector 628 Second operand selector 629 Store data selector 630 ALU circuit 631 Barrel shifter 632 Store align circuit 633 Register between threads register forward Bus 634 first general register update selector 635 second general register update selector Lector 636 Third general register update selector 637 Fourth general register update selector 638 Stack pointer update selector 639 Stack pointer signal 640 Memory store data bus 641 Operation result bus 642 Data write back bus 643 First operand bus 644 Second operand bus 645 Third operand bus 646 Condition execution control signal 647 Thread status signal 701 Thread status latch 702 Load data latch 703 Load data holding buffer 704 Local memory write data latch 705 Local memory write address latch 706 Address selector 707 Address buffer 708 Store data selector 709 Store Data buffer 711 Thread state update selector 711 Branch control unit 71 Load aligner 713 Local memory access check 714 Branch accept signal 801 X decoder 802 RAM cell 803 Sense amplifier and write buffer 804 Instruction memory access control 805 Instruction selector 806 Instruction code latch 807 Instruction decode unit 821 Thread status signal 822 Program counter 830, 831 832, 833 Instruction code 834 Decoded instruction 835 Operation unit control signal 836 Decoded instruction 837 Operation unit control signal 838 Decoded instruction 839 Operation unit control signal 840 Decoded instruction 841 Operation unit control signal 842 Decoded instruction 851 Instruction replace bus 901 Thread issue arbitration 902 Stack pointer associative memory 903 Thread State memory 904 thread start preparation flag 905 instruction cache address tag 906 local cache bank number 907 instruction cache address comparator 908 thread issue unit 909 program counter shift register 921 thread issue arbitration request 922 thread issue arbitration request 923, 924 local cache bank signal 925, 926 Program counter signal 931, 932, 933 Thread status memory word line signal 934 Local cache bank signal 935 Thread start preparation flag signal 936 Thread status signal 937 Program counter index 938 Program counter address tag 939 Instruction cache address tag 1001 Branch crossbar switch 40 Control 1002 Interrupt vector generation unit 1003 Interrupt input unit 1004 Global address tag memory 1005 Global address tag comparator 1006 Global data cache memory 1007 External bus interface unit 1008 Local memory crossbar switch 406 control 1009 Internal load store interface 1010 Load store branch receiving unit 1011 Program counter incrementer 1012 Instruction cache replacement buffer 1013 Data cache replacement buffer 1021 Address bus 1022 Data bus 1101 Data cache RAM 1102 Data cache tag RAM 1103 Address tag comparator 1104 Protection mechanism check 1111 Address bus 1 12 Data bus 1113 Read / write control signal 1114 Access violation notification signal 1201 Floating point register 1202 Floating point data path 1203 Address decode unit 1204 Instruction decode unit 1401 Thread waiting 1402 Thread running 1403 Instruction memory 1801, 1802, 1803 Thread 2301 Microprocessor 2305 Operation unit group 2308 Global access control unit 2350 Global address / data bus 2351 Global control bus 2352 Interrupt signal 2353 Local control bus 2354 Local address bus 2355 Local data bus 2401 Integer multiplication unit 2402 Floating point operation unit 2403 Floating point register unit 2404 Output bus 2501 Shared bus interface 2502 Local bus interface 2611 Local memory 2612 Local bus 2613 Shared bus 2614 Shared I / O device 2615 Local I / O device 2616 Local I / O device 2701 OS 2702 Memory space 2703, 2704, 2705 Microprocessor local Memory space 2711, 2712, 2713 Object 2901 One thread program 2902 Integer arithmetic unit array 2903 Integer arithmetic unit 2904 Inactive integer arithmetic unit 2911 Inline thread program 2911 Integer arithmetic array 2913 Integer arithmetic in use state Machine
【手続補正書】[Procedure amendment]
【提出日】平成9年12月2日[Submission date] December 2, 1997
【手続補正1】[Procedure amendment 1]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】請求項1[Correction target item name] Claim 1
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【手続補正2】[Procedure amendment 2]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】請求項9[Correction target item name] Claim 9
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【手続補正3】[Procedure amendment 3]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】請求項10[Correction target item name] Claim 10
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【手続補正4】[Procedure amendment 4]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】請求項11[Correction target item name] Claim 11
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【手続補正5】[Procedure amendment 5]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】請求項12[Correction target item name] Claim 12
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【手続補正6】[Procedure amendment 6]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】請求項13[Correction target item name] Claim 13
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【手続補正7】[Procedure amendment 7]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】請求項14[Correction target item name] Claim 14
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【手続補正8】[Procedure amendment 8]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】請求項15[Correction target item name] Claim 15
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【手続補正9】[Procedure amendment 9]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】請求項16[Correction target item name] Claim 16
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
【手続補正10】[Procedure amendment 10]
【補正対象書類名】明細書[Document name to be amended] Statement
【補正対象項目名】請求項17[Correction target item name] Claim 17
【補正方法】変更[Correction method] Change
【補正内容】[Correction contents]
Claims (17)
と、命令選択値の入力に対して命令を出力する命令選択
値格納手段と、命令の内容に応じて演算を行う演算手
段、演算結果等を格納する一時記憶手段、これらから構
成される演算処理要素を複数持ち、演算処理要素の命令
選択値と演算結果を、隣接する演算処理要素の命令選択
値格納手段、演算手段に伝達し、その演算処理要素に別
の演算処理要素をくりかえし直列に接続し、スレッドの
命令を1つづつ順に演算処理要素の接続順に格納するこ
とを特長とする情報処理装置。An instruction selection value storage means for accumulating an instruction selection value, an instruction selection value storage means for outputting an instruction in response to an instruction selection value input, an operation means for performing an operation in accordance with the content of the instruction, Temporary storage means for storing results and the like, having a plurality of arithmetic processing elements composed of these, and transmitting an instruction selection value and an operation result of the arithmetic processing element to an instruction selection value storage means and an arithmetic means of an adjacent arithmetic processing element; An information processing apparatus characterized in that another arithmetic processing element is repeatedly connected in series to the arithmetic processing element, and the instructions of the threads are stored one by one in the connection order of the arithmetic processing elements.
て、前期演算手段と同一数の記憶手段を持ち、すべての
演算手段がそれぞれに固有の記憶手段を動的に選択し
て、同時に使用することを特徴とする情報処理装置。2. An information processing apparatus according to claim 1, wherein said information processing apparatus has the same number of storage means as said arithmetic means, and all the arithmetic means dynamically select their own storage means and use them simultaneously. An information processing apparatus, comprising:
て、命令選択値を複数蓄積する命令選択値蓄積手段を有
し、空いた命令選択値格納手段に対して自動的に新規の
命令選択値を割り当てることにより、演算処理要素にス
レッドを動的に割り当てることを特徴とするスレッド発
行制御手段を持つ情報処理装置。3. An information processing apparatus according to claim 2, further comprising: an instruction selection value storage unit for storing a plurality of instruction selection values, wherein a new instruction selection is automatically performed on the empty instruction selection value storage unit. An information processing apparatus having thread issuance control means, wherein a thread is dynamically assigned to an arithmetic processing element by assigning a value.
て、直列に接続された最後尾の演算処理要素の結果を、
最前列の演算処理要素の演算手段の入力として接続し、
同時に命令選択値の更新を行い、最前列の演算処理要素
の命令選択値格納手段に伝達することを特徴とする情報
処理装置。4. The information processing apparatus according to claim 3, wherein the result of the last arithmetic processing element connected in series is
Connected as the input of the arithmetic means of the arithmetic processing element in the front row,
An information processing apparatus for updating an instruction selection value at the same time and transmitting the updated instruction selection value to an instruction selection value storage means of an arithmetic processing element in the front row.
て、演算処理要素の中のスレッドの起点の位置を自由に
設定し、演算処理要素の実行再開位置を動作中に自由に
変更できることを特長とする情報処理装置。5. The information processing apparatus according to claim 4, wherein the starting position of the thread in the processing element can be freely set, and the execution restart position of the processing element can be freely changed during operation. An information processing device that is a feature.
て、スレッド発行制御手段を複数有し、演算手段の出力
値を使用して、スレッド発行制御手段と、演算処理要素
の起点を選択することを特徴とする情報処理装置。6. The information processing apparatus according to claim 5, further comprising a plurality of thread issuance control means, and selecting a thread issuance control means and a starting point of the operation processing element by using an output value of the operation means. An information processing apparatus characterized by the above-mentioned.
て、演算手段の数と同数の記憶手段と、演算手段と記憶
手段を1対1に自在に接続して同時に伝送する情報伝達
手段と、演算が終了すると同時に隣接する演算処理要素
に演算結果とともに記憶手段の接続を渡す手段を有する
ことを特徴とする情報処理装置。7. An information processing apparatus according to claim 4, wherein the number of storage means is equal to the number of calculation means, and the information transmission means is configured to freely connect the calculation means and storage means in a one-to-one manner and simultaneously transmit the information. An information processing apparatus having means for passing the connection of the storage means together with the operation result to an adjacent operation processing element at the same time as the end of the operation.
て、複数の記憶手段を、演算手段からの出力値によって
一意に選択し、すべての演算手段が互いに異なる記憶手
段を指定する記憶手段選択値を有し、この記憶手段選択
値を演算手段の結果の転送とともに隣接する演算処理要
素に伝送するこによって、複数の演算処理要素が同時に
1つの記憶手段に接続することを防ぐことを特徴とする
情報処理装置。8. The information processing apparatus according to claim 7, wherein a plurality of storage means are uniquely selected by an output value from the arithmetic means, and all arithmetic means specify different storage means. A plurality of arithmetic processing elements are prevented from being connected to one storage means at the same time by transmitting the selected value to the adjacent arithmetic processing element together with the transfer of the result of the arithmetic means. Information processing device.
て、演算処理要素の非決定的な例外事象に対して、その
演算処理要素からすべての演算結果を記憶手段選択値を
使用して記憶手段に自動的に伝送して、中断した演算処
理要素が有する命令選択値をスレッド発行制御手段に転
送し、例外事象の終了と同時に退避していた演算結果を
情報処理要素に伝達し、演算処理手段の動作を再開する
ことを特徴とする情報処理装置。9. An information processing apparatus according to claim 6, wherein, for a non-deterministic exceptional event of an arithmetic processing element, all arithmetic results from said arithmetic processing element are stored using a storage means selection value. Automatically transferring the instruction selection value of the interrupted operation processing element to the thread issue control means, and transmitting the operation result saved at the same time as the end of the exceptional event to the information processing element; An information processing apparatus characterized by restarting the operation of (1).
報処理装置において、情報処理要素に1対1に接続され
た記憶手段以外との情報入出力が必要な場合に、例外事
象を発生させて、低速な記憶手段からの情報入出力を行
い、終了と同時に退避していた状態と入力された情報を
演算処理要素に伝達し、演算処理要素の動作を再開する
ことを特徴とする情報処理装置。10. An information processing apparatus according to claim 8, wherein an exceptional event is generated when the information processing element needs to input / output information to / from other than the storage means connected one-to-one. Information input / output from / to the low-speed storage means, transmitting the saved information and the input information to the arithmetic processing element at the same time as the termination, and restarting the operation of the arithmetic processing element. Processing equipment.
おいて、演算処理要素に1対1に接続された記憶手段以
外との情報入出力が必要な場合、情報処理手段の内部に
演算手段と独立して記憶手段と通信を行う伝送情報蓄積
手段を有し、伝送情報蓄積手段と通信を行う記憶手段が
接続された時に情報の伝達を行うことを特徴とする情報
処理装置。11. The information processing apparatus according to claim 10, wherein when the arithmetic processing element needs to input / output information to / from a unit other than the storage unit connected in a one-to-one manner, the information processing unit includes the arithmetic unit. An information processing apparatus having transmission information storage means for independently communicating with a storage means, and transmitting information when the storage means for communicating with the transmission information storage means is connected.
おいて、情報処理手段の外部の記憶手段の内容の一部
を、情報処理手段の内部の前記記憶手段に一時蓄積して
演算手段との通信を行い、一時蓄積されていない場合に
自動的に外部の記憶手段と内容を入れ替えることを特徴
とし、複数の記憶手段が内容の入れ替え動作を同時に実
行することを特徴とする情報処理装置。12. An information processing apparatus according to claim 11, wherein a part of the contents of a storage means external to said information processing means is temporarily stored in said storage means inside said information processing means, and said information is stored in said storage means. An information processing apparatus for performing communication and automatically exchanging contents with an external storage means when data is not temporarily stored, and wherein a plurality of storage means simultaneously execute an operation of exchanging contents.
いて、スレッド発行制御手段に、命令選択値とともに記
憶手段選択値を格納し、命令選択値を出力する際に、格
納された記憶手段選択値を演算処理要素の持つ記憶手段
選択値と比較し、一致する場合にのみ演算処理要素に命
令選択値を出力することを特徴とする情報処理装置。13. The information processing apparatus according to claim 8, wherein the storage means selection value is stored together with the instruction selection value in the thread issuance control means, and the stored storage means selection value is output when the instruction selection value is output. An information processing apparatus for comparing a value with a storage unit selection value of an arithmetic processing element and outputting an instruction selection value to the arithmetic processing element only when the values match.
おいて、直列に接続された演算処理要素と独立した特別
演算手段を有し、すべての演算処理要素と特別演算手段
を接続し、演算処理要素は特別演算手段の演算の要求と
共にスレッド発行制御手段に状態を伝送し、特別演算手
段の演算の終了と同時に演算処理手段の動作を再開する
ことを特徴とする情報処理装置。14. An information processing apparatus according to claim 11, further comprising special operation means independent of the operation processing elements connected in series, and connecting all the operation processing elements and the special operation means. An information processing device, wherein the element transmits a state to the thread issue control means together with a request for operation of the special operation means, and restarts the operation of the operation processing means simultaneously with the end of the operation of the special operation means.
いて、直列に接続された演算処理要素のうち、2つ以上
の演算処理要素に対し、1つの特別な演算手段を接続
し、特別な演算手段を接続されたすべての演算処理要素
から使用できることを特徴とする情報処理装置。15. An information processing apparatus according to claim 4, wherein one special operation means is connected to two or more operation processing elements among the operation processing elements connected in series. An information processing apparatus characterized in that arithmetic means can be used from all connected arithmetic processing elements.
いて、別の情報処理装置が接続する記憶手段との通信を
行うための情報処理装置間通信手段を有することを特徴
とする情報処理装置。16. An information processing apparatus according to claim 2, further comprising an information processing apparatus communication means for communicating with a storage means connected to another information processing apparatus. .
報処理装置を複数用いたシステムにおいて、情報処理装
置はそれぞれ個別の記憶手段を有し、1つの記憶手段選
択値で、複数の情報処理装置が接続するすべての記憶手
段を一意に選択して、請求項16の情報処理装置間通信
手段で転送することを特徴とする情報処理装置。17. A system using a plurality of information processing apparatuses according to claim 2 and claim 8, wherein each of the information processing apparatuses has individual storage means, and a plurality of information is stored by one storage means selection value. 17. An information processing apparatus, wherein all storage means connected to the processing apparatus are uniquely selected and transferred by the information processing apparatus communication means according to claim 16.
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP09287662A JP3099290B2 (en) | 1997-10-03 | 1997-10-03 | Information processing device using multi-thread program |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP09287662A JP3099290B2 (en) | 1997-10-03 | 1997-10-03 | Information processing device using multi-thread program |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| JPH11110215A true JPH11110215A (en) | 1999-04-23 |
| JP3099290B2 JP3099290B2 (en) | 2000-10-16 |
Family
ID=17720113
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP09287662A Expired - Fee Related JP3099290B2 (en) | 1997-10-03 | 1997-10-03 | Information processing device using multi-thread program |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JP3099290B2 (en) |
Cited By (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2002333978A (en) * | 2001-05-08 | 2002-11-22 | Nec Corp | Vliw type processor |
| JP2006039815A (en) * | 2004-07-26 | 2006-02-09 | Fujitsu Ltd | Multi-thread processor and register control method |
| JP2006174902A (en) * | 2004-12-21 | 2006-07-06 | Hitachi Medical Corp | Ultrasonic diagnostic equipment |
| JP2007004338A (en) * | 2005-06-22 | 2007-01-11 | Renesas Technology Corp | Data processor |
| JP2007213578A (en) * | 2006-02-09 | 2007-08-23 | Internatl Business Mach Corp <Ibm> | Data-cache miss prediction and scheduling |
| JP2008509493A (en) * | 2004-08-13 | 2008-03-27 | クリアスピード テクノロジー パブリック リミテッド カンパニー | Processor memory system |
| JP2009238132A (en) * | 2008-03-28 | 2009-10-15 | Nec Corp | Data processing apparatus |
| US7765250B2 (en) | 2004-11-15 | 2010-07-27 | Renesas Technology Corp. | Data processor with internal memory structure for processing stream data |
| JP2010532905A (en) * | 2008-06-26 | 2010-10-14 | ラッセル・エイチ・フィッシュ | Thread-optimized multiprocessor architecture |
| JP2011048735A (en) * | 2009-08-28 | 2011-03-10 | Ricoh Co Ltd | Simd microprocessor |
| JP2014016773A (en) * | 2012-07-09 | 2014-01-30 | Elamina Co Ltd | Cashless multiprocessor by registerless architecture |
| CN114860319A (en) * | 2022-05-12 | 2022-08-05 | 中国科学院计算技术研究所 | Interactive arithmetic device and execution method for SIMD (Single instruction multiple data) calculation instruction |
-
1997
- 1997-10-03 JP JP09287662A patent/JP3099290B2/en not_active Expired - Fee Related
Cited By (13)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2002333978A (en) * | 2001-05-08 | 2002-11-22 | Nec Corp | Vliw type processor |
| US8447959B2 (en) | 2004-07-26 | 2013-05-21 | Fujitsu Limited | Multithread processor and method of controlling multithread processor |
| JP2006039815A (en) * | 2004-07-26 | 2006-02-09 | Fujitsu Ltd | Multi-thread processor and register control method |
| JP2008509493A (en) * | 2004-08-13 | 2008-03-27 | クリアスピード テクノロジー パブリック リミテッド カンパニー | Processor memory system |
| US7765250B2 (en) | 2004-11-15 | 2010-07-27 | Renesas Technology Corp. | Data processor with internal memory structure for processing stream data |
| JP2006174902A (en) * | 2004-12-21 | 2006-07-06 | Hitachi Medical Corp | Ultrasonic diagnostic equipment |
| JP2007004338A (en) * | 2005-06-22 | 2007-01-11 | Renesas Technology Corp | Data processor |
| JP2007213578A (en) * | 2006-02-09 | 2007-08-23 | Internatl Business Mach Corp <Ibm> | Data-cache miss prediction and scheduling |
| JP2009238132A (en) * | 2008-03-28 | 2009-10-15 | Nec Corp | Data processing apparatus |
| JP2010532905A (en) * | 2008-06-26 | 2010-10-14 | ラッセル・エイチ・フィッシュ | Thread-optimized multiprocessor architecture |
| JP2011048735A (en) * | 2009-08-28 | 2011-03-10 | Ricoh Co Ltd | Simd microprocessor |
| JP2014016773A (en) * | 2012-07-09 | 2014-01-30 | Elamina Co Ltd | Cashless multiprocessor by registerless architecture |
| CN114860319A (en) * | 2022-05-12 | 2022-08-05 | 中国科学院计算技术研究所 | Interactive arithmetic device and execution method for SIMD (Single instruction multiple data) calculation instruction |
Also Published As
| Publication number | Publication date |
|---|---|
| JP3099290B2 (en) | 2000-10-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11204769B2 (en) | Memory fragments for supporting code block execution by using virtual cores instantiated by partitionable engines | |
| US9934072B2 (en) | Register file segments for supporting code block execution by using virtual cores instantiated by partitionable engines | |
| US9990200B2 (en) | Executing instruction sequence code blocks by using virtual cores instantiated by partitionable engines | |
| US11829187B2 (en) | Microprocessor with time counter for statically dispatching instructions | |
| US11954491B2 (en) | Multi-threading microprocessor with a time counter for statically dispatching instructions | |
| CN109375949B (en) | Processor with multiple cores | |
| JP2001236221A (en) | Pipe line parallel processor using multi-thread | |
| US11829762B2 (en) | Time-resource matrix for a microprocessor with time counter for statically dispatching instructions | |
| JP3099290B2 (en) | Information processing device using multi-thread program | |
| US12147812B2 (en) | Out-of-order execution of loop instructions in a microprocessor | |
| CN115437691A (en) | Physical register file allocation device for RISC-V vector and floating point register | |
| US12106114B2 (en) | Microprocessor with shared read and write buses and instruction issuance to multiple register sets in accordance with a time counter | |
| CN120508531A (en) | Dynamic reconfiguration of multi-core processor to unified core | |
| CN120508530A (en) | Dynamic reconfiguration of unified core processor to multi-core processor | |
| Kang et al. | VLSI implementation of multiprocessor system |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| LAPS | Cancellation because of no payment of annual fees |