JPH04255995A - Instruction cache - Google Patents
Instruction cacheInfo
- Publication number
- JPH04255995A JPH04255995A JP3016716A JP1671691A JPH04255995A JP H04255995 A JPH04255995 A JP H04255995A JP 3016716 A JP3016716 A JP 3016716A JP 1671691 A JP1671691 A JP 1671691A JP H04255995 A JPH04255995 A JP H04255995A
- Authority
- JP
- Japan
- Prior art keywords
- instruction
- length
- instructions
- cache
- cpu
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
Landscapes
- Advance Control (AREA)
- Static Random-Access Memory (AREA)
- Memory System Of A Hierarchy Structure (AREA)
Abstract
Description
【0001】0001
【産業上の利用分野】本発明は、コンピュータにおいて
、中央処理装置(以下、CPUという)と主記憶装置と
の間に接続されて使用される高速メモリ、いわゆる命令
キャッシュに関する。BACKGROUND OF THE INVENTION 1. Field of the Invention The present invention relates to a high-speed memory, so-called instruction cache, which is used in a computer by being connected between a central processing unit (hereinafter referred to as a CPU) and a main memory.
【0002】現在、大部分の命令を1サイクルで実行す
るCPUが商品化されてはいるが、ユーザは、演算処理
の更なる高速化を要求している。ここに、演算処理を高
速化する方法としては、動作クロックの周波数を高める
方法がある。しかし、周辺回路が構成上、これに対応で
きない状況となっている。そこで、動作クロックの周波
数を高めることなく、演算処理の高速化を図る方法が要
求される。この要求に応える方法として、複数命令を同
時に実行する方法があり、いわゆる固定長命令形式のコ
ンピュータにおいては、これを容易に行うことができる
。しかし、可変長命令形式のコンピュータにおいても、
複数命令を同時に実行できるようにするためには、CP
Uと主記憶装置との間に接続されて使用される命令キャ
ッシュの改良が必要となる。Although CPUs that execute most instructions in one cycle are currently on the market, users are demanding even faster arithmetic processing. Here, as a method of speeding up arithmetic processing, there is a method of increasing the frequency of the operating clock. However, due to the configuration of the peripheral circuits, it is not possible to support this. Therefore, there is a need for a method for increasing the speed of arithmetic processing without increasing the frequency of the operating clock. One way to meet this demand is to simultaneously execute multiple instructions, and this can be easily done in a so-called fixed-length instruction format computer. However, even in computers with variable-length instruction format,
In order to be able to execute multiple instructions simultaneously, the CP
Improvements are needed in the instruction cache used in connection between the U and the main memory.
【0003】0003
【従来の技術】図3は、従来の命令キャッシュ1を使用
してなるコンピュータの一例の要部を示す図であり、2
はCPU、3は主記憶装置である。また、命令キャッシ
ュ1において、4は主記憶装置3から取り込んだ命令を
格納する命令格納部、5はCPU2から出力される命令
アドレスと命令キャッシュ1に保持されている比較アド
レスとを比較し、キャッシュヒットしたか、キャッシュ
ミスしたかを判定するキャッシュヒット/キャッシュミ
ス判定部である。2. Description of the Related Art FIG. 3 is a diagram showing a main part of an example of a computer using a conventional instruction cache 1.
is a CPU, and 3 is a main storage device. In the instruction cache 1, numeral 4 is an instruction storage section that stores instructions fetched from the main memory 3, and 5 is an instruction storage section that compares the instruction address output from the CPU 2 with a comparison address held in the instruction cache 1, and This is a cache hit/cache miss determination unit that determines whether there is a hit or a cache miss.
【0004】かかるコンピュータにおいては、CPU2
が命令を取り込む場合、まず、CPU2から命令キャッ
シュ1にアクセスがなされ、取り込むべき命令が命令キ
ャッシュ1にあった場合(キャッシュヒットした場合)
、この命令が取り込まれる。他方、取り込むべき命令が
命令キャッシュ1になかった場合(キャッシュミスした
場合)には、主記憶装置3に対してアクセスがなされ、
該当する命令が読み出されて、これがCPU2に取り込
まれると共に命令キャッシュ1に格納される。[0004] In such a computer, the CPU 2
When fetching an instruction, first, instruction cache 1 is accessed from CPU 2, and the instruction to be fetched is in instruction cache 1 (cache hit).
, this instruction is captured. On the other hand, if the instruction to be fetched is not in the instruction cache 1 (cache miss), the main memory 3 is accessed,
The corresponding instruction is read out, taken into the CPU 2, and stored in the instruction cache 1.
【0005】[0005]
【発明が解決しようとする課題】ところで、CPU2が
複数命令を同時に実行できるように構成されている場合
において、命令が固定長である場合には、先頭命令に続
く命令の位置が固定されているので、複数命令をデコー
ドすることにより、複数命令を同時に実行し、演算処理
の高速化を図ることができる。しかしながら、命令が可
変長の場合、先頭命令に続く命令の位置が変動するため
、何らかの方法で後続命令の位置を特定できないと、複
数命令を同時に実行することができない。[Problem to be Solved by the Invention] By the way, when the CPU 2 is configured to be able to execute multiple instructions at the same time, and the instructions have a fixed length, the positions of the instructions following the first instruction are fixed. Therefore, by decoding a plurality of instructions, it is possible to simultaneously execute the plurality of instructions and speed up arithmetic processing. However, if the instructions have a variable length, the positions of the instructions following the first instruction vary, so unless the position of the subsequent instructions can be identified by some method, multiple instructions cannot be executed simultaneously.
【0006】ここに、可変長命令においても、複数命令
を同時に実行できるようにする方法として、図4に示す
ように、CPU2内に命令長を解読する命令プリデコー
ダ6を設け、命令キャッシュ1から取り込んだ可変長命
令の命令長を命令プリデコーダ6で解読し、複数命令を
命令デコーダ7に供給する方法が考えられる。なお、8
、9は命令を一時的に格納するための命令レジスタであ
る。Here, as a method for making it possible to execute multiple instructions at the same time even with variable length instructions, as shown in FIG. A possible method is to decode the instruction length of the fetched variable-length instruction using the instruction predecoder 6 and supply a plurality of instructions to the instruction decoder 7. In addition, 8
, 9 are instruction registers for temporarily storing instructions.
【0007】しかしながら、この方法においては、可変
長命令をCPU2内に取り込んでから命令長の解読をし
ており、しかも、この場合、複数の可変長命令の命令長
を同時に解読する必要があるが、複数の可変長命令を同
時に解読することは容易でなく、このため、むしろ、演
算処理が低速化してしまうという問題点があった。However, in this method, the instruction length is decoded after the variable length instruction is loaded into the CPU 2, and in this case, it is necessary to decode the instruction length of multiple variable length instructions at the same time. , it is not easy to decode a plurality of variable-length instructions at the same time, and as a result, there is a problem in that the speed of arithmetic processing is rather slowed down.
【0008】また、図5に示すように、命令キャッシュ
内に命令デコーダ10を設け、この命令デコーダ10に
よって命令をデコードし、この結果を命令格納部4に格
納する方法が考えられる。しかしながら、この場合には
、命令格納部4のデータサイズを最長命令に対応させて
大きくする必要があり、実用的ではなかった。なお、1
1は命令レジスタである。Furthermore, as shown in FIG. 5, a method can be considered in which an instruction decoder 10 is provided in the instruction cache, the instruction decoder 10 decodes instructions, and the result is stored in the instruction storage section 4. However, in this case, it was necessary to increase the data size of the instruction storage section 4 to correspond to the longest instruction, which was not practical. In addition, 1
1 is an instruction register.
【0009】本発明は、かかる点に鑑み、可変長命令で
あっても、CPUにおいて複数命令を同時にデコードし
、演算処理の高速化を図ることができるようにした命令
キャッシュを提供することを目的とする。SUMMARY OF THE INVENTION In view of the above, an object of the present invention is to provide an instruction cache that allows a CPU to decode multiple instructions at the same time, even if they are variable length instructions, thereby speeding up arithmetic processing. shall be.
【0010】0010
【課題を解決するための手段】図1は本発明の原理説明
図であって、本発明による命令キャッシュは、主記憶装
置3から取り込んだ可変長命令の命令長を解読するため
の命令長解読手段12と、主記憶装置3から取り込んだ
可変長命令を前記命令長解読手段12によって得られる
命令長情報と共に格納する命令格納手段13とを備えて
構成される。[Means for Solving the Problems] FIG. 1 is an explanatory diagram of the principle of the present invention, and the instruction cache according to the present invention is an instruction length decoder for decoding the instruction length of a variable length instruction fetched from a main storage device 3. The instruction storing means 13 stores the variable length instructions fetched from the main memory 3 together with the instruction length information obtained by the instruction length decoding means 12.
【0011】[0011]
【作用】本発明においては、可変長命令は、その命令長
情報と共に命令格納手段13に格納されるので、キャッ
シュヒットした場合には、該当する命令がその命令長情
報と共に命令格納手段13からCPU2に供給される。
したがって、可変長命令であっても、CPU2において
複数命令を同時にデコードすることができる。[Operation] In the present invention, a variable length instruction is stored in the instruction storage means 13 together with its instruction length information, so when a cache hit occurs, the corresponding instruction is transferred from the instruction storage means 13 together with its instruction length information to the CPU 2. supplied to Therefore, even if the instructions are variable length, the CPU 2 can decode multiple instructions at the same time.
【0012】なお、キャッシュミスの場合、該当する可
変長命令が主記憶装置3から本発明の命令キャッシュに
取り込まれ、命令長解読手段12によってその命令長が
解読された後、かかる可変長命令は、命令長情報と共に
CPU2に供給される。したがって、可変長命令がCP
U2においてデコードされるまでにキャッシュヒットし
た場合よりも時間を要するが、この場合でも、複数命令
を同時にデコードすることができることに変わりはない
。In the case of a cache miss, the corresponding variable length instruction is fetched from the main storage device 3 into the instruction cache of the present invention, and after its instruction length is decoded by the instruction length decoding means 12, the variable length instruction is , are supplied to the CPU 2 together with instruction length information. Therefore, variable length instructions are
Although it takes more time to decode in U2 than when there is a cache hit, even in this case, multiple instructions can still be decoded at the same time.
【0013】[0013]
【実施例】以下、図2を参照して、本発明の一実施例に
ついて説明する。なお、図2において、CPU及び主記
憶装置については図3〜図5の場合と同一の符号を付し
ている。[Embodiment] An embodiment of the present invention will be described below with reference to FIG. Note that in FIG. 2, the CPU and main storage device are given the same reference numerals as in FIGS. 3 to 5.
【0014】図2は本実施例の命令キャッシュ14を使
用してなるコンピュータの一例の要部を示すブロック図
であり、本実施例では、命令キャッシュ14は、CPU
2を設けてなるマイクロプロセッサに内蔵されている。
即ち、本実施例は命令キャッシュ14をCPU2と共に
単一の半導体基板に形成した例である。FIG. 2 is a block diagram showing the main parts of an example of a computer using the instruction cache 14 of this embodiment.
It is built into a microprocessor with 2. That is, this embodiment is an example in which the instruction cache 14 and the CPU 2 are formed on a single semiconductor substrate.
【0015】ここに、命令キャッシュ14において、1
5は命令アドレス入力部、即ち、CPU2から出力され
る命令アドレスを受け取るブロックである。Here, in the instruction cache 14, 1
Reference numeral 5 denotes an instruction address input section, that is, a block that receives an instruction address output from the CPU 2.
【0016】また、16はキャッシュヒット/キャッシ
ュミス判定部、即ち、CPU2から出力される命令アド
レスと命令キャッシュ14に保持されている比較アドレ
スとを比較し、キャッシュヒットしたか、キャッシュミ
スしたかを判定するブロックであり、比較アドレスを保
持するRAM17と、比較アドレスと命令アドレスとを
比較するアドレス比較部18とで構成されている。Further, 16 is a cache hit/cache miss determination unit, which compares the instruction address output from the CPU 2 and the comparison address held in the instruction cache 14 to determine whether there is a cache hit or a cache miss. This is a block for making decisions, and is made up of a RAM 17 that holds a comparison address, and an address comparison section 18 that compares the comparison address and the instruction address.
【0017】19は命令アドレス出力部、即ち、アドレ
ス比較部18からキャッシュミスが通知された場合、命
令アドレス入力部15に入力された命令アドレスを主記
憶装置3に出力するブロックである。Reference numeral 19 denotes an instruction address output section, that is, a block that outputs the instruction address input to the instruction address input section 15 to the main storage device 3 when a cache miss is notified from the address comparison section 18.
【0018】また、20は命令データ入力部、即ち、主
記憶装置3から出力された可変長命令を受け取るブロッ
クである。Reference numeral 20 denotes an instruction data input section, that is, a block that receives variable length instructions output from the main storage device 3.
【0019】また、21は命令プリデコーダと称される
デコーダであり、本実施例に特有のブロックであって、
主記憶装置3から取り込んだ可変長命令をCPU2に設
けられている命令デコーダに先がけてデコードし、命令
長情報を得るように構成されたデコーダである。ここに
、命令長情報は、例えば、表1に示すように構成される
。Further, 21 is a decoder called an instruction predecoder, which is a block unique to this embodiment.
This decoder is configured to decode variable length instructions taken in from the main storage device 3 prior to the instruction decoder provided in the CPU 2 to obtain instruction length information. Here, the instruction length information is configured as shown in Table 1, for example.
【0020】[0020]
【表1】[Table 1]
【0021】また、22は命令格納部、23はバイパス
部、24は命令データ出力部であり、命令格納部22は
、入出力バッファ25と、命令データ入力部20に入力
された命令及び命令プリデコーダ21によって得られる
命令長情報を格納するためのRAM26とで構成されて
いる。Further, 22 is an instruction storage section, 23 is a bypass section, and 24 is an instruction data output section. It is composed of a RAM 26 for storing instruction length information obtained by the decoder 21.
【0022】ここに、入出力バッファ25は、アドレス
比較部18からキャッシュヒットを通知された場合、該
当する命令及び命令長情報をRAM26から読み出し、
これらを命令データ出力部24に供給し、また、アドレ
ス比較部18からキャッシュミスを通知された場合には
、命令データ入力部20から命令データを受け取ると共
に、命令プリデコーダ21から命令長情報を受け取り、
これらをRAM26に書き込むように動作するものであ
る。Here, when the input/output buffer 25 is notified of a cache hit from the address comparator 18, it reads the corresponding instruction and instruction length information from the RAM 26, and
These are supplied to the instruction data output section 24, and when a cache miss is notified from the address comparison section 18, instruction data is received from the instruction data input section 20, and instruction length information is received from the instruction predecoder 21. ,
It operates to write these into the RAM 26.
【0023】また、バイパス部23は、キャッシュミス
した場合に、命令データ入力部20から可変長命令を受
け取ると共に、命令プリデコーダ21から命令長情報を
受け取り、これらを命令データ出力部24に転送するも
のである。Furthermore, when a cache miss occurs, the bypass section 23 receives a variable length instruction from the instruction data input section 20, receives instruction length information from the instruction predecoder 21, and transfers these to the instruction data output section 24. It is something.
【0024】命令データ出力部24は、命令格納部22
又はバイパス部23から供給される可変長命令及び命令
長情報をCPU2に出力するものである。The instruction data output section 24 is connected to the instruction storage section 22.
Alternatively, variable length instructions and instruction length information supplied from the bypass section 23 are output to the CPU 2.
【0025】かかる本実施例においては、主記憶装置3
から取り込まれた可変長命令は、命令プリデコーダ21
によって解読された命令長情報と共に命令格納部22に
格納されるので、キャッシュヒットした場合には、該当
する可変長命令がその命令長情報と共に命令格納部22
から命令データ出力部24を介してCPU2に供給され
る。したがって、可変長命令であっても、CPU2にお
いて複数命令を同時にデコードし、高速化を図ることが
できる。In this embodiment, the main storage device 3
The variable length instruction fetched from the instruction predecoder 21
Since the variable-length instruction is stored in the instruction storage unit 22 together with the instruction length information decoded by
is supplied to the CPU 2 via the instruction data output section 24. Therefore, even if the instructions are variable length, the CPU 2 can decode a plurality of instructions at the same time to increase the speed.
【0026】なお、キャッシュミスの場合、該当する可
変長命令が主記憶装置3から取り込まれ、命令プリデコ
ーダ21によってその命令長が解読された後、かかる可
変長命令は命令長情報と共にCPU2に供給される。し
たがって、この可変長命令がCPU2においてデコード
されるまでに、キャッシュヒットした場合よりも時間を
要するが、この場合でも、複数命令を同時にデコードす
ることができることに変わりはない。In the case of a cache miss, the corresponding variable length instruction is fetched from the main storage device 3, its instruction length is decoded by the instruction pre-decoder 21, and then the variable length instruction is supplied to the CPU 2 together with instruction length information. be done. Therefore, it takes more time for this variable length instruction to be decoded by the CPU 2 than when there is a cache hit, but even in this case, it is still possible to decode a plurality of instructions at the same time.
【0027】[0027]
【発明の効果】本発明によれば、命令長解読手段を設け
、この命令長解読手段によって得られる命令長情報を可
変長命令と共に命令格納手段に格納するように構成した
ことにより、キャッシュヒットした場合には、該当する
可変長命令及びその命令長情報を命令格納手段からCP
Uに供給できるので、可変長命令であっても、CPUに
おいて複数命令を同時にデコードし、高速化を図ること
ができる。According to the present invention, the instruction length decoding means is provided and the instruction length information obtained by the instruction length decoding means is stored in the instruction storage means together with the variable length instruction, so that cache hits can be avoided. In this case, the corresponding variable length instruction and its instruction length information are transferred from the instruction storage means to the CP
Since the instruction can be supplied to U, even if the instruction is a variable length instruction, the CPU can decode multiple instructions at the same time and increase the speed.
【図1】本発明の原理説明図である。FIG. 1 is a diagram explaining the principle of the present invention.
【図2】本発明の一実施例の要部をCPU及び主記憶装
置と共に示すブロック図である。FIG. 2 is a block diagram showing main parts of an embodiment of the present invention together with a CPU and a main storage device.
【図3】従来の命令キャッシュを使用してなるコンピュ
ータの一例の要部を示すブロック図である。FIG. 3 is a block diagram showing main parts of an example of a computer using a conventional instruction cache.
【図4】従来の命令キャッシュを使用してなるコンピュ
ータの他の例の要部を示すブロック図である。FIG. 4 is a block diagram showing the main parts of another example of a computer using a conventional instruction cache.
【図5】従来の命令キャッシュを使用してなるコンピュ
ータの更に他の例の要部を示すブロック図である。FIG. 5 is a block diagram showing the main parts of yet another example of a computer using a conventional instruction cache.
2 CPU 3 主記憶装置 12 命令長解読手段 13 命令格納手段 2 CPU 3 Main memory 12 Instruction length decoding means 13 Instruction storage means
Claims (1)
令の命令長を解読するための命令長解読手段(12)と
、前記主記憶装置(3)から取り込んだ可変長命令を前
記命令長解読手段(12)によって得られる命令長情報
と共に格納する命令格納手段(13)とを備えて構成さ
れていることを特徴とする命令キャッシュ。1. An instruction length decoding means (12) for decoding the instruction length of a variable length instruction fetched from a main memory device (3); An instruction cache comprising: instruction storage means (13) for storing instruction length information obtained by the length decoding means (12).
Priority Applications (5)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP3016716A JPH04255995A (en) | 1991-02-08 | 1991-02-08 | Instruction cache |
| DE69231011T DE69231011T2 (en) | 1991-02-08 | 1992-02-06 | Cache memory for processing command data and data processor with the same |
| EP92301024A EP0498654B1 (en) | 1991-02-08 | 1992-02-06 | Cache memory processing instruction data and data processor including the same |
| KR1019920001755A KR960011279B1 (en) | 1991-02-08 | 1992-02-07 | Data processing cache memory and data processor with it |
| US07/833,412 US5488710A (en) | 1991-02-08 | 1992-02-10 | Cache memory and data processor including instruction length decoding circuitry for simultaneously decoding a plurality of variable length instructions |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP3016716A JPH04255995A (en) | 1991-02-08 | 1991-02-08 | Instruction cache |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| JPH04255995A true JPH04255995A (en) | 1992-09-10 |
Family
ID=11923993
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP3016716A Withdrawn JPH04255995A (en) | 1991-02-08 | 1991-02-08 | Instruction cache |
Country Status (1)
| Country | Link |
|---|---|
| JP (1) | JPH04255995A (en) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH07160578A (en) * | 1993-12-13 | 1995-06-23 | Nec Corp | Information processor |
| JP2011503719A (en) * | 2007-11-02 | 2011-01-27 | クゥアルコム・インコーポレイテッド | Predecode repair cache for instructions that span instruction cache lines |
-
1991
- 1991-02-08 JP JP3016716A patent/JPH04255995A/en not_active Withdrawn
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH07160578A (en) * | 1993-12-13 | 1995-06-23 | Nec Corp | Information processor |
| JP2011503719A (en) * | 2007-11-02 | 2011-01-27 | クゥアルコム・インコーポレイテッド | Predecode repair cache for instructions that span instruction cache lines |
| JP2014044731A (en) * | 2007-11-02 | 2014-03-13 | Qualcomm Incorporated | Predecode repair cache for instructions that cross instruction cache line |
| US8898437B2 (en) | 2007-11-02 | 2014-11-25 | Qualcomm Incorporated | Predecode repair cache for instructions that cross an instruction cache line |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US6832305B2 (en) | Method and apparatus for executing coprocessor instructions | |
| JP2773471B2 (en) | Information processing device | |
| KR950033847A (en) | Method and apparatus for lazy recording of storage instructions in processor unit | |
| JP2009032257A (en) | Processor architecture selectively using finite-state-machine for control code | |
| US20220100526A1 (en) | Apparatus and method for low-latency decompression acceleration via a single job descriptor | |
| JPH03233630A (en) | Information processor | |
| JPH077356B2 (en) | Pipelined microprocessor | |
| US5421026A (en) | Data processor for processing instruction after conditional branch instruction at high speed | |
| JPS63197232A (en) | Microprocessor | |
| JPH04104350A (en) | Micro processor | |
| JPH0377137A (en) | Information processor | |
| JP4404373B2 (en) | Semiconductor integrated circuit | |
| JP4601624B2 (en) | Direct memory access unit with instruction predecoder | |
| JPH07114509A (en) | Memory access device | |
| JPH0524537B2 (en) | ||
| JP3147884B2 (en) | Storage device and information processing device | |
| JP2762798B2 (en) | Information processing apparatus of pipeline configuration having instruction cache | |
| JPH0425937A (en) | Information processor | |
| JP2762797B2 (en) | Information processing apparatus of pipeline configuration having instruction cache | |
| JP2867798B2 (en) | Advance control unit | |
| JPH07129469A (en) | Cache memory and microprocessor having the same | |
| JPH0553910A (en) | Cache storage | |
| JPH024011B2 (en) | ||
| JPH02294829A (en) | Memory operand takeout system | |
| JPH06314196A (en) | Information processing method and device |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| A300 | Application deemed to be withdrawn because no request for examination was validly filed |
Free format text: JAPANESE INTERMEDIATE CODE: A300 Effective date: 19980514 |