KR20200083123A - 로드-저장 명령 - Google Patents
로드-저장 명령 Download PDFInfo
- Publication number
- KR20200083123A KR20200083123A KR1020190056970A KR20190056970A KR20200083123A KR 20200083123 A KR20200083123 A KR 20200083123A KR 1020190056970 A KR1020190056970 A KR 1020190056970A KR 20190056970 A KR20190056970 A KR 20190056970A KR 20200083123 A KR20200083123 A KR 20200083123A
- Authority
- KR
- South Korea
- Prior art keywords
- instruction
- load
- register
- stride
- store
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/06—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons
- G06N3/063—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using electronic means
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30003—Arrangements for executing specific machine instructions
- G06F9/3004—Arrangements for executing specific machine instructions to perform operations on memory
- G06F9/30043—LOAD or STORE instructions; Clear instruction
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F17/00—Digital computing or data processing equipment or methods, specially adapted for specific functions
- G06F17/10—Complex mathematical operations
- G06F17/16—Matrix or vector computation, e.g. matrix-matrix or matrix-vector multiplication, matrix factorization
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30003—Arrangements for executing specific machine instructions
- G06F9/30007—Arrangements for executing specific machine instructions to perform operations on data operands
- G06F9/3001—Arithmetic instructions
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30098—Register arrangements
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30098—Register arrangements
- G06F9/30101—Special purpose registers
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30098—Register arrangements
- G06F9/30105—Register structure
- G06F9/30109—Register structure having multiple operands in a single register
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30098—Register arrangements
- G06F9/3012—Organisation of register space, e.g. banked or distributed register file
- G06F9/3013—Organisation of register space, e.g. banked or distributed register file according to data content, e.g. floating-point registers, address registers
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/30098—Register arrangements
- G06F9/30141—Implementation provisions of register files, e.g. ports
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/34—Addressing or accessing the instruction operand or the result ; Formation of operand address; Addressing modes
- G06F9/345—Addressing or accessing the instruction operand or the result ; Formation of operand address; Addressing modes of multiple operands or results
- G06F9/3455—Addressing or accessing the instruction operand or the result ; Formation of operand address; Addressing modes of multiple operands or results using stride
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3836—Instruction issuing, e.g. dynamic instruction scheduling or out of order instruction execution
- G06F9/3851—Instruction issuing, e.g. dynamic instruction scheduling or out of order instruction execution from multiple instruction streams, e.g. multistreaming
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3854—Instruction completion, e.g. retiring, committing or graduating
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3867—Concurrent instruction execution, e.g. pipeline or look ahead using instruction pipelines
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F9/00—Arrangements for program control, e.g. control units
- G06F9/06—Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
- G06F9/30—Arrangements for executing machine instructions, e.g. instruction decode
- G06F9/38—Concurrent instruction execution, e.g. pipeline or look ahead
- G06F9/3885—Concurrent instruction execution, e.g. pipeline or look ahead using a plurality of independent parallel functional units
- G06F9/3888—Concurrent instruction execution, e.g. pipeline or look ahead using a plurality of independent parallel functional units controlled by a single instruction for multiple threads [SIMT] in parallel
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Software Systems (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Mathematical Physics (AREA)
- Pure & Applied Mathematics (AREA)
- Mathematical Optimization (AREA)
- Mathematical Analysis (AREA)
- Computational Mathematics (AREA)
- Data Mining & Analysis (AREA)
- Computing Systems (AREA)
- Biophysics (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Biomedical Technology (AREA)
- Databases & Information Systems (AREA)
- Algebra (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Computational Linguistics (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Neurology (AREA)
- Advance Control (AREA)
- Executing Machine-Instructions (AREA)
Abstract
Description
도 1은 예시적인 멀티-스레드 프로세서의 개략적인 블록도이다.
도 2는 인터리빙된 타임 슬롯의 방식을 개략적으로 도시한다.
도 3은 수퍼바이저 스레드 및 복수의 인터리빙된 타임 슬롯에서 동작하는 복수의 워커 스레드를 개략적으로 도시한다.
도 4는 예시적인 프로세서의 논리 블록 구조를 개략적으로 도시한다.
도 5는 구성 프로세서들의 어레이를 포함하는 프로세서의 개략적인 블록도이다.
도 6은 기계 지능 알고리즘에 사용된 그래프의 개략도이다.
도 7은 하나의 타입의 로드-저장 명령을 구현하는데 사용하기 위한 어드레스 패킹 방식을 개략적으로 도시한다.
도 8은 한 세트의 스트라이드 레지스터들 내의 사전 결정된 스트라이드 값들의 배열을 개략적으로 도시한다.
도 9는 데이터의 입력 볼륨을 갖는 3D 커널(K)의 컨볼루션을 개략적으로 도시한다.
도 10은 누적 행렬 곱셈 명령의 페이즈(페이즈)들의 시퀀스에 의해 수행되는 행렬 곱셈을 개략적으로 도시한다.,
도 11은 누적 행렬 곱셈 명령의 연산을 더 도시한다.
도 12는 컨볼루션을 수행하도록 배열된 누산 행렬 곱셈 명령의 시퀀스의 일련의 루프의 예를 제공한다.
도 13은 컨볼루션 명령의 연산을 개략적으로 도시한다.
도 14는 일련의 컨볼루션 명령의 예를 보여 준다.
Claims (26)
- 실행 유닛, 메모리, 및 복수의 레지스터를 포함하는 하나 이상의 레지스터 파일을 포함하는 프로세서로서, 상기 실행 유닛은 각각이 오피코드 및 0 이상의 오퍼랜드로 구성된 기계 코드 명령의 타입들을 정의하는 명령 세트로부터 명령들의 인스턴스를 실행하도록 구성되고;
상기 실행 유닛은 로드-저장 유닛을 포함하고, 상기 명령 세트에 정의된 명령의 타입들은 하나 이상의 레지스터 파일 중 적어도 하나 내의 레지스터들 중에서, 2개의 로드 연산 각각의 개별 대상(respective destination), 저장 연산의 개별 소스 및 3개의 메모리 어드레스를 유지하도록 배열된 한 쌍의 어드레스 레지스터를 지정하는 오퍼랜드들을 갖는 로드-저장 명령을 포함하고, 상기 3개의 메모리 어드레스는 2개의 로드 연산 각각에 대한 개별 로드 어드레스 및 상기 저장 연산에 대한 각각의 저장 어드레스이고;
상기 로드-저장 명령은 2개의 로드 어드레스와 하나의 저장 어드레스 각각에 대한 개별 스트라이드(stride) 값을 각각 지정하는 3개의 즉치(immediate) 스트라이드 오퍼랜드를 더 포함하고, 상기 즉치 스트라이드 오퍼랜드 각각의 적어도 일부 가능한 값은 하나 이상의 레지스터 파일 중 하나의 스트라이드 레지스터 내의 복수의 필드 중 하나를 지정함으로써 개별 스트라이드 값을 지정하고, 상기 각 필드는 상이한 스트라이드 값을 유지하고; 그리고
상기 로드-저장 유닛은 로드-저장 명령의 오피코드에 응답하여, 상기 2개의 로드 에드레스 각각으로부터 메모리의 데이터의 개별 부분을 개별 로드 연산의 각각의 대상으로 로드하고, 상기 저장 연산의 소스로부터의 데이터의 개별 부분을 메모리의 저장 어드레스에 저장하고, 상기 각 로드 및 저장 연산 다음에 상기 개별 스트라이드 값만큼 각각의 어드레스를 증가시키도록 구성되는 것을 특징으로 하는 프로세서. - 제1항에 있어서,
상기 로드-저장 명령은,
상기 하나 이상의 레지스터 파일 중 하나내의 복수의 가능한 레지스터 중에서 스트라이드 레지스터를 지정하기 위한 스트라이드 레지스터 오퍼랜드를 더 포함하는 것을 특징으로 하는 프로세서. - 제1항 또는 제2항에 있어서,
상기 스트라이드 오퍼랜드의 하나의 가능한 값은 1 단위(unit)의 스트라이드 값을 지정하고, 상기 스트라이드 오퍼랜드의 복수의 다른 가능한 값은 스트라이드 레지스터 내의 상기 필드들의 상이한 것들을 지정하는 것을 특정하는 프로세서. - 제1항 또는 제2항에 있어서,
상기 레지스터 파일 또는 상기 어드레스 레지스터 및 스트라이드 레지스터의 파일들로부터 상기 로드-저장 유닛에 3개의 포트를 포함하고,
상기 로드-저장 유닛은 상기 스트라이드 레지스터에 액세스하기 위해 한 쌍의 어드레스 레지스터 각각 및 상기 포트들 중 하나에 대한 각각의 포트를 사용하도록 구성되는 것을 특정하는 프로세서. - 제4항에 있어서,
상기 3개의 각 포트는 액세스하는데 사용되는 각각의 어드레스 레지스터 또는 스트라이드 레지스터의 비트 폭과 동일한 비트 폭을 갖는 것을 특징으로 하는 프로세서. - 제1항 또는 제2항에 있어서,
상기 한 쌍의 어드레스 레지스터 각각은 32비트 폭이고, 상기 로드 및 저장 어드레스 각각은 21비트 폭인 것을 특징으로 하는 프로세서. - 제1항 또는 제2항에 있어서,
상기 명령 세트에 정의된 명령의 타입들은 하나 이상의 레지스터 파일 중 적어도 하나 내의 레지스터들 중에서, 제1 입력 및 제2 입력을 수신하는 소스들과 결과를 출력할 대상을 지정하는 오퍼랜드들을 취하는 산술 명령을 더 포함하고; 그리고
상기 프로세서는 로드-저장 명령의 인스턴스들과 산술 명령의 인스턴스들을 포함하는 일련의 명령들을 포함하는 프로그램을 실행하도록 프로그래밍되고, 상기 로드-저장 명령들의 적어도 일부의 소스는 산술 명령의 적어도 일부의 대상으로서 설정되고, 상기 로드-저장 명령들의 적어도 일부의 대상은 산술 명령의 적어도 일부의 소스로서 설정되는 것을 특징으로 하는 프로세서. - 제7항에 있어서,
상기 일련은 각각의 명령 쌍이 로드-저장 명령의 인스턴스 및 산술 명령의 대응하는 인스턴스로 구성된 일련의 명령 쌍들을 포함하고; 그리고
각 명령 쌍에서, 상기 로드-저장 명령의 소스는 상기 쌍들 중 하나의 선행 쌍에서 산술 명령의 대상으로서 설정되고, 상기 로드-저장 명령의 대상은 상기 쌍들 중 하나의 현재 또는 후속 쌍에서 산술 명령의 소스로서 설정되는 것을 특징으로 하는 프로세서. - 제8항에 있어서,
입력 및 결과 각각은 적어도 하나의 부동 소수점 값을 포함하고, 상기 실행 유닛은 산술 명령의 오피코드에 응답하여 산술 연산을 수행하도록 구성된 부동 소수점 산술 유닛을 포함하는 것을 특징으로 하는 프로세서. - 제8항에 있어서,
상기 산술 명령은,
벡터 내적 명령, 누적 벡터 내적 명령, 행렬 곱 명령, 누적 행렬 곱 명령 또는 컨벌루션 명령인 것을 특징으로 하는 프로세서. - 제9항에 있어서,
상기 명령 쌍들 각각은 동시에 실행될 명령 번들이고; 그리고
상기 프로세서는 첫 번째는 각 번들의 로드-저장 명령의 인스턴스를 실행하도록 배열된 로드-저장 유닛을 포함하고, 두 번째는 병렬로 산술 명령의 대응하는 인스턴스를 실행하도록 배열된 부동 소수점 산술 유닛을 포함하는 2개의 병렬 파이프 라인으로 분할되는 것을 특징으로 하는 프로세서. - 제8항에 있어서,
상기 산술 명령은 N-엘리먼트 입력 벡터에 각각이 N-엘리먼트 벡터인 M 개의 커널의 M×N 행렬을 곱하는 누적 행렬 곱 명령이고, 상기 누적 행렬 곱 명령은 N1 개의 연속 페이즈(phase)에서 하나의 페이즈를 지정하는 즉치 오퍼랜드를 더 취하고;
상기 일련의 명령은 루프내에서 반복되는 시퀀스를 포함하고, 각 루프 내의 시퀀스는 상기 명령 쌍들의 N1 개의 시퀀스를 포함하고, 상기 시퀀스 내의 각각의 연속 쌍에서 누적 행렬 곱 명령의 인스턴스는 초기 페이즈에서 N1 번째 페이즈까지 상기 시퀀스의 상이한 연속 페이즈를 지정하는 페이즈 오퍼랜드의 상이한 값을 가지며;
각 페이즈에서 상기 제1 입력은 입력 벡터의 각각의 서브 벡터이고 상기 제2 입력은 하나 이상의 부분 합의 각각의 세트이고, 각 입력 서브 벡터내의 엘리먼트의 수(N2)는 N/N1이고, 각 세트 내의 부분 합의 수(Np)는 M/N1이며;
각 루프의 각 페이즈에서의 상기 산술 연산은,
- 상기 누적 행렬 곱 명령의 각각의 인스턴스의 상기 제2 입력으로부터의 Np개의 부분 합의 각각의 세트를 다음 루프에서 사용하기 위해 임시 전파기 상태 (temporary propagator state)로 복사하는 단계와;
- 상기 M 개의 커널 각각에 대해, 상기 현재 페이즈의 상기 각각의 입력 서브 벡터와 상기 커널의 각각의 N2-엘리먼트 서브 벡터와의 내적을 수행하여, 상기 M 개의 커널 각각에 대한 대응하는 중간 결과를 생성하는 단계와;
- 상기 페이즈 오퍼랜드가 초기 페이즈를 지정하면, 상기 이전 루프로부터의 상기 대응하는 부분 곱을 상기 M 개의 중간 결과 각각에 가산하는 단계와;
- 상기 대응하는 중간 결과를 누산기 상태의 M 개의 엘리먼트 각각에 가산하는 단계와;
- 상기 누적 행렬 곱 명령의 각각의 인스턴스의 대상에 출력 결과로서 현재 또는 이전 루프로부터 상기 누산기 상태의 Np 개의 엘리먼트의 각각의 서브셋을 출력하는 단계를 포함하는 것을 특징으로 하는 프로세서. - 제12항에 있어서,
상기 프로그램은 유용한 출력 결과를 생성하기 전에 상기 루프들 중 적어도 2개의 워밍-업(warm-up) 기간을 포함하는 것을 특징으로 하는 프로세서. - 제12항에 있어서,
상기 M=8, N=16, N1=4, N2=4 및 Np=2인 것을 특징으로 하는 프로세서. - 제12항에 있어서,
컨볼루션을 수행하기 위해 루프된 시퀀스를 사용하도록 프로그래밍되고, 각 루프내의 입력 벡터는 입력 데이터의 일부의 상이한 샘플을 나타내고, 이전 루프들의 결과는 이후 루프들의 부분 합으로 사용되는 것을 특징으로 하는 프로세서. - 제15항에 있어서,
상기 입력 데이터의 일부는 3D 데이터 볼륨을 포함하고, 상기 M 개의 커널 각각은 M 개의 더 큰 3D 커널들 중 대응하는 하나의 1D 구성 부분을 나타내고, 상기 이전 루프들의 결과는 각 3D 커널과 대응 구성 1D 커널의 입력 데이터와의 컨볼루션을 설정하도록 이후 루프들의 부분 합으로 사용되는 것을 특징으로 하는 프로세서. - 제16항에 있어서,
상기 3D 커널들 각각은 신경망에서 노드의 가중치 세트를 나타내는 것을 특징으로 하는 프로세서. - 제15항에 있어서,
상기 M 개의 각 값은 입력 데이터 및 상이한 특징의 컨볼루션을 나타내는 것을 특징으로 하는 프로세서. - 제12항에 있어서,
상기 입력 벡터의 각 엘리먼트는 16 비트 부동 소수점 값이고, 각 커널의 각 엘리먼트는 16 비트 부동 소수점 값인 것을 특징으로 하는 프로세서. - 제12항에 있어서,
상기 Np 개의 출력 결과 각각은 32 비트 부동 소수점 값인 것을 특징으로 하는 프로세서. - 제1항 또는 제2항에 있어서,
상기 레지스터 파일들은 제1 레지스터 파일 및 별개의 제2 레지스터 파일을 포함하고, 상기 어드레스 레지스터들은 제1 레지스터 파일의 레지스터이고, 상기 로드-저장 명령의 소스 및 대상은 제2 레지스터 파일의 레지스터인 것을 특징으로 하는 프로세서. - 제21항에 있어서,
상기 스트라이드 레지스터는 제1 파일의 레지스터인 것을 특징으로 하는 프로세서. - 제8항에 있어서,
상기 레지스터 파일들은 제1 레지스터 파일 및 별개의 제2 레지스터 파일을 포함하고, 상기 어드레스 레지스터들은 제1 레지스터 파일의 레지스터이고, 상기 로드-저장 명령의 소스 및 대상은 제2 레지스터 파일의 레지스터이고; 그리고
상기 산술 명령의 소스 및 대상은 제2 레지스터 파일의 레지스터인 것을 특징으로 하는 프로세서. - 제12항에 있어서,
상기 레지스터 파일들은 제1 레지스터 파일 및 별개의 제2 레지스터 파일을 포함하고, 상기 어드레스 레지스터들은 제1 레지스터 파일의 레지스터이고, 상기 로드-저장 명령의 소스 및 대상은 제2 레지스터 파일의 레지스터들이고; 그리고
상기 레지스터 파일들은 상기 제1 및 제2 레지스터 파일과 별개인 제3 레지스터 파일을 더 포함하고, 상기 커널은 제3 파일에 유지되는 것을 특징으로 하는 프로세서. - 임의의 선행하는 항들 중 어느 한 항의 프로세서상에서 실행되도록 구성된 코드를 포함하고, 상기 코드는 로드-저장 명령의 하나 이상의 인스턴스를 포함하는 것을 특징으로 하는 컴퓨터 판독 가능 매체에 구현된 컴퓨터 프로그램.
- 제1항 내지 제24항 중 어느 한 항에 따라 구성된 프로세서를 동작시키는 방법으로서, 상기 방법은 실행 유닛을 통해 프로세서상에서 로드-저장 명령의 하나 이상의 인스턴스를 포함하는 프로그램을 실행하는 단계를 포함하는 것을 특징으로 하는 방법.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GB1821300.9A GB2584268B (en) | 2018-12-31 | 2018-12-31 | Load-Store Instruction |
| GB1821300.9 | 2018-12-31 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| KR20200083123A true KR20200083123A (ko) | 2020-07-08 |
| KR102201935B1 KR102201935B1 (ko) | 2021-01-12 |
Family
ID=65364673
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| KR1020190056970A Active KR102201935B1 (ko) | 2018-12-31 | 2019-05-15 | 로드-저장 명령 |
Country Status (8)
| Country | Link |
|---|---|
| US (1) | US11467833B2 (ko) |
| JP (1) | JP6944974B2 (ko) |
| KR (1) | KR102201935B1 (ko) |
| CN (1) | CN111381880B (ko) |
| CA (1) | CA3040794C (ko) |
| DE (1) | DE102019112353A1 (ko) |
| FR (1) | FR3091375B1 (ko) |
| GB (1) | GB2584268B (ko) |
Families Citing this family (22)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10671349B2 (en) * | 2017-07-24 | 2020-06-02 | Tesla, Inc. | Accelerated mathematical engine |
| US11157441B2 (en) | 2017-07-24 | 2021-10-26 | Tesla, Inc. | Computational array microprocessor system using non-consecutive data formatting |
| US11893393B2 (en) | 2017-07-24 | 2024-02-06 | Tesla, Inc. | Computational array microprocessor system with hardware arbiter managing memory requests |
| US11409692B2 (en) | 2017-07-24 | 2022-08-09 | Tesla, Inc. | Vector computational unit |
| US11561791B2 (en) | 2018-02-01 | 2023-01-24 | Tesla, Inc. | Vector computational unit receiving data elements in parallel from a last row of a computational array |
| GB2580664B (en) | 2019-01-22 | 2021-01-13 | Graphcore Ltd | Double load instruction |
| US12141915B2 (en) * | 2020-06-26 | 2024-11-12 | Advanced Micro Devices, Inc. | Load instruction for multi sample anti-aliasing |
| US12112167B2 (en) | 2020-06-27 | 2024-10-08 | Intel Corporation | Matrix data scatter and gather between rows and irregularly spaced memory locations |
| EP3979248A1 (en) * | 2020-09-30 | 2022-04-06 | Imec VZW | A memory macro |
| US11614891B2 (en) * | 2020-10-20 | 2023-03-28 | Micron Technology, Inc. | Communicating a programmable atomic operator to a memory controller |
| CN114968358B (zh) * | 2020-10-21 | 2025-04-25 | 上海壁仞科技股份有限公司 | 配置向量运算系统中的协作线程束的装置和方法 |
| US12474928B2 (en) * | 2020-12-22 | 2025-11-18 | Intel Corporation | Processors, methods, systems, and instructions to select and store data elements from strided data element positions in a first dimension from three source two-dimensional arrays in a result two-dimensional array |
| US12141683B2 (en) * | 2021-04-30 | 2024-11-12 | Intel Corporation | Performance scaling for dataflow deep neural network hardware accelerators |
| US12400108B2 (en) | 2021-06-17 | 2025-08-26 | Samsung Electronics Co., Ltd. | Mixed-precision neural network accelerator tile with lattice fusion |
| US12079630B2 (en) * | 2021-06-28 | 2024-09-03 | Silicon Laboratories Inc. | Array processor having an instruction sequencer including a program state controller and loop controllers |
| GB202112803D0 (en) * | 2021-09-08 | 2021-10-20 | Graphcore Ltd | Processing device using variable stride pattern |
| CN114090079B (zh) * | 2021-11-16 | 2023-04-21 | 海光信息技术股份有限公司 | 串操作方法、串操作装置以及存储介质 |
| US20230315459A1 (en) * | 2022-04-02 | 2023-10-05 | Intel Corporation | Synchronous microthreading |
| US20230315444A1 (en) * | 2022-04-02 | 2023-10-05 | Intel Corporation | Synchronous microthreading |
| GB2617829B (en) * | 2022-04-13 | 2024-07-10 | Advanced Risc Mach Ltd | Technique for handling data elements stored in an array storage |
| CN115525338A (zh) * | 2022-09-20 | 2022-12-27 | 平头哥(上海)半导体技术有限公司 | 处理器核、处理器、片上系统、计算装置和指令处理方法 |
| CN116126252B (zh) * | 2023-04-11 | 2023-08-08 | 南京砺算科技有限公司 | 数据加载方法及图形处理器、计算机可读存储介质 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2010044242A1 (ja) * | 2008-10-14 | 2010-04-22 | 国立大学法人奈良先端科学技術大学院大学 | データ処理装置 |
| WO2013048369A1 (en) * | 2011-09-26 | 2013-04-04 | Intel Corporation | Instruction and logic to provide vector load-op/store-op with stride functionality |
| JP2016040737A (ja) * | 2011-04-01 | 2016-03-24 | インテル・コーポレーション | 装置および方法 |
| WO2017021678A1 (en) * | 2015-07-31 | 2017-02-09 | Arm Limited | An apparatus and method for transferring a plurality of data structures between memory and a plurality of vector registers |
Family Cites Families (12)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US3417375A (en) * | 1966-03-25 | 1968-12-17 | Burroughs Corp | Circuitry for rotating fields of data in a digital computer |
| EP0442116A3 (en) * | 1990-02-13 | 1993-03-03 | Hewlett-Packard Company | Pipeline method and apparatus |
| US6795908B1 (en) * | 2000-02-16 | 2004-09-21 | Freescale Semiconductor, Inc. | Method and apparatus for instruction execution in a data processing system |
| JP5237722B2 (ja) * | 2008-08-13 | 2013-07-17 | ペンタックスリコーイメージング株式会社 | 撮像装置 |
| WO2012123061A1 (en) * | 2011-02-17 | 2012-09-20 | Hyperion Core Inc. | Parallel memory systems |
| CN103729142B (zh) * | 2012-10-10 | 2016-12-21 | 华为技术有限公司 | 内存数据的推送方法及装置 |
| US10491000B2 (en) * | 2015-02-12 | 2019-11-26 | Open Access Technology International, Inc. | Systems and methods for utilization of demand side assets for provision of grid services |
| CN106201913A (zh) * | 2015-04-23 | 2016-12-07 | 上海芯豪微电子有限公司 | 一种基于指令推送的处理器系统和方法 |
| CA2930557C (en) * | 2015-05-29 | 2023-08-29 | Wastequip, Llc | Vehicle automatic hoist system |
| CN105628377A (zh) * | 2015-12-25 | 2016-06-01 | 鼎奇(天津)主轴科技有限公司 | 一种主轴轴向静刚度测试方法及控制系统 |
| US20170249144A1 (en) * | 2016-02-26 | 2017-08-31 | Qualcomm Incorporated | Combining loads or stores in computer processing |
| CN107515004B (zh) * | 2017-07-27 | 2020-12-15 | 台州市吉吉知识产权运营有限公司 | 步长计算装置及方法 |
-
2018
- 2018-12-31 GB GB1821300.9A patent/GB2584268B/en active Active
-
2019
- 2019-02-15 US US16/276,872 patent/US11467833B2/en active Active
- 2019-04-23 CA CA3040794A patent/CA3040794C/en active Active
- 2019-05-10 DE DE102019112353.4A patent/DE102019112353A1/de active Pending
- 2019-05-15 KR KR1020190056970A patent/KR102201935B1/ko active Active
- 2019-06-07 FR FR1906123A patent/FR3091375B1/fr active Active
- 2019-06-19 JP JP2019113316A patent/JP6944974B2/ja active Active
- 2019-06-25 CN CN201910559688.XA patent/CN111381880B/zh active Active
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2010044242A1 (ja) * | 2008-10-14 | 2010-04-22 | 国立大学法人奈良先端科学技術大学院大学 | データ処理装置 |
| JP2016040737A (ja) * | 2011-04-01 | 2016-03-24 | インテル・コーポレーション | 装置および方法 |
| WO2013048369A1 (en) * | 2011-09-26 | 2013-04-04 | Intel Corporation | Instruction and logic to provide vector load-op/store-op with stride functionality |
| WO2017021678A1 (en) * | 2015-07-31 | 2017-02-09 | Arm Limited | An apparatus and method for transferring a plurality of data structures between memory and a plurality of vector registers |
Also Published As
| Publication number | Publication date |
|---|---|
| CA3040794C (en) | 2022-08-16 |
| JP6944974B2 (ja) | 2021-10-06 |
| GB2584268A (en) | 2020-12-02 |
| US20200210187A1 (en) | 2020-07-02 |
| CN111381880A (zh) | 2020-07-07 |
| CA3040794A1 (en) | 2020-06-30 |
| GB2584268B (en) | 2021-06-30 |
| JP2020109604A (ja) | 2020-07-16 |
| FR3091375A1 (fr) | 2020-07-03 |
| FR3091375B1 (fr) | 2024-04-12 |
| GB201821300D0 (en) | 2019-02-13 |
| KR102201935B1 (ko) | 2021-01-12 |
| US11467833B2 (en) | 2022-10-11 |
| DE102019112353A1 (de) | 2020-07-02 |
| CN111381880B (zh) | 2023-07-07 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| KR102201935B1 (ko) | 로드-저장 명령 | |
| CA3040896C (en) | Register files in a multi-threaded processor | |
| US20220197645A1 (en) | Repeat Instruction for Loading and/or Executing Code in a Claimable Repeat Cache a Specified Number of Times | |
| US8412917B2 (en) | Data exchange and communication between execution units in a parallel processor | |
| US11269638B2 (en) | Exposing valid byte lanes as vector predicates to CPU | |
| US7100026B2 (en) | System and method for performing efficient conditional vector operations for data parallel architectures involving both input and conditional vector values | |
| KR102549680B1 (ko) | 벡터 계산 유닛 | |
| CN114503072A (zh) | 用于向量中的区的排序的方法及设备 | |
| US5121502A (en) | System for selectively communicating instructions from memory locations simultaneously or from the same memory locations sequentially to plurality of processing | |
| US20190196825A1 (en) | Vector multiply-add instruction | |
| CN115904501B (zh) | 具有在每个维度上可选择的多维循环寻址的流引擎 | |
| US5083267A (en) | Horizontal computer having register multiconnect for execution of an instruction loop with recurrance | |
| US5036454A (en) | Horizontal computer having register multiconnect for execution of a loop with overlapped code | |
| US5276819A (en) | Horizontal computer having register multiconnect for operand address generation during execution of iterations of a loop of program code | |
| JP2019511056A (ja) | 複素数乗算命令 | |
| US5226128A (en) | Horizontal computer having register multiconnect for execution of a loop with a branch | |
| CN113853591B (zh) | 将预定义填补值插入到向量流中 | |
| Managuli et al. | VLIW processor architectures and algorithm mappings for DSP applications |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PA0109 | Patent application |
St.27 status event code: A-0-1-A10-A12-nap-PA0109 |
|
| PA0201 | Request for examination |
St.27 status event code: A-1-2-D10-D11-exm-PA0201 |
|
| PG1501 | Laying open of application |
St.27 status event code: A-1-1-Q10-Q12-nap-PG1501 |
|
| E902 | Notification of reason for refusal | ||
| PE0902 | Notice of grounds for rejection |
St.27 status event code: A-1-2-D10-D21-exm-PE0902 |
|
| P11-X000 | Amendment of application requested |
St.27 status event code: A-2-2-P10-P11-nap-X000 |
|
| P13-X000 | Application amended |
St.27 status event code: A-2-2-P10-P13-nap-X000 |
|
| E701 | Decision to grant or registration of patent right | ||
| PE0701 | Decision of registration |
St.27 status event code: A-1-2-D10-D22-exm-PE0701 |
|
| GRNT | Written decision to grant | ||
| PR0701 | Registration of establishment |
St.27 status event code: A-2-4-F10-F11-exm-PR0701 |
|
| PR1002 | Payment of registration fee |
St.27 status event code: A-2-2-U10-U11-oth-PR1002 Fee payment year number: 1 |
|
| PG1601 | Publication of registration |
St.27 status event code: A-4-4-Q10-Q13-nap-PG1601 |
|
| PR1001 | Payment of annual fee |
St.27 status event code: A-4-4-U10-U11-oth-PR1001 Fee payment year number: 4 |
|
| PR1001 | Payment of annual fee |
St.27 status event code: A-4-4-U10-U11-oth-PR1001 Fee payment year number: 5 |
|
| PR1001 | Payment of annual fee |
St.27 status event code: A-4-4-U10-U11-oth-PR1001 Fee payment year number: 6 |
|
| U11 | Full renewal or maintenance fee paid |
Free format text: ST27 STATUS EVENT CODE: A-4-4-U10-U11-OTH-PR1001 (AS PROVIDED BY THE NATIONAL OFFICE) Year of fee payment: 6 |
