KR20200069353A - 신경 네트워크 가속화를 위한 머신 러닝 런타임 라이브러리 - Google Patents
신경 네트워크 가속화를 위한 머신 러닝 런타임 라이브러리 Download PDFInfo
- Publication number
- KR20200069353A KR20200069353A KR1020207013829A KR20207013829A KR20200069353A KR 20200069353 A KR20200069353 A KR 20200069353A KR 1020207013829 A KR1020207013829 A KR 1020207013829A KR 20207013829 A KR20207013829 A KR 20207013829A KR 20200069353 A KR20200069353 A KR 20200069353A
- Authority
- KR
- South Korea
- Prior art keywords
- neural network
- processing
- packet
- library
- pipeline
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/06—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons
- G06N3/063—Physical realisation, i.e. hardware implementation of neural networks, neurons or parts of neurons using electronic means
-
- G06K9/00986—
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/0464—Convolutional networks [CNN, ConvNet]
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/08—Learning methods
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/10—Interfaces, programming languages or software development kits, e.g. for simulating neural networks
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/94—Hardware or software architectures specially adapted for image or video understanding
- G06V10/955—Hardware or software architectures specially adapted for image or video understanding using specific electronic processors
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N3/00—Computing arrangements based on biological models
- G06N3/02—Neural networks
- G06N3/04—Architecture, e.g. interconnection topology
- G06N3/045—Combinations of networks
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Biomedical Technology (AREA)
- Biophysics (AREA)
- Software Systems (AREA)
- Computing Systems (AREA)
- General Physics & Mathematics (AREA)
- General Health & Medical Sciences (AREA)
- Molecular Biology (AREA)
- Evolutionary Computation (AREA)
- General Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- Mathematical Physics (AREA)
- Computational Linguistics (AREA)
- Artificial Intelligence (AREA)
- Neurology (AREA)
- Multimedia (AREA)
- Advance Control (AREA)
- Complex Calculations (AREA)
Abstract
Description
도 1은 예에 따른, 다중 계층 신경 네트워크를 예시한다.
도 2는 예에 따른, 신경 네트워크 가속기를 신경 네트워크 어플리케이션과 인터페이싱시키기 위한 시스템이다.
도 3은 예에 따른, 신경 네트워크 가속기와 신경 네트워크 어플리케이션 사이의 통신 흐름을 예시한다.
도 4는 예에 따른, 신경 네트워크 가속기에서 실행하기 위해 신경 네트워크 애플리케이션으로부터 수신되는 태스크들을 파이프라이닝하는 것에 대한 플로차트이다.
도 5는 예에 따른, 신경 네트워크 애플리케이션에 의해 제출된 태스크들에 대한 파이프라인을 예시한다.
도 6은 예에 따른, 신경 네트워크 애플리케이션에 의해 제출된 태스크들에 대한 파이프라인의 실행을 조정하는 것에 대한 플로차트이다.
도 7은 예에 따른, 신경 네트워크 애플리케이션에 의해 제출된 태스크들을 파이프라이닝하는 것에 대응하는 타이밍 차트이다.
도 8은 예에 따른 신경 네트워크들을 구현하기 위한 시스템을 묘사하는 블록 다이어그램이다.
도 9는 예에 따른 컴퓨팅 시스템을 묘사하는 블록 다이어그램이다.
도 10은 예에 따른 가속 회로를 묘사하는 블록 다이어그램이다.
도 11은 예에 따른 프로그래밍가능 집적 회로(IC)를 묘사하는 블록 다이어그램이다.
도 12는 예에 따른 프로그래밍가능 IC의 FPGA(field programmable gate array) 구현을 예시한다.
이해를 용이하게 하기 위해, 가능한 경우, 도면들에 공통인 동일한 요소들을 표시하기 위해 동일한 참조 번호들이 사용되었다. 일 예의 요소들이 유리하게는 다른 예들에 포함될 수 있는 것이 생각되고 있다.
Claims (15)
- 신경 네트워크 가속기에 제출된 태스크들을 파이프라이닝하기 위한 방법으로서,
상기 신경 네트워크 가속기에 의해 프로세싱될 제1 태스크를 신경 네트워크 애플리케이션으로부터 수신하는 단계;
하나 이상의 컴퓨터 프로세서를 사용하여, 파이프라인 내의 다수의 스테이지들에 의해 사용되는 정보를 포함하는 패킷을 생성하는 단계;
상기 다수의 스테이지들에서 상기 패킷을 프로세싱하는 단계 - 상기 다수의 스테이지들 중 적어도 하나는 상기 신경 네트워크 가속기를 실행하는 하드웨어 시스템에 대한 호출을 수행하며, 상기 파이프라인은 상기 패킷을 프로세싱하는 것과 병렬로 제2 태스크에 대응하는 적어도 하나의 다른 패킷을 프로세싱함 -; 및
상기 파이프라인을 사용하여 상기 패킷을 프로세싱한 결과들을 상기 신경 네트워크 애플리케이션에 반환하는 단계
를 포함하는, 방법. - 제1항에 있어서,
상기 다수의 스테이지들에서 상기 패킷을 프로세싱하는 단계는:
프리-프로세싱(pre-processing) 스테이지에서 상기 패킷을 프로세싱하는 단계;
상기 프리-프로세싱 스테이지 이후에 발생하는 실행 스테이지에서 상기 패킷을 프로세싱하는 단계 - 상기 하드웨어 시스템에 대한 호출은 상기 실행 스테이지 동안 발생함 -; 및
상기 실행 스테이지 이후의 포스트-프로세싱(post-processing) 스테이지에서 상기 패킷을 프로세싱하는 단계
를 포함하는 것인, 컴퓨터-구현 방법. - 제2항에 있어서,
상기 프리-프로세싱 스테이지에서 상기 패킷을 프로세싱하는 단계는:
제1 태스크에 대응하는 데이터를 상기 신경 네트워크 애플리케이션에 의해 사용되는 제1 포맷으로부터 상기 하드웨어 시스템에 의해 사용되는 제2 포맷으로 변환하는 단계를 포함하고,
상기 포스트-프로세싱 스테이지에서 상기 패킷을 프로세싱하는 단계는,
상기 결과들을 상기 제2 포맷으로부터 상기 제1 포맷으로 변환하는 단계를 포함하는 것인, 컴퓨터-구현 방법. - 제1항에 있어서,
상기 다수의 스테이지들 각각은 다른 스레드(thread)들과 독립적으로 상기 패킷을 프로세싱하는 각자의 스레드를 포함하는 것인, 컴퓨터-구현 방법. - 제1항에 있어서,
상기 신경 네트워크 애플리케이션을 위한 할당된 메모리 블록들을 상기 하드웨어 시스템 내의 상기 신경 네트워크 가속기를 위한 할당된 메모리 블록들에 매핑하는 메모리 맵을 생성하는 단계; 및
상기 신경 네트워크 애플리케이션으로부터 수신되는 제1 메모리 어드레스들을 상기 메모리 맵에 기초하여 상기 하드웨어 시스템 내의 메모리 블록들에 대한 제2 메모리 어드레스들로 변환하는 단계
를 더 포함하는, 컴퓨터-구현 방법. - 제1항에 있어서,
신경 네트워크의 다수의 계층들을 수행하는 데 사용되는 가중치들을 행렬 포맷으로 상기 하드웨어 시스템에게 전송하는 단계; 및
새로운 태스크에 대응하는 상기 가중치들의 서브세트를 식별하는 단계; 및
상기 패킷을 프로세싱할 때 사용될 상기 가중치들의 서브세트를 지시하는 오프셋을 상기 하드웨어 시스템에게 전송하는 단계
를 더 포함하는, 컴퓨터-구현 방법. - 제1항에 있어서,
상기 하드웨어 시스템 상의 상기 신경 네트워크 가속기의 실행에 관한 메트릭(metric)을 획득하는 단계;
상기 메트릭의 시각적 표현을 디스플레이를 위해 출력하는 단계; 및
상기 신경 네트워크 가속기의 이용률을 증가시키기 위해 상기 파이프라인 내의 상기 다수의 스테이지들을 실행하는 하드웨어 자원들을 조정하는 단계
를 더 포함하는, 컴퓨터-구현 방법. - 제1항에 있어서,
상기 파이프라인 내의 상기 다수의 스테이지들은 라이브러리에 정의되어 있고, 상기 라이브러리는 상이한 유형들의 신경 네트워크 애플리케이션들이 상기 파이프라인 내의 상기 다수의 스테이지들을 사용하여 태스크들을 상기 신경 네트워크 가속기에 제출할 수 있게 해주도록 구성되는 애플리케이션 프로그램 인터페이스(application program interface; API)를 포함하는 것인, 컴퓨터-구현 방법. - 컴퓨팅 시스템으로서,
프로세서; 및
라이브러리를 포함하는 메모리
를 포함하며, 상기 라이브러리는, 상기 프로세서에 의해 실행될 때, 동작을 수행하고, 상기 동작은:
신경 네트워크 가속기에 의해 프로세싱될 제1 태스크를 신경 네트워크 애플리케이션으로부터 수신하는 것;
하나 이상의 컴퓨터 프로세서를 사용하여, 파이프라인 내의 다수의 스테이지들에 의해 사용되는 정보를 포함하는 패킷을 생성하는 것;
상기 다수의 스테이지들에서 상기 패킷을 프로세싱하는 것 - 상기 다수의 스테이지들 중 적어도 하나는 상기 신경 네트워크 가속기를 실행하는 하드웨어 시스템에 대한 호출을 수행하며, 상기 파이프라인은 상기 패킷을 프로세싱하는 것과 병렬로 제2 태스크에 대응하는 적어도 하나의 다른 패킷을 프로세싱함 -; 및
상기 파이프라인을 사용하여 상기 패킷을 프로세싱한 결과들을 상기 신경 네트워크 애플리케이션에 반환하는 것을 포함하는, 컴퓨팅 시스템. - 제9항에 있어서,
상기 다수의 스테이지들에서 상기 패킷을 프로세싱하는 동작은:
프리-프로세싱 스테이지에서 상기 패킷을 프로세싱하는 동작;
상기 프리-프로세싱 스테이지 이후에 발생하는 실행 스테이지에서 상기 패킷을 프로세싱하는 동작 - 상기 하드웨어 시스템에 대한 호출은 상기 실행 스테이지 동안 발생함 - ; 및
상기 실행 스테이지 이후의 포스트-프로세싱 스테이지에서 상기 패킷을 프로세싱하는 동작
을 포함하는 것인, 컴퓨팅 시스템. - 제10항에 있어서,
상기 프리-프로세싱 스테이지에서 상기 패킷을 프로세싱하는 동작은:
제1 태스크에 대응하는 데이터를 상기 신경 네트워크 애플리케이션에 의해 사용되는 제1 포맷으로부터 상기 하드웨어 시스템에 의해 사용되는 제2 포맷으로 변환하는 동작을 포함하고,
상기 포스트-프로세싱 스테이지에서 상기 패킷을 프로세싱하는 동작은:
상기 결과들을 상기 제2 포맷으로부터 상기 제1 포맷으로 변환하는 동작을 포함하는 것인, 컴퓨팅 시스템. - 제9항에 있어서,
상기 동작은:
상기 신경 네트워크 애플리케이션을 위한 할당된 메모리 블록들을 상기 하드웨어 시스템 내의 상기 신경 네트워크 가속기를 위한 할당된 메모리 블록들에 매핑하는 메모리 맵을 생성하는 동작; 및
상기 신경 네트워크 애플리케이션으로부터 수신되는 제1 메모리 어드레스들을 상기 메모리 맵에 기초하여 상기 하드웨어 시스템 내의 메모리 블록들에 대한 제2 메모리 어드레스들로 변환하는 동작
을 더 포함하는 것인, 컴퓨팅 시스템. - 제9항에 있어서,
상기 동작은:
신경 네트워크의 다수의 계층들을 수행하는 데 사용되는 가중치들을 행렬 포맷으로 상기 하드웨어 시스템에게 전송하는 동작; 및
새로운 태스크에 대응하는 상기 가중치들의 서브세트를 식별하는 동작; 및
상기 패킷을 프로세싱할 때 사용될 상기 가중치들의 서브세트를 지시하는 오프셋을 상기 하드웨어 시스템에게 전송하는 동작
을 더 포함하는 것인, 컴퓨팅 시스템. - 제9항에 있어서,
상기 동작은:
상기 하드웨어 시스템 상의 상기 신경 네트워크 가속기의 실행에 관한 메트릭을 획득하는 동작;
상기 메트릭의 시각적 표현을 디스플레이를 위해 출력하는 동작; 및
상기 신경 네트워크 가속기의 이용률을 증가시키기 위해 상기 파이프라인 내의 상기 다수의 스테이지들을 실행하는 하드웨어 자원들을 조정하는 동작
을 더 포함하는 것인, 컴퓨팅 시스템. - 제9항에 있어서,
상기 파이프라인 내의 상기 다수의 스테이지들은 상기 라이브러리에 정의되어 있고, 상기 라이브러리는 상이한 유형들의 신경 네트워크 애플리케이션들이 상기 파이프라인 내의 상기 다수의 스테이지들을 사용하여 태스크들을 상기 신경 네트워크 가속기에 제출할 수 있게 해주도록 구성되는 API를 포함하는 것인, 컴퓨팅 시스템.
Applications Claiming Priority (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US15/785,679 US11694066B2 (en) | 2017-10-17 | 2017-10-17 | Machine learning runtime library for neural network acceleration |
| US15/785,679 | 2017-10-17 | ||
| PCT/US2018/052833 WO2019079008A1 (en) | 2017-10-17 | 2018-09-26 | LEARNING EXECUTION LIBRARY MACHINE FOR NEURONAL NETWORK ACCELERATION |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| KR20200069353A true KR20200069353A (ko) | 2020-06-16 |
| KR102665580B1 KR102665580B1 (ko) | 2024-05-21 |
Family
ID=63858145
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| KR1020207013829A Active KR102665580B1 (ko) | 2017-10-17 | 2018-09-26 | 신경 네트워크 가속화를 위한 머신 러닝 런타임 라이브러리 |
Country Status (6)
| Country | Link |
|---|---|
| US (1) | US11694066B2 (ko) |
| EP (1) | EP3698294B1 (ko) |
| JP (1) | JP7382925B2 (ko) |
| KR (1) | KR102665580B1 (ko) |
| CN (1) | CN111247533B (ko) |
| WO (1) | WO2019079008A1 (ko) |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2022123492A1 (en) * | 2020-12-10 | 2022-06-16 | Coupang Corp. | Systems and methods for processing data for storing in a feature store and for use in machine learning |
| WO2024185925A1 (ko) * | 2023-03-06 | 2024-09-12 | 주식회사 유엑스팩토리 | 컨볼루션 신경망 시스템 |
| KR20240143338A (ko) * | 2023-03-24 | 2024-10-02 | 한국과학기술원 | 다중 신경망 가속을 위한 확장가능 벡터-어레이 이종 가속기 구조 및 스케쥴링 기법 |
| US12613742B2 (en) | 2022-06-03 | 2026-04-28 | Samsung Electronics Co., Ltd. | System and method for distributed processing of large-scale streaming data |
Families Citing this family (60)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US10609119B2 (en) * | 2017-11-03 | 2020-03-31 | Salesforce.Com, Inc. | Simultaneous optimization of multiple TCP parameters to improve download outcomes for network-based mobile applications |
| KR102615443B1 (ko) * | 2018-05-25 | 2023-12-20 | 에스케이하이닉스 주식회사 | 머신 러닝 장치 및 이를 이용한 머신 러닝 시스템 |
| US11277455B2 (en) | 2018-06-07 | 2022-03-15 | Mellanox Technologies, Ltd. | Streaming system |
| US11151769B2 (en) * | 2018-08-10 | 2021-10-19 | Intel Corporation | Graphics architecture including a neural network pipeline |
| US10846201B1 (en) * | 2018-09-21 | 2020-11-24 | Amazon Technologies, Inc. | Performance debug for networks |
| US20200106828A1 (en) * | 2018-10-02 | 2020-04-02 | Mellanox Technologies, Ltd. | Parallel Computation Network Device |
| US11044099B2 (en) * | 2018-12-28 | 2021-06-22 | Intel Corporation | Technologies for providing certified telemetry data indicative of resources utilizations |
| US11687771B2 (en) * | 2019-01-23 | 2023-06-27 | Samsung Electronics Co., Ltd. | Platform for concurrent execution of GPU operations |
| US11625393B2 (en) | 2019-02-19 | 2023-04-11 | Mellanox Technologies, Ltd. | High performance computing system |
| EP3699770B1 (en) | 2019-02-25 | 2025-05-21 | Mellanox Technologies, Ltd. | Collective communication system and methods |
| US11231961B2 (en) * | 2019-05-22 | 2022-01-25 | Fujitsu Limited | Scheduling operations |
| US11175898B2 (en) * | 2019-05-31 | 2021-11-16 | Apple Inc. | Compiling code for a machine learning model for execution on a specialized processor |
| US20210026686A1 (en) * | 2019-07-22 | 2021-01-28 | Advanced Micro Devices, Inc. | Chiplet-integrated machine learning accelerators |
| US11621808B1 (en) | 2019-10-16 | 2023-04-04 | Xilinx, Inc. | Machine learning based methodology for signal waveform, eye diagram, and bit error rate (BER) bathtub prediction |
| WO2021079168A1 (en) * | 2019-10-22 | 2021-04-29 | Mipsology SAS | Multiple locally stored artificial neural network computations |
| US11423303B1 (en) | 2019-11-21 | 2022-08-23 | Xilinx, Inc. | Machine learning based methodology for adaptative equalization |
| US11182314B1 (en) * | 2019-11-27 | 2021-11-23 | Amazon Techaologies, Inc. | Low latency neural network model loading |
| CN110991632B (zh) * | 2019-11-29 | 2023-05-23 | 电子科技大学 | 一种基于fpga的异构神经网络计算加速器设计方法 |
| KR102490539B1 (ko) * | 2019-12-30 | 2023-01-19 | 주식회사 모레 | 딥러닝을 위한 가속기용 프로그램 생성 방법 |
| WO2021137669A1 (ko) | 2019-12-30 | 2021-07-08 | 매니코어소프트주식회사 | 딥러닝을 위한 가속기용 프로그램 생성 방법 |
| US11687778B2 (en) | 2020-01-06 | 2023-06-27 | The Research Foundation For The State University Of New York | Fakecatcher: detection of synthetic portrait videos using biological signals |
| US11750699B2 (en) | 2020-01-15 | 2023-09-05 | Mellanox Technologies, Ltd. | Small message aggregation |
| US11252027B2 (en) | 2020-01-23 | 2022-02-15 | Mellanox Technologies, Ltd. | Network element supporting flexible data reduction operations |
| WO2021177976A1 (en) * | 2020-03-06 | 2021-09-10 | Google Llc | Distributed computing pipeline processing |
| KR102455310B1 (ko) * | 2020-05-08 | 2022-10-18 | 한국전자통신연구원 | 콘볼루션 신경망 양자화 추론 장치 및 방법 |
| JP2021189832A (ja) * | 2020-06-01 | 2021-12-13 | 株式会社日立製作所 | 電子制御装置 |
| US11574249B2 (en) | 2020-06-02 | 2023-02-07 | International Business Machines Corporation | Streamlining data processing optimizations for machine learning workloads |
| US11876885B2 (en) | 2020-07-02 | 2024-01-16 | Mellanox Technologies, Ltd. | Clock queue with arming and/or self-arming features |
| JP7533003B2 (ja) * | 2020-08-11 | 2024-08-14 | コニカミノルタ株式会社 | 情報処理システム、情報処理方法及びプログラム |
| CN112099943B (zh) * | 2020-08-13 | 2024-05-03 | 深圳云天励飞技术股份有限公司 | 内存分配方法及相关设备 |
| CN112101178B (zh) * | 2020-09-10 | 2023-03-24 | 电子科技大学 | 一种辅助盲人感知外界环境的智能soc终端 |
| CN113485762B (zh) * | 2020-09-19 | 2024-07-26 | 广东高云半导体科技股份有限公司 | 用可配置器件卸载计算任务以提高系统性能的方法和装置 |
| US20220101108A1 (en) * | 2020-09-30 | 2022-03-31 | International Business Machines Corporation | Memory-mapped neural network accelerator for deployable inference systems |
| KR20220049294A (ko) | 2020-10-14 | 2022-04-21 | 삼성전자주식회사 | 스케줄러, 스케줄러의 동작 방법 및 이를 포함한 전자 장치 |
| WO2022087811A1 (zh) * | 2020-10-27 | 2022-05-05 | 华为技术有限公司 | 模型推理异常处理方法及装置 |
| US20220147813A1 (en) * | 2020-11-06 | 2022-05-12 | Micron Technology, Inc. | Runtime optimization of computations of an artificial neural network compiled for execution on a deep learning accelerator |
| CN112508188B (zh) * | 2020-12-01 | 2024-06-14 | 北京奇艺世纪科技有限公司 | 一种分布式模型训练系统、方法、装置、设备和存储介质 |
| US20230325648A1 (en) * | 2020-12-10 | 2023-10-12 | Neuronix AI Labs Inc. | Neural networks processing units activation sparsity removal |
| US11556378B2 (en) | 2020-12-14 | 2023-01-17 | Mellanox Technologies, Ltd. | Offloading execution of a multi-task parameter-dependent operation to a network device |
| CN116569192A (zh) * | 2020-12-21 | 2023-08-08 | 日立数据管理有限公司 | 自学习分析解决方案核心 |
| CN112580787B (zh) * | 2020-12-25 | 2023-11-17 | 北京百度网讯科技有限公司 | 神经网络加速器的数据处理方法、装置、设备及存储介质 |
| CN114764372B (zh) * | 2021-01-15 | 2025-06-06 | 阿里巴巴集团控股有限公司 | 数据处理方法、装置、电子设备和存储介质 |
| US20220309314A1 (en) * | 2021-03-24 | 2022-09-29 | Qualcomm Incorporated | Artificial Intelligence Processor Architecture For Dynamic Scaling Of Neural Network Quantization |
| CN115222015A (zh) | 2021-04-21 | 2022-10-21 | 阿里巴巴新加坡控股有限公司 | 指令处理装置、加速单元和服务器 |
| WO2022235251A1 (en) * | 2021-05-03 | 2022-11-10 | Google Llc | Generating and globally tuning application-specific machine learning accelerators |
| US12041252B2 (en) * | 2021-06-07 | 2024-07-16 | Sony Interactive Entertainment Inc. | Multi-threaded CABAC decoding |
| US11829279B2 (en) | 2021-09-23 | 2023-11-28 | Intel Corporation | Systems, apparatus, and methods to debug accelerator hardware |
| WO2023062443A1 (en) * | 2021-10-14 | 2023-04-20 | University Of Moratuwa | A system and method for evaluating convolutional neural networks |
| US12058016B2 (en) * | 2022-01-27 | 2024-08-06 | Avago Technologies International Sales Pte. Limited | Feature extraction for inline network analysis |
| US12542788B2 (en) | 2022-01-27 | 2026-02-03 | Avago Technologies International Sales Pte. Limited | Deep learning based system and method for inline network analysis |
| CN116894467A (zh) * | 2022-03-29 | 2023-10-17 | 延世大学校 产学协力团 | 用于优化数据处理的深度神经网络加速器及其控制方法 |
| US12309070B2 (en) | 2022-04-07 | 2025-05-20 | Nvidia Corporation | In-network message aggregation for efficient small message transport |
| JP7621557B2 (ja) * | 2022-04-26 | 2025-01-24 | 三菱電機株式会社 | 情報処理装置、推論システム、及び制御方法 |
| CN117217067A (zh) * | 2022-05-31 | 2023-12-12 | 北京有竹居网络技术有限公司 | 仿真装置、仿真系统及其仿真方法、存储介质 |
| US11922237B1 (en) | 2022-09-12 | 2024-03-05 | Mellanox Technologies, Ltd. | Single-step collective operations |
| CN116050499B (zh) * | 2023-04-03 | 2023-07-18 | 合肥综合性国家科学中心人工智能研究院(安徽省人工智能实验室) | 一种模型并行训练中的自适应模型划分方法、系统及设备 |
| US12489657B2 (en) | 2023-08-17 | 2025-12-02 | Mellanox Technologies, Ltd. | In-network compute operation spreading |
| CN116962176B (zh) * | 2023-09-21 | 2024-01-23 | 浪潮电子信息产业股份有限公司 | 一种分布式集群的数据处理方法、装置、系统及存储介质 |
| JPWO2025134467A1 (ko) * | 2023-12-21 | 2025-06-26 | ||
| CN118627552B (zh) * | 2024-06-17 | 2025-03-25 | 北京海云捷迅科技股份有限公司 | Cnn加速框架以及cnn加速方法 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20100076915A1 (en) * | 2008-09-25 | 2010-03-25 | Microsoft Corporation | Field-Programmable Gate Array Based Accelerator System |
| KR20170109600A (ko) * | 2015-01-29 | 2017-09-29 | 뉴엣지 인코포레이티드 | 칩 컴퓨팅 시스템 상의 네트워크에서 프로세서들에 프로세스들의 매핑 |
Family Cites Families (17)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6346825B1 (en) | 2000-10-06 | 2002-02-12 | Xilinx, Inc. | Block RAM with configurable data width and parity for use in a field programmable gate array |
| US6947047B1 (en) * | 2001-09-20 | 2005-09-20 | Nvidia Corporation | Method and system for programmable pipelined graphics processing with branching instructions |
| US7219085B2 (en) | 2003-12-09 | 2007-05-15 | Microsoft Corporation | System and method for accelerating and optimizing the processing of machine learning techniques using a graphics processing unit |
| US20090125706A1 (en) | 2007-11-08 | 2009-05-14 | Hoover Russell D | Software Pipelining on a Network on Chip |
| JP2011034190A (ja) | 2009-07-30 | 2011-02-17 | Renesas Electronics Corp | データ処理装置 |
| WO2013003532A1 (en) | 2011-06-29 | 2013-01-03 | Verisign, Inc. | Data plane packet processing tool chain |
| US9600288B1 (en) * | 2011-07-18 | 2017-03-21 | Apple Inc. | Result bypass cache |
| US9153230B2 (en) | 2012-10-23 | 2015-10-06 | Google Inc. | Mobile speech recognition hardware accelerator |
| US9710749B2 (en) * | 2013-09-03 | 2017-07-18 | Qualcomm Incorporated | Methods and apparatus for implementing a breakpoint determination unit in an artificial nervous system |
| US9529628B2 (en) | 2014-03-21 | 2016-12-27 | Vmware, Inc. | Binary editing of applications executed by virtual machines |
| EP3218827B1 (en) | 2014-11-12 | 2020-05-27 | Xilinx, Inc. | Heterogeneous multiprocessor program compilation targeting programmable integrated circuits |
| EP3035249B1 (en) * | 2014-12-19 | 2019-11-27 | Intel Corporation | Method and apparatus for distributed and cooperative computation in artificial neural networks |
| US10769533B2 (en) | 2015-09-04 | 2020-09-08 | Baidu Usa Llc | Systems and methods for efficient neural network deployments |
| US10319374B2 (en) | 2015-11-25 | 2019-06-11 | Baidu USA, LLC | Deployed end-to-end speech recognition |
| US10621486B2 (en) | 2016-08-12 | 2020-04-14 | Beijing Deephi Intelligent Technology Co., Ltd. | Method for optimizing an artificial neural network (ANN) |
| CN106650922B (zh) * | 2016-09-29 | 2019-05-03 | 清华大学 | 硬件神经网络转换方法、计算装置、软硬件协作系统 |
| US10175980B2 (en) | 2016-10-27 | 2019-01-08 | Google Llc | Neural network compute tile |
-
2017
- 2017-10-17 US US15/785,679 patent/US11694066B2/en active Active
-
2018
- 2018-09-26 KR KR1020207013829A patent/KR102665580B1/ko active Active
- 2018-09-26 WO PCT/US2018/052833 patent/WO2019079008A1/en not_active Ceased
- 2018-09-26 CN CN201880067685.8A patent/CN111247533B/zh active Active
- 2018-09-26 EP EP18786569.6A patent/EP3698294B1/en active Active
- 2018-09-26 JP JP2020521369A patent/JP7382925B2/ja active Active
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20100076915A1 (en) * | 2008-09-25 | 2010-03-25 | Microsoft Corporation | Field-Programmable Gate Array Based Accelerator System |
| KR20170109600A (ko) * | 2015-01-29 | 2017-09-29 | 뉴엣지 인코포레이티드 | 칩 컴퓨팅 시스템 상의 네트워크에서 프로세서들에 프로세스들의 매핑 |
Non-Patent Citations (2)
| Title |
|---|
| M. Bettoni 등. "A Convolutional Neural Network Fully Implemented on FPGA for Embedded Platforms". 2017 New Generation of CAS. IEEE* * |
| V. A. Gokhale. "Nn-X - a hardware accelerator for convolutional neural networks". A Thesis submitted to the faculty of Purdue University.* * |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2022123492A1 (en) * | 2020-12-10 | 2022-06-16 | Coupang Corp. | Systems and methods for processing data for storing in a feature store and for use in machine learning |
| US12591790B2 (en) | 2020-12-10 | 2026-03-31 | Coupang Corp. | Systems and methods for processing data for storing in a feature store and for use in machine learning |
| US12613742B2 (en) | 2022-06-03 | 2026-04-28 | Samsung Electronics Co., Ltd. | System and method for distributed processing of large-scale streaming data |
| WO2024185925A1 (ko) * | 2023-03-06 | 2024-09-12 | 주식회사 유엑스팩토리 | 컨볼루션 신경망 시스템 |
| KR20240143338A (ko) * | 2023-03-24 | 2024-10-02 | 한국과학기술원 | 다중 신경망 가속을 위한 확장가능 벡터-어레이 이종 가속기 구조 및 스케쥴링 기법 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN111247533A (zh) | 2020-06-05 |
| WO2019079008A1 (en) | 2019-04-25 |
| EP3698294A1 (en) | 2020-08-26 |
| KR102665580B1 (ko) | 2024-05-21 |
| US20190114533A1 (en) | 2019-04-18 |
| EP3698294B1 (en) | 2024-07-17 |
| JP2020537784A (ja) | 2020-12-24 |
| US11694066B2 (en) | 2023-07-04 |
| CN111247533B (zh) | 2024-04-30 |
| JP7382925B2 (ja) | 2023-11-17 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| KR102665580B1 (ko) | 신경 네트워크 가속화를 위한 머신 러닝 런타임 라이브러리 | |
| JP7337053B2 (ja) | 超並列ソフトウェア定義ハードウェアシステムにおける静的ブロックスケジューリング | |
| KR102562715B1 (ko) | 다수의 프로세서들 및 뉴럴 네트워크 가속기를 갖는 뉴럴 네트워크 프로세싱 시스템 | |
| KR102578508B1 (ko) | 호스트 전달되는 병합된 가중치들 및 계층별 명령어들의 패키지를 사용한 뉴럴 네트워크 가속기에 의한 다중 계층 뉴럴 네트워크 프로세싱 | |
| US11429848B2 (en) | Host-directed multi-layer neural network processing via per-layer work requests | |
| US10949328B2 (en) | Data flow graph computation using exceptions | |
| US11204747B1 (en) | Re-targetable interface for data exchange between heterogeneous systems and accelerator abstraction into software instructions | |
| US20190130270A1 (en) | Tensor manipulation within a reconfigurable fabric using pointers | |
| US20190138373A1 (en) | Multithreaded data flow processing within a reconfigurable fabric | |
| US20190057060A1 (en) | Reconfigurable fabric data routing | |
| CN110569019A (zh) | 数值的随机修约 | |
| US9529587B2 (en) | Refactoring data flow applications without source code changes or recompilation | |
| WO2019113021A1 (en) | Tensor manipulation within a reconfigurable fabric using pointers | |
| CN113890515A (zh) | 无毛刺的多路器以及防止毛刺传播 | |
| CN114830135B (en) | Hierarchical partitioning of operators | |
| WO2021079168A1 (en) | Multiple locally stored artificial neural network computations |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PA0105 | International application |
Patent event date: 20200514 Patent event code: PA01051R01D Comment text: International Patent Application |
|
| PG1501 | Laying open of application | ||
| A201 | Request for examination | ||
| PA0201 | Request for examination |
Patent event code: PA02012R01D Patent event date: 20210819 Comment text: Request for Examination of Application |
|
| E902 | Notification of reason for refusal | ||
| PE0902 | Notice of grounds for rejection |
Comment text: Notification of reason for refusal Patent event date: 20230227 Patent event code: PE09021S01D |
|
| E90F | Notification of reason for final refusal | ||
| PE0902 | Notice of grounds for rejection |
Comment text: Final Notice of Reason for Refusal Patent event date: 20230828 Patent event code: PE09021S02D |
|
| E701 | Decision to grant or registration of patent right | ||
| PE0701 | Decision of registration |
Patent event code: PE07011S01D Comment text: Decision to Grant Registration Patent event date: 20240226 |
|
| GRNT | Written decision to grant | ||
| PR0701 | Registration of establishment |
Comment text: Registration of Establishment Patent event date: 20240508 Patent event code: PR07011E01D |
|
| PR1002 | Payment of registration fee |
Payment date: 20240508 End annual number: 3 Start annual number: 1 |
|
| PG1601 | Publication of registration |