WO2020203241A1 - 情報処理方法、プログラム、及び、情報処理装置 - Google Patents

情報処理方法、プログラム、及び、情報処理装置 Download PDF

Info

Publication number
WO2020203241A1
WO2020203241A1 PCT/JP2020/011601 JP2020011601W WO2020203241A1 WO 2020203241 A1 WO2020203241 A1 WO 2020203241A1 JP 2020011601 W JP2020011601 W JP 2020011601W WO 2020203241 A1 WO2020203241 A1 WO 2020203241A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
feature
unit
analysis
information processing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2020/011601
Other languages
English (en)
French (fr)
Inventor
晴義 米川
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Sony Semiconductor Solutions Corp
Original Assignee
Sony Semiconductor Solutions Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Sony Semiconductor Solutions Corp filed Critical Sony Semiconductor Solutions Corp
Priority to MX2021011219A priority Critical patent/MX2021011219A/es
Priority to KR1020217026682A priority patent/KR20210142604A/ko
Priority to EP20784531.4A priority patent/EP3951663B1/en
Priority to US17/442,075 priority patent/US12243320B2/en
Priority to ES20784531T priority patent/ES3010476T3/es
Priority to JP2021511392A priority patent/JP7487178B2/ja
Priority to CA3134088A priority patent/CA3134088A1/en
Publication of WO2020203241A1 publication Critical patent/WO2020203241A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00—Scenes; Scene-specific elements
    • G06V20/50—Context or environment of the image
    • G06V20/56—Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/04—Architecture, e.g. interconnection topology
    • G—PHYSICS
    • G01—MEASURING; TESTING
    • G01S—RADIO DIRECTION-FINDING; RADIO NAVIGATION; DETERMINING DISTANCE OR VELOCITY BY USE OF RADIO WAVES; LOCATING OR PRESENCE-DETECTING BY USE OF THE REFLECTION OR RERADIATION OF RADIO WAVES; ANALOGOUS ARRANGEMENTS USING OTHER WAVES
    • G01S13/00—Systems using the reflection or reradiation of radio waves, e.g. radar systems; Analogous systems using reflection or reradiation of waves whose nature or wavelength is irrelevant or unspecified
    • G01S13/88—Radar or analogous systems specially adapted for specific applications
    • G01S13/89—Radar or analogous systems specially adapted for specific applications for mapping or imaging
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06F—ELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00—Pattern recognition
    • G06F18/20—Analysing
    • G06F18/24—Classification techniques
    • G06F18/241—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches
    • G06F18/2413—Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches based on distances to training or reference patterns
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/04—Architecture, e.g. interconnection topology
    • G06N3/045—Combinations of networks
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/04—Architecture, e.g. interconnection topology
    • G06N3/0464—Convolutional networks [CNN, ConvNet]
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/08—Learning methods
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00—Computing arrangements based on biological models
    • G06N3/02—Neural networks
    • G06N3/08—Learning methods
    • G06N3/09—Supervised learning
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00—Image analysis
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T7/00—Image analysis
    • G06T7/20—Analysis of motion
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/40—Extraction of image or video features
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
    • G06V10/82—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00—Scenes; Scene-specific elements
    • G06V20/50—Context or environment of the image
    • G06V20/56—Context or environment of the image exterior to a vehicle by using sensors mounted on the vehicle
    • G06V20/58—Recognition of moving objects or obstacles, e.g. vehicles or pedestrians; Recognition of traffic objects, e.g. traffic signs, traffic lights or roads
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00—Machine learning
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00—Indexing scheme for image analysis or image enhancement
    • G06T2207/20—Special algorithmic details
    • G06T2207/20081—Training; Learning
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00—Indexing scheme for image analysis or image enhancement
    • G06T2207/20—Special algorithmic details
    • G06T2207/20084—Artificial neural networks [ANN]
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00—Indexing scheme for image analysis or image enhancement
    • G06T2207/30—Subject of image; Context of image processing
    • G06T2207/30248—Vehicle exterior or interior
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
    • G06T2207/00—Indexing scheme for image analysis or image enhancement
    • G06T2207/30—Subject of image; Context of image processing
    • G06T2207/30248—Vehicle exterior or interior
    • G06T2207/30252—Vehicle exterior; Vicinity of vehicle
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V10/00—Arrangements for image or video recognition or understanding
    • G06V10/40—Extraction of image or video features
    • G06V10/44—Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersections; Connectivity analysis, e.g. of connected components
    • G—PHYSICS
    • G06—COMPUTING OR CALCULATING; COUNTING
    • G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V20/00—Scenes; Scene-specific elements
    • G06V20/60—Type of objects
    • G06V20/64—Three-dimensional [3D] objects

Definitions

  • the present technology relates to an information processing method, a program, and an information processing apparatus, and more particularly to an information processing method, a program, and an information processing apparatus suitable for use when analyzing a model using a neural network.
  • This technology was made in view of such a situation, and makes it possible to analyze the learning situation of a model using a neural network.
  • the information processing method of one aspect of the present technology includes a feature data generation step of generating feature data that numerically represents the features of the feature map generated from the input data in a model using a neural network, and the above-mentioned multiple feature maps.
  • a feature data generation step of generating feature data that numerically represents the features of the feature map generated from the input data in a model using a neural network, and the above-mentioned multiple feature maps.
  • the program of one aspect of the present technology includes a feature data generation step of generating feature data that numerically represents the features of the feature map generated from the input data in a model using a neural network, and the feature data of a plurality of the feature maps.
  • the computer is made to execute the process including the analysis data generation step for generating the analysis data based on the above.
  • the information processing device of one aspect of the present technology includes a feature data generation unit that generates feature data that numerically represents the features of the feature map generated from the input data in a model using a neural network, and the above-mentioned multiple feature maps. It is provided with an analysis data generation unit that generates analysis data based on feature data.
  • feature data that numerically represents the features of the feature map generated from the input data in the model using the neural network is generated, and the analysis data based on the feature data of the plurality of the feature maps is generated. Will be generated.
  • FIG. 1 is a block diagram showing a schematic functional configuration example of a vehicle control system 100, which is an example of a mobile control system to which the present technology can be applied.
  • the vehicle 10 provided with the vehicle control system 100 is distinguished from other vehicles, it is referred to as a own vehicle or a own vehicle.
  • the vehicle control system 100 includes an input unit 101, a data acquisition unit 102, a communication unit 103, an in-vehicle device 104, an output control unit 105, an output unit 106, a drive system control unit 107, a drive system system 108, a body system control unit 109, and a body. It includes a system system 110, a storage unit 111, and an automatic operation control unit 112.
  • the input unit 101, the data acquisition unit 102, the communication unit 103, the output control unit 105, the drive system control unit 107, the body system control unit 109, the storage unit 111, and the automatic operation control unit 112 are connected via the communication network 121. They are interconnected.
  • the communication network 121 is, for example, from an in-vehicle communication network or bus that conforms to any standard such as CAN (Controller Area Network), LIN (Local Interconnect Network), LAN (Local Area Network), or FlexRay (registered trademark). Become. In addition, each part of the vehicle control system 100 may be directly connected without going through the communication network 121.
  • CAN Controller Area Network
  • LIN Local Interconnect Network
  • LAN Local Area Network
  • FlexRay registered trademark
  • the description of the communication network 121 shall be omitted.
  • the input unit 101 and the automatic operation control unit 112 communicate with each other via the communication network 121, it is described that the input unit 101 and the automatic operation control unit 112 simply communicate with each other.
  • the input unit 101 includes a device used by the passenger to input various data, instructions, and the like.
  • the input unit 101 includes an operation device such as a touch panel, a button, a microphone, a switch, and a lever, and an operation device capable of inputting by a method other than manual operation by voice or gesture.
  • the input unit 101 may be a remote control device using infrared rays or other radio waves, or an externally connected device such as a mobile device or a wearable device corresponding to the operation of the vehicle control system 100.
  • the input unit 101 generates an input signal based on data, instructions, and the like input by the passenger, and supplies the input signal to each unit of the vehicle control system 100.
  • the data acquisition unit 102 includes various sensors and the like that acquire data used for processing of the vehicle control system 100, and supplies the acquired data to each unit of the vehicle control system 100.
  • the data acquisition unit 102 includes various sensors for detecting the state of the own vehicle and the like.
  • the data acquisition unit 102 includes a gyro sensor, an acceleration sensor, an inertial measurement unit (IMU), an accelerator pedal operation amount, a brake pedal operation amount, a steering wheel steering angle, and an engine speed. It is equipped with a sensor or the like for detecting the rotation speed of the motor or the rotation speed of the wheels.
  • IMU inertial measurement unit
  • the data acquisition unit 102 includes various sensors for detecting information outside the own vehicle.
  • the data acquisition unit 102 includes an imaging device such as a ToF (TimeOfFlight) camera, a stereo camera, a monocular camera, an infrared camera, and other cameras.
  • the data acquisition unit 102 includes an environment sensor for detecting the weather, the weather, and the like, and a surrounding information detection sensor for detecting an object around the own vehicle.
  • the environmental sensor includes, for example, a raindrop sensor, a fog sensor, a sunshine sensor, a snow sensor, and the like.
  • the ambient information detection sensor includes, for example, an ultrasonic sensor, a radar, LiDAR (Light Detection and Ringing, Laser Imaging Detection and Ringing), a sonar, and the like.
  • the data acquisition unit 102 includes various sensors for detecting the current position of the own vehicle.
  • the data acquisition unit 102 includes a GNSS receiver or the like that receives a GNSS signal from a GNSS (Global Navigation Satellite System) satellite.
  • GNSS Global Navigation Satellite System
  • the data acquisition unit 102 includes various sensors for detecting information in the vehicle.
  • the data acquisition unit 102 includes an imaging device that images the driver, a biosensor that detects the driver's biological information, a microphone that collects sound in the vehicle interior, and the like.
  • the biosensor is provided on, for example, the seat surface or the steering wheel, and detects the biometric information of the passenger sitting on the seat or the driver holding the steering wheel.
  • the communication unit 103 communicates with the in-vehicle device 104 and various devices, servers, base stations, etc. outside the vehicle, transmits data supplied from each unit of the vehicle control system 100, and transmits the received data to the vehicle control system. It is supplied to each part of 100.
  • the communication protocol supported by the communication unit 103 is not particularly limited, and the communication unit 103 can also support a plurality of types of communication protocols.
  • the communication unit 103 wirelessly communicates with the in-vehicle device 104 by wireless LAN, Bluetooth (registered trademark), NFC (Near Field Communication), WUSB (Wireless USB), or the like. Further, for example, the communication unit 103 uses a USB (Universal Serial Bus), HDMI (registered trademark) (High-Definition Multimedia Interface), or MHL () via a connection terminal (and a cable if necessary) (not shown). Wired communication is performed with the in-vehicle device 104 by Mobile High-definition Link) or the like.
  • USB Universal Serial Bus
  • HDMI registered trademark
  • MHL Mobility Management Entity
  • the communication unit 103 is connected to a device (for example, an application server or a control server) existing on an external network (for example, the Internet, a cloud network or a network peculiar to a business operator) via a base station or an access point. Communicate. Further, for example, the communication unit 103 uses P2P (Peer To Peer) technology to connect with a terminal (for example, a pedestrian or store terminal, or an MTC (Machine Type Communication) terminal) existing in the vicinity of the own vehicle. Communicate.
  • a device for example, an application server or a control server
  • an external network for example, the Internet, a cloud network or a network peculiar to a business operator
  • the communication unit 103 uses P2P (Peer To Peer) technology to connect with a terminal (for example, a pedestrian or store terminal, or an MTC (Machine Type Communication) terminal) existing in the vicinity of the own vehicle. Communicate.
  • P2P Peer To Peer
  • a terminal for example, a pedestrian or
  • the communication unit 103 includes vehicle-to-vehicle (Vehicle to Vehicle) communication, road-to-vehicle (Vehicle to Infrastructure) communication, vehicle-to-house (Vehicle to Home) communication, and pedestrian-to-vehicle (Vehicle to Pedestrian) communication. ) Perform V2X communication such as communication. Further, for example, the communication unit 103 is provided with a beacon receiving unit, receives radio waves or electromagnetic waves transmitted from a radio station or the like installed on the road, and acquires information such as the current position, traffic congestion, traffic regulation, or required time. To do.
  • the in-vehicle device 104 includes, for example, a mobile device or a wearable device owned by a passenger, an information device carried in or attached to the own vehicle, a navigation device for searching a route to an arbitrary destination, and the like.
  • the output control unit 105 controls the output of various information to the passengers of the own vehicle or the outside of the vehicle.
  • the output control unit 105 generates an output signal including at least one of visual information (for example, image data) and auditory information (for example, audio data) and supplies it to the output unit 106 to supply the output unit 105.
  • the output control unit 105 synthesizes image data captured by different imaging devices of the data acquisition unit 102 to generate a bird's-eye view image, a panoramic image, or the like, and outputs an output signal including the generated image. It is supplied to the output unit 106.
  • the output control unit 105 generates voice data including a warning sound or a warning message for dangers such as collision, contact, and entry into a danger zone, and outputs an output signal including the generated voice data to the output unit 106.
  • Supply for example, the output control unit 105 generates voice data including a warning sound or a warning message for dangers such as collision,
  • the output unit 106 is provided with a device capable of outputting visual information or auditory information to the passengers of the own vehicle or the outside of the vehicle.
  • the output unit 106 includes a display device, an instrument panel, an audio speaker, headphones, a wearable device such as a spectacle-type display worn by a passenger, a projector, a lamp, and the like.
  • the display device included in the output unit 106 displays visual information in the driver's field of view, such as a head-up display, a transmissive display, and a device having an AR (Augmented Reality) display function, in addition to the device having a normal display. It may be a display device.
  • the drive system control unit 107 controls the drive system system 108 by generating various control signals and supplying them to the drive system system 108. Further, the drive system control unit 107 supplies a control signal to each unit other than the drive system system 108 as necessary, and notifies the control state of the drive system system 108.
  • the drive system system 108 includes various devices related to the drive system of the own vehicle.
  • the drive system system 108 includes a drive force generator for generating a drive force of an internal combustion engine or a drive motor, a drive force transmission mechanism for transmitting the drive force to the wheels, a steering mechanism for adjusting the steering angle, and the like. It is equipped with a braking device that generates braking force, ABS (Antilock Brake System), ESC (Electronic Stability Control), an electric power steering device, and the like.
  • the body system control unit 109 controls the body system 110 by generating various control signals and supplying them to the body system 110. Further, the body system control unit 109 supplies control signals to each unit other than the body system 110 as necessary, and notifies the control state of the body system 110.
  • the body system 110 includes various body devices equipped on the vehicle body.
  • the body system 110 includes a keyless entry system, a smart key system, a power window device, a power seat, a steering wheel, an air conditioner, and various lamps (for example, headlamps, back lamps, brake lamps, winkers, fog lamps, etc.).
  • various lamps for example, headlamps, back lamps, brake lamps, winkers, fog lamps, etc.
  • the storage unit 111 includes, for example, a magnetic storage device such as a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disc Drive), a semiconductor storage device, an optical storage device, an optical magnetic storage device, and the like. ..
  • the storage unit 111 stores various programs, data, and the like used by each unit of the vehicle control system 100.
  • the storage unit 111 has map data such as a three-dimensional high-precision map such as a dynamic map, a global map which is less accurate than the high-precision map and covers a wide area, and a local map including information around the own vehicle.
  • map data such as a three-dimensional high-precision map such as a dynamic map, a global map which is less accurate than the high-precision map and covers a wide area, and a local map including information around the own vehicle.
  • the automatic driving control unit 112 controls automatic driving such as autonomous driving or driving support. Specifically, for example, the automatic driving control unit 112 issues collision avoidance or impact mitigation of the own vehicle, follow-up running based on the inter-vehicle distance, vehicle speed maintenance running, collision warning of the own vehicle, lane deviation warning of the own vehicle, and the like. Collision control is performed for the purpose of realizing the functions of ADAS (Advanced Driver Assistance System) including. Further, for example, the automatic driving control unit 112 performs cooperative control for the purpose of automatic driving that autonomously travels without depending on the operation of the driver.
  • the automatic operation control unit 112 includes a detection unit 131, a self-position estimation unit 132, a situation analysis unit 133, a planning unit 134, and an operation control unit 135.
  • the detection unit 131 detects various types of information necessary for controlling automatic operation.
  • the detection unit 131 includes an outside information detection unit 141, an inside information detection unit 142, and a vehicle state detection unit 143.
  • the vehicle outside information detection unit 141 performs detection processing of information outside the own vehicle based on data or signals from each unit of the vehicle control system 100. For example, the vehicle outside information detection unit 141 performs detection processing, recognition processing, tracking processing, and distance detection processing for an object around the own vehicle. Objects to be detected include, for example, vehicles, people, obstacles, structures, roads, traffic lights, traffic signs, road markings, and the like. Further, for example, the vehicle outside information detection unit 141 performs detection processing of the environment around the own vehicle. The surrounding environment to be detected includes, for example, weather, temperature, humidity, brightness, road surface condition, and the like.
  • the vehicle outside information detection unit 141 outputs data indicating the result of the detection process to the self-position estimation unit 132, the map analysis unit 151 of the situation analysis unit 133, the traffic rule recognition unit 152, the situation recognition unit 153, and the operation control unit 135. It is supplied to the emergency situation avoidance unit 171 and the like.
  • the in-vehicle information detection unit 142 performs in-vehicle information detection processing based on data or signals from each unit of the vehicle control system 100.
  • the vehicle interior information detection unit 142 performs driver authentication processing and recognition processing, driver status detection processing, passenger detection processing, vehicle interior environment detection processing, and the like.
  • the state of the driver to be detected includes, for example, physical condition, alertness, concentration, fatigue, gaze direction, and the like.
  • the environment inside the vehicle to be detected includes, for example, temperature, humidity, brightness, odor, and the like.
  • the vehicle interior information detection unit 142 supplies data indicating the result of the detection process to the situational awareness unit 153 of the situational analysis unit 133, the emergency situation avoidance unit 171 of the motion control unit 135, and the like.
  • the vehicle state detection unit 143 performs the state detection process of the own vehicle based on the data or signals from each part of the vehicle control system 100.
  • the states of the vehicle to be detected include, for example, speed, acceleration, steering angle, presence / absence and content of abnormality, driving operation state, power seat position / tilt, door lock state, and other in-vehicle devices. The state etc. are included.
  • the vehicle state detection unit 143 supplies data indicating the result of the detection process to the situation recognition unit 153 of the situation analysis unit 133, the emergency situation avoidance unit 171 of the operation control unit 135, and the like.
  • the self-position estimation unit 132 estimates the position and attitude of the own vehicle based on data or signals from each unit of the vehicle control system 100 such as the vehicle exterior information detection unit 141 and the situation recognition unit 153 of the situation analysis unit 133. Perform processing. In addition, the self-position estimation unit 132 generates a local map (hereinafter, referred to as a self-position estimation map) used for self-position estimation, if necessary.
  • the map for self-position estimation is, for example, a highly accurate map using a technique such as SLAM (Simultaneous Localization and Mapping).
  • the self-position estimation unit 132 supplies data indicating the result of the estimation process to the map analysis unit 151, the traffic rule recognition unit 152, the situation recognition unit 153, and the like of the situation analysis unit 133. Further, the self-position estimation unit 132 stores the self-position estimation map in the storage unit 111.
  • the situation analysis unit 133 analyzes the situation of the own vehicle and the surroundings.
  • the situation analysis unit 133 includes a map analysis unit 151, a traffic rule recognition unit 152, a situation recognition unit 153, and a situation prediction unit 154.
  • the map analysis unit 151 uses data or signals from each unit of the vehicle control system 100 such as the self-position estimation unit 132 and the vehicle exterior information detection unit 141 as necessary, and the map analysis unit 151 of various maps stored in the storage unit 111. Perform analysis processing and build a map containing information necessary for automatic operation processing.
  • the map analysis unit 151 applies the constructed map to the traffic rule recognition unit 152, the situation recognition unit 153, the situation prediction unit 154, the route planning unit 161 of the planning unit 134, the action planning unit 162, the operation planning unit 163, and the like. Supply to.
  • the traffic rule recognition unit 152 determines the traffic rules around the own vehicle based on data or signals from each unit of the vehicle control system 100 such as the self-position estimation unit 132, the vehicle outside information detection unit 141, and the map analysis unit 151. Perform recognition processing. By this recognition process, for example, the position and state of the signal around the own vehicle, the content of the traffic regulation around the own vehicle, the lane in which the vehicle can travel, and the like are recognized.
  • the traffic rule recognition unit 152 supplies data indicating the result of the recognition process to the situation prediction unit 154 and the like.
  • the situation recognition unit 153 can be used for data or signals from each unit of the vehicle control system 100 such as the self-position estimation unit 132, the vehicle exterior information detection unit 141, the vehicle interior information detection unit 142, the vehicle condition detection unit 143, and the map analysis unit 151. Based on this, the situation recognition process related to the own vehicle is performed. For example, the situational awareness unit 153 performs recognition processing such as the situation of the own vehicle, the situation around the own vehicle, and the situation of the driver of the own vehicle. In addition, the situational awareness unit 153 generates a local map (hereinafter, referred to as a situational awareness map) used for recognizing the situation around the own vehicle, if necessary.
  • the situational awareness map is, for example, an occupied grid map (OccupancyGridMap).
  • the status of the own vehicle to be recognized includes, for example, the position, posture, movement (for example, speed, acceleration, moving direction, etc.) of the own vehicle, and the presence / absence and contents of an abnormality.
  • the surrounding conditions of the vehicle to be recognized include, for example, the type and position of the surrounding stationary object, the type, position and movement of the surrounding animal body (for example, speed, acceleration, moving direction, etc.), and the surrounding road.
  • the composition and road surface condition, as well as the surrounding weather, temperature, humidity, brightness, etc. are included.
  • the state of the driver to be recognized includes, for example, physical condition, arousal level, concentration level, fatigue level, eye movement, driving operation, and the like.
  • the situational awareness unit 153 supplies data indicating the result of the recognition process (including a situational awareness map, if necessary) to the self-position estimation unit 132, the situation prediction unit 154, and the like. Further, the situational awareness unit 153 stores the situational awareness map in the storage unit 111.
  • the situation prediction unit 154 performs a situation prediction process related to the own vehicle based on data or signals from each part of the vehicle control system 100 such as the map analysis unit 151, the traffic rule recognition unit 152, and the situation recognition unit 153. For example, the situation prediction unit 154 performs prediction processing such as the situation of the own vehicle, the situation around the own vehicle, and the situation of the driver.
  • the situation of the own vehicle to be predicted includes, for example, the behavior of the own vehicle, the occurrence of an abnormality, the mileage, and the like.
  • the situation around the vehicle to be predicted includes, for example, the behavior of animals around the vehicle, changes in signal conditions, changes in the environment such as weather, and the like.
  • the driver's situation to be predicted includes, for example, the driver's behavior and physical condition.
  • the situation prediction unit 154 together with the data from the traffic rule recognition unit 152 and the situation recognition unit 153, provides the data indicating the result of the prediction processing to the route planning unit 161, the action planning unit 162, and the operation planning unit 163 of the planning unit 134. And so on.
  • the route planning unit 161 plans a route to the destination based on data or signals from each unit of the vehicle control system 100 such as the map analysis unit 151 and the situation prediction unit 154. For example, the route planning unit 161 sets a route from the current position to the specified destination based on the global map. Further, for example, the route planning unit 161 appropriately changes the route based on the conditions of traffic congestion, accidents, traffic restrictions, construction, etc., and the physical condition of the driver. The route planning unit 161 supplies data indicating the planned route to the action planning unit 162 and the like.
  • the action planning unit 162 safely sets the route planned by the route planning unit 161 within the planned time based on the data or signals from each unit of the vehicle control system 100 such as the map analysis unit 151 and the situation prediction unit 154. Plan your vehicle's actions to drive. For example, the action planning unit 162 plans starting, stopping, traveling direction (for example, forward, backward, left turn, right turn, change of direction, etc.), traveling lane, traveling speed, and overtaking. The action planning unit 162 supplies data indicating the planned behavior of the own vehicle to the motion planning unit 163 and the like.
  • the motion planning unit 163 is the operation of the own vehicle for realizing the action planned by the action planning unit 162 based on the data or signals from each unit of the vehicle control system 100 such as the map analysis unit 151 and the situation prediction unit 154. Plan. For example, the motion planning unit 163 plans acceleration, deceleration, traveling track, and the like. The motion planning unit 163 supplies data indicating the planned operation of the own vehicle to the acceleration / deceleration control unit 172 and the direction control unit 173 of the motion control unit 135.
  • the motion control unit 135 controls the motion of the own vehicle.
  • the motion control unit 135 includes an emergency situation avoidance unit 171, an acceleration / deceleration control unit 172, and a direction control unit 173.
  • the emergency situation avoidance unit 171 may collide, contact, enter a danger zone, have a driver abnormality, or cause a vehicle. Performs emergency detection processing such as abnormalities.
  • the emergency situation avoidance unit 171 detects the occurrence of an emergency situation, it plans the operation of the own vehicle to avoid an emergency situation such as a sudden stop or a sharp turn.
  • the emergency situation avoidance unit 171 supplies data indicating the planned operation of the own vehicle to the acceleration / deceleration control unit 172, the direction control unit 173, and the like.
  • the acceleration / deceleration control unit 172 performs acceleration / deceleration control for realizing the operation of the own vehicle planned by the motion planning unit 163 or the emergency situation avoidance unit 171.
  • the acceleration / deceleration control unit 172 calculates a control target value of a driving force generator or a braking device for realizing a planned acceleration, deceleration, or sudden stop, and drives a control command indicating the calculated control target value. It is supplied to the system control unit 107.
  • the direction control unit 173 performs direction control for realizing the operation of the own vehicle planned by the motion planning unit 163 or the emergency situation avoidance unit 171. For example, the direction control unit 173 calculates the control target value of the steering mechanism for realizing the traveling track or the sharp turn planned by the motion planning unit 163 or the emergency situation avoidance unit 171 and controls to indicate the calculated control target value. The command is supplied to the drive system control unit 107.
  • FIG. 2 shows a configuration example of the object recognition model 201.
  • the object recognition model 201 is used, for example, in the vehicle exterior information detection unit 141 of the vehicle control system 100 of FIG.
  • a photographed image (for example, a photographed image 202) which is image data of the front of the vehicle 10 is input to the object recognition model 201 as input data. Then, the object recognition model 201 performs recognition processing of the vehicle in front of the vehicle 10 based on the captured image, and outputs an output image (for example, output image 203) which is image data showing the recognition result as output data.
  • an output image for example, output image 203
  • the object recognition model 201 is a model using a convolutional neural network (CNN), and includes a feature extraction layer 211 and a prediction layer 212.
  • CNN convolutional neural network
  • the feature extraction layer 211 has a plurality of layers, and each layer is composed of a convolution layer, a pooling layer, and the like. Each layer of the feature extraction layer 211 generates one or more feature maps showing the features of the captured image by a predetermined calculation, and supplies them to the next layer. In addition, some layers of the feature extraction layer 211 supply the feature map to the prediction layer 212.
  • the size (number of pixels) of the feature map generated in each layer of the feature extraction layer 211 is different, and gradually decreases as the layer progresses.
  • FIG. 3 shows an example of the feature map 231-1 to the feature map 231-n generated and output in the layer 221 which is the intermediate layer surrounded by the dotted square of the feature extraction layer 211 of FIG. ..
  • n feature maps 231-1 to feature maps 231-n are generated for one captured image.
  • the feature map 231-1 to the feature map 231-n are, for example, image data in which 38 vertical pixels and 38 horizontal pixels are arranged in two dimensions. Further, the feature map 231-1 to the feature map 231-n are the feature maps having the largest size among the feature maps used for the recognition process in the prediction layer 212.
  • serial numbers 1 to n are assigned to feature maps 231-1 to feature maps 231-n generated from one captured image in layer 221.
  • serial numbers 1 to N are assigned to feature maps generated from one captured image in other layers. Note that N indicates the number of feature maps generated from one captured image in the hierarchy.
  • the prediction layer 212 performs recognition processing of the vehicle in front of the vehicle 10 based on the feature map supplied from the feature extraction layer 211.
  • the prediction layer 212 outputs an output image showing the recognition result of the vehicle.
  • FIG. 4 shows a configuration example of an information processing device 301 used for learning a model (hereinafter, referred to as a learning model) using a neural network such as the object recognition model 201 of FIG.
  • the information processing device 301 includes an input unit 311, a learning unit 312, a learning situation analysis unit 313, an output control unit 314, and an output unit 315.
  • the input unit 311 is provided with an input device used for inputting various data and instructions, generates an input signal based on the input data and instructions, and supplies the input signal to the learning unit 312.
  • the input unit 311 is used for inputting teacher data for learning a learning model.
  • the learning unit 312 performs learning processing of the learning model.
  • the learning method of the learning unit 312 is not limited to a specific method. Further, the learning unit 312 supplies the feature map generated in the learning model to the learning situation analysis unit 313.
  • the learning situation analysis unit 313 analyzes the learning situation of the learning model by the learning unit 312 based on the feature map supplied from the learning unit 312.
  • the learning situation analysis unit 313 includes a feature data generation unit 321, an analysis data generation unit 322, an analysis unit 323, and a parameter setting unit 324.
  • the feature data generation unit 321 generates feature data that numerically represents the features of the feature map and supplies it to the analysis data generation unit 322.
  • the analysis data generation unit 322 generates analysis data in which feature data of a plurality of feature maps are arranged (arranged), and supplies the analysis data to the analysis unit 323 and the output control unit 314.
  • the analysis unit 323 performs analysis processing of the learning status of the learning model based on the analysis data, and supplies data indicating the analysis result to the parameter setting unit 324.
  • the parameter setting unit 324 sets various parameters for learning the learning model based on the analysis result of the learning situation of the learning model.
  • the parameter setting unit 324 supplies the learning unit 312 with data indicating the set parameters.
  • the output control unit 314 controls the output of various information by the output unit 315.
  • the output control unit 314 controls the display of analysis data by the output unit 315.
  • the output unit 315 is provided with an output device capable of outputting various information such as visual information and auditory information.
  • the output unit 315 includes a display, a speaker, and the like.
  • This process is started, for example, when the learning process of the learning model is started by the learning unit 312.
  • step S1 the feature data generation unit 321 generates feature data.
  • the feature data generation unit 321 calculates the variance indicating the degree of dispersion of the pixel values of the feature map by the following equation (1).
  • the var of the equation (1) shows the dispersion of the pixel values of the feature map.
  • I indicates the value obtained by subtracting 1 from the number of pixels in the column direction (vertical direction) of the feature map
  • J indicates the value obtained by subtracting 1 from the number of pixels in the row direction (horizontal direction) of the feature map.
  • a i, j indicates the pixel value of the coordinates (i, j) of the feature map.
  • mean represents the average of the pixel values of the feature map and is calculated by the equation (2).
  • a in FIG. 6 shows an example of a feature map having a large dispersion of pixel values
  • B in FIG. 6 shows an example of a feature map having a small dispersion of pixel values.
  • a feature map with a larger dispersion of pixel values is more likely to capture the features of a captured image
  • a feature map with a smaller dispersion of pixel values is less likely to capture the features of a captured image.
  • the learning unit 312 performs the learning process of the object recognition model 201 each time the teacher data is input, and also performs a plurality of feature maps generated from the captured images included in the teacher data in the layer 221 of the object recognition model 201. Is supplied to the feature data generation unit 321.
  • the feature data generation unit 321 calculates the variance of the pixel values of each feature map, and supplies the feature data indicating the calculated variance to the analysis data generation unit 322.
  • step S2 the analysis data generation unit 322 generates analysis data.
  • FIG. 7 shows an example in which the feature maps generated in the layer 221 of the object recognition model 201 are arranged in a row.
  • the analysis data generation unit 322 compresses the amount of information by arranging the feature data of the plurality of feature maps generated from the plurality of captured images in the layer 221 in a predetermined order, and performs one analysis. Generate data.
  • FIG. 8 shows an example in which the feature maps generated from 200 captured images are arranged two-dimensionally in the layer 221 of the object recognition model 201.
  • the vertical columns in the figure indicate the number of the captured image, and the horizontal row indicates the number of the feature map.
  • 512 feature maps generated from each captured image are arranged in numerical order from left to right.
  • 512 feature maps generated from the first captured image are arranged in numerical order from left to right.
  • the analysis data generation unit 322 generates 512 feature data (dispersion of each feature map) based on 512 feature maps generated from each captured image for each captured image in the order of the corresponding feature map numbers. By arranging them, a 512-dimensional vector (hereinafter referred to as a dispersion vector) is generated. For example, for the captured image 1 of FIG. 8, one dispersion vector having 512 feature data as elements based on the feature map 1 to the feature map 512 generated from the captured image 1 is generated. Then, the dispersion vector 1 to the dispersion vector 200 for the captured image 1 to the captured image 200 are generated.
  • FIG. 9 shows a dispersion vector and an example of imaging the dispersion vector.
  • the image of the dispersion vector is an image in which pixels showing colors corresponding to the values of each element (feature data) of the dispersion vector are arranged in order in the horizontal direction. For example, the pixel color is set to become red as the pixel value of the feature data (dispersion of the feature map) decreases, and to become blue as the value of the feature data (dispersion of the feature map) of the pixel value increases. ing.
  • the image of the dispersion vector is actually a color image, but here it is shown by a grayscale image.
  • the analysis data generation unit 322 generates analysis data composed of image data in which the element (feature data) of the dispersion vector of each captured image is a pixel.
  • FIG. 10 shows an example of imaging the analysis data generated based on the feature map of FIG.
  • the pixel color of the analysis data becomes red as the value of the feature data which is the pixel value becomes smaller, and becomes blue as the value of the feature data which is the pixel value becomes larger, as in the image of the dispersion vector of FIG. It is set to be.
  • the image of the analysis data is actually a color image, it is shown here by a grayscale image.
  • the x-axis direction (horizontal direction) of the analysis data indicates the number of the feature map
  • the y-axis direction vertical direction
  • the feature data of each feature map is arranged in the same order as the feature map of FIG. That is, in the x-axis direction (horizontal direction) of the analysis data, 512 feature data based on 512 feature maps generated from the same captured image are arranged in numerical order of the feature maps. Further, in the y-axis direction of the analysis data, feature data based on feature maps (with the same number) corresponding to each other of different captured images are arranged in the order of the photographed images.
  • the white circles in the figure indicate the positions of the pixels corresponding to the 200th feature map 200 of the 100th captured image 100.
  • the analysis data generation unit 322 supplies the generated analysis data to the analysis unit 323 and the output control unit 314.
  • the output unit 315 displays the analysis data under the control of the output control unit 314, for example.
  • step S3 the analysis unit 323 analyzes the learning situation using the analysis data.
  • FIG. 11 shows an example of how the analysis data changes as the learning progresses.
  • the analysis data of A to E in FIG. 11 are arranged in the order of learning progress, the analysis data of A in FIG. 11 is the oldest, and the analysis data of E in FIG. 11 is the newest.
  • the horizontal line is, for example, a row in the x-axis direction in which pixels having a pixel value (value of feature data) equal to or greater than a predetermined threshold value are present.
  • the vertical line is, for example, a column in the y-axis direction in which pixels having a pixel value (value of feature data) of a predetermined threshold value or more are present.
  • the number of horizontal lines and vertical lines is counted for each row in the x-axis direction and a column in the y-axis direction of the analysis data. That is, even if a plurality of horizontal lines or a plurality of vertical lines are adjacent to each other and look like one line, they are counted as different lines.
  • the ideal state is that there are no horizontal lines in the analysis data and the vertical lines converge within a predetermined range, and the learning model is properly trained. is there.
  • the highly dispersed feature map is a feature map that is likely to contribute to the recognition of an object in the learning model by extracting the features of the captured image regardless of the content of the captured image.
  • the low-dispersion feature map is a feature map that does not extract the features of the captured image regardless of the content of the captured image and is unlikely to contribute to the recognition of the object of the learning model.
  • the highly dispersed feature map does not always contribute to the recognition of the object (for example, the vehicle) to be recognized by the learning model, and may contribute to the recognition of an object other than the recognition target if the learning is not sufficient. .. However, if the learning model is properly trained, the highly distributed feature map becomes a feature map that contributes to the recognition of the object to be recognized by the learning model.
  • regularization processing is performed in order to suppress overfitting and reduce the weight of the neural network.
  • the types of features extracted from the captured image are narrowed down to some extent. That is, the number of feature maps that can extract features of captured images, that is, highly dispersed feature maps, is narrowed down to some extent.
  • the number of vertical lines of the analysis data converges within a predetermined range, which is an ideal state in which regularization is normally performed.
  • the number of vertical lines of the analysis data decreases too much or becomes 0 because the regularization is too strong and the characteristics of the captured image are sufficient. It is in a state where it cannot be extracted. Further, as shown in FIG. 15, the reason why the number of vertical lines of the analysis data does not decrease even if the learning progresses is that the regularization is too weak and the types of features to be extracted from the captured image are not completely narrowed down. Is.
  • the analysis unit 323 determines that the regularization of the learning process is normally performed. On the other hand, the analysis unit 323 determines that the regularization is too weak when the number of vertical lines of the analysis data converges to a value exceeding a predetermined range or when the number of vertical lines does not converge. Further, the analysis unit 323 determines that the regularization is too strong when the number of vertical lines of the analysis data converges to a value less than a predetermined range.
  • the number of vertical lines of the analysis data decreases and converges as the learning progresses.
  • the number of vertical lines of the analysis data converges with a large value as compared with the case where the regularization process is not performed, but the learning process is based on the number of vertical lines as in the case where the regularization process is performed. It is possible to determine whether or not is performed normally.
  • a horizontal line appears in the line corresponding to the highly dispersed captured image of the analysis data.
  • a captured image hereinafter referred to as a low-dispersion captured image
  • a horizontal line does not appear in the line corresponding to the low-dispersion captured image of the analysis data.
  • the analysis unit 323 determines that the learning process is normally performed. On the other hand, the analysis unit 323 determines that the learning process is not normally performed when the number of horizontal lines of the analysis data converges to a value equal to or more than a predetermined threshold value or when the number of horizontal lines does not converge.
  • the analysis unit 323 supplies data indicating the analysis result of the learning situation to the parameter setting unit 324.
  • step S4 the parameter setting unit 324 adjusts the parameters of the learning process based on the analysis result. For example, when the analysis unit 323 determines that the regularization is too strong, the parameter setting unit 324 makes the value of the regularization parameter used in the regularization process smaller than the current value. On the other hand, the parameter setting unit 324 increases the value of the regularization parameter to a larger value than the current value when the analysis unit 323 determines that the regularization is too weak. The larger the value of the regularization parameter, the stronger the regularization, and the smaller the value of the regularization parameter, the weaker the regularization. The parameter setting unit 324 supplies the learning unit 312 with data indicating the adjusted regularization parameter.
  • the learning unit 312 sets the value of the regularization parameter to the value set by the parameter setting unit 324. As a result, regularization can be performed more normally in the learning process of the learning model.
  • the user may adjust the learning parameters such as the regularization parameter by referring to the analysis data displayed in the output unit 315.
  • step S5 the learning situation analysis unit 313 determines whether or not the learning process by the learning unit 312 is completed. If it is determined that the learning process has not been completed, the process returns to step S1. After that, in step S5, the processes of steps S1 to S5 are repeatedly executed until it is determined that the learning process is completed.
  • step S5 if it is determined in step S5 that the learning process has been completed, the learning situation analysis process ends.
  • the learning status of the learning model by the learning unit 312 can be analyzed. Further, based on the analysis result, the parameters of the learning process can be appropriately set to improve the learning accuracy and shorten the learning time.
  • the user can easily recognize the learning status of the learning model by visually recognizing the analysis data.
  • the image data used by the learning model to be analyzed by the information processing device 301 The type and the type of the object to be recognized are not particularly limited.
  • FIG. 16 shows a configuration example of the object recognition model 401, which is a learning model for which the learning situation is analyzed.
  • the object recognition model 401 is used, for example, in the vehicle exterior information detection unit 141 of the vehicle control system 100 of FIG.
  • the object recognition model 401 includes, for example, a captured image (for example, captured image 402) that is image data captured in front of the vehicle 10 and a millimeter wave image output from a millimeter wave radar that monitors the front of the vehicle 10 (for example, a captured image 402).
  • a millimeter wave image 403 is input as input data.
  • the millimeter wave image is, for example, image data showing the distribution of the intensity of the received signal of the millimeter wave radar reflected by the object in front of the vehicle 10 from a bird's-eye view.
  • the object recognition model 401 performs recognition processing of the vehicle in front of the vehicle 10 based on the captured image and the millimeter wave image, and outputs an output image (for example, output image 404) which is image data showing the recognition result.
  • Output as.
  • the object recognition model 401 is a learning model using DSSD (Deconvolutional Single Shot Detector).
  • the object recognition model 401 includes a feature extraction layer 411, a feature extraction layer 412, a coupling portion 413, and a prediction layer 414.
  • the feature extraction layer 411 has the same configuration as the feature extraction layer 211 of the object recognition model 201 of FIG. Each layer of the feature extraction layer 411 generates a feature map showing the features of the captured image by a predetermined calculation, and supplies the feature map to the next layer. Further, a part of the feature extraction layer 411 supplies the feature map to the connecting portion 413.
  • the feature extraction layer 412 has a hierarchical structure, and each layer is composed of a convolution layer, a pooling layer, and the like. Each layer of the feature extraction layer 412 generates a feature map showing the features of the millimeter-wave image by a predetermined calculation, and supplies the feature map to the next layer. Further, a part of the feature extraction layer 412 supplies the feature map to the connecting portion 413. Further, the feature extraction layer 412 converts the millimeter wave image into an image having the same camera coordinate system as the captured image.
  • the connecting unit 413 combines the feature maps output from the corresponding layers of the feature extraction layer 411 and the feature extraction layer 412, and supplies the feature maps to the prediction layer 414.
  • the prediction layer 414 performs recognition processing of the vehicle in front of the vehicle 10 based on the feature map supplied from the joint portion 413.
  • the prediction layer 414 outputs an output image showing the recognition result of the vehicle.
  • the object recognition model 401 is divided into a camera network 421, a millimeter wave radar network 422, and a coupling network 423.
  • the camera network 421 includes the first half of the feature extraction layer 411 that generates a feature map that is not to be combined from the captured image.
  • the millimeter-wave radar network 422 includes the first half of the feature extraction layer 412, which generates a feature map that is not to be combined from a millimeter-wave image.
  • connection network 423 includes the latter half of the feature extraction layer 411 that generates the feature map to be combined, the latter half of the feature extraction layer 412 that generates the feature map to be combined, the connection portion 413, and the prediction layer 414. Including.
  • the learning situation of the object recognition model 401 was analyzed by the learning situation analysis process described above with reference to FIG. Specifically, the feature map generated and output in the layer 431 which is the intermediate layer of the feature extraction layer 411 surrounded by the thick frame in FIG. 16 and the layer 432 which is the intermediate layer of the feature extraction layer 412.
  • the feature map generated and output in was analyzed.
  • the feature map generated in the layer 431 is the largest feature map among the feature maps of the feature extraction layer 411 used for the recognition process of the prediction layer 414.
  • the feature map generated in the layer 432 is the largest feature map among the feature maps of the feature extraction layer 412 used for the recognition process of the prediction layer 414.
  • the analysis data based on the feature map generated in the feature map layer 432 of the feature extraction layer 412 remained without disappearing even if the learning process proceeded, as shown in A to E of FIG. That is, it was found that in the feature extraction layer 412, learning was not properly performed and the features of the millimeter wave image were not properly extracted.
  • the analysis data of A to E in FIG. 17 are arranged in the order of progress of learning, as in A to E of FIG.
  • FIG. 18 is an enlarged view of the analysis data of E in FIG.
  • FIGS. 19 to 21 schematically show millimeter-wave images and captured images corresponding to the rows in which the horizontal lines of the analysis data appear, which are indicated by the arrows in FIG. 19 to 21A are converted millimeter-wave images into grayscale images, and FIGS. 19 to 21B are diagrams of captured images.
  • FIG. 22 shows an example of a feature map of the row in which the horizontal line of the analysis data appears.
  • the millimeter-wave image 501 is obtained by converting the millimeter-wave image of the line in which the horizontal line of the analysis data appears into a grayscale image, and the feature map 502-1 to the feature map 502-4 are from the millimeter-wave image 501. It is part of the generated feature map.
  • the pixel values change significantly near the left and right edges in front of the vehicle 10.
  • the feature map generated from the millimeter wave image features corresponding to the left and right walls in front of the vehicle, not the vehicle in front of the vehicle 10, can be easily extracted.
  • the feature extraction layer 412 may be more suitable for recognizing the left and right walls than the vehicle in front of the vehicle 10.
  • the feature data used for the analysis data is not limited to the dispersion of the pixel values of the feature map described above, and other numerical values representing the features of the feature map can be used.
  • the norm of the feature map calculated by the following equation (3) may be used for the feature data.
  • Equation (3) indicates the norm of the feature map.
  • n and m indicate arbitrary numbers. Other symbols are the same as those in the above equation (1).
  • the norm of the feature map increases as the degree of dispersion of the pixel values of each pixel with respect to the average value of the pixel values of the feature map increases, and decreases as the degree of dispersion of the pixel values of each pixel with respect to the average value of the pixel values of the feature map decreases. Become. Therefore, the norm of the feature map indicates how the pixel values of the feature map are scattered.
  • the frequency distribution of the pixel values of the feature map may be used for the feature data.
  • FIG. 23 shows an example of a histogram (frequency distribution) of pixel values of the feature map.
  • the horizontal axis shows the class based on the pixel value of the feature map. That is, the pixel values of the feature map are classified into a plurality of classes.
  • the vertical axis shows the frequency. That is, the pixel value indicates the number of pixels of the feature map belonging to each class.
  • the analysis data shown in FIG. 24 is generated based on the plurality of feature maps generated from the plurality of input images (for example, the above-mentioned captured image or millimeter wave image).
  • the analysis data 521 of FIG. 24 is three-dimensional data in which the two-dimensional frequency map 522-1 to the frequency map 522-m generated for each class of the histogram are arranged in the z-axis direction (depth direction).
  • the x-axis direction (horizontal direction) of the frequency map 522-1 indicates the number of the feature map
  • the y-axis direction indicates the number of the input image
  • the frequencies of the first class of the histogram of the feature map of each input image are arranged in the x-axis direction (horizontal direction) in the order of the corresponding feature map numbers. Further, the frequencies corresponding to the feature maps having the same numbers of different input images are arranged in the y-axis direction (vertical direction) in the order of the numbers of the corresponding input images.
  • the frequency map 522-1 to the frequency map 522-m can be used for analyzing the learning situation based on the vertical and horizontal lines, respectively, as in the above-mentioned two-dimensional analysis data. Further, for example, the analysis data 521 can be used for analyzing the learning situation based on the line in the z-axis direction.
  • the analysis data including the extracted feature data is generated.
  • the feature data is the dispersion of the pixel values of the feature map
  • the feature data whose value is equal to or larger than a predetermined threshold is extracted from the feature data based on a plurality of feature maps generated from one image data.
  • Analysis data including feature data may be generated.
  • analysis data may be generated for one or more image data based on a plurality of feature maps generated in different layers of the neural network.
  • three-dimensional analysis data is generated by stacking the feature data of the feature map generated in each layer in two dimensions in the x-axis direction and the y-axis direction for each layer in the z-axis direction. You may try to do it. That is, in this analysis data, the feature data of the feature maps generated in the same layer are arranged in the x-axis direction and the y-axis direction, and the feature data of the feature maps generated in different layers are arranged in the z-axis direction. Lined up.
  • the learning situation analysis process can be performed after the learning process is completed.
  • the feature map generated during the learning process may be accumulated, and analysis data may be generated based on the feature map accumulated after the learning process to analyze the learning situation.
  • the learning model to be analyzed in the learning process is not limited to the above-mentioned example, and the entire learning model using the neural network can be targeted.
  • a recognition model that recognizes an object other than a vehicle and a recognition model that recognizes a plurality of objects including a vehicle are also targeted.
  • learning models whose input data is other than image data are also targeted.
  • a voice recognition model that uses voice data as input data, a sentence analysis model that uses sentence data as input data, and the like are also targeted.
  • FIG. 25 is a block diagram showing a configuration example of computer hardware that executes the above-mentioned series of processes programmatically.
  • the CPU Central Processing Unit
  • ROM Read Only Memory
  • RAM Random Access Memory
  • An input / output interface 1005 is further connected to the bus 1004.
  • An input unit 1006, an output unit 1007, a recording unit 1008, a communication unit 1009, and a drive 1010 are connected to the input / output interface 1005.
  • the input unit 1006 includes an input switch, a button, a microphone, an image sensor, and the like.
  • the output unit 1007 includes a display, a speaker, and the like.
  • the recording unit 1008 includes a hard disk, a non-volatile memory, and the like.
  • the communication unit 1009 includes a network interface and the like.
  • the drive 1010 drives a removable medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
  • the CPU 1001 loads and executes the program recorded in the recording unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004, as described above. A series of processing is performed.
  • the program executed by the computer 1000 can be recorded and provided on the removable media 1011 as a package media or the like, for example. Programs can also be provided via wired or wireless transmission media such as local area networks, the Internet, and digital satellite broadcasting.
  • the program can be installed in the recording unit 1008 via the input / output interface 1005 by mounting the removable media 1011 in the drive 1010.
  • the program can be received by the communication unit 1009 and installed in the recording unit 1008 via a wired or wireless transmission medium.
  • the program can be installed in advance in the ROM 1002 or the recording unit 1008.
  • the program executed by the computer may be a program in which processing is performed in chronological order in the order described in this specification, or in parallel or at a necessary timing such as when a call is made. It may be a program in which processing is performed.
  • the system means a set of a plurality of components (devices, modules (parts), etc.), and it does not matter whether all the components are in the same housing. Therefore, a plurality of devices housed in separate housings and connected via a network, and a device in which a plurality of modules are housed in one housing are both systems. ..
  • the embodiment of the present technology is not limited to the above-described embodiment, and various changes can be made without departing from the gist of the present technology.
  • this technology can have a cloud computing configuration in which one function is shared by a plurality of devices via a network and processed jointly.
  • each step described in the above flowchart can be executed by one device or shared by a plurality of devices.
  • one step includes a plurality of processes
  • the plurality of processes included in the one step can be executed by one device or shared by a plurality of devices.
  • the present technology can also have the following configurations.
  • a feature data generation step that generates feature data that numerically represents the features of the feature map generated from the input data in a model using a neural network, and a feature data generation step.
  • An information processing method including an analysis data generation step for generating analysis data based on the feature data of a plurality of the feature maps.
  • (3) In the analysis data the feature data of a plurality of the feature maps generated from the plurality of input data are arranged in a predetermined layer of the model that generates a plurality of the feature maps from one input data.
  • the feature data of a plurality of the feature maps generated from the same input data are arranged in the first direction of the analysis data, and different in the second direction orthogonal to the first direction.
  • the information processing method according to (3) above wherein the feature data of the feature map corresponding to each other of the input data are arranged.
  • the information processing method according to (6) above further including a parameter setting step of setting parameters for learning the model based on the analysis result of the learning situation of the model.
  • the information processing method according to (7) above wherein in the parameter setting step, regularization parameters for learning the model are set based on the number of lines in the second direction of the analysis data.
  • the feature data shows the frequency distribution of the pixel values of the feature map.
  • the information processing method according to (3) wherein in the analysis data, the feature data of a plurality of the feature maps generated from the plurality of input data in the hierarchy of the model are arranged three-dimensionally.
  • the input data is image data and The information processing method according to any one of (1) to (9) above, wherein the model performs object recognition processing.
  • the input data is image data representing the distribution of the intensity of the received signal of the millimeter wave radar by a bird's-eye view.
  • the model converts the image data into an image of a camera coordinate system.
  • the analysis data includes the feature data satisfying a predetermined condition among the feature data of the plurality of feature maps.
  • a feature data generation step that generates feature data that numerically represents the features of the feature map generated from the input data in a model using a neural network, and a feature data generation step.
  • a program for causing a computer to execute a process including an analysis data generation step for generating analysis data in which the feature data of a plurality of the feature maps are arranged.
  • a feature data generator that generates feature data that numerically represents the features of the feature map generated from the input data in a model using a neural network.
  • An information processing device including an analysis data generation unit that generates analysis data in which the feature data of a plurality of the feature maps are arranged.
  • 10 vehicles 100 vehicle control system, 141 external information detection unit, 201 object recognition model, 211 feature extraction layer, 221 hierarchy, 301 information processing device, 312 learning unit, 313 learning situation analysis unit, 321 feature data generation unit, 322 analysis Data generation unit, 323 analysis unit, 324 parameter setting unit, 401 object recognition model 411,412 feature extraction layer, 421 camera network, 422 millimeter-wave radar network, 423 coupling network, 431,432 hierarchy

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Evolutionary Computation (AREA)
  • Artificial Intelligence (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Data Mining & Analysis (AREA)
  • Multimedia (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • General Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Molecular Biology (AREA)
  • Biophysics (AREA)
  • Mathematical Physics (AREA)
  • Biomedical Technology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Medical Informatics (AREA)
  • Databases & Information Systems (AREA)
  • Remote Sensing (AREA)
  • Radar, Positioning & Navigation (AREA)
  • Evolutionary Biology (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Electromagnetism (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Traffic Control Systems (AREA)
  • Image Analysis (AREA)

Abstract

本技術は、ニューラルネットワークを用いたモデルの学習状況を解析できるようにする情報処理方法、プログラム、及び、情報処理装置に関する。 ステップS1において、ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データが生成され、ステップS2において、複数の前記特徴マップの前記特徴データに基づく解析用データが生成される。本技術は、例えば、車両の前方の車両を認識するシステムに適用できる。

Description

情報処理方法、プログラム、及び、情報処理装置
 本技術は、情報処理方法、プログラム、及び、情報処理装置に関し、特に、ニューラルネットワークを用いたモデルの解析を行う場合に用いて好適な情報処理方法、プログラム、及び、情報処理装置に関する。
 従来、画像パターン認識装置にニューラルネットワークが用いられている(例えば、特許文献1参照)。
特開平5-61976号公報
 一方、ニューラルネットワークを用いたモデルの学習状況を解析できるようにすることが望まれている。
 本技術は、このような状況に鑑みてなされたものであり、ニューラルネットワークを用いたモデルの学習状況を解析できるようにするものである。
 本技術の一側面の情報処理方法は、ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データを生成する特徴データ生成ステップと、複数の前記特徴マップの前記特徴データに基づく解析用データを生成する解析用データ生成ステップとを含む。
 本技術の一側面のプログラムは、ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データを生成する特徴データ生成ステップと、複数の前記特徴マップの前記特徴データに基づく解析用データを生成する解析用データ生成ステップとを含む処理をコンピュータに実行させる。
 本技術の一側面の情報処理装置は、ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データを生成する特徴データ生成部と、複数の前記特徴マップの前記特徴データに基づく解析用データを生成する解析用データ生成部とを備える。
 本技術の一側面においては、ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データが生成され、複数の前記特徴マップの前記特徴データに基づく解析用データが生成される。
車両制御システムの構成例を示すブロック図である。 物体認識モデルの第1の実施の形態を示す図である。 特徴マップの例を示す図である。 本技術を適用した情報処理装置の構成例を示すブロック図である。 学習状況解析処理を説明するためのフローチャートである。 画素値の分散が大きい特徴マップと画素値の分散が小さい特徴マップの例を示す図である。 特徴マップを一列に並べた例を示す図である。 特徴マップを2次元に並べた例を示す図である。 分散ベクトル及び分散ベクトルの画像の例を示す図である。 解析用データの例を示す図である。 学習の進行に伴う解析用データの変化の様子の第1の例を示す図である。 学習初期と学習末期の特徴マップの例を示す図である。 解析用データの例を示す図である。 正則化が強すぎる場合の解析用データの例を示す図である。 正則化が弱すぎる場合の解析用データの例を示す図である。 物体認識モデルの第2の実施の形態を示す図である。 学習の進行に伴う解析用データの変化の様子の第2の例を示す図である。 図17のEの解析用データを拡大した図である。 ミリ波画像と撮影画像の第1の例を示す図である。 ミリ波画像と撮影画像の第2の例を示す図である。 ミリ波画像と撮影画像の第3の例を示す図である。 ミリ波画像から生成される特徴マップの例を示す図である。 特徴マップの画素値のヒストグラムの例を示す図である。 特徴マップの画素値の度数分布に基づく解析用データの例を示す図である。 コンピュータの構成例を示す図である。
 以下、本技術を実施するための形態について説明する。説明は以下の順序で行う。
 1.実施の形態
 2.学習状況の解析事例
 3.変形例
 4.その他
 <<1.実施の形態>>
 まず、図1乃至図15を参照して、本技術の実施の形態について説明する。
  <車両制御システム100の構成例>
 図1は、本技術が適用され得る移動体制御システムの一例である車両制御システム100の概略的な機能の構成例を示すブロック図である。
 なお、以下、車両制御システム100が設けられている車両10を他の車両と区別する場合、自車又は自車両と称する。
 車両制御システム100は、入力部101、データ取得部102、通信部103、車内機器104、出力制御部105、出力部106、駆動系制御部107、駆動系システム108、ボディ系制御部109、ボディ系システム110、記憶部111、及び、自動運転制御部112を備える。入力部101、データ取得部102、通信部103、出力制御部105、駆動系制御部107、ボディ系制御部109、記憶部111、及び、自動運転制御部112は、通信ネットワーク121を介して、相互に接続されている。通信ネットワーク121は、例えば、CAN(Controller Area Network)、LIN(Local Interconnect Network)、LAN(Local Area Network)、又は、FlexRay(登録商標)等の任意の規格に準拠した車載通信ネットワークやバス等からなる。なお、車両制御システム100の各部は、通信ネットワーク121を介さずに、直接接続される場合もある。
 なお、以下、車両制御システム100の各部が、通信ネットワーク121を介して通信を行う場合、通信ネットワーク121の記載を省略するものとする。例えば、入力部101と自動運転制御部112が、通信ネットワーク121を介して通信を行う場合、単に入力部101と自動運転制御部112が通信を行うと記載する。
 入力部101は、搭乗者が各種のデータや指示等の入力に用いる装置を備える。例えば、入力部101は、タッチパネル、ボタン、マイクロフォン、スイッチ、及び、レバー等の操作デバイス、並びに、音声やジェスチャ等により手動操作以外の方法で入力可能な操作デバイス等を備える。また、例えば、入力部101は、赤外線若しくはその他の電波を利用したリモートコントロール装置、又は、車両制御システム100の操作に対応したモバイル機器若しくはウェアラブル機器等の外部接続機器であってもよい。入力部101は、搭乗者により入力されたデータや指示等に基づいて入力信号を生成し、車両制御システム100の各部に供給する。
 データ取得部102は、車両制御システム100の処理に用いるデータを取得する各種のセンサ等を備え、取得したデータを、車両制御システム100の各部に供給する。
 例えば、データ取得部102は、自車の状態等を検出するための各種のセンサを備える。具体的には、例えば、データ取得部102は、ジャイロセンサ、加速度センサ、慣性計測装置(IMU)、及び、アクセルペダルの操作量、ブレーキペダルの操作量、ステアリングホイールの操舵角、エンジン回転数、モータ回転数、若しくは、車輪の回転速度等を検出するためのセンサ等を備える。
 また、例えば、データ取得部102は、自車の外部の情報を検出するための各種のセンサを備える。具体的には、例えば、データ取得部102は、ToF(Time Of Flight)カメラ、ステレオカメラ、単眼カメラ、赤外線カメラ、及び、その他のカメラ等の撮像装置を備える。また、例えば、データ取得部102は、天候又は気象等を検出するための環境センサ、及び、自車の周囲の物体を検出するための周囲情報検出センサを備える。環境センサは、例えば、雨滴センサ、霧センサ、日照センサ、雪センサ等からなる。周囲情報検出センサは、例えば、超音波センサ、レーダ、LiDAR(Light Detection and Ranging、Laser Imaging Detection and Ranging)、ソナー等からなる。
 さらに、例えば、データ取得部102は、自車の現在位置を検出するための各種のセンサを備える。具体的には、例えば、データ取得部102は、GNSS(Global Navigation Satellite System)衛星からのGNSS信号を受信するGNSS受信機等を備える。
 また、例えば、データ取得部102は、車内の情報を検出するための各種のセンサを備える。具体的には、例えば、データ取得部102は、運転者を撮像する撮像装置、運転者の生体情報を検出する生体センサ、及び、車室内の音声を集音するマイクロフォン等を備える。生体センサは、例えば、座面又はステアリングホイール等に設けられ、座席に座っている搭乗者又はステアリングホイールを握っている運転者の生体情報を検出する。
 通信部103は、車内機器104、並びに、車外の様々な機器、サーバ、基地局等と通信を行い、車両制御システム100の各部から供給されるデータを送信したり、受信したデータを車両制御システム100の各部に供給したりする。なお、通信部103がサポートする通信プロトコルは、特に限定されるものではなく、また、通信部103が、複数の種類の通信プロトコルをサポートすることも可能である。
 例えば、通信部103は、無線LAN、Bluetooth(登録商標)、NFC(Near Field Communication)、又は、WUSB(Wireless USB)等により、車内機器104と無線通信を行う。また、例えば、通信部103は、図示しない接続端子(及び、必要であればケーブル)を介して、USB(Universal Serial Bus)、HDMI(登録商標)(High-Definition Multimedia Interface)、又は、MHL(Mobile High-definition Link)等により、車内機器104と有線通信を行う。
 さらに、例えば、通信部103は、基地局又はアクセスポイントを介して、外部ネットワーク(例えば、インターネット、クラウドネットワーク又は事業者固有のネットワーク)上に存在する機器(例えば、アプリケーションサーバ又は制御サーバ)との通信を行う。また、例えば、通信部103は、P2P(Peer To Peer)技術を用いて、自車の近傍に存在する端末(例えば、歩行者若しくは店舗の端末、又は、MTC(Machine Type Communication)端末)との通信を行う。さらに、例えば、通信部103は、車車間(Vehicle to Vehicle)通信、路車間(Vehicle to Infrastructure)通信、自車と家との間(Vehicle to Home)の通信、及び、歩車間(Vehicle to Pedestrian)通信等のV2X通信を行う。また、例えば、通信部103は、ビーコン受信部を備え、道路上に設置された無線局等から発信される電波あるいは電磁波を受信し、現在位置、渋滞、通行規制又は所要時間等の情報を取得する。
 車内機器104は、例えば、搭乗者が有するモバイル機器若しくはウェアラブル機器、自車に搬入され若しくは取り付けられる情報機器、及び、任意の目的地までの経路探索を行うナビゲーション装置等を含む。
 出力制御部105は、自車の搭乗者又は車外に対する各種の情報の出力を制御する。例えば、出力制御部105は、視覚情報(例えば、画像データ)及び聴覚情報(例えば、音声データ)のうちの少なくとも1つを含む出力信号を生成し、出力部106に供給することにより、出力部106からの視覚情報及び聴覚情報の出力を制御する。具体的には、例えば、出力制御部105は、データ取得部102の異なる撮像装置により撮像された画像データを合成して、俯瞰画像又はパノラマ画像等を生成し、生成した画像を含む出力信号を出力部106に供給する。また、例えば、出力制御部105は、衝突、接触、危険地帯への進入等の危険に対する警告音又は警告メッセージ等を含む音声データを生成し、生成した音声データを含む出力信号を出力部106に供給する。
 出力部106は、自車の搭乗者又は車外に対して、視覚情報又は聴覚情報を出力することが可能な装置を備える。例えば、出力部106は、表示装置、インストルメントパネル、オーディオスピーカ、ヘッドホン、搭乗者が装着する眼鏡型ディスプレイ等のウェアラブルデバイス、プロジェクタ、ランプ等を備える。出力部106が備える表示装置は、通常のディスプレイを有する装置以外にも、例えば、ヘッドアップディスプレイ、透過型ディスプレイ、AR(Augmented Reality)表示機能を有する装置等の運転者の視野内に視覚情報を表示する装置であってもよい。
 駆動系制御部107は、各種の制御信号を生成し、駆動系システム108に供給することにより、駆動系システム108の制御を行う。また、駆動系制御部107は、必要に応じて、駆動系システム108以外の各部に制御信号を供給し、駆動系システム108の制御状態の通知等を行う。
 駆動系システム108は、自車の駆動系に関わる各種の装置を備える。例えば、駆動系システム108は、内燃機関又は駆動用モータ等の駆動力を発生させるための駆動力発生装置、駆動力を車輪に伝達するための駆動力伝達機構、舵角を調節するステアリング機構、制動力を発生させる制動装置、ABS(Antilock Brake System)、ESC(Electronic Stability Control)、並びに、電動パワーステアリング装置等を備える。
 ボディ系制御部109は、各種の制御信号を生成し、ボディ系システム110に供給することにより、ボディ系システム110の制御を行う。また、ボディ系制御部109は、必要に応じて、ボディ系システム110以外の各部に制御信号を供給し、ボディ系システム110の制御状態の通知等を行う。
 ボディ系システム110は、車体に装備されたボディ系の各種の装置を備える。例えば、ボディ系システム110は、キーレスエントリシステム、スマートキーシステム、パワーウィンドウ装置、パワーシート、ステアリングホイール、空調装置、及び、各種ランプ(例えば、ヘッドランプ、バックランプ、ブレーキランプ、ウィンカ、フォグランプ等)等を備える。
 記憶部111は、例えば、ROM(Read Only Memory)、RAM(Random Access Memory)、HDD(Hard Disc Drive)等の磁気記憶デバイス、半導体記憶デバイス、光記憶デバイス、及び、光磁気記憶デバイス等を備える。記憶部111は、車両制御システム100の各部が用いる各種プログラムやデータ等を記憶する。例えば、記憶部111は、ダイナミックマップ等の3次元の高精度地図、高精度地図より精度が低く、広いエリアをカバーするグローバルマップ、及び、自車の周囲の情報を含むローカルマップ等の地図データを記憶する。
 自動運転制御部112は、自律走行又は運転支援等の自動運転に関する制御を行う。具体的には、例えば、自動運転制御部112は、自車の衝突回避あるいは衝撃緩和、車間距離に基づく追従走行、車速維持走行、自車の衝突警告、又は、自車のレーン逸脱警告等を含むADAS(Advanced Driver Assistance System)の機能実現を目的とした協調制御を行う。また、例えば、自動運転制御部112は、運転者の操作に拠らずに自律的に走行する自動運転等を目的とした協調制御を行う。自動運転制御部112は、検出部131、自己位置推定部132、状況分析部133、計画部134、及び、動作制御部135を備える。
 検出部131は、自動運転の制御に必要な各種の情報の検出を行う。検出部131は、車外情報検出部141、車内情報検出部142、及び、車両状態検出部143を備える。
 車外情報検出部141は、車両制御システム100の各部からのデータ又は信号に基づいて、自車の外部の情報の検出処理を行う。例えば、車外情報検出部141は、自車の周囲の物体の検出処理、認識処理、及び、追跡処理、並びに、物体までの距離の検出処理を行う。検出対象となる物体には、例えば、車両、人、障害物、構造物、道路、信号機、交通標識、道路標示等が含まれる。また、例えば、車外情報検出部141は、自車の周囲の環境の検出処理を行う。検出対象となる周囲の環境には、例えば、天候、気温、湿度、明るさ、及び、路面の状態等が含まれる。車外情報検出部141は、検出処理の結果を示すデータを自己位置推定部132、状況分析部133のマップ解析部151、交通ルール認識部152、及び、状況認識部153、並びに、動作制御部135の緊急事態回避部171等に供給する。
 車内情報検出部142は、車両制御システム100の各部からのデータ又は信号に基づいて、車内の情報の検出処理を行う。例えば、車内情報検出部142は、運転者の認証処理及び認識処理、運転者の状態の検出処理、搭乗者の検出処理、及び、車内の環境の検出処理等を行う。検出対象となる運転者の状態には、例えば、体調、覚醒度、集中度、疲労度、視線方向等が含まれる。検出対象となる車内の環境には、例えば、気温、湿度、明るさ、臭い等が含まれる。車内情報検出部142は、検出処理の結果を示すデータを状況分析部133の状況認識部153、及び、動作制御部135の緊急事態回避部171等に供給する。
 車両状態検出部143は、車両制御システム100の各部からのデータ又は信号に基づいて、自車の状態の検出処理を行う。検出対象となる自車の状態には、例えば、速度、加速度、舵角、異常の有無及び内容、運転操作の状態、パワーシートの位置及び傾き、ドアロックの状態、並びに、その他の車載機器の状態等が含まれる。車両状態検出部143は、検出処理の結果を示すデータを状況分析部133の状況認識部153、及び、動作制御部135の緊急事態回避部171等に供給する。
 自己位置推定部132は、車外情報検出部141、及び、状況分析部133の状況認識部153等の車両制御システム100の各部からのデータ又は信号に基づいて、自車の位置及び姿勢等の推定処理を行う。また、自己位置推定部132は、必要に応じて、自己位置の推定に用いるローカルマップ(以下、自己位置推定用マップと称する)を生成する。自己位置推定用マップは、例えば、SLAM(Simultaneous Localization and Mapping)等の技術を用いた高精度なマップとされる。自己位置推定部132は、推定処理の結果を示すデータを状況分析部133のマップ解析部151、交通ルール認識部152、及び、状況認識部153等に供給する。また、自己位置推定部132は、自己位置推定用マップを記憶部111に記憶させる。
 状況分析部133は、自車及び周囲の状況の分析処理を行う。状況分析部133は、マップ解析部151、交通ルール認識部152、状況認識部153、及び、状況予測部154を備える。
 マップ解析部151は、自己位置推定部132及び車外情報検出部141等の車両制御システム100の各部からのデータ又は信号を必要に応じて用いながら、記憶部111に記憶されている各種のマップの解析処理を行い、自動運転の処理に必要な情報を含むマップを構築する。マップ解析部151は、構築したマップを、交通ルール認識部152、状況認識部153、状況予測部154、並びに、計画部134のルート計画部161、行動計画部162、及び、動作計画部163等に供給する。
 交通ルール認識部152は、自己位置推定部132、車外情報検出部141、及び、マップ解析部151等の車両制御システム100の各部からのデータ又は信号に基づいて、自車の周囲の交通ルールの認識処理を行う。この認識処理により、例えば、自車の周囲の信号の位置及び状態、自車の周囲の交通規制の内容、並びに、走行可能な車線等が認識される。交通ルール認識部152は、認識処理の結果を示すデータを状況予測部154等に供給する。
 状況認識部153は、自己位置推定部132、車外情報検出部141、車内情報検出部142、車両状態検出部143、及び、マップ解析部151等の車両制御システム100の各部からのデータ又は信号に基づいて、自車に関する状況の認識処理を行う。例えば、状況認識部153は、自車の状況、自車の周囲の状況、及び、自車の運転者の状況等の認識処理を行う。また、状況認識部153は、必要に応じて、自車の周囲の状況の認識に用いるローカルマップ(以下、状況認識用マップと称する)を生成する。状況認識用マップは、例えば、占有格子地図(Occupancy Grid Map)とされる。
 認識対象となる自車の状況には、例えば、自車の位置、姿勢、動き(例えば、速度、加速度、移動方向等)、並びに、異常の有無及び内容等が含まれる。認識対象となる自車の周囲の状況には、例えば、周囲の静止物体の種類及び位置、周囲の動物体の種類、位置及び動き(例えば、速度、加速度、移動方向等)、周囲の道路の構成及び路面の状態、並びに、周囲の天候、気温、湿度、及び、明るさ等が含まれる。認識対象となる運転者の状態には、例えば、体調、覚醒度、集中度、疲労度、視線の動き、並びに、運転操作等が含まれる。
 状況認識部153は、認識処理の結果を示すデータ(必要に応じて、状況認識用マップを含む)を自己位置推定部132及び状況予測部154等に供給する。また、状況認識部153は、状況認識用マップを記憶部111に記憶させる。
 状況予測部154は、マップ解析部151、交通ルール認識部152及び状況認識部153等の車両制御システム100の各部からのデータ又は信号に基づいて、自車に関する状況の予測処理を行う。例えば、状況予測部154は、自車の状況、自車の周囲の状況、及び、運転者の状況等の予測処理を行う。
 予測対象となる自車の状況には、例えば、自車の挙動、異常の発生、及び、走行可能距離等が含まれる。予測対象となる自車の周囲の状況には、例えば、自車の周囲の動物体の挙動、信号の状態の変化、及び、天候等の環境の変化等が含まれる。予測対象となる運転者の状況には、例えば、運転者の挙動及び体調等が含まれる。
 状況予測部154は、予測処理の結果を示すデータを、交通ルール認識部152及び状況認識部153からのデータとともに、計画部134のルート計画部161、行動計画部162、及び、動作計画部163等に供給する。
 ルート計画部161は、マップ解析部151及び状況予測部154等の車両制御システム100の各部からのデータ又は信号に基づいて、目的地までのルートを計画する。例えば、ルート計画部161は、グローバルマップに基づいて、現在位置から指定された目的地までのルートを設定する。また、例えば、ルート計画部161は、渋滞、事故、通行規制、工事等の状況、及び、運転者の体調等に基づいて、適宜ルートを変更する。ルート計画部161は、計画したルートを示すデータを行動計画部162等に供給する。
 行動計画部162は、マップ解析部151及び状況予測部154等の車両制御システム100の各部からのデータ又は信号に基づいて、ルート計画部161により計画されたルートを計画された時間内で安全に走行するための自車の行動を計画する。例えば、行動計画部162は、発進、停止、進行方向(例えば、前進、後退、左折、右折、方向転換等)、走行車線、走行速度、及び、追い越し等の計画を行う。行動計画部162は、計画した自車の行動を示すデータを動作計画部163等に供給する。
 動作計画部163は、マップ解析部151及び状況予測部154等の車両制御システム100の各部からのデータ又は信号に基づいて、行動計画部162により計画された行動を実現するための自車の動作を計画する。例えば、動作計画部163は、加速、減速、及び、走行軌道等の計画を行う。動作計画部163は、計画した自車の動作を示すデータを、動作制御部135の加減速制御部172及び方向制御部173等に供給する。
 動作制御部135は、自車の動作の制御を行う。動作制御部135は、緊急事態回避部171、加減速制御部172、及び、方向制御部173を備える。
 緊急事態回避部171は、車外情報検出部141、車内情報検出部142、及び、車両状態検出部143の検出結果に基づいて、衝突、接触、危険地帯への進入、運転者の異常、車両の異常等の緊急事態の検出処理を行う。緊急事態回避部171は、緊急事態の発生を検出した場合、急停車や急旋回等の緊急事態を回避するための自車の動作を計画する。緊急事態回避部171は、計画した自車の動作を示すデータを加減速制御部172及び方向制御部173等に供給する。
 加減速制御部172は、動作計画部163又は緊急事態回避部171により計画された自車の動作を実現するための加減速制御を行う。例えば、加減速制御部172は、計画された加速、減速、又は、急停車を実現するための駆動力発生装置又は制動装置の制御目標値を演算し、演算した制御目標値を示す制御指令を駆動系制御部107に供給する。
 方向制御部173は、動作計画部163又は緊急事態回避部171により計画された自車の動作を実現するための方向制御を行う。例えば、方向制御部173は、動作計画部163又は緊急事態回避部171により計画された走行軌道又は急旋回を実現するためのステアリング機構の制御目標値を演算し、演算した制御目標値を示す制御指令を駆動系制御部107に供給する。
  <物体認識モデル201の構成例>
 図2は、物体認識モデル201の構成例を示している。物体認識モデル201は、例えば、図1の車両制御システム100の車外情報検出部141に用いられる。
 物体認識モデル201には、例えば、車両10の前方を撮影した画像データである撮影画像(例えば、撮影画像202)が入力データとして入力される。そして、物体認識モデル201は、撮影画像に基づいて、車両10の前方の車両の認識処理を行い、認識結果を示す画像データである出力画像(例えば、出力画像203)を出力データとして出力する。
 物体認識モデル201は、畳み込みニューラルネットワーク(CNN)を用いたモデルであり、特徴抽出層211及び予測層212を備える。
 特徴抽出層211は、複数の階層を備え、各階層は、畳み込み層、プーリング層等からなる。特徴抽出層211の各階層は、それぞれ所定の演算により、撮影画像の特徴を示す1以上の特徴マップを生成し、次の階層に供給する。また、特徴抽出層211の一部の階層は、特徴マップを予測層212に供給する。
 なお、特徴抽出層211の各階層において生成される特徴マップのサイズ(画素数)は異なり、階層が進むにつれて徐々に小さくなる。
 図3は、図2の特徴抽出層211の点線の四角で囲まれている中間層である階層221において生成され、出力される特徴マップ231-1乃至特徴マップ231-nの例を示している。
 階層221では、1つの撮影画像に対してn個の特徴マップ231-1乃至特徴マップ231-nが生成される。特徴マップ231-1乃至特徴マップ231-nは、例えば、縦38個×横38個の画素が2次元に並べられた画像データである。また、特徴マップ231-1乃至特徴マップ231-nは、予測層212において認識処理に用いられる特徴マップのうち、最もサイズの大きい特徴マップである。
 なお、以下、階層毎に、各撮影画像から生成される特徴マップに1から始まるシリアル番号がそれぞれ割り当てられるものとする。例えば、階層221において1つの撮影画像から生成される特徴マップ231-1乃至特徴マップ231-nに、1からnまでのシリアル番号がそれぞれ割り当てられるものとする。同様に、他の階層において1つの撮影画像から生成される特徴マップに、1からNまでのシリアル番号がそれぞれ割り当てられるものとする。なお、Nは、その階層で1つの撮影画像から生成される特徴マップの数を示す。
 予測層212は、特徴抽出層211から供給される特徴マップに基づいて、車両10の前方の車両の認識処理を行う。予測層212は、車両の認識結果を示す出力画像を出力する。
  <情報処理装置301の構成例>
 図4は、図2の物体認識モデル201等のニューラルネットワークを用いたモデル(以下、学習モデルと称する)の学習に用いられる情報処理装置301の構成例を示している。
 情報処理装置301は、入力部311、学習部312、学習状況解析部313、出力制御部314、及び、出力部315を備える。
 入力部311は、各種のデータや指示等の入力に用いる入力デバイスを備え、入力されたデータや指示等に基づいて入力信号を生成し、学習部312に供給する。例えば、入力部311は、学習モデルの学習用の教師データの入力に用いられる。
 学習部312は、学習モデルの学習処理を行う。なお、学習部312の学習方法は、特定の方法に限定されない。また、学習部312は、学習モデルにおいて生成される特徴マップを学習状況解析部313に供給する。
 学習状況解析部313は、学習部312から供給される特徴マップに基づいて、学習部312による学習モデルの学習状況の解析等を行う。学習状況解析部313は、特徴データ生成部321、解析用データ生成部322、解析部323、及び、パラメータ設定部324を備える。
 特徴データ生成部321は、特徴マップの特徴を数値で表す特徴データを生成し、解析用データ生成部322に供給する。
 解析用データ生成部322は、複数の特徴マップの特徴データを並べた(配列した)解析用データを生成し、解析部323及び出力制御部314に供給する。
 解析部323は、解析用データに基づいて、学習モデルの学習状況の解析処理を行い、解析結果を示すデータをパラメータ設定部324に供給する。
 パラメータ設定部324は、学習モデルの学習状況の解析結果に基づいて、学習モデルの学習用の各種のパラメータを設定する。パラメータ設定部324は、設定したパラメータを示すデータを学習部312に供給する。
 出力制御部314は、出力部315による各種の情報の出力を制御する。例えば、出力制御部314は、出力部315による解析用データの表示を制御する。
 出力部315は、視覚情報、聴覚情報等の各種の情報を出力可能な出力デバイスを備える。例えば、出力部315は、ディスプレイ、スピーカ等を備える。
  <学習状況解析処理>
 次に、図5のフローチャートを参照して、情報処理装置301により実行される学習状況解析処理について説明する。
 この処理は、例えば、学習部312により学習モデルの学習処理が開始されたとき、開始される。
 なお、以下、図2の物体認識モデル201の学習処理が行われ、その学習状況の解析を行う場合を例に挙げて説明する。また、以下、物体認識モデル201の階層221において、1つの撮影画像から512個の特徴マップが生成される場合を例に挙げて説明する。
 ステップS1において、特徴データ生成部321は、特徴データを生成する。例えば、特徴データ生成部321は、次式(1)により、特徴マップの画素値の散らばり具合を示す分散を算出する。
Figure JPOXMLDOC01-appb-M000001
 式(1)のvarは、特徴マップの画素値の分散を示している。Iは、特徴マップの列方向(縦方向)の画素数から1を引いた値を示し、Jは、特徴マップの行方向(横方向)の画素数から1を引いた値を示す。Ai,jは、特徴マップの座標(i,j)の画素値を示している。meanは、特徴マップの画素値の平均を示し、式(2)により算出される。
 なお、図6のAは、画素値の分散が大きい特徴マップの例を示しており、図6のBは、画素値の分散が小さい特徴マップの例を示している。画素値の分散が大きい特徴マップほど、撮影画像の特徴を捉えている可能性が高く、画素値の分散が小さい特徴マップほど、撮影画像の特徴を捉えている可能性が低くなる。
 例えば、学習部312は、教師データが入力される毎に、物体認識モデル201の学習処理を行うとともに、物体認識モデル201の階層221において教師データに含まれる撮影画像から生成された複数の特徴マップを、特徴データ生成部321に供給する。特徴データ生成部321は、各特徴マップの画素値の分散を算出し、算出した分散を示す特徴データを解析用データ生成部322に供給する。
 ステップS2において、解析用データ生成部322は、解析用データを生成する。
 図7は、物体認識モデル201の階層221において生成される特徴マップを一列に並べた例を示している。特徴マップは、有用な情報を含んでいるが、物体認識モデル201の学習状況の解析にはあまり適していない。例えば、200枚の撮影画像が学習処理に用いられた場合、階層221では、合計102,400個(=512個×200)もの特徴マップが生成される。従って、特徴マップをそのまま用いて物体認識モデル201の学習状況の解析を行うのは、情報量が多すぎて難しい。
 これに対して、解析用データ生成部322は、階層221において複数の撮影画像から生成された複数の特徴マップの特徴データを所定の順に並べることにより、情報量を圧縮して、1つの解析用データを生成する。
 図8は、物体認識モデル201の階層221において、200枚の撮影画像から生成された特徴マップを2次元に並べた例を示している。図内の縦方向の列は、撮影画像の番号を示し、横方向の行は、特徴マップの番号を示している。各行には、各撮影画像から生成された512個の特徴マップが、左から右に番号順に並べられている。例えば、1行目には、1番目の撮影画像から生成された512個の特徴マップが、左から右に番号順に並べられている。
 そして、解析用データ生成部322は、撮影画像毎に、各撮影画像から生成された512個の特徴マップに基づく512個の特徴データ(各特徴マップの分散)を、対応する特徴マップの番号順に並べることにより、512次元のベクトル(以下、分散ベクトルと称する)を生成する。例えば、図8の撮影画像1に対して、撮影画像1から生成された特徴マップ1乃至特徴マップ512に基づく512個の特徴データを要素とする1つの分散ベクトルが生成される。そして、撮影画像1乃至撮影画像200に対する分散ベクトル1乃至分散ベクトル200が生成される。
 図9は、分散ベクトル、及び、分散ベクトルを画像化した例を示している。分散ベクトルの画像は、分散ベクトルの各要素(特徴データ)の値に応じた色を示す画素を横方向に順番に並べたものである。例えば、画素の色は、画素値である特徴データの値(特徴マップの分散)が小さくなるほど赤くなり、画素値である特徴データの値(特徴マップの分散)が大きくなるほど青くなるように設定されている。なお、分散ベクトルの画像は、実際にはカラーの画像であるが、ここではグレースケールの画像により示されている。
 そして、解析用データ生成部322は、各撮影画像の分散ベクトルの要素(特徴データ)を画素とする画像データからなる解析用データを生成する。
 図10は、図8の特徴マップに基づいて生成された解析用データを画像化した例を示している。なお、解析用データの画素の色は、例えば、図9の分散ベクトルの画像と同様に、画素値である特徴データの値が小さくなるほど赤くなり、画素値である特徴データの値が大きくなるほど青くなるように設定されている。なお、解析用データの画像は、実際にはカラーの画像であるが、ここではグレースケールの画像により示されている。また、解析用データのx軸方向(横方向)は、特徴マップの番号を示し、y軸方向(縦方向)は、撮影画像の番号を示している。
 図10の解析用データでは、図8の特徴マップと同じ並び順に、各特徴マップの特徴データが並べられている。すなわち、解析用データのx軸方向(横方向)には、同じ撮影画像から生成された512個の特徴マップにそれぞれ基づく512個の特徴データが、特徴マップの番号順に並べられている。また、解析用データのy軸方向には、異なる撮影画像の互いに対応する(同じ番号の)特徴マップに基づく特徴データが、撮影画像の番号順に並べられている。図内の白丸は、100番目の撮影画像100の200番目の特徴マップ200に対応する画素の位置を示している。
 解析用データ生成部322は、生成した解析用データを解析部323及び出力制御部314に供給する。
 出力部315は、例えば、出力制御部314の制御の下に、解析用データを表示する。
 ステップS3において、解析部323は、解析用データを用いて、学習状況の解析を行う。
 図11は、学習の進行に伴う解析用データの変化の様子の例を示している。図11のA乃至Eの解析用データは、学習の進行順に並べられており、図11のAの解析用データが最も古く、図11のEの解析用データが最も新しい。
 学習の初期段階の図11のAの解析用データでは、横線(x軸方向のライン)及び縦線(y軸方向のライン)が多い。一方、学習が進むにつれて、横線及び縦線が減少し、図11のEの解析用データでは、横線は存在せず、縦線も10本程度になっている。
 ここで、横線は、例えば、画素値(特徴データの値)が所定の閾値以上の画素が所定の閾値以上存在するx軸方向の行とする。縦線は、例えば、画素値(特徴データの値)が所定の閾値以上の画素が所定の閾値以上存在するy軸方向の列とする。また、横線及び縦線の数は、解析用データのx軸方向の行及びy軸方向の列毎にカウントされる。すなわち、複数の横線又は複数の縦線が隣接していて1本の線のように見えても、それぞれ別の線としてカウントされる。
 そして、この例に示されるように、解析用データの横線がなく、縦線が所定の範囲内に収束するのが理想的な状態であり、学習モデルの学習が適切に行われている状態である。
 具体的には、撮影画像の内容に関わらず画素値の分散が大きくなる特徴マップ(以下、高分散特徴マップと称する)が存在する場合、解析用データの高分散特徴マップに対応する列に縦線が現れる。高分散特徴マップは、撮影画像の内容に関わらず撮影画像の特徴を抽出し、学習モデルの物体の認識に寄与する可能性が高い特徴マップである。
 一方、撮影画像の内容に関わらず画素値の分散が小さくなる特徴マップ(以下、低分散特徴マップと称する)が存在する場合、解析用データの低分散特徴マップに対応する列には縦線が現れない。低分散特徴マップは、撮影画像の内容に関わらず撮影画像の特徴を抽出せず、学習モデルの物体の認識に寄与する可能性が低い特徴マップである。
 従って、高分散特徴マップが多くなるほど、解析用データの縦線の数が多くなり、高分散特徴マップが少なくなるほど、解析用データの縦線の数が少なくなる。
 なお、高分散特徴マップは、必ずしも学習モデルの認識対象となる物体(例えば、車両)の認識に寄与するとは限らず、学習が十分でない場合は認識対象以外の物体の認識に寄与する場合もある。ただし、学習モデルが適切に学習されていれば、高分散特徴マップは、学習モデルの認識対象となる物体の認識に寄与する特徴マップとなる。
 ここで、一般的に、ニューラルネットワークを用いた学習モデルの学習処理において、過学習の抑制やニューラルネットワークの軽量化等のために、正則化処理が行われる。この正則化処理により、撮影画像から抽出する特徴量の種類が、ある程度絞られる。すなわち、撮影画像の特徴を抽出することが可能な特徴マップ、すなわち、高分散特徴マップの数が、ある程度絞られる。
 例えば、図12に示されるように、学習の初期においては、高分散特徴マップが数多く存在するが、学習の末期には、高分散特徴マップの数が絞り込まれる。
 従って、例えば、図13の矢印で示されるように、解析用データの縦線の数が所定の範囲内に収束するのが、正則化が正常に行われている理想的な状態である。
 一方、図14に示されるように、学習が進むにつれて、解析用データの縦線の数が減少しすぎたり、0になったりするのは、正則化が強すぎて、撮影画像の特徴を十分に抽出できていない状態である。また、図15に示されるように、学習が進んでも、解析用データの縦線の数が減少しないのは、正則化が弱すぎて、撮影画像から抽出する特徴の種類を絞り切れていない状態である。
 そこで、解析部323は、解析用データの縦線の数が所定の範囲内に収束した場合、学習処理の正則化が正常に行われていると判定する。一方、解析部323は、解析用データの縦線の数が所定の範囲を超える値に収束した場合、又は、縦線の数が収束しない場合、正則化が弱すぎると判定する。また、解析部323は、解析用データの縦線の数が所定の範囲未満の値に収束した場合、正則化が強すぎると判定する。
 なお、正則化処理を行わない場合であっても、学習が進むにつれて、解析用データの縦線の数は減少し、収束する。この場合、解析用データの縦線の数は、正則化処理を行わない場合と比較して大きな値で収束するが、正則化処理を行う場合と同様に、縦線の数に基づいて学習処理が正常に行われているか否かを判定することが可能である。
 また、ほとんどの特徴マップの画素値の分散が大きくなる撮影画像(以下、高分散撮影画像と称する)が存在する場合、解析用データの高分散撮影画像に対応する行に横線が現れる。一方、ほとんどの特徴マップの画素値の分散が小さくなる撮影画像(以下、低分散撮影画像と称する)が存在する場合、解析用データの低分散撮影画像に対応する行に横線は現れない。
 従って、高分散撮影画像と低分散撮影画像が混在している場合、解析用データに横線が現れる。これは、各特徴マップの役割分担が十分にできておらず、すなわち、各特徴マップが抽出する特徴が明確に分かれておらず、撮影画像の内容により、各特徴マップが抽出する特徴にバラツキがある状態を示している。
 従って、解析用データの横線の数は少ないほどよく、横線の数が0となるのが最も理想的な状態である。
 そこで、解析部323は、解析用データの横線の数が所定の閾値未満の値に収束した場合、学習処理が正常に行われていると判定する。一方、解析部323は、解析用データの横線の数が所定の閾値以上の値に収束した場合、又は、横線の数が収束しない場合、学習処理が正常に行われていないと判定する。
 そして、解析部323は、学習状況の解析結果を示すデータをパラメータ設定部324に供給する。
 ステップS4において、パラメータ設定部324は、解析結果に基づいて、学習処理のパラメータを調整する。例えば、パラメータ設定部324は、解析部323により正則化が強すぎると判定された場合、正則化処理に用いる正則化パラメータの値を現在の値より小さくする。一方、パラメータ設定部324は、解析部323により正則化が弱すぎると判定された場合、正則化パラメータの値を現在の値より大きくする。なお、正則化パラメータの値が大きくなるほど、正則化が強くなり、正則化パラメータの値が小さくなるほど、正則化が弱くなる。パラメータ設定部324は、調整後の正則化パラメータを示すデータを学習部312に供給する。
 学習部312は、正則化パラメータの値をパラメータ設定部324により設定された値に設定する。これにより、学習モデルの学習処理において、正則化がより正常に行われるようになる。
 なお、例えば、ユーザが、出力部315において表示された解析用データを参照して、正則化パラメータ等の学習用のパラメータの調整を行うようにしてもよい。
 ステップS5において、学習状況解析部313は、学習部312による学習処理が終了したか否かを判定する。まだ学習処理が終了していないと判定された場合、処理はステップS1に戻る。その後、ステップS5において、学習処理が終了したと判定されるまで、ステップS1乃至ステップS5の処理が繰り返し実行される。
 一方、ステップS5において、学習処理が終了したと判定された場合、学習状況解析処理は終了する。
 以上のようにして、学習部312による学習モデルの学習状況を解析することができる。また、解析結果に基づいて、学習処理のパラメータを適切に設定し、学習精度を高めたり、学習時間を短縮したりすることができる。
 また、ユーザは、解析用データを視認することにより、学習モデルの学習状況を容易に認識することができる。
  <<2.学習状況の解析事例>>
 次に、図16乃至図22を参照して、図4の情報処理装置301を用いて学習モデルの学習状況を解析した事例について説明する。
 なお、以上の説明では、撮影画像に基づいて車両の認識処理を行う物体認識モデル201の学習状況を解析する例を示したが、情報処理装置301の解析対象となる学習モデルが用いる画像データの種類や認識対象となる物体の種類は、特に限定されない。以下では、車両10の前方を撮影した撮影画像、及び、車両10の前方を監視するミリ波レーダから出力されるミリ波画像に基づいて、車両10の前方の車両の認識処理を行う物体認識モデルの学習状況を解析した事例について説明する。
  <物体認識モデル401の構成例>
 図16は、学習状況の解析を行う対象となった学習モデルである物体認識モデル401の構成例を示している。物体認識モデル401は、例えば、図1の車両制御システム100の車外情報検出部141に用いられる。
 物体認識モデル401には、例えば、車両10の前方を撮影した画像データである撮影画像(例えば、撮影画像402)、及び、車両10の前方を監視するミリ波レーダから出力されるミリ波画像(例えば、ミリ波画像403)が入力データとして入力される。なお、ミリ波画像は、例えば、車両10の前方の物体により反射されたミリ波レーダの受信信号の強度の分布を鳥瞰図により表した画像データである。そして、物体認識モデル401は、撮影画像及びミリ波画像に基づいて、車両10の前方の車両の認識処理を行い、認識結果を示す画像データである出力画像(例えば、出力画像404)を出力データとして出力する。
 物体認識モデル401は、DSSD(Deconvolutional Single Shot Detector)を用いた学習モデルである。物体認識モデル401は、特徴抽出層411、特徴抽出層412、結合部413、及び、予測層414を備える。
 特徴抽出層411は、図2の物体認識モデル201の特徴抽出層211と同様の構成を有している。特徴抽出層411の各階層は、それぞれ所定の演算により、撮影画像の特徴を示す特徴マップを生成し、次の階層に供給する。また、特徴抽出層411の一部の階層は、特徴マップを結合部413に供給する。
 特徴抽出層412は、階層構造を有しており、各階層は、畳み込み層、プーリング層等からなる。特徴抽出層412の各階層は、それぞれ所定の演算により、ミリ波画像の特徴を示す特徴マップを生成し、次の階層に供給する。また、特徴抽出層412の一部の階層は、特徴マップを結合部413に供給する。さらに、特徴抽出層412は、ミリ波画像を撮影画像と同じカメラ座標系の画像に変換する。
 結合部413は、特徴抽出層411と特徴抽出層412の互いに対応する階層から出力される特徴マップを結合して、予測層414に供給する。
 予測層414は、結合部413から供給される特徴マップに基づいて、車両10の前方の車両の認識処理を行う。予測層414は、車両の認識結果を示す出力画像を出力する。
 なお、物体認識モデル401は、カメラ用ネットワーク421、ミリ波レーダ用ネットワーク422、及び、結合用ネットワーク423に分かれる。
 カメラ用ネットワーク421は、結合対象とならない特徴マップを撮影画像から生成する、特徴抽出層411の前半部分を含む。
 ミリ波レーダ用ネットワーク422は、結合対象とならない特徴マップをミリ波画像から生成する、特徴抽出層412の前半部分を含む。
 結合用ネットワーク423は、結合対象となる特徴マップを生成する特徴抽出層411の後半部分、結合対象となる特徴マップを生成する特徴抽出層412の後半部分、結合部413、及び、予測層414を含む。
 ここで、物体認識モデル401の学習処理を行い、学習後の物体認識モデル401を用いて車両10の前方の車両の認識処理を行ったところ、何も認識できずに、車両の認識に失敗する結果に終わった。
 そこで、図5を参照して上述した学習状況解析処理により、物体認識モデル401の学習状況の解析を行った。具体的には、図16の太い枠線で囲まれている特徴抽出層411の中間層である階層431で生成され、出力される特徴マップ、及び、特徴抽出層412の中間層である階層432で生成され、出力される特徴マップの解析処理を行った。なお、階層431で生成される特徴マップは、予測層414の認識処理に用いられる特徴抽出層411の特徴マップのうち、最もサイズの大きい特徴マップである。階層432で生成される特徴マップは、予測層414の認識処理に用いられる特徴抽出層412の特徴マップのうち、最もサイズの大きい特徴マップである。
 そして、特徴抽出層411の階層431で生成される特徴マップに基づく解析用データは、先に示した図11に示されるように、学習処理が進むにつれて、横線が消え、縦線の数が所定の範囲内に収束した。
 一方、特徴抽出層412の階層432で生成される特徴マップに基づく解析用データは、図17のA乃至Eに示されるように、学習処理が進んでも、横線が消えずに残った。すなわち、特徴抽出層412において、学習が適切に行われておらず、ミリ波画像の特徴が適切に抽出されていないことが分かった。なお、図17のA乃至Eの解析用データは、図11のA乃至Eと同様に、学習の進行順に並べられている。図18は、図17のEの解析用データを拡大したものである。
 そこで、解析用データに横線が現れる原因を探るために、横線が現れた行に対応する撮影画像及びミリ波画像が調査された。
 図19乃至図21は、図18の矢印で示される、解析用データの横線が現れた行に対応するミリ波画像及び撮影画像を模式的に示している。図19乃至図21のAは、ミリ波画像をグレースケールの画像に変換したものであり、図19乃至図21のBは、撮影画像を線図にしたものである。
 そして、図19乃至図21のAの撮影画像の例に示されるように、解析用データの横線が現れている行では、全て車両10が高速道路の中央車線を走行していることが分かった。
 また、図22は、解析用データの横線が現れている行の特徴マップの例を示している。ミリ波画像501は、解析用データの横線が現れている行のミリ波画像をグレースケールの画像に変換したものであり、特徴マップ502-1乃至特徴マップ502-4は、ミリ波画像501から生成された特徴マップの一部である。
 特徴マップ502-1乃至特徴マップ502-4では、車両10の前方の左右の端付近において画素値が大きく変化している。これにより、ミリ波画像から生成された特徴マップでは、車両10の前方の車両ではなく、車両の前方の左右の壁等に対応する特徴が抽出されやすいことが分かる。
 これにより、特徴抽出層412は、車両10の前方の車両よりも、左右の壁等の認識に適している可能性があることが分かった。
 このように、解析用データを用いることにより、物体認識モデル401が車両の認識に失敗する原因を容易に特定することができた。
 <<3.変形例>>
 以下、上述した本技術の実施の形態の変形例について説明する。
 解析用データに用いる特徴データは、上述した特徴マップの画素値の分散に限定されるものではなく、特徴マップの特徴を表す他の数値を用いることが可能である。
 例えば、特徴マップの画素値の平均値、最大値、中央値等を特徴データに用いることが可能である。
 また、例えば、次式(3)により計算される特徴マップのノルムを特徴データに用いても良い。
Figure JPOXMLDOC01-appb-M000002
 式(3)のnormは、特徴マップのノルムを示している。n、mは、任意の数を示している。他の記号は、上述した式(1)と同様である。
 特徴マップのノルムは、特徴マップの画素値の平均値に対する各画素の画素値の散らばり具合が大きくなるほど大きくなり、特徴マップの画素値の平均値に対する各画素の画素値の散らばり具合が小さくなるほど小さくなる。従って、特徴マップのノルムは、特徴マップの画素値の散らばり具合を示す。
 さらに、例えば、特徴マップの画素値の度数分布を特徴データに用いてもよい。
 図23は、特徴マップの画素値のヒストグラム(度数分布)の例を示している。横軸は、特徴マップの画素値に基づく階級を示している。すなわち、特徴マップの画素値が複数の階級に分類されている。縦軸は、度数を示している。すなわち、画素値が各階級に属する特徴マップの画素の個数を示している。
 そして、複数の入力画像(例えば、上述した撮影画像又はミリ波画像等)から生成された複数の特徴マップに基づいて、図24に示される解析用データが生成される。
 図24の解析用データ521は、ヒストグラムの階級毎に生成される2次元の度数マップ522-1乃至度数マップ522-mをz軸方向(奥行き方向)に並べた3次元のデータである。
 具体的には、度数マップ522-1のx軸方向(横方向)は、特徴マップの番号を示し、y軸方向は、入力画像の番号を示している。
 度数マップ522-1では、各入力画像の特徴マップのヒストグラムの1番目の階級の度数が、対応する特徴マップの番号順にx軸方向(横方向)に並べられている。また、異なる入力画像の同じ番号の特徴マップに対応する度数が、対応する入力画像の番号順にy軸方向(縦方向)に並べられている。
 度数マップ522-2乃至度数マップ522-mについても同様に、各撮影画像の特徴マップのヒストグラムのi番目(i=2~m)の階級の度数が、x軸方向及びy軸方向に並べられている。
 度数マップ522-1乃至度数マップ522-mは、例えば、それぞれ上述した2次元の解析用データと同様に、縦線及び横線に基づいて、学習状況の解析に用いることができる。また、例えば、解析用データ521は、z軸方向の線に基づいて、学習状況の解析に用いることができる。
 また、例えば、1以上の画像データから生成される複数の特徴マップに基づく特徴データのうち、所定の条件を満たす特徴データのみを抽出し、抽出した特徴データを含む解析用データを生成するようにしてもよい。例えば、特徴データが特徴マップの画素値の分散である場合、1つの画像データから生成された複数の特徴マップに基づく特徴データのうち、値が所定の閾値以上の特徴データを抽出し、抽出した特徴データを含む解析用データを生成するようにしてもよい。この場合、解析用データには、画素値の分散が所定の閾値以上の特徴マップの特徴データ(=画素値の分散)のみが含まれる。
 さらに、例えば、1以上の画像データに対してニューラルネットワークの異なる階層において生成された複数の特徴マップに基づいて解析用データを生成するようにしてもよい。例えば、各階層において生成された特徴マップの特徴データをx軸方向及びy軸方向の2次元に配列した階層毎のデータを、z軸方向に積層することにより、3次元の解析用データを生成するようにしてもよい。すなわち、この解析用データにおいては、x軸方向及びy軸方向に、同じ階層において生成された特徴マップの特徴データが並べられ、z軸方向に、異なる階層において生成された特徴マップの特徴データが並べられる。
 また、以上の説明では、学習処理と並行して学習状況の解析処理を行う例を示したが、例えば、学習状況の解析処理は、学習処理の終了後に行うことも可能である。例えば、学習処理中に生成される特徴マップを蓄積しておき、学習処理後に蓄積された特徴マップに基づいて、解析用データを生成して、学習状況の解析を行うようにしてもよい。
 さらに、学習処理の解析対象となる学習モデルは、上述した例に限定されず、ニューラルネットワークを用いた学習モデル全般を対象とすることができる。例えば、車両以外の物体を認識する認識モデル、及び、車両を含む複数の物体を認識する認識モデルも対象となる。また、入力データが画像データ以外の学習モデルも対象となる。例えば、音声データを入力データとする音声認識モデルや、文章データを入力データとする文章解析モデル等も対象となる。
 <<4.その他>>
  <コンピュータの構成例>
 上述した一連の処理は、ハードウェアにより実行することもできるし、ソフトウェアにより実行することもできる。一連の処理をソフトウェアにより実行する場合には、そのソフトウェアを構成するプログラムが、コンピュータにインストールされる。ここで、コンピュータには、専用のハードウェアに組み込まれているコンピュータや、各種のプログラムをインストールすることで、各種の機能を実行することが可能な、例えば汎用のパーソナルコンピュータなどが含まれる。
 図25は、上述した一連の処理をプログラムにより実行するコンピュータのハードウェアの構成例を示すブロック図である。
 コンピュータ1000において、CPU(Central Processing Unit)1001,ROM(Read Only Memory)1002,RAM(Random Access Memory)1003は、バス1004により相互に接続されている。
 バス1004には、さらに、入出力インタフェース1005が接続されている。入出力インタフェース1005には、入力部1006、出力部1007、記録部1008、通信部1009、及びドライブ1010が接続されている。
 入力部1006は、入力スイッチ、ボタン、マイクロフォン、撮像素子などよりなる。出力部1007は、ディスプレイ、スピーカなどよりなる。記録部1008は、ハードディスクや不揮発性のメモリなどよりなる。通信部1009は、ネットワークインタフェースなどよりなる。ドライブ1010は、磁気ディスク、光ディスク、光磁気ディスク、又は半導体メモリなどのリムーバブルメディア1011を駆動する。
 以上のように構成されるコンピュータ1000では、CPU1001が、例えば、記録部1008に記録されているプログラムを、入出力インタフェース1005及びバス1004を介して、RAM1003にロードして実行することにより、上述した一連の処理が行われる。
 コンピュータ1000(CPU1001)が実行するプログラムは、例えば、パッケージメディア等としてのリムーバブルメディア1011に記録して提供することができる。また、プログラムは、ローカルエリアネットワーク、インターネット、デジタル衛星放送といった、有線または無線の伝送媒体を介して提供することができる。
 コンピュータ1000では、プログラムは、リムーバブルメディア1011をドライブ1010に装着することにより、入出力インタフェース1005を介して、記録部1008にインストールすることができる。また、プログラムは、有線または無線の伝送媒体を介して、通信部1009で受信し、記録部1008にインストールすることができる。その他、プログラムは、ROM1002や記録部1008に、あらかじめインストールしておくことができる。
 なお、コンピュータが実行するプログラムは、本明細書で説明する順序に沿って時系列に処理が行われるプログラムであっても良いし、並列に、あるいは呼び出しが行われたとき等の必要なタイミングで処理が行われるプログラムであっても良い。
 また、本明細書において、システムとは、複数の構成要素(装置、モジュール(部品)等)の集合を意味し、すべての構成要素が同一筐体中にあるか否かは問わない。したがって、別個の筐体に収納され、ネットワークを介して接続されている複数の装置、及び、1つの筐体の中に複数のモジュールが収納されている1つの装置は、いずれも、システムである。
 さらに、本技術の実施の形態は、上述した実施の形態に限定されるものではなく、本技術の要旨を逸脱しない範囲において種々の変更が可能である。
 例えば、本技術は、1つの機能をネットワークを介して複数の装置で分担、共同して処理するクラウドコンピューティングの構成をとることができる。
 また、上述のフローチャートで説明した各ステップは、1つの装置で実行する他、複数の装置で分担して実行することができる。
 さらに、1つのステップに複数の処理が含まれる場合には、その1つのステップに含まれる複数の処理は、1つの装置で実行する他、複数の装置で分担して実行することができる。
  <構成の組み合わせ例>
 本技術は、以下のような構成をとることもできる。
(1)
 ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データを生成する特徴データ生成ステップと、
 複数の前記特徴マップの前記特徴データに基づく解析用データを生成する解析用データ生成ステップと
 を含む情報処理方法。
(2)
 前記解析用データには、複数の前記特徴マップの前記特徴データが並べられている
 前記(1)に記載の情報処理方法。
(3)
 前記解析用データには、1つの前記入力データから複数の前記特徴マップを生成する前記モデルの所定の階層において複数の前記入力データから生成された複数の前記特徴マップの前記特徴データが並べられている
 前記(2)に記載の情報処理方法。
(4)
 前記解析用データの第1の方向には、同じ前記入力データから生成された複数の前記特徴マップの前記特徴データが並べられ、前記第1の方向と直交する第2の方向には、異なる前記入力データの互いに対応する前記特徴マップの前記特徴データが並べられている
 前記(3)に記載の情報処理方法。
(5)
 前記特徴データは、前記特徴マップの画素値の散らばり具合を示す
 前記(4)に記載の情報処理方法。
(6)
 前記解析用データの前記第1の方向の線及び前記第2の方向の線に基づいて、前記モデルの学習状況の解析を行う解析ステップを
 さらに含む前記(5)に記載の情報処理方法。
(7)
 前記モデルの学習状況の解析結果に基づいて、前記モデルの学習用のパラメータを設定するパラメータ設定ステップを
 さらに含む前記(6)に記載の情報処理方法。
(8)
 前記パラメータ設定ステップにおいて、前記解析用データの前記第2の方向の線の数に基づいて、前記モデルの学習用の正則化パラメータが設定される
 前記(7)に記載の情報処理方法。
(9)
 前記特徴データは、前記特徴マップの画素値の度数分布を示し、
 前記解析用データには、前記モデルの前記階層において複数の前記入力データから生成された複数の前記特徴マップの前記特徴データが3次元に並べられている
 前記(3)に記載の情報処理方法。
(10)
 前記入力データは、画像データであり、
 前記モデルは、物体の認識処理を行う
 前記(1)乃至(9)のいずれかに記載の情報処理方法。
(11)
 前記モデルは、車両の認識処理を行う
 前記(10)に記載の情報処理方法。
(12)
 前記入力データは、ミリ波レーダの受信信号の強度の分布を鳥瞰図により表す画像データである
 前記(11)に記載の情報処理方法。
(13)
 前記モデルは、前記画像データをカメラ座標系の画像に変換する
 前記(12)に記載の情報処理方法。
(14)
 前記解析用データは、複数の前記特徴マップの前記特徴データのうち所定の条件を満たす前記特徴データを含む
 前記(1)に記載の情報処理方法。
(15)
 ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データを生成する特徴データ生成ステップと、
 複数の前記特徴マップの前記特徴データを並べた解析用データを生成する解析用データ生成ステップと
 を含む処理をコンピュータに実行させるためのプログラム。
(16)
 ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データを生成する特徴データ生成部と、
 複数の前記特徴マップの前記特徴データを並べた解析用データを生成する解析用データ生成部と
 を備える情報処理装置。
 なお、本明細書に記載された効果はあくまで例示であって限定されるものではなく、他の効果があってもよい。
 10 車両, 100 車両制御システム, 141 車外情報検出部, 201 物体認識モデル, 211 特徴抽出層, 221 階層, 301 情報処理装置, 312 学習部, 313 学習状況解析部, 321 特徴データ生成部, 322 解析用データ生成部, 323 解析部, 324 パラメータ設定部, 401 物体認識モデル 411,412 特徴抽出層, 421 カメラ用ネットワーク, 422 ミリ波レーダ用ネットワーク, 423 結合用ネットワーク, 431,432 階層

Claims (16)

  1.  ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データを生成する特徴データ生成ステップと、
     複数の前記特徴マップの前記特徴データに基づく解析用データを生成する解析用データ生成ステップと
     を含む情報処理方法。
  2.  前記解析用データには、複数の前記特徴マップの前記特徴データが並べられている
     請求項1に記載の情報処理方法。
  3.  前記解析用データには、1つの前記入力データから複数の前記特徴マップを生成する前記モデルの所定の階層において複数の前記入力データから生成された複数の前記特徴マップの前記特徴データが並べられている
     請求項2に記載の情報処理方法。
  4.  前記解析用データの第1の方向には、同じ前記入力データから生成された複数の前記特徴マップの前記特徴データが並べられ、前記第1の方向と直交する第2の方向には、異なる前記入力データの互いに対応する前記特徴マップの前記特徴データが並べられている
     請求項3に記載の情報処理方法。
  5.  前記特徴データは、前記特徴マップの画素値の散らばり具合を示す
     請求項4に記載の情報処理方法。
  6.  前記解析用データの前記第1の方向の線及び前記第2の方向の線に基づいて、前記モデルの学習状況の解析を行う解析ステップを
     さらに含む請求項5に記載の情報処理方法。
  7.  前記モデルの学習状況の解析結果に基づいて、前記モデルの学習用のパラメータを設定するパラメータ設定ステップを
     さらに含む請求項6に記載の情報処理方法。
  8.  前記パラメータ設定ステップにおいて、前記解析用データの前記第2の方向の線の数に基づいて、前記モデルの学習用の正則化パラメータが設定される
     請求項7に記載の情報処理方法。
  9.  前記特徴データは、前記特徴マップの画素値の度数分布を示し、
     前記解析用データには、前記モデルの前記階層において複数の前記入力データから生成された複数の前記特徴マップの前記特徴データが3次元に並べられている
     請求項3に記載の情報処理方法。
  10.  前記入力データは、画像データであり、
     前記モデルは、物体の認識処理を行う
     請求項1に記載の情報処理方法。
  11.  前記モデルは、車両の認識処理を行う
     請求項10に記載の情報処理方法。
  12.  前記入力データは、ミリ波レーダの受信信号の強度の分布を鳥瞰図により表す画像データである
     請求項11に記載の情報処理方法。
  13.  前記モデルは、前記画像データをカメラ座標系の画像に変換する
     請求項12に記載の情報処理方法。
  14.  前記解析用データは、複数の前記特徴マップの前記特徴データのうち所定の条件を満たす前記特徴データを含む
     請求項1に記載の情報処理方法。
  15.  ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データを生成する特徴データ生成ステップと、
     複数の前記特徴マップの前記特徴データに基づく解析用データを生成する解析用データ生成ステップと
     を含む処理をコンピュータに実行させるためのプログラム。
  16.  ニューラルネットワークを用いたモデルにおいて入力データから生成される特徴マップの特徴を数値で表す特徴データを生成する特徴データ生成部と、
     複数の前記特徴マップの前記特徴データに基づく解析用データを生成する解析用データ生成部と
     を備える情報処理装置。
PCT/JP2020/011601 2019-03-29 2020-03-17 情報処理方法、プログラム、及び、情報処理装置 Ceased WO2020203241A1 (ja)

Priority Applications (7)

Application Number Priority Date Filing Date Title
MX2021011219A MX2021011219A (es) 2019-03-29 2020-03-17 Metodo de procesamiento de informacion, programa y dispositivo de procesamiento de informacion.
KR1020217026682A KR20210142604A (ko) 2019-03-29 2020-03-17 정보 처리 방법, 프로그램 및 정보 처리 장치
EP20784531.4A EP3951663B1 (en) 2019-03-29 2020-03-17 Information processing method, program, and information processing device
US17/442,075 US12243320B2 (en) 2019-03-29 2020-03-17 Information processing method, program, and information processing apparatus
ES20784531T ES3010476T3 (en) 2019-03-29 2020-03-17 Information processing method, program, and information processing device
JP2021511392A JP7487178B2 (ja) 2019-03-29 2020-03-17 情報処理方法、プログラム、及び、情報処理装置
CA3134088A CA3134088A1 (en) 2019-03-29 2020-03-17 Information processing method, program, and information processing apparatus

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2019065378 2019-03-29
JP2019-065378 2019-03-29

Publications (1)

Publication Number Publication Date
WO2020203241A1 true WO2020203241A1 (ja) 2020-10-08

Family

ID=72668318

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2020/011601 Ceased WO2020203241A1 (ja) 2019-03-29 2020-03-17 情報処理方法、プログラム、及び、情報処理装置

Country Status (8)

Country Link
US (1) US12243320B2 (ja)
EP (1) EP3951663B1 (ja)
JP (1) JP7487178B2 (ja)
KR (1) KR20210142604A (ja)
CA (1) CA3134088A1 (ja)
ES (1) ES3010476T3 (ja)
MX (1) MX2021011219A (ja)
WO (1) WO2020203241A1 (ja)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114863472A (zh) * 2022-03-28 2022-08-05 深圳海翼智新科技有限公司 多级行人检测方法、装置和存储介质

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US12392628B2 (en) * 2020-06-30 2025-08-19 Lyft, Inc. Localization based on multi-collect fusion
US12087004B2 (en) 2020-06-30 2024-09-10 Lyft, Inc. Multi-collect fusion
US12523478B2 (en) 2022-10-13 2026-01-13 Woven By Toyota, Inc. Map data compression methods implementing machine learning
CN116739437A (zh) * 2023-07-14 2023-09-12 鱼快创领智能科技(南京)有限公司 一种基于车联网数据的综合运力分级方法

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2018520404A (ja) * 2015-04-28 2018-07-26 クゥアルコム・インコーポレイテッドQualcomm Incorporated ニューラルネットワークのためのトレーニング基準としてのフィルタ特異性
WO2018165753A1 (en) * 2017-03-14 2018-09-20 University Of Manitoba Structure defect detection using machine learning algorithms

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3058951B2 (ja) 1991-09-03 2000-07-04 浜松ホトニクス株式会社 画像パターン認識装置
KR102592076B1 (ko) * 2015-12-14 2023-10-19 삼성전자주식회사 딥러닝 기반 영상 처리 장치 및 방법, 학습 장치
WO2017192889A1 (en) 2016-05-04 2017-11-09 Intel Corporation Antenna panel switching and beam indication
WO2018212606A1 (ko) 2017-05-17 2018-11-22 엘지전자(주) 무선 통신 시스템에서 하향링크 채널을 수신하는 방법 및 이를 위한 장치
KR102419136B1 (ko) * 2017-06-15 2022-07-08 삼성전자주식회사 다채널 특징맵을 이용하는 영상 처리 장치 및 방법
US20190026588A1 (en) * 2017-07-19 2019-01-24 GM Global Technology Operations LLC Classification methods and systems
JP7160040B2 (ja) * 2017-08-22 2022-10-25 ソニーグループ株式会社 信号処理装置、および信号処理方法、プログラム、移動体、並びに、信号処理システム
EP3617947B1 (en) * 2018-08-30 2026-04-22 Nokia Technologies Oy Apparatus and method for processing image data

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2018520404A (ja) * 2015-04-28 2018-07-26 クゥアルコム・インコーポレイテッドQualcomm Incorporated ニューラルネットワークのためのトレーニング基準としてのフィルタ特異性
WO2018165753A1 (en) * 2017-03-14 2018-09-20 University Of Manitoba Structure defect detection using machine learning algorithms

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
MATTHEW D ZEILER; ROB FERGUS: "Visualizing and Understanding Convolutional Networks", EUROPEAN CONFERENCE ON COMPUTER VISION (ECCV), 2014, ECCV, 2014, pages 818 - 833, XP055566692, Retrieved from the Internet <URL:https://cs.nyu.edu/~fergus/papers/zeilerECCV2014.pdf> *
See also references of EP3951663A4 *

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114863472A (zh) * 2022-03-28 2022-08-05 深圳海翼智新科技有限公司 多级行人检测方法、装置和存储介质

Also Published As

Publication number Publication date
MX2021011219A (es) 2021-10-22
CA3134088A1 (en) 2020-10-08
EP3951663A1 (en) 2022-02-09
ES3010476T3 (en) 2025-04-03
US12243320B2 (en) 2025-03-04
US20220172484A1 (en) 2022-06-02
EP3951663B1 (en) 2024-12-25
KR20210142604A (ko) 2021-11-25
JP7487178B2 (ja) 2024-05-20
JPWO2020203241A1 (ja) 2020-10-08
EP3951663A4 (en) 2022-05-04

Similar Documents

Publication Publication Date Title
JP7757353B2 (ja) 情報処理装置、情報処理方法、及び、情報処理システム
US11531354B2 (en) Image processing apparatus and image processing method
US11450026B2 (en) Information processing apparatus, information processing method, and mobile object
US12511762B2 (en) Learning model generation method, information processing device, and information processing system
US20210116930A1 (en) Information processing apparatus, information processing method, program, and mobile object
JPWO2019077999A1 (ja) 撮像装置、画像処理装置、及び、画像処理方法
JP7487178B2 (ja) 情報処理方法、プログラム、及び、情報処理装置
JPWO2020116194A1 (ja) 情報処理装置、情報処理方法、プログラム、移動体制御装置、及び、移動体
WO2021241189A1 (ja) 情報処理装置、情報処理方法、およびプログラム
JPWO2019082670A1 (ja) 情報処理装置、情報処理方法、プログラム、及び、移動体
JPWO2019188391A1 (ja) 制御装置、制御方法、並びにプログラム
US20240290108A1 (en) Information processing apparatus, information processing method, learning apparatus, learning method, and computer program
WO2021193099A1 (ja) 情報処理装置、情報処理方法、及びプログラム
JPWO2020116204A1 (ja) 情報処理装置、情報処理方法、プログラム、移動体制御装置、及び、移動体
US20240257508A1 (en) Information processing device, information processing method, and program
US20230206596A1 (en) Information processing device, information processing method, and program
US20230121905A1 (en) Information processing apparatus, information processing method, and program
WO2021033574A1 (ja) 情報処理装置、および情報処理方法、並びにプログラム
JP7676407B2 (ja) 情報処理装置、情報処理方法、及び、プログラム
WO2021145227A1 (ja) 情報処理装置、情報処理方法、及び、プログラム
US20250172950A1 (en) Information processing apparatus, information processing method, information processing program, and mobile apparatus
WO2024009829A1 (ja) 情報処理装置、情報処理方法および車両制御システム
US12633133B2 (en) Recognition processing device, recognition processing method, and recognition processing system
US20240386724A1 (en) Recognition processing device, recognition processing method, and recognition processing system
US20250128732A1 (en) Information processing apparatus, information processing method, and moving apparatus

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 20784531

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2021511392

Country of ref document: JP

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 3134088

Country of ref document: CA

NENP Non-entry into the national phase

Ref country code: DE

ENP Entry into the national phase

Ref document number: 2020784531

Country of ref document: EP

Effective date: 20211029

WWG Wipo information: grant in national office

Ref document number: MX/A/2021/011219

Country of ref document: MX

WWG Wipo information: grant in national office

Ref document number: 17442075

Country of ref document: US

WWW Wipo information: withdrawn in national office

Ref document number: 1020217026682

Country of ref document: KR