WO2019187834A1 - 情報処理装置、情報処理方法、およびプログラム - Google Patents
情報処理装置、情報処理方法、およびプログラム Download PDFInfo
- Publication number
- WO2019187834A1 WO2019187834A1 PCT/JP2019/006580 JP2019006580W WO2019187834A1 WO 2019187834 A1 WO2019187834 A1 WO 2019187834A1 JP 2019006580 W JP2019006580 W JP 2019006580W WO 2019187834 A1 WO2019187834 A1 WO 2019187834A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- target sound
- autonomous mobile
- mobile body
- target
- sound
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B25—HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
- B25J—MANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
- B25J13/00—Controls for manipulators
- B25J13/003—Controls for manipulators by means of an audio-responsive input
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63H—TOYS, e.g. TOPS, DOLLS, HOOPS OR BUILDING BLOCKS
- A63H11/00—Self-movable toy figures
- A63H11/18—Figure toys which perform a realistic walking motion
- A63H11/20—Figure toys which perform a realistic walking motion with pairs of legs, e.g. horses
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/083—Recognition networks
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/20—Speech recognition techniques specially adapted for robustness in adverse environments, e.g. in noise, of stress induced speech
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63H—TOYS, e.g. TOPS, DOLLS, HOOPS OR BUILDING BLOCKS
- A63H11/00—Self-movable toy figures
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63H—TOYS, e.g. TOPS, DOLLS, HOOPS OR BUILDING BLOCKS
- A63H11/00—Self-movable toy figures
- A63H11/18—Figure toys which perform a realistic walking motion
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63H—TOYS, e.g. TOPS, DOLLS, HOOPS OR BUILDING BLOCKS
- A63H2200/00—Computerized interactive toys, e.g. dolls
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63H—TOYS, e.g. TOPS, DOLLS, HOOPS OR BUILDING BLOCKS
- A63H30/00—Remote-control arrangements specially adapted for toys, e.g. for toy vehicles
- A63H30/02—Electrical arrangements
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
- G10L2015/223—Execution procedure of a spoken command
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/03—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
- G10L25/18—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/78—Detection of presence or absence of voice signals
- G10L25/84—Detection of presence or absence of voice signals for discriminating voice from noise
Definitions
- the present disclosure relates to an information processing apparatus, an information processing method, and a program.
- Patent Literature 1 discloses a technique for moving a robot apparatus in a direction in which a user's utterance or face is recognized.
- Patent Document 1 does not consider the presence of sounds other than the user's utterance, that is, noise. For this reason, when the robot apparatus is simply brought close to the estimated user's direction, the noise input level increases, and it may be difficult to recognize the user's utterance.
- the present disclosure proposes a new and improved information processing apparatus, information processing method, and program capable of causing an autonomous mobile body to execute an operation for further improving the accuracy of voice recognition.
- an operation control unit that controls the operation of an autonomous moving body that performs an action based on a recognition process, and the operation control unit detects a target sound that is a target voice of the voice recognition process
- An information processing apparatus is provided that moves the autonomous mobile body to a position where the input level of the non-target sound that is not the target sound is further reduced around the approach target determined based on the target sound.
- the processor includes controlling the operation of the autonomous mobile body that performs an action based on the recognition process, and the control detects a target sound that is a target voice of the voice recognition process. And moving the autonomous mobile body to a position where the input level of the non-target sound that is not the target voice is further reduced around the approach target determined based on the target sound. Is provided.
- the computer includes an operation control unit that controls an operation of the autonomous mobile body that performs an action based on the recognition process, and the operation control unit is a target sound that is a target voice of the voice recognition process. Is detected, and functions as an information processing apparatus that moves the autonomous mobile body to a position where the input level of the non-target sound that is not the target sound is further reduced around the approach target determined based on the target sound A program is provided.
- the autonomous mobile body it is possible to cause the autonomous mobile body to execute an operation for further improving the accuracy of voice recognition.
- a target sound for example, a user's utterance voice
- a non-target sound that is not a target voice.
- the power ratio that is, the SN ratio (Signal-to-Noise Ratio).
- Patent Document 1 does not consider non-target sounds, and only moves the robot apparatus in the direction in which the user's speech or face is recognized. For this reason, in the technique described in Patent Document 1, as a result of simultaneously approaching the user and a noise source that emits a non-target sound, the S / N ratio is lowered and the voice recognition accuracy is also lowered.
- the operation of the robot apparatus is controlled so as to approach the user by using the user's utterance or face as a trigger. For this reason, the robot apparatus described in Patent Literature 1 is likely to always follow a user existing in the vicinity, and it may be assumed that the user feels bothersome.
- An information processing apparatus, an information processing method, and a program according to an embodiment of the present disclosure have been conceived by focusing on the above points, and cause an autonomous mobile body to perform an operation that further improves the accuracy of voice recognition. Make it possible.
- the autonomous mobile body 10 is an information processing apparatus that executes situation estimation based on collected sensor information and autonomously selects and executes various operations according to the situation.
- One feature of the autonomous mobile body 10 is that, unlike a robot that simply performs an operation in accordance with a user instruction command, the autonomous mobile body 10 autonomously executes an operation that is estimated to be optimal for each situation.
- the autonomous mobile body 10 avoids input of a non-target sound that is not the target sound, for example, when the target sound that is the target sound of the voice recognition process, that is, the user's speech is not detected. Autonomous operation may be performed. According to this operation, it is possible to effectively increase the possibility that the accuracy of voice recognition related to the utterance is improved without always following the user and when the user's utterance is detected.
- the autonomous mobile body 10 when the target sound is detected, the autonomous mobile body 10 according to the embodiment of the present disclosure is located at a position where the input level of the non-target sound is further reduced around the approach target determined based on the target sound. You may move. That is, the autonomous mobile body 10 according to an embodiment of the present disclosure improves the SN ratio and effectively improves the voice recognition accuracy related to the user's utterance by performing a moving operation in consideration of the non-target sound. Is possible.
- the autonomous mobile body 10 determines and executes an autonomous operation by comprehensively judging its own state, surrounding environment, and the like, similarly to animals including humans.
- the autonomous mobile body 10 according to an embodiment of the present disclosure is clearly different from a passive device that executes a corresponding operation or process based on an instruction.
- the autonomous mobile body 10 may be an autonomous mobile robot that performs autonomous posture control in space and executes various operations.
- the autonomous mobile body 10 may be, for example, an autonomous mobile robot having a shape imitating an animal such as a human or a dog or an operation capability.
- the autonomous mobile body 10 may be a device such as a vehicle or an unmanned aerial vehicle having communication capability with a user, for example.
- the level of the shape, ability, desire, and the like of the autonomous mobile body 10 according to an embodiment of the present disclosure can be appropriately designed according to the purpose and role.
- FIG. 1 is a diagram illustrating a hardware configuration example of an autonomous mobile body 10 according to an embodiment of the present disclosure.
- the autonomous mobile body 10 is a dog-type quadruped walking robot having a head, a trunk, four legs, and a tail.
- the autonomous mobile body 10 includes two displays 510 on the head.
- the autonomous mobile body 10 includes various sensors.
- the autonomous mobile body 10 includes, for example, a microphone 515, a camera 520, a ToF (Time of Flight) sensor 525, a human sensor 530, a distance sensor 535, a touch sensor 540, an illuminance sensor 545, a foot button 550, and an inertial sensor 555.
- a microphone 515 a camera 520
- a ToF (Time of Flight) sensor 525 a human sensor 530
- a distance sensor 535 a touch sensor 540
- an illuminance sensor 545 a foot button 550
- an inertial sensor 555 an inertial sensor 555.
- the microphone 515 has a function of collecting ambient sounds.
- the above sounds include, for example, user's utterances and surrounding environmental sounds.
- the autonomous mobile body 10 may include four microphones on the head. By providing a plurality of microphones 515, it is possible to collect sound generated in the surroundings with high sensitivity and realize localization of the sound source.
- the camera 520 has a function of imaging the user and the surrounding environment.
- the autonomous mobile body 10 may include two wide-angle cameras at the nose tip and the waist.
- the wide-angle camera arranged at the tip of the nose captures an image corresponding to the front field of view of the autonomous mobile body 10 (that is, the field of view of the dog), and the wide-angle camera of the waist section Take an image.
- the autonomous mobile body 10 can extract a feature point of a ceiling based on an image captured by a wide-angle camera placed on the waist, and can realize SLAM (Simultaneous Localization and Mapping).
- SLAM Simultaneous Localization and Mapping
- the ToF sensor 525 has a function of detecting a distance from an object existing in front of the head.
- the ToF sensor 525 is provided at the nose of the head. According to the ToF sensor 525, it is possible to detect the distance to various objects with high accuracy, and it is possible to realize an operation according to the relative position with respect to an object or an obstacle including the user.
- the human sensor 530 has a function of detecting the location of a user or a pet raised by the user.
- the human sensor 530 is disposed on the chest, for example. According to the human sensor 530, it is possible to realize various operations on the moving object, for example, operations corresponding to emotions such as interest, fear, and surprise by detecting the moving object existing in the front. .
- the distance measuring sensor 535 has a function of acquiring the state of the front floor surface of the autonomous mobile body 10.
- the distance measuring sensor 535 is disposed on the chest, for example. According to the distance measuring sensor 535, the distance from the object existing on the front floor surface of the autonomous mobile body 10 can be detected with high accuracy, and an operation according to the relative position with the object can be realized.
- the touch sensor 540 has a function of detecting contact by the user.
- the touch sensor 540 is disposed at a site where the user is likely to touch the autonomous mobile body 10, such as the top of the head, under the chin, or the back.
- the touch sensor 540 may be, for example, a capacitance type or pressure sensitive type touch sensor. According to the touch sensor 540, it is possible to detect a contact action such as touching, stroking, hitting, and pressing by a user, and it is possible to perform an operation according to the contact action.
- the illuminance sensor 545 detects the illuminance of the space where the autonomous mobile body 10 is located.
- the illuminance sensor 545 may be arranged at the base of the tail on the back of the head. According to the illuminance sensor 545, it is possible to detect ambient brightness and perform an operation according to the brightness.
- the sole button 550 has a function of detecting whether or not the leg bottom surface of the autonomous mobile body 10 is in contact with the floor.
- the sole button 550 is disposed at a portion corresponding to the paws of the four legs. According to the sole button 550, contact or non-contact between the autonomous mobile body 10 and the floor surface can be detected, and for example, it is possible to grasp that the autonomous mobile body 10 is lifted by the user. .
- Inertial sensor 555 is a six-axis sensor that detects physical quantities such as the speed, acceleration, and rotation of the head and torso. In other words, inertial sensor 555 detects the X-axis, Y-axis, and Z-axis accelerations and angular velocities. Inertial sensors 555 are disposed on the head and the trunk, respectively. According to the inertial sensor 555, it is possible to detect the movements of the head and torso of the autonomous mobile body 10 with high accuracy and to realize operation control according to the situation.
- the autonomous mobile body 10 may further include various communication devices including a temperature sensor, a geomagnetic sensor, and a GNSS (Global Navigation Satellite System) signal receiver, for example.
- GNSS Global Navigation Satellite System
- FIG. 2 is a configuration example of the actuator 570 provided in the autonomous mobile body 10 according to an embodiment of the present disclosure.
- the autonomous mobile body 10 according to an embodiment of the present disclosure has a total of 22 rotational degrees of freedom, one for the ear and the tail, and one for the mouth, in addition to the rotation locations shown in FIG.
- the autonomous mobile body 10 has both the movement of tilting and tilting the neck by having three degrees of freedom in the head.
- the autonomous mobile body 10 can realize a natural and flexible operation closer to a real dog by reproducing the swing motion of the waist by the actuator 570 provided in the waist.
- the autonomous mobile body 10 may realize the above-described 22 rotational degrees of freedom by combining, for example, a uniaxial actuator and a biaxial actuator.
- a uniaxial actuator may be employed at the elbow or knee portion of the leg
- a biaxial actuator may be employed at the base of the shoulder or thigh.
- FIG. 3 and 4 are diagrams for describing the operation of the actuator 570 provided in the autonomous mobile body 10 according to the embodiment of the present disclosure.
- the actuator 570 can drive the movable arm 590 at an arbitrary rotational position and rotational speed by rotating the output gear by the motor 575.
- an actuator 570 includes a rear cover 571, a gear BOX cover 572, a control board 573, a gear BOX base 574, a motor 575, a first gear 576, a second gear 577, and an output gear. 578, a detection magnet 579, and two bearings 580 are provided.
- the actuator 570 may be, for example, a magnetic svGMR (spin-valve giant magnetoresistive).
- the control board 573 rotates the motor 575 based on the control by the main processor, whereby power is transmitted to the output gear 578 via the first gear 576 and the second gear 577 and the movable arm 590 is driven. Is possible.
- the position sensor provided on the control board 573 detects the rotation angle of the detection magnet 579 that rotates in synchronization with the output gear 578, thereby detecting the rotation angle of the movable arm 590, that is, the rotation position with high accuracy. Can do.
- the magnetic svGMR is excellent in durability because it is a non-contact type, and has an advantage that it is less affected by signal fluctuation due to distance fluctuation of the detection magnet 579 and the position sensor when used in the GMR saturation region.
- the configuration example of the actuator 570 provided in the autonomous mobile body 10 according to an embodiment of the present disclosure has been described above. According to said structure, it becomes possible to control the bending operation of the joint part with which the autonomous mobile body 10 is provided with high precision, and to detect the rotation position of a joint part correctly.
- FIG. 5 is a diagram for describing functions of the display 510 included in the autonomous mobile body 10 according to an embodiment of the present disclosure.
- the display 510 has a function of visually expressing eye movements and emotions of the autonomous mobile body 10. As shown in FIG. 5, the display 510 can express eyeball, pupil, and eyelid movements according to emotions and actions. The display 510 does not display characters, symbols, images that are not related to eye movements, and the like, thereby producing a natural motion close to an animal such as a real dog.
- the autonomous mobile body 10 includes two displays 510r and 510l corresponding to the right eye and the left eye, respectively.
- the displays 510r and 510l are realized by, for example, two independent OLEDs (Organic Light Emitting Diode). OLED makes it possible to reproduce the curved surface of the eyeball, as compared to the case where a pair of eyeballs are represented by a single flat display or the case where two eyeballs are each represented by two independent flat displays. Thus, a more natural exterior can be realized.
- the displays 510r and 510l can express the line of sight and emotion of the autonomous mobile body 10 as shown in FIG. 5 with high accuracy and flexibility.
- the user can intuitively grasp the state of the autonomous mobile body 10 from the movement of the eyeball displayed on the display 510.
- FIG. 6 is a diagram illustrating an operation example of the autonomous mobile body 10 according to an embodiment of the present disclosure.
- the operation of the joint unit and the eyeball of the autonomous mobile body 10 will be described. Therefore, the external structure of the autonomous mobile body 10 is shown in a simplified manner.
- the external structure of the autonomous mobile body 10 may be shown in a simplified manner, but the hardware configuration and exterior of the autonomous mobile body 10 according to an embodiment of the present disclosure are shown in the drawings. It is not limited to an example, It can design suitably.
- FIG. 7 is a diagram illustrating a functional configuration example of the autonomous mobile body 10 according to an embodiment of the present disclosure.
- an autonomous moving body 10 according to an embodiment of the present disclosure includes an input unit 110, a recognition unit 120, a surrounding environment estimation unit 130, a surrounding environment holding unit 140, an operation control unit 150, a driving unit 160, and an output. Part 170 is provided.
- the input unit 110 has a function of collecting various information related to the user and the surrounding environment.
- the input unit 110 collects, for example, user's utterances and environmental sounds generated around the user, image information related to the users and the surrounding environment, and various sensor information.
- the input unit 110 includes various sensors shown in FIG.
- the recognition unit 120 has a function of performing various recognitions regarding the state of the user, surrounding objects, and the autonomous mobile body 10 based on various information collected by the input unit 110.
- the recognition unit 120 performs human identification, face recognition, facial expression and line of sight recognition, voice recognition, object recognition, color recognition, shape recognition, marker recognition, obstacle recognition, step recognition, brightness recognition, and the like. Good.
- the surrounding environment estimation unit 130 has a function of generating and updating a noise map indicating a non-target sound occurrence state based on the sensor information collected by the input unit 110 and the recognition result by the recognition unit 120.
- the functions of the surrounding environment estimation unit 130 will be described in detail separately.
- the surrounding environment holding unit 140 has a function of holding the noise map generated and updated by the surrounding environment estimation unit 130.
- the operation control unit 150 performs an action plan based on the recognition result of the recognition unit 120 and the noise map held by the surrounding environment holding unit 140, and controls the operations of the drive unit 160 and the output unit 170 based on the action plan. .
- the operation control unit 150 performs, for example, rotation control of the actuator 570, display control of the display 510, audio output control by a speaker, and the like based on the above action plan.
- the functions of the operation control unit 150 according to an embodiment of the present disclosure will be described in detail separately.
- the drive unit 160 has a function of bending and stretching a plurality of joints included in the autonomous mobile body 10 based on control by the operation control unit 150. More specifically, the drive unit 160 drives the actuator 570 provided in each joint unit based on control by the operation control unit 150.
- the output unit 170 has a function of outputting visual information and sound information based on control by the operation control unit 150.
- the output unit 170 includes a display 510 and a speaker.
- the functional configuration of the autonomous mobile body 10 according to an embodiment of the present disclosure has been described above.
- the structure shown in FIG. 7 is an example to the last, and the function structure of the autonomous mobile body 10 which concerns on this embodiment is not limited to the example which concerns.
- the autonomous mobile body 10 according to an embodiment of the present disclosure may include, for example, a communication unit that communicates with an information processing server or another autonomous mobile body.
- the functions of the recognition unit 120, the surrounding environment estimation unit 130, the operation control unit 150, and the like may be realized as the functions of the information processing server.
- the information processing server executes various recognition processes, generation or update of a noise map, and action plan based on the sensor information collected by the input unit 110 of the autonomous mobile body 10, and the driving unit 160 of the autonomous mobile body 10 And the output unit 170 can be controlled.
- the functional configuration of the autonomous mobile body 10 according to an embodiment of the present disclosure can be flexibly modified according to specifications and operations.
- the autonomous mobile body 10 is autonomous so that the SN ratio related to the target sound and the non-target sound is improved in order to improve the accuracy of speech recognition related to the target sound. Perform the action.
- the S / N ratio has the strongest influence of the physical distance between the target sound source and the non-target sound source (hereinafter also referred to as noise source).
- the autonomous mobile body 10 effectively improves the S / N ratio by moving away from the non-target sound source as much as possible while approaching the target sound source instead of simply approaching the target sound source. be able to.
- FIG. 8A and FIG. 8B are diagrams for explaining an outline of the operation of the autonomous mobile body 10 according to the present embodiment.
- FIG. 8A shows an example of movement control by the comparison method according to the present embodiment.
- the comparison device 90 sets the shortest distance to the user U so that the input level (sound pressure level) of the target sound increases. It is approaching.
- the comparison device 90 does not consider the presence of the noise source NS that emits the non-target sound, and thus approaches the user U and simultaneously approaches the noise source NS.
- the input level of the non-target sound increases as well as the input level of the target sound.
- the effect of improving the S / N ratio is diminished, and the voice recognition accuracy related to the utterance of the user U may be reduced. Arise.
- FIG. 8B shows an example of movement control by the information processing method according to the present embodiment.
- the autonomous mobile body 10 according to the present embodiment moves in consideration of the presence of a noise source that emits a non-target sound when the voice of the user U, that is, a target sound is detected.
- the motion control unit 150 according to the present embodiment autonomously moves to a position where the input level of the non-target sound is further decreased around the approach target determined based on the target sound.
- the body 10 may be moved.
- the approach target may be the user U who has emitted the target sound, that is, the speech sound. That is, when the user U's utterance is detected, the operation control unit 150 according to the present embodiment further increases the input level related to the user U's utterance around the user U that is the approach target, and the noise sound source NS is emitted.
- the autonomous mobile body 10 can be moved to a position where the input level of the non-target sound is further lowered.
- the motion control unit 150 moves the autonomous moving body 10 to the position where the non-target sound is generated, that is, the side opposite to the noise source NS, centering on the user U that is the approach target. I am letting.
- the autonomous mobile body 10 can be moved so as to be further away from the noise sound source NS and closer to the user U that is the approach target. Further, when the autonomous mobile body 10 is moved to the opposite side of the noise source NS across the user U, the input level of the non-target sound emitted from the noise source NS is further effectively reduced. Expected to be effective.
- the input level of the target sound can be increased, the input level of the non-target sound can be decreased, or the rate of increase of the input level can be reduced.
- the operation control unit 150 according to the present embodiment only the moving function that the autonomous mobile body 10 originally has is used without performing signal processing on the input signal and beam forming using the directional microphone.
- the SN ratio can be greatly improved.
- the motion control unit 150 causes the autonomous mobile body 10 to perform an operation considering the presence of the non-target sound in addition to the target sound, thereby improving the S / N ratio and the sound related to the target sound. Recognition accuracy can be improved effectively.
- the target sound according to the present embodiment can be defined as a target voice for voice recognition by the recognition unit 120.
- the target voice may be general voice.
- the recognizing unit 120 can detect the overall sound as described above as the target sound, for example, by comparing the pitches while paying attention to the overtone structure of a human voice.
- the autonomous mobile body 10 can perform some action according to the sounds output from the television device, for example.
- the target sound according to the present embodiment may be, for example, only a predetermined user's voice registered in advance among the above voices.
- the recognizing unit 120 performs speaker identification based on a user's voice characteristics registered in advance or a person's face existing in the direction of arrival of the input signal, so that only the utterance voice of a predetermined user is the target sound. Can be detected.
- the target sound according to the present embodiment may be only a specific word related to a specific keyword or an action instruction among uttered voices uttered by a predetermined user.
- the recognition unit 120 can detect only a predetermined user's utterance voice including a specific keyword or word as a target sound by performing voice recognition based on the input signal.
- non-target sound according to the present embodiment can be defined as all sounds other than the target sound.
- examples of the non-target sound according to the present embodiment include work sounds in a kitchen and various non-sounds emitted by devices such as a ventilation fan, a refrigerator, and a car.
- FIG. 9 is a diagram for explaining the operation control of the autonomous mobile body 10 when the target sound according to the present embodiment is not detected.
- FIG. 9 shows an example of a noise map held by the surrounding environment holding unit 140.
- the noise map according to the present embodiment is a map that indicates a generation state of non-target sounds that is generated and updated by the surrounding environment estimation unit 130.
- the noise map according to the present embodiment includes, for example, a noise source that exists in a space where the autonomous mobile body 10 is present, and a noise that is a region where the input level of a non-target sound generated by the noise source is strong (for example, a threshold value or more). Information related to the region is included.
- the noise map includes information on the noise sources NS1 and NS2 and noise regions NR1 and NR2 corresponding to the noise sources NS1 and NS2, respectively.
- One feature of the operation control unit 150 according to the present embodiment is that the operation of the autonomous mobile body 10 is controlled based on a noise map including the above information. For example, when the target sound is not detected, the operation control unit 150 according to the present embodiment may control the operation of the autonomous mobile body 10 so as to avoid the input of the non-target sound based on the noise map.
- the motion control unit 150 moves the movement range of the autonomous mobile body 10 in an area where the input level of the non-target sound is equal to or less than the threshold based on the noise map. Can be limited.
- the operation control unit 150 prevents the autonomous mobile body 10 from entering the noise regions NR1 and NR2 so that the autonomous mobile body 10 does not enter the noise regions NR1 and NR2. You may limit 10 movement ranges.
- the operation control unit 150 randomly moves the autonomous mobile body 10 in the area or places the autonomous mobile body 10 at a position P min where the sound pressure of the non-target sound is expected to be minimum. You may move it.
- the autonomous mobile body 10 is operated so that the input of the non-target sound is suppressed as much as possible.
- the operation control unit 150 it is possible to effectively improve the accuracy of speech recognition related to the target sound.
- the operation control based on the noise map when the target sound according to the present embodiment is detected will be described in detail.
- the motion control unit 150 according to the present embodiment is located at a position where the input level of the target sound is further increased and the input level of the non-target sound is further decreased in the vicinity of the approach target.
- One feature is that the autonomous mobile body 10 is moved.
- the operation control unit 150 according to the present embodiment can realize the above-described operation control with high accuracy by referring to the noise map.
- FIG. 10 is a diagram for explaining the operation control based on the noise map when the target sound according to the present embodiment is detected.
- the operation control unit 150 approaches the autonomous mobile body 10 to the user U based on the detection of the speech UO1 of the user U who calls the name of the autonomous mobile body 10, that is, the target sound. I am letting.
- the operation control unit 150 according to the present embodiment refers to the noise map and controls the movement of the autonomous mobile body 10 in consideration of the noise source NS and the noise region NR included in the noise map.
- the operation control unit 150 refers to the noise map so that the autonomous mobile body 10 does not enter or stop within the noise region NR and is further away from the noise source NS. It is possible to move the autonomous mobile body 10 to the position. More specifically, the operation control unit 150 bypasses the noise region NR as indicated by a solid line in the figure, and moves the autonomous mobile body 10 on the opposite side of the noise source NS around the user U that is the approach target. You may let me.
- the operation control unit 150 accurately grasps the noise source and the noise region by referring to the noise map, the input level of the target sound increases, and the input level of the non-target sound. It is possible to move the autonomous mobile body 10 to a position where the lowering is. According to the above-described operation control by the operation control unit 150 according to the present embodiment, it is possible to improve the SN ratio and effectively improve the voice recognition accuracy related to the target sound.
- the motion control unit 150 does not necessarily have to move the autonomous mobile body 10 to the side opposite to the noise source around the approach target.
- a wall exists on a straight line connecting the noise source NS and the user U.
- the operation control unit 150 may stop the autonomous mobile body 10 at a position close to the opposite side and not entering the noise region NR. . Even in this case, it is possible to achieve both an increase in the input level of the target sound and a decrease in the input level of the non-target sound, and an effect of improving the SN ratio.
- the operation control unit 150 may grasp the presence of an obstacle based on information such as walls and furniture included in the noise map, as described above based on the presence of the obstacle recognized by the recognition unit 120. Operation control may be performed.
- FIG. 12 is a diagram for describing operation control when the approach target according to the present embodiment is not a utterance user.
- the operation control unit 150 sets the approach target to the user U2 based on the utterance voice UO2 recognized by the recognition unit 120, and controls the operation of the autonomous mobile body 10 so as to approach the user U2. To do.
- the approach target according to the present embodiment is not only a utterance user who utters a voice utterance, but a moving object such as another user identified by a voice recognition process based on the utterance voice, a fixed object such as a charging station, Or any position may be sufficient.
- the motion control unit 150 refers to the noise map in the same manner even when the approach target is not the utterance user, and performs the movement considering the noise source NS and the noise region NR to the autonomous mobile body 10. Can be made.
- the operation control unit 150 moves the autonomous mobile body 10 to the opposite side of the noise source NS with the user U2 as the center. According to the above-described operation control by the operation control unit 150 according to the present embodiment, it is possible to effectively improve the accuracy of speech recognition related to the uttered voice of the user U2, which is expected to be issued later, that is, the target sound. It becomes possible.
- the operation control unit 150 refers to the noise map held by the surrounding environment holding unit 140, so that not only the target sound input level but also the non-target sound input level is considered.
- the operation control unit 150 has been described as a main example in which the autonomous mobile body 10 is controlled to move to the approach target based on the detection of the target sound.
- the movement trigger in the present embodiment is not limited to such an example.
- the operation control unit 150 according to the present embodiment moves the autonomous mobile body 10 to an approach target based on, for example, recognition of the user's face or recognition of a gesture related to a movement instruction by the user. Control may be performed. Even in this case, by referring to the noise map and moving the autonomous mobile body to a position where the input level of the non-target sound is further lowered, the accuracy of speech recognition related to the target sound that is expected to be emitted later is improved. It is possible to increase.
- the surrounding environment estimation unit 130 can generate the noise map as described above based on the results of sound source direction estimation and sound pressure measurement, for example.
- FIG. 13 is a diagram for describing generation of a noise map based on sound source direction estimation according to the present embodiment.
- the surrounding environment estimation unit 130 In generating a noise map based on sound source direction estimation, the surrounding environment estimation unit 130 first performs sound source localization at an arbitrary point to estimate the direction of the sound source. In the example shown in FIG. 13, the surrounding environment estimation unit 130 estimates the directions of the noise sources NS1 and NS2 at the point P1. At this time, the directions of the noise sources NS1 and NS2 can be estimated, but the distance between the autonomous mobile body 10 and the noise sources NS1 and NS2 is unknown.
- the surrounding environment estimation unit 130 moves to a point different from the previously estimated sound source direction, performs sound source localization again, and estimates the direction of the sound source.
- the surrounding environment estimation unit 130 estimates the directions of the noise sources NS1 and NS2 again at the point P2.
- the surrounding environment estimation unit 130 can estimate the positions of the noise sources NS1 and NS2 in space from the moving distance of the autonomous mobile body 10 and the intersection of the directions estimated at the points P1 and P2.
- the surrounding environment estimation unit 130 can improve the estimation accuracy related to the sound source position by repeating the sound source direction estimation at another point.
- the surrounding environment estimation unit 130 estimates the directions of the noise sources NS1 and NS2 again at the point P3.
- the surrounding environment estimation unit 130 estimates the positions of the noise sources NS1 and NS2 in space with high accuracy by repeatedly estimating the sound source direction at a plurality of points. It is possible to generate a noise map in which regions at a predetermined distance are set as noise regions NR1 and NR2.
- FIG. 14 is a diagram for describing generation of a noise map based on sound source direction estimation according to the present embodiment.
- the surrounding environment estimation unit 130 In generating a noise map based on sound source direction estimation, the surrounding environment estimation unit 130 first measures the sound pressure level at an arbitrary point. In the example shown in FIG. 14, the surrounding environment estimation unit 130 measures the sound pressure level at the point P4. Subsequently, the surrounding environment estimation unit 130 repeatedly performs measurement of the sound pressure level at another point different from the measured point. In the example illustrated in FIG. 14, the surrounding environment estimation unit 130 performs sound pressure level measurement at the points P5 and P6.
- the surrounding environment estimation unit 130 repeats the sound pressure measurement at a plurality of points, thereby estimating the isobars related to the sound pressure level as shown in FIG.
- a certain area can be set as a noise area.
- the surrounding environment estimation unit 130 can also estimate a point having the highest sound pressure level in the noise region as the position of the noise source.
- the surrounding environment estimation unit 130 sets the noise regions NR1 and NR2 based on the estimated isobaric lines, and estimates the positions of the noise sources NS1 and NS2.
- the surrounding environment estimation unit 130 even if the autonomous mobile body 10 includes only a single microphone, it is possible to perform accuracy by repeatedly performing sound pressure measurement at a plurality of points. It is possible to generate a high noise map. Note that when noise map generation based on sound pressure measurement is performed, it is necessary to separate the target sound and the non-target sound, but the separation can be realized by the function of the recognition unit 120 described above.
- the noise map according to the present embodiment may include information such as the type of noise source.
- FIG. 15 is a diagram illustrating an example of a noise map including noise source type information according to the present embodiment.
- the noise map information includes that the noise sources NS ⁇ b> 1 and NS ⁇ b> 2 are a kitchen and a television device, respectively.
- the surrounding environment estimation unit 130 can generate a noise map including noise source type information as shown in FIG. 15 based on the result of object recognition by the recognition unit 120, for example. .
- the operation control unit 150 can perform more accurate operation control according to the specification of the noise source.
- the update of the noise map may be performed dynamically at all times.
- non-target sound generated in the surroundings can be detected without omission, and all information such as sudden sound that is not useful for the operation control by the operation control unit 150 is also included as noise map information.
- the calculation amount becomes enormous, and a high-performance processor or the like is required.
- the surrounding environment estimation unit 130 may execute the noise map generation process and update process only under conditions that are efficient with respect to the collection of non-target sounds.
- the above-mentioned conditions with good efficiency are situations where a large number of non-target sounds can occur.
- a situation in which a large number of non-target sounds can occur is a situation in which the user is active in space.
- the surrounding environment estimation part 130 which concerns on this embodiment may perform the production
- the surrounding environment estimation unit 130 estimates the absence or presence of the user based on the user's schedule and various types of sensor information, and the noise map is only used under conditions where the user is highly likely to exist. Generation processing and update processing can be executed.
- FIG. 16 is a setting example of execution conditions related to the noise map generation processing and update processing according to the present embodiment.
- factors such as the user's schedule (absence or at home), detection of key sound, detection of door opening / closing sound, detection of utterance such as “Imaima”, detection of moving object by human sensor, etc. Whether or not to perform generation processing and update processing is set for each combination.
- FIG. 16 shows an example in which the process is executed only when it is determined that the user exists in the same space as the autonomous mobile body 10 in all the above elements.
- the execution feasibility as described above may be dynamically set according to the characteristics or situation of the autonomous mobile body 10.
- the noise map is generated and updated based on the non-target sound collected in the time zone in which the user exists in the surrounding environment. It is possible to maintain a high noise map.
- the surrounding environment estimation unit 130 can dynamically update the noise map based on the non-target sound collected in the time zone where the user exists in the surrounding environment.
- the generated non-target sound is different at the timing when the sound is collected. For this reason, when the noise map is simply overwritten based on the latest collected sound data, information on the non-target sound, which has a natural influence such as sudden sound, is included in the noise map, and the accuracy of the operation control by the operation control unit 150 is increased. Can also be a factor in the decline.
- the surrounding environment estimation unit 130 does not overwrite the existing noise map based on the latest sound collection data, but integrates the latest sound collection data into the existing noise map,
- the noise map may be updated.
- FIG. 17 is a diagram for explaining the noise map integration processing according to the present embodiment.
- the surrounding environment estimation unit 130 performs noise map integration based on three times of sound collection.
- the non-target sound related to the kitchen and the television device is detected, and in the second sound collection, the non-purpose sound related to the window and the television device is detected, and the third sound collection.
- the third sound collection Suppose that a non-target sound relating to the kitchen and the television device is detected.
- the surrounding environment estimation unit 130 may update the noise map by integrating the collected sound data for three times, for example, by averaging.
- the occurrence frequency of each non-target sound can be reflected in the noise map, and the influence of the non-target sound with a low occurrence frequency such as sudden sound is generated. Can be suppressed.
- the occurrence frequency of the non-target sound is indicated by the density of hatching density (the higher the density, the higher the occurrence frequency).
- the operation control unit 150 may control the movement of the autonomous mobile body 10 so as to avoid the television apparatus that frequently generates non-target sounds more seriously.
- the generation and update of the noise map according to the present embodiment has been described.
- the method described above is merely an example, and the generation and update of the noise map according to the present embodiment is not limited to such an example.
- the surrounding environment estimation unit 130 may generate or update a noise map based on information input by the user, for example.
- 18 and 19 are diagrams for explaining generation and update of a noise map based on a user input according to the present embodiment.
- the surrounding environment estimation unit 130 generates and updates a noise map based on furniture arrangement information input by the user via the information processing terminal 20 or the like. May be.
- the surrounding environment estimation unit 130 requests the display unit included in the information processing terminal 20 to arrange the icon IC corresponding to each piece of furniture in the input area IA that is assumed to be a user's room.
- the noise map can be generated or updated based on the above.
- the surrounding environment estimation unit 130 determines the noise source NS based on a gesture such as pointing performed by the user U and an utterance voice UO3 related to the noise source teaching. It can be identified and reflected in the noise map.
- FIG. 20 is a diagram illustrating an example of a situation where it is difficult to avoid a noise region according to the present embodiment.
- a noise map based on sound source direction estimation is illustrated on the left side of FIG. 20, and a noise map based on sound pressure measurement is illustrated on the right side of FIG.
- the autonomous mobile body 10 is surrounded by the noise regions NR1 to NR4 and it is difficult to move to other places.
- the operation control unit 150 based on the avoidance priority associated with the noise sources NS1 to NS4, moves to an autonomous mobile body so as to move to a noise region corresponding to a noise source with a lower avoidance priority. 10 may be controlled.
- the avoidance priority according to the present embodiment can be determined, for example, by the type and characteristics of the non-target sound emitted by the noise source.
- the non-target sound according to the present embodiment includes various sounds other than the target sound.
- the influence of the non-target sound on the accuracy of speech recognition related to the target sound varies depending on the characteristics of the non-target sound.
- the surrounding environment estimation unit 130 classifies the non-target sounds based on the magnitude of the degree of influence on the speech recognition accuracy, and sets a higher avoidance priority in descending order of the degree of influence. May be generated.
- category 1 may be a non-target sound that is relatively loud and is a human voice but not a target sound.
- Category 1 includes, for example, audio output by a television device, radio, and other devices, music including vocals, conversations between third parties who are not users, and the like.
- Category 1 may be a non-target sound having the highest influence on voice recognition accuracy and the highest avoidance priority among the four categories.
- category 2 may be a non-target sound that is difficult to obtain the effect of noise suppression because the occurrence is non-stationary and the volume is relatively large.
- Category 2 includes, for example, work sounds such as dishwashing and cooking, outdoor sounds that flow in when windows are opened, and the like.
- Category 2 may be a non-target sound having the second highest impact on the speech recognition accuracy and the second highest avoidance priority among the four categories.
- category 3 may be a non-target sound that is regularly generated and relatively easy to obtain a noise suppression effect.
- Category 3 includes, for example, sounds generated by air conditioners, ventilation fans, PC fans, and the like.
- Category 2 may be a non-target sound that has the third highest impact on voice recognition accuracy and the third highest avoidance priority among the four categories.
- category 4 since category 4 is suddenly emitted, it may be a non-target sound that only affects instantaneously.
- Category 4 includes, for example, door opening sounds, footsteps, sounds emitted by microwave ovens, and the like.
- Category 4 may be a non-target sound having the lowest influence on the speech recognition accuracy and the lowest avoidance priority among the four categories.
- the surrounding environment estimation unit 130 can generate a noise map in which avoidance priority is set according to the characteristics of the non-target sound.
- the noise source avoidance priority according to the present embodiment may be set based on some acoustic quantitative indicator related to the non-target sound.
- the quantitative index include an index that indicates the degree of soundness and an index that indicates the degree of stationarity.
- the target sound that is the target of speech recognition that is, the user's uttered speech is “unsteady” “speech”.
- “non-stationary” “sound” includes, for example, non-target sounds such as a conversation between third parties and a sound output from a television device or a radio. For this reason, in order to improve the accuracy of speech recognition related to the target sound that is “non-stationary” “speech”, it is “non-stationary” “speech” that is difficult to separate from the target sound. It is important to avoid non-target sounds.
- non-stationary “non-speech” includes work sounds related to dishwashing and cooking
- “steady” “non-speech” includes sounds emitted from air conditioners, ventilation fans, PC fans, and the like. Although such a non-target sound is relatively easy to separate from the target sound, it can be said that it has a lower avoidance priority than the above-mentioned non-target sound that is “non-stationary” “speech”. .
- “steady” “speech” corresponds to the case where the same sound is spoken for a long time, such as “Ahhh”, but such a sound is very unlikely to occur in daily life, so it is ignored. May be.
- the surrounding environment estimation unit 130 uses the non-target sound for the speech recognition related to the target sound based on the index ⁇ indicating the degree of speech likelihood and the index ⁇ indicating the degree of continuity.
- the degree of influence may be calculated, and the avoidance priority may be set based on the calculated value.
- FIG. 21 and FIG. 22 are diagrams for explaining calculation of the index ⁇ indicating the degree of speech likeness and the index ⁇ indicating the degree of continuity according to the present embodiment.
- FIG. 21 shows a calculation flow when the autonomous mobile body 10 according to the present embodiment includes a plurality of microphones.
- the surrounding environment estimation unit 130 performs beamforming in the direction of the noise sources NS1 to NS4 in order to calculate the index ⁇ and the index ⁇ as shown in FIG. According to such a method, it is possible to calculate the index ⁇ and the index ⁇ related to the noise sources NR1 to NR4 without entering the noise regions NR1 to NR4. Even if it exists, the precision of the speech recognition which concerns on the said target sound can be maintained.
- FIG. 22 shows a calculation flow when the autonomous mobile body 10 according to the present embodiment includes a single microphone.
- the surrounding environment estimation unit 130 may calculate the index ⁇ and the index ⁇ around the noise sources NS1 to NS4 as shown in FIG.
- the operation control unit 150 may control the autonomous mobile body 10 to quickly escape from the noise regions NR1 to NR4 when the calculation of the index ⁇ and the index ⁇ related to each noise source is completed.
- the surrounding environment estimation unit 130 calculates the degree of influence of each noise source based on the index ⁇ and the index ⁇ calculated as described above, and sets the avoidance priority based on the degree of influence. Is possible. For example, the surrounding environment estimation unit 130 may set the total value of the index ⁇ and the index ⁇ as the high degree of influence, and set the avoidance priority higher in descending order of the total value.
- the surrounding environment estimation unit 130 may calculate the index ⁇ indicating the degree of sound quality based on, for example, the spectroentropy of sound.
- the spectroentropy of sound is an index that is also used in VAD (Voice Activity Detection) technology, and human speech tends to show a lower value than other sounds.
- the surrounding environment estimation unit 130 can calculate the spectroentropy of the sound, that is, the index ⁇ , for example, by the following mathematical formula (1).
- f indicates a certain frequency
- S f indicates an amplitude spectrum of the frequency f of the observation signal.
- P f in the formula (1) is defined by the following formula (2).
- the surrounding environment estimation unit 130 may calculate the index ⁇ indicating the degree of continuity based on, for example, the kurtosis of the sound.
- the kurtosis of a sound is an index that is often used to discriminate between stationary and non-stationary sounds, and can be calculated by the following mathematical formula (3).
- T in the following mathematical formula (3) indicates a sound section length for performing kurtosis calculation, and a length of 3 to 5 seconds, for example, may be set.
- t in Formula (3) indicates a certain time
- x (t) indicates a speech waveform at the time t.
- the autonomous mobile body 10 can preferentially avoid a non-target sound that affects voice recognition related to the target sound.
- FIG. 23 is a flowchart showing a flow of noise map update according to the present embodiment.
- the surrounding environment estimation unit 130 estimates the surrounding environment based on the sensor information collected by the input unit 110 and the recognition result by the recognition unit 120 (S1101). Specifically, the surrounding environment estimation unit 130 estimates a noise source and a noise region.
- the surrounding environment estimation unit 130 determines whether or not an existing noise map is held in the surrounding environment holding unit 140 (S1102).
- the surrounding environment estimation unit 130 when there is no noise map held in the surrounding environment holding unit 140 (S1102: NO), the surrounding environment estimation unit 130 generates a noise map based on the surrounding environment estimated in step S1101, and the surrounding environment is calculated.
- the data is stored in the holding unit 140 (S1107).
- the surrounding environment estimation unit 130 subsequently changes the number of noise sources between the estimated surrounding environment and the existing noise map. It is determined whether or not (S1103).
- the surrounding environment estimation unit 130 integrates the noise maps based on the surrounding environment estimated in step S1101 (S1106), and the integrated noise map. Is stored in the surrounding environment holding unit 140 (S1107).
- the surrounding environment estimation unit 130 subsequently determines whether the position of the noise source has changed between the estimated surrounding environment and the existing noise map. Is determined (S1104).
- the surrounding environment estimation unit 130 integrates the noise map based on the surrounding environment estimated in step S1101 (S1106), and the integrated noise map Is stored in the surrounding environment holding unit 140 (S1107).
- the surrounding environment estimation unit 130 subsequently determines the sound pressure of the non-target sound generated by the noise source between the estimated surrounding environment and the existing noise map. It is determined whether or not there is a change (S1105).
- the surrounding environment estimation unit 130 integrates the noise map based on the surrounding environment estimated in step S1101 (S1106).
- the integrated noise map is stored in the surrounding environment holding unit 140 (S1107).
- the surrounding environment estimation unit 130 does not update the noise map and maintains the existing noise map in the surrounding environment holding unit 140.
- FIG. 24 is a flowchart showing a flow of operation control according to the present embodiment.
- the operation control unit 150 first reads a noise map held by the surrounding environment holding unit 140 (S1201).
- the operation control unit 150 causes the autonomous mobile body 10 to perform an autonomous action avoiding the noise region based on the noise map read in step S1201 (S1202).
- the operation control unit 150 continuously determines whether or not the target sound is detected during the autonomous action in step S1202 (S1203).
- the operation control unit 150 is located at a position where the input level of the non-target sound is further reduced around the approach target based on the noise map read in step S1201.
- the autonomous mobile body 10 is moved (S1204).
- the motion control unit 150 causes the autonomous mobile body 10 to perform a corresponding motion based on the speech recognition result of the target sound (S1205).
- the autonomous mobile body 10 according to the present embodiment has been described above.
- the autonomous mobile body 10 according to the present embodiment has been described focusing on the movement considering the input levels of the target sound and the non-target sound in order to improve the SN ratio.
- the method for improving the S / N ratio according to the present embodiment is not limited to such an example, and for example, signal processing or beam forming technology may be used in combination.
- the operation control unit 150 may control the autonomous moving body 10 to move between the approaching object and the noise source and to perform beam forming in the direction of the approaching object.
- the motion control unit 150 may perform control so that beam forming is performed with an elevation angle corresponding to the height of the face of the user who is the approach target. In this case, an effect of effectively eliminating the non-target sound reaching from the horizontal direction and effectively improving the SN ratio is expected.
- the operation control unit 150 may cause the autonomous mobile body 10 to perform an operation of guiding the user, for example, in order to avoid a noise region.
- the motion control unit 150 causes the autonomous mobile body 10 to perform an operation of guiding the user to move away from the noise region and approach the autonomous mobile body 10, thereby It is possible to increase the input level of the target sound without entering.
- the above-described guidance can be realized by, for example, operations such as barking, stopping in front of the noise region, and wandering.
- the autonomous mobile body 10 has a communication function by language, such as a humanoid robot device, for example, it may be explicitly communicated by voice that it is desired to move away from the noise region.
- the autonomous mobile body 10 that is an example of an information processing apparatus according to an embodiment of the present disclosure includes the operation control unit 150 that controls the operation of the autonomous mobile body 10 based on the recognition process.
- the motion control unit 150 when the target sound that is the target voice of the voice recognition process is detected, is a non-target voice that is not the target voice in the vicinity of the approach target that is determined based on the target sound.
- the autonomous mobile body 10 is moved to a position where the sound input level is further lowered. According to this configuration, it is possible to cause the autonomous mobile body to perform an operation that further improves the accuracy of voice recognition.
- each step related to the processing of the autonomous mobile body 10 in this specification does not necessarily have to be processed in time series in the order described in the flowchart.
- each step related to the processing of the autonomous mobile body 10 may be processed in an order different from the order described in the flowchart, or may be processed in parallel.
- An operation control unit that controls the operation of an autonomous mobile body that performs actions based on recognition processing; With When the target sound that is the target voice of the speech recognition process is detected, the operation control unit is a position where the input level of the non-target sound that is not the target voice is further reduced around the approach target determined based on the target sound. To move the autonomous mobile body, Information processing device. (2) When the target sound is detected, the operation control unit has a position where the input level of the target sound is further increased and the input level of the non-target sound is further decreased in the vicinity of the approach target determined based on the target sound. To move the autonomous mobile body, The information processing apparatus according to (1).
- the operation control unit moves the autonomous mobile body away from a noise source that emits the non-target sound and moves closer to the approach target.
- the information processing apparatus according to (1) or (2).
- the operation control unit moves the autonomous mobile body to the side opposite to the noise source that emits the non-target sound around the approach target.
- the information processing apparatus according to any one of (1) to (3).
- the target sound is an utterance voice of the user
- the approach target is an utterance user who utters the utterance voice.
- the information processing apparatus according to any one of (1) to (4).
- the approach target is a moving object, a fixed object, or a position specified by the voice recognition process based on a user's voice.
- the information processing apparatus controls the operation of the autonomous mobile body based on a noise map indicating the occurrence state of the non-target sound in the surrounding environment.
- the information processing apparatus according to any one of (1) to (6).
- the noise map includes information of a noise source that emits the non-target sound,
- the operation control unit controls the operation of the autonomous mobile body based on the avoidance priority related to the noise source.
- the information processing apparatus according to (7).
- the avoidance priority related to the noise source is determined based on the type of the noise source.
- the information processing apparatus according to (8).
- the avoidance priority related to the noise source is determined based on the degree of influence of the non-target sound emitted by the noise source on the speech recognition process.
- the information processing apparatus according to (8). (11) The degree of influence is calculated based on at least one of an index indicating the degree of speech likeness of the non-target sound and an index indicating the degree of stationarity. The information processing apparatus according to (10). (12) The operation control unit controls the operation of the autonomous mobile body to avoid the input of the non-target sound based on the noise map when the target sound is not detected. The information processing apparatus according to any one of (7) to (10). (13) When the target sound is not detected, the operation control unit limits a movement range of the autonomous mobile body within an area where an input level of the non-target sound is equal to or less than a threshold based on the noise map. The information processing apparatus according to any one of (7) to (12).
- the surrounding environment estimation unit generates the noise map based on direction estimation or sound pressure measurement related to a noise source that emits the non-target sound.
- the surrounding environment estimation unit dynamically updates the noise map based on the collected non-target sound.
- the surrounding environment estimation unit dynamically updates the noise map based on a change in the number, position, or sound pressure of noise sources that emit the non-target sound.
- the surrounding environment estimation unit generates or updates the noise map based on the non-target sound collected in a time zone in which the user exists in the surrounding environment.
- the information processing apparatus according to (16) or (17).
- the processor controls the movement of the autonomous mobile body that performs actions based on the recognition process; Including The control means that, when a target sound that is a target voice of voice recognition processing is detected, a position where an input level of a non-target sound that is not the target voice is further reduced around an approach target determined based on the target sound.
- the control Further including Information processing method.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Physics & Mathematics (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Mechanical Engineering (AREA)
- Robotics (AREA)
- Manipulator (AREA)
- Toys (AREA)
Abstract
【課題】自律移動体に音声認識の精度をより向上させる動作を実行させる。 【解決手段】認識処理に基づいて行動を行う自律移動体の動作を制御する動作制御部、を備え、前記動作制御部は、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、情報処理装置が提供される。また、プロセッサが、認識処理に基づいて行動を行う自律移動体の動作を制御すること、を含み、前記制御することは、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させること、をさらに含む、情報処理方法が提供される。
Description
本開示は、情報処理装置、情報処理方法、およびプログラムに関する。
近年、ユーザの発話などの音源の方向を推定し、当該音源の方向に応じた動作を実行する種々の装置が開発されている。上記の装置には、例えば、推定した音源方向に基づいて、自律移動を実行する自律移動体が含まれる。例えば、特許文献1には、ユーザの発話や顔が認識された方向にロボット装置を移動させる技術が開示されている。
しかし、特許文献1に記載の技術では、ユーザの発話以外の音、すなわちノイズの存在が考慮されていない。このため、推定されたユーザの方向にロボット装置を単に接近させた場合、ノイズの入力レベルが上昇し、ユーザの発話の認識が困難となる可能性がある。
そこで、本開示では、自律移動体に音声認識の精度をより向上させる動作を実行させることが可能な、新規かつ改良された情報処理装置、情報処理方法、およびプログラムを提案する。
本開示によれば、認識処理に基づいて行動を行う自律移動体の動作を制御する動作制御部、を備え、前記動作制御部は、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、情報処理装置が提供される。
また、本開示によれば、プロセッサが、認識処理に基づいて行動を行う自律移動体の動作を制御すること、を含み、前記制御することは、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させること、をさらに含む、情報処理方法が提供される。
また、本開示によれば、コンピュータを、認識処理に基づいて行動を行う自律移動体の動作を制御する動作制御部、を備え、前記動作制御部は、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、情報処理装置、として機能させるためのプログラムが提供される。
以上説明したように本開示によれば、自律移動体に音声認識の精度をより向上させる動作を実行させることが可能となる。
なお、上記の効果は必ずしも限定的なものではなく、上記の効果とともに、または上記の効果に代えて、本明細書に示されたいずれかの効果、または本明細書から把握され得る他の効果が奏されてもよい。
以下に添付図面を参照しながら、本開示の好適な実施の形態について詳細に説明する。なお、本明細書及び図面において、実質的に同一の機能構成を有する構成要素については、同一の符号を付することにより重複説明を省略する。
なお、説明は以下の順序で行うものとする。
1.構成
1.1.自律移動体10の概要
1.2.自律移動体10のハードウェア構成例
1.3.自律移動体10の機能構成例
2.実施形態
2.1.概要
2.2.動作制御の詳細
2.3.ノイズマップの生成と更新
2.4.ノイズ源の回避優先度に基づく動作制御
2.5.動作の流れ
3.まとめ
1.構成
1.1.自律移動体10の概要
1.2.自律移動体10のハードウェア構成例
1.3.自律移動体10の機能構成例
2.実施形態
2.1.概要
2.2.動作制御の詳細
2.3.ノイズマップの生成と更新
2.4.ノイズ源の回避優先度に基づく動作制御
2.5.動作の流れ
3.まとめ
<1.構成>
<<1.1.自律移動体10の概要>>
上述したように、近年においては、ユーザの発話などを認識し、認識結果に基づく動作を実行する種々の装置が開発されている。上記のような装置には、例えば、ユーザの発話や、周囲の環境などに応じて振る舞いを変化させる自律移動体が挙げられる。
<<1.1.自律移動体10の概要>>
上述したように、近年においては、ユーザの発話などを認識し、認識結果に基づく動作を実行する種々の装置が開発されている。上記のような装置には、例えば、ユーザの発話や、周囲の環境などに応じて振る舞いを変化させる自律移動体が挙げられる。
ここで、一般に精度の高い音声認識を実現するためには、マイクロフォンを通じて得られる音の信号において、音声認識の対象音声である目的音(例えば、ユーザの発話音声)と対象音声ではない非目的音のパワーの比、すなわちSN比(Signal-to-Noise Ratio)を高めることが重要となる。特に、音声認識機能を有する自律移動体にあっては、SN比が向上する位置に移動を行うことで音声認識精度を高めることが望ましい。
しかし、特許文献1に記載の技術では、非目的音を考慮しておらず、ユーザの発話や顔が認識された方向にロボット装置を移動させるに留まっている。このため、特許文献1に記載の技術では、ユーザに接近すると共に非目的音を発するノイズ源にも同時に接近した結果、SN比が低下し、音声認識精度が低下する状況も想定される。
また、特許文献1に記載の技術では、ユーザの発話や顔が認識されたことをトリガーとし、ユーザに接近するようロボット装置の動作を制御している。このため、特許文献1に記載のロボット装置は、周辺に存在するユーザに常に追従する可能性が高く、ユーザに煩わしさを感じさせてしまうことも想定される。
本開示の一実施形態に係る情報処理装置、情報処理方法、およびプログラムは、上記の点に着目して発想されたものであり、自律移動体に音声認識の精度をより向上させる動作を実行させることを可能とする。
ここで、まず、本開示の一実施形態に係る自律移動体10の概要について説明する。本開示の一実施形態に係る自律移動体10は、収集したセンサ情報に基づく状況推定を実行し、状況に応じた種々の動作を自律的に選択し実行する情報処理装置である。自律移動体10は、単にユーザの指示コマンドに従った動作を行うロボットとは異なり、状況ごとに最適であると推測した動作を自律的に実行することを特徴の一つとする。
本開示の一実施形態に係る自律移動体10は、例えば、音声認識処理の対象音声である目的音、すなわちユーザの発話が検出されない場合には、上記対象音声ではない非目的音の入力が回避されるよう自律動作を行ってよい。当該動作によれば、ユーザに常につきまとうことなく、またユーザの発話が検出された際には、当該発話に係る音声認識の精度が向上する可能性を効果的に高めることが可能である。
また、本開示の一実施形態に係る自律移動体10は、目的音が検出された場合には、当該目的音に基づいて定まる接近対象の周辺において非目的音の入力レベルがより低下する位置に移動してよい。すなわち、本開示の一実施形態に係る自律移動体10は、非目的音を考慮した移動動作を行うことで、SN比を向上させ、ユーザの発話に係る音声認識精度を効果的に向上させることが可能である。
このように、本開示の一実施形態に係る自律移動体10は、ヒトを含む動物と同様に、自己の状態や周囲の環境などを総合的に判断して自律動作を決定、実行する。上記の点において、本開示の一実施形態に係る自律移動体10は、指示に基づいて対応する動作や処理を実行する受動的な装置とは明確に相違する。
本開示の一実施形態に係る自律移動体10は、空間内において自律的な姿勢制御を行い、種々の動作を実行する自律移動型ロボットであってよい。自律移動体10は、例えば、ヒトやイヌなどの動物を模した形状や、動作能力を有する自律移動型ロボットであってもよい。また、自律移動体10は、例えば、ユーザとのコミュニケーション能力を有する車両や無人航空機などの装置であってもよい。本開示の一実施形態に係る自律移動体10の形状、能力、また欲求などのレベルは、目的や役割に応じて適宜設計され得る。
<<1.2.自律移動体10のハードウェア構成例>>
次に、本開示の一実施形態に係る自律移動体10のハードウェア構成例について説明する。なお、以下では、自律移動体10がイヌ型の四足歩行ロボットである場合を例に説明する。
次に、本開示の一実施形態に係る自律移動体10のハードウェア構成例について説明する。なお、以下では、自律移動体10がイヌ型の四足歩行ロボットである場合を例に説明する。
図1は、本開示の一実施形態に係る自律移動体10のハードウェア構成例を示す図である。図1に示すように、自律移動体10は、頭部、胴部、4つの脚部、および尾部を有するイヌ型の四足歩行ロボットである。また、自律移動体10は、頭部に2つのディスプレイ510を備える。
また、自律移動体10は、種々のセンサを備える。自律移動体10は、例えば、マイクロフォン515、カメラ520、ToF(Time of Flight)センサ525、人感センサ530、測距センサ535、タッチセンサ540、照度センサ545、足裏ボタン550、慣性センサ555を備える。
(マイクロフォン515)
マイクロフォン515は、周囲の音を収集する機能を有する。上記の音には、例えば、ユーザの発話や、周囲の環境音が含まれる。自律移動体10は、例えば、頭部に4つのマイクロフォンを備えてもよい。複数のマイクロフォン515を備えることで、周囲で発生する音を感度高く収集すると共に、音源の定位を実現することが可能となる。
マイクロフォン515は、周囲の音を収集する機能を有する。上記の音には、例えば、ユーザの発話や、周囲の環境音が含まれる。自律移動体10は、例えば、頭部に4つのマイクロフォンを備えてもよい。複数のマイクロフォン515を備えることで、周囲で発生する音を感度高く収集すると共に、音源の定位を実現することが可能となる。
(カメラ520)
カメラ520は、ユーザや周囲環境を撮像する機能を有する。自律移動体10は、例えば、鼻先と腰部に2つの広角カメラを備えてもよい。この場合、鼻先に配置される広角カメラは、自律移動体10の前方視野(すなわち、イヌの視野)に対応した画像を撮像し、腰部の広角カメラは、上方を中心とする周囲領域の画像を撮像する。自律移動体10は、例えば、腰部に配置される広角カメラにより撮像された画像に基づいて、天井の特徴点などを抽出し、SLAM(Simultaneous Localization and Mapping)を実現することができる。
カメラ520は、ユーザや周囲環境を撮像する機能を有する。自律移動体10は、例えば、鼻先と腰部に2つの広角カメラを備えてもよい。この場合、鼻先に配置される広角カメラは、自律移動体10の前方視野(すなわち、イヌの視野)に対応した画像を撮像し、腰部の広角カメラは、上方を中心とする周囲領域の画像を撮像する。自律移動体10は、例えば、腰部に配置される広角カメラにより撮像された画像に基づいて、天井の特徴点などを抽出し、SLAM(Simultaneous Localization and Mapping)を実現することができる。
(ToFセンサ525)
ToFセンサ525は、頭部前方に存在する物体との距離を検出する機能を有する。ToFセンサ525は、頭部の鼻先に備えられる。ToFセンサ525によれば、種々の物体との距離を精度高く検出することができ、ユーザを含む対象物や障害物などとの相対位置に応じた動作を実現することが可能となる。
ToFセンサ525は、頭部前方に存在する物体との距離を検出する機能を有する。ToFセンサ525は、頭部の鼻先に備えられる。ToFセンサ525によれば、種々の物体との距離を精度高く検出することができ、ユーザを含む対象物や障害物などとの相対位置に応じた動作を実現することが可能となる。
(人感センサ530)
人感センサ530は、ユーザやユーザが飼育するペットなどの所在を検知する機能を有する。人感センサ530は、例えば、胸部に配置される。人感センサ530によれば、前方に存在する動物体を検知することで、当該動物体に対する種々の動作、例えば、興味、恐怖、驚きなどの感情に応じた動作を実現することが可能となる。
人感センサ530は、ユーザやユーザが飼育するペットなどの所在を検知する機能を有する。人感センサ530は、例えば、胸部に配置される。人感センサ530によれば、前方に存在する動物体を検知することで、当該動物体に対する種々の動作、例えば、興味、恐怖、驚きなどの感情に応じた動作を実現することが可能となる。
(測距センサ535)
測距センサ535は、自律移動体10の前方床面の状況を取得する機能を有する。測距センサ535は、例えば、胸部に配置される。測距センサ535によれば、自律移動体10の前方床面に存在する物体との距離を精度高く検出することができ、当該物体との相対位置に応じた動作を実現することができる。
測距センサ535は、自律移動体10の前方床面の状況を取得する機能を有する。測距センサ535は、例えば、胸部に配置される。測距センサ535によれば、自律移動体10の前方床面に存在する物体との距離を精度高く検出することができ、当該物体との相対位置に応じた動作を実現することができる。
(タッチセンサ540)
タッチセンサ540は、ユーザによる接触を検知する機能を有する。タッチセンサ540は、例えば、頭頂、あご下、背中など、ユーザが自律移動体10に対し触れる可能性が高い部位に配置される。タッチセンサ540は、例えば、静電容量式や感圧式のタッチセンサであってよい。タッチセンサ540によれば、ユーザによる触れる、撫でる、叩く、押すなどの接触行為を検知することができ、当該接触行為に応じた動作を行うことが可能となる。
タッチセンサ540は、ユーザによる接触を検知する機能を有する。タッチセンサ540は、例えば、頭頂、あご下、背中など、ユーザが自律移動体10に対し触れる可能性が高い部位に配置される。タッチセンサ540は、例えば、静電容量式や感圧式のタッチセンサであってよい。タッチセンサ540によれば、ユーザによる触れる、撫でる、叩く、押すなどの接触行為を検知することができ、当該接触行為に応じた動作を行うことが可能となる。
(照度センサ545)
照度センサ545は、自律移動体10が位置する空間の照度を検出する。照度センサ545は、例えば、頭部背面において尾部の付け根などに配置されてもよい。照度センサ545によれば、周囲の明るさを検出し、当該明るさに応じた動作を実行することが可能となる。
照度センサ545は、自律移動体10が位置する空間の照度を検出する。照度センサ545は、例えば、頭部背面において尾部の付け根などに配置されてもよい。照度センサ545によれば、周囲の明るさを検出し、当該明るさに応じた動作を実行することが可能となる。
(足裏ボタン550)
足裏ボタン550は、自律移動体10の脚部底面が床と接触しているか否かを検知する機能を有する。このために、足裏ボタン550は、4つの脚部の肉球に該当する部位にそれぞれ配置される。足裏ボタン550によれば、自律移動体10と床面との接触または非接触を検知することができ、例えば、自律移動体10がユーザにより抱き上げられたことなどを把握することが可能となる。
足裏ボタン550は、自律移動体10の脚部底面が床と接触しているか否かを検知する機能を有する。このために、足裏ボタン550は、4つの脚部の肉球に該当する部位にそれぞれ配置される。足裏ボタン550によれば、自律移動体10と床面との接触または非接触を検知することができ、例えば、自律移動体10がユーザにより抱き上げられたことなどを把握することが可能となる。
(慣性センサ555)
慣性センサ555は、頭部や胴部の速度や加速度、回転などの物理量を検出する6軸センサである。すなわち、慣性センサ555は、X軸、Y軸、Z軸の加速度および角速度を検出する。慣性センサ555は、頭部および胴部にそれぞれ配置される。慣性センサ555によれば、自律移動体10の頭部および胴部の運動を精度高く検出し、状況に応じた動作制御を実現することが可能となる。
慣性センサ555は、頭部や胴部の速度や加速度、回転などの物理量を検出する6軸センサである。すなわち、慣性センサ555は、X軸、Y軸、Z軸の加速度および角速度を検出する。慣性センサ555は、頭部および胴部にそれぞれ配置される。慣性センサ555によれば、自律移動体10の頭部および胴部の運動を精度高く検出し、状況に応じた動作制御を実現することが可能となる。
以上、本開示の一実施形態に係る自律移動体10が備えるセンサの一例について説明した。なお、図1を用いて説明した上記の構成はあくまで一例であり、自律移動体10が備え得るセンサの構成は係る例に限定されない。自律移動体10は、上記の構成のほか、例えば、温度センサ、地磁気センサ、GNSS(Global Navigation Satellite System)信号受信機を含む各種の通信装置などをさらに備えてよい。自律移動体10が備えるセンサの構成は、仕様や運用に応じて柔軟に変形され得る。
続いて、本開示の一実施形態に係る自律移動体10の関節部の構成例について説明する。図2は、本開示の一実施形態に係る自律移動体10が備えるアクチュエータ570の構成例である。本開示の一実施形態に係る自律移動体10は、図2に示す回転箇所に加え、耳部と尾部に2つずつ、口に1つの合計22の回転自由度を有する。
例えば、自律移動体10は、頭部に3自由度を有することで、頷きや首を傾げる動作を両立することができる。また、自律移動体10は、腰部に備えるアクチュエータ570により、腰のスイング動作を再現することで、より現実のイヌに近い自然かつ柔軟な動作を実現することが可能である。
なお、本開示の一実施形態に係る自律移動体10は、例えば、1軸アクチュエータと2軸アクチュエータを組み合わせることで、上記の22の回転自由度を実現してもよい。例えば、脚部における肘や膝部分においては1軸アクチュエータを、肩や大腿の付け根には2軸アクチュエータをそれぞれ採用してもよい。
図3および図4は、本開示の一実施形態に係る自律移動体10が備えるアクチュエータ570の動作について説明するための図である。図3を参照すると、アクチュエータ570は、モータ575により出力ギアを回転させることで、可動アーム590を任意の回転位置および回転速度で駆動させることができる。
図4を参照すると、本開示の一実施形態に係るアクチュエータ570は、リアカバー571、ギアBOXカバー572、制御基板573、ギアBOXベース574、モータ575、第1ギア576、第2ギア577、出力ギア578、検出用マグネット579、2個のベアリング580を備える。
本開示の一実施形態に係るアクチュエータ570は、例えば、磁気式svGMR(spin-valve Giant Magnetoresistive)であってもよい。制御基板573が、メインプロセッサによる制御に基づいて、モータ575を回転させることで、第1ギア576および第2ギア577を介して出力ギア578に動力が伝達され、可動アーム590を駆動させることが可能である。
また、制御基板573に備えられる位置センサが、出力ギア578に同期して回転する検出用マグネット579の回転角を検出することで、可動アーム590の回転角度、すなわち回転位置を精度高く検出することができる。
なお、磁気式svGMRは、非接触方式であるため耐久性に優れるとともに、GMR飽和領域において使用することで、検出用マグネット579や位置センサの距離変動による信号変動の影響が少ないという利点を有する。
以上、本開示の一実施形態に係る自律移動体10が備えるアクチュエータ570の構成例について説明した。上記の構成によれば、自律移動体10が備える関節部の屈伸動作を精度高く制御し、また関節部の回転位置を正確に検出することが可能となる。
続いて、図5を参照して、本開示の一実施形態に係る自律移動体10が備えるディスプレイ510の機能について説明する。図5は、本開示の一実施形態に係る自律移動体10が備えるディスプレイ510の機能について説明するための図である。
(ディスプレイ510)
ディスプレイ510は、自律移動体10の目の動きや感情を視覚的に表現する機能を有する。図5に示すように、ディスプレイ510は、感情や動作に応じた眼球、瞳孔、瞼の動作を表現することができる。ディスプレイ510は、文字や記号、また眼球運動とは関連しない画像などを敢えて表示しないことで、実在するイヌなどの動物に近い自然な動作を演出する。
ディスプレイ510は、自律移動体10の目の動きや感情を視覚的に表現する機能を有する。図5に示すように、ディスプレイ510は、感情や動作に応じた眼球、瞳孔、瞼の動作を表現することができる。ディスプレイ510は、文字や記号、また眼球運動とは関連しない画像などを敢えて表示しないことで、実在するイヌなどの動物に近い自然な動作を演出する。
また、図5に示すように、自律移動体10は、右眼および左眼にそれぞれ相当する2つのディスプレイ510rおよび510lを備える。ディスプレイ510rおよび510lは、例えば、独立した2つのOLED(Organic Light Emitting Diode)により実現される。OLEDによれば、眼球の曲面を再現することが可能となり、1枚の平面ディスプレイにより一対の眼球を表現する場合や、2枚の独立した平面ディスプレイにより2つの眼球をそれぞれ表現する場合と比較して、より自然な外装を実現することができる。
以上述べたように、ディスプレイ510rおよび510lによれば、図5に示すような自律移動体10の視線や感情を高精度かつ柔軟に表現することが可能となる。また、ユーザはディスプレイ510に表示される眼球の動作から、自律移動体10の状態を直観的に把握することが可能となる。
以上、本開示の一実施形態に係る自律移動体10のハードウェア構成例について説明した。上記の構成によれば、図6に示すように、自律移動体10の関節部や眼球の動作を精度高くまた柔軟に制御することで、より実在の生物に近い動作および感情表現を実現することが可能となる。なお、図6は、本開示の一実施形態に係る自律移動体10の動作例を示す図であるが、図6では、自律移動体10の関節部および眼球の動作について着目して説明を行うため、自律移動体10の外部構造を簡略化して示している。同様に、以下の説明においては、自律移動体10の外部構造を簡略化して示す場合があるが、本開示の一実施形態に係る自律移動体10のハードウェア構成および外装は、図面により示される例に限定されず、適宜設計され得る。
<<1.3.自律移動体10の機能構成例>>
次に、本開示の一実施形態に係る自律移動体10の機能構成例について説明する。図7は、本開示の一実施形態に係る自律移動体10の機能構成例を示す図である。図7を参照すると、本開示の一実施形態に係る自律移動体10は、入力部110、認識部120、周辺環境推定部130、周辺環境保持部140、動作制御部150、駆動部160、出力部170を備える。
次に、本開示の一実施形態に係る自律移動体10の機能構成例について説明する。図7は、本開示の一実施形態に係る自律移動体10の機能構成例を示す図である。図7を参照すると、本開示の一実施形態に係る自律移動体10は、入力部110、認識部120、周辺環境推定部130、周辺環境保持部140、動作制御部150、駆動部160、出力部170を備える。
(入力部110)
入力部110は、ユーザや周囲環境に係る種々の情報を収集する機能を有する。入力部110は、例えば、ユーザの発話や周囲で発生する環境音、ユーザや周囲環境に係る画像情報、および種々のセンサ情報を収集する。このために、入力部110は、図1に示す各種のセンサを備える。
入力部110は、ユーザや周囲環境に係る種々の情報を収集する機能を有する。入力部110は、例えば、ユーザの発話や周囲で発生する環境音、ユーザや周囲環境に係る画像情報、および種々のセンサ情報を収集する。このために、入力部110は、図1に示す各種のセンサを備える。
(認識部120)
認識部120は、入力部110が収集した種々の情報に基づいて、ユーザや周囲の物体、また自律移動体10の状態に係る種々の認識を行う機能を有する。一例としては、認識部120は、人識別、顔認識、表情や視線の認識、音声認識、物体認識、色認識、形認識、マーカー認識、障害物認識、段差認識、明るさ認識などを行ってよい。
認識部120は、入力部110が収集した種々の情報に基づいて、ユーザや周囲の物体、また自律移動体10の状態に係る種々の認識を行う機能を有する。一例としては、認識部120は、人識別、顔認識、表情や視線の認識、音声認識、物体認識、色認識、形認識、マーカー認識、障害物認識、段差認識、明るさ認識などを行ってよい。
(周辺環境推定部130)
周辺環境推定部130は、入力部110が収集したセンサ情報や、認識部120による認識結果に基づいて、非目的音の発生状況を示すノイズマップを生成、更新する機能を有する。周辺環境推定部130が有する機能については別途詳細に説明する。
周辺環境推定部130は、入力部110が収集したセンサ情報や、認識部120による認識結果に基づいて、非目的音の発生状況を示すノイズマップを生成、更新する機能を有する。周辺環境推定部130が有する機能については別途詳細に説明する。
(周辺環境保持部140)
周辺環境保持部140は、周辺環境推定部130が生成、更新したノイズマップを保持する機能を有する。
周辺環境保持部140は、周辺環境推定部130が生成、更新したノイズマップを保持する機能を有する。
(動作制御部150)
動作制御部150は、認識部120による認識結果や周辺環境保持部140が保持するノイズマップに基づいて行動計画を行い、当該行動計画に基づいて、駆動部160および出力部170の動作を制御する。動作制御部150は、例えば、上記の行動計画に基づいて、アクチュエータ570の回転制御や、ディスプレイ510の表示制御、スピーカによる音声出力制御などを行う。本開示の一実施形態に係る動作制御部150が有する機能については別途詳細に説明する。
動作制御部150は、認識部120による認識結果や周辺環境保持部140が保持するノイズマップに基づいて行動計画を行い、当該行動計画に基づいて、駆動部160および出力部170の動作を制御する。動作制御部150は、例えば、上記の行動計画に基づいて、アクチュエータ570の回転制御や、ディスプレイ510の表示制御、スピーカによる音声出力制御などを行う。本開示の一実施形態に係る動作制御部150が有する機能については別途詳細に説明する。
(駆動部160)
駆動部160は、動作制御部150による制御に基づいて、自律移動体10が有する複数の関節部を屈伸させる機能を有する。より具体的には、駆動部160は、動作制御部150による制御に基づき、各関節部が備えるアクチュエータ570を駆動させる。
駆動部160は、動作制御部150による制御に基づいて、自律移動体10が有する複数の関節部を屈伸させる機能を有する。より具体的には、駆動部160は、動作制御部150による制御に基づき、各関節部が備えるアクチュエータ570を駆動させる。
(出力部170)
出力部170は、動作制御部150による制御に基づいて、視覚情報や音情報の出力を行う機能を有する。このために、出力部170は、ディスプレイ510やスピーカを備える。
出力部170は、動作制御部150による制御に基づいて、視覚情報や音情報の出力を行う機能を有する。このために、出力部170は、ディスプレイ510やスピーカを備える。
以上、本開示の一実施形態に係る自律移動体10の機能構成について説明した。なお、図7に示す構成はあくまで一例であり、本実施形態に係る自律移動体10の機能構成は係る例に限定されない。本開示の一実施形態に係る自律移動体10は、例えば、情報処理サーバや他の自律移動体と通信を行う通信部などを備えてよい。また、認識部120、周辺環境推定部130、動作制御部150などが有する機能は、上記の情報処理サーバの機能として実現されてもよい。この場合、情報処理サーバは、自律移動体10の入力部110が収集したセンサ情報に基づいて各種の認識処理、ノイズマップの生成または更新、行動計画を実行し、自律移動体10の駆動部160と出力部170の制御を行うことが可能である。本開示の一実施形態に係る自律移動体10の機能構成は仕様や運用に応じて柔軟に変形され得る。
<2.実施形態>
<<2.1.概要>>
次に、本開示の実施形態の概要について説明する。上述したように、本開示の一実施形態に係る自律移動体10は、目的音に係る音声認識の精度を向上させるために、目的音と非目的音とに係るSN比が向上するように自律動作を行う。
<<2.1.概要>>
次に、本開示の実施形態の概要について説明する。上述したように、本開示の一実施形態に係る自律移動体10は、目的音に係る音声認識の精度を向上させるために、目的音と非目的音とに係るSN比が向上するように自律動作を行う。
ここで、SN比を向上させるための手法としては、入力信号に信号処理を施す手法(マルチマイク信号処理、シングルマイク信号処理)や、指向性マイクロフォンなどを用いる手法も想定される。しかし、SN比に対しては、目的音源や非目的音源(以下、ノイズ源、とも称する)との物理的な距離の影響が最も強いといえる。
このために、本実施形態に係る自律移動体10は、単純に目的音源へ近づくのではなく、目的音源へ近づきながらも非目的音源から可能な限り遠ざかることで、SN比を効果的に向上させることができる。
図8Aおよび図8Bは、本実施形態に係る自律移動体10の動作概要について説明するための図である。図8Aには、本実施形態に係る比較手法による移動制御の一例が示されている。図8Aに示す一例の場合、比較装置90は、ユーザUの音声、すなわち目的音が検出された場合、当該目的音の入力レベル(音圧レベル)が上昇するように、ユーザUに最短距離を以って接近している。しかし、この際、比較装置90は、非目的音を発するノイズ源NSの存在を考慮していないため、ユーザUに接近すると同時にノイズ源NSにも接近している。このように、比較手法の場合、目的音の入力レベルと共に非目的音の入力レベルも上昇し、結果としてSN比の向上効果が薄れ、ユーザUの発話に係る音声認識精度が低下する可能性が生じる。
一方、図8Bは、本実施形態に係る情報処理方法による移動制御の一例が示されている。図8Bに示すように、本実施形態に係る自律移動体10は、ユーザUの音声、すなわち目的音が検出された場合、非目的音を発するノイズ源の存在を考慮して移動を行う。具体的には、本実施形態に係る動作制御部150は、目的音が検出された場合、目的音に基づいて定まる接近対象の周辺において非目的音の入力レベルがより低下する位置に、自律移動体10を移動させてよい。
ここで、上記の接近対象は、目的音、すなわち発話音声を発したユーザUであってよい。すなわち、本実施形態に係る動作制御部150は、ユーザUの発話が検出された場合、接近対象であるユーザUの周辺においてユーザUの発話に係る入力レベルがより上昇し、ノイズ音源NSが発する非目的音の入力レベルがより低下する位置に、自律移動体10を移動させることができる。
図8Bに示す一例の場合、本実施形態に係る動作制御部150は、接近対象であるユーザUを中心に非目的音の発生位置、すなわちノイズ源NSとは反対側に自律移動体10を移動させている。このように、本実施形態に係る動作制御部150によれば、ノイズ音源NSからより遠ざかり、かつ接近対象であるユーザUにより近づくように、自律移動体10を移動させることができる。また、自律移動体10をユーザUを挟んでノイズ源NSとは反対側に移動させる場合、ユーザUが壁の役割を果たしノイズ源NSが発する非目的音の入力レベルがさらに効果的に低下する効果が期待される。
本実施形態に係る動作制御部150が有する上記の機能によれば、目的音の入力レベルを上昇させると共に、非目的音の入力レベルを低下、または入力レベルの上昇率を軽減することができ、結果としてSN比を効果的に向上させることが可能となる。このように、本実施形態に係る動作制御部150によれば、入力信号に対する信号処理や指向性マイクロフォンを用いたビームフォーミングを行わずとも、自律移動体10が本来有する移動機能だけを利用してSN比を大幅に改善することができる。
<<2.2.動作制御の詳細>>
次に、本実施形態に係る動作制御部150による自律移動体10の動作制御についてより詳細に説明する。上述したように、本実施形態に係る動作制御部150は、目的音に加え非目的音の存在を考慮した動作を自律移動体10に実行させることで、SN比を改善し目的音に係る音声認識精度を効果的に向上させることができる。
次に、本実施形態に係る動作制御部150による自律移動体10の動作制御についてより詳細に説明する。上述したように、本実施形態に係る動作制御部150は、目的音に加え非目的音の存在を考慮した動作を自律移動体10に実行させることで、SN比を改善し目的音に係る音声認識精度を効果的に向上させることができる。
ここで、本実施形態に係る目的音とは、認識部120による音声認識の対象音声と定義できる。上記対象音声は、音声全般であってもよい。例えば、自律移動体10がテレビジョン装置やラジオなどから出力される音声や、ユーザまたは第三者の発話音声のすべてを音声認識の対象とする場合、上記のような音声は、すべて目的音といえる。この場合、認識部120は、例えば、人の声の倍音構造などに着目しピッチを比較するなどして、上記のような音声全般を目的音として検出することが可能である。なお、上記のような音声をすべて目的音とする場合、自律移動体10が例えば、テレビジョン装置から出力される音声に応じて何らかのアクションを行うことなどが可能となる。
一方、本実施形態に係る目的音は、上記のような音声のうち、例えば、事前に登録された所定のユーザの音声のみであってもよい。この場合、認識部120は、予め登録されたユーザの音声特徴に基づく話者識別や、入力信号の到来方向に存在する人物の顔識別を行うことで、所定のユーザの発話音声のみを目的音として検出可能である。
他方、本実施形態に係る目的音は、所定ユーザが発する発話音声のうち、特定のキーワードや、行動指示などに係る特定の文言のみであってもよい。この場合、認識部120は、入力信号に基づく音声認識を行うことで特定のキーワードや文言を含む所定のユーザの発話音声のみを目的音として検出することが可能である。
また、本実施形態に係る非目的音とは、目的音以外のすべての音として定義できる。本実施形態に係る非目的音には、例えば、キッチンにおける作業音や、換気扇や冷蔵庫、車などの装置が発する種々の非音声が挙げられる。
以上、本実施形態に係る目的音と非目的音について詳細に説明した。続いて、図9を参照して、目的音が検出されない場合における動作制御について説明する。図9は、本実施形態に係る目的音の未検出時における自律移動体10の動作制御について説明するための図である。図9には、周辺環境保持部140が保持するノイズマップの一例が示されている。
ここで、本実施形態に係るノイズマップとは、周辺環境推定部130により生成、更新される、非目的音の発生状況を示すマップである。本実施形態に係るノイズマップには、例えば、自律移動体10が存在する空間に存在するノイズ源や、当該ノイズ源が発する非目的音の入力レベルが強い(例えば、閾値以上)領域であるノイズ領域に係る情報が含まれる。図9に示す一例の場合、ノイズマップには、ノイズ源NS1およびNS2と、ノイズ源NS1およびNS2のそれぞれに対応するノイズ領域NR1およびNR2の情報が含まれている。
本実施形態に係る動作制御部150は、上記のような情報を含むノイズマップに基づいて、自律移動体10の動作を制御することを特徴の一つとする。例えば、本実施形態に係る動作制御部150は、目的音が検出されない場合、ノイズマップに基づいて非目的音の入力を回避するように自律移動体10の動作を制御してよい。
より具体的には、本実施形態に係る動作制御部150は、目的音が検出されない場合、ノイズマップに基づいて非目的音の入力レベルが閾値以下となるエリア内に自律移動体10の移動範囲を限定することができる。例えば、図9に示す一例の場合、動作制御部150は、目的音が検出されていない場合、ノイズ領域NR1およびNR2に自律移動体10が進入しないよう、両領域を除いたエリアに自律移動体10の移動範囲を限定してよい。この際、動作制御部150は、例えば、上記エリア内において自律移動体10をランダムに移動させたり、非目的音の音圧が最小であると予想される位置Pminなどに自律移動体10を移動させてよい。
本実施形態に係る動作制御部150による上記の制御によれば、目的音が検出されていない場合であっても、非目的音の入力がなるべく抑えられるように自律移動体10を動作させることで、ユーザからの呼びかけなどがあった場合、すなわち目的音が検出された際に、当該目的音に係る音声認識の精度を効果的に向上させることが可能である。
続いて、本実施形態に係る目的音が検出された場合におけるノイズマップに基づく動作制御について詳細に説明する。上述したように、本実施形態に係る動作制御部150は、目的音が検出された場合、接近対象の周辺において目的音の入力レベルがより上昇し非目的音の入力レベルがより低下する位置に、自律移動体10を移動させることを特徴の一つとする。この際、本実施形態に係る動作制御部150は、ノイズマップを参照することで上記の動作制御を精度高く実現することが可能である。
図10は、本実施形態に係る目的音が検出された場合におけるノイズマップに基づく動作制御について説明するための図である。図10に示す一例の場合、動作制御部150は、自律移動体10の名前を呼ぶユーザUの発話音声UO1、すなわち目的音が検出されたことに基づいて、自律移動体10をユーザUに接近させている。この際、本実施形態に係る動作制御部150は、ノイズマップを参照し、ノイズマップに含まれるノイズ源NSおよびノイズ領域NRを考慮して自律移動体10の移動を制御する。
例えば、図10が示す状況において、図中に二点鎖線で示すように、自律移動体10を最短距離でユーザUに接近させた場合、自律移動体10がノイズ領域NR内を移動することとなる。しかし、本実施形態に係る動作制御部150は、ノイズマップを参照することで、自律移動体10がノイズ領域NRに進入またはノイズ領域NR内で停止することなく、かつノイズ源NSからより遠ざかる位置に自律移動体10を移動させることが可能である。より具体的には、動作制御部150は、図中に実線で示すようにノイズ領域NRを迂回させ、接近対象であるユーザUを中心にノイズ源NSとは反対側に自律移動体10を移動させてよい。
このように、本実施形態に係る動作制御部150は、ノイズマップを参照することにより、ノイズ源やノイズ領域を正確に把握し、目的音の入力レベルが上昇し、かつ非目的音の入力レベルが低下する位置に自律移動体10を移動させることが可能である。本実施形態に係る動作制御部150による上記の動作制御によれば、SN比を向上させ、目的音に係る音声認識精度を効果的に向上させることが可能となる。
なお、本実施形態に係る動作制御部150は、必ずしも、接近対象を中心にノイズ源とは反対側に自律移動体10を移動させなくてもよい。例えば、図11に示す一例の場合、ノイズ源NSとユーザUを結ぶ直線上には壁が存在している。このように、ノイズ源と接近対象を結ぶ直線上に障害物がある場合、動作制御部150は、上記反対側に近い位置かつノイズ領域NRに進入しない位置で自律移動体10を停止させてよい。この場合であっても、目的音の入力レベル上昇と非目的音の入力レベル低下を両立し、SN比を向上させる効果を得ることが可能である。なお、動作制御部150は、ノイズマップに含まれる壁や家具などの情報に基づいて障害物の存在を把握してもよい、認識部120が認識した障害物の存在に基づいて上記のような動作制御を行ってもよい。
次に、本実施形態に係る接近対象が発話ユーザではない場合の動作制御について説明する。図12は、本実施形態に係る接近対象が発話ユーザではない場合における動作制御について説明するための図である。
図12に示す一例の場合、ユーザU1は、図10や図11に示した一例とは異なり、自身のところではなく、ユーザU2のもとへ移動を指示する旨の音声発話UO2を行っている。この際、本実施形態に係る動作制御部150は、認識部120が認識した発話音声UO2に基づいて、接近対象をユーザU2に設定し、ユーザU2に接近するよう自律移動体10の動作を制御する。
このように、本実施形態に係る接近対象は、音声発話を発した発話ユーザのみではなく、当該発話音声に基づく音声認識処理により特定された他のユーザなどの動体、充電ステーションなどの固定物体、または任意の位置であってもよい。
本実施形態に係る動作制御部150は、接近対象が発話ユーザではない場合であっても、同様にノイズマップを参照し、ノイズ源NSやノイズ領域NRを考慮した移動を自律移動体10に実行させることができる。図12に示す一例の場合、動作制御部150は、ユーザU2を中心にノイズ源NSとは反対側に自律移動体10を移動させている。本実施形態に係る動作制御部150による上記の動作制御によれば、この後発せられることが予想されるユーザU2の発話音声、すなわち目的音に係る音声認識の精度を効果的に向上させることが可能となる。
以上、本実施形態に係るノイズマップに基づく動作制御について説明した。上述したように、本実施形態に係る動作制御部150は、周辺環境保持部140が保持するノイズマップを参照することで、目的音の入力レベルのみではなく非目的音の入力レベルを考慮した自律移動体10の移動を実現し、SM比を向上させることで精度の高い音声認識を実現することができる。
なお、上記では、本実施形態に係る動作制御部150が、目的音が検出されたことに基づいて、接近対象へ移動するよう自律移動体10を制御する場合を主な例に述べたが、本実施形態に移動のトリガーは係る例に限定されない。本実施形態に係る動作制御部150は、例えば、ユーザの顔が認識されたことや、ユーザによる移動指示に係るジェスチャが認識されたことことに基づいて、自律移動体10を接近対象に移動させる制御を行ってもよい。この場合であっても、ノイズマップを参照し、非目的音の入力レベルがより低下する位置に自律移動体を移動させることで、後に発せられると予想される目的音に係る音声認識の精度を高めることが可能である。
<<2.3.ノイズマップの生成と更新>>
次に、本実施形態に係るノイズマップの生成と更新について詳細に説明する。本実施形態に係る周辺環境推定部130は、例えば、音源方向推定や音圧測定の結果に基づいて上述したようなノイズマップを生成することができる。
次に、本実施形態に係るノイズマップの生成と更新について詳細に説明する。本実施形態に係る周辺環境推定部130は、例えば、音源方向推定や音圧測定の結果に基づいて上述したようなノイズマップを生成することができる。
まず、本実施形態に係る音源方向推定に基づくノイズマップの生成について説明する。図13は、本実施形態に係る音源方向推定に基づくノイズマップの生成について説明するための図である。
音源方向推定に基づくノイズマップの生成において、周辺環境推定部130は、まず任意の地点において音源定位を行い音源の方向を推定する。図13に示す一例の場合、周辺環境推定部130は、地点P1において、ノイズ源NS1およびNS2の方向をそれぞれ推定している。なお、この時点においては、ノイズ源NS1およびNS2の方向は推定可能であるが、自律移動体10とノイズ源NS1およびNS2との距離は不明である。
続いて、周辺環境推定部130は、前回推定された音源方向とは異なる地点に移動し、再び音源定位を行い音源の方向を推定する。図13に示す一例の場合、周辺環境推定部130は、地点P2において、再びノイズ源NS1およびNS2の方向をそれぞれ推定している。この際、周辺環境推定部130は、自律移動体10移動距離と、地点P1およびP2において推定した方向の交点とから、空間上におけるノイズ源NS1およびNS2の位置を推定することができる。
この後、周辺環境推定部130は、また別の地点において音源方向推定を繰り返すことで、音源位置に係る推定精度を向上させることが可能である。図13に示す一例の場合、周辺環境推定部130は、地点P3において再度ノイズ源NS1およびNS2の方向を推定している。
このように、本実施形態に係る周辺環境推定部130は、複数地点において音源方向の推定を繰り返すことで、空間上におけるノイズ源NS1およびNS2の位置を精度高く推定し、例えば、それぞれの推定位置から所定距離にある領域をノイズ領域NR1およびNR2として設定したノイズマップを生成することが可能である。
続いて、本実施形態に係る音圧測定に基づくノイズマップの生成について説明する。上述した音源方向推定に基づくノイズマップの生成は、自律移動体10が、同時に発生している音源数よりも多い数のマイクロフォンを備えていない場合、困難である。一方、本実施形態に係る周辺環境推定部130は、自律移動体10が単一のマイクロフォンのみしか備えていない場合であっても、下記に説明する音圧測定に基づいてノイズマップを生成することが可能である。図14は、本実施形態に係る音源方向推定に基づくノイズマップの生成について説明するための図である。
音源方向推定に基づくノイズマップの生成において、周辺環境推定部130は、まず任意の地点において音圧レベルの測定を実行する。図14に示す一例の場合、周辺環境推定部130は、地点P4において音圧レベルを測定している。続いて、周辺環境推定部130は、測定済の地点とは異なる他の地点において音圧レベルの測定を繰り返し実行する。図14に示す一例の場合、周辺環境推定部130は、地点P5およびP6において音圧レベルの測定を実行している。
このように、本実施形態に係る周辺環境推定部130は、複数地点において音圧測定を繰り返すことで、図14に示すように音圧レベルに係る等圧線を推定し、音圧レベルが閾値以上であるエリアをノイズ領域として設定することができる。また、周辺環境推定部130は、ノイズ領域において最も音圧レベルが高い地点をノイズ源の位置として推定することも可能である。図14に示す一例の場合、周辺環境推定部130は、推定した等圧線に基づいてノイズ領域NR1およびNR2を設定し、またノイズ源NS1およびNS2の位置を推定している。
このように、本実施形態に係る周辺環境推定部130によれば、自律移動体10が単一のマイクロフォンのみを備える場合であっても、複数地点において音圧測定を繰り返し実行することで、精度の高いノイズマップを生成することが可能である。なお、音圧測定に基づくノイズマップ生成を行う場合には、目的音と非目的音を切り分けることが必要となるが、当該切り分けは、上述した認識部120の機能により実現可能である。
また、本実施形態に係るノイズマップは、ノイズ源の種別などの情報を含んでもよい。図15は、本実施形態に係るノイズ源の種別情報を含むノイズマップの一例を示す図である。図15に示す一例の場合、ノイズ源NS1およびNS2がそれぞれキッチンおよびテレビジョン装置であることがノイズマップの情報として含まれていることがわかる。
本実施形態に係る周辺環境推定部130は、例えば、認識部120による物体認識の結果に基づいて、図15に示すようなノイズ源の種別情報を含んだノイズマップを生成することが可能である。本実施形態に係る周辺環境推定部130が有する上記の機能によれば、動作制御部150が、ノイズ源の特定に応じてより精度の高い動作制御を行うことが可能となる。
以上、本実施形態に係るノイズマップ生成の一例について説明した。続いて、本実施形態に係るノイズマップの生成または更新に係るタイミングについて説明する。
例えば、ノイズマップの更新については、常時動的に行われる場合も想定される。この場合、周囲で発生する非目的音を漏れなく検出できる一方、動作制御部150による動作制御に対し有用ではない突発音などの情報もすべてノイズマップの情報として含まれることとなる。また、ノイズマップの更新を常時動的に実行する場合、計算量が膨大となるため、高性能のプロセッサなどが必要となる。
このため、本実施形態に係る周辺環境推定部130は、非目的音の収集に関し効率が良い条件においてのみノイズマップの生成処理や更新処理を実行してよい。ここで、上記の効率が良い条件とは、非目的音が多数発生し得る状況である。また、非目的音が多数発生し得る状況とは、すなわちユーザが空間上において活動を行っている状況といる。このため、本実施形態に係る周辺環境推定部130は、ユーザが自律移動体10が設置される空間に存在するタイミングで、ノイズマップの生成処理や更新処理を実行してよい。
この際、本実施形態に係る周辺環境推定部130は、ユーザの予定や各種のセンサ情報に基づいて、ユーザの不在あるいは存在を推定し、ユーザが存在する可能性が高い条件においてのみノイズマップの生成処理や更新処理を実行することができる。
図16は、本実施形態に係るノイズマップの生成処理および更新処理に係る実行条件の設定例である。図16に示す一例の場合、ユーザの予定(不在または在宅)、鍵音の検出有無、ドア開閉音の検出有無、「ただいま」などの発話検出有無、人感センサによる動体の検出有無などの要素の組み合わせごとに、生成処理や更新処理を実施するか否かが設定されている。なお、図16では、上記すべての要素において、ユーザが自律移動体10と同一の空間に存在すると判定された場合にのみ、処理を実行する場合の一例が示されている。
なお、上記のような実行可否は、自律移動体10の特性や状況などに応じて動的に設定可能であってもよい。このように、本実施形態に係る周辺環境推定部130によれば、周辺環境にユーザが存在する時間帯において収集された非目的音に基づいてノイズマップの生成や更新を行うことで、より精度の高いノイズマップを保持することが可能となる。
次に、本実施形態に係るノイズマップの更新処理について詳細に説明する。上述したように、本実施形態に係る周辺環境推定部130は、周辺環境にユーザが存在する時間帯において収集された非目的音に基づいてノイズマップを動的に更新することが可能である。
しかし、この場合、集音が行われるタイミングでは、発生する非目的音が異なる場合も想定される。このため、単純に最新の集音データに基づいてノイズマップを上書きした場合、突発音など本来影響の少ない非目的音の情報がノイズマップに含まれることとなり、動作制御部150による動作制御の精度が低下する要因ともなりかねない。
このため、本実施形態に係る周辺環境推定部130は、最新の集音データに基づいて既存のノイズマップを上書きするのではなく、最新の集音データを既存のノイズマップに統合することで、当該ノイズマップの更新を行ってよい。
図17は、本実施形態に係るノイズマップの統合処理について説明するための図である。図17に示す一例の場合、本実施形態に係る周辺環境推定部130は、3回の集音に基づいて、ノイズマップの統合を行っている。ここで、1回目の集音において、キッチンおよびテレビジョン装置に係る非目的音が検出され、2回目の集音において、窓およびテレビジョン装置に係る非目的音が検出され、3回目の集音において、キッチンおよびテレビジョン装置に係る非目的音が検出された場合を想定する。
この場合、本実施形態に係る周辺環境推定部130は、3回分の集音データを、例えば平均化するなどして統合することで、ノイズマップの更新を行ってよい。本実施形態に係る周辺環境推定部130によれば図17に示すように、各非目的音の発生頻度などをノイズマップに反映することができ、突発音など発生頻度の少ない非目的音の影響を抑えることが可能となる。なお、図17に示す一例の場合、非目的音の発生頻度は、ハッチングの濃度の濃淡(濃度が濃いほど発生頻度が高い)により示されている。この場合、動作制御部150は、非目的音を発する頻度が高いテレビジョン装置を、より重点的に回避するように、自律移動体10の移動を制御してよい。
以上、本実施形態に係るノイズマップの生成および更新について説明した。なお、上記で示した手法はあくまで一例であり、本実施形態に係るノイズマップの生成および更新は、係る例に限定されない。
本実施形態に係る周辺環境推定部130は、例えば、ユーザが入力した情報に基づいてノイズマップの生成や更新を行ってもよい。図18および図19は、本実施形態に係るユーザ入力に基づくノイズマップの生成および更新について説明するための図である。
例えば、本実施形態に係る周辺環境推定部130は、図18に示すように、情報処理端末20などを介してユーザが入力した家具の配置情報などに基づいて、ノイズマップの生成や更新を行ってもよい。この際、周辺環境推定部130は、情報処理端末20が備える表示部において、ユーザの部屋に見立てた入力領域IAに、各家具に対応するアイコンICを配置するように依頼し、入力された情報に基づいて、ノイズマップの生成や更新を実行することができる。
また、例えば、本実施形態に係る周辺環境推定部130は、図19に示すように、ユーザUが行う指差しなどのジェスチャやノイズ源の教示に係る発話音声UO3に基づいて、ノイズ源NSを特定し、ノイズマップに反映することが可能である。
<<2.4.ノイズ源の回避優先度に基づく動作制御>>
次に、本実施形態に係るノイズ源の回避優先度に基づく動作制御について説明する。上記では、本実施形態に係る動作制御部150が、ノイズマップを参照し、ノイズ領域を避けるように自律移動体10の移動を制御することについて述べた。
次に、本実施形態に係るノイズ源の回避優先度に基づく動作制御について説明する。上記では、本実施形態に係る動作制御部150が、ノイズマップを参照し、ノイズ領域を避けるように自律移動体10の移動を制御することについて述べた。
しかし、状況によっては、ノイズ領域を回避して移動することが困難な場合も想定される。図20は、本実施形態に係るノイズ領域の回避が困難な状況の一例を示す図である。図20の左側には、音源方向推定に基づくノイズマップが、図20の右側には音圧測定に基づくノイズマップがそれぞれ例示されている。ここで、両ノイズマップに着目すると、いずれのノイズマップにおいても、自律移動体10がノイズ領域NR1~NR4に囲まれており、他の場所への移動が困難であることがわかる。
このような場合、本実施形態に係る動作制御部150は、ノイズ源NS1~NS4に係る回避優先度に基づいて、より回避優先度が低いノイズ源に対応するノイズ領域に移動するよう自律移動体10を制御してよい。
ここで、本実施形態に係る回避優先度は、例えば、ノイズ源が発する非目的音の種別や特性により決定され得る。上述したように、本実施形態に係る非目的音は、目的音以外の種々の音を含む。一方、非目的音が目的音に係る音声認識の精度に与える影響は、非目的音の特性に応じて異なる。
このため、本実施形態に係る周辺環境推定部130は、音声認識精度に対する影響度の大きさに基づいて、非目的音を分類し、当該影響度の大きい順に回避優先度を高く設定したノイズマップを生成してよい。
ここでは、非目的を4つのカテゴリ1~4に分類する例を示す。例えば、カテゴリ1は、比較的音量が大きく、人間の音声でありながら目的音ではない非目的音であってよい。カテゴリ1には、例えば、テレビジョン装置やラジオ、その他の装置が出力する音声や、ヴォーカルを含む楽曲、ユーザではない第三者同士による会話などが含まれる。カテゴリ1は、4つのカテゴリ中、音声認識精度に与える影響が最も高く、かつ回避優先度が最も高い非目的音であってよい。
また、カテゴリ2は、発生が非定常的かつ比較的音量が大きいためノイズ抑圧の効果が十分に得にくい非目的音であってよい。カテゴリ2には、例えば、食器洗いや調理などの作業音、窓を解放した際に流入する屋外音などが含まれる。カテゴリ2は、4つのカテゴリ中、音声認識精度に与える影響が2番目に高く、かつ回避優先度が2番目に高い非目的音であってよい。
また、カテゴリ3は、発生が定常的であり、比較的にノイズ抑圧の効果が得やすい非目的音であってよい。カテゴリ3には、例えば、エアコンディショナーや換気扇、PCのファンなどが発する音が含まれる。カテゴリ2は、4つのカテゴリ中、音声認識精度に与える影響が3番目に高く、かつ回避優先度が3番目に高い非目的音であってよい。
また、カテゴリ4は、突発的に発せられるため、瞬時的にしか影響のない非目的音であってよい。カテゴリ4には、例えば、ドアの開扉音、足音、電子レンジが発する音などが含まれる。カテゴリ4は、4つのカテゴリ中、音声認識精度に与える影響が最も低く、かつ回避優先度が最も低い非目的音であってよい。
このように、本実施形態に係る周辺環境推定部130は、非目的音の特性に応じた回避優先度を設定したノイズマップを生成することが可能である。
また、本実施形態に係るノイズ源の回避優先度は、非目的音に係る音響的な何らかの定量的な指標に基づいて設定されてもよい。上記の定量的な指標には、例えば、音声らしさの度合いを示す指標や、定常性の度合いを示す指標が挙げられる。
一般に、音声認識の対象となる目的音、すなわちユーザの発話音声は、「非定常的」な「音声」である。一方、「非定常的」な「音声」には、例えば、第三者同士による会話や、テレビジョン装置やラジオが出力する音声などの非目的音も含まれる。このため、「非定常的」な「音声」である目的音に係る音声認識の精度を向上させるためには、当該目的音との分離が困難である「非定常的」な「音声」である非目的音を回避することが重要となる。
一方、「非定常的」な「非音声」としては、食器洗いや調理に係る作業音などが、「定常的」な「非音声」には、エアコンディショナー、換気扇、PCのファンなどが発する音が挙げられるが、このような非目的音は、目的音との分離が比較的容易であるため、上記の「非定常的」な「音声」である非目的音と比べ回避優先度は低いといえる。
また、「定常的」な「音声」としては、例えば、「あーーーー」、など同じ音を長く発話し続ける場合が該当するが、このような音は日常において生じる可能性が著しく低いため、無視してもよい。
以上のことから、本実施形態に係る周辺環境推定部130は、音声らしさの度合いを示す指標αと定常性の度合いを示す指標βとに基づいて、非目的音が目的音に係る音声認識に与える影響度を算出し、算出した値に基づいて回避優先度を設定してよい。
図21および図22は、本実施形態に係る音声らしさの度合いを示す指標αと定常性の度合いを示す指標βの算出について説明するための図である。図21には、本実施形態に係る自律移動体10が、複数のマイクロフォンを備える場合における算出の流れが示されている。
自律移動体10が複数のマイクロフォンを備える場合、周辺環境推定部130は、図21に示すように、ノイズ源NS1~NS4の方向へ順にビームフォーミングを実行し、指標αおよび指標βを算出する。係る手法によれば、ノイズ領域NR1~NR4に進入することなくノイズ源NR1~NR4に係る指標αおよび指標βを算出することが可能であることから、算出中に目的音が検出された場合であっても、当該目的音に係る音声認識の精度を維持することができる。
また、図22は、本実施形態に係る自律移動体10が、単一のマイクロフォンを備える場合における算出の流れが示されている。自律移動体10が単一のマイクロフォンを備える場合、周辺環境推定部130は、図22に示すように、ノイズ源NS1~NS4の周辺において、それぞれ指標αおよび指標βを算出してよい。なお、動作制御部150は、各ノイズ源に係る指標αおよび指標βの算出が完了した場合、速やかにノイズ領域NR1~NR4から脱出するよう自律移動体10を制御してよい。
以上、本実施形態に係る指標αおよび指標βの算出の流れについて説明した。本実施形態に係る周辺環境推定部130は、上記のように算出した指標αおよび指標βに基づいて、各ノイズ源の影響度を算出し、当該影響度に基づいて回避優先度を設定することが可能である。周辺環境推定部130は、例えば、指標αおよび指標βの合計値を影響度の高さとし、当該合計値が高い順に回避優先度を高く設定してもよい。
なお、本実施形態に係る周辺環境推定部130は、音声らしさの度合いを示す指標αを、例えば、音のスペクトロエントロピーに基づいて算出してもよい。音のスペクトロエントロピーは、VAD(Voice Activity Detection)技術にも用いられる指標であり、人間の音声は他の音と比べて低い値を示す傾向にある。
本実施形態に係る周辺環境推定部130は、音のスペクトロエントロピー、すなわち指標αを例えば、下記の数式(1)により算出することができる。なお、数式(1)におけるfは、ある周波数を示し、Sfは、観測信号の周波数fの振幅スペクトルを示す。また、数式(1)におけるPfは、下記の数式(2)により定義される。
また、本実施形態に係る周辺環境推定部130は、定常性の度合いを示す指標βを、例えば、音の尖度に基づいて算出してもよい。音の尖度とは、音の定常性、非定常性を判別するためにしばしば用いられる指標であり、下記の数式(3)により算出され得る。なお、下記の数式(3)におけるTは、尖度計算を行う音区間長を示し、例えば、3~5秒などの長さが設定されてよい。また、数式(3)におけるtは、ある時刻を示し、x(t)は、時刻tにおける音声波形を示す。
以上、本実施形態に係るノイズ源の回避優先度の設定について説明した。本実施形態に係る回避優先度の設定によれば、自律移動体10が、目的音に係る音声認識に影響を与える非目的音を優先的に回避することが可能となる。
<<2.5.動作の流れ>>
次に、本実施形態に係る自律移動体10の動作の流れについて詳細に説明する。まず、本実施形態に係るノイズマップの更新の流れについて説明する。図23は、本実施形態に係るノイズマップ更新の流れを示すフローチャートである。
次に、本実施形態に係る自律移動体10の動作の流れについて詳細に説明する。まず、本実施形態に係るノイズマップの更新の流れについて説明する。図23は、本実施形態に係るノイズマップ更新の流れを示すフローチャートである。
図23を参照すると、まず、周辺環境推定部130が、入力部110が収集したセンサ情報や認識部120による認識結果に基づいて、周辺環境の推定を行う(S1101)。具体的には、周辺環境推定部130は、ノイズ源やノイズ領域の推定を行う。
次に、周辺環境推定部130は、周辺環境保持部140に既存のノイズマップが保持されているか否かを判定する(S1102)。
ここで、周辺環境保持部140に保持されているノイズマップが存在しない場合(S1102:NO)、周辺環境推定部130は、ステップS1101において推定した周辺環境に基づいてノイズマップを生成し、周辺環境保持部140に記憶させる(S1107)。
一方、周辺環境保持部140に既存のノイズマップが存在する場合(S1102:YES)、周辺環境推定部130は、続いて、推定した周辺環境と既存のノイズマップとでノイズ源の数が変化しているか否かを判定する(S1103)。
ここで、ノイズ源の数が変化している場合(S1103:YES)、周辺環境推定部130は、ステップS1101において推定した周辺環境に基づいてノイズマップを統合し(S1106)、統合後のノイズマップを周辺環境保持部140に記憶させる(S1107)。
一方、ノイズ源の数が変化していない場合(S1103:NO)、周辺環境推定部130は、続いて、推定した周辺環境と既存のノイズマップとでノイズ源の位置が変化しているか否かを判定する(S1104)。
ここで、ノイズ源の位置が変化している場合(S1104:YES)、周辺環境推定部130は、ステップS1101において推定した周辺環境に基づいてノイズマップを統合し(S1106)、統合後のノイズマップを周辺環境保持部140に記憶させる(S1107)。
一方、ノイズ源の位置が変化していない場合(S1104:NO)、周辺環境推定部130は、続いて、推定した周辺環境と既存のノイズマップとでノイズ源が発する非目的音の音圧が変化しているか否かを判定する(S1105)。
ここで、ノイズ源が発する非目的音の音圧が変化している場合(S1105:YES)、周辺環境推定部130は、ステップS1101において推定した周辺環境に基づいてノイズマップを統合し(S1106)、統合後のノイズマップを周辺環境保持部140に記憶させる(S1107)。
一方、ノイズ源が発する非目的音の音圧が変化していない場合(S1105:NO)、周辺環境推定部130は、ノイズマップを更新せず、周辺環境保持部140に既存のノイズマップを維持させる。
次に、本実施形態に係る動作制御の流れについて詳細に説明する。図24は、本実施形態に係る動作制御の流れを示すフローチャートである。
図24を参照すると、動作制御部150は、まず、周辺環境保持部140が保持するノイズマップの読み込みを行う(S1201)。
続いて、動作制御部150は、ステップS1201において読み込んだノイズマップに基づいて、自律移動体10にノイズ領域を回避した自律行動を行わせる(S1202)。
また、動作制御部150は、ステップS1202における自律行動中に、目的音が検出されたか否かを継続的に判定する(S1203)。
ここで、目的音が検出された場合(S1203:YES)、動作制御部150は、ステップS1201において読み込んだノイズマップに基づいて、接近対象の周辺において非目的音の入力レベルがより低下する位置に自律移動体10を移動させる(S1204)。
次に、動作制御部150は、目的音の音声認識結果に基づいて、対応する動作を自律移動体10に実行させる(S1205)。
以上、本実施形態に係る自律移動体10の動作の流れについて説明した。なお、上記では、本実施形態に係る自律移動体10がSN比を向上させるために、目的音と非目的音の入力レベルを考慮した移動を行うことを中心に述べた。しかし、本実施形態に係るSN比の向上手法は係る例に限定されず、例えば、信号処理やビームフォーミング技術が併用して用いられてもよい。
例えば、本実施形態に係る動作制御部150は、自律移動体10を、接近対象とノイズ源の間に移動させ、接近対象の方向に対しビームフォーミングを張るよう制御してもよい。自律移動体10が犬型のロボット装置である場合、動作制御部150は、接近対象であるユーザの顔の高さに応じた仰角を以ってビームフォーミングが張られるよう制御してよい。この場合、水平方向から到達する非目的音を効果的に排除し、SN比を効果的に向上させる効果が期待される。
また、本実施形態に係る動作制御部150は、ノイズ領域を回避するために、例えば、ユーザを誘導する動作を自律移動体10に行わせてもよい。動作制御部150は、例えば、接近対象であるユーザがノイズ領域に居る場合、ユーザがノイズ領域から離れ自律移動体10に近づくよう誘導する動作を自律移動体10に行わせることで、ノイズ領域に進入することなく目的音の入力レベルを上昇させることが可能である。上記の誘導は、例えば、吠える、ノイズ領域手前で止まる、うろつくなどの動作により実現され得る。また、自律移動体10が、例えば、人型のロボット装置など、言語によるコミュニケーション機能を有する場合には、ノイズ領域から離れてほしい旨を明示的に音声で伝えてもよい。
<3.まとめ>
以上説明したように、本開示の一実施形態に係る情報処理装置の一例である自律移動体10は、認識処理に基づいて自律移動体10の動作を制御する動作制御部150を備える。また、本開示の一実施形態に係る動作制御部150は、音声認識処理の対象音声である目的音が検出された場合、当該目的音に基づいて定まる接近対象の周辺において対象音声ではない非目的音の入力レベルがより低下する位置に、自律移動体10を移動させること、を特徴の一つとする。係る構成によれば、自律移動体に音声認識の精度をより向上させる動作を実行させることが可能となる。
以上説明したように、本開示の一実施形態に係る情報処理装置の一例である自律移動体10は、認識処理に基づいて自律移動体10の動作を制御する動作制御部150を備える。また、本開示の一実施形態に係る動作制御部150は、音声認識処理の対象音声である目的音が検出された場合、当該目的音に基づいて定まる接近対象の周辺において対象音声ではない非目的音の入力レベルがより低下する位置に、自律移動体10を移動させること、を特徴の一つとする。係る構成によれば、自律移動体に音声認識の精度をより向上させる動作を実行させることが可能となる。
以上、添付図面を参照しながら本開示の好適な実施形態について詳細に説明したが、本開示の技術的範囲はかかる例に限定されない。本開示の技術分野における通常の知識を有する者であれば、請求の範囲に記載された技術的思想の範疇内において、各種の変更例または修正例に想到し得ることは明らかであり、これらについても、当然に本開示の技術的範囲に属するものと了解される。
また、本明細書に記載された効果は、あくまで説明的または例示的なものであって限定的ではない。つまり、本開示に係る技術は、上記の効果とともに、または上記の効果に代えて、本明細書の記載から当業者には明らかな他の効果を奏しうる。
また、本明細書の自律移動体10の処理に係る各ステップは、必ずしもフローチャートに記載された順序に沿って時系列に処理される必要はない。例えば、自律移動体10の処理に係る各ステップは、フローチャートに記載された順序と異なる順序で処理されても、並列的に処理されてもよい。
なお、以下のような構成も本開示の技術的範囲に属する。
(1)
認識処理に基づいて行動を行う自律移動体の動作を制御する動作制御部、
を備え、
前記動作制御部は、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、
情報処理装置。
(2)
前記動作制御部は、前記目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記目的音の入力レベルがより上昇し、前記非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、
前記(1)に記載の情報処理装置。
(3)
前記動作制御部は、前記目的音が検出された場合、前記非目的音を発するノイズ源からより遠ざかり、前記接近対象により近づく位置に、前記自律移動体を移動させる、
前記(1)または(2)に記載の情報処理装置。
(4)
前記動作制御部は、前記目的音が検出された場合、前記接近対象を中心に前記非目的音を発するノイズ源とは反対側に前記自律移動体を移動させる、
前記(1)~(3)のいずれかに記載の情報処理装置。
(5)
前記目的音は、ユーザの発話音声であり、
前記接近対象は、前記発話音声を発した発話ユーザである、
前記(1)~(4)のいずれかに記載の情報処理装置。
(6)
前記接近対象は、ユーザの発話音声に基づく前記音声認識処理により特定された動体、固定物体、または位置である、
前記(1)~(5)のいずれかに記載の情報処理装置。
(7)
前記動作制御部は、周囲環境における前記非目的音の発生状況を示すノイズマップに基づいて、前記自律移動体の動作を制御する、
前記(1)~(6)のいずれかに記載の情報処理装置。
(8)
前記ノイズマップは、前記非目的音を発するノイズ源の情報を含み、
前記動作制御部は、前記ノイズ源に係る回避優先度に基づいて、前記自律移動体の動作を制御する、
前記(7)に記載の情報処理装置。
(9)
前記ノイズ源に係る回避優先度は、前記ノイズ源の種別に基づいて決定される、
前記(8)に記載の情報処理装置。
(10)
前記ノイズ源に係る回避優先度は、前記ノイズ源が発する前記非目的音の前記音声認識処理に対する影響度に基づいて決定される、
前記(8)に記載の情報処理装置。
(11)
前記影響度は、前記非目的音の音声らしさの度合いを示す指標または定常性の度合いを示す指標のうち少なくともいずれかに基づいて算出される、
前記(10)に記載の情報処理装置。
(12)
前記動作制御部は、前記目的音が検出されない場合、前記ノイズマップに基づいて、前記非目的音の入力を回避するよう前記自律移動体の動作を制御する、
前記(7)~(10)のいずれかに記載の情報処理装置。
(13)
前記動作制御部は、前記目的音が検出されない場合、前記ノイズマップに基づいて、前記非目的音の入力レベルが閾値以下となるエリア内に前記自律移動体の移動範囲を制限する、
前記(7)~(12)のいずれかに記載の情報処理装置。
(14)
前記ノイズマップを生成する周辺環境推定部、
をさらに備える、
前記(7)~(13)のいずれかに記載の情報処理装置。
(15)
前記周辺環境推定部は、前記非目的音を発するノイズ源に係る方向推定または音圧測定に基づいて、前記ノイズマップを生成する、
前記(14)に記載の情報処理装置。
(16)
前記周辺環境推定部は、収集された前記非目的音に基づいて、前記ノイズマップを動的に更新する、
前記(14)または(15)に記載の情報処理装置。
(17)
前記周辺環境推定部は、前記非目的音を発するノイズ源の数、位置、または音圧の変化に基づいて、前記ノイズマップを動的に更新する、
前記(16)に記載の情報処理装置。
(18)
前記周辺環境推定部は、周辺環境にユーザが存在する時間帯において収集された前記非目的音に基づいて、前記ノイズマップの生成または更新を行う、
前記(16)または(17)に記載の情報処理装置。
(19)
プロセッサが、認識処理に基づいて行動を行う自律移動体の動作を制御すること、
を含み、
前記制御することは、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させること、
をさらに含む、
情報処理方法。
(20)
コンピュータを、
認識処理に基づいて行動を行う自律移動体の動作を制御する動作制御部、
を備え、
前記動作制御部は、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、
情報処理装置、
として機能させるためのプログラム。
(1)
認識処理に基づいて行動を行う自律移動体の動作を制御する動作制御部、
を備え、
前記動作制御部は、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、
情報処理装置。
(2)
前記動作制御部は、前記目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記目的音の入力レベルがより上昇し、前記非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、
前記(1)に記載の情報処理装置。
(3)
前記動作制御部は、前記目的音が検出された場合、前記非目的音を発するノイズ源からより遠ざかり、前記接近対象により近づく位置に、前記自律移動体を移動させる、
前記(1)または(2)に記載の情報処理装置。
(4)
前記動作制御部は、前記目的音が検出された場合、前記接近対象を中心に前記非目的音を発するノイズ源とは反対側に前記自律移動体を移動させる、
前記(1)~(3)のいずれかに記載の情報処理装置。
(5)
前記目的音は、ユーザの発話音声であり、
前記接近対象は、前記発話音声を発した発話ユーザである、
前記(1)~(4)のいずれかに記載の情報処理装置。
(6)
前記接近対象は、ユーザの発話音声に基づく前記音声認識処理により特定された動体、固定物体、または位置である、
前記(1)~(5)のいずれかに記載の情報処理装置。
(7)
前記動作制御部は、周囲環境における前記非目的音の発生状況を示すノイズマップに基づいて、前記自律移動体の動作を制御する、
前記(1)~(6)のいずれかに記載の情報処理装置。
(8)
前記ノイズマップは、前記非目的音を発するノイズ源の情報を含み、
前記動作制御部は、前記ノイズ源に係る回避優先度に基づいて、前記自律移動体の動作を制御する、
前記(7)に記載の情報処理装置。
(9)
前記ノイズ源に係る回避優先度は、前記ノイズ源の種別に基づいて決定される、
前記(8)に記載の情報処理装置。
(10)
前記ノイズ源に係る回避優先度は、前記ノイズ源が発する前記非目的音の前記音声認識処理に対する影響度に基づいて決定される、
前記(8)に記載の情報処理装置。
(11)
前記影響度は、前記非目的音の音声らしさの度合いを示す指標または定常性の度合いを示す指標のうち少なくともいずれかに基づいて算出される、
前記(10)に記載の情報処理装置。
(12)
前記動作制御部は、前記目的音が検出されない場合、前記ノイズマップに基づいて、前記非目的音の入力を回避するよう前記自律移動体の動作を制御する、
前記(7)~(10)のいずれかに記載の情報処理装置。
(13)
前記動作制御部は、前記目的音が検出されない場合、前記ノイズマップに基づいて、前記非目的音の入力レベルが閾値以下となるエリア内に前記自律移動体の移動範囲を制限する、
前記(7)~(12)のいずれかに記載の情報処理装置。
(14)
前記ノイズマップを生成する周辺環境推定部、
をさらに備える、
前記(7)~(13)のいずれかに記載の情報処理装置。
(15)
前記周辺環境推定部は、前記非目的音を発するノイズ源に係る方向推定または音圧測定に基づいて、前記ノイズマップを生成する、
前記(14)に記載の情報処理装置。
(16)
前記周辺環境推定部は、収集された前記非目的音に基づいて、前記ノイズマップを動的に更新する、
前記(14)または(15)に記載の情報処理装置。
(17)
前記周辺環境推定部は、前記非目的音を発するノイズ源の数、位置、または音圧の変化に基づいて、前記ノイズマップを動的に更新する、
前記(16)に記載の情報処理装置。
(18)
前記周辺環境推定部は、周辺環境にユーザが存在する時間帯において収集された前記非目的音に基づいて、前記ノイズマップの生成または更新を行う、
前記(16)または(17)に記載の情報処理装置。
(19)
プロセッサが、認識処理に基づいて行動を行う自律移動体の動作を制御すること、
を含み、
前記制御することは、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させること、
をさらに含む、
情報処理方法。
(20)
コンピュータを、
認識処理に基づいて行動を行う自律移動体の動作を制御する動作制御部、
を備え、
前記動作制御部は、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、
情報処理装置、
として機能させるためのプログラム。
10 自律移動体
110 入力部
120 認識部
130 周辺環境推定部
140 周辺環境保持部
150 動作制御部
160 駆動部
170 出力部
110 入力部
120 認識部
130 周辺環境推定部
140 周辺環境保持部
150 動作制御部
160 駆動部
170 出力部
Claims (20)
- 認識処理に基づいて行動を行う自律移動体の動作を制御する動作制御部、
を備え、
前記動作制御部は、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、
情報処理装置。 - 前記動作制御部は、前記目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記目的音の入力レベルがより上昇し、前記非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、
請求項1に記載の情報処理装置。 - 前記動作制御部は、前記目的音が検出された場合、前記非目的音を発するノイズ源からより遠ざかり、前記接近対象により近づく位置に、前記自律移動体を移動させる、
請求項1に記載の情報処理装置。 - 前記動作制御部は、前記目的音が検出された場合、前記接近対象を中心に前記非目的音を発するノイズ源とは反対側に前記自律移動体を移動させる、
請求項1に記載の情報処理装置。 - 前記目的音は、ユーザの発話音声であり、
前記接近対象は、前記発話音声を発した発話ユーザである、
請求項1に記載の情報処理装置。 - 前記接近対象は、ユーザの発話音声に基づく前記音声認識処理により特定された動体、固定物体、または位置である、
請求項1に記載の情報処理装置。 - 前記動作制御部は、周囲環境における前記非目的音の発生状況を示すノイズマップに基づいて、前記自律移動体の動作を制御する、
請求項1に記載の情報処理装置。 - 前記ノイズマップは、前記非目的音を発するノイズ源の情報を含み、
前記動作制御部は、前記ノイズ源に係る回避優先度に基づいて、前記自律移動体の動作を制御する、
請求項7に記載の情報処理装置。 - 前記ノイズ源に係る回避優先度は、前記ノイズ源の種別に基づいて決定される、
請求項8に記載の情報処理装置。 - 前記ノイズ源に係る回避優先度は、前記ノイズ源が発する前記非目的音の前記音声認識処理に対する影響度に基づいて決定される、
請求項8に記載の情報処理装置。 - 前記影響度は、前記非目的音の音声らしさの度合いを示す指標または定常性の度合いを示す指標のうち少なくともいずれかに基づいて算出される、
請求項10に記載の情報処理装置。 - 前記動作制御部は、前記目的音が検出されない場合、前記ノイズマップに基づいて、前記非目的音の入力を回避するよう前記自律移動体の動作を制御する、
請求項7に記載の情報処理装置。 - 前記動作制御部は、前記目的音が検出されない場合、前記ノイズマップに基づいて、前記非目的音の入力レベルが閾値以下となるエリア内に前記自律移動体の移動範囲を制限する、
請求項7に記載の情報処理装置。 - 前記ノイズマップを生成する周辺環境推定部、
をさらに備える、
請求項7に記載の情報処理装置。 - 前記周辺環境推定部は、前記非目的音を発するノイズ源に係る方向推定または音圧測定に基づいて、前記ノイズマップを生成する、
請求項14に記載の情報処理装置。 - 前記周辺環境推定部は、収集された前記非目的音に基づいて、前記ノイズマップを動的に更新する、
請求項14に記載の情報処理装置。 - 前記周辺環境推定部は、前記非目的音を発するノイズ源の数、位置、または音圧の変化に基づいて、前記ノイズマップを動的に更新する、
請求項16に記載の情報処理装置。 - 前記周辺環境推定部は、周辺環境にユーザが存在する時間帯において収集された前記非目的音に基づいて、前記ノイズマップの生成または更新を行う、
請求項16に記載の情報処理装置。 - プロセッサが、認識処理に基づいて行動を行う自律移動体の動作を制御すること、
を含み、
前記制御することは、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させること、
をさらに含む、
情報処理方法。 - コンピュータを、
認識処理に基づいて行動を行う自律移動体の動作を制御する動作制御部、
を備え、
前記動作制御部は、音声認識処理の対象音声である目的音が検出された場合、前記目的音に基づいて定まる接近対象の周辺において前記対象音声ではない非目的音の入力レベルがより低下する位置に、前記自律移動体を移動させる、
情報処理装置、
として機能させるためのプログラム。
Priority Applications (5)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP19775537.4A EP3778151B1 (en) | 2018-03-30 | 2019-02-21 | Information processing device, information processing method, and program |
| JP2020510435A JP7259843B2 (ja) | 2018-03-30 | 2019-02-21 | 情報処理装置、情報処理方法、およびプログラム |
| US16/976,493 US11468891B2 (en) | 2018-03-30 | 2019-02-21 | Information processor, information processing method, and program |
| CN201980016177.1A CN111788043B (zh) | 2018-03-30 | 2019-02-21 | 信息处理装置、信息处理方法和程序 |
| US17/943,205 US12230265B2 (en) | 2018-03-30 | 2022-09-13 | Information processor, information processing method, and program |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2018-069788 | 2018-03-30 | ||
| JP2018069788 | 2018-03-30 |
Related Child Applications (2)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US16/976,493 A-371-Of-International US11468891B2 (en) | 2018-03-30 | 2019-02-21 | Information processor, information processing method, and program |
| US17/943,205 Continuation US12230265B2 (en) | 2018-03-30 | 2022-09-13 | Information processor, information processing method, and program |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2019187834A1 true WO2019187834A1 (ja) | 2019-10-03 |
Family
ID=68059750
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2019/006580 Ceased WO2019187834A1 (ja) | 2018-03-30 | 2019-02-21 | 情報処理装置、情報処理方法、およびプログラム |
Country Status (5)
| Country | Link |
|---|---|
| US (2) | US11468891B2 (ja) |
| EP (1) | EP3778151B1 (ja) |
| JP (1) | JP7259843B2 (ja) |
| CN (1) | CN111788043B (ja) |
| WO (1) | WO2019187834A1 (ja) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2022515307A (ja) * | 2019-11-28 | 2022-02-18 | 北京市商▲湯▼科技▲開▼▲發▼有限公司 | インタラクティブオブジェクト駆動方法、装置、電子デバイス及び記憶媒体 |
Families Citing this family (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11957991B2 (en) * | 2020-03-06 | 2024-04-16 | Moose Creative Management Pty Limited | Balloon toy |
| US11501794B1 (en) * | 2020-05-15 | 2022-11-15 | Amazon Technologies, Inc. | Multimodal sentiment detection |
| US12172537B2 (en) * | 2020-12-22 | 2024-12-24 | Boston Dynamics, Inc. | Robust docking of robots with imperfect sensing |
| CN117279746A (zh) * | 2021-05-19 | 2023-12-22 | 发那科株式会社 | 机器人系统 |
| CN114516061B (zh) * | 2022-02-25 | 2024-03-05 | 杭州萤石软件有限公司 | 一种机器人控制方法、机器人系统及一种机器人 |
| CN115816480A (zh) * | 2022-08-01 | 2023-03-21 | 北京可以科技有限公司 | 一种机器人及机器人的情感表达方法 |
Citations (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2002140092A (ja) * | 2000-10-31 | 2002-05-17 | Nec Corp | 音声認識ロボット |
| JP2004130427A (ja) | 2002-10-09 | 2004-04-30 | Sony Corp | ロボット装置及びロボット装置の動作制御方法 |
| JP2005103679A (ja) * | 2003-09-29 | 2005-04-21 | Toshiba Corp | ロボット装置 |
| JP2006181651A (ja) * | 2004-12-24 | 2006-07-13 | Toshiba Corp | 対話型ロボット、対話型ロボットの音声認識方法および対話型ロボットの音声認識プログラム |
| JP2007152443A (ja) * | 2005-11-30 | 2007-06-21 | Mitsubishi Heavy Ind Ltd | 片付けロボット |
| JP2007164379A (ja) * | 2005-12-12 | 2007-06-28 | Honda Motor Co Ltd | インターフェース装置およびそれを備えた移動ロボット |
| JP2007264472A (ja) * | 2006-03-29 | 2007-10-11 | Toshiba Corp | 位置検出装置、自律移動装置、位置検出方法および位置検出プログラム |
| US20120130716A1 (en) * | 2010-11-22 | 2012-05-24 | Samsung Electronics Co., Ltd. | Speech recognition method for robot |
| WO2017169826A1 (ja) * | 2016-03-28 | 2017-10-05 | Groove X株式会社 | お出迎え行動する自律行動型ロボット |
Family Cites Families (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP4797330B2 (ja) * | 2004-03-08 | 2011-10-19 | 日本電気株式会社 | ロボット |
| KR100586893B1 (ko) * | 2004-06-28 | 2006-06-08 | 삼성전자주식회사 | 시변 잡음 환경에서의 화자 위치 추정 시스템 및 방법 |
| JP2009080309A (ja) * | 2007-09-26 | 2009-04-16 | Toshiba Corp | 音声認識装置、音声認識方法、音声認識プログラム、及び音声認識プログラムを記録した記録媒体 |
| JP5075664B2 (ja) * | 2008-02-15 | 2012-11-21 | 株式会社東芝 | 音声対話装置及び支援方法 |
| JP5555987B2 (ja) * | 2008-07-11 | 2014-07-23 | 富士通株式会社 | 雑音抑圧装置、携帯電話機、雑音抑圧方法及びコンピュータプログラム |
| RU2635046C2 (ru) * | 2012-07-27 | 2017-11-08 | Сони Корпорейшн | Система обработки информации и носитель информации |
| CN105845135A (zh) * | 2015-01-12 | 2016-08-10 | 芋头科技(杭州)有限公司 | 一种机器人系统的声音识别系统及方法 |
| US9848269B2 (en) * | 2015-12-29 | 2017-12-19 | International Business Machines Corporation | Predicting harmful noise events and implementing corrective actions prior to noise induced hearing loss |
| TR201620208A1 (tr) * | 2016-12-30 | 2018-07-23 | Ford Otomotiv Sanayi As | Kompakt ti̇treşi̇m ve gürültü hari̇talandirma si̇stemi̇ ve yöntemi̇ |
-
2019
- 2019-02-21 JP JP2020510435A patent/JP7259843B2/ja active Active
- 2019-02-21 EP EP19775537.4A patent/EP3778151B1/en active Active
- 2019-02-21 WO PCT/JP2019/006580 patent/WO2019187834A1/ja not_active Ceased
- 2019-02-21 US US16/976,493 patent/US11468891B2/en active Active
- 2019-02-21 CN CN201980016177.1A patent/CN111788043B/zh active Active
-
2022
- 2022-09-13 US US17/943,205 patent/US12230265B2/en active Active
Patent Citations (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2002140092A (ja) * | 2000-10-31 | 2002-05-17 | Nec Corp | 音声認識ロボット |
| JP2004130427A (ja) | 2002-10-09 | 2004-04-30 | Sony Corp | ロボット装置及びロボット装置の動作制御方法 |
| JP2005103679A (ja) * | 2003-09-29 | 2005-04-21 | Toshiba Corp | ロボット装置 |
| JP2006181651A (ja) * | 2004-12-24 | 2006-07-13 | Toshiba Corp | 対話型ロボット、対話型ロボットの音声認識方法および対話型ロボットの音声認識プログラム |
| JP2007152443A (ja) * | 2005-11-30 | 2007-06-21 | Mitsubishi Heavy Ind Ltd | 片付けロボット |
| JP2007164379A (ja) * | 2005-12-12 | 2007-06-28 | Honda Motor Co Ltd | インターフェース装置およびそれを備えた移動ロボット |
| JP2007264472A (ja) * | 2006-03-29 | 2007-10-11 | Toshiba Corp | 位置検出装置、自律移動装置、位置検出方法および位置検出プログラム |
| US20120130716A1 (en) * | 2010-11-22 | 2012-05-24 | Samsung Electronics Co., Ltd. | Speech recognition method for robot |
| WO2017169826A1 (ja) * | 2016-03-28 | 2017-10-05 | Groove X株式会社 | お出迎え行動する自律行動型ロボット |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP3778151A4 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2022515307A (ja) * | 2019-11-28 | 2022-02-18 | 北京市商▲湯▼科技▲開▼▲發▼有限公司 | インタラクティブオブジェクト駆動方法、装置、電子デバイス及び記憶媒体 |
| JP7267411B2 (ja) | 2019-11-28 | 2023-05-01 | 北京市商▲湯▼科技▲開▼▲發▼有限公司 | インタラクティブオブジェクト駆動方法、装置、電子デバイス及び記憶媒体 |
| US11769499B2 (en) | 2019-11-28 | 2023-09-26 | Beijing Sensetime Technology Development Co., Ltd. | Driving interaction object |
Also Published As
| Publication number | Publication date |
|---|---|
| EP3778151A4 (en) | 2021-06-16 |
| CN111788043A (zh) | 2020-10-16 |
| JP7259843B2 (ja) | 2023-04-18 |
| CN111788043B (zh) | 2024-06-14 |
| EP3778151B1 (en) | 2025-03-26 |
| US12230265B2 (en) | 2025-02-18 |
| EP3778151A1 (en) | 2021-02-17 |
| JPWO2019187834A1 (ja) | 2021-07-15 |
| US11468891B2 (en) | 2022-10-11 |
| US20210050011A1 (en) | 2021-02-18 |
| US20230005481A1 (en) | 2023-01-05 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7259843B2 (ja) | 情報処理装置、情報処理方法、およびプログラム | |
| US10225656B1 (en) | Mobile speaker system for virtual reality environments | |
| JP7747032B2 (ja) | 情報処理装置及び情報処理方法 | |
| JP7120254B2 (ja) | 情報処理装置、情報処理方法、およびプログラム | |
| JP7559900B2 (ja) | 情報処理装置、情報処理方法、およびプログラム | |
| JP7375748B2 (ja) | 情報処理装置、情報処理方法、およびプログラム | |
| JP5411789B2 (ja) | コミュニケーションロボット | |
| US11433546B1 (en) | Non-verbal cuing by autonomous mobile device | |
| US12067971B2 (en) | Information processing apparatus and information processing method | |
| EP3832421A1 (en) | Information processing device, action determination method, and program |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 19775537 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2020510435 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2019775537 Country of ref document: EP |

