WO2016103710A1 - Dispositif de traitement de la voix - Google Patents
Dispositif de traitement de la voix Download PDFInfo
- Publication number
- WO2016103710A1 WO2016103710A1 PCT/JP2015/006448 JP2015006448W WO2016103710A1 WO 2016103710 A1 WO2016103710 A1 WO 2016103710A1 JP 2015006448 W JP2015006448 W JP 2015006448W WO 2016103710 A1 WO2016103710 A1 WO 2016103710A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- sound
- source
- voice
- processing unit
- audio
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B60—VEHICLES IN GENERAL
- B60R—VEHICLES, VEHICLE FITTINGS, OR VEHICLE PARTS, NOT OTHERWISE PROVIDED FOR
- B60R16/00—Electric or fluid circuits specially adapted for vehicles and not otherwise provided for; Arrangement of elements of electric or fluid circuits specially adapted for vehicles and not otherwise provided for
- B60R16/02—Electric or fluid circuits specially adapted for vehicles and not otherwise provided for; Arrangement of elements of electric or fluid circuits specially adapted for vehicles and not otherwise provided for electric constitutive elements
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/28—Constructional details of speech recognition systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0272—Voice signal separating
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
Definitions
- the present invention relates to a voice processing device.
- Various devices are provided in vehicles such as automobiles. Operations on these various devices are performed, for example, by operating operation buttons, operation panels, and the like.
- Patent Documents 1 to 3 Recently, voice recognition technology has also been proposed (Patent Documents 1 to 3).
- An object of the present invention is to provide a good speech processing apparatus capable of improving the certainty of speech recognition.
- a plurality of microphones arranged in a vehicle, and a sound source that determines an orientation of a sound source that is a sound source included in a sound reception signal acquired by each of the plurality of microphones
- An audio processing apparatus is provided that performs the beamforming in the direction of the specified audio source.
- the present invention by performing a predetermined action, it is possible to reliably specify a voice source to be a target of voice recognition. For this reason, according to this invention, the favorable audio processing apparatus which can improve the reliability of audio
- FIG. 1 is a schematic diagram showing a configuration of a vehicle.
- a driver's seat 40 that is a driver's seat and a passenger's seat 44 that is a passenger's seat are arranged at the front of a vehicle body (cabinet) 46 of a vehicle (automobile). ing.
- the driver's seat 40 is located on the right side of the passenger compartment 46, for example.
- a steering wheel (handle) 78 is disposed in front of the driver seat 40.
- the passenger seat 44 is located on the left side of the passenger compartment 46, for example.
- the driver seat 40 and the passenger seat 44 constitute a front seat.
- an audio source 72a when the driver emits audio is located.
- an audio source 72b when the passenger seat makes a sound is located.
- a rear seat 70 is disposed at the rear of the vehicle body 46.
- reference numeral 72 is used when the description is made without distinguishing between the individual sound sources, and reference numerals 72a and 72b are used when the description is made with the individual sound sources distinguished.
- a plurality of microphones 22 (22a to 22c), that is, microphone arrays are arranged in front of the front seats 40 and 44.
- reference numeral 22 is used when the description is made without distinguishing the individual microphones, and reference numerals 22a to 22c are used when the description is made with the individual microphones distinguished.
- the microphone 22 may be disposed on the dashboard 42 or may be disposed on a portion close to the roof.
- the distance between the sound source 72 of the front seats 40 and 44 and the microphone 22 is often about several tens of centimeters. However, the distance between the microphone 22 and the audio source 72 can be less than a few tens of centimeters. Also, the distance between the microphone 22 and the audio source 72 can exceed 1 m.
- a speaker (loud speaker) 76 constituting a speaker system of an on-vehicle acoustic device (car audio device) 84 (see FIG. 2) is arranged.
- Music (music) emitted from the speaker 76 can be noise when performing speech recognition.
- the vehicle body 46 is provided with an engine 80 for driving the vehicle.
- the sound emitted from the engine 80 can be noise when performing speech recognition.
- the noise generated in the passenger compartment 46 by the road surface stimulus during the traveling of the vehicle can also be a noise when performing voice recognition.
- wind noise generated when the vehicle travels can also be a noise source in performing speech recognition.
- the noise source 82 may exist outside the vehicle body 46. The sound emitted from the external noise source 82 can also be noise in performing speech recognition.
- the user's voice instruction is recognized using, for example, an automatic voice recognition device 68 (see FIG. 2).
- the speech processing apparatus contributes to improvement of speech recognition accuracy in the automatic speech recognition apparatus 68.
- FIG. 2 is a block diagram showing a system configuration of the speech processing apparatus according to the present embodiment.
- the speech processing apparatus includes a pre-processing unit 10, a processing unit 12, a post-processing unit 14, a speech source direction determination unit 16, an adaptive algorithm determination unit 18, and a noise model.
- a determination unit 20 and a designated input processing unit 86 are included.
- the voice processing device may further include an automatic voice recognition device 68, and the voice processing device according to the present embodiment and the automatic voice recognition device 68 may be separate devices.
- a device including these components and the automatic speech recognition device 68 can be referred to as a speech processing device or an automatic speech recognition device.
- a signal acquired by each of the plurality of microphones 22a to 22c, that is, a sound reception signal is input to the preprocessing unit 10.
- the microphone 22 for example, an omnidirectional microphone is used.
- FIG. 3A and 3B are schematic diagrams showing examples of microphone arrangement.
- FIG. 3A shows a case where the number of microphones 22 is three.
- FIG. 3B shows a case where the number of microphones 22 is two.
- the plurality of microphones 22 are arranged so as to be positioned on a straight line.
- the sound reaching the microphone 22 is handled as a plane wave, and the direction (direction) of the sound source 72, that is, the sound source direction (DOA: DirectionDirectOf Arrival) is determined. it can.
- DOA DirectionDirectOf Arrival
- the sound source 72 When the sound source 72 is located in the near field, it is preferable to determine the direction of the sound source 72 by treating the sound reaching the microphone 22 as a spherical wave.
- the distance L1 between the microphone 22a and the microphone 22b is set to be relatively long so as to be suitable for a relatively low frequency sound.
- the distance L2 between the microphone 22b and the microphone 22c is set to be relatively short so as to be suitable for a relatively high frequency sound.
- sound reception signals acquired by the plurality of microphones 22 are input to the preprocessing unit 10.
- sound field correction is performed.
- tuning is performed in consideration of the acoustic characteristics of the vehicle compartment 46 that is an acoustic space.
- the preprocessing unit 10 When the sound reception signal acquired by the microphone 22 includes music, the preprocessing unit 10 removes the music from the sound reception signal acquired by the microphone 22.
- a reference music signal (reference signal) is input to the preprocessing unit 10.
- the preprocessing unit 10 removes music included in the sound reception signal acquired by the microphone 22 using the reference music signal.
- the sound source direction determination unit 16 determines the direction of the sound source.
- the speed of sound is c [m / s]
- the distance between microphones is d [m]
- the arrival time difference is ⁇ [seconds]
- the direction ⁇ [degree] of the sound source 72 is expressed by the following equation (1). Represented by The sound speed c is about 340 [m / s].
- the output signal of the voice source direction determination unit 16, that is, the signal indicating the direction of the voice source 72 is input to the adaptive algorithm determination unit 18.
- the adaptive algorithm determination unit 18 determines an adaptive algorithm based on the orientation of the audio source 72.
- a signal indicating the adaptation algorithm determined by the adaptation algorithm determination unit 18 is input from the adaptation algorithm determination unit 18 to the processing unit 12.
- the processing unit 12 performs adaptive beamforming, which is signal processing that adaptively forms directivity (adaptive beamformer).
- the processing unit 12 not only functions as an adaptive beamformer that adaptively performs beamforming, but also controls the entire speech processing apparatus according to the present embodiment.
- the beam former for example, a Frost beam former or the like can be used.
- the beam forming is not limited to the Frost beamformer, and various beamformers can be applied as appropriate.
- the processing unit 12 performs beam forming based on the adaptive algorithm determined by the adaptive algorithm determination unit 18. In this embodiment, the beam forming is performed in order to reduce the sensitivity other than the arrival direction of the target sound while securing the sensitivity to the arrival direction of the target sound.
- the target sound is, for example, a sound emitted from the driver.
- the position of the sound source 72a can change.
- the arrival direction of the target sound changes according to the change in the position of the sound source 72a.
- the beam former is sequentially updated so as to suppress sound from an azimuth range other than the azimuth range including the azimuth.
- the voice source 72b to be subjected to voice recognition is located in the passenger seat 44, sound coming from an azimuth range other than the azimuth range including the azimuth of the passenger seat 44 is suppressed. Good.
- FIG. 4 is a diagram showing a beamformer algorithm.
- the received sound signals acquired by the microphones 22a to 22c are input to the window function / fast Fourier transform processing units 48a to 48c provided in the processing unit 12 via the preprocessing unit 10 (see FIG. 2). It is like that.
- the window function / fast Fourier transform processing units 48a to 48c perform window function processing and fast Fourier transform processing. In this embodiment, the window function process and the fast Fourier transform process are performed because the calculation in the frequency domain is faster than the calculation in the time domain.
- the output signal X1 , k of the window function / fast Fourier transform processing unit 48a and the beamformer weight tensor W1 , k * are multiplied at the multiplication point 50a.
- the output signal X2 , k of the window function / fast Fourier transform processor 48b and the beamformer weight tensor W2 , k * are multiplied at the multiplication point 50b.
- the output signal X 3, k of the window function / fast Fourier transform processing unit 48c and the beamformer weight tensor W 3, k * are multiplied at the multiplication point 50c.
- the signals multiplied at the multiplication points 50 a to 50 c are added at the addition point 52.
- the signal Y k added at the addition point 52 is input to an inverse fast Fourier transform / superimposition addition processing unit 54 provided in the processing unit 12.
- the inverse fast Fourier transform / superimposition addition processing unit 54 performs an inverse fast Fourier transform process and a process based on an overlay addition (OLA: OverLap-Add) method. By performing processing by the superposition addition method, the frequency domain signal is returned to the time domain signal. A signal subjected to the inverse fast Fourier transform process and the superposition addition method is input from the inverse fast Fourier transform / superimposition addition processing unit 54 to the post-processing unit 14.
- OVA OverLap-Add
- FIG. 5 is a diagram showing the directivity of the beamformer and the angle characteristics of the audio source direction determination cancellation process.
- the solid line indicates the directivity of the beamformer.
- the alternate long and short dash line indicates the angle characteristic of the audio source direction determination cancellation process.
- the output signal power becomes minimum at the azimuth angle ⁇ 1 degree and the azimuth angle ⁇ 2. It is sufficiently suppressed between the azimuth angle ⁇ 1 and the azimuth angle ⁇ 2. If a directional beamformer as shown in FIG. 5 is used, the sound arriving from the passenger seat can be sufficiently suppressed. On the other hand, the voice coming from the driver's seat reaches the microphone 22 with almost no suppression.
- the direction of the audio source 72 is determined. Suspend (voice source direction determination cancellation process). For example, when the beamformer is set to acquire the voice from the driver, if the voice from the passenger seat is larger than the voice from the driver, the direction of the voice source is estimated. Interrupt. In this case, the sound reception signal acquired by the microphone 22 is sufficiently suppressed. For example, when a voice arriving from a direction smaller than ⁇ 1 or a voice arriving from a direction larger than ⁇ 2, for example, is larger than the voice from the driver, a voice source direction determination canceling process is performed.
- the beamformer is set so as to acquire the voice from the driver has been described as an example, but the beamformer may be set so as to acquire the voice from the passenger. .
- the voice from the driver is louder than the voice from the passenger, the estimation of the direction of the voice source is interrupted.
- a signal in which sound coming from an azimuth range other than the azimuth range including the azimuth of the audio source 72 is suppressed is output from the processing unit 12.
- An output signal from the processing unit 12 is input to the post-processing unit 14.
- noise is removed.
- noise includes engine noise, road noise, wind noise, and the like.
- the engine noise model determination unit 20 generates a reference noise signal by performing noise modeling processing.
- the reference noise signal output from the noise model determination unit 20 is a reference signal for removing noise from a signal including noise.
- the reference engine noise signal is input to the post-processing unit 14.
- the post-processing unit 14 uses the reference engine noise signal to remove noise from the signal including noise.
- the post-processing unit 14 outputs a signal from which noise has been removed.
- the post-processing unit 14 also performs distortion reduction processing. Note that noise removal is not performed only in the post-processing unit 14. Noise is removed from a sound acquired via the microphone 22 by a series of processes performed in the preprocessing unit 10, the processing unit 12, and the postprocessing unit 14.
- a signal that has been post-processed by the post-processing unit 14 is output to the automatic speech recognition device 68. Since a good target sound in which sounds other than the target sound are suppressed is input to the automatic speech recognition device 68, the automatic speech recognition device 68 can improve the accuracy of speech recognition. Based on the voice recognition result by the automatic voice recognition device 68, the operation on the device mounted on the vehicle is automatically performed.
- the voice recognition result by the automatic voice recognition device 68 is also input to the designated input processing unit 86.
- the designation input processing unit 86 is for the user to designate a voice source 72 that is a target of voice recognition when a user (occupant) performs a predetermined action. Examples of the predetermined action include utterance of a predetermined word. A user who has issued a predetermined word is designated as the voice source 72 to be subjected to voice recognition.
- the sound source 72 designated by performing a predetermined action is referred to as a designated sound source.
- the designated input processing unit 68 determines whether or not a predetermined word has been issued based on the voice recognition result by the automatic voice recognition device 68.
- a signal indicating whether or not a predetermined word has been issued is input from the designated input processing unit 86 to the processing unit 12.
- the processing unit 12 performs beam forming so as to suppress sound coming from an azimuth range other than the azimuth range including the azimuth of the sound source 72 that issued the predetermined word. Note that the direction of the sound source 72 that has issued the predetermined word is determined by the sound source direction determination unit 16.
- FIG. 6 is a flowchart showing the operation of the speech processing apparatus according to the present embodiment.
- step S1 the sound processor is turned on (step S1).
- step S2 when the user has issued a predetermined word (YES in step S2), the audio source 72 that has issued the predetermined word is designated as the designated audio source (step S3). If the predetermined word is not issued (NO in step S2), step S2 is repeated.
- the designated voice source is a voice source 72 that is a target of voice recognition. Since the direction of the sound source 72 that has issued the predetermined word is determined by the sound source direction determination unit 16, it is possible to determine which seat the user has issued the predetermined word from. In this way, the sound source 72 that has issued the predetermined word is determined, and the designated sound source 72 to be subjected to speech recognition is designated.
- step S4 the orientation of the designated audio source 72 is determined (step S4).
- the direction of the designated audio source 72 is determined by the audio source direction determining unit 16.
- the directivity of the beamformer is set according to the direction of the designated audio source 72 (step S5).
- the setting of the beamformer directivity is performed by the adaptive algorithm determination unit 18, the processing unit 12, and the like as described above.
- step S5 When the volume of sound coming from an azimuth range other than the predetermined azimuth range including the azimuth of designated voice source 72 is equal to or greater than the magnitude of voice coming from designated voice source 72 (YES in step S5), the voice The determination of the source 72 is interrupted (step S7).
- step S4 when the magnitude of the sound coming from the azimuth range other than the predetermined azimuth range including the azimuth of the voice source 72 is not greater than the magnitude of the voice coming from the voice source 72 (NO in step S6), step S4 , S5 is repeated.
- the beamformer is adaptively set according to the change in the position of the designated sound source 72, and the sound other than the sound from the designated sound source 72, that is, the sound other than the target sound is surely suppressed.
- the present embodiment it is possible to reliably specify the voice source 72 to be subjected to voice recognition by issuing a predetermined word. For this reason, according to the present embodiment, it is possible to provide a good speech processing apparatus that can improve the certainty of speech recognition.
- FIG. 7 is a block diagram showing the system configuration of the speech processing apparatus according to the present embodiment.
- the same components as those of the speech processing apparatus according to the first embodiment shown in FIGS. 1 to 6 are denoted by the same reference numerals, and description thereof is omitted or simplified.
- the predetermined action for the user to specify the voice source 72 that is the target of voice recognition is an operation or gesture of the switches 90 and 92.
- the speech processing apparatus includes a pre-processing unit 10, a processing unit 12, a post-processing unit 14, a speech source direction determination unit 16, an adaptive algorithm determination unit 18, engine noise, and the like.
- a model determining unit 20 The speech processing apparatus according to the present embodiment also includes a learning processing unit 88, a driver seat side switch 90, a passenger seat side switch 92, a camera 94, a switch designation input processing unit 96, and an image designation input processing unit. 98.
- a driver's seat side switch 90 is arranged in the vicinity of the driver's seat 40.
- a passenger seat side switch 92 is disposed in the vicinity of the passenger seat 44.
- the driver seat side switch 90 and the passenger seat side switch 92 are connected to the switch designation input processing unit 96.
- the switch designation input processing unit 96 is for the user to designate the voice source 72 that is the target of voice recognition by the user operating the switches 90 and 92.
- the voice source 72a located in the driver's seat is designated as the designated voice source that is the target of voice recognition.
- the voice source 72b located in the passenger seat is designated as the designated voice source to be recognized.
- a signal indicating that the driver's seat side switch 90 has been operated is input from the switch designation input processing unit 96 to the processing unit 12.
- the processing unit 12 performs beam forming so as to suppress sound coming from an azimuth range other than the azimuth range including the azimuth of the sound source 72a located at the driver's seat 40. .
- a signal indicating that the passenger seat side switch 92 has been operated is input from the switch designation input processing unit 96 to the processing unit 12.
- the processing unit 12 performs beam forming so as to suppress sound coming from an azimuth range other than the azimuth range including the azimuth of the sound source 72b located at the passenger seat 44. .
- a camera 94 is disposed on the vehicle 46.
- An image acquired by the camera 94 is input to the image designation input processing unit 98.
- the image designation input processing unit 98 is for the user to designate a voice source 72 that is a target of voice recognition when a user (occupant) performs a predetermined action. Examples of the predetermined action include a predetermined gesture (gesture, pose).
- a user who has performed a predetermined gesture is designated as a voice source (designated voice source) 72 to be a target of voice recognition.
- the image designation input processing unit 98 determines whether a predetermined gesture has been performed based on the image acquired by the camera 94. A signal indicating whether or not a predetermined gesture has been performed is input from the image designation input processing unit 98 to the processing unit 12.
- the processing unit 12 performs beam forming so as to suppress sound coming from an azimuth range other than the azimuth range including the azimuth of the audio source 72a located at the driver's seat 40.
- the processing unit 12 performs beam forming so as to suppress sound coming from an azimuth range other than the azimuth range including the azimuth of the audio source 72b located at the passenger seat 44 when a predetermined gesture is performed by the passenger. I do.
- a learning processing unit 88 is connected to the processing unit 12.
- the learning processing unit 88 learns beam forming suitable for each of the sound sources 72a and 72b for each of the sound sources 72a and 72b.
- the learning processing unit 88 is provided for the following reason. That is, in the present embodiment, the predetermined action for the user to specify the voice source 72 that is the target of voice recognition is an operation or gesture of the switches 90 and 92. That is, in the present embodiment, the voice source 72 that is the target of voice recognition is designated by means other than voice. For this reason, when the sound source 72 to be subjected to speech recognition is designated, the sound from the designated sound source 72 is not necessarily obtained via the microphone 22.
- beam forming suitable for the designated sound source 72 is learned in advance, and the designated sound source 72 Preferably, beam forming suitable for 72 is applied.
- a learning processing unit 88 is provided.
- the learning processing unit 88 learns beam forming suitable for acquiring the sound from the sound source 72a when the sound is emitted from the sound source 72a.
- the learning processing unit 88 learns beamforming suitable for acquiring the sound from the sound source 72b when the sound is emitted from the sound source 72b.
- the beam forming learned as the beam forming suitable for the sound source 72a located in the driver's seat 40 is applied.
- the beam forming learned as the beam forming suitable for the sound source 72b located in the passenger seat 44 is applied.
- the signal that has been post-processed by the post-processing unit 14 is output as an audio output.
- FIG. 8 is a flowchart showing the operation of the speech processing apparatus according to the present embodiment.
- the sound processor is turned on (step S10).
- step S11 beam forming learning is performed (step S11).
- the learning processing unit 88 learns beamforming suitable for the sound source 72a located at the driver's seat 40.
- the learning processing unit 88 learns beamforming suitable for the sound source 72b located in the passenger seat 44.
- driver's seat side switch 90 When driver's seat side switch 90 is operated, specifically, when driver's seat side switch 90 is turned on (YES in step S12), beam forming suitable for audio source 72a located in driver's seat 40 is performed. The beam forming learned by the learning processing unit 88 is applied (step S13).
- step S14 If the driver's seat side switch 90 has not been operated (NO in step S12), it is confirmed whether or not the passenger's seat side switch 92 has been operated (step S14).
- the passenger seat side switch is operated, specifically, when the passenger seat side switch 92 is turned on (YES in step S14), beam forming suitable for the sound source 72b located in the passenger seat 44 is performed. Beam forming learned by the learning processing unit 88 is applied (step S15).
- step S16 If the passenger seat side switch 92 is not operated (NO in step S14), it is confirmed whether or not a predetermined gesture is performed by the driver (step S16).
- a predetermined gesture is performed by the driver (YES in step S16)
- the beam forming learned by the learning processing unit 88 is applied as the beam forming suitable for the sound source 72a located in the driver seat 40 ( Step S17).
- step S18 If the predetermined gesture is not performed by the driver (NO in step S16), it is confirmed whether or not the predetermined gesture is performed by the passenger seat (step S18).
- the beamforming learned by the learning processing unit 88 is applied as the beamforming suitable for the sound source 72b located in the passenger seat 44. (Step S19).
- step S21 when sound is emitted from the designated sound source 72, the direction of the designated sound source 72 is determined (step S21).
- the orientation of the designated audio source 72 is performed by the audio source orientation determining unit 16 as described above.
- the directivity of the beamformer is set according to the direction of the designated audio source 72 (step S22).
- the setting of the beamformer directivity is performed by the adaptive algorithm determination unit 18, the processing unit 12, and the like as described above.
- step S21 when the magnitude of sound coming from an azimuth range other than the predetermined azimuth range including the azimuth of voice source 72 is not greater than the magnitude of voice coming from voice source 72 (NO in step S23), step S21 , S22 is repeated.
- the beamformer is adaptively set according to the change in the position of the designated sound source 72, and the sound other than the sound from the designated sound source 72, that is, the sound other than the target sound is surely suppressed.
- the predetermined action for the user to specify the voice source 72 to be subjected to voice recognition may be an operation of the switches 90 and 92, a gesture, or the like.
- the case where the number of the microphones 22 is three has been described as an example, but the number of the microphones 22 is not limited to three, and may be four or more. If many microphones 22 are used, the direction of the sound source 72 can be determined with higher accuracy.
- the sound source 72 is located in the driver seat 40 or the passenger seat 44 .
- the position of the sound source 72 is not limited to the driver seat 40 or the passenger seat 44.
- the present invention is also applicable when the audio source 72 is located in the rear seat 70.
- a learning processing unit 88 may be further provided.
- the case where the output of the speech processing apparatus according to the present embodiment is input to the automatic speech recognition apparatus 68 that is, the case where the output of the speech processing apparatus according to the present embodiment is used for speech recognition will be described as an example.
- the present invention is not limited to this.
- the output of the speech processing apparatus according to the present embodiment may not be used for automatic speech recognition.
- the voice processing device according to the present embodiment may be applied to voice processing in a telephone conversation.
- the sound processing apparatus according to the present embodiment may be used to suppress sounds other than the target sound and transmit good sound. If the voice processing device according to the present embodiment is applied to telephone conversation, it is possible to realize a voice conversation.
- whether or not a predetermined gesture has been performed is determined based on an image acquired by the camera 94, but the present invention is not limited to this.
- a motion sensor or the like may be used to determine whether a predetermined gesture has been performed.
- the case where a plurality of microphones 22 are arranged linearly has been described as an example.
- the arrangement of three or more microphones 22 is not limited to this.
- the plurality of microphones 22 may be arranged on the same plane, or the plurality of microphones 22 may be arranged three-dimensionally.
Landscapes
- Engineering & Computer Science (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Multimedia (AREA)
- Acoustics & Sound (AREA)
- Computational Linguistics (AREA)
- Signal Processing (AREA)
- Theoretical Computer Science (AREA)
- Quality & Reliability (AREA)
- General Health & Medical Sciences (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Mechanical Engineering (AREA)
- Circuit For Audible Band Transducer (AREA)
- Fittings On The Vehicle Exterior For Carrying Loads, And Devices For Holding Or Mounting Articles (AREA)
Abstract
L'invention concerne un dispositif de traitement de la voix qui comprend : une pluralité de microphones (22) placés dans un véhicule ; une unité de détermination de direction de source vocale déterminant la direction d'une source vocale qui est la source d'une voix incluse dans un signal de réception de son acquis par chacun des microphones ; et une unité de traitement de formation de faisceau qui effectue une formation de faisceau pour supprimer les sons arrivant de plages de direction hors de la plage de direction comportant la direction de la source vocale. L'unité de traitement de formation de faisceau effectue une formation de faisceau dans la direction de la source vocale désignée par une action prédéfinie.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2014-263921 | 2014-12-26 | ||
| JP2014263921A JP2016126022A (ja) | 2014-12-26 | 2014-12-26 | 音声処理装置 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2016103710A1 true WO2016103710A1 (fr) | 2016-06-30 |
Family
ID=56149768
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2015/006448 Ceased WO2016103710A1 (fr) | 2014-12-26 | 2015-12-24 | Dispositif de traitement de la voix |
Country Status (2)
| Country | Link |
|---|---|
| JP (1) | JP2016126022A (fr) |
| WO (1) | WO2016103710A1 (fr) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108674344A (zh) * | 2018-03-30 | 2018-10-19 | 斑马网络技术有限公司 | 基于方向盘的语音处理系统及其应用 |
| CN112911465A (zh) * | 2021-02-01 | 2021-06-04 | 杭州海康威视数字技术股份有限公司 | 信号发送方法、装置及电子设备 |
Families Citing this family (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP6643720B2 (ja) | 2016-06-24 | 2020-02-12 | ミツミ電機株式会社 | レンズ駆動装置、カメラモジュール及びカメラ搭載装置 |
| JP6755843B2 (ja) | 2017-09-14 | 2020-09-16 | 株式会社東芝 | 音響処理装置、音声認識装置、音響処理方法、音声認識方法、音響処理プログラム及び音声認識プログラム |
| JP6872710B2 (ja) * | 2017-10-26 | 2021-05-19 | パナソニックIpマネジメント株式会社 | 指向性制御装置および指向性制御方法 |
| CN108597507A (zh) * | 2018-03-14 | 2018-09-28 | 百度在线网络技术(北京)有限公司 | 远场语音功能实现方法、设备、系统及存储介质 |
| JP7223561B2 (ja) * | 2018-03-29 | 2023-02-16 | パナソニックホールディングス株式会社 | 音声翻訳装置、音声翻訳方法及びそのプログラム |
| KR102208536B1 (ko) * | 2019-05-07 | 2021-01-27 | 서강대학교산학협력단 | 음성인식 장치 및 음성인식 장치의 동작방법 |
| JP6888851B1 (ja) * | 2020-04-06 | 2021-06-16 | 山内 和博 | 自動運転車 |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2001296891A (ja) * | 2000-04-14 | 2001-10-26 | Mitsubishi Electric Corp | 音声認識方法および装置 |
| JP2004109361A (ja) * | 2002-09-17 | 2004-04-08 | Toshiba Corp | 指向性設定装置、指向性設定方法及び指向性設定プログラム |
| JP2014153663A (ja) * | 2013-02-13 | 2014-08-25 | Sony Corp | 音声認識装置、および音声認識方法、並びにプログラム |
| JP2014203031A (ja) * | 2013-04-09 | 2014-10-27 | 小島プレス工業株式会社 | 音声認識制御装置 |
Family Cites Families (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP3484112B2 (ja) * | 1999-09-27 | 2004-01-06 | 株式会社東芝 | 雑音成分抑圧処理装置および雑音成分抑圧処理方法 |
| JP4097219B2 (ja) * | 2004-10-25 | 2008-06-11 | 本田技研工業株式会社 | 音声認識装置及びその搭載車両 |
| KR100959983B1 (ko) * | 2005-08-11 | 2010-05-27 | 아사히 가세이 가부시키가이샤 | 음원 분리 장치, 음성 인식 장치, 휴대 전화기, 음원 분리방법, 및, 프로그램 |
| GB0906269D0 (en) * | 2009-04-09 | 2009-05-20 | Ntnu Technology Transfer As | Optimal modal beamformer for sensor arrays |
| JP5962038B2 (ja) * | 2012-02-03 | 2016-08-03 | ソニー株式会社 | 信号処理装置、信号処理方法、プログラム、信号処理システムおよび通信端末 |
-
2014
- 2014-12-26 JP JP2014263921A patent/JP2016126022A/ja active Pending
-
2015
- 2015-12-24 WO PCT/JP2015/006448 patent/WO2016103710A1/fr not_active Ceased
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2001296891A (ja) * | 2000-04-14 | 2001-10-26 | Mitsubishi Electric Corp | 音声認識方法および装置 |
| JP2004109361A (ja) * | 2002-09-17 | 2004-04-08 | Toshiba Corp | 指向性設定装置、指向性設定方法及び指向性設定プログラム |
| JP2014153663A (ja) * | 2013-02-13 | 2014-08-25 | Sony Corp | 音声認識装置、および音声認識方法、並びにプログラム |
| JP2014203031A (ja) * | 2013-04-09 | 2014-10-27 | 小島プレス工業株式会社 | 音声認識制御装置 |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN108674344A (zh) * | 2018-03-30 | 2018-10-19 | 斑马网络技术有限公司 | 基于方向盘的语音处理系统及其应用 |
| CN108674344B (zh) * | 2018-03-30 | 2024-04-02 | 斑马网络技术有限公司 | 基于方向盘的语音处理系统及其应用 |
| CN112911465A (zh) * | 2021-02-01 | 2021-06-04 | 杭州海康威视数字技术股份有限公司 | 信号发送方法、装置及电子设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| JP2016126022A (ja) | 2016-07-11 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| WO2016103709A1 (fr) | Dispositif de traitement vocal | |
| JP2016126022A (ja) | 音声処理装置 | |
| WO2016143340A1 (fr) | Dispositif de traitement de la parole, et dispositif de commande | |
| CN110691299B (zh) | 音频处理系统、方法、装置、设备及存储介质 | |
| JP5913340B2 (ja) | マルチビーム音響システム | |
| JP4779748B2 (ja) | 車両用音声入出力装置および音声入出力装置用プログラム | |
| CN105592384B (zh) | 用于控制车内噪声的系统和方法 | |
| US9953641B2 (en) | Speech collector in car cabin | |
| CN111489750B (zh) | 声音处理设备和声音处理方法 | |
| CN105810203B (zh) | 消除噪声的设备和方法、声音识别设备和配备其的车辆 | |
| JP2004109361A (ja) | 指向性設定装置、指向性設定方法及び指向性設定プログラム | |
| CN110120217A (zh) | 一种音频数据处理方法及装置 | |
| JP7692069B2 (ja) | 信号処理装置及び信号処理方法 | |
| US12039965B2 (en) | Audio processing system and audio processing device | |
| JP2002351488A (ja) | ノイズキャンセル装置および車載システム | |
| JP2009073417A (ja) | 騒音制御装置および方法 | |
| JP2007180896A (ja) | 音声信号処理装置および音声信号処理方法 | |
| GB2560498A (en) | System and method for noise cancellation | |
| JP6606921B2 (ja) | 発声方向特定装置 | |
| JP2016149014A (ja) | 対話装置 | |
| JP4508147B2 (ja) | 車載ハンズフリー装置 | |
| WO2016135964A1 (fr) | Dispositif de commande de volume sonore, procédé de commande de volume sonore, et programme de commande de volume sonore |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 15872281 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 15872281 Country of ref document: EP Kind code of ref document: A1 |