RS49859B - SYSTEM AND PROCEDURE FOR LOCATION OF SPEAKERS BY MICROPHONE - Google Patents

SYSTEM AND PROCEDURE FOR LOCATION OF SPEAKERS BY MICROPHONE

Info

Publication number
RS49859B
RS49859B RSP-2006/0642A RSP20060642A RS49859B RS 49859 B RS49859 B RS 49859B RS P20060642 A RSP20060642 A RS P20060642A RS 49859 B RS49859 B RS 49859B
Authority
RS
Serbia
Prior art keywords
fact
cross
speaker
correlation
block
Prior art date
Application number
RSP-2006/0642A
Other languages
Serbian (sr)
Inventor
dr. Zoran Šarić
dr. Slobodan Jovičić
dr. Vladimir Kovačević
dr. Nikola Teslić
dr. Dragan Kukolj
Original Assignee
Micronasnit,
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Micronasnit, filed Critical Micronasnit,
Priority to RSP-2006/0642A priority Critical patent/RS49859B/en
Publication of RS20060642A publication Critical patent/RS20060642A/en
Publication of RS49859B publication Critical patent/RS49859B/en

Links

Landscapes

  • Circuit For Audible Band Transducer (AREA)
  • Telephonic Communication Services (AREA)
  • Obtaining Desirable Characteristics In Audible-Bandwidth Transducers (AREA)

Abstract

Sistem za lociranje govornika pomoću mikrofonskog niza karakterisan time, što sadrži: mikrofonski niz od M mikrofona u odnosu na čiju simetralu se određuje ugao azimuta, odnosno položaj govornika u horizontalnoj ravni; blok za predprocesiranje mikrofonskih signala i konverziju u digitalnu formu i frekvencijskidomen; blok za kroskorelacionu analizu mikrofonskih signala i njenu optimizaciju na bazi fazne transformacije; blok za određivanje filterske funkcije na bazi prozodijskih karakteristika govornog signala, pomoću koje se vrši optimizacija kroskorelacione PHAT analize; blok za detekciju aktivnosti govora (VAD) zasnovan na superdirektivnom usmerivaču (SD-BF) koji obezbeđuje prostorno filtriranje govornika; blok za estimaciju ugla azimuta na bazi maksimuma interpoliranih kroskorelacionih funkcija.A system for locating a speaker by means of a microphone array, characterized in that it comprises: a microphone array of M microphones with respect to whose symmetry determines the azimuth angle or position of the speaker in a horizontal plane; block for preprocessing of microphone signals and conversion to digital form and frequency domain; block for cross-correlation analysis of microphone signals and its optimization based on phase transformation; a block for determining a filter function based on the prosodic characteristics of the speech signal, which optimizes the cross-correlation PHAT analysis; a speech-based activity detection unit (VAD) based on a superdirectional router (SD-BF) that provides spatial filtering for speakers; block for azimuth angle estimation based on maximum of interpolated cross-correlation functions.

Description

OBLAST TEHNIKE NA KOJU SE PRONALAZAK ODNOSITECHNICAL FIELD TO WHICH THE INVENTION RELETS

Pronalazak pripada oblasti obrade akustičkog signala, ili preciznije, metodama lociranja govornika primenom mikrofonskog niza u akustičkom ambijentu sa prisutnim šumom i reverberacijom. The invention belongs to the field of acoustic signal processing, or more precisely, methods of locating speakers using a microphone array in an acoustic environment with noise and reverberation.

TEHNIČKI PROBLEMTECHNICAL PROBLEM

Lokalizacija govornika u prostoru je veoma važan tehnički problem u sistemima koji se baziraju na govornoj komunikaciji na relaciji čovek-čovek ili čovek-mašina. On nastaje kao potreba da komunikacija bude što razumljivija uprkos mnogim smetnjama koje se mogu pojaviti u prostoru a koje maskiraju govorni signal. U sistemima kao što su telekonferencijski sistemi ili spikerfoni u prostoriji ili kolima, pored razumljivosti je od primarne važnosti i kvalitet komunikacije. Isti atributi govorne komunikacije su važni i u komunikaciji na primer čoveka i robota, gde robot mora tačno da prepozna govornu komandu. Nešto drugačiji problem se pojavljuje kod upravljanja video kamere, gde kamera treba da se usmeri ka aktuelnom govorniku apstrahujući ostale izvore zvuka u datom ambijentu. Dakle, ne postavlja se problem razumljivosti govora, već separacije govornog signala i ostalih akustičkih signala. Localization of speakers in space is a very important technical problem in systems based on human-human or human-machine voice communication. It arises as a need for communication to be as intelligible as possible despite the many disturbances that may appear in space and that mask the speech signal. In systems such as teleconference systems or speakerphones in a room or in a car, in addition to intelligibility, the quality of communication is also of primary importance. The same attributes of speech communication are important in communication, for example, between a human and a robot, where the robot must correctly recognize a spoken command. A slightly different problem appears with the control of the video camera, where the camera should be directed towards the current speaker, abstracting other sound sources in the given environment. Therefore, the problem is not the intelligibility of speech, but the separation of the speech signal and other acoustic signals.

Navedeni primeri govorne komunikacije u prostoru, ili prostoriji, definišu osnovni problem u vidu lokalizacije aktuelnog govornika, odnosno usmeravanje mikrofonskog sistema ka njemu. Mikrofonski sistem može biti usmereni mikrofon ili više mikrofona u odgovarajućem fizičkom rasporedu. Pošto u akustičkom ambijentu pored izvora korisnog signala postoje smetnje veoma različitog porekla, čiji izvori u prostoru mogu biti proizvoljno raspoređeni, mikrofonski sistem mora imati usmerenu karakteristiku osetljivosti i mora se usmeriti ka željenom izvoru signala, tj. aktuelnom govorniku. Drugačije rečeno, mikrofonski sistem mora locirati govornika u horizontalnoj ravni i odrediti ugao azimuta u odnosu na svoje koordinate u prostoru. The given examples of speech communication in a space, or room, define the basic problem in the form of localization of the current speaker, that is, directing the microphone system towards him. A microphone system can be a directional microphone or multiple microphones in a suitable physical arrangement. Since in the acoustic environment, next to the source of the useful signal, there are disturbances of very different origins, whose sources can be arbitrarily distributed in space, the microphone system must have a directional sensitivity characteristic and must be directed towards the desired signal source, i.e. current speaker. In other words, the microphone system must locate the speaker in the horizontal plane and determine the azimuth angle relative to its coordinates in space.

Tehnički problem nastaje kada se u ambijentu pojavi veći broj izvora smetnji, kada su ove smetnje nestacionarne, kada se izvori smetnji kreću u prostoru ili kada se aktuelni govornik kreće u prostoru. Postupak određivanja ugla azimuta govornika mora da reši tri osnovna problema: (1) detekciju govorne aktivnosti aktuelnog govornika, pri tome treba imati u vidu da se u posmatranom prostoru može pojaviti veći broj govornika, (2) separaciju aktuelnog govornika u odnosu na sve ostale izvore smetnji, što podrazumeva potiskivanje signala smetnji a isticanje korisnog signala, i (3) adaptivno praćenje aktuelnog govornika u pokretu, pri čemu se mora uzeti u obzir da i ostali izvori zvuka mogu biti pokretni. Ovaj treći problem se tiče pravilnog usmeravanja sistema (robota, kamere) ka aktuelnom govorniku. A technical problem arises when a large number of interference sources appear in the environment, when these interferences are non-stationary, when the interference sources move in space or when the current speaker moves in space. The procedure for determining the azimuth angle of the speaker must solve three basic problems: (1) detection of the speaking activity of the current speaker, while keeping in mind that a larger number of speakers may appear in the observed space, (2) separation of the current speaker in relation to all other sources of interference, which implies suppression of the interference signal and emphasis of the useful signal, and (3) adaptive tracking of the current speaker in motion, where it must be taken into account that other sound sources may also be moving. This third problem concerns the correct orientation of the system (robot, camera) towards the current speaker.

Dodatni problemi se pojavljuju kod lokalizacije govornika u prostoriji sa izraženom reverberacijom. Signali refleksija od zidova ili objekata u prostoriji mogu biti, u zavisnosti od položaja govornika, izvora smetnji i mikrofonskog sistema, znatno jači od direktnog zvučnog talasa aktuelnog govornika. Additional problems arise with speaker localization in a room with pronounced reverberation. The signals of reflections from the walls or objects in the room can be, depending on the position of the speaker, the source of interference and the microphone system, significantly stronger than the direct sound wave of the current speaker.

Iz izloženog se vidi da su tehnički problemi u rešenju lokalizacije govornika u prostoru veoma složeni i da zahtevaju kompleksan pristup u optimizaciji rešenja, posebno kada se ima u vidu rad sistema u realnom vremenu na bazi komercijalne platforme digitalnog procesora signala (DSP). It can be seen from the above that the technical problems in the solution of speaker localization in space are very complex and require a complex approach in the optimization of the solution, especially when considering the operation of the system in real time based on a commercial digital signal processor (DSP) platform.

STANJE TEHNIKESTATE OF THE ART

Lociranje govornika u uslovima prisustva akustičkih smetnji i reverberacije prostorije predstavlja složen problem. U uslovima kada se spektri korisnog govornog signala preklapaju sa spektrima prisutnih smetnji lociranje govornika se može rešiti na pouzdan način primenom mikrofonskog niza i odgovarajuće obrade signala koja uzima u obzir specifičnosti uslova primene sistema za lociranje. Teorijske osnove u primeni mikrofonskih nizova za lokalizaciju govornika date su u M.S. Brandstein, D.B. Ward (Eds.),Microphone Arrays: Signal Processing Techniques and Applications,Springer, Berlin 2001; i u Y. Huang, J. Benestv,Audio signal processing for next generation multimedia communication systems,Kluwer Academic Publishers Publ., 2004. Locating speakers in the presence of acoustic disturbances and room reverberation is a complex problem. In conditions when the spectrum of the useful speech signal overlaps with the spectrum of the present disturbances, locating the speaker can be solved in a reliable way by using a microphone array and appropriate signal processing that takes into account the specifics of the application conditions of the locating system. Theoretical foundations in the application of microphone arrays for speaker localization are given in M.S. Brandstein, D.B. Ward (Eds.), Microphone Arrays: Signal Processing Techniques and Applications, Springer, Berlin 2001; and in Y. Huang, J. Benestv, Audio signal processing for next generation multimedia communication systems, Kluwer Academic Publishers Publ., 2004.

Postoji veliki broj patentiranih rešenja na bazi mikrofonskih nizova kao što su na primer: U.S. objavljena patentna prijava 2003/0051532 Al, prijavljena 15. avgusta 2002., sa naslovom „Robust talker localization in reverberant environment", daje rešenje lokalizacije govornika u reverberantnoj prostoriji na bazi mikrofonskog sistema cirkularne konfiguracije i na bazi energetske detekcije direktnog zvučnog talasa; zatim U.S. patent 6,970,796 B2, prijavljen 1. marta 2004., sa naslovom "System and method for improving the precision of localization estimates", daje rešenje koje, pored konvencijalnog određivanja DO A sa mikrofonskim nizom, ima sistem za post-procesiranje na bazi statističkog klasterovanja inicijalnih estimacija lokacija i dobijanja finalne estimacije lokacije sa povećanom preciznošću i pouzdanošću; zatim U.S. patent 6,999,593 B2, prijavljena 28. maja 2003., sa naslovom "Svstem and process for robust sound source localization", daje rešenje za lokalizaciju govornika na bazi mikrofonskog niza kombinovanjem težinske kros-korelacije i podešene usmerene karakteristike parova mikrofonskog niza; zatim U.S. objavljena patentna prijava 2005/0080619 Al, prijavljena 13. oktobra 2004, sa naslovom „Method and apparatus for robust speaker localization and automatic camera steering svstem emploving the same", daje rešenje koje lokalizaciju govornika određuje pomoću MUSIC tehnologije. There are a number of patented solutions based on microphone arrays such as: U.S. published patent application 2003/0051532 Al, filed on August 15, 2002, with the title "Robust talker localization in reverberant environment", provides a solution for speaker localization in a reverberant room based on a circular configuration microphone system and on the basis of direct sound wave energy detection; then the U.S. patent 6,970,796 B2, filed on March 1, 2004, with the title "System and method for improving the precision of localization estimates", provides a solution that, in addition to conventional DO A determination with a microphone array, has a post-processing system based on statistical clustering of initial location estimates and obtaining a final location estimate with increased precision and reliability; then the U.S. patent 6,999,593 B2, filed May 28, 2003, entitled "System and process for robust sound source localization", provides a solution for microphone array-based speaker localization by combining weighted cross-correlation and tuned directional characteristics of microphone array pairs; then the U.S. published patent application 2005/0080619 Al, filed on October 13, 2004, entitled "Method and apparatus for robust speaker localization and automatic camera steering system employing the same", provides a solution that determines speaker localization using MUSIC technology.

Generalno, metode estimacije pravca lociranja govornika (izvora zvuka) se mogu podeliti u tri osnovne grupe: metode na bazi superdirektivne karakteristike usmerenosti mikrofonskog niza, metode na bazi kompleksne estimacije spektra mikrofonskih signala i metode na bazi vremenskog kašnjenja zvučnih talasa do mikrofonskog niza TDOA{ Time Delay of Arrival).Metode iz prve i druge grupe su osetljive na frekvencijske karakteristike svih izvora zvuka u analiziranoj prostoriji i bez a priornog znanja ne pružaju zadovoljavajuću tačnost. U praksi se najčešće koriste TDOA metode (P. Julian et al., A comparative study of sound localization algorithms for energy aware sensor networknodes,ZE££ Trans. Circuits and Systems,Vol. 51, No. 4, pp. 640-648, Apr. 2004.). U konvencionalnom postupku u prvom koraku vrši se estimacija TDOA za svaki par mikrofona u mikrofonskom nizu. Estimacija se zasniva na kroskorelacionoj analizi koja se u drugom koraku ponderiše težinskom funkcijom PHAT( Phase Trans/ orni),koja povećava robusnost algoritma procene dolaznih pravaca na prisustvo šuma i reverberacije u prostoriji. Međutim, reverberacija prostorije, prisustvo više izvora zvuka i šuma predstavljaju i dalje veliki problem za ove metode lokalizacije govornika. In general, the methods of estimating the direction of locating the speaker (sound source) can be divided into three basic groups: methods based on the superdirective characteristic of the directionality of the microphone array, methods based on the complex estimation of the spectrum of microphone signals and methods based on the time delay of sound waves to the microphone array TDOA{Time Delay of Arrival).The methods from the first and second groups are sensitive to the frequency characteristics of all sound sources in the analyzed room and without prior knowledge do not provide satisfactory accuracy. TDOA methods are most often used in practice (P. Julian et al., A comparative study of sound localization algorithms for energy aware sensor network nodes, ZE££ Trans. Circuits and Systems, Vol. 51, No. 4, pp. 640-648, Apr. 2004.). In the conventional procedure, in the first step TDOA is estimated for each pair of microphones in the microphone array. The estimation is based on cross-correlation analysis, which in the second step is weighted by the PHAT (Phase Trans/orni) weighting function, which increases the robustness of the algorithm for estimating incoming directions for the presence of noise and reverberation in the room. However, room reverberation, the presence of multiple sound sources and noise are still a major problem for these speaker localization methods.

IZLAGANJE SUŠTINE PRONALASKADISCLOSURE OF THE ESSENCE OF THE INVENTION

Predmet ovog pronalaska je sistem i postupak za lociranje govornika pomoću mikrofonskog niza u složenom akustičkom ambijentu koji pored aktuelnog govornika sadrži mnoge signale smetnji kao što su: ambijentalna buka, izvori akustičkih smetnji, reverberacija prostorije i drugi govornici. Kao takav, sistem može naći široku primenu u sistemima za govornu komunikaciju kao u sistemima za kontrolu i upravljanje putem glasa. The subject of this invention is a system and procedure for locating a speaker using a microphone array in a complex acoustic environment which, in addition to the actual speaker, contains many interference signals such as: ambient noise, sources of acoustic interference, room reverberation and other speakers. As such, the system can find wide application in speech communication systems as well as in voice control and management systems.

Sistem, koji je predmet pronalaska, sadrži M mikrofona raspoređenih u linijskoj strukturi i na jednakim rastojanjima, blok za predobradu i digitalizaciju mikrofonskih signala i blok za digitalnu obradu mikrofonskih signala. Sistem se postavlja u horizontalnu ravan i određuje ugao azimuta govornika u odnosu na simetralu sistema. Sistem može biti povezan na konferencijski sistem, ili biti deo njega, ili na sistem upravljanja ili kontrole, kao što su robot ili video kamera. The system, which is the subject of the invention, contains M microphones arranged in a linear structure and at equal distances, a block for preprocessing and digitization of microphone signals and a block for digital processing of microphone signals. The system is placed in the horizontal plane and determines the azimuth angle of the speaker in relation to the bisector of the system. The system can be connected to, or be part of, a conferencing system, or a management or control system, such as a robot or video camera.

Suština pronalaska jeste u specifičnoj obradi govornog signala koji se snima u akustičkom ambijentu prostorije u kojoj se nalazi sistem i govornik. Mikrofonski niz snima sve signale u prostoriji: koristan signal kao direktan talas koji stiže od govornika do mikrofona i signale smetnji koji mogu biti raznovrsni. Kao signali smetnje pojavljuju se direktni talasi od jednog ili više izvora šumova ili izvora drugih smetnji koji se mogu naći u prostoriji i svi reflektovani talasi (eho prostorije) koji potiču od svih izvora zvukova, uključujući i aktuelnog govornika, a koji nastaju usled reverberacije prostorije. Treba naglasiti da izvori zvukova u prostoriji mogu biti stacionarni ili nestacionarni, što je najčešći slučaj, kako po svojim karakteristikama tako i po lokaciji u prostoriji (pokretni izvori zvukova). The essence of the invention lies in the specific processing of the speech signal that is recorded in the acoustic environment of the room where the system and the speaker are located. The microphone array records all the signals in the room: the useful signal as a direct wave arriving from the speaker to the microphone and interference signals that can be varied. Disturbance signals include direct waves from one or more noise sources or sources of other disturbances that can be found in the room and all reflected waves (room echoes) originating from all sound sources, including the actual speaker, which are caused by room reverberation. It should be emphasized that the sources of sounds in the room can be stationary or non-stationary, which is the most common case, both according to their characteristics and their location in the room (moving sources of sounds).

Mikrofonski signali iz mikrofonskog niza se obrađuju u digitalnoj formi u frekvencijskom domenu. Ovaj domen omogućava određene prednosti u pogledu brzine obrade i broja računskih operacija, što je veoma važno za realizaciju sistema u realnom vremenu. Microphone signals from the microphone array are processed in digital form in the frequency domain. This domain provides certain advantages in terms of processing speed and the number of computational operations, which is very important for the realization of real-time systems.

Specifičan aspekt pronalaska se nalazi u optimizaciji kroskorelacione analize mikrofonskih signala kroz dva aspekta: prvo, generalizacijom kroskorelacije koja se u literaturi označava kao fazna transformacija PHAT( Phase Transform),a koja podrazumeva normalizaciju kroskorelacije na svoj moduo kada se gubi informacija o snazi signala, a ostaje samo informacija o fazi u kojoj je sadržano relativno vremensko kašnjenje signala i drugo, ponderisanjem PHAT transformacije filterskom funkcijomW( n)koja sadrži osnovne prozodijske karakteristike govornog signala, pre svega energetsku dinamiku formantnih struktura vokala. A specific aspect of the invention is in the optimization of the cross-correlation analysis of microphone signals through two aspects: first, by generalizing the cross-correlation, which is referred to in the literature as phase transformation PHAT (Phase Transform), which implies the normalization of the cross-correlation to its modulus when the information about the signal strength is lost, and only the information about the phase that contains the relative time delay of the signal remains, and second, by weighting the PHAT transformation with a filter function W(n) that contains the basic prosodic characteristics of the speech signal, before above all, the energetic dynamics of the vocal formant structures.

Sledeća specifičnost pronalaska jeste određivanje filterske funkcijeW( n)na bazi analize mikrofonskih signala u tri domena: energetskom, frekvencijskom i vremenskom. Cilj ove analize je da se generalizovana kroskorelaciona analiza odvija pod kontrolom prozodijskih karakteristika govornog signala, što predstavlja na specifičan način separaciju govornog signala u odnosu na ostale signale ambijentalnih smetnji i reverberaciju, i što u krajnjem slučaju daje pouzdaniju estimaciju lokacije govornika. The next specificity of the invention is the determination of the filter function W(n) based on the analysis of microphone signals in three domains: energy, frequency and time. The goal of this analysis is that the generalized cross-correlation analysis takes place under the control of the prosodic characteristics of the speech signal, which represents in a specific way the separation of the speech signal in relation to other signals of ambient disturbances and reverberation, and which ultimately provides a more reliable estimate of the speaker's location.

Specifičnost pronalaska jeste i realizacija detektora aktivnosti govora (VAD) na bazi superdirektivnog usmerivača (SD-BF), koji obezbeđuje veći indeks usmerenosti mikrofonskog niza i time efikasnije prostorno filtriranje govornog signala u odnosu na ambijentalne smetnje. The specificity of the invention is the implementation of a speech activity detector (VAD) based on a superdirective router (SD-BF), which provides a higher directivity index of the microphone array and thus more efficient spatial filtering of the speech signal in relation to ambient interference.

Inventivnost u ovom pronalasku se nalazi u načinu realizacije svake od navedenih specifičnosti, ali i u postupku integrisanja svih algoritama u jedinstvenu celinu koja funkcioniše stabilno i kvalitetno. Algoritamske procedure su optimizirane korišćenjem zajedničkih resursa, posebno ako se ima u vidu realizacija u spektralnom i multidimenzionalnom domenu (multimikrofonski sistem). Inventiveness in this invention is found in the way of realization of each of the mentioned specificities, but also in the process of integrating all algorithms into a unique unit that functions stably and with quality. Algorithmic procedures are optimized by using common resources, especially if one considers the realization in the spectral and multidimensional domain (multimicrophone system).

Ovi i drugi aspekti, specifičnosti i benefiti ovog pronalaska biće očigledniji nakon uvida u detaljan opis pronalaska, patentne zahteve i pripadajuće crteže. These and other aspects, specificities and benefits of the present invention will be more apparent upon review of the detailed description of the invention, patent claims and accompanying drawings.

KRATAK OPIS SLIKA I NACRTABRIEF DESCRIPTION OF THE IMAGES AND DRAWINGS

Slika 1- prikazuje ambijentalne uslove primene sistema za lociranje govornika pomoću mikrofonskog niza. Figure 1 - shows the ambient conditions of application of the system for locating speakers using a microphone array.

Slika 2- prikazuje osnovni blok dijagram sistema za lociranje govornika. Figure 2- shows the basic block diagram of the system for locating speakers.

Slika3 - prikazuje blok dijagram podsistema za estimaciju filterske funkcije Figure 3 - shows the block diagram of the filter function estimation subsystem

W( n).W(n).

Slika 4- prikazuje blok dijagram podsistema za estimaciju ugla azimuta0.Figure 4- shows the block diagram of the azimuth angle estimation subsystem0.

Slika 5- prikazuje blok dijagram podsistema VAD za detekciju aktivnosti govora aktuelnog govornika. Figure 5- shows a block diagram of the VAD subsystem for detecting the speech activity of the current speaker.

DETALJAN OPIS PRONALASKADETAILED DESCRIPTION OF THE INVENTION

Ovaj pronalazak opisuje sistem i postupak za lokalizaciju govornika pomoću mikrofonskog niza u akustičkom ambijentu kakav je prostorija, sa prisutnim stacionarnim i/ili nestacionarnim smetnjama. This invention describes a system and method for localizing a speaker using a microphone array in an acoustic environment such as a room, with stationary and/or non-stationary disturbances present.

Slika 1 prikazuje ambijentalne uslove u kojima se sistem, koji je predmet ovog pronalaska, može naći. Naime, u prostoriji100nalazi se aktuelni govornik101u horizontalnoj ravni na pravcu102pod uglom9u odnosu na simetralu mikrofonskog niza103.Mikrofonski niz sadrži M mikrofona koji snimljene signale prosleđuju u blok104gde se vrši obrada signala u cilju određivanja estimacije ugla azimuta6.Informacija o estimiranom uglu azimuta može da se koristi, na primer za kontrolu robota105ili video kamere 106, ili za komunikacione potrebe kao što su govorna komunikacija preko intemeta107ili preko telekonferencijskog sistema108.U drugom slučaju uglom azimuta upravlja se karakteristikom usmerenosti mikrofonskog niza 103, koja se usmerava prema aktuelnom govorniku. Figure 1 shows the ambient conditions in which the system, which is the subject of the present invention, can be found. Namely, in the room 100 there is the current speaker 101 in the horizontal plane in the direction 102 at an angle 9 in relation to the bisector of the microphone array 103. The microphone array contains M microphones that transmit the recorded signals to the block 104 where the signal is processed in order to determine the estimation of the azimuth angle 6. Information about the estimated azimuth angle can be used, for example, to control a robot 105 or a video camera 106, or for communication needs such as voice communication over the Internet107 or through a teleconference system108. In the second case, the azimuth angle is controlled by the directionality characteristic of the microphone array 103, which is directed towards the current speaker.

Osnovni problem u estimaciji ugla azimuta6čine smetnje u prostoriji koje direktno utiču na tačnost i preciznost estimacije. Osnovni izvor smetnje može biti izvor šuma, govora, muzike, itd., 109, sa direktnim zvučnim talasom110,ali i reflektovanim zvučnim talasima o zidove prostorije, kao što je talas110a.Naravno, i aktuelni govornik jeste izvor reflektovanih talasa,102ai102b,koji predstavljaju smetnju. Ako se mikrofonski niz103koristi za komunikacione potrebe, slučajevi 107 i108,kod tzv. „hands-free" komunikacija, tada se pojavljuje veoma ozbiljna smetnja u vidu akustičkog eha111,koja ima svoje akustičke refleksijelila.The main problem in the estimation of the azimuth angle is the interference in the room that directly affects the accuracy and precision of the estimation. The main source of disturbance can be a source of noise, speech, music, etc., 109, with a direct sound wave110, but also reflected sound waves on the walls of the room, such as wave 110a. If the microphone array 103 is used for communication purposes, cases 107 and 108, in the so-called "hands-free" communication, then a very serious disturbance appears in the form of an acoustic echo111, which has its own acoustic reflection elements.

Prema tome, tačnost i preciznost određivanja ugla azimuta u velikoj meri zavisi od ambijentalnih uslova u kojima se sistem, koji je predmet ovog patenta, koristi. Dodatni problem se pojavljuje ukoliko se aktuelni govornik ili izvori smetnji kreću u prostoriji, čime se postavlja zahtev adaptivnog praćenja pozicije aktuelnog govornika. Therefore, the accuracy and precision of determining the azimuth angle largely depends on the ambient conditions in which the system, which is the subject of this patent, is used. An additional problem arises if the current speaker or the sources of interference move in the room, which requires adaptive tracking of the position of the current speaker.

Na slici 2 prikazana je blok šema sistema za lokalizaciju govornika pomoću mikrofonskog niza. Signali iz mikrofonaxidoxmmikrofonskog niza103ulaze u blok201u kome se vrši njihova predobrada, odnosno pojačanje, filtriranje, digitalizacija i konverzija u frekvencijski domen pomoću diskretne Fourierove transformacije (DFT). Predobrada se vrši na nivou segmenata dužine N odmeraka, sa preklapanjem 50% i sa primenjenim Hammingovim prozorom i FFT reda N. Figure 2 shows a block diagram of a speaker localization system using a microphone array. Signals from the microphone and the microphone array 103 enter the block 201 in which their pre-processing is performed, i.e. amplification, filtering, digitization and conversion into the frequency domain using discrete Fourier transformation (DFT). Preprocessing is performed at the level of segments of length N measurements, with an overlap of 50% and with an applied Hamming window and FFT of order N.

Izlaz bloka201jesu Fourierove transformacijeX/doXM.Na ovim signalima vrši se kroskorelaciona analiza prvog mikrofona sa svim ostalim mikrofonima. Na izlazu bloka201dobijaju se estimacije kroskorelacije između signal provog i svih ostalih mikrofona,G\ j( n)doGi,m(")rekurzivnim usrednjavanjem prema relaciji: The output of block 201 are Fourier transformations X/doXM. Cross-correlation analysis of the first microphone with all other microphones is performed on these signals. At the output of block 201, estimates of the cross-correlation between the signal and all other microphones are obtained, G\ j( n)doGi,m(") by recursive averaging according to the relation:

Konstante a+ i a. se biraju tako da ispunjavaju nejednakost 0.5 < a+ < a. < 1 i pod tim uslovom favorizuje se uticaj članovaXi( t, f) Xk'( t, f)sa većim modulom. SignaliGu(/i)doGiM( n)ulaze u blokove 203 i205.Constants a+ and a. are chosen to satisfy the inequality 0.5 < a+ < a. < 1 and under that condition the influence of members Xi( t, f) Xk'( t, f) with a larger modulus is favored. Signals Gu(/i) doGiM( n) enter blocks 203 and 205.

U bloku203sa oznakomPHATrealizuje se generalizovana kroskorelacija u literaturi često označena kao fazna transformacija. Naime, normalizacijom kroskorelacije na svoj moduo gubi se informacija o snazi signala, a ostaje samo informacija o fazi u kojoj je sadržano relativno vremensko kašnjenje signala. In block 203 labeled PHA, generalized cross-correlation, often labeled as phase transformation in the literature, is implemented. Namely, by normalizing the cross-correlation to its modulus, the information about the signal strength is lost, and only the information about the phase in which the relative time delay of the signal is contained remains.

U obradi generalizovanih kroskorelacionih funkcijaG12 Pkatučestvuje filterska funkcijaW( n)koja se generiše u bloku 204. FunkcijaW( n)se dobija obradom mikrofonskih signalaX\doXu,koja će kasnije biti detaljnije opisana, a čiji je cilj da osnovne prozodijske karakteristike govornog signala, pre svega energetsku dinamiku formantnih struktura vokala, iskoristi za pouzdaniju ocenu ugla azimuta, odnosno lokaciju govornika u prostoriji. The filter function W(n) generated in block 204 takes part in the processing of generalized cross-correlation functions G12 P. The function W(n) is obtained by processing microphone signals X\doXu, which will be described in more detail later, and whose goal is to use the basic prosodic characteristics of the speech signal, primarily the energy dynamics of the vocal formant structures, for a more reliable assessment of the azimuth angle, i.e. the location of the speaker in the room.

U bloku205vrši se određivanje estimacije ugla azimuta0na bazi maksimuma generalizovanih kroskorelacionih funkcija. Validnost date estimacije kontroliše blok206,sa oznakomVAD,koji vrši detekciju aktivnosti aktuelnog govornika, i kada je govornik aktivan validna je tekuća estimacija ugla azimuta, u suprotnom usvaja se estimacija dobijena za vreme poslednje njegove aktivnosti. In block 205, the estimation of the azimuth angle is determined based on the maximum of the generalized cross-correlation functions. The validity of the given estimate is controlled by block 206, labeled VAD, which detects the activity of the current speaker, and when the speaker is active, the current estimate of the azimuth angle is valid, otherwise the estimate obtained during his last activity is adopted.

Na slici 3 prikazana je blok šema podsistema za određivanje filterske funkcijeW{ n). Poštogovorni signal ima formantnu strukturu, zbog čega svi frekvencijski binovi nemaju istu snagu, potrebno je selektovati binove sa najvećom snagom i njih iskoristiti za određivanje kroskorelacione funkcije. U tom cilju se u bloku 301 vrši računanje srednje snage mikrofonskih signalaXjdoXupo svakom DFT binu unutar blokan,tj. trenutne snage kanala prema relaciji: Figure 3 shows the block diagram of the subsystem for determining the filter function W{n). The post-speech signal has a formant structure, which is why all frequency bins do not have the same power, it is necessary to select the bins with the highest power and use them to determine the cross-correlation function. To this end, block 301 calculates the mean power of microphone signals XjdoXupo each DFT bin within the block, ie. current channel strengths according to the relationship:

U bloku302određuje se težinska funkcijaW( n)kojom se favorizuju binovi kod kojih postoji rast trenutne snage signala. Razlog izbora ovakvog rešenja je taj što je na delu signala sa naglim rastom snage veći udeo direktnog talasa nego na delu sa padom snage, gde dominiraju refleksije talasa, odnosno reverberacija prostorije. Ovaj pristup se realizuje relacijom: In block 302, the weighting function W(n) is determined, which favors the bins where there is an increase in the current signal strength. The reason for choosing such a solution is that the part of the signal with a sudden increase in power has a larger share of direct waves than the part with a drop in power, where wave reflections dominate, i.e. room reverberation. This approach is realized by the relation:

U bloku 303 vrši se dalja obrada kanalskih trenutnih snaga glačanjem( smoothing,engl.) snagaP( n)po frekvenciji, snagaP( n),a zatim usrednjavanjem po vremenu, snagaP( ri).Glačanje snageP( n)vrši se nekauzalnim IIR filtrom prvog reda (nulto fazno kašnjenje se postiže dvostrukim filtriranjem unapređ i unazad), tako da se dobija snagaP( n),dok se usrednjavanje ove snage po vremenu vrši nelinearnim IIR filtrom prvog reda sa dva koeficijenta usrednjavanja, jedan za rast i drugi za pad snage signala. Ovaj nelinearni filtar se opisuje relacijama: In block 303, further processing of channel current powers is performed by smoothing (smoothing, English) powerP(n) by frequency, powerP(n), and then averaging by time, powerP(ri). Smoothing of powerP(n) is performed by a non-causal IIR filter of the first order (zero phase delay is achieved by double forward and backward filtering), so that powerP(n) is obtained, while the averaging of this power by time is performed by nonlinear A first-order IIR filter with two averaging coefficients, one for increasing and the other for decreasing signal power. This nonlinear filter is described by the relations:

VeličinaP( n)koristi se za definisanje praga odluke za izdvajanje binova sa najvećom snagom u bloku 304. Postupak se sastoji u poređenju veličineP( n)i kroskorelacionih funkcija Cn,2(«) do Gi,m(h) sa binarnom odlukom na izlazu za svaki bin. To znači da se na izlazu bloka 304 dobija M-l binarnih nizova dužine N. SizeP(n) is used to define a decision threshold for extracting bins with the highest power in block 304. The procedure consists of comparing sizeP(n) and cross-correlation functions Cn,2(«) to Gi,m(h) with a binary decision at the output for each bin. This means that at the output of block 304, M-l binary strings of length N are obtained.

Množenjem binarnog izlaza iz bloka 304 i težinske funkcijeW( n)iz bloka 302 dobija se filterska funkcijaW{ n)na uzlazu bloka 305, kojom se ponderišu binovi fazne transformacijeGl kPha,( n)u bloku 203, slika 2. Fazno transformisane kroskorelacione funkcije se dodatno filtriraju IIR filtrom u vremenu kako bi se umanjila varijansa estimacije korelacionih funkcija. Ovo se opisuje relacijom: By multiplying the binary output from block 304 and the weighting function W(n) from block 302, a filter function W{n) is obtained at the output of block 305, which weights the bins of the phase transformation Gl kPha,(n) in block 203, Figure 2. The phase-transformed cross-correlation functions are additionally filtered with an IIR filter in time in order to reduce the variance of the estimation of the correlation functions. This is described by the relation:

Na slici 4 prikazana je detaljna blok šema bloka 205 sa slike 2, u kome se vrši određivanje estimacije ugla azimuta9.Fazno transformisane kroskorelacione funkcije Figure 4 shows a detailed block diagram of block 205 from Figure 2, in which the estimation of the azimuth angle is determined9. Phase-transformed cross-correlation functions

G, kPhatse u bloku401pomoću inverzne Fourierove transformacije (IFFT) transformišu iz frekvencijskog u vremensi domen u kroskorelacije /?i,2(x) doR], m( x),Pre IFFT transformacije primenjuje se u bloku402apriorno odbacivanje binova koji se nalaze izvan opsega od interesa. Kriterijum ovog odbacivanja je izbor opsega frekvencija za koji je snaga govornog signala dovoljno velika a da za najveću frekvenciju opsega ne dolazi do alijasinga u prostornom domenu. G, kPhats in block 401 are transformed from the frequency to the time domain into cross-correlations /?i,2(x) doR], m(x) using the inverse Fourier transform (IFFT). Before the IFFT transformation, a priori rejection of bins that are outside the range of interest is applied in block 402. The criterion for this rejection is the selection of a frequency range for which the power of the speech signal is high enough without aliasing occurring in the spatial domain for the highest frequency of the range.

U bloku403vrši se vremensko usklađivanje kroskorelacionih funkcija R\ j.( t) doRi, m( x)primenom odgovarajućih faktora interpolacije, koje se zatim usrednjavaju i na njihovoj srednjoj vrednostiR, r( T)se određuje maksimum u bloku404,čija apscisa predstavlja estimaciju vremenskog kašnjenja f . In block 403, cross-correlation functions R\ j.(t) to Ri, m(x) are time-matched by applying appropriate interpolation factors, which are then averaged and at their mean valueR, r(T), the maximum is determined in block 404, whose abscissa represents the estimate of the time delay f.

U bloku405vrši se preračunavanje vremenskog kašnjenjaiu upadni ugao6Rdirektnog talasa aktivnog govornika. Estimacija dolaznog pravca ima smisla kada je govornik aktivan; kada nije aktivan za validnu estimaciju se usvaja estimacija dobijena za vreme poslednje njegove aktivnosti. U tu svrhu u bloku406se pod kontrolom signala VAD definitivno određuje validnost estimacije0R,tako da se na izlazu dobija konačna vrednost estimiranog ugla azimuta9.In block 405, the time delay and incident angle 6R of the direct wave of the active speaker are recalculated. Incoming direction estimation makes sense when the speaker is active; when he is not active, the estimate obtained during his last activity is used for a valid estimate. For this purpose, the validity of the estimation 0R is definitively determined under the control of the VAD signal in block 406, so that the final value of the estimated azimuth angle 9 is obtained at the output.

U cilju detekcije aktivnosti govornika koriste se: a) informacija iz bloka301o srednjoj snazi mikrofonskih signalaP( n),slika 3, i b) informacijasbfiz bloka501,blokSD-BFsuperdirektivni usmerivač, slika 5. Na osnovu ovih informacija u bloku502se donosi odluka o aktivnosti bliskog govornika. In order to detect speaker activity, the following are used: a) information from block 301 about the mean power of microphone signals P(n), Figure 3, and b) information from Phys block 501, block SD-BF superdirective router, Figure 5. Based on this information, a decision is made in block 502 about the activity of a nearby speaker.

Formiranje superdirektivnog prostornog filtra vrši se u bloku501.On obezbeđuje veći indeks usmerenosti u odnosu na prostorni konvencionalni filter koji sadrži samo kompenzaciju kašnjenja i sumiranje. The formation of the superdirective spatial filter is performed in block 501. It provides a higher directivity index compared to a conventional spatial filter that contains only delay compensation and summation.

Za prostoriju sa reverberacijom se obično usvaja model difuznog polja šuma, što podrazumeva da šum dolazi iz svih pravaca sa približno istim intenzitetom. Za takav model polja šuma pokazuje se da je koherencija između dva mikrofona realan broj jednak: For a room with reverberation, a diffuse noise field model is usually adopted, which implies that the noise comes from all directions with approximately the same intensity. For such a noise field model, it is shown that the coherence between two microphones is a real number equal to:

gde je/uČestanost,dtj jerastojanjemikrofona i ij, a cbrzina zvuka. Koherencije parova mikrofonartj( f)formiraju matricu koherencijaTd.Koristeći ovako definisanu matricu koherencijaTd,koeficijenti superdirektivnog mikrofonskog niza se odredjuju u bloku504prema relaciji: gde je Ce vektor usmerenja na pravac odabranog govornika definisan estimiranim uglom azimuta9.Ovaj vektor se određuje u bloku503prema relaciji: where /u is the frequency, dtj is the microphone distance and ij, and c is the sound speed. Coherences of microphone pairs (f) form the coherence matrix Td. Using the coherence matrix Td defined in this way, the coefficients of the superdirective microphone array are determined in block 504 according to the relation: where Ce is the direction vector to the direction of the selected speaker defined by the estimated azimuth angle 9. This vector is determined in block 503 according to the relation:

Veličinad jerastojanje dva susedna mikrofona. Sized is the distance between two adjacent microphones.

Na izlazu bloka 501 dobija se estimacija govoraSbfaktuelnog govornika na bazi relacije: At the output of block 501, an estimate of the actual speaker's speech is obtained based on the relation:

Prema tome, u blok502dolaze dve informacije: srednja snaga mikrofonskih signalaP( n),koja pored aktuelnog govornika sadrži i sve signale smetnji u prostoriji, i signal estimacije govorasbfaktuelnog govornika na pravcu estimiranog ugla azimuta9.U bloku502se vrši binarna odluka o aktivnosti aktuelnog govornika na bazi komparativne analize prispelih informacija i binarni signal VAD odlučuje o izlaznoj vrednosti ugla azimuta6,odnosno na izlaz sistema se prosleđuje trenutna estimacija dolaznog pravca ako je aktivan aktuelni govornik, u suprotnom se prosleđuje poslednja validna estimacija pravca. Therefore, two pieces of information arrive in block 502: the mean power of the microphone signals P(n), which, in addition to the current speaker, also contains all noise signals in the room, and the speech estimation signal b of the actual speaker in the direction of the estimated azimuth angle 9. In block 502, a binary decision is made about the activity of the current speaker based on a comparative analysis of the received information and the binary signal VAD decides on the output value of the azimuth angle 6, that is, the current estimate is sent to the output of the system incoming direction if the current speaker is active, otherwise the last valid direction estimate is passed.

U ovom pronalasku opisan je postupak obrade akustičkih i govornih signala u cilju lokacije govornika u prostoru u odnosu na sistem za lokaciju govornika koncipiranog na bazi mikrofonskog niza. Opisanim sistemom se govornik može locirati u zatvorenom ili otvorenom prostoru, a sistem se može primeniti u kontroli i upravljanju robota, video kamere ili procesa koji zahtevaju interaktivnu informaciju o lokaciji govornika, ili u „hands-free" komunikacionim sistemima kao što su telekonferencijski sistemi, video konferencijski sistemi, spikerfoni, itd. This invention describes the process of processing acoustic and speech signals in order to locate the speaker in space in relation to the speaker location system designed on the basis of a microphone array. The described system can locate the speaker in a closed or open space, and the system can be applied in the control and management of robots, video cameras or processes that require interactive information about the location of the speaker, or in "hands-free" communication systems such as teleconference systems, video conference systems, speakerphones, etc.

Postupci i tehnike obrade akustičkih i govornih signala u ovom pronalasku su nezavisne od broja mikrofona u nizu a nalaze se pod kontrolom većeg broja parametara koji omogućavaju optimizaciju rešenja za različite aplikacije. The procedures and techniques for processing acoustic and speech signals in this invention are independent of the number of microphones in the array and are under the control of a number of parameters that enable the optimization of solutions for different applications.

Postupci i tehnike obrade akustičkih i govornih signala u ovom pronalasku mogu se implementirati na različite načine. Na primer, ove tehnike mogu biti implementirane u hardveru, softveru ili kombinovano. U hardverskoj implementaciji mogu se koristiti specifična integrisana kola (ASIC), procesori za digitalnu obradu signala (DSP), programabilna logička kola (PLD ili FPGA) i druga elektronska kola projektovana tako da mogu izvršiti opisane funkcije u ovom pronalasku. The acoustic and speech signal processing methods and techniques of the present invention can be implemented in a variety of ways. For example, these techniques can be implemented in hardware, software, or a combination. A hardware implementation may use specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic circuits (PLDs or FPGAs), and other electronic circuits designed to perform the functions described in this invention.

Postupci i tehnike obrade akustičkih i govornih signala u ovom pronalasku mogu se implementirati i softverski, tako da se programski kodovi mogu memoristi u memorijskim jedinicama i izvršavati pomoću procesora kao što su PC, PDA, DSP, itd. The acoustic and speech signal processing methods and techniques of the present invention can also be implemented in software, so that program codes can be stored in memory units and executed by processors such as PCs, PDAs, DSPs, etc.

Detalji ovog pronalaska opisani ovde omogućavaju bilo kom stručnjaku u ovoj oblasti da generičke principe ovog pronalaska može implementirati u drugim sistemima čime se ne izlazi iz okvira ovog pronalaska. The details of the invention described herein allow any person skilled in the art to implement the generic principles of the invention in other systems without departing from the scope of the invention.

Claims (29)

1. Sistem za lociranje govornika pomoću mikrofonskog nizakarakterisan time,što sadrži: mikrofonski niz od M mikrofona u odnosu na čiju simetralu se određuje ugao azimuta, odnosno položaj govornika u horizontalnoj ravni; blok za predprocesiranje mikrofonskih signala i konverziju u digitalnu formu i frekvencijski domen; blok za kroskorelacionu analizu mikrofonskih signala i njenu optimizaciju na bazi fazne transformacije; blok za određivanje fdterske funkcije na bazi prozodijskih karakteristika govornog signala, pomoću koje se vrši optimizacija kroskorelacione PHAT analize; blok za detekciju aktivnosti govora (VAD) zasnovan na superdirektivnom usmerivaču (SD-BF) koji obezbeđuje prostorno filtriranje govornika; blok za estimaciju ugla azimuta na bazi maksimuma interpoliranih kroskorelacionih funkcija.1. A system for locating the speaker using a microphone array, characterized by the fact that it contains: a microphone array of M microphones in relation to whose bisector the angle is determined azimuth, that is, the position of the speaker in the horizontal plane; block for preprocessing of microphone signals and conversion into digital form i frequency domain; block for cross-correlation analysis of microphone signals and its optimization at based on phase transformation; block for determining the phdter function on the basis of prosodic characteristics speech signal, which is used to optimize the cross-correlation PHAT analysis; superdirective-based speech activity detection (VAD) block the router (SD-BF) which provides spatial filtering of speakers; block for estimation of the azimuth angle based on the interpolated maxima cross-correlation functions. 2. Sistem prema zahtevu 1karakterisan time,što sadrži blokove koji vrše detekciju govorne aktivnosti aktuelnog govornika, koji vrše separaciju aktuelnog govornika u odnosu na sve ostale izvore smetnji i koji vrše adaptivno praćenje aktuelnog govornika u pokretu.2. The system according to claim 1, characterized by the fact that it contains blocks that detect the speech activity of the current speaker, that perform the separation of the current speaker in relation to all other sources of interference and that perform adaptive monitoring of the current speaker in motion. 3. Sistem prema zahtevu 1karakterisan time,što sadrži mikrofonski niz od M mikrofona i što broj mikrofona u nizu nije ograničavajući faktor.3. The system according to claim 1, characterized by the fact that it contains a microphone array of M microphones and that the number of microphones in the array is not a limiting factor. 4. Sistem prema zahtevu 2karakterisan time,što se mikrofonski niz nalazi u horizontalnoj ravni i što se lociranje govornika određuje pomoću ugla azimuta u odnosu na simetralu mikrofonskog niza.4. The system according to claim 2, characterized by the fact that the microphone array is located in a horizontal plane and that the location of the speaker is determined by the azimuth angle in relation to the bisector of the microphone array. 5. Sistem prema zahtevu 1karakterisan time,što se obrada signala odvija u frekvencijskom domenu i što se ista može realizovati u realnom vremenu.5. The system according to claim 1, characterized by the fact that signal processing takes place in the frequency domain and that it can be realized in real time. 6. Sistem prema bilo kom od prethodnih zahtevakarakterisan time,što sadrži blok kroskorelacione analize mikrofonskih signala koji određuje vremensko kašnjenje zvučnih talasa od izvora zvuka do mikrofonskog niza (TDOA) i koji vrši njenu optimizaciju kroskorelacione analize na bazi fazne transformacije (PHAT).6. The system according to any of the previous requirements, characterized by the fact that it contains a block of cross-correlation analysis of microphone signals that determines the time delay of sound waves from the sound source to the microphone array (TDOA) and performs its optimization by cross-correlation analysis based on phase transformation (PHAT). 7. Sistem prema zahtevu 6karakterisan time,što sadrži blok za određivanje filterske funkcije na bazi prozodijskih karakteristika govornog signala koji omogućava optimizaciju kroskorelacione PHAT analize prilagođenu karakteristikama govornog signala.7. The system according to claim 6, characterized by the fact that it contains a block for determining the filter function based on the prosodic characteristics of the speech signal, which enables the optimization of cross-correlation PHAT analysis adapted to the characteristics of the speech signal. 8. Sistem prema zahtevima 1 do 5karakterisan time,što sadrži blok za detekciju aktivnosti govora (VAD) u uslovima ambijentalnih smetnji i reverberacije, koji odlučuje o konačnoj vrednosti ugla azimuta.8. The system according to claims 1 to 5, characterized by the fact that it contains a speech activity detection block (VAD) in conditions of ambient disturbances and reverberation, which decides on the final value of the azimuth angle. 9. Sistem prema zahtevu 8karakterisan time,što osnovu bloka za detekciju aktivnosti govora (VAD) čini superdirektivi usmerivač (SD-BF) koji obezbeđuje separaciju aktuelnog govornika od ostalih izvora zvuka u prostoriji na bazi prostornog filtriranja.9. The system according to claim 8, characterized by the fact that the basis of the speech activity detection block (VAD) is a superdirective router (SD-BF) that ensures the separation of the current speaker from other sound sources in the room based on spatial filtering. 10. Sistem prema zahtevima 6 do 7karakterisan time,što sadrži blok za estimaciju ugla azimuta na bazi maksimuma interpoliranih i usrednjenih M-l kroskorelacionih funkcija.10. The system according to claims 6 to 7, characterized by the fact that it contains a block for estimating the azimuth angle based on the maximum of the interpolated and averaged M-1 cross-correlation functions. 11. Sistem prema bilo kom od prethodnih zahtevakarakterisan time,što se može primeniti za kontrolu uređaja, sistema ili procesa putem glasa.11. A system according to any of the previous requirements, characterized by the fact that it can be applied to control devices, systems or processes by voice. 12. Sistem prema bilo kom od prethodnih zahtevakarakterisan time,što se može primeniti u „hands-free" komunikacionim sistemima za slobodnu govornu komunikaciju u cilju poboljšanja kvaliteta i razumljivosti komunikacije u akustičkom ambijentu.12. The system according to any of the previous requirements, characterized by the fact that it can be applied in "hands-free" communication systems for free speech communication in order to improve the quality and intelligibility of communication in an acoustic environment. 13. Postupak za lociranje govornika pomoću mikrofonskog nizakarakterisantime, što sadrži: kroskorelacionu analizu koja vrši analizu vremenskog kašnjenja zvučnih talasa od izvora zvuka do mikrofonskog niza; generalizaciju kroskorelacione analize, odnosno njenu faznu transformaciju (PHAT), uz primenu adaptivnog ponderisanja filterskom funkcijomW ( n) ;adaptivno određivanje filterske funkcijeW( ri)na bazi prozodijskih karakteristika govornog signala; adaptivnu detekciju aktivnosti govornika (VAD) na bazi superdirektivnog usmerivača (SD-BF); interpolaciju kroskorelacionih funkcija i određivanje estimacije ugla azimuta.13. A method for locating a speaker using a microphone array of characteristics, which includes: cross-correlation analysis that analyzes the time delay of sound waves from the sound source to the microphone array; generalization of cross-correlation analysis, i.e. its phase transformation (PHAT), with the application of adaptive weighting by the filter function W ( n); adaptive determination of the filter function W( ri) based on prosodic speech signal characteristic; adaptive speaker activity detection (VAD) based on superdirectivity router (SD-BF); interpolation of cross-correlation functions and determination of azimuth angle estimation. 14. Postupak prema zahtevu 13karakterisan time,što se kroskorelacija vrši između prvog mikrofonskog signala i svih ostalih mikrofonskih signala, tako da se izvršava M-l kroskorelacija.14. The method according to claim 13, characterized by the fact that cross-correlation is performed between the first microphone signal and all other microphone signals, so that M-1 cross-correlation is performed. 15. Postupak prema zahtevu 14karakterisan time,što se generalizacija kroskorelacione analize vrši normalizacijom kroskorelacije na svoj moduo pri čemu se gubi informacija o snazi signala, a ostaje samo informacija o fazi u kojoj je sadržano relativno vremensko kašnjenje između analiziranih signala.15. The procedure according to claim 14, characterized by the fact that the generalization of the cross-correlation analysis is carried out by normalizing the cross-correlation to its modulus, whereby the information about the signal strength is lost, and only the information about the phase that contains the relative time delay between the analyzed signals remains. 16. Postupak prema zahtevu 15karakterisan time,što se fazno transformisane kroskorelacione funkcije dodatno adaptivno filtriraju IIR filtrom u vremenu kako bi se umanjile varijanse estimacije korelacionih funkcija.16. The method according to claim 15, characterized by the fact that the phase-transformed cross-correlation functions are additionally adaptively filtered with an IIR filter in time in order to reduce the variances of the estimation of the correlation functions. 17. Postupak prema zahtevu 16karakterisantime, što se dodatno filtriranje fazno transformisanih kroskorelacionih funkcija vrši adaptivnim ponderisanjem binova fazne transformacije filterskom funkcijomW( ri).17. The method according to claim 16, characterized in that the additional filtering of the phase-transformed cross-correlation functions is performed by adaptive weighting of the phase transformation bins with the filter function W(ri). 18. Postupak prema zahtevu 13karakterisan time,što se filterska funkcijaW{ n)određuje na bazi prozodijskih karakteristika govornog signala detektovanih u trenutnoj snazi mikrofonskih signala.18. The method according to claim 13, characterized by the fact that the filter function W{n) is determined based on the prosodic characteristics of the speech signal detected in the current strength of the microphone signals. 19. Postupak prema zahtevu 18karakterisan time,što se filterskom funkcijomW( n)selektuju binovi sa najvećom snagom i oni koriste za određivanje kroskorelacionih funkcija.19. The procedure according to claim 18, characterized by the fact that bins with the highest power are selected by the filter function W(n) and used to determine the cross-correlation functions. 20. Postupak prema zahtevu 18 i 19karakterisan time,što se u određivanju filterske funkcijeW( ri)izračunavaju trajektorije snaga mikrofonskih signala usrednjavanjem po frekvenciji i po vremenu.20. The method according to claim 18 and 19, characterized by the fact that in determining the filter function W(ri), the power trajectories of the microphone signals are calculated by averaging over frequency and over time. 21. Postupak prema zahtevu 18 do 20karakterisan time,što se u određivanju filterske funkcijeW( n)favorizuju binovi kod kojih postoji rast trenutne snage signala, iz razloga što je na delu signala sa naglim rastom snage veći udeo direktnog talasa nego na delu sa padom snage, gde dominiraju refleksije talasa, odnosno reverberacija prostorije.21. The procedure according to claim 18 to 20, characterized by the fact that in the determination of the filter function W(n), the bins are favored where there is an increase in the current signal strength, due to the fact that the part of the signal with a sudden increase in power has a greater share of direct waves than the part with a drop in power, where wave reflections dominate, i.e. room reverberation. 22. Postupak prema zahtevu 13karakterisan time,što se u detektoru aktivnosti govora (VAD) donosi odluka o aktivnosti bliskog govornika.22. The method according to claim 13, characterized by the fact that a decision is made in the speech activity detector (VAD) about the activity of a nearby speaker. 23. Postupak prema zahtevu 13 i 22karakterisan time,što se VAD bazira na superdirektivnom usmerivaču (SD-BF) koji obradom mikrofonskih signala obezbeđuje usmerenu karakteristiku osetljivosti mikrofonskog niza.23. The method according to claim 13 and 22, characterized by the fact that the VAD is based on a superdirective router (SD-BF) which, by processing microphone signals, ensures the directional characteristic of the sensitivity of the microphone array. 24. Postupak prema zahtevu 23 karakterisan time, što superdirektivni usmerivač vrši prostorno filtriranje kojim ističe signal aktuelnog govornika i potiskuje signale ambijentalnih smetnji.24. The method according to claim 23, characterized by the fact that the superdirective router performs spatial filtering, which emphasizes the signal of the current speaker and suppresses signals of ambient interference. 25. Postupak prema zahtevima 23 i 24 karakterisan time, što se karakteristikom usmerenosti superdirektivnog usmerivača (SD-BF) upravlja estimiranim uglom azimuta9.25. The method according to claims 23 and 24, characterized in that the directionality characteristic of the superdirective beam (SD-BF) is controlled by the estimated azimuth angle9. 26. Postupak prema zahtevu 13 karakterisan time, što se estimacija ugla azimuta6određuje na bazi maksimuma usklađenih i usrednjenih M-l kroskorelacionih funkcija.26. The method according to claim 13, characterized by the fact that the estimation of the azimuth angle is determined based on the maximum of the matched and averaged M-1 cross-correlation functions. 27. Postupak prema zahtevu 26 karakterisan time, što se usklađivanje kroskorelacionih funkcija vrši postupkom interpolacije.27. The method according to claim 26, characterized in that the matching of cross-correlation functions is performed by an interpolation method. 28. Postupak prema zahtevima 13,26 i 27 karakterisan time, što se estimacija ugla azimuta9dobija preračunavanjem vremenskog kašnjenja f na kome se nalazi maksimum usklađenih i usrednjenih M-l kroskorelacionih funkcija u upadni ugao6direktnog talasa aktuelnog govornika.28. The procedure according to claims 13, 26 and 27, characterized by the fact that the estimation of the azimuth angle is obtained by recalculating the time delay f at which the maximum of the matched and averaged M-l cross-correlation functions is located in the angle of incidence of the direct wave of the current speaker. 29. Postupak prema zahtevima 13 i 28 karakterisan time, što se ugao azimuta6određuje na bazi aktivnosti aktuelnog govornika, pa se na izlaz sistema prosleđuje trenutna estimacija dolaznog pravca9u slučaju aktivnosti aktuelnog govornika, u suprotnom kada nije aktivan prosleđuje se poslednja validna estimacija azimuta6.29. The procedure according to claims 13 and 28, characterized by the fact that the azimuth angle6 is determined based on the activity of the current speaker, so the current estimate of the incoming direction9 is sent to the system output in the case of the activity of the current speaker, otherwise when it is not active, the last valid estimate of the azimuth6 is sent.
RSP-2006/0642A 2006-11-21 2006-11-21 SYSTEM AND PROCEDURE FOR LOCATION OF SPEAKERS BY MICROPHONE RS49859B (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
RSP-2006/0642A RS49859B (en) 2006-11-21 2006-11-21 SYSTEM AND PROCEDURE FOR LOCATION OF SPEAKERS BY MICROPHONE

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
RSP-2006/0642A RS49859B (en) 2006-11-21 2006-11-21 SYSTEM AND PROCEDURE FOR LOCATION OF SPEAKERS BY MICROPHONE

Publications (2)

Publication Number Publication Date
RS20060642A RS20060642A (en) 2007-06-04
RS49859B true RS49859B (en) 2008-08-07

Family

ID=43646405

Family Applications (1)

Application Number Title Priority Date Filing Date
RSP-2006/0642A RS49859B (en) 2006-11-21 2006-11-21 SYSTEM AND PROCEDURE FOR LOCATION OF SPEAKERS BY MICROPHONE

Country Status (1)

Country Link
RS (1) RS49859B (en)

Also Published As

Publication number Publication date
RS20060642A (en) 2007-06-04

Similar Documents

Publication Publication Date Title
US10957338B2 (en) 360-degree multi-source location detection, tracking and enhancement
CN110770827B (en) Near field detector based on correlation
CN113113034B (en) Multi-source tracking and speech activity detection for planar microphone arrays
US5208864A (en) Method of detecting acoustic signal
EP3566461B1 (en) Method and apparatus for audio capture using beamforming
US10638224B2 (en) Audio capture using beamforming
US10887691B2 (en) Audio capture using beamforming
CN110610718B (en) Method and device for extracting expected sound source voice signal
CN106251877A (en) Voice Sounnd source direction method of estimation and device
RS49875B (en) SYSTEM AND PROCEDURE FOR FREE SPEECH COMMUNICATION WITH A MICROPHONE STRIP
Brutti et al. Multiple source localization based on acoustic map de-emphasis
US10049685B2 (en) Integrated sensor-array processor
Rodemann et al. Real-time sound localization with a binaural head-system using a biologically-inspired cue-triple mapping
CN112951261B (en) Sound source positioning method and device and voice equipment
Rascon et al. Lightweight multi-DOA tracking of mobile speech sources
CN109212481A (en) A method of auditory localization is carried out using microphone array
Ba et al. Enhanced MVDR beamforming for arrays of directional microphones
CN111157949A (en) Voice recognition and sound source positioning method
Marti et al. Real time speaker localization and detection system for camera steering in multiparticipant videoconferencing environments
Bechler et al. Three different reliability criteria for time delay estimates
RS20060642A (en) System and technique for speaker localization using microphone array
US10204638B2 (en) Integrated sensor-array processor
TWI517143B (en) A method for noise reduction and speech enhancement
US12143783B1 (en) Sound source localization with reflection detection
Potamitis et al. Speech separation of multiple moving speakers using multisensor multitarget techniques