CN105682000B - A kind of audio-frequency processing method and system - Google Patents

A kind of audio-frequency processing method and system Download PDF

Info

Publication number
CN105682000B
CN105682000B CN201610017000.1A CN201610017000A CN105682000B CN 105682000 B CN105682000 B CN 105682000B CN 201610017000 A CN201610017000 A CN 201610017000A CN 105682000 B CN105682000 B CN 105682000B
Authority
CN
China
Prior art keywords
binaural
signal
audio
signals
rotation angle
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Active
Application number
CN201610017000.1A
Other languages
Chinese (zh)
Other versions
CN105682000A (en
Inventor
张晨
孙学京
刘皓
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
China Film Science And Technology Research Institute Film Technology Quality Inspection Institute Of Central Propaganda Department
Original Assignee
Beijing Tuoling Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Tuoling Inc filed Critical Beijing Tuoling Inc
Priority to CN201610017000.1A priority Critical patent/CN105682000B/en
Publication of CN105682000A publication Critical patent/CN105682000A/en
Application granted granted Critical
Publication of CN105682000B publication Critical patent/CN105682000B/en
Active legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/302Electronic adaptation of stereophonic sound system to listener position or orientation
    • H04S7/303Tracking of listener position or orientation
    • H04S7/304For headphones
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/008Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/305Electronic adaptation of stereophonic audio signals to reverberation of the listening space
    • H04S7/306For headphones

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Multimedia (AREA)
  • Mathematical Physics (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Stereophonic System (AREA)

Abstract

本发明涉及一种云端音频处理方法,服务器和系统,针对不同格式的音频信号,根据客户端的头部旋转角度,分别对所述不同格式的音频信号进行双耳转码,生成相应格式的双声道音频信号;对所述相应格式的双声道信号叠加,得到音频双耳输出虚拟环绕声信号。本发明的音频处理是在云端服务器上进行的,很好的适应了现有的基于云架构音频处理和存储,从而减少了移动终端生成虚拟环绕声质量不高、运算量大的问题。另外,针对上述在服务器上执行可能带来的延迟,本发明还包括对于角度进行平滑处理,以消除延迟。

The present invention relates to a cloud audio processing method, a server and a system. For audio signals of different formats, according to the head rotation angle of the client, binaural transcoding is performed on the audio signals of different formats to generate binaural sounds of corresponding formats. channel audio signal; superimpose the two-channel signal of the corresponding format to obtain the audio binaural output virtual surround sound signal. The audio processing of the present invention is carried out on the cloud server, which is well adapted to the existing audio processing and storage based on the cloud architecture, thereby reducing the problems of low quality and heavy calculation of virtual surround sound generated by the mobile terminal. In addition, for the possible delay caused by the above-mentioned execution on the server, the present invention also includes smoothing the angle to eliminate the delay.

Description

一种音频处理方法和系统An audio processing method and system

技术领域technical field

本发明涉及信号处理技术领域,特别涉及一种音频处理的方法,服务器以及系统。The present invention relates to the technical field of signal processing, in particular to an audio processing method, server and system.

背景技术Background technique

在利用虚拟现实头戴设备(head-mounted display,HMD)向用户呈现内容时,采用虚拟3D音频技术,通过立体声耳机向用户播放音频内容,一种提高临场感的方法是跟踪用户头部动作(head tracking),对声音进行相应的处理。比如,如果原始声音被用户感知为来自正前方,当用户向左转头90度后,声音应被处理使得用户感知声音来自正右方90度。在这里虚拟现实设备可以有很多种类,比如带头部跟踪的显示设备,或者只是一部带头部跟踪传感器的立体声耳机。When using a virtual reality head-mounted display (HMD) to present content to users, virtual 3D audio technology is used to play audio content to users through stereo headphones. One way to improve the sense of presence is to track user head movements ( head tracking) to process the sound accordingly. For example, if the original sound is perceived by the user as coming from the front, when the user turns his head 90 degrees to the left, the sound should be processed so that the user perceives the sound as coming from the right 90 degrees. The virtual reality device here can be of many kinds, such as a display device with head tracking, or just a stereo headset with head tracking sensors.

实现头部跟踪也有多种方法。比较常见的是使用多种运动传感器。运动传感器套件通常包括加速度计、陀螺仪和磁力传感器。在运动跟踪和绝对方向方面每种传感器都有自己固有的强项和弱点。因此常用做法是采用传感器“融合”(sensor fusion)将来自各传感器的信号组合在一起,产生一个更加精确的运动检测结果。There are also multiple ways to implement head tracking. It is more common to use a variety of motion sensors. Motion sensor kits typically include accelerometers, gyroscopes, and magnetic sensors. Each sensor has its own inherent strengths and weaknesses when it comes to motion tracking and absolute orientation. It is therefore common practice to use sensor "fusion" (sensor fusion) to combine the signals from the various sensors together to produce a more accurate motion detection result.

在得到头部旋转角度后,需要对声音进行相应的变化。一种方式是将声音转到Ambisonic域,然后再通过使用旋转矩阵对信号做变换。Ambisonic信号通常是多于两个声道,而常见的媒体播放器只支持立体声两声道,这对直接播放Ambisonic或其他多声道的音频信号带来困难。After getting the head rotation angle, you need to change the sound accordingly. One way is to convert the sound to the ambisonic domain and then transform the signal by using a rotation matrix. Ambisonic signals usually have more than two channels, and common media players only support two stereo channels, which makes it difficult to directly play ambisonic or other multi-channel audio signals.

有鉴于此,在本领域需要一种有效且高质量的虚拟环绕声生成和播放的解决方案。In view of this, there is a need in the art for an effective and high-quality solution for generating and playing virtual surround sound.

发明内容Contents of the invention

为了克服现有技术的上述缺陷,本发明的目的在于提供一种云端音频处理方法,服务器和系统,其能有效且高质量地生成虚拟环绕声,主要用于配合虚拟现实头戴设备进行音频的立体声耳机播放,并且所述虚拟环绕声的生成是在云端服务器上进行的,很好的适应了现有的基于云架构的网络类型,由服务器执行虚拟环绕声的生成和存储,从而解决了现有客户端无法播放各种3603D audio,尤其是适用于虚拟现实应用的音频的问题。In order to overcome the above-mentioned defects of the prior art, the object of the present invention is to provide a cloud audio processing method, server and system, which can effectively and high-quality generate virtual surround sound, and is mainly used for audio processing with virtual reality headsets Stereo headphones are played, and the generation of the virtual surround sound is carried out on the cloud server, which is well adapted to the existing network type based on the cloud architecture, and the generation and storage of the virtual surround sound are performed by the server, thus solving the current problem There is a problem that the client cannot play various 3603D audio, especially the audio suitable for virtual reality applications.

为了实现上述目的,本发明提供一种云端音频处理方法,所述音频处理方法包括以下步骤,获取用户头部旋转的旋转角度;获取不同格式的音频信号,根据所述旋转角度,分别对所述不同格式的音频信号进行双耳转码,生成相应格式的双声道音频信号;对所述相应格式的双声道信号叠加,得到音频双耳输出虚拟环绕声信号。In order to achieve the above object, the present invention provides a cloud audio processing method. The audio processing method includes the following steps: acquiring the rotation angle of the user's head; Audio signals in different formats are binaurally transcoded to generate binaural audio signals in corresponding formats; the binaural signals in corresponding formats are superimposed to obtain audio binaural output virtual surround sound signals.

优选地,所述不同格式的音频信号包括双耳录音信号,Ambisonic录音信号和音频对象信号。Preferably, the audio signals in different formats include binaural recording signals, ambisonic recording signals and audio object signals.

优选地,对所述不同格式的音频信号进行双耳转码,生成相应格式的双耳转码音频信号具体包括:Preferably, performing binaural transcoding on the audio signals in different formats, and generating binaural transcoding audio signals in corresponding formats specifically includes:

对所述双耳录音信号,根据所述旋转角度进行插值,生成双耳录音双声道信号;For the binaural recording signal, interpolation is performed according to the rotation angle to generate a binaural recording binaural signal;

对所述Ambisonic录音信号,根据所述旋转角度对所述Ambisonic录音信号进行调整,对所述调整后的Ambisonic录音信号双耳转码生成Ambisonic录音双声道信号;For the ambisonic recording signal, the ambisonic recording signal is adjusted according to the rotation angle, and the adjusted ambisonic recording signal is binaurally transcoded to generate an ambisonic recording binaural signal;

对所述音频对象信号,根据所述旋转角度对所述音频对象信号调整,对所述调整后的音频对象信号双耳转码生成音频对象双声道信号。For the audio object signal, the audio object signal is adjusted according to the rotation angle, and the adjusted audio object signal is binaurally transcoded to generate an audio object binaural signal.

优选地,如需要较高的空间精度,将音频对象信号根据旋转角度进行旋转,将旋转后的音频对象信号编码为高阶B格式音频对象信号,经双耳转码后生成高阶B格式音频对象双声道信号,与Ambisonic录音双声道信号、双耳录音双声道信号进行叠加;Preferably, if higher spatial precision is required, the audio object signal is rotated according to the rotation angle, the rotated audio object signal is encoded into a high-order B-format audio object signal, and the high-order B-format audio is generated after binaural transcoding Object binaural signal, superimposed with ambisonic recording binaural signal and binaural recording binaural signal;

如需要低复杂度低延迟,将音频对象信号编码为一阶B格式音频对象信号,与其他一阶Ambisonic录音信号叠加,然后根据旋转角度对所述叠加后的混合信号进行双耳转码,生成音频对象与Ambisonic录音信号的混合双声道信号,与所述双耳录音双声道信号进行叠加。If low complexity and low delay are required, the audio object signal is encoded into a first-order B-format audio object signal, superimposed with other first-order ambisonic recording signals, and then the superimposed mixed signal is binaurally transcoded according to the rotation angle to generate The mixed binaural signal of the audio object and the ambisonic recording signal is superimposed on the binaural signal of the binaural recording.

优选地,所获取用户头部旋转的旋转角度具体为获取用户头部旋转的旋转角度,对所述旋转角度进行平滑处理。Preferably, the acquired rotation angle of the user's head rotation is specifically acquiring the rotation angle of the user's head rotation, and performing smoothing processing on the rotation angle.

本发明还提供了一种云端音频处理服务器,所述服务器包括:获取单元,获取用户头部旋转的旋转角度;采集单元,采集不同格式的音频信号;双耳转码单元,分别与所述获取单元和采集单元相连接,根据所述旋转角度,分别对所述不同格式的音频信号进行双耳转码,生成相应格式的双声道音频信号;叠加单元,与所述双耳转码单元连接,对所述相应格式的双声道信号叠加,得到音频双耳输出虚拟环绕声信号。The present invention also provides a cloud audio processing server. The server includes: an acquisition unit, which acquires the rotation angle of the user's head; an acquisition unit, which acquires audio signals in different formats; The unit is connected to the acquisition unit, and according to the rotation angle, binaural transcoding is performed on the audio signals in different formats to generate binaural audio signals in corresponding formats; the superposition unit is connected to the binaural transcoding unit , superimposing the binaural signals of the corresponding formats to obtain binaural audio output virtual surround sound signals.

优选地,所述不同格式的音频信号包括双耳录音信号,Ambisonic录音信号和音频对象信号。Preferably, the audio signals in different formats include binaural recording signals, ambisonic recording signals and audio object signals.

优选地,双耳转码单元对所述不同格式的音频信号进行双耳转码,生成相应格式的双耳转码音频信号具体包括:Preferably, the binaural transcoding unit performs binaural transcoding on the audio signals of different formats, and generating a binaural transcoding audio signal of a corresponding format specifically includes:

对所述双耳录音信号,根据所述旋转角度进行插值,生成双耳录音双声道信号;For the binaural recording signal, interpolation is performed according to the rotation angle to generate a binaural recording binaural signal;

对所述Ambisonic录音信号,根据所述旋转角度对所述Ambisonic录音信号进行调整,对所述调整后的Ambisonic录音信号双耳转码生成Ambisonic录音双声道信号;For the ambisonic recording signal, the ambisonic recording signal is adjusted according to the rotation angle, and the adjusted ambisonic recording signal is binaurally transcoded to generate an ambisonic recording binaural signal;

对所述音频对象信号,根据所述旋转角度对所述音频对象信号调整,对所述调整后的音频对象信号双耳转码生成音频对象双声道信号。For the audio object signal, the audio object signal is adjusted according to the rotation angle, and the adjusted audio object signal is binaurally transcoded to generate an audio object binaural signal.

优选地,如需要较高的空间精度,双耳转码单元将音频对象信号根据旋转角度进行旋转,将旋转后的音频对象信号编码为高阶B格式音频对象信号,经双耳转码后生成高阶B格式音频对象双声道信号,叠加单元对双耳转码单元生成的高阶B格式音频对象双声道信号,Ambisonic录音双声道信号、双耳录音双声道信号进行叠加;Preferably, if higher spatial precision is required, the binaural transcoding unit rotates the audio object signal according to the rotation angle, and encodes the rotated audio object signal into a high-order B-format audio object signal, which is generated after binaural transcoding The high-order B-format audio object binaural signal, the superposition unit superimposes the high-order B-format audio object binaural signal generated by the binaural transcoding unit, the ambisonic recording binaural signal, and the binaural recording binaural signal;

如需要低复杂度低延迟,双耳转码单元将音频对象信号编码为一阶B格式音频对象信号,与其他一阶Ambisonic录音信号叠加,然后根据旋转角度对所述叠加后的混合信号进行双耳转码,生成音频对象与Ambisonic录音信号的混合双声道信号,叠加单元对双耳转码单元生成的与所述混合双声道信号、双耳录音双声道信号进行叠加。If low complexity and low delay are required, the binaural transcoding unit encodes the audio object signal into a first-order B-format audio object signal, superimposes it with other first-order ambisonic recording signals, and then double-binds the superimposed mixed signal according to the rotation angle. Ear transcoding generates a mixed binaural signal of an audio object and an ambisonic recording signal, and the superposition unit superimposes the mixed binaural signal and the binaural recording binaural signal generated by the binaural transcoding unit.

优选地,所述云端服务器还包括平滑单元,分别与所述双耳转码单元和所述获取单元连接,平滑单元从获取单元接收用户头部旋转的旋转角度,对所述旋转角度进行平滑处理。Preferably, the cloud server further includes a smoothing unit connected to the binaural transcoding unit and the acquisition unit respectively, the smoothing unit receives the rotation angle of the user's head rotation from the acquisition unit, and performs smoothing processing on the rotation angle .

本发明还提供了一种音频播放系统,所述系统包括云端音频处理服务器,以及客户端;所述客户端包括头部跟踪装置,所述头部跟踪装置抓取头部旋转角度,通过网络上传至所述云端音频处理服务器,所述云端音频处理器接收所述旋转角度,生成音频双耳输出虚拟环绕声信号后,通过所述网络传输至客户端。The present invention also provides an audio playback system, the system includes a cloud audio processing server, and a client; the client includes a head tracking device, and the head tracking device captures the rotation angle of the head and uploads a To the cloud audio processing server, the cloud audio processor receives the rotation angle, generates an audio binaural output virtual surround sound signal, and transmits it to the client through the network.

根据本发明的云端音频处理方法,服务器和系统,有效且高质量地生成虚拟环绕声,主要用于配合虚拟现实头戴设备进行音频的立体声耳机播放,并且所述虚拟环绕声的生成是在云端服务器上进行的,很好的适应了现有的基于云架构的网络类型,由云端服务器执行音频处理和存储,从而解决了现有客户端无法播放各种3603D audio,尤其是适用于虚拟现实应用的音频的问题。According to the cloud audio processing method, server and system of the present invention, effectively and high-quality generation of virtual surround sound is mainly used for stereo earphone playback of audio in conjunction with virtual reality headsets, and the generation of the virtual surround sound is in the cloud It is carried out on the server, which is well adapted to the existing cloud-based network type, and the cloud server performs audio processing and storage, thus solving the problem that the existing client cannot play various 3603D audio, especially for virtual reality applications audio problem.

采用本发明的云端音频处理技术,在多人语音通讯中会大大提升临场感,用户可以随意转头来关注某一方向的声音,更加逼近现实中的多人交谈场景。特别在使用流媒体的场景中,通过实时调整空间声,音频的方位,可以提升用户的音频体验。如果辅助虚拟现实视频内容,则会更好的提升用户体验。Using the cloud audio processing technology of the present invention can greatly enhance the sense of presence in multi-person voice communication, and users can turn their heads at will to pay attention to the sound in a certain direction, which is closer to the real multi-person conversation scene. Especially in scenarios where streaming media is used, the user's audio experience can be improved by adjusting spatial sound and audio orientation in real time. If the virtual reality video content is assisted, the user experience will be better improved.

附图说明Description of drawings

图1是本发明的云端音频处理方法一个实施例的原理框图;Fig. 1 is a functional block diagram of an embodiment of the cloud audio processing method of the present invention;

图2a-c是本发明的云端音频处理方法另一个实施例的原理框图;2a-c are schematic block diagrams of another embodiment of the cloud audio processing method of the present invention;

图3是本发明的音频处理服务器的一个实施例的结构示意图;Fig. 3 is a schematic structural diagram of an embodiment of the audio processing server of the present invention;

图4是本发明的音频处理系统的另一个实施例的结构示意图;Fig. 4 is a schematic structural diagram of another embodiment of the audio processing system of the present invention;

具体实施方式detailed description

实施例一:如图1所示,一种对音频对象处理包括如下处理步骤:Embodiment one: as shown in Figure 1, a kind of audio object processing comprises the following processing steps:

通过头部跟踪装置获取用户头部旋转角度;Obtain the rotation angle of the user's head through the head tracking device;

根据所述旋转角度,将音频对象编码到高阶(优选为2阶或3阶)Ambisonic B-格式信号;Encoding the audio object into a higher order (preferably 2nd or 3rd order) Ambisonic B-format signal according to said rotation angle;

将所述Ambisonic B-格式信号转换成虚拟扬声器阵列信号;以一个一阶B-格式信号[W1 X1 Y1 Z1]T为例,转换成虚拟扬声器阵列信号[L1 L2 … LN]T的过程就是进行下列运算:Convert the Ambisonic B-format signal into a virtual loudspeaker array signal; take a first-order B-format signal [W 1 X 1 Y 1 Z 1 ] T as an example, convert it into a virtual loudspeaker array signal [L 1 L 2 ... L The process of N ] T is to perform the following operations:

其中,N为虚拟扬声器拓扑结构中包括的虚拟扬声器的数目。上式中所用的G矩阵为ambisonic解码矩阵,可以通过求伪逆矩阵来得出。Wherein, N is the number of virtual speakers included in the virtual speaker topology. The G matrix used in the above formula is an ambisonic decoding matrix, which can be obtained by calculating the pseudo-inverse matrix.

对音频对象的所述虚拟扬声器阵列信号基于双耳房间脉冲响应(BRIR)进行双耳转码(通常是3维,即包含高度信息),得到音频对象的双耳输出虚拟环绕声信号。具体是:从虚拟扬声器信号转到耳机信号对应的二路立体声BRIR矩阵,将该二路立体声矩阵和虚拟扬声器阵列信号进行矩阵乘法,得到虚拟环绕声。The virtual speaker array signal of the audio object is binaurally transcoded based on the binaural room impulse response (BRIR) (usually 3D, ie contains height information), to obtain the binaural output virtual surround sound signal of the audio object. Specifically: transfer the virtual speaker signal to the two-way stereo BRIR matrix corresponding to the headphone signal, perform matrix multiplication on the two-way stereo matrix and the virtual speaker array signal, and obtain virtual surround sound.

BRIR矩阵为则虚拟环绕声为所述音频信号可以为一个或多个。The BRIR matrix is Then the virtual surround sound is There may be one or more audio signals.

所述双耳房间脉冲响应优选为离线生成,可以采用真实测量或由专门的软件生成,因此不必像现有技术下采用在线生成方式时需要存储大量的BRIR,减少了内存消耗。The binaural room impulse response is preferably generated offline, which can be generated by real measurement or by special software. Therefore, it is not necessary to store a large amount of BRIR as in the prior art when online generation is used, reducing memory consumption.

将音频对象编码到Ambisonic B-格式信号时,水平方向阶数优选大于或等于垂直方向阶数,例如,水平方向编码优选为3阶Ambisonic B-格式信号时,垂直方向编码优选为2阶或1阶Ambisonic B-格式信号,分别用H3V2、H3V1表示。由于人对高度感知低于平面角度的分辨率,因此采用以上适当在某个特定方向上降低阶数的方法,减少了运算量,但又不明显降低用户对声音的感知效果。When encoding an audio object into an Ambisonic B-format signal, the order in the horizontal direction is preferably greater than or equal to the order in the vertical direction. For example, when the encoding in the horizontal direction is preferably a 3-order Ambisonic B-format signal, the encoding in the vertical direction is preferably 2-order or 1 The second-order Ambisonic B-format signal is represented by H3V2 and H3V1 respectively. Since the height perception of human beings is lower than the resolution of the plane angle, the above method of appropriately reducing the order in a specific direction reduces the amount of calculation, but does not significantly reduce the user's perception of sound.

对声场信号即环境声进行处理包括如下步骤:Processing the sound field signal, that is, the ambient sound, includes the following steps:

将环境声转换成环境声的双耳输出虚拟环绕声信号,再将所述音频对象(此时的音频对象主要是指环境声之外的声音内容)和所述环境声各自的双耳输出虚拟环绕声信号对应混音并双耳输出。图1所示为该方法的一个实施例的原理框图。其中,所述将环境声(即图1中的声场信号)转换成环境声的双耳输出虚拟环绕声信号优选包括如下步骤:The ambient sound is converted into the binaural output virtual surround sound signal of the ambient sound, and then the audio object (the audio object at this time mainly refers to the sound content other than the ambient sound) and the respective binaural output of the ambient sound virtual Surround sound signals are mixed and binaurally output. Figure 1 shows a schematic block diagram of one embodiment of the method. Wherein, the binaural output virtual surround sound signal that described environmental sound (being the sound field signal in Fig. 1) is converted into environmental sound preferably comprises the following steps:

获取环境声的1阶Ambisonic B-格式信号;Obtain the 1st order Ambisonic B-format signal of the ambient sound;

根据所述旋转角度,将环境声的所述Ambisonic B-格式信号旋转得到旋转后的Ambisonic B-格式信号;具体来说,是根据所述旋转角度生成旋转矩阵,再根据所述旋转矩阵,对环境声的所述Ambisonic B-格式信号(即待调整信号)进行旋转。所谓旋转,即将旋转矩阵与待调整信号矩阵相乘,旋转不改变音频信号矩阵分量的大小,只改变分量的方向。旋转矩阵的阶数与音频信号矩阵相适应。例如,当待调整信号矩阵为[W2 X2 Y2]T时,旋转矩阵为当待调整信号矩阵为[W2 X2 Y2 Z2]T时,旋转矩阵为 According to the rotation angle, the ambisonic B-format signal of the ambient sound is rotated to obtain the rotated ambisonic B-format signal; specifically, a rotation matrix is generated according to the rotation angle, and then according to the rotation matrix, the The ambisonic B-format signal of the ambient sound (ie the signal to be adjusted) is rotated. The so-called rotation refers to multiplying the rotation matrix by the signal matrix to be adjusted. The rotation does not change the size of the audio signal matrix components, but only changes the direction of the components. The order of the rotation matrix is adapted to the audio signal matrix. For example, when the signal matrix to be adjusted is [W 2 X 2 Y 2 ] T , the rotation matrix is When the signal matrix to be adjusted is [W 2 X 2 Y 2 Z 2 ] T , the rotation matrix is

将环境声的所述旋转后的Ambisonic B-格式信号转换成虚拟扬声器阵列信号;对环境声的所述虚拟扬声器阵列信号基于头相关变换函数(HRTF)进行双耳转码(通常是2维,即不包含高度信息),得到环境声的双耳输出虚拟环绕声信号。HRTF在时间域所对应的名称是HRIR(Head Related Impulse Response)。The Ambisonic B-format signal after the rotation of the ambient sound is converted into a virtual speaker array signal; the virtual speaker array signal of the ambient sound is binaurally transcoded based on a head-related transform function (HRTF) (usually 2-dimensional, That is, the height information is not included), and the binaural output of the ambient sound is obtained to output the virtual surround sound signal. The name corresponding to HRTF in the time domain is HRIR (Head Related Impulse Response).

需要指出的是对于音频对象或环境声都可以根据需要使用BRIR或HRIR进行滤波。由于BRIR通常包含房间模型和一组描述声音方位的HRIR/HRTF组成,所以如果输入信号已带有房间或环境的信息则使用HRIR就可满足需求。It should be pointed out that BRIR or HRIR can be used to filter audio objects or ambient sounds as needed. Since BRIR usually consists of a room model and a set of HRIR/HRTF describing the sound orientation, if the input signal already has room or environment information, using HRIR can meet the requirements.

所述生成虚拟环绕声的方法在实施运算时优选基于以下假定:虚拟扬声器阵列具有左右对称性,用户在房间的中轴线上,用户对应的所述双耳房间脉冲响应和头相关变换函数也具有左右对称性。基于该假设,可以利用高阶Ambisonic B-格式对称性优化方法,显著减少运算量,提高运算效率。The method for generating virtual surround sound is preferably based on the following assumptions when performing calculations: the virtual loudspeaker array has left-right symmetry, the user is on the central axis of the room, and the binaural room impulse response and head-related transfer function corresponding to the user also have Left-right symmetry. Based on this assumption, the high-order Ambisonic B-format symmetry optimization method can be used to significantly reduce the amount of computation and improve the computation efficiency.

下面描述了如何将音频对象编码到ambisonic域。The following describes how to encode an audio object into the ambisonic domain.

将音频对象编码到一阶ambisonic信号:Encode an audio object to a first-order ambisonic signal:

sisi是第i个音频对象,i=1..k,k是音频对象的个数。θiθi是平面上的角度(方位角),φiφi是垂直方向上的角度。W声道信号表示全方向声波,X声道信号、Y声道信号和Z声道信号分别表示沿空间三个互相垂直取向X、Y、Z的声波。s i s i is the i-th audio object, i=1..k, k is the number of audio objects. θ i θ i is the angle (azimuth angle) on the plane, and φ i φ i is the angle in the vertical direction. The W channel signal represents an omnidirectional sound wave, and the X channel signal, Y channel signal and Z channel signal respectively represent three mutually vertically oriented sound waves of X, Y, and Z along the space.

一阶Ambisonic B-格式信号表示为 A first-order ambisonic B-format signal is expressed as

同理,将音频对象编码到2阶或3阶Ambisonic B-格式信号优选依照下表定义进行:Likewise, encoding an audio object into a 2nd-order or 3rd-order Ambisonic B-format signal is preferably performed as defined in the following table:

上表中的三角函数对于方位角θ是偶函数的,则相应Ambisonic B-格式信号的相应分量是左右对称的,如果上表中的三角函数对于方位角θ是奇函数,则相应Ambisonic B-格式信号的相应分量是左右相反的。以一阶Ambisonic B-格式信号为例,从物理意义和坐标来看,w,x,z不分左右,所以如果听着的位置左右对称,并且假定相应的HRTF系数也近似左右对称,那么w,x,z对应的双耳输出的分量对于输出的左右通道是相同的。而y对于左右正好反向。所以y对应的双耳输出的分量对于左右通道是相反的。对于具有对称性的分量,可以采用快速算法,即运算过程中的对称性优化,可进一步降低运算量。The trigonometric function in the above table is an even function for the azimuth angle θ, then the corresponding component of the corresponding Ambisonic B-format signal is left-right symmetric, if the trigonometric function in the above table is an odd function for the azimuth angle θ, then the corresponding Ambisonic B- The corresponding components of the format signal are inverse left and right. Taking the first-order ambisonic B-format signal as an example, from the perspective of physical meaning and coordinates, w, x, and z are not divided into left and right, so if the listening position is left-right symmetrical, and assuming that the corresponding HRTF coefficients are also approximately left-right symmetrical, then w , the components of the binaural output corresponding to x, z are the same for the left and right channels of the output. And y is exactly the opposite for left and right. So the components of the binaural output corresponding to y are opposite for the left and right channels. For components with symmetry, a fast algorithm can be used, that is, symmetry optimization in the operation process, which can further reduce the amount of calculation.

另外,由于服务器对音频文件的处理可能存延迟,采取的解决方案是获取用户头部旋转的旋转角度,对所述旋转角度进行平滑处理。因此小的角度变化可不做新的旋转方向处理,有效解决了服务器处理的时延问题。In addition, since there may be a delay in the processing of the audio file by the server, the solution adopted is to obtain the rotation angle of the user's head rotation, and perform smoothing processing on the rotation angle. Therefore, small angle changes do not require new rotation direction processing, which effectively solves the delay problem of server processing.

实施例二:Embodiment two:

图2a-c描述基于云端多路音频传输用来提升沉浸式体验效果的实施例。需要注意的是本发明涵盖两种应用场景(1)音频实时通讯(会议场景),如图2b所示;(2)音频下载,如图2c所示;Figures 2a-c describe an embodiment of enhancing the effect of immersive experience based on cloud-based multi-channel audio transmission. It should be noted that the present invention covers two application scenarios (1) audio real-time communication (conference scenario), as shown in Figure 2b; (2) audio downloading, as shown in Figure 2c;

针对两种场景,输入有三种形式:单独的音频对象,声场输入(wxy形式),双耳录音信号。For the two scenarios, there are three forms of input: a separate audio object, sound field input (wxy form), and binaural recording signals.

如图2b所示,针对音频下载场景:As shown in Figure 2b, for the audio download scenario:

存储服务器存储有双耳录音信号,Ambisonic录音信号(声场信号),和/或音频对象,双耳转码服务器从存储服务器获取上述信号,在双耳转码服务器端将音频对象转成Ambisonic信号,例如,一阶水平方向B格式信号,即wxy,并与其他wxy信号(声场信号)相加。双耳转码服务器根据客户端头部跟踪装置传来的角度通过使用旋转矩阵对wxy信号进行旋转,将wxy信号转成双声道,再与双耳录音双声道信号叠加生成音频下载文件。通常需要压缩来降低传输带宽。然后客户端下载压缩后的双声道音频。这种做法会更加高效,但缺点是如果音频对象只用一阶B格式,空间定位的分辨率会有所下降,但是基于云服务的优选做法如果双耳化过程放置在客户端,则客户端从服务器下载wxy信号,则旋转操作无需经过服务器。The storage server stores binaural recording signals, ambisonic recording signals (sound field signals), and/or audio objects, and the binaural transcoding server obtains the above-mentioned signals from the storage server, and converts the audio objects into ambisonic signals on the binaural transcoding server side, For example, a first-order horizontal direction B-format signal, ie, wxy, is summed with other wxy signals (sound field signals). The binaural transcoding server rotates the wxy signal by using the rotation matrix according to the angle transmitted by the head tracking device of the client, converts the wxy signal into binaural, and then superimposes the binaural recording binaural signal to generate an audio download file. Compression is often required to reduce transmission bandwidth. Then the client downloads the compressed binaural audio. This approach will be more efficient, but the disadvantage is that if the audio object only uses the first-order B format, the resolution of spatial positioning will be reduced, but the preferred method based on cloud services If the binauralization process is placed on the client, the client If the wxy signal is downloaded from the server, the rotation operation does not need to go through the server.

如需要较高的空间精度,双耳转码服务器先将音频对象根据旋转角度进行旋转,将旋转后的音频对象信号编码为高阶B格式(例如三3阶),与其他B格式信号在双声道域叠加:经双耳转码后生成高阶B格式音频对象双声道信号,与Ambisonic录音双声道信号、双耳录音双声道信号进行叠加生成音频文件。If higher spatial precision is required, the binaural transcoding server first rotates the audio object according to the rotation angle, and encodes the rotated audio object signal into a high-order B format (such as three-order three), and other B-format signals in the binaural Channel domain superimposition: After binaural transcoding, a high-order B-format audio object binaural signal is generated, which is superimposed with the Ambisonic recording binaural signal and binaural recording binaural signal to generate an audio file.

在这里我们要注意的是头部跟踪只是一种形式,不排除其他动作参数,如挥手等。本发明同样适用。Here we should note that head tracking is only a form and does not exclude other motion parameters, such as waving, etc. The present invention is equally applicable.

如图2c所示,针对音频实时通讯(会议场景):As shown in Figure 2c, for audio real-time communication (conference scenario):

双耳转码服务器直接获取双耳录音麦克风阵列,Ambisonic麦克风阵列,单独声源或音频对象,在双耳转码服务器端将双耳录音麦克风阵列,Ambisonic麦克风阵列,单独声源或音频对象执行上述相似的处理过程。The binaural transcoding server directly obtains the binaural recording microphone array, ambisonic microphone array, individual sound source or audio object, and performs the above-mentioned operations on the binaural recording microphone array, ambisonic microphone array, individual sound source or audio object on the binaural transcoding server side similar process.

实施例三:Embodiment three:

如图3所示,一种云端音频处理服务器,获取单元,获取客户端中的头部跟踪装置传送的用户头部旋转的旋转角度;采集单元,分别采集双耳录音信号,Ambisonic录音信号,音频对象;双耳转码单元,分别与所述获取单元和采集单元相连接,根据所述旋转角度,分别对所述不同格式的音频信号进行双耳转码,其中对于双耳录音信号,根据所述旋转角度进行插值,生成双耳录音双声道信号;而如需要较高的空间精度的情况下,双耳转码单元将音频对象信号根据旋转角度进行旋转,将旋转后的音频对象信号编码为高阶B格式音频对象信号,经双耳转码后生成高阶B格式音频对象双声道信号,叠加单元对双耳转码单元生成的高阶B格式音频对象双声道信号,Ambisonic录音双声道信号、双耳录音双声道信号进行叠加;As shown in Figure 3, a cloud audio processing server, the acquisition unit acquires the rotation angle of the user's head rotation transmitted by the head tracking device in the client; the acquisition unit collects binaural recording signals, ambisonic recording signals, audio Object: a binaural transcoding unit, which is respectively connected to the acquisition unit and the acquisition unit, and performs binaural transcoding on the audio signals of different formats according to the rotation angle, wherein for the binaural recording signal, according to the The above rotation angle is interpolated to generate a binaural recording binaural signal; and if higher spatial accuracy is required, the binaural transcoding unit rotates the audio object signal according to the rotation angle, and encodes the rotated audio object signal It is a high-order B-format audio object signal. After binaural transcoding, a high-order B-format audio object binaural signal is generated. The superposition unit generates a high-order B-format audio object binaural signal generated by the binaural transcoding unit. Ambisonic recording Superposition of binaural signals and binaural recording binaural signals;

如需要低复杂度低延迟,双耳转码单元将音频对象信号编码为一阶B格式音频对象信号,与其他一阶Ambisonic录音信号叠加,然后根据旋转角度对所述叠加后的混合信号进行双耳转码,生成音频对象与Ambisonic录音信号的混合双声道信号,叠加单元对双耳转码单元生成的与所述混合双声道信号、双耳录音双声道信号进行叠加,得到音频双耳输出虚拟环绕声信号。If low complexity and low delay are required, the binaural transcoding unit encodes the audio object signal into a first-order B-format audio object signal, superimposes it with other first-order ambisonic recording signals, and then double-binds the superimposed mixed signal according to the rotation angle. Ear transcoding generates a mixed binaural signal of an audio object and an ambisonic recording signal, and the superposition unit superimposes the mixed binaural signal and the binaural recording binaural signal generated by the binaural transcoding unit to obtain an audio binaural The ear outputs a virtual surround sound signal.

本实施例利用云端服务器来解决支持头部跟踪的多声道音频传输和播放的问题。In this embodiment, a cloud server is used to solve the problem of multi-channel audio transmission and playback supporting head tracking.

实施例四:Embodiment four:

如图4所示,本发明的一种音频处理系统主要包含客户端,存储服务器,云端音频处理服务器;客户端包括头部跟踪模块,存储服务器端存有多声道音频文件,以特定方式存放。客户端头部跟踪模块获取用户头部动作如头部旋转角度,将参数经互联网上传到服务器端的一台或多台云端音频处理服务器,对多声道音频文件进行相应处理:云端音频处理服务器从存储服务器提取不同格式的音频信号,并根据接收的旋转角度生成音频双耳输出虚拟环绕声信号,将经过双耳转码后的音频文件通过所述网络传输至客户端。As shown in Figure 4, an audio processing system of the present invention mainly includes a client, a storage server, and a cloud audio processing server; the client includes a head tracking module, and the storage server stores multi-channel audio files and stores them in a specific way . The client head tracking module obtains the user's head movement such as the head rotation angle, and uploads the parameters to one or more cloud audio processing servers on the server side via the Internet, and performs corresponding processing on the multi-channel audio files: the cloud audio processing server from The storage server extracts audio signals in different formats, generates audio binaural output virtual surround sound signals according to the received rotation angle, and transmits the binaurally transcoded audio files to the client through the network.

客户端下载上述处理后的音频文件,优选的,以双声道立体声格式播放。The client downloads the above-mentioned processed audio file, and preferably plays it in a two-channel stereo format.

以上结合附图详细描述了本发明的优选实施方式,但是,本发明并不限于上述实施方式中的具体细节,在本发明的技术构思范围内,可以对本发明的技术方案进行多种简单变型,这些简单变型均属于本发明的保护范围。The preferred embodiment of the present invention has been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the specific details of the above embodiment, within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solution of the present invention, These simple modifications all belong to the protection scope of the present invention.

另外需要说明的是,在上述具体实施方式中所描述的各个具体技术特征,在不矛盾的情况下,可以通过任何合适的方式进行组合。为了避免不必要的重复,本发明对各种可能的组合方式不再另行说明。In addition, it should be noted that the various specific technical features described in the above specific implementation manners may be combined in any suitable manner if there is no contradiction. In order to avoid unnecessary repetition, various possible combinations are not further described in the present invention.

此外,本发明的各种不同的实施方式之间也可以进行任意组合,只要其不违背本发明的思想,其同样应当视为本发明所公开的内容。In addition, various combinations of different embodiments of the present invention can also be combined arbitrarily, as long as they do not violate the idea of the present invention, they should also be regarded as the disclosed content of the present invention.

Claims (10)

1.一种云端音频处理方法,其特征在于:所述音频处理方法包括以下步骤,1. A cloud audio processing method, characterized in that: the audio processing method comprises the following steps, 获取用户头部旋转的旋转角度;Obtain the rotation angle of the user's head rotation; 获取不同格式的音频信号,根据所述旋转角度,分别对所述不同格式的音频信号进行双耳转码,生成相应格式的双声道音频信号;Acquiring audio signals in different formats, performing binaural transcoding on the audio signals in different formats according to the rotation angle, to generate binaural audio signals in corresponding formats; 对所述相应格式的双声道信号叠加,得到音频双耳输出虚拟环绕声信号;Superimposing the binaural signals of the corresponding formats to obtain audio binaural output virtual surround sound signals; 所获取用户头部旋转的旋转角度具体为获取用户头部旋转的旋转角度,对所述旋转角度进行平滑处理。The acquired rotation angle of the user's head rotation is specifically acquiring the rotation angle of the user's head rotation, and smoothing the rotation angle. 2.根据权利要求1所述的云端音频处理方法,其特征在于:2. The cloud audio processing method according to claim 1, characterized in that: 所述不同格式的音频信号包括双耳录音信号,Ambisonic录音信号和音频对象信号。The audio signals in different formats include binaural recording signals, ambisonic recording signals and audio object signals. 3.根据权利要求2所述的云端音频处理方法,其特征在于:3. The cloud audio processing method according to claim 2, characterized in that: 对所述不同格式的音频信号进行双耳转码,生成相应格式的双耳转码音频信号具体包括:Performing binaural transcoding on the audio signals of different formats, and generating binaural transcoding audio signals of corresponding formats specifically includes: 对所述双耳录音信号,根据所述旋转角度进行插值,生成双耳录音双声道信号;For the binaural recording signal, interpolation is performed according to the rotation angle to generate a binaural recording binaural signal; 对所述Ambisonic录音信号,根据所述旋转角度对所述Ambisonic录音信号进行调整,对所述调整后的Ambisonic录音信号双耳转码生成Ambisonic录音双声道信号;For the ambisonic recording signal, the ambisonic recording signal is adjusted according to the rotation angle, and the adjusted ambisonic recording signal is binaurally transcoded to generate an ambisonic recording binaural signal; 对所述音频对象信号,根据所述旋转角度对所述音频对象信号调整,对所述调整后的音频对象信号双耳转码生成音频对象双声道信号。For the audio object signal, the audio object signal is adjusted according to the rotation angle, and the adjusted audio object signal is binaurally transcoded to generate an audio object binaural signal. 4.根据权利要求3所述的云端音频处理方法,其特征在于:4. The cloud audio processing method according to claim 3, characterized in that: 如需要较高的空间精度,将音频对象信号根据旋转角度进行旋转,将旋转后的音频对象信号编码为高阶B格式音频对象信号,经双耳转码后生成高阶B格式音频对象双声道信号,与Ambisonic录音双声道信号、双耳录音双声道信号进行叠加;If higher spatial precision is required, the audio object signal is rotated according to the rotation angle, and the rotated audio object signal is encoded into a high-order B-format audio object signal, and the high-order B-format audio object binaural is generated after binaural transcoding channel signal, superimposed with ambisonic recording binaural signal and binaural recording binaural signal; 如需要低复杂度低延迟,将音频对象信号编码为一阶B格式音频对象信号,与其他一阶Ambisonic录音信号叠加,然后根据旋转角度对所述叠加后的混合信号进行双耳转码,生成音频对象与Ambisonic录音信号的混合双声道信号,与所述双耳录音双声道信号进行叠加。If low complexity and low delay are required, the audio object signal is encoded into a first-order B-format audio object signal, superimposed with other first-order ambisonic recording signals, and then the superimposed mixed signal is binaurally transcoded according to the rotation angle to generate The mixed binaural signal of the audio object and the ambisonic recording signal is superimposed on the binaural signal of the binaural recording. 5.一种云端音频处理服务器,其特征在于,所述服务器包括:5. A cloud audio processing server, characterized in that the server comprises: 获取单元,获取用户头部旋转的旋转角度;The acquisition unit acquires the rotation angle of the user's head rotation; 采集单元,采集不同格式的音频信号;The acquisition unit collects audio signals in different formats; 双耳转码单元,分别与所述获取单元和采集单元相连接,根据所述旋转角度,分别对所述不同格式的音频信号进行双耳转码,生成相应格式的双声道音频信号;The binaural transcoding unit is respectively connected to the acquisition unit and the acquisition unit, and according to the rotation angle, performs binaural transcoding on the audio signals of different formats to generate a binaural audio signal of a corresponding format; 叠加单元,与所述双耳转码单元连接,对所述相应格式的双声道信号叠加,得到音频双耳输出虚拟环绕声信号;A superposition unit, connected to the binaural transcoding unit, superimposes the binaural signals of the corresponding format to obtain audio binaural output virtual surround sound signals; 所述云端音频处理服务器还包括平滑单元,分别与所述双耳转码单元和所述获取单元连接,平滑单元从获取单元接收用户头部旋转的旋转角度,对所述旋转角度进行平滑处理。The cloud audio processing server further includes a smoothing unit connected to the binaural transcoding unit and the acquisition unit respectively, the smoothing unit receives the rotation angle of the user's head rotation from the acquisition unit, and performs smoothing processing on the rotation angle. 6.根据权利要求5所述的云端音频处理服务器,其特征在于:6. The cloud audio processing server according to claim 5, characterized in that: 所述不同格式的音频信号包括双耳录音信号,Ambisonic录音信号和音频对象信号。The audio signals in different formats include binaural recording signals, ambisonic recording signals and audio object signals. 7.根据权利要求6所述的云端音频处理服务器,其特征在于:7. The cloud audio processing server according to claim 6, characterized in that: 双耳转码单元对所述不同格式的音频信号进行双耳转码,生成相应格式的双耳转码音频信号具体包括:The binaural transcoding unit performs binaural transcoding on the audio signals of different formats, and generates binaural transcoding audio signals of corresponding formats specifically including: 对所述双耳录音信号,根据所述旋转角度进行插值,生成双耳录音双声道信号;For the binaural recording signal, interpolation is performed according to the rotation angle to generate a binaural recording binaural signal; 对所述Ambisonic录音信号,根据所述旋转角度对所述Ambisonic录音信号进行调整,对所述调整后的Ambisonic录音信号双耳转码生成Ambisonic录音双声道信号;For the ambisonic recording signal, the ambisonic recording signal is adjusted according to the rotation angle, and the adjusted ambisonic recording signal is binaurally transcoded to generate an ambisonic recording binaural signal; 对所述音频对象信号,根据所述旋转角度对所述音频对象信号调整,对所述调整后的音频对象信号双耳转码生成音频对象双声道信号。For the audio object signal, the audio object signal is adjusted according to the rotation angle, and the adjusted audio object signal is binaurally transcoded to generate an audio object binaural signal. 8.根据权利要求7所述的云端音频处理服务器,其特征在于:8. The cloud audio processing server according to claim 7, characterized in that: 如需要较高的空间精度,双耳转码单元将音频对象信号根据旋转角度进行旋转,将旋转后的音频对象信号编码为高阶B格式音频对象信号,经双耳转码后生成高阶B格式音频对象双声道信号,叠加单元对双耳转码单元生成的高阶B格式音频对象双声道信号,Ambisonic录音双声道信号、双耳录音双声道信号进行叠加;If higher spatial precision is required, the binaural transcoding unit rotates the audio object signal according to the rotation angle, encodes the rotated audio object signal into a high-order B format audio object signal, and generates a high-order B format after binaural transcoding Format audio object binaural signal, the superimposition unit superimposes the high-order B-format audio object binaural signal generated by the binaural transcoding unit, the ambisonic recording binaural signal, and the binaural recording binaural signal; 如需要低复杂度低延迟,双耳转码单元将音频对象信号编码为一阶B格式音频对象信号,与其他一阶Ambisonic录音信号叠加,然后根据旋转角度对所述叠加后的混合信号进行双耳转码,生成音频对象与Ambisonic录音信号的混合双声道信号,叠加单元对双耳转码单元生成的与所述混合双声道信号、双耳录音双声道信号进行叠加。If low complexity and low delay are required, the binaural transcoding unit encodes the audio object signal into a first-order B-format audio object signal, superimposes it with other first-order ambisonic recording signals, and then double-binds the superimposed mixed signal according to the rotation angle. Ear transcoding generates a mixed binaural signal of an audio object and an ambisonic recording signal, and the superposition unit superimposes the mixed binaural signal and the binaural recording binaural signal generated by the binaural transcoding unit. 9.一种音频播放系统,其特征在于:所述系统包括权利要求5-8任一所述云端音频处理服务器,以及客户端;所述客户端包括头部跟踪装置,所述头部跟踪装置抓取头部旋转角度,通过网络上传至所述云端音频处理服务器,所述云端音频处理器获取不同格式的音频信号,并根据所述旋转角度生成音频双耳输出虚拟环绕声信号后,通过所述网络传输至客户端。9. An audio playback system, characterized in that: the system includes the cloud audio processing server described in any one of claims 5-8, and a client; the client includes a head tracking device, and the head tracking device Grab the head rotation angle, upload it to the cloud audio processing server through the network, the cloud audio processor obtains audio signals in different formats, and generates audio binaural output virtual surround sound signals according to the rotation angle, and passes through the The above network is transmitted to the client. 10.根据权利要求9所述的音频播放系统,其特征在于:所述系统还包括存储服务器,存储不同格式的音频信号,当用户请求下载音频文件下载时,所述云端音频处理服务器从所述存储服务器提取所述音频信号。10. The audio playback system according to claim 9, characterized in that: the system also includes a storage server that stores audio signals in different formats, and when a user requests to download an audio file to download, the cloud audio processing server starts from the A storage server retrieves the audio signal.
CN201610017000.1A 2016-01-11 2016-01-11 A kind of audio-frequency processing method and system Active CN105682000B (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
CN201610017000.1A CN105682000B (en) 2016-01-11 2016-01-11 A kind of audio-frequency processing method and system

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
CN201610017000.1A CN105682000B (en) 2016-01-11 2016-01-11 A kind of audio-frequency processing method and system

Publications (2)

Publication Number Publication Date
CN105682000A CN105682000A (en) 2016-06-15
CN105682000B true CN105682000B (en) 2017-11-07

Family

ID=56300173

Family Applications (1)

Application Number Title Priority Date Filing Date
CN201610017000.1A Active CN105682000B (en) 2016-01-11 2016-01-11 A kind of audio-frequency processing method and system

Country Status (1)

Country Link
CN (1) CN105682000B (en)

Families Citing this family (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN106210990B (en) * 2016-07-13 2018-08-10 北京时代拓灵科技有限公司 A kind of panorama sound audio processing method
CN106851482A (en) * 2017-03-24 2017-06-13 北京时代拓灵科技有限公司 A kind of panorama sound loudspeaker body-sensing real-time interaction system and exchange method
CN108877815B (en) 2017-05-16 2021-02-23 华为技术有限公司 Stereo signal processing method and device
CN108495234B (en) * 2018-04-19 2020-01-07 北京微播视界科技有限公司 Multi-channel audio processing method, apparatus and computer-readable storage medium
CN113709654B (en) * 2021-08-27 2024-08-06 维沃移动通信(杭州)有限公司 Recording method, recording device, recording apparatus, and readable storage medium

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105376690A (en) * 2015-11-04 2016-03-02 北京时代拓灵科技有限公司 Method and device of generating virtual surround sound

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPS5442102A (en) * 1977-09-10 1979-04-03 Victor Co Of Japan Ltd Stereo reproduction system
US7876903B2 (en) * 2006-07-07 2011-01-25 Harris Corporation Method and apparatus for creating a multi-dimensional communication space for use in a binaural audio system
WO2009125046A1 (en) * 2008-04-11 2009-10-15 Nokia Corporation Processing of signals
EP2428813B1 (en) * 2010-09-08 2014-02-26 Harman Becker Automotive Systems GmbH Head Tracking System with Improved Detection of Head Rotation
CN105120421B (en) * 2015-08-21 2017-06-30 北京时代拓灵科技有限公司 A kind of method and apparatus for generating virtual surround sound

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105376690A (en) * 2015-11-04 2016-03-02 北京时代拓灵科技有限公司 Method and device of generating virtual surround sound

Also Published As

Publication number Publication date
CN105682000A (en) 2016-06-15

Similar Documents

Publication Publication Date Title
CN111466124B (en) Method, processor system and computer readable medium for rendering an audiovisual recording of a user
Rafaely et al. Spatial audio signal processing for binaural reproduction of recorded acoustic scenes–review and challenges
TWI713866B (en) Apparatus and method for generating an enhanced sound field description, computer program and storage medium
JP7038725B2 (en) Audio signal processing method and equipment
US10349197B2 (en) Method and device for generating and playing back audio signal
CN105872940B (en) A kind of virtual reality sound field generation method and system
Davis et al. High order spatial audio capture and its binaural head-tracked playback over headphones with HRTF cues
CN107533843B (en) System and method for capturing, encoding, distributing and decoding immersive audio
CN114067810B (en) Audio signal rendering method and apparatus
US20150189455A1 (en) Transformation of multiple sound fields to generate a transformed reproduced sound field including modified reproductions of the multiple sound fields
CN109410912B (en) Audio processing method and device, electronic equipment and computer readable storage medium
CN106210990B (en) A kind of panorama sound audio processing method
CN108476367A (en) The synthesis of signal for immersion audio playback
CN105682000A (en) Audio processing method and system
WO2018026963A1 (en) Head-trackable spatial audio for headphones and system and method for head-trackable spatial audio for headphones
CN106331977B (en) A kind of virtual reality panorama acoustic processing method of network K songs
WO2022262758A1 (en) Audio rendering system and method and electronic device
Sun Immersive audio, capture, transport, and rendering: A review
WO2022262750A1 (en) Audio rendering system and method, and electronic device
Enzner et al. Advanced system options for binaural rendering of ambisonic format
US20240312468A1 (en) Spatial Audio Upscaling Using Machine Learning
CN113347530A (en) Panoramic audio processing method for panoramic camera
Suzuki et al. 3D spatial sound systems compatible with human's active listening to realize rich high-level kansei information
CN113674751A (en) Audio processing method and device, electronic equipment and storage medium
Salvador et al. Enhancement of Spatial Sound Recordings by Adding Virtual Microphones to Spherical Microphone Arrays.

Legal Events

Date Code Title Description
C06 Publication
PB01 Publication
C10 Entry into substantive examination
SE01 Entry into force of request for substantive examination
GR01 Patent grant
GR01 Patent grant
TR01 Transfer of patent right
TR01 Transfer of patent right

Effective date of registration: 20251218

Address after: 100086 Beijing city Haidian District Shuangyushu Academy Road No. 44

Patentee after: China Film Science and Technology Research Institute (Film Technology Quality Inspection Institute of the Central Propaganda Department)

Country or region after: China

Address before: Room 0014-32, floor 01, No. 26, Shangdi Information Road, Haidian District, Beijing 100085

Patentee before: BEIJING TUOLING Inc.

Country or region before: China