EP4648441A1 - Procédé et appareil de rendu audio spatial - Google Patents

Procédé et appareil de rendu audio spatial

Info

Publication number
EP4648441A1
EP4648441A1 EP23919449.1A EP23919449A EP4648441A1 EP 4648441 A1 EP4648441 A1 EP 4648441A1 EP 23919449 A EP23919449 A EP 23919449A EP 4648441 A1 EP4648441 A1 EP 4648441A1
Authority
EP
European Patent Office
Prior art keywords
head movement
audio
movement position
rendering
historical
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23919449.1A
Other languages
German (de)
English (en)
Other versions
EP4648441A4 (fr
Inventor
Peng Qin
Yiwei KOU
Yu Yang
Ke Jia
Yuanpeng LIN
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Huawei Technologies Co Ltd
Original Assignee
Huawei Technologies Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Huawei Technologies Co Ltd filed Critical Huawei Technologies Co Ltd
Publication of EP4648441A1 publication Critical patent/EP4648441A1/fr
Publication of EP4648441A4 publication Critical patent/EP4648441A4/fr
Pending legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/302Electronic adaptation of stereophonic sound system to listener position or orientation
    • H04S7/303Tracking of listener position or orientation
    • H04S7/304For headphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/305Electronic adaptation of stereophonic audio signals to reverberation of the listening space
    • H04S7/306For headphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S2400/00Details of stereophonic systems covered by H04S but not provided for in its groups
    • H04S2400/11Positioning of individual sound objects, e.g. moving airplane, within a sound field
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S2420/00Techniques used stereophonic systems covered by H04S but not provided for in its groups
    • H04S2420/01Enhancing the perception of the sound image or of the spatial distribution using head related transfer functions [HRTF's] or equivalents thereof, e.g. interaural time difference [ITD] or interaural level difference [ILD]

Definitions

  • This application relates to audio processing technologies, and in particular, to a spatial audio rendering method and apparatus.
  • sound may have two characteristics: a sense of orientation and a sense of space.
  • a spatial audio technology is usually a technology that simulates immersive audio experience with the foregoing two characteristics on a head-mounted play device such as a headset.
  • People's judgment on a sense of orientation for sound is mainly affected by a time difference, a sound level difference, human body filter effect learned during growth, head shaking, and other factors.
  • the time difference, the sound level difference, and the human body filter effect may be comprehensively expressed as a head-related transfer function (Head-Related Transfer Function, HRTF).
  • HRTF head-Related Transfer Function
  • the head shaking in all directions is of great help for determining a position of a sound source.
  • An indoor sound field may include direct sound, early reflected sound (Early Reflection, ER), and late reverberant sound (Late Reverb, LR). People's sense of space for sound is built based on the ER and the LR.
  • Sound sources in different formats such as stereo, multi-channel surround sound, and a specially produced multi-object sound source, may be rendered into a binaural signal by using the spatial audio technology, so that real and immersive auditory experience can be enjoyed through a headset.
  • head tracking may be performed through a gyroscope or another sensor in a headset, and then a corresponding rotation or displacement change is incorporated in spatial rendering of a sound source.
  • spatial audio rendering usually needs to be supported by high computing power, and a low delay is needed for head movement tracking. If spatial audio rendering is performed at an audio content provider end (for example, a mobile phone, a tablet computer, or a PC), although a computing power requirement can be met, a long delay occurs, and overall auditory experience is affected, especially in a head movement scenario.
  • an audio content provider end for example, a mobile phone, a tablet computer, or a PC
  • This application provides a spatial audio rendering method and apparatus, to eliminate a delay, so that rendering effect exactly matches a head movement position of a user, to achieve real and immersive auditory experience.
  • this application provides a spatial audio rendering method, including: obtaining delay information of an audio content provider end; obtaining a predicted head movement position based on the delay information; and sending first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position.
  • a head movement position of a user is predicted to obtain a predicted head movement position, and after the predicted head movement position is sent to the audio content provider end, the audio content provider end may render an audio frame based on the predicted head movement position.
  • an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position.
  • predicted time may offset a delay caused by rendering, link transmission, and the like. In this way, when the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.
  • the delay information indicates a delay status of the audio content provider end, and may include the following cases.
  • the delay information includes a delay sent by the audio content provider end.
  • the audio content provider end may package the delay into first information to be sent to the audio signal play end, and send the first information to the audio signal play end. In this way, the audio signal play end can directly extract the delay of the audio content provider end from the first information that comes from the audio content provider end.
  • the delay information includes a second historical head movement position sent by the audio content provider end.
  • the audio signal play end (for example, a headset) may periodically detect, through a sensor (for example, a gyroscope or a gravity sensor), a current position (referred to as a measured head movement position in this specification) of a head of a user wearing the headset. In this way, the audio signal play end can send the measured head movement position obtained through measurement to the audio content provider end.
  • a sensor for example, a gyroscope or a gravity sensor
  • a current position referred to as a measured head movement position in this specification
  • the audio content provider end After receiving the measured head movement position, the audio content provider end does not process the measured head movement position, but still performs related processing on an audio frame according to a specified process.
  • the measured head movement position is carried in the first information. In this case, because a period of time has elapsed, the measured head movement position becomes a historical head movement position (referred to as the second historical head movement position in this specification). It should be understood that, in a mathematical sense, the measured head movement position is the same as the second historical head movement position.
  • the audio signal play end may obtain, based on time at which the second historical head movement position is received and time at which the measured head movement position is sent, a time difference between the time at which the second historical head movement position is received and the time at which the measured head movement position is sent (that is, calculate a time difference between receiving time and sending time of a same head movement position), to indirectly obtain the delay of the audio content provider end.
  • the delay information includes a second historical head movement position and a first uncertainty coefficient that are sent by the audio content provider end.
  • the first uncertainty coefficient is an important parameter in a Kalman filter, and helps improve accuracy of a prediction result.
  • sound sources in different formats such as stereo, multi-channel surround sound, and a specially produced multi-object sound source, may be rendered into a binaural signal by using a spatial audio technology.
  • a head of a user rotates or moves, an absolute position of a sound source does not change, but a relative direction between the sound source and the head changes. For example, a guitar is being played in front of a user, and if the user turns to the right, sound of the guitar is correspondingly shifted to the left of the user.
  • the audio content provider end may provide high computing power to implement the foregoing rendering algorithm with head movement effect
  • a long delay occurs, causing disorder of important spatial information such as an orientation.
  • an acoustic image of an object is designed to be directly in front of space. If a head of a user rotates, the acoustic image first moves to be directly in front of a face of the user, and then is restored to be directly in front of space. A clear delay and "sense of damping" occur, affecting auditory experience.
  • the audio signal play end may predict a head movement position of a user, and then send the predicted head movement position to the audio content provider end, so that the audio content provider end can render an audio frame based on the predicted head movement position.
  • an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position.
  • predicted time may offset a delay caused by rendering, link transmission, and the like. In this way, when the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.
  • the audio signal play end may first obtain a first historical head movement position corresponding to the delay information, and then obtain a measured head movement position through the sensor, and then obtain the predicted head movement position based on the first historical head movement position and the measured head movement position.
  • the audio signal play end may obtain the first historical head movement position by using three methods.
  • the delay information includes the delay sent by the audio content provider end (the first case)
  • the first historical head movement position corresponding to the delay is extracted from a cache, where a correspondence between a plurality of delays and a plurality of historical head movement positions is prestored in the cache.
  • the audio signal play end may periodically detect, through the sensor, a measured head movement position of a user wearing the audio signal play end. For ease of subsequent use, a correspondence between a measured head movement position and detection time may be stored on the audio signal play end. In this way, when the foregoing delay is obtained, a current measured head movement position (referred to as the first historical head movement position in this specification) corresponding to the delay may be obtained by querying the correspondence.
  • the delay information includes the second historical head movement position sent by the audio content provider end (the second case)
  • the second historical head movement position is used as the first historical head movement position.
  • the audio content provider end directly packages the second historical head movement position into the first information and sends the first information to the audio signal play end.
  • the audio signal play end can directly obtain a current measured head movement position (the second historical head movement position) corresponding to the delay. Therefore, the audio signal play end can directly determine the second historical head movement position as the first historical head movement position.
  • the delay information includes the second historical head movement position and the first uncertainty coefficient that are sent by the audio content provider end (the third case)
  • the second historical head movement position is used as the first historical head movement position.
  • the audio signal play end may alternatively directly determine the second historical head movement position as the first historical head movement position.
  • the audio signal play end may obtain a measured head movement position through the sensor, where the measured head movement position is a newly detected head movement position of the user.
  • obtaining the predicted head movement position based on the first historical head movement position and the measured head movement position may include the following two algorithms.
  • the audio signal play end may obtain a difference between the first historical head movement position and the measured head movement position; obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; and obtain the predicted head movement position based on the difference and the head movement change rate.
  • the difference between the first historical head movement position and the measured head movement position may be a distance between the first historical head movement position and the measured head movement position.
  • a head movement position may be represented by using coordinate axes x, y, and z. Therefore, the foregoing difference may be a distance between a coordinate value of the first historical head movement position and a coordinate value of the measured head movement position.
  • a head movement position may alternatively be represented in another manner. This is not specifically limited in this application.
  • the audio signal play end may periodically detect, through the sensor, a measured head movement position of a user wearing the audio signal play end. Therefore, the audio signal play end may obtain a head movement change rate based on N measured head movement positions that are previously obtained, where the N measured head movement positions may include N measured head movement positions obtained through counting forward from a current measured head movement position.
  • the audio signal play end may predict a head movement change value based on the difference and the head movement change rate through cubic spline interpolation or by another means, and then add up the head movement change value and a current measured head movement position to obtain a predicted head movement position.
  • the audio signal play end obtains a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; obtains a second uncertainty coefficient based on the head movement change rate and the first uncertainty coefficient; and inputs the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position to a Kalman filter to obtain the predicted head movement position.
  • the second uncertainty coefficient may be obtained based on the head movement change rate and the first uncertainty coefficient.
  • the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position are input to the Kalman filter.
  • a Kalman gain coefficient may be determined based on the second uncertainty coefficient and estimated uncertainty of a previous iteration, and the Kalman gain coefficient is used as a latest gain coefficient. In this way, the Kalman filter can output the predicted head movement position.
  • the predicted head movement position may alternatively be obtained based on the first historical head movement position and the measured head movement position by using another algorithm. This is not specifically limited herein.
  • the audio signal play end sends the first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position, so that the audio content provider end can perform rendering with head movement effect on an audio frame based on the predicted head movement position.
  • the audio content provider end can perform rendering with head movement effect on an audio frame based on the predicted head movement position.
  • an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position.
  • predicted time may offset a delay caused by rendering, link transmission, and the like.
  • the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.
  • the audio content provider end when a first rendering mode is used, the audio content provider end obtains a current frame.
  • an audio format of an audio source is a non-stereo format
  • the audio content provider end obtains a predicted head movement position.
  • the audio content provider end renders the current frame based on the predicted head movement position to obtain a first binaural signal.
  • an audio format of an audio source is a stereo format
  • the audio content provider end does not perform spatial rendering on the audio source.
  • the audio content provider end sends first information to the audio signal play end, where the first information includes a first audio stream, the audio format of the audio source, and the delay information.
  • the audio format indicates that the audio source is in the stereo format
  • the audio signal play end obtains a measured head movement position through the sensor.
  • the audio signal play end renders the current frame based on the measured head movement position to obtain a second binaural signal.
  • the audio signal play end plays audio based on a target binaural signal, where the target binaural signal includes the first binaural signal or the second binaural signal.
  • the first rendering mode may be a non-low-delay link mode.
  • the audio content provider end may perform rendering with head movement effect on an audio frame based on the predicted head movement position.
  • a user may perform an operation on an application (application, APP) installed on the audio content provider end, to choose whether to select the first rendering mode.
  • this application further includes a second rendering mode.
  • the audio content provider end performs rendering without head movement effect on an audio frame.
  • the user may perform an operation on the APP installed on the audio content provider end, to choose whether to select the second rendering mode.
  • the operation of the user refer to the following embodiments.
  • the current frame may be a frame of the audio source, and usually, may be an audio frame that is currently being processed by the audio content provider end.
  • the audio format of the audio source includes the stereo format or the non-stereo format, where the non-stereo format may include but is not limited to audio in a multi-channel or multi-object format such as 5.1, 7.1, or 3DA.
  • the audio content provider end may perform high-computing-power rendering on the audio in the multi-channel or multi-object format (non-stereo format) such as 5.1, 7.1, or 3DA, so that computing power advantages of the audio content provider end can be fully utilized.
  • the audio content provider end may perform rendering with head movement effect on an audio frame. Therefore, the predicted head movement position needs to be obtained.
  • the audio content provider end may obtain the predicted head movement position from rendering information that is previously received from the audio signal play end.
  • the rendering information may be latest received rendering information. In this way, accuracy of the predicted head movement position can be improved.
  • For obtaining of the predicted head movement position refer to an embodiment shown in FIG. 3 . Details are not described herein.
  • the audio content provider end may separately perform rendering with head movement effect on a direct sound part, an early reflected sound (Early Reflection, ER) part, and a late reverberant sound (Late Reverb, LR) part in the current frame based on the predicted head movement position, to obtain the first binaural signal.
  • ER early reflected sound
  • LR late reverberant sound
  • direct sound with head movement effect and reverberation (obtained by mixing ER and LR) with head movement effect are separately rendered by an audio effector and then mixed, and an output result obtained through mixing is rendered by the audio effector again and then transmitted in a form of the first binaural signal.
  • the audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.
  • the audio content provider end does not need to render an audio source in a stereo format, and the audio signal play end performs rendering with head movement effect on the audio source. Therefore, the audio content provider end may perform high-computing-power rendering on the audio in the multi-channel or multi-object format (non-stereo format) such as 5.1, 7.1, or 3DA, so that computing power advantages of the audio content provider end can be fully utilized. Based on this, when the audio format of the audio source is the non-stereo format, the audio content provider end may perform rendering with head movement effect on an audio frame. Therefore, the audio content provider end only needs to package the audio source into the first audio stream.
  • the first audio stream includes the first binaural signal rendered by the audio content provider end; or when the audio format indicates that the audio source is in the stereo format, the first audio stream includes the audio source.
  • the audio signal play end For an audio source in a stereo format, the audio signal play end performs rendering with a head movement position on the audio source. Because an IMU data transmission link of the audio signal play end has a low delay, the audio signal play end can obtain a current head movement position (referred to as the measured head movement position in this specification) in real time through the sensor.
  • the audio signal play end may perform rendering with head movement effect on the direct sound part in the current frame based on the measured head movement position, and perform rendering without head movement effect on the ER and the LR in the current frame, to obtain the second binaural signal.
  • direct sound with head movement effect and reverberation (obtained by mixing ER and LR) without head movement effect are separately rendered by an audio effector and then mixed, and an output result obtained through mixing is rendered by the audio effector again and then mixed again, to obtain the second binaural signal.
  • the audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.
  • the audio signal play end obtains the target binaural signal through processing in the foregoing steps, where the target binaural signal may be the first binaural signal (the audio source is in the non-stereo format) or the second binaural signal (the audio source is in the stereo format).
  • the target binaural signal may be the first binaural signal (the audio source is in the non-stereo format) or the second binaural signal (the audio source is in the stereo format).
  • the audio content provider end when the second rendering mode is used, obtains a current frame, performs rendering without head movement effect on the current frame to obtain a second binaural signal (the second binaural signal is different from the foregoing second binaural signal, and is referred to as a third binaural signal below for differentiation), and sends second information to the audio signal play end, where the second information includes a second audio stream, and the second audio stream includes the third binaural signal.
  • the audio signal play end obtains a measured head movement position through the sensor, and renders a current frame (the current frame is an audio frame obtained by the audio content provider end by performing rendering without head movement effect) based on the measured head movement position to obtain a fourth binaural signal.
  • the audio signal play end plays audio based on a target binaural signal, where the target binaural signal includes the fourth binaural signal.
  • the audio content provider end performs rendering without head movement effect on an audio frame, so that computing power advantages of the audio content provider end can be fully utilized.
  • a rendering delay can be shortened without head movement effect.
  • the audio signal play end performs low-computing-power rendering with head movement effect on the audio frame. Compared with the first rendering mode, this achieves a lower delay, but rendering effect of head movement effect is poor.
  • this application provides a spatial audio rendering apparatus, including: an obtaining module, configured to: obtain delay information of an audio content provider end; a prediction module, configured to: obtain a predicted head movement position based on the delay information; and a sending module, configured to: send first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position.
  • the prediction module is specifically configured to: obtain a first historical head movement position corresponding to the delay information; obtain a measured head movement position through a sensor; and obtain the predicted head movement position based on the first historical head movement position and the measured head movement position.
  • the delay information includes a delay sent by the audio content provider end; and the prediction module is specifically configured to: extract, from a cache, the first historical head movement position corresponding to the delay, where a correspondence between a plurality of delays and a plurality of historical head movement positions is prestored in the cache.
  • the delay information includes a second historical head movement position sent by the audio content provider end, the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, and the second rendering information is sent earlier than the first rendering information; and the prediction module is specifically configured to: use the second historical head movement position as the first historical head movement position.
  • the prediction module is specifically configured to: obtain a difference between the first historical head movement position and the measured head movement position; obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; and obtain the predicted head movement position based on the difference and the head movement change rate.
  • the delay information includes a second historical head movement position and a first uncertainty coefficient that are sent by the audio content provider end
  • the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, the first uncertainty coefficient comes from the second rendering information, and the second rendering information is sent earlier than the first rendering information
  • the prediction module is specifically configured to: use the second historical head movement position as the first historical head movement position.
  • the prediction module is specifically configured to: obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; obtain a second uncertainty coefficient based on the head movement change rate and the first uncertainty coefficient; and input the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position to a Kalman filter to obtain the predicted head movement position.
  • the apparatus further includes: a receiving module, configured to: receive first information sent by the audio content provider end, where the first information includes an audio format, a first audio stream, and the delay information, the audio format indicates that an audio source is in a stereo format or a non-stereo format, and when the audio format indicates that the audio source is in the non-stereo format, the first audio stream includes a first binaural signal rendered by the audio content provider end, or when the audio format indicates that the audio source is in the stereo format, the first audio stream includes the audio source; a rendering module, configured to: when the audio format indicates that the audio source is in the stereo format, obtain a measured head movement position through the sensor; and render a current frame based on the measured head movement position to obtain a second binaural signal, where the current frame is a frame of the audio source; and a playing module, configured to: play audio based on a target binaural signal, where the target binaural signal includes the first binaural signal or the second binaural signal.
  • a receiving module
  • the rendering module is specifically configured to: perform rendering with head movement effect on a direct sound part in the current frame based on the measured head movement position, and perform rendering without head movement effect on early reflected sound ER and late reverberant sound LR in the current frame, to obtain the second binaural signal.
  • this application provides a spatial audio rendering apparatus, including: an obtaining module, configured to: obtain a current frame when a first rendering mode is used, where the current frame is a frame of an audio source; and obtain a predicted head movement position when an audio format of the audio source is a non-stereo format; a rendering module, configured to: render the current frame based on the predicted head movement position to obtain a first binaural signal; and a sending module, configured to: send first information to an audio signal play end, where the first information includes a first audio stream, the audio format of the audio source, and delay information, and the first audio stream includes the first binaural signal.
  • the rendering module is specifically configured to: separately perform rendering with head movement effect on a direct sound part, an early reflected sound ER part, and a late reverberant sound LR part in the current frame based on the predicted head movement position, to obtain the first binaural signal.
  • the obtaining module is specifically configured to: obtain the predicted head movement position from rendering information sent by the audio signal play end.
  • the delay information includes a delay.
  • the rendering information further includes a measured head movement position in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information includes the measured head movement position.
  • the rendering information further includes a measured head movement position and a first uncertainty coefficient in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information includes the measured head movement position and the first uncertainty coefficient.
  • the first audio stream when an audio format of the audio source is a stereo format, the first audio stream includes the audio source.
  • the obtaining module is further configured to: obtain the current frame when a second rendering mode is used; the rendering module is further configured to: perform rendering without head movement effect on the current frame to obtain a second binaural signal; and the sending module is further configured to: send second information to the audio signal play end, where the second information includes a second audio stream, and the second audio stream includes the second binaural signal.
  • this application provides an audio signal play device, including: one or more processors; and a memory, configured to store one or more programs, where when the one or more programs is/are executed by the one or more processors, the one or more processors is/are enabled to implement the method implemented by the audio signal play end according to any one of the implementations of the first aspect.
  • this application provides an audio content providing device, including: one or more processors; and a memory, configured to store one or more programs, where when the one or more programs is/are executed by the one or more processors, the one or more processors is/are enabled to implement the method implemented by the audio content provider end according to any one of the implementations of the first aspect.
  • this application provides a computer-readable storage medium, including a computer program.
  • the computer program When the computer program is executed on a computer, the computer is enabled to perform the method according to any one of the implementations of the first aspect.
  • this application provides a computer program product.
  • the computer program product includes computer program code.
  • the computer program code When the computer program code is run on a computer, the computer is enabled to perform the method according to any one of the implementations of the first aspect.
  • the terms “first”, “second”, and the like are merely intended for differentiation in descriptions, but shall not be construed as indicating or implying relative importance or indicating or implying a sequence.
  • the terms “include”, “have”, and any variant thereof are intended to cover non-exclusive inclusion, for example, include a series of steps or units.
  • a method, system, product, or device is not necessarily limited to those expressly listed steps or units, but may include other steps or units that are not expressly listed or that are inherent to such a process, method, product, or device.
  • At least one of a, b, or c may indicate a, b, c, "a and b", “a and c", “b and c", or "a, b, and c", where a, b, and c may be in a singular form or a plural form.
  • sound may have two characteristics: a sense of orientation and a sense of space.
  • a spatial audio technology is usually a technology that simulates immersive audio experience with the foregoing two characteristics on a head-mounted play device such as a headset.
  • People's judgment on a sense of orientation for sound is mainly affected by a time difference, a sound level difference, human body filter effect learned during growth, head shaking, and other factors.
  • the time difference, the sound level difference, and the human body filter effect may be comprehensively expressed as a head-related transfer function (Head-Related Transfer Function, HRTF).
  • HRTF head-Related Transfer Function
  • the head shaking in all directions is of great help for determining a position of a sound source.
  • An indoor sound field may include direct sound, early reflected sound (Early Reflection, ER), and late reverberant sound (Late Reverb, LR). People's sense of space for sound is built based on the ER and the LR.
  • Sound sources in different formats such as stereo, multi-channel surround sound, and a specially produced multi-object sound source, may be rendered into a binaural signal by using the spatial audio technology, so that real and immersive auditory experience can be enjoyed through a headset.
  • a head of a user rotates or moves, an absolute position of a sound source does not change, but a relative direction between the sound source and the head changes.
  • a guitar is being played in front of a user, and if the user turns to the right, sound of the guitar is correspondingly shifted to the left of the user.
  • there is a guitar on a left side of a stage and there is a saxophone on a right side.
  • head tracking may be performed through a gyroscope or another sensor in a headset, and then a corresponding rotation or displacement change is incorporated in spatial rendering of a sound source.
  • spatial audio rendering usually needs to be supported by high computing power, and a low delay is needed for head movement tracking. If spatial audio rendering is performed at an audio content provider end (for example, a mobile phone, a tablet computer, or a PC), although a computing power requirement can be met, a long delay occurs.
  • an audio content provider end for example, a mobile phone, a tablet computer, or a PC
  • a computing power requirement can be met, a long delay occurs.
  • important spatial information such as an orientation is disordered, leading to degradation of overall quality of spatial audio rendering and affecting overall auditory experience.
  • this application provides a spatial audio rendering method and apparatus.
  • the following embodiments describe the technical solutions of this application.
  • FIG. 1A and FIG. 1B are a diagram of an application scenario according to this application.
  • the scenario includes an audio content provider end and an audio signal play end, and an interconnection mode between the audio content provider end and the audio signal play end includes but is not limited to a Bluetooth technology.
  • the audio content provider end may include but is not limited to a mobile phone, a tablet computer, a notebook computer, a desktop computer, and the like, and may provide high computing power, but may have a long delay.
  • the audio signal play end may include but is not limited to a true wireless stereo (true wireless stereo, TWS) headset, a wireless head-mounted headset, a wireless neckband headset, and the like, and has a low delay, but provides low computing power.
  • a gyroscope or another sensor is disposed in a device serving as the audio signal play end, to capture head movement information of a user.
  • a rendering algorithm includes two parts.
  • One part of the algorithm is deployed at the audio content provider end, and is used to perform high-computing-power rendering on audio in a multi-channel or multi-object format such as 5.1, 7.1, or 3DA, so that computing power advantages of the audio content provider end can be fully utilized.
  • the other part of the algorithm is deployed at the audio signal play end, and is used to perform low-computing-power rendering on audio in a stereo format.
  • An inertial measurement unit (Inertial Measurement Unit, IMU) data transmission link of the audio signal play end has a low delay, and may also run independently, without relying on a specific audio content provider end, to meet a requirement of a user for connecting to a device of any audio content provider end.
  • IMU Inertial Measurement Unit
  • this application further provides an algorithm for predicting a head movement position of a user, to predict a future head movement position of the user.
  • the predicted head movement position is used to provide assistance for the rendering algorithm at the audio content provider end, to reduce impact of a delay of the audio content provider end.
  • FIG. 1A and FIG. 1B is an example, but this should not constitute any limitation on this application.
  • An application scenario of the spatial audio rendering method is not specifically limited in this application either.
  • FIG. 2 is a diagram of a structure of an audio content provider end 200 according to this application. It should be understood that the audio content provider end 200 shown in FIG. 2 is merely an example, and the audio content provider end 200 may have more or fewer components than those shown in the figure, two or more components may be combined, or there may be different component configurations.
  • the components shown in FIG. 2 may be implemented in hardware including one or more signal processing and/or application-specific integrated circuits, software, or a combination of hardware and software.
  • the audio content provider end 200 may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (universal serial bus, USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headset jack 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display 294, a subscriber identity module (subscriber identity module, SIM) card interface 295, and the like.
  • a processor 210 an external memory interface 220, an internal memory 221, a universal serial bus (universal serial bus, USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an
  • the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, an optical proximity sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, and the like.
  • the processor 210 may include one or more processing units.
  • the processor 210 may include an application processor (application processor, AP), a modem processor, a graphics processing unit (graphics processing unit, GPU), an image signal processor (image signal processor, ISP), a controller, a memory, a video codec, a digital signal processor (digital signal processor, DSP), a baseband processor, and/or a neural-network processing unit (neural-network processing unit, NPU).
  • Different processing units may be independent components, or may be integrated into one or more processors.
  • the controller may be a nerve center and a command center of the audio content provider end 200.
  • the controller may generate an operation control signal based on an instruction operation code and a time sequence signal, to control instruction reading and instruction execution.
  • a memory may be further disposed in the processor 210 to store instructions and data.
  • the memory in the processor 210 is a cache.
  • the memory may store instructions or data that has been used or is cyclically used by the processor 210. If the processor 210 needs to use the instructions or the data again, the processor may directly invoke the instructions or the data from the memory. This avoids repeated access, reduces waiting time of the processor 210, and therefore improves system efficiency.
  • the processor 210 may include one or more interfaces.
  • the interface may include an inter-integrated circuit (inter-integrated circuit, I2C) interface, an inter-integrated circuit sound (inter-integrated circuit sound, I2S) interface, a pulse code modulation (pulse code modulation, PCM) interface, a universal asynchronous receiver/transmitter (universal asynchronous receiver/transmitter, UART) interface, a mobile industry processor interface (mobile industry processor interface, MIPI), a general-purpose input/output (general-purpose input/output, GPIO) interface, a subscriber identity module (subscriber identity module, SIM) interface, a universal serial bus (universal serial bus, USB) interface, and/or the like.
  • I2C inter-integrated circuit
  • I2S inter-integrated circuit sound
  • PCM pulse code modulation
  • PCM pulse code modulation
  • UART universal asynchronous receiver/transmitter
  • MIPI mobile industry processor interface
  • GPIO general-purpose input/output
  • the I2C interface is a two-way synchronous serial bus, and includes a serial data line (serial data line, SDA) and a serial clock line (serial clock line, SCL).
  • the processor 210 may include a plurality of groups of I2C buses.
  • the processor 210 may be separately coupled to the touch sensor 280K, a charger, a flash, the camera 293, and the like through different I2C bus interfaces.
  • the processor 210 may be coupled to the touch sensor 280K through the I2C interface, so that the processor 210 communicates with the touch sensor 280K through the I2C bus interface, to implement a touch function of the audio content provider end 200.
  • the I2S interface may be used for audio communication.
  • the processor 210 may include a plurality of groups of I2S buses.
  • the processor 210 may be coupled to the audio module 270 through the I2S bus, to implement communication between the processor 210 and the audio module 270.
  • the audio module 270 may transmit an audio signal to the wireless communication module 260 through the I2S interface, to implement a function of answering a call through a Bluetooth headset.
  • the UART interface is a universal serial data bus, and is used for asynchronous communication.
  • the bus may be a two-way communication bus.
  • the bus converts to-be-transmitted data between serial communication and parallel communication.
  • the UART interface is usually configured to connect the processor 210 to the wireless communication module 260.
  • the processor 210 communicates with a Bluetooth module in the wireless communication module 260 through the UART interface, to implement a Bluetooth function.
  • the audio module 270 may transmit an audio signal to the wireless communication module 260 through the UART interface, to implement a function of playing music through a Bluetooth headset.
  • the MIPI interface may be configured to connect the processor 210 to a peripheral component such as the display 294 or the camera 293.
  • the MIPI interface includes a camera serial interface (camera serial interface, CSI), a display serial interface (display serial interface, DSI), and the like.
  • the processor 210 communicates with the camera 293 through the CSI interface, to implement an image shooting function of the audio content provider end 200.
  • the processor 210 communicates with the display 294 through the DSI interface, to implement a display function of the audio content provider end 200.
  • the GPIO interface may be configured by software.
  • the GPIO interface may be configured as a control signal or a data signal.
  • the GPIO interface may be configured to connect the processor 210 to the camera 293, the display 294, the wireless communication module 260, the audio module 270, the sensor module 280, and the like.
  • the GPIO interface may alternatively be configured as an I2C interface, an I2S interface, a UART interface, an MIPI interface, or the like.
  • the USB interface 230 is an interface that conforms to a USB standard specification, and may be specifically a mini USB interface, a micro USB interface, a USB Type-C interface, or the like.
  • the USB interface 230 may be used for connecting a charger to charge the audio content provider end 200, or may be configured to transmit data between the audio content provider end 200 and a peripheral device, or may be used for connecting a headset for playing audio through the headset.
  • the interface may alternatively be used for connecting other user equipment, for example, an AR device.
  • an interface connection relationship between the modules shown in this embodiment of this application is merely an example for description, and does not constitute a limitation on a structure of the audio content provider end 200.
  • the audio content provider end 200 may alternatively use an interface connection mode different from that in the foregoing embodiment, or use a combination of a plurality of interface connection modes.
  • the charging management module 240 is configured to receive charging input from the charger.
  • the charger may be a wireless charger or a wired charger.
  • the charging management module 240 may receive charging input from a wired charger through the USB interface 230.
  • the charging management module 240 may receive wireless charging input through a wireless charging coil of the audio content provider end 200.
  • the charging management module 240 may further supply power to user equipment through the power management module 241.
  • the power management module 241 is configured to connect to the battery 242, the charging management module 240, and the processor 210.
  • the power management module 241 receives input from the battery 242 and/or the charging management module 240, and supplies power to the processor 210, the internal memory 221, an external memory, the display 294, the camera 293, the wireless communication module 260, and the like.
  • the power management module 241 may be further configured to monitor parameters such as a battery capacity, a quantity of battery cycles, and a battery health status (electric leakage and impedance).
  • the power management module 241 may alternatively be disposed in the processor 210.
  • the power management module 241 and the charging management module 240 may alternatively be disposed in a same component.
  • a wireless communication function of the audio content provider end 200 may be implemented by the antenna 1, the antenna 2, the mobile communication module 250, the wireless communication module 260, the modem processor, the baseband processor, and the like.
  • the antenna 1 and the antenna 2 are configured to transmit and receive an electromagnetic wave signal.
  • Each antenna in the audio content provider end 200 may be configured to cover one or more communication frequency bands. Different antennas may be further multiplexed to improve antenna utilization.
  • the antenna 1 may be multiplexed as a diversity antenna of a wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.
  • the mobile communication module 250 may provide a solution applied to the audio content provider end 200 for wireless communication such as 2G/3G/4G/5G.
  • the mobile communication module 250 may include at least one filter, a switch, a power amplifier, a low noise amplifier (low noise amplifier, LNA), and the like.
  • the mobile communication module 250 may receive an electromagnetic wave through the antenna 1, perform processing such as filtering or amplification on the received electromagnetic wave, and transmit a processed electromagnetic wave to the modem processor for demodulation.
  • the mobile communication module 250 may further amplify a signal modulated by the modem processor, and convert an amplified signal into an electromagnetic wave for radiation through the antenna 1.
  • at least some functional modules of the mobile communication module 250 may be disposed in the processor 210.
  • at least some functional modules of the mobile communication module 250 may be disposed in a same component as at least some modules of the processor 210.
  • the modem processor may include a modulator and a demodulator.
  • the modulator is configured to modulate a to-be-sent low-frequency baseband signal into a medium-high frequency signal.
  • the demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. Then the demodulator transmits the low-frequency baseband signal obtained through demodulation to the baseband processor for processing.
  • the low-frequency baseband signal is processed by the baseband processor and then transmitted to the application processor.
  • the application processor outputs a sound signal through an audio device (not limited to the speaker 270A, the receiver 270B, and the like), or displays an image or a video through the display 294.
  • the modem processor may be an independent component.
  • the modem processor may be independent of the processor 210, and is disposed in a same component as the mobile communication module 250 or another functional module.
  • the wireless communication module 260 may provide a solution applied to the audio content provider end 200 for wireless communication such as a wireless local area network (wireless local area network, WLAN) (for example, a wireless fidelity (wireless fidelity, Wi-Fi) network), Bluetooth (Bluetooth, BT), a global navigation satellite system (global navigation satellite system, GNSS), frequency modulation (frequency modulation, FM), a near field communication (near field communication, NFC) technology, or an infrared (infrared, IR) technology.
  • the wireless communication module 260 may be one or more devices integrating at least one communication processing module.
  • the wireless communication module 260 receives an electromagnetic wave through the antenna 2, performs frequency modulation and filtering on an electromagnetic wave signal, and sends a processed signal to the processor 210.
  • the wireless communication module 260 may further receive a to-be-sent signal from the processor 210, perform frequency modulation and amplification on the signal, and convert a processed signal into an electromagnetic wave for radiation through the antenna 2.
  • the antenna 1 of the audio content provider end 200 is coupled to the mobile communication module 250, and the antenna 2 is coupled to the wireless communication module 260, so that the audio content provider end 200 can communicate with a network and another device by using a wireless communication technology.
  • the wireless communication technology may include a global system for mobile communications (global system for mobile communications, GSM), a general packet radio service (general packet radio service, GPRS), code division multiple access (code division multiple access, CDMA), wideband code division multiple access (wideband code division multiple access, WCDMA), time-division code division multiple access (time-division code division multiple access, TD-SCDMA), long term evolution (long term evolution, LTE), BT, a GNSS, a WLAN, NFC, FM, an IR technology, and/or the like.
  • GSM global system for mobile communications
  • GPRS general packet radio service
  • code division multiple access code division multiple access
  • CDMA wideband code division multiple access
  • WCDMA wideband code division multiple access
  • the GNSS may include a global positioning system (global positioning system, GPS), a global navigation satellite system (global navigation satellite system, GLONASS), a BeiDou navigation satellite system (BeiDou navigation satellite system, BDS), a quasi-zenith satellite system (quasi-zenith satellite system, QZSS), and/or a satellite-based augmentation system (satellite-based augmentation system, SBAS).
  • GPS global positioning system
  • GLONASS global navigation satellite system
  • BeiDou navigation satellite system BeiDou navigation satellite system
  • BDS BeiDou navigation satellite system
  • QZSS quasi-zenith satellite system
  • SBAS satellite-based augmentation system
  • the audio content provider end 200 implements a display function through the GPU, the display 294, the application processor, and the like.
  • the GPU is a microprocessor for image processing, and is connected to the display 294 and the application processor.
  • the GPU is configured to perform mathematical and geometric computation, and render an image.
  • the processor 210 may include one or more GPUs that execute program instructions to generate or change displayed information.
  • the display 294 is configured to display an image, a video, or the like.
  • the display 294 includes a display panel.
  • the display panel may be a liquid crystal display (liquid crystal display, LCD), an organic light-emitting diode (organic light-emitting diode, OLED), an active-matrix organic light-emitting diode (active-matrix organic light-emitting diode, AMOLED), a flexible light-emitting diode (flex light-emitting diode, FLED), a mini-LED, a micro-LED, a micro-OLED, a quantum dot light-emitting diode (quantum dot light-emitting diode, QLED), or the like.
  • the audio content provider end 200 may include one or N displays 294, where N is a positive integer greater than 1.
  • the audio content provider end 200 may implement an image shooting function through the ISP, the camera 293, the video codec, the GPU, the display 294, the application processor, and the like.
  • the ISP is configured to process data fed back by the camera 293. For example, during photographing, a shutter is pressed, and light is transmitted to a photosensitive element of the camera through a lens. An optical signal is converted into an electrical signal, and the photosensitive element of the camera transmits the electrical signal to the ISP for processing, to convert the electrical signal into a visible image.
  • the ISP may further perform algorithm optimization on noise, brightness, and complexion of the image.
  • the ISP may further optimize parameters such as exposure and color temperature of an image shooting scene.
  • the ISP may be disposed in the camera 293.
  • the camera 293 is configured to capture a static image or a video. An optical image of an object is generated through the lens, and is projected onto the photosensitive element.
  • the photosensitive element may be a charge coupled device (charge coupled device, CCD) or a complementary metal-oxide-semiconductor (complementary metal-oxide-semiconductor, CMOS) phototransistor.
  • CCD charge coupled device
  • CMOS complementary metal-oxide-semiconductor
  • the photosensitive element converts an optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert the electrical signal into a digital image signal.
  • the ISP outputs the digital image signal to the DSP for processing.
  • the DSP converts the digital image signal into an image signal in a standard format, for example, RGB or YUV.
  • the audio content provider end 200 may include one or N cameras 293, where N is a positive integer greater than 1.
  • the video codec is configured to compress or decompress a digital video.
  • the audio content provider end 200 may support one or more types of video codecs. In this way, the audio content provider end 200 can play or record videos in a plurality of coding formats, for example, moving picture experts group (moving picture experts group, MPEG)-1, MPEG-2, MPEG-3, and MPEG-4.
  • MPEG moving picture experts group
  • the NPU is a neural-network (neural-network, NN) computing processor.
  • the NPU quickly processes input information with reference to a structure of a biological neural network, for example, a mode of transfer between human brain neurons, and may further continuously perform self-learning.
  • Intelligent cognition applications such as image recognition, facial recognition, speech recognition, and text understanding, of the audio content provider end 200 may be implemented through the NPU.
  • the external memory interface 220 may be used for connecting an external memory card, for example, a microSD card, to extend a storage capability of the audio content provider end 200.
  • the external memory card communicates with the processor 210 through the external memory interface 220, to implement a data storage function. For example, files such as music and videos are stored in the external storage card.
  • the internal memory 221 may be configured to store computer-executable program code, and the executable program code includes instructions.
  • the processor 210 runs the instructions stored in the internal memory 221 to implement various function applications and data processing of the audio content provider end 200.
  • the internal memory 221 may include a program storage area and a data storage area.
  • the program storage area may store an operating system, an application for at least one function (for example, a sound play function or an image play function), and the like.
  • the data storage area may store data (for example, audio data and an address book) created during use of the audio content provider end 200, and the like.
  • the internal memory 221 may include a high-speed random access memory, or may include a non-volatile memory, for example, at least one magnetic disk storage device, a flash memory, or a universal flash storage (universal flash storage, UFS).
  • the audio content provider end 200 may implement an audio function, for example, music playing or recording, through the audio module 270, the speaker 270A, the receiver 270B, the microphone 270C, the headset jack 270D, the application processor, and the like.
  • an audio function for example, music playing or recording
  • the audio module 270 is configured to convert digital audio information into an analog audio signal for output, and is also configured to convert analog audio input into a digital audio signal.
  • the audio module 270 may be further configured to encode and decode an audio signal.
  • the audio module 270 may be disposed in the processor 210, or some functional modules of the audio module 270 are disposed in the processor 210.
  • the speaker 270A also referred to as a "loudspeaker" is configured to convert an electrical audio signal into a sound signal.
  • the audio content provider end 200 may be used to listen to music or answer a call in a hands-free mode through the speaker 270A.
  • the receiver 270B also referred to as an "earpiece" is configured to convert an electrical audio signal into a sound signal.
  • the receiver 270B may be put close to a human ear to listen to a voice.
  • the pressure sensor 280A is configured to sense a pressure signal, and may convert the pressure signal into an electrical signal.
  • the pressure sensor 280A may be disposed on the display 294.
  • There are many types of pressure sensors 280A for example, a resistive pressure sensor, an inductive pressure sensor, and a capacitive pressure sensor.
  • the capacitive pressure sensor may include at least two parallel plates made of conductive materials.
  • the barometric pressure sensor 280C is configured to measure barometric pressure.
  • the audio content provider end 200 calculates an altitude based on a value of the barometric pressure measured by the barometric pressure sensor 280C, to provide assistance for positioning and navigation.
  • the magnetic sensor 280D includes a Hall effect sensor.
  • the audio content provider end 200 may detect opening and closing of a flip leather case through the magnetic sensor 280D.
  • the audio content provider end 200 may detect opening and closing of a flip cover based on the magnetic sensor 280D.
  • a feature such as automatic unlocking of the flip cover is set based on a detected opening/closing state of the leather case or a detected opening/closing state of the flip cover.
  • the acceleration sensor 280E may detect accelerations of the audio content provider end 200 in various directions (usually on three axes), may detect a magnitude and a direction of gravity when the audio content provider end 200 is still, and may be further configured to recognize an attitude of user equipment and used in applications such as landscape/portrait mode switching and a pedometer.
  • the fingerprint sensor 280H is configured to capture a fingerprint.
  • the audio content provider end 200 may implement fingerprint-based unlocking, application lock access, fingerprint-based photographing, fingerprint-based call answering, and the like by using a feature of the captured fingerprint.
  • the temperature sensor 280J is configured to detect temperature.
  • the audio content provider end 200 executes a temperature processing policy based on the temperature detected by the temperature sensor 280J. For example, when the temperature reported by the temperature sensor 280J exceeds a threshold, the audio content provider end 200 degrades performance of a processor near the temperature sensor 280J, to reduce power consumption for thermal protection.
  • the audio content provider end 200 heats the battery 242 to avoid abnormal shutdown of the audio content provider end 200 due to low temperature.
  • the audio content provider end 200 boosts an output voltage of the battery 242 to avoid abnormal shutdown due to low temperature.
  • the touch sensor 280K is also referred to as a "touch panel”.
  • the touch sensor 280K may be disposed on the display 294.
  • the touch sensor 280K and the display 294 constitute a touchscreen, which is also referred to as a "touch control screen”.
  • the touch sensor 280K is configured to detect a touch operation performed on or near the touch sensor.
  • the touch sensor may transmit the detected touch operation to the application processor to determine a type of a touch event.
  • the display 294 may provide visual output related to the touch operation.
  • the touch sensor 280K may alternatively be disposed on a surface of the audio content provider end 200, and a position of the touch sensor 280K is different from a position of the display 294.
  • the bone conduction sensor 280M may obtain a vibration signal.
  • the bone conduction sensor 280M may obtain a vibration signal of a vibration bone of a human vocal-cord part.
  • the bone conduction sensor 280M may also be in contact with a body pulse to receive a blood pressure beating signal.
  • the bone conduction sensor 280M may alternatively be disposed in a headset, to constitute a bone conduction headset.
  • the audio module 270 may obtain a speech signal through parsing based on the vibration signal, obtained by the bone conduction sensor 280M, of the vibration bone of the vocal-cord part, to implement a speech function.
  • the application processor may parse heart rate information based on the blood pressure beating signal obtained by the bone conduction sensor 280M, to implement a heart rate detection function.
  • the button 290 includes a power button, a volume button, and the like.
  • the button 290 may be a mechanical button or a touch button.
  • the audio content provider end 200 may receive key input, and generate key signal input related to user settings and function control of the audio content provider end 200.
  • the motor 291 may generate a vibration prompt.
  • the motor 291 may be configured to provide an incoming call vibration prompt or a touch vibration feedback.
  • touch operations performed on different applications may correspond to different vibration feedback effect.
  • the motor 291 may also correspond to different vibration feedback effect for touch operations performed on different areas of the display 294.
  • Different application scenarios for example, a time reminder, information receiving, an alarm clock, and a game
  • Touch vibration feedback effect may be further customized.
  • the indicator 292 may be an indicator light, and may be configured to indicate a charging status and a battery level change, or may be configured to indicate a message, a missed call, a notification, and the like.
  • the audio content provider end 200 interacts with a network through the SIM card, to implement functions such as calling and data communication.
  • the audio content provider end 200 uses an eSIM, namely, an embedded SIM card.
  • the eSIM card may be embedded in the audio content provider end 200, and cannot be separated from the audio content provider end 200.
  • the structure shown in this embodiment of the present invention does not constitute a specific limitation on a controlling device.
  • the controlling device may include more or fewer components than those shown in the figure, or some components may be combined, or some components may be split, or different component layouts may be used.
  • the components shown in the figure may be implemented by hardware, software, or a combination of software and hardware.
  • Step 301 Obtain delay information of the audio content provider end.
  • the delay information indicates a delay status of the audio content provider end, and may include the following cases.
  • the delay information includes a delay sent by the audio content provider end.
  • the audio content provider end may package the delay into first information to be sent to the audio signal play end, and send the first information to the audio signal play end. In this way, the audio signal play end can directly extract the delay of the audio content provider end from the first information that comes from the audio content provider end.
  • the delay information includes a second historical head movement position sent by the audio content provider end.
  • the audio signal play end (for example, a headset) may periodically detect, through a sensor (for example, a gyroscope or a gravity sensor), a current position (referred to as a measured head movement position in this specification) of a head of a user wearing the headset. In this way, the audio signal play end can send the measured head movement position obtained through measurement to the audio content provider end.
  • a sensor for example, a gyroscope or a gravity sensor
  • a current position referred to as a measured head movement position in this specification
  • the audio content provider end After receiving the measured head movement position, the audio content provider end does not process the measured head movement position, but still performs related processing on an audio frame according to a specified process.
  • the measured head movement position is carried in the first information. In this case, because a period of time has elapsed, the measured head movement position becomes a historical head movement position (referred to as the second historical head movement position in this specification). It should be understood that, in a mathematical sense, the measured head movement position is the same as the second historical head movement position.
  • the audio signal play end may obtain, based on time at which the second historical head movement position is received and time at which the measured head movement position is sent, a time difference between the time at which the second historical head movement position is received and the time at which the measured head movement position is sent (in other words, calculate a time difference between receiving time and sending time of a same head movement position), to indirectly obtain the delay of the audio content provider end.
  • the delay information includes a second historical head movement position and a first uncertainty coefficient that are sent by the audio content provider end.
  • the first uncertainty coefficient is an important parameter in a Kalman filter, and helps improve accuracy of a prediction result.
  • This case is related to an algorithm for obtaining a predicted head movement position in step 302.
  • This case is related to an algorithm for obtaining a predicted head movement position in step 302.
  • Step 302 Obtain a predicted head movement position based on the delay information.
  • sound sources in different formats such as stereo, multi-channel surround sound, and a specially produced multi-object sound source, may be rendered into a binaural signal by using a spatial audio technology.
  • a head of a user rotates or moves, an absolute position of a sound source does not change, but a relative direction between the sound source and the head changes. For example, a guitar is being played in front of a user, and if the user turns to the right, sound of the guitar is correspondingly shifted to the left of the user.
  • the audio content provider end may provide high computing power to implement the foregoing rendering algorithm with head movement effect
  • a long delay occurs, causing disorder of important spatial information such as an orientation.
  • an acoustic image of an object is designed to be directly in front of space. If a head of a user rotates, the acoustic image first moves to be directly in front of a face of the user, and then is restored to be directly in front of space. A clear delay and "sense of damping" occur, affecting auditory experience.
  • the audio signal play end may predict a head movement position of a user, and then send the predicted head movement position to the audio content provider end, so that the audio content provider end can render an audio frame based on the predicted head movement position.
  • an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position.
  • predicted time may offset a delay caused by rendering, link transmission, and the like. In this way, when the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.
  • the audio signal play end may first obtain a first historical head movement position corresponding to the delay information, and then obtain a measured head movement position through the sensor, and then obtain the predicted head movement position based on the first historical head movement position and the measured head movement position.
  • the audio signal play end may obtain the first historical head movement position by using three methods.
  • the audio signal play end may alternatively directly determine the second historical head movement position as the first historical head movement position.
  • the audio signal play end may obtain a measured head movement position through the sensor, where the measured head movement position is a newly detected head movement position of the user.
  • obtaining the predicted head movement position based on the first historical head movement position and the measured head movement position may include the following two algorithms.
  • the audio signal play end may obtain a difference between the first historical head movement position and the measured head movement position; obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; and obtain the predicted head movement position based on the difference and the head movement change rate.
  • the difference between the first historical head movement position and the measured head movement position may be a distance between the first historical head movement position and the measured head movement position.
  • a head movement position may be represented by using coordinate axes x, y, and z. Therefore, the foregoing difference may be a distance between a coordinate value of the first historical head movement position and a coordinate value of the measured head movement position.
  • a head movement position may alternatively be represented in another manner. This is not specifically limited in this application.
  • the audio signal play end may periodically detect, through the sensor, a measured head movement position of a user wearing the audio signal play end. Therefore, the audio signal play end may obtain a head movement change rate based on N measured head movement positions that are previously obtained, where the N measured head movement positions may include N measured head movement positions obtained through counting forward from a current measured head movement position.
  • the audio signal play end may predict a head movement change value based on the difference and the head movement change rate through cubic spline interpolation or by another means, and then add up the head movement change value and a current measured head movement position to obtain a predicted head movement position.
  • the audio signal play end obtains a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; obtains a second uncertainty coefficient based on the head movement change rate and the first uncertainty coefficient; and inputs the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position to a Kalman filter to obtain the predicted head movement position.
  • the second uncertainty coefficient may be obtained based on the head movement change rate and the first uncertainty coefficient.
  • the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position are input to the Kalman filter.
  • a Kalman gain coefficient may be determined based on the second uncertainty coefficient and estimated uncertainty of a previous iteration, and the Kalman gain coefficient is used as a latest gain coefficient. In this way, the Kalman filter can output the predicted head movement position.
  • the predicted head movement position may alternatively be obtained based on the first historical head movement position and the measured head movement position by using another algorithm. This is not specifically limited herein.
  • Step 303 Send first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position.
  • the audio signal play end sends the first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position obtained in step 302, so that the audio content provider end can perform rendering with head movement effect on an audio frame based on the predicted head movement position.
  • the audio content provider end can perform rendering with head movement effect on an audio frame based on the predicted head movement position.
  • an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position.
  • predicted time may offset a delay caused by rendering, link transmission, and the like.
  • the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.
  • a head movement position of a user is predicted to obtain a predicted head movement position, and after the predicted head movement position is sent to the audio content provider end, the audio content provider end may render an audio frame based on the predicted head movement position.
  • an obtained rendered audio frame can provide auditory effect corresponding to the predicted head movement position.
  • predicted time may offset a delay caused by rendering, link transmission, and the like. In this way, when the audio frame is played at the audio signal play end, the audio frame exactly matches the head movement position of the user, to achieve the real and immersive auditory experience described above.
  • FIG. 4 is a flowchart of a process 400 of a spatial audio rendering method according to this application.
  • the process 400 may be applied to the communication system shown in FIG. 1A and FIG. 1B , and is jointly performed by an audio content provider end and an audio signal play end, to complete rendering of an audio source, and play audio to a user through the audio signal play end.
  • the process 400 is described as a series of steps or operations. It should be understood that the process 400 may be performed in various sequences and/or simultaneously, and is not limited to an execution sequence shown in FIG. 4 .
  • the process 400 includes the following steps.
  • Step 401 When a first rendering mode is used, the audio content provider end obtains a current frame.
  • the first rendering mode may be a non-low-delay link mode.
  • the audio content provider end may perform rendering with head movement effect on an audio frame based on a predicted head movement position.
  • a user may perform an operation on an application (application, APP) installed on the audio content provider end, to choose whether to select the first rendering mode.
  • this application further includes a second rendering mode.
  • the audio content provider end performs rendering without head movement effect on an audio frame.
  • the user may perform an operation on the APP installed on the audio content provider end, to choose whether to select the second rendering mode.
  • the operation of the user refer to the following embodiments.
  • the current frame may be a frame of the audio source, and usually, may be an audio frame that is currently being processed by the audio content provider end.
  • Step 402 When an audio format of an audio source is a non-stereo format, the audio content provider end obtains a predicted head movement position.
  • the audio format of the audio source includes a stereo format or the non-stereo format, where the non-stereo format may include but is not limited to audio in a multi-channel or multi-object format such as 5.1, 7.1, or 3DA.
  • the audio content provider end may perform high-computing-power rendering on the audio in the multi-channel or multi-object format (non-stereo format) such as 5.1, 7.1, or 3DA, so that computing power advantages of the audio content provider end can be fully utilized.
  • the audio content provider end may perform rendering with head movement effect on an audio frame. Therefore, the predicted head movement position needs to be obtained.
  • the audio content provider end may obtain the predicted head movement position from rendering information that is previously received from the audio signal play end.
  • the rendering information may be latest received rendering information. In this way, accuracy of the predicted head movement position can be improved.
  • For obtaining of the predicted head movement position refer to the embodiment shown in FIG. 3 . Details are not described herein.
  • Step 403 The audio content provider end renders the current frame based on the predicted head movement position to obtain a first binaural signal.
  • the audio content provider end may separately perform rendering with head movement effect on a direct sound part, an early reflected sound (Early Reflection, ER) part, and a late reverberant sound (Late Reverb, LR) part in the current frame based on the predicted head movement position, to obtain the first binaural signal.
  • ER early reflected sound
  • LR late reverberant sound
  • direct sound with head movement effect and reverberation (obtained by mixing ER and LR) with head movement effect are separately rendered by an audio effector and then mixed, and an output result obtained through mixing is rendered by the audio effector again and then transmitted in a form of the first binaural signal.
  • the audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.
  • a first audio stream may include the first binaural signal.
  • Step 404 When an audio format of an audio source is a stereo format, the audio content provider end does not perform spatial rendering on the audio source.
  • the audio content provider end does not need to render an audio source in a stereo format, and the audio signal play end performs rendering with head movement effect on the audio source. Therefore, the audio content provider end may perform high-computing-power rendering on the audio in the multi-channel or multi-object format (non-stereo format) such as 5.1, 7.1, or 3DA, so that computing power advantages of the audio content provider end can be fully utilized. Based on this, when the audio format of the audio source is the non-stereo format, the audio content provider end may perform rendering with head movement effect on an audio frame. Therefore, the audio content provider end only needs to package the audio source into the first audio stream.
  • Step 405 The audio content provider end sends first information to the audio signal play end, where the first information includes the first audio stream, the audio format of the audio source, and delay information.
  • the first audio stream includes the first binaural signal rendered by the audio content provider end; or when the audio format indicates that the audio source is in the stereo format, the first audio stream includes the audio source.
  • Step 406 When the audio format indicates that the audio source is in the stereo format, the audio signal play end obtains a measured head movement position through a sensor.
  • the audio signal play end For an audio source in a stereo format, the audio signal play end performs rendering with a head movement position on the audio source. Because an IMU data transmission link of the audio signal play end has a low delay, the audio signal play end can obtain a current head movement position (referred to as the measured head movement position in this specification) in real time through the sensor.
  • Step 407 The audio signal play end renders the current frame based on the measured head movement position to obtain a second binaural signal.
  • the audio signal play end may perform rendering with head movement effect on the direct sound part in the current frame based on the measured head movement position, and perform rendering without head movement effect on the ER and the LR in the current frame, to obtain the second binaural signal.
  • direct sound with head movement effect and reverberation (obtained by mixing ER and LR) without head movement effect are separately rendered by an audio effector and then mixed, and an output result obtained through mixing is rendered by the audio effector again and then mixed again, to obtain the second binaural signal.
  • the audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.
  • Step 408 The audio signal play end plays audio based on a target binaural signal, where the target binaural signal includes the first binaural signal or the second binaural signal.
  • the audio signal play end obtains the target binaural signal through processing in the foregoing steps, where the target binaural signal may be the first binaural signal (the audio source is in the non-stereo format) or the second binaural signal (the audio source is in the stereo format).
  • the target binaural signal may be the first binaural signal (the audio source is in the non-stereo format) or the second binaural signal (the audio source is in the stereo format).
  • the audio content provider end when the second rendering mode is used, obtains a current frame, performs rendering without head movement effect on the current frame to obtain a second binaural signal (the second binaural signal is different from the foregoing second binaural signal, and is referred to as a third binaural signal below for differentiation), and sends second information to the audio signal play end, where the second information includes a second audio stream, and the second audio stream includes the third binaural signal.
  • the audio signal play end obtains a measured head movement position through the sensor, and renders a current frame (the current frame is an audio frame obtained by the audio content provider end by performing rendering without head movement effect) based on the measured head movement position to obtain a fourth binaural signal.
  • the audio signal play end plays audio based on a target binaural signal, where the target binaural signal includes the fourth binaural signal.
  • the audio content provider end performs rendering without head movement effect on an audio frame, so that computing power advantages of the audio content provider end can be fully utilized.
  • a rendering delay can be shortened without head movement effect.
  • the audio signal play end performs low-computing-power rendering with head movement effect on the audio frame. Compared with the first rendering mode, this achieves a lower delay, but rendering effect of head movement effect is poor.
  • an audio content provider end may be a mobile phone, and an audio signal play end may be a headset.
  • An audio source may alternatively be a sound source or an input signal.
  • Rendering may alternatively be (binaural) spatial audio rendering.
  • a head movement position may alternatively be IMU data.
  • a rendering algorithm deployed on the mobile phone may alternatively be a first rendering part of a binaural spatial audio rendering algorithm.
  • a rendering algorithm deployed on the headset may alternatively be a second rendering part of the binaural spatial audio rendering algorithm.
  • An audio format may alternatively be a format flag.
  • a delay may alternatively be a dynamic delay of a spatial audio link.
  • FIG. 5 is a diagram of an overall process of a spatial audio rendering method according to this application. As shown in FIG. 5 , this application is applied to a spatial audio processing scenario in which a mobile phone and a headset perform joint rendering in a first rendering mode.
  • a joint rendering policy in this embodiment may be designed as follows: If the mobile phone determines that a sound source is stereo, audio is sent to the headset for binaural spatial audio rendering; or if the mobile phone determines that a sound source is non-stereo, audio is rendered on the mobile phone, and after the rendering is completed, rendered audio is sent to the headset, and the headset directly outputs the rendered audio to a speaker.
  • FIG. 6 is a diagram of a specific process of a spatial audio rendering method according to this application. As shown in FIG. 6 , the process in this embodiment includes the following steps.
  • a mobile phone receives an input signal, and detects a format of the input signal, and a headset delivers predicted IMU data to a first rendering part of a binaural spatial audio rendering algorithm of the mobile phone.
  • the first rendering part of the binaural spatial audio rendering algorithm includes direct sound (Direct) rendering, early reflection (ER) rendering, and late reverberation (LR) rendering.
  • Head movement effect processing is performed, by using received IMU data, on a reverberation part obtained by mixing a direct sound part with ER and LR.
  • Direct sound with head movement effect and reverberation with head movement effect are separately rendered by an audio effector, and then rendered audio enters a first mixing module (Mixer 1).
  • An output result of the first mixing module is rendered by an audio effector again, and then rendered audio is transmitted to Bluetooth in a dual-channel form.
  • the audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.
  • S4 Package the audio stream and the format flag of the first mixing module and a dynamic delay of a spatial audio link of the mobile phone, and transmit a package to the headset through Bluetooth.
  • the headset stores measured IMU data, selects corresponding historical measured IMU data and measured IMU data of the headset based on the delay reported by the mobile phone, transmits the historical measured IMU data and the measured IMU data of the headset to a prediction module, and transmits the measured IMU data of the headset to a second rendering part of the binaural spatial audio rendering algorithm.
  • the prediction module generates predicted IMU data based on the historical measured IMU data and the measured IMU data of the headset, and delivers the predicted IMU data to the mobile phone again with reference to S1.
  • FIG. 7 is a schematic flowchart of an algorithm for predicting a head movement position. As shown in FIG. 7 , the process includes the following steps.
  • a headset stores historical IMU data in a cache, extracts corresponding historical IMU data from the historical cache based on a received mobile phone delay, calculates a difference between measured IMU data and the historical IMU data, and sends the difference to a prediction unit.
  • S5.2 Calculate a change rate of previous N pieces of measured IMU data (for example, including a yaw (yaw), a pitch (pitch), a roll (roll), or four elements) of the headset, and transmit a change coefficient to the prediction unit.
  • a change rate of previous N pieces of measured IMU data for example, including a yaw (yaw), a pitch (pitch), a roll (roll), or four elements
  • S5.3 Predict a head movement change value through cubic spline interpolation or by another means based on the change rate and the difference calculated in S5.1.
  • the headset detects a format of a sound source, and if the format of the sound source is stereo, renders an input signal based on a second part of a binaural spatial audio rendering algorithm by using the measured IMU data of the headset, and sends a rendered audio stream to a second mixing module.
  • the second rendering part of the binaural spatial audio rendering algorithm includes direct sound (Direct) rendering and basic low-computing-power reverberation rendering.
  • Head movement effect processing is performed on a direct sound part by using received IMU data.
  • Direct sound with head movement effect and basic reverberation without head movement effect are separately rendered by an audio effector and then mixed.
  • An overall output result obtained through mixing is rendered by an audio effector again, and then rendered audio enters the second mixing module (Mixer 2).
  • the audio effector includes but is not limited to an equalization effector, a dynamic compression effector, a low frequency enhancement effector, and the like.
  • S8 Send the second mixing module to a headset speaker for playing.
  • Embodiment 1 high computing power of the mobile phone and a low delay of the headset are fully utilized, and the headset may run independently.
  • an end-to-end delay prediction algorithm is designed to greatly reduce a head movement delay (by more than 50% as predicted).
  • Embodiment 2 Compared with Embodiment 1, a difference in Embodiment 2 mainly focuses on content transmitted through Bluetooth in S1 and S4 and an algorithm for predicting a head movement position in S5.
  • S1 is changed as follows: A mobile phone receives an input signal, and detects a format of the input signal, and a headset delivers IMU data of the headset and predicted IMU data to a first rendering part of a binaural spatial audio rendering algorithm of the mobile phone.
  • S4 is changed as follows: Package the IMU data, the audio stream, and the format flag of the first mixing module, and transmit a package to the headset through Bluetooth.
  • S5 is changed as follows:
  • the headset transmits, to a prediction module, historical IMU data reported by the mobile phone, transmits measured IMU data to the prediction module, and transmits the measured IMU data to a second rendering part of the binaural spatial audio rendering algorithm.
  • the prediction module generates predicted IMU data based on the historical IMU data and the measured IMU data, and delivers the measured IMU data and the predicted IMU data to the mobile phone with reference to S1.
  • FIG. 8a and FIG. 8b are a schematic flowchart of an algorithm for predicting a head movement position. As shown in FIG. 8a and FIG. 8b , the process includes the following steps.
  • S5.1 Calculate a difference between measured IMU data and historical IMU data reported by a mobile phone, and send the difference to a prediction algorithm.
  • S5.2 Calculate a change rate of previous N pieces of measured IMU data (for example, including a yaw, a pitch, a roll, or four elements) of a headset, and transmit a change coefficient to a prediction unit.
  • N pieces of measured IMU data for example, including a yaw, a pitch, a roll, or four elements
  • S5.3 Predict a head movement change value through cubic spline interpolation or by another means based on the change rate and the difference calculated in S5.1.
  • Embodiment 3 Compared with Embodiment 1, a difference in Embodiment 3 mainly focuses on content transmitted through Bluetooth in S1 and S4 and a prediction algorithm in S5.
  • S1 is changed as follows: A mobile phone receives an input signal, and detects a format of the input signal, and a headset delivers an uncertainty coefficient Rn and predicted IMU data to a first rendering part of a binaural spatial audio rendering algorithm of the mobile phone.
  • S4 is changed as follows: Package the IMU data, the uncertainty coefficient Rn, the audio stream, and the format flag of the first mixing module, and transmit a package to the headset through Bluetooth.
  • S5 is changed as follows:
  • the headset transmits, to a prediction module, IMU data reported by the mobile phone and the uncertainty coefficient Rn, transmits measured IMU data to the prediction module, and transmits the measured IMU data to a second rendering part of the binaural spatial audio rendering algorithm.
  • the prediction module generates a predicted value of IMU data based on the IMU data reported by the mobile phone, the uncertainty coefficient Rn, and IMU data of the headset at this time, and delivers the measured IMU data and the predicted IMU data to the mobile phone with reference to S1.
  • FIG. 9a and FIG. 9b are a schematic flowchart of an algorithm for predicting a head movement position. As shown in FIG. 9a and FIG. 9b , the process includes the following steps.
  • S5.1 Calculate a change rate of previous N pieces of measured IMU data (for example, including a yaw, a pitch, a roll, or four elements) of a headset, determine a measurement uncertainty coefficient Rn+1 based on a historical measurement uncertainty coefficient Rn, send Rn+1 to a Kalman filter, and send, to the Kalman filter, a historical head movement position reported by a mobile phone and a measured head movement position measured by a sensor of the headset.
  • N pieces of measured IMU data for example, including a yaw, a pitch, a roll, or four elements
  • S5.2 Determine a Kalman gain coefficient Kn+1 based on the uncertainty coefficient Rn+1 and estimated uncertainty Pn+1 of a previous iteration, and send the gain coefficient to a current status update module.
  • FIG. 10 is a schematic framework flowchart according to this application.
  • this embodiment provides a compatibility solution for a spatial audio processing scenario in which a mobile phone and a headset perform joint rendering in a non-low-delay case (a first rendering mode) and in a low-delay case (a second rendering mode).
  • the headset In the second rendering mode, the headset no longer performs reverberation rendering, and cannot independently perform spatial audio rendering without the mobile phone, but this mode is characterized by a low requirement for computing power and a simple structure.
  • S1 Add a low-delay mode to a user interface (User Interface, UI) of an APP supported by spatial audio, and deliver, based on selection of a user, a flag indicating whether the low-delay mode is used.
  • UI User Interface
  • the mobile phone performs preprocessing, including downmixing and reverberation processing, on all types of audio signals.
  • S3 Transmit a preprocessed signal to the headset for direct sound rendering in a binaural spatial audio rendering algorithm, and output a rendered signal to a headset speaker.
  • the headset determines an input format; and performs direct sound rendering on a stereo sound source, and outputs a rendered sound source to the headset speaker; or directly outputs a non-stereo sound source to the headset speaker.
  • FIG. 11 is a diagram of a process in a low-delay mode according to this application.
  • the mobile phone preprocesses a received audio source in a format of stereo, 5.1, 7.1, 3DA, or the like.
  • the preprocessing includes downmixing an audio source in a multi-channel format such as 5.1 or 7.1 or in a multi-object format such as 3DA, to unify the audio source into a stereo format.
  • the audio source is transmitted to a first rendering part of the binaural spatial audio rendering algorithm for basic reverberation rendering without head movement effect and audio effector processing.
  • a reverberation part is mixed with downmixed input obtained through an audio effector, and then mixed audio is transmitted to the headset through Bluetooth.
  • the headset receives stereo with reverberation effect that is transmitted by the mobile phone, and performs direct sound rendering by using a second part of the binaural spatial audio rendering algorithm based on IMU head movement data input by a headset sensor.
  • the second part of the binaural spatial audio rendering algorithm includes a direct sound orientation rendering module and an audio effector module.
  • FIG. 12 is a diagram of a process in a non-low-delay mode according to this application.
  • the mobile phone After receiving input sources in various formats, the mobile phone performs direct sound rendering on a non-stereo sound source and reverberation rendering on sound sources in all formats by using a first part of the binaural spatial audio rendering algorithm. Both the direct sound rendering and the reverberation rendering in this part include head movement effect, and an implementation solution in which ER and LR are separately rendered and then added up is used for reverberation.
  • a head movement delay may be optimized with reference to the prediction algorithm in Embodiment 1 to Embodiment 3.
  • the prediction algorithm in Embodiment 2 is used as an example for illustration.
  • the mobile phone processes a direct sound rendering result or stereo raw input and a reverberation rendering result through a corresponding audio effector, mixes audio into a stereo format, and transmits the audio to the headset through Bluetooth.
  • the headset determines a format of raw input; and if the raw input is in a non-stereo format, directly outputs the raw input to the speaker; or if the raw input is in a stereo format, performs direct sound rendering and audio effector processing on the raw input by using a second part of the binaural spatial audio rendering algorithm based on IMU head movement data input by a headset sensor, and then outputs processed audio to the speaker for playing.
  • a UI design is added, and a joint rendering solution is selected based on the user's concern about a head movement delay and spatial audio rendering quality.
  • a rendering algorithm is divided into two parts: the mobile phone and the headset.
  • the mobile phone renders a multi-channel or multi-object part such as 5.1, 7.1, or 3DA, and performs reverberation effect rendering, to fully utilize computing power advantages of the mobile phone.
  • the headset performs direct sound rendering on stereo, to fully utilize advantages of a short IMU data transmission link and a low delay.
  • Head movement position prediction To resolve a problem that the mobile phone has a long head movement delay, a head movement delay prediction algorithm is designed. The prediction algorithm receives, on the headset, a link delay of the mobile phone, a delay of the headset, and a current IMU value, and generates predicted IMU data. Measured IMU data of the headset is transmitted to the second part of the binaural spatial audio rendering algorithm, and the predicted IMU data is transmitted to the first part of the binaural spatial audio rendering algorithm. In this solution, head movement position prediction is implemented, and impact of a link delay can be reduced by more than 1/2.
  • FIG. 13 is a diagram of an example structure of a spatial audio rendering apparatus 1300 according to this application. As shown in FIG. 13 , the spatial audio rendering apparatus 1300 in this embodiment may be used at an audio signal play end.
  • the spatial audio rendering apparatus 1300 may include an obtaining module 1301, a prediction module 1302, a sending module 1303, a receiving module 1304, a rendering module 1305, and a playing module 1306.
  • the obtaining module 1301 is configured to: obtain delay information of an audio content provider end.
  • the prediction module 1302 is configured to: obtain a predicted head movement position based on the delay information.
  • the sending module 1303 is configured to: send first rendering information to the audio content provider end, where the first rendering information includes the predicted head movement position.
  • the prediction module 1302 is specifically configured to: obtain a first historical head movement position corresponding to the delay information; obtain a measured head movement position through a sensor; and obtain the predicted head movement position based on the first historical head movement position and the measured head movement position.
  • the delay information includes a delay sent by the audio content provider end; and the prediction module 1302 is specifically configured to: extract, from a cache, the first historical head movement position corresponding to the delay, where a correspondence between a plurality of delays and a plurality of historical head movement positions is prestored in the cache.
  • the delay information includes a second historical head movement position sent by the audio content provider end, the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, and the second rendering information is sent earlier than the first rendering information; and the prediction module 1302 is specifically configured to: use the second historical head movement position as the first historical head movement position.
  • the prediction module 1302 is specifically configured to: obtain a difference between the first historical head movement position and the measured head movement position; obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; and obtain the predicted head movement position based on the difference and the head movement change rate.
  • the delay information includes a second historical head movement position and a first uncertainty coefficient that are sent by the audio content provider end
  • the second historical head movement position is a measured head movement position in a case in which second rendering information is sent, the first uncertainty coefficient comes from the second rendering information, and the second rendering information is sent earlier than the first rendering information
  • the prediction module 1302 is specifically configured to: use the second historical head movement position as the first historical head movement position.
  • the prediction module 1302 is specifically configured to: obtain a head movement change rate, where the head movement change rate is obtained based on N measured head movement positions that are previously obtained, and N>1; obtain a second uncertainty coefficient based on the head movement change rate and the first uncertainty coefficient; and input the second uncertainty coefficient, the first historical head movement position, and the measured current head movement position to a Kalman filter to obtain the predicted head movement position.
  • the receiving module 1304 is configured to: receive first information sent by the audio content provider end, where the first information includes an audio format, a first audio stream, and the delay information, the audio format indicates that an audio source is in a stereo format or a non-stereo format, and when the audio format indicates that the audio source is in the non-stereo format, the first audio stream includes a first binaural signal rendered by the audio content provider end, or when the audio format indicates that the audio source is in the stereo format, the first audio stream includes the audio source; the rendering module 1305 is configured to: when the audio format indicates that the audio source is in the stereo format, obtain a measured head movement position through the sensor; and render a current frame based on the measured head movement position to obtain a second binaural signal, where the current frame is a frame of the audio source; and the playing module 1306 is configured to: play audio based on a target binaural signal, where the target binaural signal includes the first binaural signal or the second binaural signal.
  • the rendering module 1305 is specifically configured to: perform rendering with head movement effect on a direct sound part in the current frame based on the measured head movement position, and perform rendering without head movement effect on early reflected sound ER and late reverberant sound LR in the current frame, to obtain the second binaural signal.
  • the apparatus in this embodiment may be configured to perform the technical solution performed by the client in the method embodiment shown in FIG. 3 or FIG. 4 .
  • An implementation principle and technical effect of the apparatus are similar. Details are not described herein again.
  • FIG. 14 is a diagram of an example structure of a spatial audio rendering apparatus 1400 according to this application. As shown in FIG. 14 , the spatial audio rendering apparatus 1400 in this embodiment may be used at an audio content provider end.
  • the spatial audio rendering apparatus 1400 may include an obtaining module 1401, a rendering module 1402, and a sending module 1403.
  • the obtaining module 1401 is configured to: obtain a current frame when a first rendering mode is used, where the current frame is a frame of an audio source; and obtain a predicted head movement position when an audio format of the audio source is a non-stereo format.
  • the rendering module 1402 is configured to: render the current frame based on the predicted head movement position to obtain a first binaural signal.
  • the sending module 1403 is configured to: send first information to an audio signal play end, where the first information includes a first audio stream, the audio format of the audio source, and delay information, and the first audio stream includes the first binaural signal.
  • the rendering module 1402 is specifically configured to: separately perform rendering with head movement effect on a direct sound part, an early reflected sound ER part, and a late reverberant sound LR part in the current frame based on the predicted head movement position, to obtain the first binaural signal.
  • the obtaining module 1401 is specifically configured to: obtain the predicted head movement position from rendering information sent by the audio signal play end.
  • the delay information includes a delay.
  • the rendering information further includes a measured head movement position in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information includes the measured head movement position.
  • the rendering information further includes a measured head movement position and a first uncertainty coefficient in a case in which the audio signal play end sends the rendering information; and correspondingly, the delay information includes the measured head movement position and the first uncertainty coefficient.
  • the first audio stream when an audio format of the audio source is a stereo format, the first audio stream includes the audio source.
  • the obtaining module 1401 is further configured to: obtain the current frame when a second rendering mode is used; the rendering module 1402 is further configured to: perform rendering without head movement effect on the current frame to obtain a second binaural signal; and the sending module 1403 is further configured to: send second information to the audio signal play end, where the second information includes a second audio stream, and the second audio stream includes the second binaural signal.
  • the apparatus in this embodiment may be configured to perform the technical solution performed by the client in the method embodiment shown in FIG. 3 or FIG. 4 .
  • An implementation principle and technical effect of the apparatus are similar. Details are not described herein again.
  • the steps in the foregoing method embodiments may be performed by a hardware integrated logic circuit in a processor or by using instructions in a form of software.
  • the processor may be a general-purpose processor, a digital signal processor (digital signal processor, DSP), an application-specific integrated circuit (application-specific integrated circuit, ASIC), a field programmable gate array (field programmable gate array, FPGA) or another programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component.
  • the general-purpose processor may be a microprocessor, or the processor may be any conventional processor or the like.
  • the steps of the methods disclosed in embodiments of this application may be directly performed by a hardware encoding processor, or performed by a combination of hardware and a software module in an encoding processor.
  • the software module may be located in a mature storage medium in the art, for example, a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, or a register.
  • the storage medium is located in a memory, and the processor reads information in the memory and performs the steps of the foregoing methods based on hardware of the processor.
  • the memory in the foregoing embodiments may be a volatile memory or a non-volatile memory, or may include both a volatile memory and a non-volatile memory.
  • the non-volatile memory may be a read-only memory (read-only memory, ROM), a programmable read-only memory (programmable ROM, PROM), an erasable programmable read-only memory (erasable PROM, EPROM), an electrically erasable programmable read-only memory (electrically EPROM, EEPROM), or a flash memory.
  • the volatile memory may be a random access memory (random access memory, RAM), and serves as an external cache.
  • RAMs in many forms may be used, for example, a static random access memory (static RAM, SRAM), a dynamic random access memory (dynamic RAM, DRAM), a synchronous dynamic random access memory (synchronous DRAM, SDRAM), a double data rate synchronous dynamic random access memory (double data rate SDRAM, DDR SDRAM), an enhanced synchronous dynamic random access memory (enhanced SDRAM, ESDRAM), a synchronous link dynamic random access memory (synchlink DRAM, SLDRAM), and a direct rambus random access memory (direct rambus RAM, DR RAM).
  • static random access memory static random access memory
  • DRAM dynamic random access memory
  • DRAM dynamic random access memory
  • SDRAM synchronous dynamic random access memory
  • double data rate SDRAM double data rate SDRAM
  • DDR SDRAM double data rate SDRAM
  • ESDRAM enhanced synchronous dynamic random access memory
  • SLDRAM synchronous link dynamic random access memory
  • direct rambus RAM direct rambus RAM
  • the disclosed system, apparatus, and method may be implemented in other manners.
  • the described apparatus embodiments are merely examples.
  • division into the units is merely logical function division and may be other division during actual implementation.
  • a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not performed.
  • the shown or discussed mutual couplings or direct couplings or communication connections may be implemented through some interfaces.
  • the indirect couplings or communication connections between the apparatuses or units may be implemented in electrical, mechanical, or other forms.
  • the units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, to be specific, may be located in one place, or may be distributed on a plurality of network units. Some or all of the units may be selected according to actual requirements to achieve the objectives of the solutions of embodiments.
  • the functions When the functions are implemented in a form of a software functional unit and sold or used as an independent product, the functions may be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions in this application essentially, or the part contributing to the conventional technology, or some of the technical solutions may be implemented in a form of a software product.
  • the computer software product is stored in a storage medium, and includes several instructions for instructing a computer device (a personal computer, a server, a network device, or the like) to perform all or some of the steps of the methods in embodiments of this application.
  • the storage medium includes any medium that can store program code, for example, a USB flash drive, a removable hard disk drive, a read-only memory (read-only memory, ROM), a random access memory (random access memory, RAM), a magnetic disk, or a compact disc.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Multimedia (AREA)
  • Stereophonic System (AREA)
EP23919449.1A 2023-01-30 2023-11-22 Procédé et appareil de rendu audio spatial Pending EP4648441A4 (fr)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310108164.5A CN118413802A (zh) 2023-01-30 2023-01-30 空间音频渲染方法和装置
PCT/CN2023/133413 WO2024159885A1 (fr) 2023-01-30 2023-11-22 Procédé et appareil de rendu audio spatial

Publications (2)

Publication Number Publication Date
EP4648441A1 true EP4648441A1 (fr) 2025-11-12
EP4648441A4 EP4648441A4 (fr) 2026-05-20

Family

ID=91981953

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23919449.1A Pending EP4648441A4 (fr) 2023-01-30 2023-11-22 Procédé et appareil de rendu audio spatial

Country Status (3)

Country Link
EP (1) EP4648441A4 (fr)
CN (1) CN118413802A (fr)
WO (1) WO2024159885A1 (fr)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120935502B (zh) * 2025-10-14 2026-01-30 歌尔股份有限公司 空间音频处理方法、设备、存储介质及计算机程序产品

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111194561B (zh) * 2017-09-27 2021-10-29 苹果公司 预测性的头部跟踪的双耳音频渲染
CN114531640A (zh) * 2018-12-29 2022-05-24 华为技术有限公司 一种音频信号处理方法及装置
CN112380989B (zh) * 2020-11-13 2023-01-24 歌尔科技有限公司 一种头戴显示设备及其数据获取方法、装置和主机
CN114173256B (zh) * 2021-12-10 2024-04-19 中国电影科学技术研究所 一种还原声场空间及姿态追踪的方法、装置和设备

Also Published As

Publication number Publication date
EP4648441A4 (fr) 2026-05-20
WO2024159885A1 (fr) 2024-08-08
CN118413802A (zh) 2024-07-30

Similar Documents

Publication Publication Date Title
US12596195B2 (en) Electronic device and sensor control method for controlling light sensor based on status of time of flight sensor
CN113228701B (zh) 音频数据的同步方法及设备
US11750926B2 (en) Video image stabilization processing method and electronic device
US12200587B2 (en) Method for implementing functions by using NFC tag, electronic device, and system
US20240048651A1 (en) Message transmission method and corresponding terminal
WO2021043219A1 (fr) Procédé de reconnexion bluetooth et appareil associé
CN114489533A (zh) 投屏方法、装置、电子设备及计算机可读存储介质
US12063607B2 (en) Simultaneous response method and device
US12333206B2 (en) Projection method and related apparatus
WO2022262313A1 (fr) Procédé de traitement d'image à base d'incrustation d'image, dispositif, support de stockage, et produit de programme
US12513471B2 (en) Method for identifying earbud wearing error and related device
US20240289088A1 (en) Volume Adjustment Method, Electronic Device, and System
US20250039599A1 (en) Wearable device, sound pickup method, and apparatus
CN111182140A (zh) 马达控制方法及装置、计算机可读介质及终端设备
CN114466107A (zh) 音效控制方法、装置、电子设备及计算机可读存储介质
CN114449393A (zh) 一种声音增强方法、耳机控制方法、装置及耳机
CN115145517A (zh) 一种投屏方法、电子设备和系统
CN114339429A (zh) 音视频播放控制方法、电子设备和存储介质
US20250168576A1 (en) Signal processing method, apparatus, and device control method and apparatus
CN114257920B (zh) 一种音频播放方法、系统和电子设备
CN113596320B (zh) 视频拍摄变速录制方法、设备、存储介质
WO2024159885A1 (fr) Procédé et appareil de rendu audio spatial
CN113593567A (zh) 视频声音转文本的方法及相关设备
WO2025092283A1 (fr) Procédé de traitement audio, puce et dispositif électronique
CN113923351B (zh) 多路视频拍摄的退出方法、设备和存储介质

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250808

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)