CN117242517A - Audio signal processing method and device, communication equipment, communication system, storage medium - Google Patents

Audio signal processing method and device, communication equipment, communication system, storage medium Download PDF

Info

Publication number
CN117242517A
CN117242517A CN202380010554.7A CN202380010554A CN117242517A CN 117242517 A CN117242517 A CN 117242517A CN 202380010554 A CN202380010554 A CN 202380010554A CN 117242517 A CN117242517 A CN 117242517A
Authority
CN
China
Prior art keywords
window function
value
audio signal
variable
processing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
CN202380010554.7A
Other languages
Chinese (zh)
Inventor
邱云
王宾
刘勇
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Xiaomi Mobile Software Co Ltd
Original Assignee
Beijing Xiaomi Mobile Software Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Xiaomi Mobile Software Co Ltd filed Critical Beijing Xiaomi Mobile Software Co Ltd
Publication of CN117242517A publication Critical patent/CN117242517A/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/02Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
    • G10L19/022Blocking, i.e. grouping of samples in time; Choice of analysis windows; Overlap factoring
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/21Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being power information
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/45Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of analysis window

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Computational Linguistics (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)
  • Telephonic Communication Services (AREA)

Abstract

本公开提出一种音频信号处理方法及装置、通信设备、通信系统、存储介质,该方法包括:确定处理第一音频信号所需的窗函数长度值,并确定第一窗函数,其中,第一窗函数包括:与窗函数长度值相关的参考变量值,参考变量值小于窗函数长度值,参考变量值用于确定恢复第一音频信号所需要的时延;以及根据第一窗函数处理第一音频信号,得到与第一音频信号相关的谱系数。本公开的方法,实现有效地降低音频信号处理的延时。

The present disclosure proposes an audio signal processing method and device, communication equipment, communication system, and storage medium. The method includes: determining a window function length value required to process the first audio signal, and determining the first window function, wherein the first The window function includes: a reference variable value related to the window function length value, the reference variable value is less than the window function length value, the reference variable value is used to determine the time delay required to restore the first audio signal; and processing the first audio signal according to the first window function audio signal, and obtain the spectral coefficients related to the first audio signal. The method of the present disclosure can effectively reduce the delay of audio signal processing.

Description

Audio signal processing method and device, communication equipment, communication system and storage medium
Technical Field
The disclosure relates to the technical field of communication, and in particular relates to an audio signal processing method and device, a communication system and a storage medium.
Background
Digitized audio signals are often large in data size and are not suitable for storage and transmission applications. Therefore, compression coding techniques for digital audio are typically used to compression encode the audio signal. Compression coding technology of digital audio has become a very important audio signal processing technology.
Disclosure of Invention
The embodiment of the disclosure provides an audio signal processing method, an apparatus, a device, a chip system, a storage medium, a computer program and a computer program product, which can be applied to the technical field of communication, and is used for firstly dividing an audio signal into audio signals with a predetermined time period length during audio signal processing, so as to solve the technical problem that delay is introduced in the compression encoding process of the audio signal, thereby influencing the timeliness of audio signal recovery.
The disclosure provides an audio signal processing method and device, a communication system and a storage medium.
According to a first aspect of an embodiment of the present disclosure, there is provided an audio signal processing method, including: determining a window function length value required to process the first audio signal; determining a first window function, wherein the first window function comprises: a reference variable value associated with the window function length value, the reference variable value being less than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal; and processing the first audio signal according to the first window function to obtain spectral coefficients associated with the first audio signal.
According to a second aspect of an embodiment of the present disclosure, there is provided an audio signal processing method, including: acquiring spectral coefficients associated with the first audio signal; determining a first window function, wherein the first window function comprises: a reference variable value associated with a window function length value used to process the first audio signal, the reference variable value being less than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal; and processing the spectral coefficients according to a first window function to obtain a first audio signal.
According to a third aspect of the embodiments of the present disclosure, there is provided an audio signal processing method, including: a first communication device for determining a window function length value required to process a first audio signal and determining a first window function, wherein the first window function comprises: a reference variable value associated with the window function length value, the reference variable value being less than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal, and processing the first audio signal according to the first window function to obtain a spectral coefficient associated with the first audio signal; and the second communication device is used for acquiring spectral coefficients related to the first audio signal and processing the spectral coefficients according to a first window function to obtain the first audio signal.
According to a fourth aspect of the embodiments of the present disclosure, there is provided an audio signal processing apparatus including: a processing module for determining a window function length value required for processing the first audio signal and determining a first window function, wherein the first window function comprises: and a reference variable value associated with the window function length value, the reference variable value being smaller than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal, and processing the first audio signal according to the first window function to obtain spectral coefficients associated with the first audio signal.
According to a fifth aspect of the embodiments of the present disclosure, there is provided an audio signal processing apparatus including: a processing module for obtaining spectral coefficients associated with a first audio signal and determining a first window function, wherein the first window function comprises: and processing the spectral coefficients according to the first window function to obtain the first audio signal.
According to a sixth aspect of the embodiments of the present disclosure, there is provided a communication device, including: one or more processors; wherein the processor is configured to invoke instructions to cause the communication device to perform the audio signal processing method of any of the first, second and third aspects.
According to a seventh aspect of the embodiments of the present disclosure, a communication system is proposed, which is characterized by comprising a first communication device and a second communication device, wherein the first communication device is configured to implement the audio signal processing method of the first aspect, and the second communication device is configured to implement the audio signal processing method of the second aspect.
According to an eighth aspect of an embodiment of the present disclosure, there is provided a storage medium storing instructions, characterized in that the instructions, when executed on a communication device, cause the communication device to perform the audio signal processing method as in any one of the first, second, and third aspects.
Drawings
In order to more clearly illustrate the technical solutions in the embodiments or the background of the present disclosure, the following description will explain the drawings that are required to be used in the embodiments or the background of the present disclosure.
Fig. 1 is a schematic architecture diagram of a communication system shown in accordance with an embodiment of the present disclosure;
fig. 2 is a schematic diagram of windowing a signal in an MDCT block in AVS 3;
FIG. 3A is an interactive schematic diagram of an audio signal processing method according to an embodiment of the disclosure;
FIG. 3B is an interactive schematic diagram of an audio signal processing method according to another embodiment of the disclosure;
fig. 4A is an interactive schematic diagram illustrating an audio signal processing method according to another embodiment of the present disclosure;
fig. 4B is an interactive schematic diagram illustrating an audio signal processing method according to another embodiment of the present disclosure;
fig. 5A is an interactive schematic diagram illustrating an audio signal processing method according to another embodiment of the present disclosure;
Fig. 5B is an interactive schematic diagram illustrating an audio signal processing method according to another embodiment of the present disclosure;
FIG. 6A is a flow chart illustrating a time domain signal encoding process according to one embodiment of the present disclosure;
FIG. 6B is a schematic diagram of a function of a low latency window in an embodiment of the present disclosure;
FIG. 6C is a flow diagram of a decoding process according to an embodiment of the disclosure;
FIG. 6D is a schematic diagram of signal overlap-add in an embodiment of the disclosure;
FIG. 7A is a waveform schematic diagram of Wav test File one in an embodiment of the present disclosure;
FIG. 7B is a schematic diagram of the output signal 1 corresponding to the waveform of Wav test File one in an embodiment of the disclosure;
FIG. 7C is a schematic diagram of the output signal 2 corresponding to the waveform of Wav test File one in an embodiment of the disclosure;
FIG. 7D is a waveform schematic diagram of Wav test File two in an embodiment of the disclosure;
FIG. 7E is a schematic diagram of the output signal 1 corresponding to the waveform of Wav test File two in an embodiment of the disclosure;
FIG. 7F is a schematic diagram of the output signal 2 corresponding to the waveform of Wav test File two in an embodiment of the disclosure;
FIG. 7G is a waveform schematic of Wav test File III in an embodiment of the disclosure;
FIG. 7H is a schematic diagram of the output signal 1 corresponding to the waveform of Wav test File III in an embodiment of the disclosure;
FIG. 7I is a schematic diagram of output signal 2 corresponding to the waveform of Wav test File III in an embodiment of the disclosure;
fig. 8A is a schematic structural diagram of an audio signal processing apparatus according to an embodiment of the present disclosure;
fig. 8B is a schematic structural diagram of an audio signal processing apparatus according to an embodiment of the present disclosure;
fig. 9A is a schematic structural diagram of a communication device according to an embodiment of the present disclosure;
fig. 9B is a schematic structural diagram of a chip according to an embodiment of the disclosure.
Detailed Description
The embodiment of the disclosure provides an audio signal processing method and device, communication equipment, a communication system and a storage medium. In some embodiments, terms of an audio signal processing method and an information processing method, a communication method, and the like may be replaced with each other, terms of an audio signal processing apparatus and an information processing apparatus, a communication apparatus, and the like may be replaced with each other, and terms of an information processing system, a communication system, and the like may be replaced with each other.
The embodiments of the present disclosure are not intended to be exhaustive, but rather are exemplary of some embodiments and are not intended to limit the scope of the disclosure. In the case of no contradiction, each step in a certain embodiment may be implemented as an independent embodiment, and the steps may be arbitrarily combined, for example, a scheme in which part of the steps are removed in a certain embodiment may also be implemented as an independent embodiment, the order of the steps in a certain embodiment may be arbitrarily exchanged, and further, alternative implementations in a certain embodiment may be arbitrarily combined; furthermore, various embodiments may be arbitrarily combined, for example, some or all steps of different embodiments may be arbitrarily combined, and an embodiment may be arbitrarily combined with alternative implementations of other embodiments.
In the various embodiments of the disclosure, terms and/or descriptions of the various embodiments are consistent throughout the various embodiments and may be referenced to each other in the absence of any particular explanation or logic conflict, and features from different embodiments may be combined to form new embodiments in accordance with their inherent logic relationships.
The terminology used in the embodiments of the disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure.
In the presently disclosed embodiments, elements that are referred to in the singular, such as "a," "an," "the," "said," etc., may mean "one and only one," or "one or more," "at least one," etc., unless otherwise indicated. For example, where an article (article) is used in translation, such as "a," "an," "the," etc., in english, a noun following the article may be understood as a singular expression or as a plural expression.
In the presently disclosed embodiments, "plurality" refers to two or more.
In some embodiments, terms such as "at least one of", "one or more of", "multiple of" and the like may be substituted for each other.
Description modes such as at least one of A, B, C … …, A and/or B and/or C … … include any single case of A, B, C … … and any combination case of any plurality of A, B, C … …, and each case may exist independently; for example, "at least one of A, B, C" includes the cases of a alone, B alone, C, A and B in combination, a and C in combination, B and C in combination, a and B and C in combination; for example, a and/or B includes the case of a alone, a combination of a alone B, A and B.
In some embodiments, "in a case a, in another case B", "in response to a case a", "in response to another case B", and the like, the following technical solutions may be included according to the circumstances: a is performed independently of B, i.e., a in some embodiments; b is performed independently of a, i.e., in some embodiments B; a and B are selectively performed, i.e., in some embodiments selected from a and B; both a and B are performed, i.e., a and B in some embodiments. Similar to that described above when there are more branches such as A, B, C.
The prefix words "first", "second", etc. in the embodiments of the present disclosure are only for distinguishing different description objects, and do not limit the location, order, priority, number, content, etc. of the description objects, and the statement of the description object refers to the claims or the description of the embodiment context, and should not constitute unnecessary limitations due to the use of the prefix words. For example, if the description object is a "field", the ordinal words before the "field" in the "first field" and the "second field" do not limit the position or the order between the "fields", and the "first" and the "second" do not limit whether the "fields" modified by the "first" and the "second" are in the same message or not. For another example, describing an object as "level", ordinal words preceding "level" in "first level" and "second level" do not limit priority between "levels". As another example, the number of descriptive objects is not limited by ordinal words, and may be one or more, taking "first device" as an example, where the number of "devices" may be one or more. Furthermore, objects modified by different prefix words may be the same or different, e.g., the description object is "a device", then "a first device" and "a second device" may be the same device or different devices, and the types may be the same or different; for another example, the description object is "information", and the "first information" and the "second information" may be the same information or different information, and the contents thereof may be the same or different.
In some embodiments, "comprising a", "containing a", "for indicating a", "carrying a", may be interpreted as carrying a directly, or as indicating a indirectly.
In some embodiments, terms "responsive to … …", "responsive to determination … …", "in the case of … …", "at … …", "when … …", "if … …", "if … …", and the like may be interchanged.
In some embodiments, terms "greater than", "greater than or equal to", "not less than", "more than or equal to", "not less than", "above" and the like may be interchanged, and terms "less than", "less than or equal to", "not greater than", "less than or equal to", "not more than", "below", "lower than or equal to", "no higher than", "below" and the like may be interchanged.
In some embodiments, an apparatus or the like may be interpreted as an entity, or may be interpreted as a virtual, and the names thereof are not limited to the names described in the embodiments, "apparatus," "device," "circuit," "network element," "node," "function," "unit," "section," "system," "network," "chip system," "entity," "body," and the like may be replaced with each other.
In some embodiments, a "network" may be interpreted as an apparatus (e.g., access network device, core network device, etc.) contained in a network.
In some embodiments, "access network device (access network device, AN device)", "radio access network device (radio access network device, RAN device)", "Base Station (BS)", "radio base station (radio base station)", "fixed station (fixed station)", "node (node)", "access point (access point)", "transmit point (transmission point, TP)", "Receive Point (RP)", "transmit receive point (transmit/receive point), the terms TRP), panel, antenna array, cell, macrocell, microcell, femtocell, sector, cell group, carrier, component carrier, bandwidth part, BWP, etc. may be replaced with each other.
In some embodiments, "terminal," terminal device, "" user equipment, "" user terminal, "" mobile station, "" mobile terminal, MT) ", subscriber station (subscriber station), mobile unit (mobile unit), subscriber unit (subscriber unit), wireless unit (wireless unit), remote unit (remote unit), mobile device (mobile device), wireless device (wireless device), wireless communication device (wireless communication device), remote device (remote device), mobile subscriber station (mobile subscriber station), access terminal (access terminal), mobile terminal (mobile terminal), wireless terminal (wireless terminal), remote terminal (remote terminal), handheld device (handset), user agent (user agent), mobile client (mobile client), client (client), and the like may be substituted for each other.
In some embodiments, the acquisition of data, information, etc. may comply with laws and regulations of the country of locale.
In some embodiments, data, information, etc. may be obtained after user consent is obtained.
Furthermore, each element, each row, or each column in the tables of the embodiments of the present disclosure may be implemented as a separate embodiment, and any combination of elements, any rows, or any columns may also be implemented as a separate embodiment.
The correspondence relationships shown in the tables in the present disclosure may be configured or predefined. The values of the information in each table are merely examples, and may be configured as other values, and the present disclosure is not limited thereto. In the case of the correspondence between the configuration information and each parameter, it is not necessarily required to configure all the correspondence shown in each table. For example, in the table in the present disclosure, the correspondence shown by some rows may not be configured. For another example, appropriate morphing adjustments, e.g., splitting, merging, etc., may be made based on the tables described above. The names of the parameters indicated in the tables may be other names which are understood by the communication device, and the values or expressions of the parameters may be other values or expressions which are understood by the communication device. When the tables are implemented, other data structures may be used, for example, an array, a queue, a container, a stack, a linear table, a pointer, a linked list, a tree, a graph, a structure, a class, a heap, a hash table, or a hash table.
Predefined in this disclosure may be understood as defining, predefining, storing, pre-negotiating, pre-configuring, curing, or pre-sintering.
Fig. 1 is a schematic architecture diagram of a communication system shown in accordance with an embodiment of the present disclosure. As shown in fig. 1, a communication system 100 may include a terminal (terminal) 101, a network device 102. The network device 102 may include at least one of an access network device and a core network device (core network device).
In some embodiments, the terminal 101 includes at least one of a mobile phone (mobile phone), a wearable device, an internet of things device, a communication enabled car, a smart car, a tablet (Pad), a wireless transceiver enabled computer, a Virtual Reality (VR) terminal device, an augmented reality (augmented reality, AR) terminal device, a wireless terminal device in industrial control (industrial control), a wireless terminal device in unmanned (self-driving), a wireless terminal device in teleoperation (remote medical surgery), a wireless terminal device in smart grid (smart grid), a wireless terminal device in transportation security (transportation safety), a wireless terminal device in smart city (smart city), a wireless terminal device in smart home (smart home), for example, but is not limited thereto.
In some embodiments, the access network device is, for example, a node or device that accesses a terminal to a wireless network, and the access network device may include at least one of an evolved NodeB (eNB), a next generation evolved NodeB (next generation eNB, ng-eNB), a next generation NodeB (next generation NodeB, gNB), a NodeB (node B, NB), a Home NodeB (HNB), a home NodeB (home evolved nodeB, heNB), a wireless backhaul device, a radio network controller (radio network controller, RNC), a base station controller (base station controller, BSC), a base transceiver station (base transceiver station, BTS), a baseband unit (BBU), a mobile switching center, a base station in a 6G communication system, an Open base station (Open RAN), a Cloud base station (Cloud RAN), a base station in other communication systems, an access node in a WiFi system, but is not limited thereto.
In some embodiments, the technical solutions of the present disclosure may be applied to an Open RAN architecture, where an access network device or an interface in an access network device according to the embodiments of the present disclosure may become an internal interface of the Open RAN, and flow and information interaction between these internal interfaces may be implemented by using software or a program.
In some embodiments, the access network device may be composed of a Central Unit (CU) and a Distributed Unit (DU), where the CU may also be referred to as a control unit (control unit), and the structure of the CU-DU may be used to split the protocol layers of the access network device, where functions of part of the protocol layers are centrally controlled by the CU, and functions of the rest of all the protocol layers are distributed in the DU, and the DU is centrally controlled by the CU, but is not limited thereto.
In some embodiments, the core network device may be a device, including one or more network elements, or may be a plurality of devices or groups of devices, each including all or part of one or more network elements. The network element may be virtual or physical. The core network comprises, for example, at least one of an evolved packet core (Evolved Packet Core, EPC), a 5G core network (5G Core Network,5GCN), a next generation core (Next Generation Core, NGC).
It may be understood that, the communication system described in the embodiments of the present disclosure is for more clearly describing the technical solutions of the embodiments of the present disclosure, and is not limited to the technical solutions provided in the embodiments of the present disclosure, and those skilled in the art can know that, with the evolution of the system architecture and the appearance of new service scenarios, the technical solutions provided in the embodiments of the present disclosure are applicable to similar technical problems.
The embodiments of the present disclosure described below may be applied to the communication system 100 shown in fig. 1, or a part of the main body, but are not limited thereto. The respective bodies shown in fig. 1 are examples, and the communication system may include all or part of the bodies in fig. 1, or may include other bodies than fig. 1, and the number and form of the respective bodies are arbitrary, and the connection relationship between the respective bodies is examples, and the respective bodies may be not connected or may be connected, and the connection may be arbitrary, direct connection or indirect connection, or wired connection or wireless connection.
The embodiments of the present disclosure may be applied to long term evolution (Long Term Evolution, LTE), LTE-Advanced (LTE-a), LTE-Beyond (LTE-B), upper 3G, IMT-Advanced, fourth generation mobile communication system (4th generation mobile communication system,4G)), fifth generation mobile communication system (5th generation mobile communication system,5G), 5G New air (New Radio, NR), future wireless access (Future Radio Access, FRA), new wireless access technology (New-Radio Access Technology, RAT), new wireless (New Radio, NR), new wireless access (New Radio access, NX), future generation wireless access (Future generation Radio access, FX), global System for Mobile communications (GSM (registered trademark)), CDMA2000, ultra mobile broadband (Ultra Mobile Broadband, UMB), IEEE 802.11 (registered trademark), IEEE 802.16 (WiMAX (registered trademark)), IEEE 802.20, ultra WideBand (Ultra-wide bandwidth, UWB), bluetooth (Bluetooth) mobile communication network (Public Land Mobile Network, PLMN, device-D-Device, device-M, device-M, internet of things system, internet of things (internet of things), machine-2, device-M, device-M, internet of things (internet of things), system (internet of things), internet of things 2, device (internet of things), machine (internet of things), etc. In addition, a plurality of system combinations (e.g., LTE or a combination of LTE-a and 5G, etc.) may be applied.
Digitized audio signals are often large in data size and are not suitable for storage and transmission applications. Therefore, compression coding techniques for digital audio are typically used to compression encode the audio signal. Compression coding technology of digital audio has become a very important audio signal processing technology.
Alternatively, compression coding techniques for digital audio, which may be for example modified discrete cosine transforms (Modified Discrete Cosine Transform, MDCT), are an algorithm for multimedia signal compression. MDCT is a theory based on time-frequency decomposition for dividing a frequency domain signal into a plurality of subbands for analysis and processing. The time domain signal may be divided into a series of time segments of predetermined time segment length using a window function and each time segment signal is discrete cosine transformed (Discrete Cosine Transform, DCT) and then formed into a frequency spectrum. Due to the use of the window function, the audio artifacts caused by DCT abrupt change can be reduced, and the quality of the audio can be improved.
Alternatively, advanced audio coding (Advanced Audio Coding, AAC), moving picture experts compression standard audio layer 3 (Moving Picture Experts Group Audio Layer III, MP 3), AC-3 (AC-3 is a digital multimedia technology), and so on, all use MDCT algorithms for audio signal compression. And MDCT technology is widely used for video signal compression in video coding standards such as h.264 (h.264 is a highly compressed digital video codec standard) and high efficiency video coding (High Efficiency Video Coding, HEVC). In addition, MDCT has been widely used in the fields of spectral analysis, artificial speech synthesis, images, and the like.
As shown in fig. 2, fig. 2 is a schematic diagram of windowing a signal in an MDCT block in AVS3, wherein AVS3 is an audio and video source coding standard For N point sequence X (N) The mathematical expression of MDCT is
Wherein the method comprises the steps of x (n) represents the input signal
The mathematical expression of IMDCT is:
wherein n=0, 1,...
The window functions commonly used in audio codec are:
where w (n) represents a window function.
Alternatively, most audio encoders window the signal first and then MDCT, exemplified by AVS3, with mainly 4 types of windows, where the long and short windows are switching windows, also called transition windows. When the information is abrupt, the MDCT process may be performed using a short window, and at this time, a transition may be made from a long window to a short window, and a transition window may be used. There are also two types of windows, namely a long window and a short window, which are the most commonly used windows in audio coding, the long window being typically 2048 points and the short window being 256 points. Taking a long window as an example, the signal is subjected to windowing and MDCT processing at an encoding end, IMDCT and windowing at a decoding end, and 50% overlap addition is performed on the signal by using a time domain aliasing cancellation technique (Time Domain Aliasing Cancellation, TDAC), so that the signal is perfectly reconstructed.
Alternatively, the signals may be overlap-added in order to be able to reconstruct the signal perfectly. For example, 2N points may be used to perform MDCT processing to obtain N spectral coefficients. Wherein the composition of 2N points is N points of the previous frame and N points of the current frame, as shown in fig. 2, when the decoding end recovers the signal, N points are obtained by performing overlap-add on the N points of the previous frame and the N points of the current frame, and the window length is 2N, that is, there is 50% overlap-add. However, when 2N points of the current frame are processed, only N points of the current frame can be restored, and the remaining N points are restored in the next frame. At this point, the resulting algorithm delay is N. Thus, in AVS3, an inherent algorithmic delay N is created by windowing. That is, the compression encoding process of the audio signal introduces delay, thereby affecting the timeliness of the audio signal recovery.
According to the method, the time delay introduced in the compression encoding process of the audio signal is reduced through the changed window shape and parameters, so that the timeliness of the audio signal recovery can be effectively improved.
The method shown in the embodiment of the disclosure can be applied to various fields of audio coding, video coding, audio synthesis and the like, and can be applied to application scenes such as analysis and processing of signals and images, modulation and demodulation of communication, mathematical analysis and the like.
Fig. 3A is an interactive schematic diagram illustrating an audio signal processing method according to an embodiment of the present disclosure. As shown in fig. 3A, the embodiment of the present disclosure relates to an audio signal processing method, which may be used in the communication system 100, where the embodiment may be applied to a process of encoding a first audio signal, where the terminal 101 and/or the network device 102 in the communication system 100 may perform encoding processing on the first audio signal. Of course, any other possible device, apparatus, component, module, etc. having an encoding processing function may also be used to encode the first audio signal using the audio signal processing method provided in the embodiments of the present disclosure. The method comprises the following steps:
in step S3101, an energy threshold is determined.
In some embodiments, the energy threshold refers to a threshold that determines the energy level of each sub-signal block in the signal. For example, if the energy of the sub-signal block is greater than or equal to the energy threshold, it indicates that the energy of the sub-signal block is relatively large, and if the energy of the sub-signal block is less than the energy threshold, it indicates that the energy of the sub-signal block is relatively small.
In some embodiments, the energy threshold may be adaptively adjusted according to audio signal processing requirements, or may be based on protocol conventions.
Step S3102, an energy value corresponding to each sub-signal block in the first audio signal is determined.
In some embodiments, the first audio signal refers to a signal to be processed. The first audio signal may be divided to obtain a plurality of sub-signal blocks, and the energy condition of each sub-signal block may be analyzed separately. The quantized energy value may be used to express the energy condition of each sub-signal block. The energy value of each sub-signal block may be used to determine a window function length value required to process the first audio signal, which window function length value determines whether to process the first audio signal with a long window or with a short window.
In some embodiments, the energy value of each sub-signal block may be analyzed based on any possible energy analysis method.
In some embodiments, the sub-signal blocks may also be referred to as sub-blocks.
Step S3103, determining a window function length value according to the energy threshold and the plurality of energy values.
In some embodiments, the window function length value may be represented by N, and the window function length value may be used to represent the size of the window.
In some embodiments, after determining the energy threshold and the energy value for each sub-signal block, the window function length value may be determined with reference to the energy threshold and the plurality of energy values.
In some embodiments, the magnitude of each energy value and energy threshold may be compared, and an appropriate window function length value may be selected based on the magnitude comparison.
In some embodiments, the size of the window may be determined. By way of example, the size of the window may be derived by comparing the energy threshold to the size of each sub-block energy (an alternative example of the energy value of a sub-signal block). For example, when one sub-block energy is greater than the energy threshold, a short window is determined, and when all sub-block energies are less than or equal to the energy threshold, a long window is determined, wherein the short window and the long window may be distinguished by different window function length values. For example, the window function length value of the long window may be a first window function length value, the window function length value of the short window may be a second window function length value, and the second window function length value may be substantially smaller than the first window function length value.
In some embodiments, if the first audio signal is divided into a plurality of sub-signal blocks, where a plurality of energy values corresponding to the plurality of sub-signal blocks are all less than or equal to the energy threshold, determining the window function length value as the first window function length value indicates that the first audio signal can be processed based on a long window. In some embodiments, if the first audio signal is divided into a plurality of sub-signal blocks, where at least one energy value of a plurality of energy values respectively corresponding to the plurality of sub-signal blocks is greater than an energy threshold, determining that the window function length value is a second window function length value, where the second window function length value is smaller than the first window function length value, that is, indicates that the first audio signal can be processed based on a short window.
Therefore, through carrying out energy analysis on the first audio signal and determining a proper window function length value for the processing of the first audio signal, the purpose of selecting a proper window to cut off the first audio signal is achieved, and the accuracy and the coding processing effect of the coding processing of the first audio signal are improved.
Step S3104, determining a first window function according to the window function length value, the reference variable value associated with the window function length value, the first variable, and the second window function.
In some embodiments, the first window function is a window function used in embodiments of the present disclosure to truncate the first audio signal.
In some embodiments, the first window function includes: and a reference variable value associated with the window function length value, the reference variable value being smaller than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal.
In some embodiments, the second window function may be, for example, an existing window function. The second window function may specifically be, for example, the window function commonly used in audio codec:
where w (n) represents the second window function.
It will be appreciated that after the first audio signal is truncated using a window function commonly used in audio codec and transformed in the frequency domain, if the first audio signal is to be recovered, there will typically be a certain delay, typically the delay is associated with N, the greater the delay, and the smaller the N, the smaller the delay. In order to effectively reduce the time delay required to recover the first audio signal, certain improvements to the window function described above may be made in embodiments of the present disclosure. For example, the window function length value, the reference variable value related to the window function length value, and the first variable may be referred to for performing a certain improvement on the second window function, and the window function obtained by the improvement is taken as the first window function.
In some embodiments, the window function length value may be represented by N, the reference variable value associated with the window function length value may be represented by X, the first variable may be represented by N, the first variable may also be referred to as a sequence number of the discrete time axis, the reference variable value is smaller than the window function length value (i.e., X is smaller than N), and the reference variable value is used to determine the time delay required to recover the first audio signal. That is, the improved window function in the embodiments of the present disclosure, the reference variable value X is related to the delay required to recover the first audio signal (i.e., the delay required to recover the first audio signal is determined by the reference variable value X rather than the window function length value N), and since the reference variable value X is less than the window function length value N. As the time delay required to recover the first audio signal is reduced. Thus, the method is applicable to a variety of applications. By changing the shape and parameters of the window, the delay introduced in the compression encoding process of the audio signal is reduced, so that the timeliness of the audio signal recovery can be effectively improved.
In the embodiment of the disclosure, a certain improvement treatment is performed on a second window function through a reference window function length value, a reference variable value related to the window function length value and a first variable, and the window function obtained through the improvement treatment is used as a first window function. The realization can effectively avoid introducing excessive operation expenditure while supporting the reduction of the time delay required for recovering the first audio signal as much as possible,
In some embodiments, the first window function satisfies a first condition, the first condition comprising: the sum of the first value and the second value is 1; wherein the first value is the square of the value obtained by processing the first variable based on the first window function, the second value is the square of the value obtained by processing the second variable based on the first window function, and the second variable is the sum between the first variable and the window function length value.
In some embodiments, the first condition is a condition that can support reducing the delay required to recover the first audio signal and that supports the first window function that the first audio signal can be effectively recovered needs to satisfy. The first variable is N, the second variable is n+n, the first value is the square of the value obtained by processing the first variable based on the first window function, the second value is the square of the value obtained by processing the second variable based on the first window function, and the sum of the first value and the second value needs to be 1.
In some embodiments, by configuring the first condition, the first window function can be flexibly configured and processed based on the first condition, so that the method is effectively applicable to personalized audio signal processing scenes, and signal coding processing flexibility and coding processing effect are improved.
In some embodiments, the first condition is the following formula: w (n) 2 +W(n+N) 2 =1; where N represents a first variable, N represents a window function length value, and W () represents a first window function, so that the configuration convenience of the first condition can be improved.
In some embodiments, the first window function includes at least one of: a first function portion corresponding to a first variable, wherein the first variable belongs to a first value range; a second function portion corresponding to the first variable, wherein the first variable belongs to a second range of values; a third function portion corresponding to the first variable, wherein the first variable belongs to a third range of values; a fourth function portion corresponding to the first variable, wherein the first variable belongs to a fourth range of values; wherein, there is no intersection part between different value ranges, and the value range is obtained by dividing the reference value range according to the window function length value and the reference variable value, the minimum value of the reference value range is zero, the maximum value is twice the window function length, different function parts are different, at least one function part is related to the second window function, and at least one function part is a constant value. After the first audio signal is truncated based on the improved first window function, the time delay introduced in the compression encoding process of the audio signal is reduced, and therefore timeliness of audio signal recovery can be effectively improved.
In some embodiments, the first window function may be a piecewise function, including different functional portions, where the first variable n is located in a different range of values, and the different functional portions may be correspondingly used. A first value range, e.g., 0.ltoreq.n < N-X, i.e., the first variable N is greater than or equal to 0 and less than the difference between the window function length value and the reference variable value; a second value range, e.g., N-x.ltoreq.n < N, i.e., the first variable N is greater than or equal to the difference between the window function length value and the reference variable value and less than the window function length value; a third range of values, for example, n.ltoreq.n <2N-X, i.e., the first variable N is greater than or equal to the window function length value and less than twice the difference between the window function length value and the reference variable value; a fourth range of values, e.g., 2N-X.ltoreq.n <2N, i.e., the first variable N is greater than or equal to twice the difference between the window function length value and the reference variable value and less than twice the window function length value; correspondingly, the first window function may specifically correspond to different function portions corresponding to the different value ranges. Illustratively, the first function portion may be a constant value of 0 corresponding to the first range of values; corresponding to the second range of values, the second function portion may be w (N-n+x); corresponding to the third range of values, the third function portion may be a constant value of 1; the fourth function portion may be w (n-2n+2x) corresponding to the fourth range of values.
In some embodiments, the first window function may be specifically, for example, the following formula:
wherein N represents a first variable, N represents a window function length value, X represents a reference variable value, W D () The first window function is represented, and w () represents the second window function, so that the first window function can be accurately and clearly expressed, so that the first window function can be accurately used for performing a truncation process on the first audio signal.
In some embodiments, the reference variable value is related to the window function length value, that is, the reference variable value may be determined with reference to the window function length value required for processing the first audio signal, and the relation between the reference variable value and the window function length value may be expressed by the following formula:wherein N represents a window function length value, X represents a reference variable value, and m is a positive integer. />
As can be seen from the above formula, the reference variable value is smaller than the window function length value, and the existing second window function is modified to obtain the first window function, so that the first window function contains the reference variable value, and the delay required for recovering the first audio signal is determined by the reference variable value. Therefore, the time delay introduced in the compression encoding process of the audio signal is reduced, and the timeliness of the audio signal recovery can be effectively improved.
In step S3105, the first audio signal is processed according to the first window function, resulting in an intermediate audio signal.
In some embodiments, after determining the first window function, the first audio signal may be truncated using the first window function, and the resulting signal may be referred to as an intermediate audio signal. For example, the first window function and the first audio signal may be multiplied to obtain an intermediate audio signal.
In some embodiments, the first audio signal is processed according to a first window function, which may also be referred to as windowing. For example, the input signal (an alternative example of the first audio signal) is multiplied by a window function (an alternative example of the first window function), i.e. X (n) =w (n) ×input (n), where W (n) is the first window function, input (n) is the input signal, and X (n) is the windowed signal (an alternative example of the intermediate audio signal).
Step S3106, processing the intermediate audio signal according to the first processing information to obtain a spectral coefficient related to the first audio signal, where the first processing information includes: reference variable values.
In some embodiments, the intermediate audio signal may be processed using first processing information to obtain spectral coefficients associated with the first audio signal, the first processing information being used to perform a modified discrete cosine transform process on the intermediate audio signal.
In some embodiments, the first processing information refers to information required to perform a modified discrete cosine transform process.
Since the first window function is a modified window function, the first window function comprises reference variable values, which may also be included in the first processing information in order to efficiently determine spectral coefficients associated with the first audio signal from the intermediate audio signal, such that the first processing information may be efficiently adapted to a transform process of the intermediate audio signal.
In some embodiments, the first process information is the following formula:
where k=0, 1..n-1, N represents the first variable, N represents the window function length value, X represents the reference variable value, X (N) represents the intermediate audio signal, and X (k) represents the spectral coefficient. Thereby, the reference variable value is contained in the modified discrete cosine transform formula, so that the modified discrete cosine transform formula can be effectively utilizedIs suitable for performing a modified discrete cosine transform process on the intermediate audio signal.
In some embodiments, after processing the intermediate audio signal according to the first processing information to obtain the spectral coefficients associated with the first audio signal, the encoding side may transmit the spectral coefficients associated with the first audio signal to the decoding side, which may recover the first audio signal according to the spectral coefficients associated with the first audio signal.
The audio signal processing method according to the embodiment of the present disclosure may include at least one of step S3101 to step S3106. For example, step S3101 may be implemented as a separate embodiment, step S3102 may be implemented as a separate embodiment, and so on, but is not limited thereto. Step s3101+s3102 may be implemented as a separate embodiment and step s3101+s3102+s3103 may be implemented as a separate embodiment, but is not limited thereto.
In this embodiment mode or example, the steps may be independently, arbitrarily combined, or exchanged in order, and the alternative modes or examples may be arbitrarily combined, and may be arbitrarily combined with any steps of other embodiment modes or other examples without contradiction.
In this embodiment, by determining a window function length value required for processing the first audio signal, and determining a first window function, the first window function includes: and a reference variable value associated with the window function length value, the reference variable value being smaller than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal, and processing the first audio signal according to the first window function to obtain spectral coefficients associated with the first audio signal. By changing the shape and parameters of the window, the delay introduced in the compression encoding process of the audio signal is reduced, so that the timeliness of the audio signal recovery can be effectively improved.
Fig. 3B is an interactive schematic diagram illustrating an audio signal processing method according to another embodiment of the present disclosure. As shown in fig. 3B, the embodiment of the present disclosure relates to an audio signal processing method, which may be used in the communication system 100, where the present embodiment may be applied to a process of decoding a first audio signal, where the terminal 101 and/or the network device 102 in the communication system 100 may perform decoding processing on the first audio signal. Of course, any other possible device, apparatus, component, module, etc. having a decoding processing function may also perform decoding processing on the first audio signal using the audio signal processing method provided in the embodiments of the present disclosure. In addition, the encoding process and the decoding process may be performed by the same apparatus or may be performed by different apparatuses, which is not limited thereto. The method comprises the following steps:
in step S3201, spectral coefficients associated with the first audio signal are acquired.
In some embodiments, spectral coefficients obtained by the encoding end in the encoding process of the first audio signal may be obtained. For example, the encoding end may transmit the spectral coefficients related to the first audio signal to the decoding end after processing the intermediate audio signal according to the first processing information, thereby the decoding end may acquire the spectral coefficients related to the first audio signal from the encoding end and restore the first audio signal according to the spectral coefficients related to the first audio signal.
In step S3202, a first window function is determined according to a window function length value used for processing the first audio signal, a reference variable value associated with the window function length value, a first variable, and a second window function.
In some embodiments, the reference variable value is less than the window function length value, the reference variable value being used to determine the time delay required to recover the first audio signal. In this embodiment, the description of the window function length value, the reference variable value related to the window function length value, the first variable, and the second window function used for processing the first audio signal, and the description of determining the first window function may be specifically referred to the above embodiments, which are not repeated herein.
Step S3203, processing the spectral coefficients according to second processing information to obtain an intermediate audio signal, wherein the second processing information comprises: reference variable values.
In some embodiments, the second processing information refers to information required to perform an inverse modified discrete cosine transform process.
In some embodiments, the spectral coefficients may be processed using second processing information to obtain the intermediate audio signal, the second processing information being used to perform an inverse modified discrete cosine transform process on the spectral coefficients.
Since the first window function is a modified window function, the first window function comprises reference variable values, in order to effectively recover the intermediate audio signal from the spectral coefficients associated with the first audio signal, the reference variable values may also be included in the second processing information such that the second processing information can be effectively adapted to the transform processing of the spectral coefficients associated with the first audio signal.
In some embodiments, the second process information is the following formula:
where n=0, 1,..2N-1, N represents the first variable, N represents the window function length value, X represents the reference variable value, X (N) represents the intermediate audio signal, and X (k) represents the spectral coefficient. Thereby, the inclusion of the reference variable value in the formula of the inverse modified discrete cosine transform is achieved such that the formula of the inverse modified discrete cosine transform can be effectively adapted to the inverse modified discrete cosine transform processing of spectral coefficients associated with the first audio signal.
In step S3204, the intermediate audio signal is processed according to the first window function, resulting in a first audio signal.
In some embodiments, after processing spectral coefficients associated with the first audio signal using the second processing information described above, the intermediate audio signal may be windowed to recover the first audio signal. Illustratively, the intermediate audio signal is windowed, i.e. the input signal (e.g. the intermediate audio signal) is multiplied by a window function (an alternative example of a first window function), the formula being as follows: i.e., X (n) =w (n) ×input (n), where W (n) is the first window function, input (n) is the input signal, and X (n) is the windowed signal.
In some embodiments, the windowed signals may be overlap-added to recover the first audio signal. For example, the X points of the previous frame may be overlap-added to the N-X to N points of the current frame to recover X points, the N-X points of the current frame are added (the N-X points are not overlap-added), and a total of N points are recovered, while the X points of the 2N-X to 2N portion of the current frame are recovered to the next frame, so the delay is X.
The audio signal processing method according to the embodiment of the present disclosure may include at least one of step S3201 to step S3204. For example, step S3201 may be implemented as a separate embodiment, step S3202 may be implemented as a separate embodiment, and so on, but is not limited thereto. The steps s3201+s3202 may be implemented as an independent embodiment, and the steps s3201+s3202+s3203 may be implemented as an independent embodiment, but is not limited thereto.
In this embodiment mode or example, the steps may be independently, arbitrarily combined, or exchanged in order, and the alternative modes or examples may be arbitrarily combined, and may be arbitrarily combined with any steps of other embodiment modes or other examples without contradiction.
In this embodiment, the spectral coefficients associated with the first audio signal are obtained, and a first window function is determined, where the first window function includes: and processing the spectral coefficients according to the first window function to obtain the first audio signal. By changing the shape and parameters of the window, the delay introduced in the compression encoding process of the audio signal is reduced, so that the timeliness of the audio signal recovery can be effectively improved.
Fig. 4A is an interactive schematic diagram illustrating an audio signal processing method according to another embodiment of the present disclosure. As shown in fig. 4A, an embodiment of the present disclosure relates to an audio signal processing method, which may be used in an encoding process, where the method includes:
in step S4101, a window function length value required for processing the first audio signal is determined.
Step S4102, determining a first window function, wherein the first window function comprises: and a reference variable value associated with the window function length value, the reference variable value being smaller than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal.
In step S4103, the first audio signal is processed according to the first window function, resulting in spectral coefficients associated with the first audio signal.
The audio signal processing method according to the embodiment of the present disclosure may include at least one of step S4101 to step S4103. For example, step S4101 may be implemented as a separate embodiment, step S4102 may be implemented as a separate embodiment, and so on, but is not limited thereto. Step S4101+s4102 may be implemented as a separate embodiment, and step S4101+s4102+s4103 may be implemented as a separate embodiment, but is not limited thereto.
In this embodiment mode or example, the steps may be independently, arbitrarily combined, or exchanged in order, and the alternative modes or examples may be arbitrarily combined, and may be arbitrarily combined with any steps of other embodiment modes or other examples without contradiction.
Fig. 4B is an interactive schematic diagram illustrating an audio signal processing method according to another embodiment of the present disclosure. As shown in fig. 4B, an embodiment of the present disclosure relates to an audio signal processing method, which may be used in an encoding process, where the method includes:
in step S4201, an energy threshold is determined.
In step S4202, an energy value corresponding to each sub-signal block in the first audio signal is determined.
In step S4203, a window function length value is determined based on the energy threshold and the plurality of energy values.
Step S4204, determining a first window function according to the window function length value, the reference variable value associated with the window function length value, the first variable, and the second window function.
In step S4205, the first audio signal is processed according to the first window function to obtain an intermediate audio signal.
Step S4206, processing the intermediate audio signal according to the first processing information to obtain spectral coefficients related to the first audio signal, wherein the first processing information includes: reference variable values.
The audio signal processing method according to the embodiment of the present disclosure may include at least one of step S4201 to step S4206. For example, step S4201 may be implemented as a stand-alone embodiment, step S4202 may be implemented as a stand-alone embodiment, and so on, but is not limited thereto. Step S4201+s4202 may be implemented as a separate embodiment, and step S4201+s4202+s4203 may be implemented as a separate embodiment, but is not limited thereto.
In this embodiment mode or example, the steps may be independently, arbitrarily combined, or exchanged in order, and the alternative modes or examples may be arbitrarily combined, and may be arbitrarily combined with any steps of other embodiment modes or other examples without contradiction.
Fig. 5A is an interactive schematic diagram illustrating an audio signal processing method according to another embodiment of the present disclosure. As shown in fig. 5A, an embodiment of the present disclosure relates to an audio signal processing method, which may be used in a decoding process, where the method includes:
in step S5101, spectral coefficients associated with the first audio signal are acquired.
Step S5102, determining a first window function, wherein the first window function includes: and a reference variable value associated with a window function length value used to process the first audio signal, the reference variable value being smaller than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal.
In step S5103, the spectral coefficients are processed according to the first window function, so as to obtain a first audio signal.
The audio signal processing method according to the embodiment of the present disclosure may include at least one of step S5101 to step S5103. For example, step S5101 may be implemented as a separate embodiment, step S5102 may be implemented as a separate embodiment, and so on, but is not limited thereto. Step S5101+s5102 may be implemented as an independent embodiment, and step S5101+s5102+s5103 may be implemented as an independent embodiment, but is not limited thereto.
In this embodiment mode or example, the steps may be independently, arbitrarily combined, or exchanged in order, and the alternative modes or examples may be arbitrarily combined, and may be arbitrarily combined with any steps of other embodiment modes or other examples without contradiction.
Fig. 5B is an interactive schematic diagram illustrating an audio signal processing method according to another embodiment of the present disclosure. As shown in fig. 5B, an embodiment of the present disclosure relates to an audio signal processing method, which may be used in a decoding process, where the method includes:
in step S5201, spectral coefficients associated with the first audio signal are acquired.
In step S5202, the first window function is determined according to the window function length value used for processing the first audio signal, the reference variable value related to the window function length value, the first variable, and the second window function, wherein the reference variable value is smaller than the window function length value, and the reference variable value is used for determining the time delay required for recovering the first audio signal.
Step S5203, processing the spectral coefficients according to second processing information to obtain an intermediate audio signal, where the second processing information includes: reference variable values.
In step S5204, the intermediate audio signal is processed according to the first window function to obtain a first audio signal.
The audio signal processing method according to the embodiment of the present disclosure may include at least one of step S5201 to step S5204. For example, step S5201 may be implemented as a separate embodiment, step S5202 may be implemented as a separate embodiment, and so on, but is not limited thereto. The steps S5201+s5202 may be implemented as an independent embodiment, and the steps S5201+s5202+s5203 may be implemented as an independent embodiment, but are not limited thereto.
In this embodiment mode or example, the steps may be independently, arbitrarily combined, or exchanged in order, and the alternative modes or examples may be arbitrarily combined, and may be arbitrarily combined with any steps of other embodiment modes or other examples without contradiction.
The following is an exemplary description of the above method.
Alternative embodiments:
as shown in fig. 6A, fig. 6A is a schematic flow chart of a time domain signal encoding process according to an embodiment of the disclosure. As shown in fig. 6A, in an embodiment of the present disclosure, an encoding end flow is shown. An example is given in which the encoding side is an encoder.
In some embodiments, the encoder input reads N points at a time, taking 1024 points as an example, changing the input data from shaped data to floating point data.
In some embodiments, the encoder processes the first audio signal in units of frames, and calculates energy of the read first audio signal in a frame according to the following calculation formula:
blockEnergy[i]+=input[i*nblockSize+j]*input[i*nblockSize+j];
where i=0, 1,..7 denotes dividing one frame of the first audio signal into 8 sub-blocks (an alternative example of a sub-signal block), block energy [ i ] denotes calculating the energy of each sub-block, j=0, 1,..127 denotes the size of each sub-block being 128 points, nblock size denotes the length of a sub-block, input denotes the input signal, the above formula divides 1024 points into exactly 8 sub-blocks, and calculates the energy of each sub-block.
In some embodiments, the size of the window may be determined by comparing the energy of the threshold (an alternative example of the energy threshold) to the energy of each sub-block (an alternative example of the energy value), i.e., the size of the window (an alternative example of the window function length value), and when the energy of one sub-block is greater than the threshold, the window is determined to be a short window, and otherwise the window is determined to be a long window.
In some embodiments, the window function is exemplified as a long window. In the embodiments of the present disclosure, the delay may be reduced by changing the length of the overlap-add portion of the window function. The improved window function needs to meet the perfect reconstruction of the signal. The formula for the long window is as follows:
wherein X is a variable parameter(an alternative example of the above-mentioned reference variable value) which takes on the value ofWherein m is a positive integer. The window function in the disclosed embodiment is divided into four parts altogether, the first part is zeroed, N ranges from 0 to N-X, the second part is w (), N ranges from N-X to N, the third part is a constant gain 1, N ranges from N to 2N-X, and the fourth part N ranges from 2N-X to 2N. The window shape is shown in fig. 6B below, and fig. 6B is a schematic diagram of a function of a low-latency window in an embodiment of the present disclosure.
In some embodiments, the first audio signal may be windowed. For example, the input signal (an alternative example of the first audio signal) is multiplied by a window function (an alternative example of the first window function), i.e. X (n) =w (n) ×input (n), where W (n) is the first window function, input (n) is the input signal, and X (n) is the windowed signal (an alternative example of the intermediate audio signal).
In some embodiments, the windowed signal may be subjected to an MDCT transform, where the formula for the MDCT is as follows:
where k=0, 1..n-1, N represents the first variable, N represents the window function length value, X represents the reference variable value, X (N) represents the intermediate audio signal, and X (k) represents the spectral coefficient. That is, the MDCT transform is performed on 2N points, resulting in MDCT coefficients for N points.
As shown in fig. 6C, fig. 6C is a schematic flow chart of a decoding process according to an embodiment of the disclosure. As shown in fig. 6C, in the embodiment of the present disclosure, a decoding end flow is shown.
In some embodiments, the MDCT coefficients (an alternative example of the spectral coefficients of the first audio signal) obtained at the encoding end may be subjected to an inverse modified discrete cosine transform (Inverse Modified Discrete Cosine Transform, IMDCT) to obtain the time-domain signal. The formula of IMDCT is as follows:
Where n=0, 1,..2N-1, N represents the first variable, N represents the window function length value, X represents the reference variable value, X (N) represents the intermediate audio signal, and X (k) represents the spectral coefficient. I.e. the frequency domain coefficients of the N points are converted into time domain signals of 2N points.
In some embodiments, the window type may be determined by reading the information transmitted from the encoding side, and then windowing the signal, i.e. multiplying the input signal by a window function, for example, windowing the intermediate audio signal, i.e. multiplying the input signal (e.g. the intermediate audio signal) by a window function (an alternative example of a first window function), as follows: i.e., X (n) =w (n) ×input (n), where W (n) is the first window function, input (n) is the input signal, and X (n) is the windowed signal.
In some embodiments, the overlap-add is performed after the window is applied, and the X points of the previous frame are overlapped with the N-X to N points of the current frame to recover X points, and the N-X points of the current frame are added (without overlap-add), so that N points are recovered altogether, and the X points of the 2N-X to 2N portion of the current frame can be recovered to the next frame, so the delay is X. An illustration of overlap-add is shown in fig. 6D, fig. 6D being an illustration of signal overlap-add in an embodiment of the present disclosure.
In the embodiment of the disclosure, the original long window may be replaced by the above-mentioned low-delay window (an optional example of the first window function), and whether the low-delay window can be perfectly reconstructed is first verified, and x=n/4 is selected for verification. Experiments prove that the method used by the embodiment of the disclosure can completely recover the original signal, namely perfect reconstruction is realized, but the time delay is changed from the original N to X (N/4).
Experimental protocol: the original audio signal is encoded and decoded through the existing long window to obtain an output signal 1, then the original audio signal is encoded and decoded through the improved low-delay window to obtain an output signal 2, and the waveforms of the two signals are compared. Take x=n/32.
As shown in fig. 7A-7C, fig. 7A is a schematic waveform diagram of Wav test file one in the embodiment of the disclosure. Fig. 7B is a schematic diagram of the output signal 1 corresponding to the waveform of the Wav test file one in the embodiment of the disclosure. Fig. 7C is a schematic diagram of the output signal 2 corresponding to the waveform of the Wav test file one in the embodiment of the disclosure.
As shown in fig. 7D-7F, fig. 7D is a schematic waveform diagram of Wav test file two in the embodiment of the disclosure. Fig. 7E is a schematic diagram of the output signal 1 corresponding to the waveform of the Wav test file two in the embodiment of the disclosure. Fig. 7F is a schematic diagram of the output signal 2 corresponding to the waveform of the Wav test file two in the embodiment of the disclosure.
As shown in fig. 7G-7I, fig. 7G is a waveform schematic diagram of Wav test file three in the embodiment of the disclosure. Fig. 7H is a schematic diagram of the output signal 1 corresponding to the waveform of the Wav test file three in the embodiment of the disclosure. Fig. 7I is a schematic diagram of the output signal 2 corresponding to the waveform of the Wav test file three in the embodiment of the disclosure.
As can be seen from fig. 7A to fig. 7I, the output signal 1 and the output signal 2 are substantially identical to the original signal, and the method provided in the embodiments of the present disclosure can effectively ensure the quality of audio processing while reducing the algorithm delay.
The embodiment of the disclosure also provides an audio signal processing method, wherein the first communication device determines a window function length value required for processing a first audio signal, and determines a first window function, and the first window function comprises: a reference variable value associated with the window function length value, the reference variable value being less than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal, and processing the first audio signal according to the first window function to obtain a spectral coefficient associated with the first audio signal; the second communication device obtains spectral coefficients associated with the first audio signal and processes the spectral coefficients according to a first window function to obtain the first audio signal.
The first communication device is, for example, a communication device performing an encoding process, the second communication device is, for example, a communication device performing a decoding process, the first communication device may be a terminal or a network device, the second communication device may be a terminal or a network device, and the first communication device and the second communication device may be the same or different communication devices.
Fig. 8A is a schematic structural diagram of an audio signal processing apparatus according to an embodiment of the present disclosure. As shown in fig. 8A, the audio signal processing apparatus includes: a processing module for determining a window function length value required for processing the first audio signal and determining a first window function, wherein the first window function comprises: and a reference variable value associated with the window function length value, the reference variable value being smaller than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal, and processing the first audio signal according to the first window function to obtain spectral coefficients associated with the first audio signal.
Fig. 8B is a schematic structural diagram of an audio signal processing apparatus according to an embodiment of the present disclosure. As shown in fig. 8B, the audio signal processing apparatus includes: a processing module for obtaining spectral coefficients associated with a first audio signal and determining a first window function, wherein the first window function comprises: and processing the spectral coefficients according to the first window function to obtain the first audio signal.
It should be understood that the division of each unit or module in the above apparatus is merely a division of a logic function, and may be fully or partially integrated into one physical entity or may be physically separated when actually implemented. Furthermore, units or modules in the apparatus may be implemented in the form of processor-invoked software: the device comprises, for example, a processor, the processor being connected to a memory, the memory having instructions stored therein, the processor invoking the instructions stored in the memory to perform any of the methods or to perform the functions of the units or modules of the device, wherein the processor is, for example, a general purpose processor, such as a central processing unit (Central Processing Unit, CPU) or microprocessor, and the memory is internal to the device or external to the device. Alternatively, the units or modules in the apparatus may be implemented in the form of hardware circuits, and part or all of the functions of the units or modules may be implemented by designing hardware circuits, which may be understood as one or more processors; for example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC), and the functions of some or all of the units or modules are implemented by designing the logic relationships of elements in the circuit; for another example, in another implementation, the above hardware circuit may be implemented by a programmable logic device (programmable logic device, PLD), for example, a field programmable gate array (Field Programmable Gate Array, FPGA), which may include a large number of logic gates, and the connection relationship between the logic gates is configured by a configuration file, so as to implement the functions of some or all of the above units or modules. All units or modules of the above device may be realized in the form of invoking software by a processor, or in the form of hardware circuits, or in part in the form of invoking software by a processor, and in the rest in the form of hardware circuits.
As shown in fig. 9A, fig. 9A is a schematic structural diagram of a communication device according to an embodiment of the present disclosure, where the communication device 9100 includes one or more processors 9101. The processor 9101 may be a general-purpose processor or a special-purpose processor, and may be, for example, a baseband processor or a central processing unit. The baseband processor may be used to process communication protocols and communication data, and the central processor may be used to control communication devices (e.g., base stations, baseband chips, terminal devices, terminal device chips, DUs or CUs, etc.), execute programs, and process data for the programs. The processor 9101 is configured to invoke instructions to cause the communication device 9100 to perform any of the above methods.
In some embodiments, communication device 9100 also includes one or more memories 9102 for storing instructions. Alternatively, all or part of the memory 9102 may be external to the communication device 9100.
In some embodiments, the communication device 9100 further comprises one or more transceivers 9103. When the communication device 9100 includes one or more transceivers 9103, communication steps such as transmission and reception in the above method are performed by the transceivers 9103, and other steps are performed by the processors 9101.
In some embodiments, the transceiver may include a receiver and a transmitter, which may be separate or integrated. Alternatively, terms such as transceiver, transceiver unit, transceiver circuit, etc. may be replaced with each other, terms such as transmitter, transmitter circuit, etc. may be replaced with each other, and terms such as receiver, receiving unit, receiver, receiving circuit, etc. may be replaced with each other.
Optionally, the communication device 9100 further comprises one or more interface circuits 9104, where the interface circuits 9104 are connected to the memory 9102, and the interface circuits 9104 can be used to receive signals from the memory 9102 or other apparatus, and can be used to send signals to the memory 9102 or other apparatus. For example, the interface circuit 9104 may read instructions stored in the memory 9102 and send the instructions to the processor 9101.
The communication device 9100 in the above embodiment description may be a network device or a terminal, but the scope of the communication device 9100 described in the present disclosure is not limited thereto, and the structure of the communication device 9100 may not be limited by fig. 9A. The communication device may be a stand-alone device or may be part of a larger device. For example, the communication device may be: 1) A stand-alone integrated circuit IC, or chip, or a system-on-a-chip or subsystem; (2) A set of one or more ICs, optionally including storage means for storing data, programs; (3) an ASIC, such as a Modem (Modem); (4) modules that may be embedded within other devices; (5) A receiver, a terminal device, an intelligent terminal device, a cellular phone, a wireless device, a handset, a mobile unit, a vehicle-mounted device, a network device, a cloud device, an artificial intelligent device, and the like; (6) others, and so on.
Fig. 9B is a schematic structural diagram of a chip according to an embodiment of the disclosure. For the case where the communication device 9100 may be a chip or a chip system, a schematic structural diagram of the chip 9200 shown in fig. 9B may be referred to, but is not limited thereto. The chip 9200 includes one or more processors 9201, the processors 9201 are configured to call instructions to cause the chip 9200 to perform any of the above methods.
In some embodiments, the chip 9200 further includes one or more interface circuits 9202, the interface circuits 9202 are connected to the memory 9203, the interface circuits 9202 may be used to receive signals from the memory 9203 or other devices, and the interface circuits 9202 may be used to transmit signals to the memory 9203 or other devices. For example, the interface circuit 9202 may read an instruction stored in the memory 9203 and send the instruction to the processor 9201. Alternatively, the terms interface circuit, interface, transceiver pin, transceiver, etc. may be interchanged.
In some embodiments, the chip 9200 further includes one or more memories 9203 for storing instructions. Alternatively, all or part of the memory 9203 may be external to the chip 9200.
The present disclosure also proposes a storage medium having stored thereon instructions that, when executed on the communication device 9100, cause the communication device 9100 to perform any of the above methods. Optionally, the storage medium is an electronic storage medium. Alternatively, the storage medium described above is a computer-readable storage medium, but is not limited thereto, and it may be a storage medium readable by other devices. Alternatively, the above-described storage medium may be a non-transitory (non-transitory) storage medium, but is not limited thereto, and it may also be a transitory storage medium.
In the disclosed embodiments, the processor is a circuit with audio signal processing capability, in one implementation, the processor may be a circuit with instruction reading and running capability, such as a central processing unit (Central Processing Unit, CPU), microprocessor, graphics processor (graphics processing unit, GPU) (which may be understood as a microprocessor), or digital audio signal processor (digital signal processor, DSP), etc.; in another implementation, the processor may implement a function through a logical relationship of hardware circuits that are fixed or reconfigurable, e.g., a hardware circuit implemented as an application-specific integrated circuit (ASIC) or a programmable logic device (programmable logic device, PLD), such as an FPGA. In the reconfigurable hardware circuit, the processor loads the configuration document, and the process of implementing the configuration of the hardware circuit may be understood as a process of loading instructions by the processor to implement the functions of some or all of the above units or modules. Furthermore, hardware circuits designed for artificial intelligence may be used, which may be understood as ASICs, such as neural network processing units (Neural Network Processing Unit, NPU), tensor processing units (Tensor Processing Unit, TPU), deep learning processing units (Deep learning Processing Unit, DPU), etc.
The present disclosure also proposes a program product that, when executed by the communication device 9100, causes the communication device 9100 to perform any of the above methods. Optionally, the above-described program product is a computer program product.
The present disclosure also proposes a computer program which, when run on a computer, causes the computer to perform any of the above methods.
In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented in software, may be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer programs. When the computer program is loaded and executed on a computer, the flow or functions described in accordance with the embodiments of the present disclosure are produced in whole or in part. The computer may be a general purpose computer, a special purpose computer, a computer network, or other programmable apparatus. The computer program may be stored in or transmitted from one computer readable storage medium to another, for example, by wired (e.g., coaxial cable, optical fiber, digital subscriber line (digital subscriber line, DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means from one website, computer, server, or data center. The computer readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that contains an integration of one or more available media. The usable medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a high-density digital video disc (digital video disc, DVD)), or a semiconductor medium (e.g., a Solid State Disk (SSD)), or the like.
Those of ordinary skill in the art will appreciate that the various illustrative elements and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, or combinations of computer software and electronic hardware. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the solution. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
It will be clear to those skilled in the art that, for convenience and brevity of description, specific working procedures of the above-described systems, apparatuses and units may refer to corresponding procedures in the foregoing method embodiments, and are not repeated herein.
The foregoing is merely specific embodiments of the disclosure, but the protection scope of the disclosure is not limited thereto, and any person skilled in the art can easily think about changes or substitutions within the technical scope of the disclosure, and it is intended to cover the scope of the disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims (28)

1. A method of audio signal processing, the method comprising:
determining a window function length value required to process the first audio signal;
determining a first window function, wherein the first window function comprises: a reference variable value associated with the window function length value, the reference variable value being less than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal; and
and processing the first audio signal according to the first window function to obtain a spectrum coefficient related to the first audio signal.
2. The method of claim 1, wherein determining the window function length value required to process the first audio signal comprises:
determining an energy threshold;
determining an energy value corresponding to each sub-signal block in the first audio signal;
and determining the length value of the window function according to the energy threshold value and a plurality of energy values.
3. The method of claim 2, wherein said determining said window function length value based on said energy threshold and a plurality of said energy values comprises:
determining the window function length value as a first window function length value if a plurality of the energy values are all less than or equal to the energy threshold;
And if at least one of the energy values is greater than the energy threshold, determining that the window function length value is a second window function length value, wherein the second window function length value is smaller than the first window function length value.
4. A method according to any of claims 1-3, wherein the first window function satisfies a first condition, the first condition comprising: the sum of the first value and the second value is 1; wherein the first value is the square of the value obtained by processing a first variable based on the first window function, the second value is the square of the value obtained by processing a second variable based on the first window function, and the second variable is the sum between the first variable and the window function length value.
5. The method of claim 4, wherein the first condition is the following formula:
W(n) 2 +W(n+N) 2 =1;
where N represents a first variable, N represents the window function length value, and W () represents the first window function.
6. The method of any of claims 1-4, wherein the determining a first window function comprises:
the first window function is determined from the window function length value, the reference variable value, a first variable, and a second window function.
7. The method of claim 6, wherein the first window function comprises at least one of:
a first function portion corresponding to the first variable, wherein the first variable belongs to a first value range;
a second function portion corresponding to the first variable, wherein the first variable belongs to a second range of values;
a third function portion corresponding to the first variable, wherein the first variable belongs to a third range of values;
a fourth function portion corresponding to the first variable, wherein the first variable belongs to a fourth range of values;
the value range is obtained by dividing a reference value range according to the window function length value and the reference variable value, the minimum value of the reference value range is zero, the maximum value is twice the window function length, the function parts are different, at least one function part is related to the second window function, and at least one function part is a constant value.
8. The method of claim 6 or 7, wherein the first window function is of the formula:
Wherein N represents the first variable, N represents the window function length value, X represents the reference variable value, W D () Representing the first window function, w () represents the second window function.
9. The method according to any of claims 1-8, wherein said processing the first audio signal according to the first window function to obtain spectral coefficients related to the first audio signal comprises:
processing the first audio signal according to the first window function to obtain an intermediate audio signal;
processing the intermediate audio signal according to first processing information to obtain the spectral coefficients related to the first audio signal, wherein the first processing information comprises: the reference variable value.
10. The method of claim 9, wherein the first processing information is used to perform a modified discrete cosine transform process on the intermediate audio signal.
11. The method according to claim 9 or 10, wherein the first processing information is the following formula:
where k=0, 1,..n-1, N represents the first variable, N represents the window function length value, X represents the reference variable value, X (N) represents the intermediate audio signal, and X (k) represents the spectral coefficient.
12. The method of any one of claims 1-11, wherein the reference variable value is determined by the following formula:
wherein N represents the window function length value, X represents the reference variable value, and m is a positive integer.
13. A method of audio signal processing, the method comprising:
acquiring spectral coefficients associated with the first audio signal;
determining a first window function, wherein the first window function comprises: a reference variable value associated with a window function length value used to process a first audio signal, the reference variable value being smaller than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal;
and processing the spectral coefficients according to the first window function to obtain the first audio signal.
14. The method of claim 13, wherein the first window function satisfies a first condition, the first condition comprising: the sum of the first value and the second value is 1; wherein the first value is the square of the value obtained by processing a first variable based on the first window function, the second value is the square of the value obtained by processing a second variable based on the first window function, and the second variable is the sum between the first variable and the window function length value.
15. The method of claim 14, wherein the first condition is the following formula:
W(n) 2 +W(n+N) 2 =1;
where N represents a first variable, N represents the window function length value, and W () represents the first window function.
16. The method of any of claims 13-15, wherein the determining a first window function comprises:
the first window function is determined from the window function length value, the reference variable value, a first variable, and a second window function.
17. The method of claim 16, wherein the first window function comprises at least one of:
a first function portion corresponding to the first variable, wherein the first variable belongs to a first value range;
a second function portion corresponding to the first variable, wherein the first variable belongs to a second range of values;
a third function portion corresponding to the first variable, wherein the first variable belongs to a third range of values;
a fourth function portion corresponding to the first variable, wherein the first variable belongs to a fourth range of values;
the value range is obtained by dividing a reference value range according to the window function length value and the reference variable value, the minimum value of the reference value range is zero, the maximum value is twice the window function length, the function parts are different, at least one function part is related to the second window function, and at least one function part is a constant value.
18. The method of claim 16 or 17, wherein the first window function is of the formula:
wherein N represents the first variable, N represents the window function length value, X represents the reference variable value, W D () Representing the first window function, w () represents the second window function.
19. The method according to any of claims 13-18, wherein said processing the spectral coefficients according to the first window function to obtain the first audio signal comprises:
processing the spectral coefficients according to second processing information to obtain an intermediate audio signal, wherein the second processing information comprises: the reference variable value;
and processing the intermediate audio signal according to the first window function to obtain the first audio signal.
20. The method of claim 19, wherein the second processing information is used to perform an inverse modified discrete cosine transform process on the spectral coefficients.
21. The method of claim 19 or 20, wherein the second processing information is the following formula:
where n=0, 1..2N-1, N represents the first variable, N represents the window function length value, X represents the reference variable value, X (N) represents the intermediate audio signal, and X (k) represents the spectral coefficient.
22. The method of any one of claims 1-21, wherein the reference variable value is determined by the following formula:
wherein N represents the window function length value, X represents the reference variable value, and m is a positive integer.
23. A method of audio signal processing, the method comprising:
a first communication device for determining a window function length value required for processing a first audio signal and determining a first window function, wherein the first window function comprises: a reference variable value associated with the window function length value, the reference variable value being smaller than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal, and to process the first audio signal according to the first window function to obtain spectral coefficients associated with the first audio signal;
and the second communication equipment is used for acquiring the spectral coefficients related to the first audio signal and processing the spectral coefficients according to the first window function to obtain the first audio signal.
24. An audio signal processing apparatus, the apparatus comprising:
a processing module, configured to determine a window function length value required for processing a first audio signal, and determine a first window function, where the first window function includes: and a reference variable value associated with the window function length value, the reference variable value being smaller than the window function length value, the reference variable value being used to determine a time delay required to recover the first audio signal, and processing the first audio signal according to the first window function to obtain spectral coefficients associated with the first audio signal.
25. An audio signal processing apparatus, the apparatus comprising:
a processing module, configured to obtain spectral coefficients associated with a first audio signal, and determine a first window function, where the first window function includes: and processing the spectral coefficients according to the first window function to obtain the first audio signal.
26. A communication device, comprising:
one or more processors;
wherein the processor is configured to invoke instructions to cause the communication device to perform the method of any of claims 1-23.
27. A communication system comprising a first communication device configured to implement the method of any of claims 1-12 and a second communication device configured to implement the method of any of claims 13-22.
28. A storage medium storing instructions which, when executed on a communications device, cause the communications device to perform the method of any one of claims 1-23.
CN202380010554.7A 2023-08-09 2023-08-09 Audio signal processing method and device, communication equipment, communication system, storage medium Pending CN117242517A (en)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2023/112078 WO2025030449A1 (en) 2023-08-09 2023-08-09 Audio signal processing method and apparatus, communication device, communication system, and storage medium

Publications (1)

Publication Number Publication Date
CN117242517A true CN117242517A (en) 2023-12-15

Family

ID=89091651

Family Applications (1)

Application Number Title Priority Date Filing Date
CN202380010554.7A Pending CN117242517A (en) 2023-08-09 2023-08-09 Audio signal processing method and device, communication equipment, communication system, storage medium

Country Status (2)

Country Link
CN (1) CN117242517A (en)
WO (1) WO2025030449A1 (en)

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP2372703A1 (en) * 2010-03-11 2011-10-05 Fraunhofer-Gesellschaft zur Förderung der Angewandten Forschung e.V. Signal processor, window provider, encoded media signal, method for processing a signal and method for providing a window
CN109215667B (en) * 2017-06-29 2020-12-22 华为技术有限公司 Time delay estimation method and device
US11380343B2 (en) * 2019-09-12 2022-07-05 Immersion Networks, Inc. Systems and methods for processing high frequency audio signal
CN113035223B (en) * 2021-03-12 2023-11-14 北京字节跳动网络技术有限公司 Audio processing method, device, equipment and storage medium
CN115116475B (en) * 2022-06-13 2024-02-02 北京邮电大学 Voice depression automatic detection method and device based on time delay neural network

Also Published As

Publication number Publication date
WO2025030449A1 (en) 2025-02-13

Similar Documents

Publication Publication Date Title
ES2375192T3 (en) CODIFICATION FOR IMPROVED SPEECH TRANSFORMATION AND AUDIO SIGNALS.
RU2718421C1 (en) Audio decoding device, audio coding device, audio decoding method, audio coding method, audio decoding program and audio coding program
CN113870872B (en) Speech quality enhancement method, device and system based on deep learning
CN103368682B (en) Method and device for encoding and decoding signals
KR101837191B1 (en) Prediction method and coding/decoding device for high frequency band signal
JPWO2007088853A1 (en) Speech coding apparatus, speech decoding apparatus, speech coding system, speech coding method, and speech decoding method
CN113994425B (en) Device and method for encoding and decoding scene-based audio data
KR102915663B1 (en) Audio signal encoding method, decoding method, encoding device and decoding device
CN114008705B (en) Perform psychoacoustic audio encoding and decoding based on operating conditions
WO2015136078A1 (en) Audio coding method and apparatus
CN117831548A (en) Training method, encoding method, decoding method and device of audio codec system
EP4135330A1 (en) Data processing method and system
CN116434760A (en) An audio coding method, device, electronic equipment and storage medium
KR20070070174A (en) Scalable coding apparatus, scalable decoding apparatus and scalable coding method
WO2025145380A1 (en) Signal coding method, signal decoding method, coding device and decoding device, and storage medium
CN104981868B (en) Method for encoding and decoding audio signals and device for encoding and decoding audio signals
KR102763821B1 (en) Audio encoding method and device and audio decoding method and device
WO2025030449A1 (en) Audio signal processing method and apparatus, communication device, communication system, and storage medium
WO2022012677A1 (en) Audio encoding method, audio decoding method, related apparatus and computer-readable storage medium
JP7407110B2 (en) Encoding device and encoding method
CN118160034A (en) Coding and decoding method, device and storage medium
CN114945981A (en) Audio signal processing method and device
CN120435737A (en) Coding and decoding method, device and storage medium
WO2025020009A1 (en) Model initialization method and apparatus
WO2025091293A1 (en) Grouping method, encoder, decoder, and storage medium

Legal Events

Date Code Title Description
PB01 Publication
PB01 Publication
SE01 Entry into force of request for substantive examination
SE01 Entry into force of request for substantive examination