EP4604120A1 - Appareil, procédé et programme informatique pour traitement de signal audio sur la base d'une différence de niveau entre canaux et d'une manipulation de composante de signal latérale - Google Patents

Appareil, procédé et programme informatique pour traitement de signal audio sur la base d'une différence de niveau entre canaux et d'une manipulation de composante de signal latérale

Info

Publication number
EP4604120A1
EP4604120A1 EP24157976.2A EP24157976A EP4604120A1 EP 4604120 A1 EP4604120 A1 EP 4604120A1 EP 24157976 A EP24157976 A EP 24157976A EP 4604120 A1 EP4604120 A1 EP 4604120A1
Authority
EP
European Patent Office
Prior art keywords
audio signal
signal
inter
channel
input audio
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24157976.2A
Other languages
German (de)
English (en)
Inventor
Christian Uhle
Philipp Weber
Matthias Lang
Maximilian Mayer
Dominik Zehnder
Sophia Emmert
Patrick Gampp
Thomas Bachmann
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Fraunhofer Gesellschaft zur Foerderung der Angewandten Forschung eV
Original Assignee
Fraunhofer Gesellschaft zur Foerderung der Angewandten Forschung eV
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Fraunhofer Gesellschaft zur Foerderung der Angewandten Forschung eV filed Critical Fraunhofer Gesellschaft zur Foerderung der Angewandten Forschung eV
Priority to EP24157976.2A priority Critical patent/EP4604120A1/fr
Priority to PCT/EP2025/054116 priority patent/WO2025172586A1/fr
Publication of EP4604120A1 publication Critical patent/EP4604120A1/fr
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L21/00Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
    • G10L21/02Speech enhancement, e.g. noise reduction or echo cancellation
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/008Multichannel audio signal coding or decoding using interchannel correlation to reduce redundancy, e.g. joint-stereo, intensity-coding or matrixing
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L19/00Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
    • G10L19/02Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders
    • G10L19/0204Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using spectral analysis, e.g. transform vocoders or subband vocoders using subband decomposition
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/03Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters
    • G10L25/18Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 characterised by the type of extracted parameters the extracted parameters being spectral information of each sub-band
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S3/00Systems employing more than two channels, e.g. quadraphonic
    • H04S3/008Systems employing more than two channels, e.g. quadraphonic in which the audio signals are in digital form, i.e. employing more than two discrete digital channels
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field

Definitions

  • Embodiments according to the invention comprise apparatuses, methods and computer programs for audio signal processing based on inter-channel-level-difference and side signal component manipulation.
  • Embodiments according to the invention comprise audio processing systems, methods and computer programs for audio signal processing based on transient enhancement, stage width enhancement and ambience enhancement.
  • Embodiments according to the invention comprise apparatuses, systems, methods and computer programs for performing an acoustic boost.
  • a quality of the audio reproduction is additionally dependent on a surrounding of a respective listener, e.g. whether the listener is located in an opera hall or within a more confined space, such as a car.
  • Embodiments of the invention will be presented according to a first and second aspect. This structuring of embodiments is provided in order to facilitate understanding the different details, functionalities and features of embodiments.
  • an audio processing system may comprise any of the apparatuses as disclosed in the context of the first aspect, for example for applying the stage width enhancement.
  • Embodiments according to the first aspect of the invention comprise an apparatus for processing an audio signal, e.g. a stage width enhancement or a stage width enhancement block.
  • the apparatus is configured to provide, on the basis of an input audio signal (e.g. on the basis of a multi-channel input audio signal; e.g. on the basis of a stereo input in a time domain; e.g. on the basis of a multi-channel input audio signal in a frequency domain; e.g. on the basis of a combined audio signal obtained using a transient enhancement and an ambience enhancement) a first modified audio signal (e.g. a first modified multi-channel audio signal; e.g.
  • the apparatus is further configured to provide, on the basis of the input audio signal, e.g. on the basis of the stereo input in the time domain, a second modified audio signal, e.g. a second modified multi-channel audio signal, in which side signal components, e.g. corresponding audio signal components in multiple channels having a comparatively large inter-channel-level-difference, are emphasized relative to centered signal components (e.g. corresponding audio signal components in multiple channels having no inter-channel-level-difference or having a comparatively small inter-channel-level-difference) when compared to the input audio signal, and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal, e.g. using a side signal emphasizer or using a side signal extraction, and/or in which only side signal components are included while centered signal components are suppressed.
  • a second modified audio signal e.g. a second modified multi-channel audio signal
  • side signal components e.g. corresponding audio
  • the apparatus is further configured to combine, e.g. using a linear combination; e.g. using a weighted linear combination, the first modified audio signal and the second modified audio signal, e.g. in order to obtain a stage width enhanced multi-channel output (audio) signal or, for example, in order to obtain a stage width enhanced stereo output (audio) signal.
  • a linear combination e.g. using a weighted linear combination
  • the first modified audio signal and the second modified audio signal e.g. in order to obtain a stage width enhanced multi-channel output (audio) signal or, for example, in order to obtain a stage width enhanced stereo output (audio) signal.
  • This embodiment is based on the finding that boosting or increasing inter-channel-level-differences allows manipulating an audio scene, so as to off center a perceived position of an audio source in the audio scene, e.g. stereo image, further to sides of the audio scene and hence away from the center of the audio scene.
  • the inventors recognized that based on a manipulation of the inter-channel-level-differences, this may be achieved, for example, without modifying the loudness of the sources and, for example, without modifying the timbre, loudness and/or position of sources panned to the center.
  • a manipulation of positional sound cues for a listener may be performed without altering characteristics, or at least some characteristics of the sounds themselves.
  • manipulation of side signal components relative to centered signal components allows a manipulation of components so as to increase the perceived width of the audio scene for the listener.
  • increasing inter-channel-level-difference and side signal extraction and manipulation may both achieve an increase in perceived width of the audio scene.
  • embodiments may not only allow addressing a wide plurality of different audio signals (e.g. to enhance a perceived width thereof), but as well allow synergistically reducing respective artifacts in the perceived audio scene. Therefore, a processing of inter-channel-level-difference and side signal extraction to provide first and second signal components may be performed in parallel or consecutively, and it has been found to provide high audio quality for many different types of signals.
  • the apparatus is configured to determine inter-channel-level differences (e.g. inter-channel level difference values D x (n,k) (also referred to as D(n,k or as D x (f,k) or D(f,k))); wherein is should be noted that, at least in some embodiments, indices (f,m) may be used instead of indices (n,k), wherein f and n are frequency indices (e.g. bin (f) and band (n) indices), and wherein indices m and k are time indices) for a plurality of corresponding time-frequency bins of the input audio signal, e.g. between corresponding time-frequency bins of different channels of the input audio signal.
  • inter-channel-level differences e.g. inter-channel level difference values D x (n,k) (also referred to as D(n,k or as D x (f,k) or D(f,k))
  • indices (f,m) may be used
  • band-wise and bin-wise processing may be performed.
  • embodiments discussed with regard to a bin-wise processing may as well be processed band-wise and vice versa.
  • indices n and f may be used interchangeably.
  • the apparatus is further configured to scale spectral bin values, e.g. X 1 (n,k), X 2 (n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined inter-channel level differences, in order to obtain corresponding scaled spectral bin values of the first modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal) such that an inter channel level difference between corresponding spectral bin values of the first modified audio signal, e.g. Y 1 (n,k),Y 2 (n,k), is increased when compared to an inter-channel level difference between corresponding spectral bin values of the input audio signals, e.g. X 1 (n,k),X 2 (n,k), associated with a same frequency, e.g. having frequency index n, and a same time portion, e.g. having time index k.
  • spectral bin values e.g.
  • the inventors recognized that adapting the inter-channel-level differences frequency-wise and optionally portion-wise, hence, for example per set of corresponding frequency bins and for example per set of corresponding time indices, allows manipulating spatial cues of the audio scene efficiently, for example, for providing a particularly immersive experience of the audio scene. Accordingly, the inventive approach may be applied on different frequency bands independently.
  • the apparatus is configured to scale spectral bin values, e.g. X 1 (n,k), X 2 (n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined inter-channel level differences, in order to obtain corresponding scaled spectral bin values of the first modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal) such that a total energy in corresponding spectral bins of the first modified audio signal, e.g. Y 1 (n,k), Y 2 (n,k), is, for example at least substantially, e.g.
  • X 1 (n,k), X 2 (n,k) associated with a same frequency, e.g. having frequency index n, and a same time portion, e.g. having time index k.
  • the inventors recognized that a quality of the perceived audio scene may be improved by maintaining the total energy in corresponding spectral bins, so as to change a relationship of the channels in the form of level differences between the channels, without changing the overall loudness, and hence also without changing the perceived loudness.
  • the apparatus is configured to scale spectral bin values, e.g. X 1 (n,k), X 2 (n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in such a manner that an inter-channel level difference, when considered on a logarithmic scale, is scaled, e.g. linearly scaled; e.g. scaled by multiplication, by a predetermined value.
  • spectral bin values e.g. X 1 (n,k), X 2 (n,k)
  • corresponding time frequency bins e.g. of corresponding time frequency bins of different channels of the input audio signal
  • This may, for example be performed, such that an inter-channel level difference between corresponding time frequency bins of the first modified audio signal (e.g. between corresponding time frequency bins of different channels of the first modified audio signal) on a logarithmic scale is linearly scaled (e.g. multiplied by a predetermined factor, e.g. alpha) when compared to an inter-channel level difference between corresponding time frequency bins of input audio signal.
  • a predetermined factor e.g. alpha
  • the inventors recognized that a scaling performed on a logarithmic scale may allow for a precisely selectable (because of the scale) impact of a audio scene width enhancement.
  • the apparatus is configured to determine spectral weights, e.g. G 1 (n,k), G 2 (n,k), in dependence on inter-channel-level differences, e.g. inter-channel level difference values D x (n,k), for a plurality of corresponding time-frequency bins of the input audio signal, e.g. between corresponding time-frequency bins of different channels of the input audio signal.
  • spectral weights e.g. G 1 (n,k), G 2 (n,k
  • inter-channel-level differences e.g. inter-channel level difference values D x (n,k)
  • the apparatus may be further configured to scale spectral bin values, e.g. X 1 (n,k), X 2 (n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined spectral weights, in order to obtain corresponding scaled spectral bin values of the first modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal, e.g. Y 1 (n,k), Y 2 (n,k)).
  • spectral bin values e.g. X 1 (n,k), X 2 (n,k)
  • the apparatus is configured to obtain a logarithmic representation, e.g. D x (n,k), of an inter-channel level difference between corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal (e.g. using a computation of powers or energies P 1 (n,k) and P 2 (n,k) the corresponding time frequency bins of the input audio signal and using a formation of a logarithm of a ratio between the computed powers or energies). Furthermore, the apparatus is configured to multiply the logarithmic representation, e.g.
  • the apparatus is configured to derive one or more intermediate scaling values in a linear domain, e.g. K 1 (n,k) or K 2 (n,k), from the one or more intermediate scaling values in the logarithmic domain (e.g.
  • the apparatus is configured to apply an energy normalization, in order to obtain the spectral weights, e.g. G 1 (n,k), G 2 (n,k), in dependence on the intermediate scaling value in the linear domain, e.g. K 1 (n,k) and/or K 2 (n,k).
  • This approach may allow for a good compromise between computational complexity and audio scene enhancement.
  • the apparatus is configured to derive a factor, by which the inter-channel level difference is to be changed, e.g. K 1 (n,k) or K 2 (n,k), from an inter-channel level difference value (e.g. from a logarithmic representation of a ratio between energies in corresponding time frequency bins of two channels of the input audio signal; e.g. from D x (n,k), e.g. such that a logarithmic representation of factor, by which the inter-channel level difference is to be changed, e.g. H 1 (f,m) or H 2 (f,m), is proportional to the logarithmic representation of a ratio between energies in corresponding time frequency bins of two channels of the input audio signal, e.g. to D x (f,m)).
  • an inter-channel level difference value e.g. from a logarithmic representation of a ratio between energies in corresponding time frequency bins of two channels of the input audio signal; e.g. from D x
  • the apparatus is configured to provide spectral weights, e.g. G 1 (n,k), G 2 (n,k), for scaling corresponding time frequency bins of two channels of the input audio signal, such that ratio of the spectral weights is only determined by the factor. This may allow determining the spectral weights with low computational effort.
  • the apparatus is configured to derive spectral weights for scaling corresponding time frequency bins of two channels of the input audio signal, such that ratio of the spectral weights is equal to a potency, e.g. having a real-valued exponent, of a ratio of powers in the corresponding time frequency bins.
  • a potency e.g. having a real-valued exponent
  • the apparatus is configured to map, e.g. directly, e.g. without a computation of a panning index or the like, a value representing the inter-channel level difference, e.g. D x (n,k), (e.g. in a logarithmic domain, describing a ratio between energies in corresponding time-frequency bins of two channels in a logarithmized form) onto one or more values representing a change of the inter-channel level difference, e.g. in a logarithmic domain, using a linear or piecewise-linear mapping.
  • a value representing the inter-channel level difference e.g. D x (n,k)
  • D x (n,k) e.g. in a logarithmic domain, describing a ratio between energies in corresponding time-frequency bins of two channels in a logarithmized form
  • this may be performed, such that the value representing the inter-channel level difference is proportional to the value representing the change of the inter-channel level difference at least for inter-channel level differences having a first sign.
  • such an approach may be performed using a piecewise linear mapping having a first proportionality factor between a value representing the inter-channel level difference and a (e.g. first) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is larger than an energy in a second channel, and having a second proportionality factor between the value representing the inter-channel level difference and a (e.g. second) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is smaller than an energy in a second channel.
  • the inventors recognized that the linear or piecewise-linear mapping may allow for an efficient determination of signal weights for inter channel level adjustment.
  • the apparatus is configured to, for example directly, derive spectral weights for scaling corresponding time frequency bin without computing a panning index. This may allow increasing an efficiency of the audio processing.
  • the apparatus is configured to obtain the first modified audio signal using an, optionally selective (e.g. total energy maintaining or total power maintaining) opposite modification of intensities of signal components in corresponding time frequency bins of different channels of the input audio signal, which increases inter-channel-level differences.
  • an, optionally selective (e.g. total energy maintaining or total power maintaining) opposite modification of intensities of signal components in corresponding time frequency bins of different channels of the input audio signal which increases inter-channel-level differences.
  • This may allow providing a balanced adaptation of inter channel levels, e.g. so as to maintain an overall loudness of the audio signal.
  • the apparatus is configured to obtain the second modified audio signal using an increase of energies of off-center signal components (e.g. of corresponding signal components having a comparatively larger inter-channel level difference or of corresponding signal components having a comparatively large inter-channel time difference or of corresponding signal components having a comparatively large inter-channel-phase difference) in corresponding time frequency bins of different channels of the input audio signal.
  • off-center signal components e.g. of corresponding signal components having a comparatively larger inter-channel level difference or of corresponding signal components having a comparatively large inter-channel time difference or of corresponding signal components having a comparatively large inter-channel-phase difference
  • the apparatus is configured to obtain the second modified audio signal using a reduction of energies of centered signal components (e.g. of corresponding signal components having a comparatively small inter-channel level difference or of corresponding signal components having a comparatively small inter-channel time difference or of corresponding signal components having a comparatively small inter-channel-phase difference) in corresponding time frequency bins of different channels of the input audio signal.
  • centered signal components e.g. of corresponding signal components having a comparatively small inter-channel level difference or of corresponding signal components having a comparatively small inter-channel time difference or of corresponding signal components having a comparatively small inter-channel-phase difference
  • an increase in signal energy differences between center and off center components may allow a manipulating in accordance with a signal manipulation based on inter channel level differences (for the first modified audio signal), but with a different or even complementary behavior regarding artifacts, so as to supplement the first modified audio signal.
  • the apparatus is configured to obtain the second modified audio signal using a processing in which a ratio between energies of off-center signal components and centered signal components is increased, e.g. by more than 20 percent, or by more than 50 percent.
  • the apparatus is configured to obtain the first modified audio signal using a processing in which a ratio between energies of off-center signal components and centered signal components is maintained, within a tolerance of +/-10 percent.
  • stage width enhancement a combination of energy conserving and energy altering approaches, for a same or similar type of manipulation, e.g. stage width enhancement, may be combined so as to address diverse input audio signals (e.g. to achieve a good width enhancement for many types of signals) and/or to cancel out approach-specific artifact behavior and/or to shift an emphasis on the one or the other approach depending on the type or the characteristics of the input signal.
  • the apparatus is configured to scale corresponding signal components in corresponding time frequency bins of different channels of the input audio signal in dependence on an inter-channel-level difference between the respective corresponding signal components, and/or in dependence on an inter-channel time difference between the respective corresponding signal components, and/or in dependence on an inter-channel phase difference between the respective corresponding signal components, in order obtain the second modified audio signal.
  • the apparatus is configured to compute a gain mask (e.g. a gain mask in a time-frequency domain, e.g. a gain mask which extracts side signal components) for a scaling of time-frequency bins of the input audio signal on the basis of the input audio signal (e.g. in order to obtain the second modified audio signal, e.g. on the basis of a time-frequency domain representation of the input audio signal).
  • a gain mask e.g. a gain mask in a time-frequency domain, e.g. a gain mask which extracts side signal components
  • the apparatus is configured to scale time-frequency bins of the input audio signal using respective entries of the gain mask, in order to obtain the second modified audio signal, e.g. in order to obtain a time-frequency domain representation of the second modified audio signal.
  • the apparatus comprises an inter-channel level difference modifier, wherein the inter-channel level difference modifier is configured to determine inter-channel-level differences (e.g. inter-channel level difference values D x (n,k); wherein it should be noted that, in some embodiments, indices (f,m) may be used instead of indices (n,k), wherein f and n are frequency indices, and wherein indices m and k are time indices) for a plurality of corresponding time-frequency bins of an input audio signal (e.g. between corresponding time-frequency bins of different channels of the input audio signal).
  • inter-channel-level differences e.g. inter-channel level difference values D x (n,k)
  • indices (f,m) may be used instead of indices (n,k), wherein f and n are frequency indices, and wherein indices m and k are time indices
  • the inter-channel level difference modifier is configured to scale spectral bin values, e.g. X 1 (n,k), X 2 (n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined inter-channel level differences, in order to obtain corresponding scaled spectral bin values of a modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal) such that an inter channel level difference between corresponding spectral bin values of the first modified audio signal, e.g.
  • Y 1 (n,k), Y 2 (n,k) is modified when compared to an inter-channel level difference between corresponding spectral bin values of the input audio signals, e.g. X 1 (n,k), X 2 (n,k), associated with a same frequency (e.g. having frequency index n) and a same time portion (e.g. having time index k)
  • the same may be performed, such that the value representing the inter-channel level difference is proportional to the value representing the change of the inter-channel level difference at least for inter-channel level differences having a first sign.
  • the approach may comprise using a piecewise linear mapping having a first proportionality factor between a value representing the inter-channel level difference and a (e.g. first) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is larger than an energy in a second channel, and having a second proportionality factor between the value representing the inter-channel level difference and a (e.g. second) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is smaller than an energy in a second channel.
  • inter-channel-level-differences e.g. between corresponding audio signal components in different channels
  • inter-channel-level-differences e.g. between corresponding audio signal components in different channels
  • an intensity relationship between signal components having a comparatively small inter-channel-level-difference and signal components having a comparatively large inter-channel-level-difference is not changed by more than 10 percent
  • an intensity relationship between signal components having a comparatively small inter-channel-level-difference and signal components having a comparatively large inter-channel-level-difference is not changed by more than 10 percent
  • inter-channel-level-difference increaser e.g. using an inter-channel-level difference booster
  • the method further comprises providing, on the basis of the input audio signal, e.g. on the basis of the stereo input in the time domain, a second modified audio signal, e.g. a second modified multi-channel audio signal, in which side signal components (e.g. corresponding audio signal components in multiple channels having a comparatively large inter-channel-level-difference) are emphasized relative to centered signal components (e.g. corresponding audio signal components in multiple channels having no inter-channel-level-difference or having a comparatively small inter-channel-level-difference) when compared to the input audio signal, and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal, e.g. using a side signal emphasizer or using a side signal extraction, and/or in which only side signal components are included while centered signal components are suppressed.
  • a second modified audio signal e.g. a second modified multi-channel audio signal
  • side signal components e.g. corresponding audio signal components in
  • the method further comprises combining (e.g. using a linear combination; e.g. using a weighted linear combination) the first modified audio signal and the second modified audio signal (e.g. in order to obtain a stage width enhanced multi-channel output (audio) signal or in order to obtain a stage width enhanced stereo output (audio) signal).
  • combining e.g. using a linear combination; e.g. using a weighted linear combination
  • the first modified audio signal and the second modified audio signal e.g. in order to obtain a stage width enhanced multi-channel output (audio) signal or in order to obtain a stage width enhanced stereo output (audio) signal.
  • inventions according to the first aspect comprise a computer program for performing a method according to an embodiment according to the first aspect, when the computer program runs on a computer.
  • Embodiments according to the second aspect comprise an audio processing system for obtaining a processed audio signal on the basis of an input audio signal, wherein the audio processing system is configured, in order to obtain the processed audio signal, to apply a transient enhancement, which emphasizes transients in the input audio signal, or in a processed version of the input audio signal, relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal (e.g. by increasing energies of transient audio signal components, and/or by reducing energies of non-transient audio signal components).
  • a transient enhancement which emphasizes transients in the input audio signal, or in a processed version of the input audio signal, relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal (e.g. by increasing energies of transient audio signal components, and/or by reducing energies of non-transient audio signal components).
  • the audio processing system is configured to apply a stage width enhancement, which increases inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, or between corresponding audio signal components of different channels of a processed version of the input audio signal, and/or which emphasizes audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene (e.g. having a comparatively large inter-channel level difference) relative to audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to a center of the audio scene (e.g. having a comparatively small inter-channel level difference, e.g.
  • a stage width enhancement which increases inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, or between corresponding audio signal components of different channels of a processed version of the input audio signal, and/or which emphasizes audio signal components of the input audio signal, or of a processed version of the input audio signal, panned
  • a side signal extraction which selectively extracts audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene (e.g. having a comparatively large inter-channel level difference, e.g. while suppressing audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference).
  • a side signal extraction which selectively extracts audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene (e.g. having a comparatively large inter-channel level difference, e.g. while suppressing audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference).
  • the audio processing system is configured to apply an ambience enhancement, which provides decorrelated audio signal components on the basis of the input audio signal or on the basis of a processed version of the input audio signal.
  • transient enhancement e.g. having different characteristics
  • ambience enhancement allows improving a wide variety of audio signals, e.g. having different characteristics, and providing an improved quality of a respective rendered audio scene, for example with reduced artifacts.
  • the inventors recognized that in such a combination of audio processing techniques, in contrast to conventional approaches, emphasizing transients may allow maintaining the sonic characteristic of the input signal (e.g. punch) while still allowing for further modifications such as width enhancement.
  • the transient enhancement, stage width enhancement and ambience enhancement synergistically introduce degrees of freedom for improvement of a respective rendered audio scene, while cancelling, at least partially, shortcomings and/or artifacts of each other.
  • the ambience enhancement may cancel artifacts introduced by the transient enhancement, for example comprising a transient boost, and vice versa. It was further recognized that such a supplementary effect may be further improved for audio signals which are to be perceptually widened, hence using a stage width enhancement.
  • the audio processing system is configured, for applying the ambience enhancement, to obtain, on the basis of an input audio signal, e.g. a multi-channel input audio signal, a first modified audio signal, e.g. a first modified multi-channel audio signal, using a transient suppression (e.g. such that transients that are included in the input audio signal are suppressed in the first modified audio signal, e.g. such that both direct signal components, having a comparatively higher inter-channel correlation, and ambience signal components, having a comparatively lower inter-channel correlation, are included in the first modified audio signal, e.g. independent from an inter-channel correlation of signal components of the input audio signal).
  • a transient suppression e.g. such that transients that are included in the input audio signal are suppressed in the first modified audio signal, e.g. such that both direct signal components, having a comparatively higher inter-channel correlation, and ambience signal components, having a comparatively lower inter-channel correlation, are included in the first modified audio signal, e.g.
  • the audio processing system is further configured to combine the transient enhanced audio signal and the decorrelated audio signal, in order to obtain a combined audio signal, and the audio processing system is configured to apply the stage width enhancement, to increase inter-channel-level differences between corresponding audio signal components of different channels of the combined audio signal, and/or to emphasize audio signal components of the combined audio signal panned to one of the sides of an audio scene relative to audio signal components of the combined audio signal panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference, in order to obtain a stage-width-enhanced audio signal.
  • stage width enhanced audio signal may constitute the processed audio signal, or the audio processing system may be configured to derive the processed audio signal from the stage width enhanced audio signal using a post-processing.
  • the audio processing system is configured to apply the stage width enhancement, to increase inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, and/or to emphasize audio signal components of the input audio signal panned to one of the sides of an audio scene relative to audio signal components of the input audio signal panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference, in order to obtain a stage-width-enhanced audio signal.
  • the transient enhancement is configured to increase inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, or between corresponding audio signal components of different channels of a processed version of the input audio signal, in order to obtain an inter-channel-level-difference enhanced audio signal. Furthermore, the transient enhancement is configured to emphasize transients in the inter-channel-level-difference-enhanced audio signal, relative to non-transient audio signal components of the inter-channel-level-difference-enhanced audio signal, in order to obtain a transient enhanced audio signal.
  • the audio processing system is configured to apply a post-processing, e.g. to the combined audio signal, or to the stage-width-enhanced audio signal.
  • the post processing comprises one or more of the following functionalities: a dynamic range compression; an equalization; a loudness compensation.
  • the post-processing is configured to split-up a signal to be post-processed into a plurality of frequency ranges (e.g. into a low frequency range, a mid frequency range and a high frequency range, e.g. using a crossover filter bank).
  • the post-processing may further be configured to separately post-process the different frequency ranges, e.g. using a respective dynamic range compression and a respective equalization.
  • the audio processing system is configured to determine inter-channel-level differences (e.g. inter-channel level difference values D x (n,k); wherein is should be noted that, in some embodiments, indices (f,m) may be used instead of indices (n,k), wherein f and n are frequency indices, and wherein indices m and k are time indices) for a plurality of corresponding time-frequency bins of an input audio signal (e.g. between corresponding time-frequency bins of different channels of the input audio signal).
  • inter-channel-level differences e.g. inter-channel level difference values D x (n,k)
  • the audio processing system is further configured to scale spectral bin values, e.g. X 1 (n,k), X 2 (n,k), of corresponding time frequency bins (e.g. of corresponding time frequency bins of different channels of the input audio signal) of the input audio signal in dependence on the determined inter-channel level differences, in order to obtain corresponding scaled spectral bin values of a modified audio signal (e.g. corresponding time frequency bins of different channels of the first modified audio signal) such that an inter channel level difference between corresponding spectral bin values of the first modified audio signal, e.g.
  • the audio processing system is further configured to map (e.g. directly, e.g. without a computation of a panning index or the like) a value representing the inter-channel level difference, e.g. D x (n,k), (e.g. in a logarithmic domain, describing a ratio between energies in corresponding time-frequency bins of two channels in a logarithmized form) onto one or more values representing a change of the inter-channel level difference, e.g. in a logarithmic domain) using a linear or piecewise-linear mapping.
  • a value representing the inter-channel level difference e.g. D x (n,k)
  • D x (n,k) e.g. in a logarithmic domain, describing a ratio between energies in corresponding time-frequency bins of two channels in a logarithmized form
  • this may be performed, such that the value representing the inter-channel level difference is proportional to the value representing the change of the inter-channel level difference at least for inter-channel level differences having a first sign.
  • this approach may be performed using a piecewise linear mapping having a first proportionality factor between a value representing the inter-channel level difference and a (e.g. first) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is larger than an energy in a second channel, and having a second proportionality factor between the value representing the inter-channel level difference and a (e.g. second) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is smaller than an energy in a second channel.
  • the audio processing system is configured to provide, on the basis of an input audio signal (e.g. on the basis of a multi-channel input audio signal; e.g. on the basis of a stereo input in a time domain; e.g. on the basis of a multi-channel input audio signal in a frequency domain; e.g. on the basis of a combined audio signal obtained using a transient enhancement and an ambience enhancement) a first modified audio signal (e.g. a first modified multi-channel audio signal; e.g. a stage width enhanced stereo output (audio) signal in a time domain; e.g.
  • an input audio signal e.g. on the basis of a multi-channel input audio signal; e.g. on the basis of a stereo input in a time domain; e.g. on the basis of a multi-channel input audio signal in a frequency domain; e.g. on the basis of a combined audio signal obtained using a transient enhancement and an ambience enhancement
  • a first modified audio signal e.g
  • inter-channel-level-differences e.g. between corresponding audio signal components in different channels
  • inter-channel-level-differences e.g. between corresponding audio signal components in different channels
  • an intensity relationship between signal components having a comparatively small inter-channel-level-difference and signal components having a comparatively large inter-channel-level-difference is not changed by more than 10 percent
  • an intensity relationship between signal components having a comparatively small inter-channel-level-difference and signal components having a comparatively large inter-channel-level-difference is not changed by more than 10 percent
  • inter-channel-level-difference increaser e.g. using an inter-channel-level difference booster
  • such an audio processing system is configured to provide, on the basis of the input audio signal (e.g. on the basis of the stereo input in the time domain), a second modified audio signal, e.g. a second modified multi-channel audio signal, in which side signal components (e.g. corresponding audio signal components in multiple channels having a comparatively large inter-channel-level-difference) are emphasized relative to centered signal components (e.g. corresponding audio signal components in multiple channels having no inter-channel-level-difference or having a comparatively small inter-channel-level-difference) when compared to the input audio signal, and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal, e.g. using a side signal emphasizer or using a side signal extraction, and/or in which only side signal components are included while centered signal components are suppressed.
  • side signal components e.g. corresponding audio signal components in multiple channels having a comparatively large inter-channel-level-difference
  • such an audio processing system is configured to combine, e.g. using a linear combination; e.g. using a weighted linear combination, the first modified audio signal and the second modified audio signal (e.g. in order to obtain a stage width enhanced multi-channel output (audio) signal or in order to obtain a stage width enhanced stereo output (audio) signal).
  • embodiments according to the second aspect of the invention may comprise the functionalities, details and/or features (and/or even an apparatus, e.g. as a portion of a respective audio processing system) of any of the embodiments of the first aspect, both individually or taken in combination, for example, for applying the stage width enhancement.
  • the audio processing system is configured to obtain, on the basis of an input audio signal, e.g. a multi-channel input audio signal, a first modified audio signal, e.g. a first modified multi-channel audio signal, using a transient suppression (e.g. such that transients that are included in the input audio signal are suppressed in the first modified audio signal, e.g. such that both direct signal components, having a comparatively higher inter-channel correlation, and ambience signal components, having a comparatively lower inter-channel correlation, are included in the first modified audio signal, e.g. independent from an inter-channel correlation of signal components of the input audio signal).
  • a transient suppression e.g. such that transients that are included in the input audio signal are suppressed in the first modified audio signal, e.g. such that both direct signal components, having a comparatively higher inter-channel correlation, and ambience signal components, having a comparatively lower inter-channel correlation, are included in the first modified audio signal, e.g. independent from an inter-channel correlation
  • such an audio processing system is further configured to obtain, on the basis of the input audio signal, a second modified audio signal using an ambience extraction (which prefers audio signal components having a comparatively smaller inter-channel-correlation over audio signal components having a comparatively larger inter-channel correlation, for example which provides the second modified audio signal in dependence on inter-channel-correlation characteristics of audio signal components of the input audio signal).
  • an ambience extraction which prefers audio signal components having a comparatively smaller inter-channel-correlation over audio signal components having a comparatively larger inter-channel correlation, for example which provides the second modified audio signal in dependence on inter-channel-correlation characteristics of audio signal components of the input audio signal.
  • such an audio processing system is further configured to combine, e.g. using a weighted combination, the first modified audio signal, or a post-processed version of the first modified audio signal, and the second modified audio signal, or a post-processed version of the second modified audio signal, in order to obtain a combined audio signal.
  • such an audio processing system is further configured to apply a decorrelation to the combined audio signal, in order to obtain a processed audio signal, e.g. an ambience-enhanced signal.
  • the previously discussed decomposition of the input signal in a transient portion and an ambience portion may be supplemented by a decorrelation for further improvement of the perceived quality, e.g. with regard to immersion, of the respective rendered audio signal.
  • the audio processing system is configured to extract from the input audio signal ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation, e.g. left-right correlation, (e.g. an inter-channel correlation which is below a predetermined threshold value, or an inter-channel correlation which is smaller than an inter-channel correlation of direct signal components, e.g. such that direct signal components (and also non-transient direct signal components) are at least partially suppressed in the second modified audio signal, while ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation are non-suppressed or enhanced in the second modified audio signal).
  • a comparatively low inter-channel correlation e.g. left-right correlation
  • direct signal components and also non-transient direct signal components
  • such an audio processing system is further configured to emphasize ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation over, e.g. relative to, direct signal components and/or signal components having a comparatively high, or for example higher, inter-channel correlation
  • such an audio processing system is further configured to attenuate direct signal components and/or signal components having a comparatively high inter-channel correlation relative to ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation.
  • the audio processing system is configured to compute a covariance matrix on the basis of a frequency domain representation of the input audio signal (e.g. on the basis of a stereo input signal in a frequency domain, e.g. on the basis of X(f,k) or X(f,m) or X(n,k)), to compute a multichannel parametric Wiener filter on the basis of the covariance matrix, to estimate a direct signal using the multichannel parametric Wiener filter, and to subtract the estimated direct signal from the input audio signal, in order to obtain the second modified audio signal, e.g. Y(f,k) or Y(f,m) or Y(n,k), e.g. to thereby obtain the second modified audio signal such that signal components of the input audio signal having a comparatively higher inter-channel correlation are reduced when compared to signal components having a comparatively lower inter-channel correlation.
  • a covariance matrix on the basis of a frequency domain representation of the input audio signal (e.g. on the basis of a
  • the audio processing system is configured to apply a transient suppression to the second modified audio signal, in order to obtain a post-processed version of the second modified audio signal for combination with the first modified audio signal.
  • the transient suppression in the ambience enhancement may counteract artifacts from the transient enhancement and vice versa, whilst still allowing to improve an acoustic foreground (e.g. via the transient enhancement) and an acoustic background (e.g. via ambience enhancement).
  • the audio processing system is configured to obtain the first modified audio signal using a spectral weighting process with real weights (e.g. real valued weights).
  • the audio processing system is configured to obtain an estimate (e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate) of a, for example non-transient, sustained signal component, e.g. S(f,m), (e.g. an estimate of a temporal evolution of the sustained signal component, e.g. an estimate of an energy or an intensity of the sustained signal component) on the basis of one or more subband signals, e.g.
  • an estimate e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate
  • a, for example non-transient, sustained signal component e.g. S(f,m)
  • an estimate of a temporal evolution of the sustained signal component e.g. an estimate of an energy or an intensity of the sustained signal component
  • such an audio processing system is configured to obtain the first modified audio signal in dependence on the estimate of the sustained signal component.
  • the inventors recognized that a subband-wise sustained signal estimation provides good audio scene enhancement, for example with regard to its acoustic background.
  • the audio processing system is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency-band wise estimate, of a transient signal component, e.g. T ⁇ (f,m) or T ⁇ (n,k), (e.g. an estimate of a temporal evolution of the transient signal component, e.g. an estimate of an energy or an intensity of the transient signal component) on the basis of one or more subband signals (e.g.
  • a transient signal component e.g. T ⁇ (f,m) or T ⁇ (n,k)
  • subband signals e.g.
  • the audio processing system is configured to combine a plurality of frequency bins of a time-frequency-domain representation of the input audio signal, in order to obtain frequency band values associated with a plurality of frequency bands (wherein the frequency band values are examples of the subband values), and to obtain the estimate of the sustained signal component and/or the estimate of the transient signal component in dependence on the frequency band values.
  • the audio processing system is configured to obtain spectral weights, e.g. G g (f,m) or G g (n,k), associated with frequency bins, e.g. generally speaking with frequency subbands, on the basis of the frequency band values (e.g. using frequency-band-wise estimates of sustained signal components and/or of transient signal components derived from the frequency band values).
  • spectral weights e.g. G g (f,m) or G g (n,k
  • such an audio processing system is configured to obtain the first modified audio signal in dependence on the temporally smoothened (e.g. purged from peaks) subband signals or the temporally smoothened frequency band signals or the temporally smoothened subband signal envelopes or the temporally smoothened frequency band signal envelopes.
  • the envelope of the signal may be low pass filtered.
  • the transformed, e.g. STFT, magnitudes coefficients of frequency bins corresponding to one frequency band e.g. the bin indices 11-15
  • the audio processing system is configured to obtain sustained signal envelopes, e.g. ⁇ (f,m) or ⁇ (n,k), using a temporal smoothing, e.g. low pass filtering, of subband time trajectories or frequency band time trajectories (wherein frequency band time trajectories are a special case of subband time trajectories, wherein the term "subband time trajectories" comprises time trajectories comprising only a single frequency bin and time trajectories comprising a plurality of frequency bins) of the input audio signal, e.g.
  • a temporal smoothing e.g. low pass filtering
  • ⁇ (f,m) ⁇ X(f,m) or ⁇ (n,k) ⁇ magnitude(X(n,k))
  • respective sustained signal envelopes e.g. ⁇ (f,m) or ⁇ (n,k)
  • respective sustained signal envelopes e.g. ⁇ (f,m) or ⁇ (n,k)
  • respective sustained signal envelopes e.g. ⁇ (f,m) or ⁇ (n,k)
  • the subband time trajectories may, for example, represent a temporal evolution of energies in respective subbands of the input audio signal, wherein subbands may, for example, be frequency bands of a filterbank or may, for example, be frequency bins of a time-domain-to-spectral-domain transform or of a time-domain-to-frequency-domain transform.
  • the audio processing system is configured to obtain transient signal envelopes, e.g. T ⁇ (f,m) or T ⁇ (n,k), using a subtraction of sustained signal envelopes, e.g. ⁇ (f,m) or ⁇ (n,k), from respective subband signals, e.g. X(f,m), or from magnitudes or respective subband signals or from frequency band signals of the input audio signal or from magnitudes of frequency band signals of the input audio signal or from or the subband time trajectories, e.g. of magnitudes, of the input audio signal or from frequency band time trajectories, e.g. of magnitudes, of the input audio signal.
  • transient signal envelopes e.g. T ⁇ (f,m) or T ⁇ (n,k)
  • respective subband signals e.g. X(f,m)
  • the inventors recognized that efficiency may be improved by estimating sustained signal envelopes and determining transient signal envelopes by subtraction, e.g. instead of directly estimating transient signal envelopes. Furthermore, sustained signal envelopes may be estimated more robustly.
  • the audio processing system is configured to obtain spectral weights, e.g. G g (f,m) or G g (n,k), in dependence on sustained signal envelope values (e.g. in dependence on the sustained signal envelope values mentioned before, e.g. ⁇ (n,k)).
  • the audio processing system is configured to obtain spectral weights, e.g. G g (f,m) or G g (n,k), in dependence on transient signal envelope values (e.g. in dependence on the transient signal envelope values mentioned before, e.g. T ⁇ (n,k)).
  • an alternative weight determination may be provided, for example, for signal having strong transients, wherein it may me more efficient and/or robust to directly estimate the transients.
  • the audio processing system is configured to obtain respective spectral weights, e.g. G g (f,m), in dependence on a ratio between a weighted sum of respective sustained signal envelope values, e.g. ⁇ (f,m) or ⁇ (n,k), or an exponentiated version thereof, e.g. ⁇ (f,m) alpha or ⁇ (n,k) alpha , and of respective transient signal envelope values, e.g. T ⁇ (f,m) or T ⁇ (n,k), or an exponentiated version thereof, e.g.
  • T ⁇ (f,m) alpha or T ⁇ (n,k) alpha and magnitudes of respective subband signals, e.g. X(f,m) or X(n,k), or frequency band signals, or an exponentiated version thereof, e.g.
  • Such a determination of weights may allow for a computationally inexpensive audio scene enhancement.
  • the audio processing system is configured to obtain the sustained signal envelopes ⁇ (n,k) using a low pass filtering of subband time trajectories of magnitudes of a time-frequency-domain representation, e.g. X(f,m) or X(n,k), of the input audio signal, and using an application of a constraint ⁇ (n,k) ⁇
  • or of a constraint ⁇ (n,k) ⁇
  • , to obtain the transient signal envelopes T ⁇ (n,k) according to T ⁇ n , k X n , k ⁇ S ⁇ n , k .
  • the audio processing system is configured to obtain an estimate (e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate) of a, for example non-transient, sustained signal component, e.g. S(f,m), (e.g. an estimate of a temporal evolution of the sustained signal component, e.g. an estimate of an energy or an intensity of the sustained signal component) on the basis of one or more subband signals (e.g.
  • an estimate e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate
  • a, for example non-transient, sustained signal component e.g. S(f,m)
  • an estimate of a temporal evolution of the sustained signal component e.g. an estimate of an energy or an intensity of the sustained signal component
  • subband signals may be signals associated with a single frequency bin, or signals associated with a frequency band comprising a plurality of frequency bins).
  • such an audio processing system is configured to obtain the modified audio signal, in which transient signal components are enhanced or reduced, in dependence on the estimate of the sustained signal component.
  • the audio processing system is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency-band wise estimate, of a transient signal component, e.g. T ⁇ (f,m) or T ⁇ (n,k), (e.g. an estimate of a temporal evolution of the transient signal component, e.g. an estimate of an energy or an intensity of the transient signal component) on the basis of one or more subband signals (e.g.
  • a transient signal component e.g. T ⁇ (f,m) or T ⁇ (n,k)
  • subband signals e.g.
  • ambience enhancement and transient enhancement may share a common transient determination, therefore reducing the computational complexity, since intermediate results, such as the estimated transients, may be used by multiple (enhancement) modules.
  • Embodiments according to the second aspect of the invention comprise a method for obtaining a processed audio signal on the basis of an input audio signal, wherein the method comprises, in order to obtain the processed audio signal, applying a transient enhancement, which emphasizes transients in the input audio signal, or in a processed version of the input audio signal, relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal (e.g. by increasing energies of transient audio signal components, and/or by reducing energies of non-transient audio signal components).
  • a transient enhancement which emphasizes transients in the input audio signal, or in a processed version of the input audio signal, relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal (e.g. by increasing energies of transient audio signal components, and/or by reducing energies of non-transient audio signal components).
  • a side signal extraction and/or which selectively extracts audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene (e.g. having a comparatively large inter-channel level difference, e.g. while suppressing audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to a center of the audio scene, e.g. having a comparatively small inter-channel level difference).
  • the method further comprises, in order to obtain the processed audio signal, applying an ambience enhancement, which provides decorrelated audio signal components on the basis of the input audio signal or on the basis of a processed version of the input audio signal.
  • the method as described above is based on the same considerations as the above-described audio processing system.
  • the method can, by the way, be completed with all features and functionalities, which are also described with regard to the audio processing system.
  • elements having same ending numerals or same names may comprise same, according, or corresponding functionalities or details, or may be examples of each other, e.g. such as inputs 101 and 201 (corresponding numbers) or stage width enhancement unit 100 and stage width enhancement 620 (corresponding names).
  • Fig. 1 shows a schematic view of an apparatus for processing an audio signal according to an embodiment.
  • Fig. 1 shows apparatus 100, which may, for example, be a stage width enhancement apparatus or a stage width enhancement block, e.g. of an audio processing system.
  • the apparatus comprises an Inter Channel Level Difference Booster 110, a Side Signal Extraction unit 120 and a combiner 130.
  • the apparatus is provided with an input audio signal 101.
  • the Inter Channel Level Difference Booster 110 is configured to provide, on the basis of an input audio signal 101 a first modified audio signal 111 in which inter-channel-level-differences are increased when compared to inter-channel-level-differences in the input audio signal 101.
  • the Side Signal Extraction unit 120 is configured to provide, on the basis of the input audio signal, a second modified audio signal 121 in which side signal components are emphasized relative to centered signal components when compared to the input audio signal 101, and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal 101, and/or in which only side signal components are included while centered signal components are suppressed.
  • the combiner 130 is configured to combine the first modified audio signal 111 and the second modified audio signal 121.
  • the combined signal is indicated, as an example, as a Stage Width Enhances Signal 102, as for the optional form of the apparatus 100 as a stage width enhancement block or apparatus.
  • combiner 130 is indicated as a summation, however, embodiments are not limited to a specific form of combination. For example, besides addition, weighted addition and/or even subtraction may be performed.
  • Fig. 2 shows a schematic view of an apparatus for processing an audio signal with additional optional details, according to an embodiment.
  • Fig. 2 shows apparatus 200 comprising an Inter Channel Level Difference Booster 210, a Side Signal Extraction unit 220, a combiner 230, a transformer 240 and inverse transformers 250.
  • the apparatus is provided with an input audio signal 201.
  • the optional transformer 240 is shown, as an example, as a short term Fourier transformer. However, it is to be noted that other forms of transforms or for example filterbanks, see e.g. Fig. 15 , may be used as well (and accordingly for corresponding inverse transformers 250 or inverse filterbanks).
  • the input signal which is shown, as an optional example, as a stereo signal in time domain, may be time-domain-to-spectral-domain transformed or time-domain-to-frequency-domain transformed by transformer 240. In particular, a frequency band-wise or bin-wise processing may be performed.
  • the Inter Channel Level Difference Booster 210 is configured to provide, on the basis of the transformed input audio signal 201 a first modified audio signal in which inter-channel-level-differences are increased when compared to inter-channel-level-differences in the input audio signal 201 (and hence, for example, its transformed counterpart).
  • the Side Signal Extraction unit 220 is configured to provide, on the basis of the transformed input audio signal, a second modified audio signal, in which side signal components are emphasized relative to centered signal components when compared to the input audio signal 201 (and hence, for example, its transformed counterpart), and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal 201 (and hence, for example, its transformed counterpart), and/or in which only side signal components are included while centered signal components are suppressed.
  • the Side Signal Extraction unit 220 may be configured to increase or boost energies of off-center signal components and/or to attenuate or reduce energies of centered signal components.
  • the inter-channel-level-difference boost as well as the side signal extraction may be performed in corresponding time frequency bins (or time frequency bands) of different channels of the input audio signal.
  • Results of the Inter Channel Level Difference Booster 210 and the Side Signal Extraction unit 220 are provided to inverse transformers 250 for performing an inverse transformation, e.g. frequency-domain-to-time-domain-transformation or spatial-domain-to-time-domain transformation.
  • an inverse transformation e.g. frequency-domain-to-time-domain-transformation or spatial-domain-to-time-domain transformation.
  • the modified and re-transformed signals may then be combined in combiner 230 to provide an output signal 202.
  • the output signal is shown, for the example of the apparatus 200 being a stage width enhancement block or apparatus, as a stage width enhanced stereo output signal.
  • Fig. 3 shows a schematic view of an Inter Channel Level Difference Booster according to embodiments of the invention.
  • FIG. 3 shows Inter Channel Level Difference Booster 300 comprising an inter-channel-level-difference calculation unit 310, a gain calculation unit 320, comprising, as optional features, a subunit 320l for calculation a gain 321l for a left channel 301l of the input signal and a subunit 320r for calculation a gain 321r for a right channel 301r of the input signal and multiplication units 330.
  • the Inter Channel Level Difference Booster 300 is provided with an input signal, wherein, as shown as an optional feature in Fig. 3 , said input signal may be provided in frequency domain.
  • the input signal may referred to as X(f, k).
  • f may refer to a frequency bin.
  • X(n,k) may be used, for example, interchangeably.
  • f may denote the frequency bin index
  • k the time index
  • n optionally the frequency band index or for some embodiments n may be used interchangeably for f.
  • the input signal may comprise two (or more) input channels, hence a left and right input signal 3011 and 301r.
  • the input signals 301l and 301r are provided to inter-channel-level-difference calculation unit 310.
  • the inter-channel-level-difference calculation may be performed per frequency bin, e.g. as indicated by index f.
  • embodiments, and in particular the embodiment according to Fig. 3 are not limited to such an approach.
  • inter-channel-level-difference calculation unit 310 the inter-channel-level-difference, referred to as D(f,k) (e.g. also referred to as D x (f,k) or respectively D x (n,k)) is obtained and provided to the gain calculation unit 320. It is to be noted that inter-channel-level differences may be determined for a plurality of corresponding time-frequency bins of the input audio signal.
  • the gain calculation is performed channel-wise in respective subunits (and, for example, as another optional feature per frequency bin).
  • gain values 321l also referred to as G 1 (f, k)
  • 321r also referred to as G 2 (f, k)
  • These gains are used to scale the respective input signals 301l and 301r (wherein, for example, signals 301l and 301r correspond to each other regarding their indices (f,k)), using multiplication units 330, so as to provide output signals 302l (e.g. Y 1 (f,k)) and 302r (e.g. Y 2 (f,k)) in frequency domain, (also referred to as Y(f,k)).
  • the weight determination and scaling may, for example, be performed for a plurality of corresponding time-frequency bins of the input audio signal. Such a processing may be performed consecutively (e.g. for streaming) or in parallel (e.g. when the whole audio signal to be processed is fully available).
  • the inter channel level difference between corresponding spectral bin values of such modified audio signals 302l, 302r may be increased and hence boosted, when compared to an inter-channel level difference between corresponding spectral bin values of the input audio signals, e.g. X 1 (f,k), X 2 (f,k), for example associated with a same frequency (referring to index f or n) and a same time portion, e.g. having time index k.
  • the first modified audio signal 111 may be or may comprise audio signals 302l, 302r.
  • the weight calculator 320 may be configured to determine the weights (here as an example G 1 and G 2 (embodiments are not limited to stereo signals), so as to maintain a total energy of the channels from input 301(l+r) to output 302(l+r), for example associated with a same frequency (referring to index f or n) and a same time portion, e.g. having time index k.
  • the weights may be calculated in a direct manner, for example without computing a panning index, and hence with good computational efficiency.
  • Inter Channel Level Difference Booster 300 may be an example for Inter Channel Level Difference Boosters 110 and/or 210.
  • Fig. 4 shows a schematic view of a side signal extraction unit according to embodiments.
  • Fig. 4 shows side signal extraction unit 400 comprising a gain calculation unit 420 and a multiplication unit 430.
  • Unit 400 may correspond to side signal extraction unit 210.
  • f may denote the frequency bin index, and k the time index.
  • Input signal 401 (here shown as an example as stereo signal), referred to as X(f,k) is provided to the gain unit 420, for determining gain values 411, referred to as G(f,k). These gain values are used to scale the input signal 401 using multiplication unit 430 to obtain an output signal 402 (here accordingly shown as an example as stereo signal), referred to as Y(f, k).
  • the gains 411 may be determined so as to maintain a ratio between energies of off-center signal components and centered signal components, within a tolerance of +/-10 percent or for example +/-5 percent or for example +/-1 percent.
  • the gain unit 420 may be configured to determine the gain values 411 for scaling based on corresponding signal components in corresponding time frequency bins of different channels of the input audio signal in dependence on an inter-channel-level difference between the respective corresponding signal components, and/or in dependence on an inter-channel time difference between the respective corresponding signal components, and/or in dependence on an inter-channel phase difference between the respective corresponding signal components.
  • signal 402 may be an example for the second modified audio signal 121.
  • Fig. 5 shows a schematic block diagram of a method according to embodiments of the first aspect of the invention.
  • Fig. 5 shows method 500 comprising providing, 510, on the basis of an input audio signal, a first modified audio signal in which inter-channel-level-differences are increased when compared to inter-channel-level-differences in the input audio signal, providing, 520, on the basis of the input audio signal, a second modified audio signal in which side signal components are emphasized relative to centered signal components when compared to the input audio signal, and/or in which centered signal components are attenuated relative to side signal components when compared to the input audio signal and/or in which only side signal components are included while centered signal components are suppressed and combining, 530 the first modified audio signal and the second modified audio signal.
  • Fig. 6 a)-c) show schematic views of different variations (e.g. variants) of audio processing systems according to embodiments.
  • Fig. 6a )-c) may show overview block diagrams shows, hence block diagrams of various ways how the basic algorithms according to embodiments may be combined.
  • Fig. 6 a)-c) show audio processing systems 600a, 600b, 600c comprising a transient enhancement unit 610, a stage width enhancement unit 620, and an ambience enhancement unit 630.
  • the audio processing systems 600a, 600b, 600c comprise combination units 640 and a post processing units 650.
  • the audio processing systems 600a, 600b, 600c are configured to receive an input audio signal 601.
  • the audio processing systems 600a, 600b, 600c are configured to apply a transient enhancement, using transient enhancement unit 610, which emphasizes transients in the input audio signal (see Variants A and B), or in a processed version of the input audio signal (see Variant C), relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal.
  • transient enhancement unit 610 which emphasizes transients in the input audio signal (see Variants A and B), or in a processed version of the input audio signal (see Variant C), relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal.
  • the audio processing systems 600a, 600b, 600c are further configured to apply a stage width enhancement, using stage width enhancement unit 620, which increases inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal (see Variants A and C), or between corresponding audio signal components of different channels of a processed version of the input audio signal (see Variant B), and/or which emphasizes audio signal components of the input audio signal (see Variants A and C), or of a processed version of the input audio signal (see Variant B), panned to one of the sides of an audio scene relative to audio signal components of the input audio signal (see Variants A and C), or of a processed version of the input audio signal (see Variant B), panned to a center of the audio scene, and/or which selectively extracts audio signal components of the input audio signal (see Variants A and C), or of a processed version of the input audio signal (see Variant B), panned to one of the sides of an audio scene.
  • stage width enhancement unit 620 which increases inter-
  • the audio processing systems 600a, 600b, 600c are further configured to apply an ambience enhancement, using ambience enhancement unit 630, which provides decorrelated audio signal components on the basis of the input audio signal (see Variants A and B) or on the basis of a processed version of the input audio signal (see Variant C).
  • the audio processing systems 600a, 600b, 600c are configured to obtain a processed audio signal 602 on the basis of the input audio signal 601.
  • combining units 640 as an example as adding units, are shown. However, embodiments are not limited to a specific kind of combining. As examples, summation, weighted summating and/or even subtraction may be performed.
  • the audio processing systems 600a, 600b, 600c comprise a post processing unit 650, for further improvement of the enhanced and combined signal (based on enhancements 610, 620, 630).
  • any of the units 610, 620 and 630 may receive a version of the input signal which is preprocessed by any of the other units.
  • the preprocessing may even comprise applying more than one enhancement and then combining the intermediate signals (see e.g. Variant B).
  • the enhancements 610, 620 and 630 may be performed in parallel or consecutively or a combination of parallel enhancement and consecutive enhancement (see Variants B and C) may be performed.
  • stage width enhancement unit 620 may comprise any or all of the features as the stage width enhancement units as discussed in the context of Fig. 1 to 5 .
  • apparatus 100 as well as apparatus 200 may be examples of stage width enhancement unit 620.
  • the stage width enhancement unit may hence comprise any or all of the further optional features as disclosed in the context of Fig. 3 and 4 .
  • Fig. 7 shows a schematic view of a transient enhancement unit, e.g. a transient enhancement block, according to embodiments.
  • transformer 710 and the inverse transformer 750 are optional units. As shown, the transformation used may be a short term Fourier transform, however, as previously discussed, embodiments are not limited to the same.
  • input audio signal 701 (e.g. a preprocessed audio signal) is provided to the Inter Channel Level Difference Booster 720, which is configured to increase inter-channel-level differences between corresponding (e.g. with regard to (f, k), e.g. with regard to (n, k)) audio signal components of different channels of the input audio signal (e.g. as in Variant A, B of Fig. 6 ), or between corresponding audio signal components of different channels of a processed version of the input audio signal (e.g. as in Variant C of Fig. 6 ), in order to obtain an inter-channel-level-difference enhanced audio signal.
  • the input audio signal 701 is therefore shown as a stereo signal, hence having two channels.
  • a result of the Inter Channel Level Difference Booster 720 is provided to the transient enhancement subunit 730, which is configured to emphasize transients in the inter-channel-level-difference-enhanced audio signal, relative to non-transient audio signal components of the inter-channel-level-difference-enhanced audio signal, in order to obtain a transient enhanced audio signal.
  • the output signal may be provided based on a subtraction between a result of the transient enhancement subunit 730 and the Inter Channel Level Difference Booster 720, and the inverse transformation of unit 750.
  • Fig. 8 shows a transient enhancement subunit according to an embodiment of the invention.
  • the transient enhancement subunit may, for example be a transient enhancement / transient suppression module (e.g. depending on a parametrization for the weight computation).
  • Transient enhancement subunit 800 comprises as combiner 810, weight computation units 820 1 , ..., 820 n , a weight calculation unit 830 and a multiplication unit 840.
  • Transient enhancement subunit 800 may be an example for transient enhancement subunit 730.
  • f may denote the frequency bin index, k the time index and n optionally the frequency band index (or for example interchangeably a bin index). It is to be noted that embodiments, in general, may optionally comprise frequency bin-wise processing, and/or frequency band-wise processing, e.g. if several bins are combined to a frequency band. However, a single bin may as well be understood according to some embodiments as a frequency band. Hence, f and n may be used interchangeably according to some embodiments or f as bin index and n as band index.
  • Transient enhancement subunit 800 is provided with input signal 801.
  • the input signal may be a frequency domain signal, e.g. referred to as X(f, k).
  • the combiner 810 may be configured to combine frequency bins of the input signal 801 to frequency bands.
  • Band-wise combined inputs X(1, k) to X(n, k) may then be provided to respective weight computation units 820 1 , ..., 820 n to determine band-wise weights, which may then be used to determine, based thereon, bin-wise weights, using weight calculation unit 830.
  • bin-wise weights are then used to scale input signal 801, using multiplication unit 840, so as to obtain an output signal in frequency domain 802, referred to as Y(f, k).
  • Fig. 9 shows a schematic view of a post processing unit according to embodiments of the invention.
  • Fig. 9 shows post processing unit 900 comprising a dynamic range compression unit 910, an equalizer 920 and a loudness compensation unit 930.
  • the post processing unit 900 may comprise an arbitrary selection of the units 910, 920, 930.
  • the sequence of blocks can be altered.
  • Post processing unit 900 may be an example of post processing unit 650. Hence, the post-processing may be applied to a combined audio signal 901, or to a stage-width-enhanced audio signal 901, in order to obtain a post processed signal 902.
  • Fig. 10 shows a schematic view of a post processing unit with additional, optional features, according to embodiments of the invention.
  • Fig. 10 shows post processing unit 1000, which may be an example of post processing unit 900 or 650, with a loudness normalization unit 1010, a crossover filter bank 1020 and dynamic range compression units 1030l, 1030m, 1030h and equalizers 1040l, 1040m, 1040h, as well as a combination unit 1050.
  • the separately processed signals may be combined using combination unit 1050, for example as a summation or weighted summation, in order to obtain the output signal 1002.
  • the post processing may be performed in time domain, e.g. using a time domain input signal 1001 to provide a time domain output signal 1002.
  • the signals are indicated as stereo signals.
  • Fig. 10 it is to be noted that the sequence of blocks and the number of frequency ranges can be altered. Hence, embodiments may comprise more, less or different blocks an frequency ranges.
  • the loudness normalization 1010 is optional as well.
  • Fig. 11 shows a schematic view of an example of an ambience enhancement unit according to embodiments of the invention.
  • Fig. 11 shows ambience enhancement unit 1100, which may be an example for ambience enhancement unit 630 comprising, as optional features, a transient suppression unit 1110, an ambience extraction unit 1120 and another optional transient suppression unit 1110, as well as a combination unit 1130 and an optional decorrelation unit 1140.
  • the input signal 1101 may be provided to the ambience extraction unit 1120, for example, in order to extract audio signal components having a comparatively smaller inter-channel-correlation over audio signal components having a comparatively larger inter-channel correlation.
  • a result thereof for example for providing an ambience signal processing
  • input signal 1101 for example for providing a background signal processing
  • transient suppression units 1110 These may comprise the details and functionalities as shown in Fig. 8 , however with weights chosen, so as to attenuate transients instead of boosting the same.
  • Respective results may be combined using combination unit 1130, e.g. for performing an addition or a weighted summation.
  • the combined result may be decorrelated using the decorrelation unit 1140.
  • Figs. 1 , 9 and 11 may show detailed block diagram, hence showing more detailed block diagrams explaining the basic algorithms according to embodiments, e.g. with more optional details in comparison to Fig. 6a )-c).
  • Fig. 12 shows a schematic view of an example of an ambience enhancement unit (for example also referred to as ambience enhancement block), with additional, optional features, according to embodiments of the invention.
  • Fig. 12 shows ambience enhancement unit 1200 comprising, as optional features, transform units 1210, as an example in the form of short term Fourier transform units (e.g. having functionalities as previously discussed), optional transient suppression units 1220 an ambience extraction unit 1230, an optional Inter Channel Level Difference Booster 1240 and optional inverse transform units 1250, as an example in the form of inverse short term Fourier transform units (e.g. having functionalities as previously discussed).
  • ambience enhancement unit 1200 comprises a combination unit 1260 and as an optional feature a decorrelation unit 1270.
  • the background signal processing may, as well as the ambience signal processing, be performed in frequency or spatial domain. Furthermore, additionally inter channel level differences may be boosted for the background signal, using unit 1240.
  • Fig. 13 shows a schematic view of an ambience extraction unit according to embodiments of the invention.
  • Fig. 13 shows ambience extraction unit 1300 comprising a covariance computation unit 1310, a filtering unit 1320 and a direct signal estimator 1330, as well as a subtraction unit 1340.
  • the ambience extraction unit 1300 may be an example for the ambience extraction unit 1230.
  • the input signal 1301 is shown, as an example, as a stereo input signal in frequency domain and may be referred to as X(f, k). Again, frequency band-wise processing may be performed as well.
  • a covariance matrix on the basis of the multi-channel frequency domain representation 1301 of the input audio signal may be determined, for a subsequent Wiener filtering (which may be computed using unit 1320) in order to estimate a direct signal using direct signal estimator 1330.
  • the ambience extraction unit 1300 may be configured to subtract the estimated direct signal from the input audio signal 1301 (e.g. the frequency domain representation thereof) to obtain output signal 1302, e.g. in the shown form of an Ambient output signal in frequency domain.
  • FIG. 14 shows a schematic view of a detailed block diagram of an audio processing system (e.g. as a whole system) according to embodiments of the invention.
  • Fig. 14 shows audio processing system 1400 comprising the subsystems 200, 700, 1000 and 1200 as previously discussed in Fig. 2 , 7 , 10 and 12 .
  • System 1400 may show an example for system 600a shown in Fig. 6a .
  • the input signal is indicated as stereo input signal 1401, e.g. corresponding to input signal 601 and the output signal is indicated as stereo output signal 1402, e.g. corresponding to input signal 602.
  • Intermediate results of the transient enhancement block 700, the stage width enhancement block 200 and the ambience enhancement block 1200 may be combined using combination unit 640 (e.g. as discussed regarding Fig. 6a ).
  • the blocks 200, 700 and 1200 may share a common transform unit 1410, e.g. corresponding to units 240, 710 and 1210.
  • Fig. 3 , 4 , 8 and 13 may be Module Block diagrams, for example constituting portions of such a system 1400.
  • one aim of methods according to embodiments is to enhance an audio signal with respect to listener preference and general sound quality.
  • algorithms are, for example, combined (e.g. individually or in combination) that modify spatial characteristics, dynamic characteristics, and/or tonal characteristics.
  • These may, for example, comprise dedicated modules transient enhancement, ambience enhancement, and stage width enhancement.
  • One aim of at least some methods according to embodiments is to process an audio signal such that it is preferred over the unprocessed signal when rating its sound quality.
  • Audio signals like musical recordings, audio books or movie sounds are created by recording, synthesizing, processing and mixing (adding) sound sources and finally processing the mixture.
  • the processing of the audio signals is done in analog or digital domain and comprises the application of various devices (analog) or algorithms (digital), where digital processing is more often used than analog.
  • the processing of sound sources and mix aims to create a final mix of high sound quality where all sound sources are (or at least should be) audible and intelligible.
  • the applied tools, procedures and best practices have been improved in the history of music and audio production.
  • the preferences of human listeners are also evolving over time. Consequently, musical recordings of same genres but different eras have different characteristics.
  • One aim of at least some methods according to embodiments is to modify characteristics of older music productions such that they sound as having been produced more recently.
  • the characteristics of audio signals comprise spatial characteristics, dynamic characteristics, tonal characteristics, loudness, timbral attributes and for example more.
  • loudness level where loudness denotes a perceptual attribute in contrast to the physical level.
  • this preference can, for example only, be achieved with restrictions (e.g. one or more of the following) that loudness should not exceed armful levels (exposure to sounds at high levels can cause hearing loss), all reproduced signals should have consistent levels, and reproduced sound should not mask other relevant sounds in the environment, e.g. conversations or alarms.
  • Human hearing processes sound in frequency bands with non-linear (compressive) mechanisms and temporal integration on time-variant stimuli.
  • Computational models of loudness may determine specific loudness in each frequency band and may accumulate these to compute the overall loudness. Due to the compressive effect, stimuli that are dense in the time-frequency representation may be perceived as being louder than spare signals when reproduced at same SPL.
  • Stereophonic sound reproduction aims to emulate a rich spatial sound image by a small number of signal channels and loudspeakers. This may be achieved by spatial cues of sound signals at both ears. Sound sources reproduced at different levels by each transducer may be perceived as coming from different locations and form a stereo image. Uncorrelated sounds, e.g., wind noise, reverberations, of signals reproduced with slight variations at different positions may not be localized as coming from one direction and may thereby contribute to a natural and aesthetically pleasing experience. Therefore, spatial characteristics are important for listening preference. Enhancing decorrelation and inter-channel level differences may yield richer stereophonic images and can increase listener preference.
  • Another aim of methods according to at least some embodiments is to compensate for limitations of the environment in which reproduced sound is presented.
  • the ideal listening environment for loudspeaker reproduction is, for example, quiet (e.g. loudness of external sound sources is below or near threshold of hearing), has walls, floor and ceiling reflecting sound waves and thereby contributing uncorrelated and diffuse sounds.
  • the size of the room should be such that early reflections reach the ears within a time span such that reflections are not heard as echoes and not merged with the direct sounds and reduce their clarity.
  • the geometry of the room and the positions of listeners and loudspeakers should not result in large scaling of energy at different frequencies, that is, the effect of sound reproduction in the room than modelled a linear time-invariant system should have a smooth magnitude transfer function.
  • the material of walls, floor, ceiling should reflect sound to some extend without amplifying particular frequencies.
  • Vehicles for example, have limitations in these respects, and at least some methods according to embodiments may compensate them.
  • noise is emitted from engine, tires on roads and air turbulences.
  • Environmental noise partially masks the reproduced program and thereby reduces its overall loudness and modifies the tonal balance.
  • Masking follows the same principles as loudness perception, it is frequency selective and nonlinear. Since softer sounds are masked more than louder sounds, a dynamic processing that reduces the level differences between soft and loud sounds is beneficial in this respect.
  • the geometry of the cabin and material of the interior may be less appropriate for music listening as compared to a living room, for example. Therefore, the contribution of a great listening room to spatial cues of program material may be missing.
  • the same may hold for positioning of loudspeakers and this may result in spatial cues of less quality as intended during production.
  • a further aspect according to embodiments relates to the development of mobile sound reproduction technology, in particular mobile phones, laptops, small wireless loudspeakers. These devices are often used outside and in noisy environments.
  • the dissemination of this technology has influenced listening habits (e.g., listening more in noisy environments) and also practices in music production, in particular the use of dynamic range compression in mastering. Modern music productions are more dynamic range compressed, and presumably because this 1) improves masking of environmental noise, 2) is perceived louder, and 3) is simply preferred by artists and listeners.
  • methods according to embodiments may comprise one or more of a transient processing, a stage width enhancement, an inter-channel level difference enhancement, a side signal extraction, an ambience enhancement and/or a post-processing.
  • Fig. 6a )-c) may show various ways how the above basic algorithms according to embodiments are combined, Fig. 1 , 9 , 11 may show more detailed block diagrams explaining such algorithms.
  • Fig. 15 shows a schematic view of a module for processing an input signal according to embodiments of the invention.
  • Module 1500 may be configured to boost Inter-channel level differences for audio rendering and/or to process transients.
  • Module 1500 is provided with an input signal 1501, which is processed by a filterbank 1510.
  • the input signal is indicated as x(t) in time domain, however an optional previous transformation may be present.
  • Filterbank 1510 may be configured, as shown to provide frequency bin-wise (or optionally band-wise e.g. subband-wise) input signals, here as an example referred to as X(1,k), ..., X(n,k).
  • Bin-wise (or band-wise) weights G may be determined, using weight computation units 1520, based on the filter outputs, for scaling the filter outputs using multiplication units 1530, in order to provide band-wise output signals Y(1,k), ..., Y(n,k). These may be retransformed to a time domain output signal y(t) using an inverse processing of filterbank 1540.
  • Fig. 15 may show a module and respectively method for boosting inter-channel level differences for audio rendering according to embodiments is disclosed.
  • Module 1500 may correspond to Inter Channel Level Difference Booster 300, with additional filterbank 1510 for time to frequency transformation.
  • computation shown in Fig. 3 may correspond to the computation shown in Fig. 15 .
  • such a method for manipulating inter-channel level differences of an audio signal having two channels may have the aim to widen the stereo image of the signal, which may be useful for many audio rendering systems, e.g. upmixing and surround sound.
  • Such an inventive module e.g. module 1500, may, for example, increase inter-channel level differences of an audio signal in time-frequency domain. This may, for example, result in the audible effect that the perceived position in the stereo image of sources panned off center are further moved away from the center of the stereo image. This may, for example, be achieved without modifying the loudness of the sources and/or without modifying the timbre, loudness and/or position of sources panned to the center.
  • This may, for example, be applied for upmixing of two-channel stereo signals to surround signals, and/or stereophonic enhancement for rendering applications.
  • Algorithmic description (2) of at least some embodiments for manipulating inter-channel level differences of an audio signal The processing may, for example, work on frequency bands that are processed independently of each other. Therefore, a filterbank (or, alternatively, frequency transform) may, for example, be used or even required, optionally together with its inverse processing for computing the broad-band output signal (or time signal). This can, for example, be implemented by means of short-term Fourier transform (STFT) and its inverse.
  • STFT short-term Fourier transform
  • the block diagram of Fig. 15 may illustrate the spectral weighting.
  • the module may, for example, be implemented as a spectral weighting of STFT coefficients with real-valued weights.
  • the approach may hence be performed for all frequency bins, e.g. 1, ..., n.
  • Spectral weights G 1 (n,k), G 2 (n,k) may, for example, be computed as function of inter-channel level differences (ICLD), i.e. differences of levels of both channels in each time-frequency bin.
  • ICLD inter-channel level differences
  • Spectral weights G 1 (n,k), G 2 (n,k) may, for example, be computed such that ICLD of output is larger than ICLD of input for all coefficient pairs with non-zero ICLD.
  • Spectral weights G 1 (n,k), G 2 (n,k) may, for example, be computed such that total level in each time-frequency bin of output equals the total level of input.
  • module 1500 may show a schematic view of a block diagram of the spectral weighting.
  • the computation of the spectral weights may be performed using weight computation units 1520 using the above and following formulas.
  • module 1500 may correspond to transient enhancement 800, with additional filterbank 1510 for time to frequency transformation.
  • the shown in Fig. 8 may correspond to the computation shown in Fig. 15 .
  • module 1500 may correspond to transient suppression 1110, 1220, with additional filterbank 1510 for time to frequency transformation.
  • transient boost or transient suppression may be achieved.
  • the input signal may be assumed to be an additive mixture of a sustained signal and a transient signal.
  • the spectral envelope of the sustained signal may, for example, be estimated first, and the spectral envelope of the transient signal may, for example, be obtained using the first result and the signal model.
  • Spectral weights may, for example, be computed from the estimated spectral envelopes of the sustained signal and the transient signal, for example, by means of Wiener filtering and/or a generalized spectral weighting method, such that the sustained signal or the transient signal are attenuated or amplified.
  • the module may, for example, attenuate or amplify transient signal components.
  • Transient signal components are, for example, the attack part of a note, a percussive hit, a clicking sound or the first 100 ms of a gun shot recording.
  • the attenuation of transients may, for example, be applied for processing surround signals, artificial reverberation and decorrelation and/or stereophonic enhancement.
  • the amplification of transients can be desired when it is mixed with a surround signal in order to maintain the sonic characteristic of the input signal (e.g. punch) or as a tool for sound design.
  • the counterpart of the transient may be a stationary or steady-state and/or sustained signal component.
  • Sustained signal components may, for example, relate to sound sources with slowly evolving temporal envelopes.
  • Algorithmic description (2) of at least some embodiments for transient processing The processing works (may hence, according to embodiments, optionally be performed) on frequency bands that are processed (e.g. at least mainly) independently of each other. Therefore, a filterbank (or, alternatively, frequency transform) may be used or even required, for example, together with its inverse processing for computing the broad-band output signal (or time signal). More specifically, the module may, for example, be implemented as a spectral weighting processing, e.g. with real-valued weights.
  • the block diagram of Fig. 15 illustrates an example for the spectral weighting.
  • transient processing two basic approaches to transient processing are feasible and may hence be performed according to embodiments.
  • One is to detect first the occurrence of a transient sound and then to manipulate it.
  • the second is to replace the binary detection of a transient by an estimation of the "transientness” or transient presence probability (TPP) which is then used to control the subsequent processing.
  • TPP transient presence probability
  • the processing is a scaling of the signal which is inversely proportional to the TPP, transients can be suppressed.
  • TPP transient presence probability
  • the input time domain signal x(t) may, for example, be assumed to be an additive mixture of a sustained signal s(t) and a transient signal t(t).
  • x t s t + t t
  • the input signal may, for example, be processed in the frequency domain, for example by using the STFT or another transform or a filterbank, leading to sub-band signals for each frequency band.
  • the matrix X(n, k) may, for example, be represent the sub-band signals of the input signal x m or x(t), with subsampled time index k and frequency band index n.
  • the matrices of sub-band signals are denoted by uppercase letters and are represented by real-valued non-negative quantities, i.e. in case of complex-valued sub-band signals only their magnitudes respectively powers may, for example, be modified whereas their phases may, for example, not be modified.
  • the transient signals and/or sustained signals may, for example, be modified by means of a spectral weighting method, for example each sub-band signal may, for example, be multiplied with a gain factor G(n,k).
  • G(n,k) a gain factor
  • Y n k G n k X n k
  • the output time signal y(t) (e.g. also referred to as y m ) may, for example, be obtained from the sub-band signals Y(n, k) by using the inverse processing corresponding to the STFT or transform or filterbank used for computing the input sub-band signals.
  • the spectral weights G(n,k) may, for example, be computed using estimates of the sub-band signals of sustained signal S(n,k) and transient signal T(n,k) (e.g. using weigh computation units 1520).
  • the transient signal envelopes T ⁇ ( n , k ) can, for example, then be estimated by taking the signal model in Equation (1) into account.
  • T ⁇ n k X n k ⁇ S ⁇ n k
  • This gaining rule is based on a spectral subtraction and Wiener filtering by choosing ⁇ and ⁇ accordingly.
  • Fig. 16 shows a schematic view of a weight calculation unit according to embodiments of the invention.
  • weight calculation unit 1600 is shown for transient suppression and/or transient enhancement, and may hence be an example of weight computation units 1520.
  • Weight calculation unit 1600 comprises a sustained signal estimator 1610, which may be configured to receive an input signal 1601, in order to determine, e.g. as explained above, based on a filtering, a sustained signal component 1611 (e.g. an estimate thereof) of the input signal.
  • a sustained signal estimator 1610 which may be configured to receive an input signal 1601, in order to determine, e.g. as explained above, based on a filtering, a sustained signal component 1611 (e.g. an estimate thereof) of the input signal.
  • a transient signal component 1621 of the input signal (e.g. an estimate thereof) may be determined.
  • Input signal 1601, sustained signal component 1611 and transient signal component 1621 are provided to the weight calculator 1630 in order to determine the weights for the transient enhancement or suppression, e.g. according to eqn. (4).
  • Fig. 17 shows a schematic block diagram of a method for obtaining a processed audio signal on the basis of an input audio signal according to embodiments of the invention.
  • Method (1700) comprises applying (1710) a transient enhancement, which emphasizes transients in the input audio signal, or in a processed version of the input audio signal, relative to non-transient audio signal components of the input audio signal or of a processed version of the input audio signal, applying (1720) a stage width enhancement, which increases inter-channel-level differences between corresponding audio signal components of different channels of the input audio signal, or between corresponding audio signal components of different channels of a processed version of the input audio signal, and/or which emphasizes audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to one of the sides of an audio scene relative to audio signal components of the input audio signal, or of a processed version of the input audio signal, panned to a center of the audio scene, and/or which selectively extracts audio signal components of the input audio
  • a first further embodiment comprises an apparatus for processing an audio signal, e.g. an ambience enhancement block, wherein the apparatus is configured to obtain, on the basis of an input audio signal, e.g. a multi-channel input audio signal, a first modified audio signal, e.g. a first modified multi-channel audio signal, using a transient suppression (e.g. such that transients that are included in the input audio signal are suppressed in the first modified audio signal, e.g. such that both direct signal components, having a comparatively higher inter-channel correlation, and ambience signal components, having a comparatively lower inter-channel correlation, are included in the first modified audio signal, e.g. independent from an inter-channel correlation of signal components of the input audio signal).
  • a transient suppression e.g. such that transients that are included in the input audio signal are suppressed in the first modified audio signal, e.g. such that both direct signal components, having a comparatively higher inter-channel correlation, and ambience signal components, having a comparatively lower inter-channel
  • the apparatus is further configured to obtain, on the basis of the input audio signal, a second modified audio signal using an ambience extraction (which, for example, prefers audio signal components having a comparatively smaller inter-channel-correlation over audio signal components having a comparatively larger inter-channel correlation; and/or which, for example, provides the second modified audio signal in dependence on inter-channel-correlation characteristics of audio signal components of the input audio signal).
  • an ambience extraction which, for example, prefers audio signal components having a comparatively smaller inter-channel-correlation over audio signal components having a comparatively larger inter-channel correlation
  • the apparatus is further configured to combine, e.g. using a weighted combination, the first modified audio signal, or a post-processed version of the first modified audio signal, and the second modified audio signal, or a post-processed version of the second modified audio signal, in order to obtain a combined audio signal; and the apparatus is configured to apply a decorrelation to the combined audio signal, in order to obtain a processed audio signal, e.g. an ambience-enhanced signal.
  • a processed audio signal e.g. an ambience-enhanced signal.
  • a second further embodiment comprises the apparatus according to the first further embodiment, wherein the apparatus, or for example the ambience extraction, is configured to extract from the input audio signal ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation (e.g. left-right correlation; e.g. an inter-channel correlation which is below a predetermined threshold value, or an inter-channel correlation which is smaller than an inter-channel correlation of direct signal components; e.g. such that direct signal components (and also non-transient direct signal components) are at least partially suppressed in the second modified audio signal, while ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation are non-suppressed or enhanced in the second modified audio signal).
  • a comparatively low inter-channel correlation e.g. left-right correlation; e.g. an inter-channel correlation which is below a predetermined threshold value, or an inter-channel correlation which is smaller than an inter-channel correlation of direct signal components; e.g. such that direct signal
  • the apparatus or for example the ambience extraction, is configured to emphasize ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation over, e.g. relative to, direct signal components and/or signal components having a comparatively high, or for example higher, inter-channel correlation.
  • the apparatus or for example the ambience extraction, is configured to attenuate direct signal components and/or signal components having a comparatively high inter-channel correlation relative to ambience signal components and/or diffuse signal components and/or signal components having a comparatively low inter-channel correlation.
  • a third further embodiment comprises the apparatus according to the first or second further embodiment, wherein the apparatus is configured to compute a covariance matrix on the basis of a frequency domain representation of the input audio signal, e.g. on the basis of a stereo input signal in a frequency domain, e.g. on the basis of X(f,k) or X(f,m) or X(n,k) (which may be used herein, in general optionally interchangeably). Furthermore the apparatus according to the third further embodiment is configured to compute a multichannel parametric Wiener filter on the basis of the covariance matrix, to estimate a direct signal using the multichannel parametric Wiener filter, and to subtract the estimated direct signal from the input audio signal, in order to obtain the second modified audio signal, e.g.
  • a fourth further embodiment comprises the apparatus according to one of the first to third further embodiments, wherein the apparatus is configured to apply a transient suppression to the second modified audio signal, in order to obtain a post-processed version of the second modified audio signal for combination with the first modified audio signal.
  • a sixth further embodiment comprises the apparatus according to one of the first to fourth further embodiments, wherein the apparatus is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate, of a, non-transient, sustained signal component, e.g. S(f,m), (e.g. an estimate of a temporal evolution of the sustained signal component, e.g. an estimate of an energy or an intensity of the sustained signal component) on the basis of one or more subband signals (e.g.
  • an estimate e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate
  • a, non-transient, sustained signal component e.g. S(f,m)
  • an estimate of a temporal evolution of the sustained signal component e.g. an estimate of an energy or an intensity of the sustained signal component
  • a seventh further embodiment comprises the apparatus according to one of the first to sixth further embodiments, wherein the apparatus is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency-band wise estimate, of a transient signal component (e.g. T ⁇ (f,m) or T ⁇ (n,k); e.g. an estimate of a temporal evolution of the transient signal component; e.g. an estimate of an energy or an intensity of the transient signal component) on the basis of one or more subband signals (e.g.
  • a transient signal component e.g. T ⁇ (f,m) or T ⁇ (n,k
  • e.g. an estimate of a temporal evolution of the transient signal component e.g. an estimate of an energy or an intensity of the transient signal component
  • the apparatus is configured to obtain the first modified audio signal in dependence on the estimate of the transient signal component.
  • An eighth further embodiment comprises the apparatus according to one of the first to seventh further embodiments, wherein the apparatus is configured to combine a plurality of frequency bins of a time-frequency-domain representation of the input audio signal, in order to obtain frequency band values associated with a plurality of frequency bands; (wherein the frequency band values are optionally examples of the subband values), and wherein the apparatus is configured to obtain the estimate of the sustained signal component and/or the estimate of the transient signal component in dependence on the frequency band values.
  • a ninth further embodiment comprises the apparatus according to one of the first to eighth further embodiments, wherein the apparatus is configured to obtain spectral weights, e.g. G g (f,m) or G g (n,k), associated with frequency bins, e.g. generally speaking with frequency subbands, on the basis of the frequency band values (e.g. using frequency-band-wise estimates of sustained signal components and/or of transient signal components derived from the frequency band values).
  • spectral weights e.g. G g (f,m) or G g (n,k
  • a tenth further embodiment comprises the apparatus according to one of the first to ninth further embodiments, wherein the apparatus is configured to obtain temporally smoothened (e.g. low-pass filtered subband signals and/or subband signals purged from peaks) subband signals or temporally smoothened frequency-band signals (e.g. using a combination of frequency bins (frequency subbands) to frequency bands) or temporally smoothened subband signal envelopes or temporally smoothened frequency band signal envelopes (e.g. low-pass filtered subband signal envelopes and/or subband signal envelopes purged from peaks; e.g. low-pass filtered frequency-band signal envelopes or frequency band signal envelopes purged from peaks) on the basis of subband signals representing the input audio signal.
  • temporally smoothened e.g. low-pass filtered subband signals and/or subband signals purged from peaks
  • subband signals or temporally smoothened frequency-band signals e.g. using a combination of frequency bins (frequency subband
  • frequency-band signals may be a special example of subband signals, wherein the term "subband signals" comprises both signals associated with a single frequency bins and signals associated with a frequency band comprising a plurality of frequency bins.
  • the apparatus according to the tenth further embodiments is configured to obtain the first modified audio signal in dependence on the temporally smoothened (e.g. low-pass filtered and/or purged from peaks) subband signals or the temporally smoothened frequency band signals or the temporally smoothened subband signal envelopes or the temporally smoothened frequency band signal envelopes.
  • temporally smoothened e.g. low-pass filtered and/or purged from peaks
  • An eleventh further embodiment comprises the apparatus according to one of the first to tenth further embodiments, wherein the apparatus is configured to obtain sustained signal envelopes, e.g. ⁇ (f,m) or ⁇ (n,k), using a temporal smoothing, e.g. low pass filtering, of subband time trajectories or frequency band time trajectories (wherein frequency band time trajectories are a special case of subband time trajectories, wherein the term "subband time trajectories" comprises time trajectories comprising only a single frequency bin and time trajectories comprising a plurality of frequency bins) of the input audio signal (e.g.
  • respective sustained signal envelopes e.g. ⁇ (f,m) or S ⁇ (n,k)
  • respective sustained signal envelopes e.g. ⁇ (f,m) or S ⁇ (n,k)
  • respective sustained signal envelopes e.g. ⁇ (f,m) or S ⁇ (n,k)
  • the subband time trajectories may, for example, represent a temporal evolution of energies in respective subbands of the input audio signal, wherein subbands may, for example, be frequency bands of a filterbank or may, for example, be frequency bins of a time-domain-to-spectral-domain transform or of a time-domain-to-frequency-domain transform.
  • a twelfth further embodiment comprises the apparatus according to one of the first to eleventh further embodiments, wherein the apparatus is configured to obtain transient signal envelopes, e.g. T ⁇ (f,m) or T ⁇ (n,k), using a subtraction of sustained signal envelopes, e.g. ⁇ (f,m) or ⁇ (n,k), from respective subband signals, e.g. X(f,m), or from magnitudes or respective subband signals or from frequency band signals of the input audio signal or from magnitudes of frequency band signals of the input audio signal or from or the subband time trajectories, e.g. of magnitudes, of the input audio signal or from frequency band time trajectories, e.g. of magnitudes, of the input audio signal.
  • transient signal envelopes e.g. T ⁇ (f,m) or T ⁇ (n,k)
  • respective subband signals e.g. X(f,m)
  • a thirteenth further embodiment comprises the apparatus according to one of the first to twelfth further embodiments, wherein the apparatus is configured to obtain spectral weights, e.g. G g (f,m)or G g (n,k), in dependence on sustained signal envelope values (e.g. in dependence on the sustained signal envelope values mentioned before, e.g. ⁇ (n,k)).
  • spectral weights e.g. G g (f,m)or G g (n,k)
  • sustained signal envelope values e.g. in dependence on the sustained signal envelope values mentioned before, e.g. ⁇ (n,k)
  • a thirteenth further embodiment comprises the apparatus according to one of the first to twelfth further embodiments, wherein the apparatus is configured to obtain spectral weights, e.g. G g (f,m) or G g (n,k), in dependence on transient signal envelope values (e.g. in dependence on the transient signal envelope values mentioned before, e.g. T ⁇ (n,k)).
  • spectral weights e.g. G g (f,m) or G g (n,k)
  • transient signal envelope values e.g. in dependence on the transient signal envelope values mentioned before, e.g. T ⁇ (n,k)
  • a fourteenth further embodiment comprises the apparatus according to one of the first to thirteenth further embodiments, wherein the apparatus is configured to obtain respective spectral weights, e.g. G g (f,m), in dependence on a ratio between a weighted sum of respective sustained signal envelope values, e.g. ⁇ (f,m) or ⁇ (n,k), or an exponentiated version thereof, e.g. ⁇ (f,m) alpha or ⁇ (n,k) alpha , and of respective transient signal envelope values, e.g. T ⁇ (f,m) or T ⁇ (n,k), or an exponentiated version thereof, e.g.
  • T ⁇ (f,m) alpha or T ⁇ (n,k) alpha and magnitudes of respective subband signals, e.g. X(f,m) or X(n,k), or frequency band signals, or an exponentiated version thereof, e.g.
  • a sixteenth further embodiment comprises the apparatus according to one of the first to fifteenth further embodiments, wherein the apparatus is configured to obtain the sustained signal envelopes ⁇ (n,k) using a low pass filtering of subband time trajectories of magnitudes of a time-frequency-domain representation, e.g.
  • a first additional embodiment comprises an transient signal processor, e.g. a transient enhancement or a transient suppression, wherein the transient signal processor is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency band wise estimate, of a, for example non-transient, sustained signal component, e.g. S(f,m), (e.g. an estimate of a temporal evolution of the sustained signal component, e.g. an estimate of an energy or an intensity of the sustained signal component) on the basis of one or more subband signals (e.g.
  • the transient signal processor is configured to obtain the modified audio signal, in which transient signal components are enhanced or reduced, in dependence on the estimate of the sustained signal component.
  • a second additional embodiment comprises the transient signal processor according to one of the first or second additional embodiments, wherein the transient signal processor is configured to obtain an estimate, e.g. a time-frequency-bin-wise estimate or a frequency-band wise estimate, of a transient signal component, e.g. T ⁇ (f,m) or T ⁇ (n,k), e.g. an estimate of a temporal evolution of the transient signal component, e.g. an estimate of an energy or an intensity of the transient signal component, on the basis of one or more subband signals, (e.g.
  • the transient signal processor is configured to obtain the modified audio signal in dependence on the estimate of the transient signal component.
  • inter-channel level difference modifier comprises an inter-channel level difference modifier, wherein the inter-channel level difference modifier is configured to determine inter-channel-level differences (e.g. inter-channel level difference values Dx(n,k); wherein is should be noted that, in some embodiments, indices (f,m) may be used instead of indices (n,k), wherein f and n are frequency indices, and wherein indices m and k are time indices) for a plurality of corresponding time-frequency bins of an input audio signal (e.g. between corresponding time-frequency bins of different channels of the input audio signal).
  • the inter-channel level difference modifier is configured to scale spectral bin values, e.g.
  • inter-channel level difference modifiers is configured to map (e.g. directly, e.g. without a computation of a panning index or the like) a value representing the inter-channel level difference, e.g. D x (n,k), (e.g. in a logarithmic domain, describing a ratio between energies in corresponding time-frequency bins of two channels in a logarithmized form) onto one or more values representing a change of the inter-channel level difference, e.g. in a logarithmic domain, using a linear or piecewise-linear mapping.
  • a value representing the inter-channel level difference e.g. D x (n,k)
  • D x (n,k) e.g. in a logarithmic domain, describing a ratio between energies in corresponding time-frequency bins of two channels in a logarithmized form
  • This may, for example, be performed such that the value representing the inter-channel level difference is proportional to the value representing the change of the inter-channel level difference at least for inter-channel level differences having a first sign.
  • the above may be performed using a piecewise linear mapping having a first proportionality factor between a value representing the inter-channel level difference and a (e.g. first) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is larger than an energy in a second channel, and having a second proportionality factor between the value representing the inter-channel level difference and a (e.g. second) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is smaller than an energy in a second channel.
  • a piecewise linear mapping having a first proportionality factor between a value representing the inter-channel level difference and a (e.g. first) value representing the change of the inter-channel level difference for a case in which an energy in a first channel is larger than an energy
  • Another embodiment according to the invention comprises a consumer-side audio signal processor, for processing an input audio signal for a playback to a listener, wherein the audio signal processor comprises a transient enhancement configured to emphasize transients in an input audio signal, in order to obtain a processed audio signal for playback to the listener.
  • the audio signal processor comprises a transient enhancement configured to emphasize transients in an input audio signal, in order to obtain a processed audio signal for playback to the listener.
  • aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus.
  • Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus.
  • embodiments of the invention can be implemented in hardware or in software.
  • the implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
  • Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
  • embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer.
  • the program code may for example be stored on a machine readable carrier.
  • inventions comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
  • an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
  • a further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein.
  • the data carrier, the digital storage medium or the recorded medium are typically tangible and/or non-transitionary.
  • a further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein.
  • the data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.
  • a further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
  • a processing means for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
  • a further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
  • a further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver.
  • the receiver may, for example, be a computer, a mobile device, a memory device or the like.
  • the apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
  • a programmable logic device for example a field programmable gate array
  • a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein.
  • the methods are preferably performed by any hardware apparatus.
  • the apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
  • the apparatus described herein, or any components of the apparatus described herein, may be implemented at least partially in hardware and/or in software.

Landscapes

  • Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Quality & Reliability (AREA)
  • Signal Processing (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Stereophonic System (AREA)
  • Tone Control, Compression And Expansion, Limiting Amplitude (AREA)
EP24157976.2A 2024-02-15 2024-02-15 Appareil, procédé et programme informatique pour traitement de signal audio sur la base d'une différence de niveau entre canaux et d'une manipulation de composante de signal latérale Pending EP4604120A1 (fr)

Priority Applications (2)

Application Number Priority Date Filing Date Title
EP24157976.2A EP4604120A1 (fr) 2024-02-15 2024-02-15 Appareil, procédé et programme informatique pour traitement de signal audio sur la base d'une différence de niveau entre canaux et d'une manipulation de composante de signal latérale
PCT/EP2025/054116 WO2025172586A1 (fr) 2024-02-15 2025-02-14 Appareil, procédé et programme informatique de traitement de signaux audio basé sur la différence de niveau entre les canaux et la manipulation de composant du signal latéral

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
EP24157976.2A EP4604120A1 (fr) 2024-02-15 2024-02-15 Appareil, procédé et programme informatique pour traitement de signal audio sur la base d'une différence de niveau entre canaux et d'une manipulation de composante de signal latérale

Publications (1)

Publication Number Publication Date
EP4604120A1 true EP4604120A1 (fr) 2025-08-20

Family

ID=89977507

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24157976.2A Pending EP4604120A1 (fr) 2024-02-15 2024-02-15 Appareil, procédé et programme informatique pour traitement de signal audio sur la base d'une différence de niveau entre canaux et d'une manipulation de composante de signal latérale

Country Status (2)

Country Link
EP (1) EP4604120A1 (fr)
WO (1) WO2025172586A1 (fr)

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1565036A2 (fr) * 2004-02-12 2005-08-17 Agere System Inc. Synthèse de scènes audio basée sur réverbérations retardées
US20140369504A1 (en) * 2013-06-12 2014-12-18 Anthony Bongiovi System and method for stereo field enhancement in two-channel audio systems
US20170055096A1 (en) * 2015-08-21 2017-02-23 Sonos, Inc. Manipulation of Playback Device Response Using Signal Processing

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1565036A2 (fr) * 2004-02-12 2005-08-17 Agere System Inc. Synthèse de scènes audio basée sur réverbérations retardées
US20140369504A1 (en) * 2013-06-12 2014-12-18 Anthony Bongiovi System and method for stereo field enhancement in two-channel audio systems
US20170055096A1 (en) * 2015-08-21 2017-02-23 Sonos, Inc. Manipulation of Playback Device Response Using Signal Processing

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
LEE HYUNKOOK ET AL: "Level and Time Panning of Phantom Images for Musical Sources", JAES, AES, 60 EAST 42ND STREET, ROOM 2520 NEW YORK 10165-2520, USA, vol. 61, no. 12, 20 December 2013 (2013-12-20), pages 978 - 988, XP040636949 *

Also Published As

Publication number Publication date
WO2025172586A1 (fr) 2025-08-21

Similar Documents

Publication Publication Date Title
JP6637014B2 (ja) 音声信号処理のためのマルチチャネル直接・環境分解のための装置及び方法
US8588427B2 (en) Apparatus and method for extracting an ambient signal in an apparatus and method for obtaining weighting coefficients for extracting an ambient signal and computer program
US10242692B2 (en) Audio coherence enhancement by controlling time variant weighting factors for decorrelated signals
JP5149968B2 (ja) スピーチ信号処理を含むマルチチャンネル信号を生成するための装置および方法
EP2730102B1 (fr) Procédé et appareil pour décomposer un enregistrement stéréo à l'aide d'un traitement dans le domaine fréquentiel employant un générateur de poids spectraux
CN105284133B (zh) 基于信号下混比进行中心信号缩放和立体声增强的设备和方法
JP2016048927A (ja) 少なくとも2つの出力チャネルを有する出力信号を生成するための装置および方法
WO2025171880A1 (fr) Système de traitement audio, procédé et programme informatique pour l'amélioration des transitoires, l'amélioration de la largeur de la scène et l'amélioration de l'ambiance basées sur un traitement de signal audio
WO2025172586A1 (fr) Appareil, procédé et programme informatique de traitement de signaux audio basé sur la différence de niveau entre les canaux et la manipulation de composant du signal latéral
HK1197782A (en) Method and apparatus for decomposing a stereo recording using frequency-domain processing employing a spectral subtractor
HK1197782B (en) Method and apparatus for decomposing a stereo recording using frequency-domain processing employing a spectral subtractor
HK1197959B (en) Method and apparatus for decomposing a stereo recording using frequency-domain processing employing a spectral weights generator
HK1237528B (en) Apparatus and method for enhancing an audio signal, sound enhancing system
HK1237528A1 (en) Apparatus and method for enhancing an audio signal, sound enhancing system

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION HAS BEEN PUBLISHED

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR