WO2021148009A1 - 音频处理方法及电子设备 - Google Patents
音频处理方法及电子设备 Download PDFInfo
- Publication number
- WO2021148009A1 WO2021148009A1 PCT/CN2021/073380 CN2021073380W WO2021148009A1 WO 2021148009 A1 WO2021148009 A1 WO 2021148009A1 CN 2021073380 W CN2021073380 W CN 2021073380W WO 2021148009 A1 WO2021148009 A1 WO 2021148009A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- parameter value
- intensity parameter
- reverberation
- reverberation intensity
- frequency domain
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/36—Accompaniment arrangements
- G10H1/361—Recording/reproducing of accompaniment for use with an external source, e.g. karaoke systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/36—Accompaniment arrangements
- G10H1/361—Recording/reproducing of accompaniment for use with an external source, e.g. karaoke systems
- G10H1/366—Recording/reproducing of accompaniment for use with an external source, e.g. karaoke systems with means for modifying or correcting the external signal, e.g. pitch correction, reverberation, changing a singer's voice
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H1/00—Details of electrophonic musical instruments
- G10H1/0008—Associated control or indicating means
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2210/00—Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
- G10H2210/005—Musical accompaniment, i.e. complete instrumental rhythm synthesis added to a performed melody, e.g. as output by drum machines
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2210/00—Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
- G10H2210/031—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal
- G10H2210/076—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal for extraction of timing, tempo; Beat detection
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2210/00—Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
- G10H2210/031—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal
- G10H2210/091—Musical analysis, i.e. isolation, extraction or identification of musical elements or musical parameters from a raw acoustic signal or from an encoded audio signal for performance evaluation, i.e. judging, grading or scoring the musical qualities or faithfulness of a performance, e.g. with respect to pitch, tempo or other timings of a reference performance
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10H—ELECTROPHONIC MUSICAL INSTRUMENTS; INSTRUMENTS IN WHICH THE TONES ARE GENERATED BY ELECTROMECHANICAL MEANS OR ELECTRONIC GENERATORS, OR IN WHICH THE TONES ARE SYNTHESISED FROM A DATA STORE
- G10H2210/00—Aspects or methods of musical processing having intrinsic musical character, i.e. involving musical theory or musical parameters or relying on musical knowledge, as applied in electrophonic musical tools or instruments
- G10H2210/155—Musical effects
- G10H2210/265—Acoustic effect simulation, i.e. volume, spatial, resonance or reverberation effects added to a musical sound, usually by appropriate filtering or delays
- G10H2210/281—Reverberation or echo
Definitions
- the present disclosure relates to the technical field of signal processing, and in particular to an audio processing method and electronic equipment.
- the K song sound effect refers to the audio processing of the collected human voice and background music, making the processed human voice more pleasant than the one before processing, and at the same time, it can mask the pitch inaccuracy of a part of the human voice. And other issues.
- the present disclosure provides an audio processing method and electronic device, which can make the sound output by the electronic device more full and beautiful.
- the technical solutions of the present disclosure are as follows:
- an audio processing method including:
- Target reverberation intensity parameter value of the collected accompaniment audio signal, where the target reverberation intensity parameter value is used to indicate at least one of the rhythm speed, the accompaniment type, and the singer's singing score of the current music to be processed;
- the determining the target reverberation intensity parameter value of the collected accompaniment audio signal includes:
- the target reverberation intensity parameter value is determined based on the first reverberation intensity parameter value, the second type reverberation intensity parameter value, and the third type reverberation intensity parameter value.
- the determining the first reverberation intensity parameter value of the collected accompaniment audio signal includes:
- the frequency domain richness coefficient is used to indicate the frequency domain richness of the amplitude information of each frame of accompaniment audio, and the frequency domain richness reflects the accompaniment type of the current music to be processed;
- the first reverberation intensity parameter value is determined based on the frequency domain rich coefficient of each frame of accompaniment audio.
- the determining the first reverberation intensity parameter value based on the frequency domain rich coefficient of each frame of accompaniment audio includes:
- a first ratio between the global frequency domain richness coefficient and the maximum frequency domain richness coefficient is obtained, and the smallest of the first ratio and the target value is determined as the first reverberation intensity parameter value.
- the determining the first reverberation intensity parameter value based on the frequency domain rich coefficient of each frame of accompaniment audio includes:
- the smallest of the second ratio and the target value is determined as the first reverberation intensity parameter value.
- the determining the second reverberation intensity parameter value of the collected accompaniment audio signal includes:
- the smallest of the third ratio and the target value is determined as the second reverberation intensity parameter value.
- the determining the third reverberation intensity parameter value of the collected accompaniment audio signal includes:
- the target reverberation intensity parameter value is determined based on the first reverberation intensity parameter value, the second type reverberation intensity parameter value, and the third type reverberation intensity parameter value ,include:
- the performing reverberation processing on the collected human voice signal based on the target reverberation intensity parameter value includes:
- the method further includes:
- an audio processing device including:
- the collection module is configured to collect the accompaniment audio signal and the vocal signal of the current music to be processed
- the determining module is configured to determine the target reverberation intensity parameter value of the collected accompaniment audio signal, and the target reverberation intensity parameter value is used to indicate the rhythm speed, accompaniment type, and the singer’s singing score of the current music to be processed. At least one
- the processing module is configured to perform reverberation processing on the collected human voice signal based on the target reverberation intensity parameter value.
- the determining module is further configured to determine a first reverberation intensity parameter value of the collected accompaniment audio signal, and the first reverberation intensity parameter value is used to indicate the accompaniment type of the currently to-be-processed music ; Determine the second reverberation intensity parameter value of the collected accompaniment audio signal, the second reverberation intensity parameter value is used to indicate the rhythm speed of the current music to be processed; determine the third reverberation intensity of the collected accompaniment audio signal Parameter value, the third reverberation intensity parameter value is used to indicate the singing score of the singer of the currently to-be-processed music; based on the first reverberation intensity parameter value, the second-type reverberation intensity parameter value, and the The third type of reverberation intensity parameter value determines the target reverberation intensity parameter value.
- the determining module is further configured to transform the collected accompaniment audio signal from the time domain to the time-frequency domain to obtain an accompaniment audio frame sequence; obtain the amplitude information of each frame of accompaniment audio; based on each frame of accompaniment The amplitude information of the audio determines the frequency domain richness coefficient of each frame of accompaniment audio; wherein, the frequency domain richness coefficient is used to indicate the frequency domain richness of the amplitude information of each frame of accompaniment audio, and the frequency domain richness reflects the current waiting Processing the accompaniment type of the music; determining the first reverberation intensity parameter value based on the frequency domain rich coefficient of each frame of accompaniment audio.
- the determining module is further configured to determine the global frequency domain rich coefficients of the current music piece to be processed based on the frequency domain rich coefficients of each frame of accompaniment audio; obtain the global frequency domain rich coefficients and frequency domain rich coefficients The first ratio between the maximum value of the coefficient, the smallest of the first ratio and the target value is determined as the first reverberation intensity parameter value.
- the determining module is further configured to generate a waveform graph indicating the richness of the frequency domain based on the frequency domain rich coefficient of each frame of accompaniment audio; perform smoothing processing on the generated waveform graph based on the smoothed Determine the frequency domain rich coefficients of different parts of the current music to be processed; obtain the second ratio between the frequency domain rich coefficients of the different parts and the maximum value of the frequency domain rich coefficient; for each second The ratio, the smallest of the second ratio and the target value is determined as the first reverberation intensity parameter value.
- the determining module is further configured to obtain the number of beats of the collected accompaniment audio signal in a prescribed duration; determine the third ratio between the obtained number of beats and the maximum value of the number of beats; The smallest of the third ratio and the target value is determined as the second reverberation intensity parameter value.
- the determining module is further configured to obtain the audio singing score of the singer of the current music to be processed, and determine the third reverberation intensity parameter value based on the audio singing score.
- the determining module is further configured to obtain a basic reverberation intensity parameter value, a first weight value, a second weight value, and a third weight value; determine the first weight value and the first weight value.
- the first sum value between the reverberation intensity parameter values determine the second sum value between the second weight value and the second reverberation intensity parameter value; determine the third weight value and the third The third sum value between the reverberation intensity parameter values; obtaining the basic reverberation intensity parameter value, the first sum value, and the fourth sum value between the second sum value and the third sum value , Determining the smallest of the fourth ratio and the target value as the target reverberation intensity parameter value.
- the processing module is further configured to adjust the total reverberation gain of the collected human voice signal based on the target reverberation strength parameter value; or, based on the target reverberation strength parameter value Value to adjust at least one reverberation algorithm parameter of the collected human voice signal.
- the processing module is further configured to perform reverberation processing on the collected vocal signals, and then perform mixing processing on the collected accompaniment audio signals and the vocal signals after the reverberation processing. , Output the audio signal after mixing processing.
- an electronic device including:
- a memory for storing executable instructions of the processor
- the processor is configured to execute the instructions to implement the audio processing method described above.
- a storage medium is provided, and instructions in the storage medium are executed by a processor of an electronic device, so that the electronic device can execute the aforementioned audio processing method.
- a computer program product is provided, and instructions in the computer program product are executed by a processor of an electronic device, so that the electronic device can execute the audio processing method as described above.
- Fig. 1 is a schematic diagram showing an implementation environment involved in an audio processing method according to an embodiment.
- Fig. 2 is a flowchart of an audio processing method according to an embodiment.
- Fig. 3 is a flowchart showing an audio processing method according to an embodiment.
- Fig. 4 is an overall system block diagram showing an audio processing method according to an embodiment.
- Fig. 5 is a flowchart showing an audio processing method according to an embodiment.
- Fig. 6 is a waveform diagram showing the richness of the frequency domain according to an embodiment.
- Fig. 7 is a smoothed waveform diagram with respect to the richness of the frequency domain according to an embodiment.
- Fig. 8 is a block diagram showing an audio processing device according to an embodiment.
- Fig. 9 is a block diagram showing an electronic device according to an embodiment.
- Fig. 10 is a block diagram showing another electronic device according to an embodiment.
- the user information involved in this disclosure is information authorized by the user or fully authorized by all parties.
- at least one of A, B, and C includes the following situations: A alone, B alone, C alone, A and B, A and C, B and C, and A, B, and C.
- K song sound effect refers to the audio processing of the collected vocals and background music, making the processed vocals more pleasing than the vocals before processing, and at the same time, it can mask the pitch inaccuracy of a part of the vocals, etc. problem.
- K song sound effects are used to modify the collected human voice.
- BGM Background Music, accompaniment music or background music
- accompaniment music soundtrack for short.
- BGM usually refers to a kind of music used to adjust the atmosphere in TV dramas, movies, animations, video games, and websites. It can be inserted into the dialogue to enhance the expression of emotions and achieve an immersive experience for the audience. Feelings.
- the music played in some public places is also called background music.
- BGM refers to song accompaniment.
- STFT Short-Time Fourier Transform, Short-Time Fourier Transform
- STFT Short-Time Fourier Transform
- STFT Short-Time Fourier Transform
- It is a mathematical transformation related to the Fourier Transform, used to determine the frequency and phase of the sine wave in a local area of a time-varying signal. That is, the long non-stationary signal is regarded as the superposition of a series of short-term stationary signals, and the short-term stationary signal is realized by a windowing function, that is, multiple segments of signals are intercepted and Fourier transformed respectively. Its time-frequency analysis characteristics are shown in: expressing the characteristics of a certain moment through a period of signal in the time window.
- Reverberation When sound waves propagate indoors, they will be reflected by obstacles such as walls, ceilings, or floors, and each reflection will be absorbed by the obstacles. In this way, when the sound source stops sounding, the sound wave has to undergo multiple reflections and absorptions in the room before it disappears. The human ear will feel that there are several sound waves mixed for a period of time after the sound source stops sounding, that is, the sound source stops sounding. The phenomenon of sound continuity still exists after that, and this phenomenon is called reverberation.
- reverberation is mainly used to sing karaoke, increase the delay of the microphone sound, generate an appropriate amount of echo, make the singing sound more round and beautiful, and the singing voice is not so dry. That is, for the singing of karaoke, in order to make the sound less dry and weak, the reverberation is usually added artificially in the later stage to make the sound more full and beautiful.
- the implementation environment includes: an electronic device 101 for audio processing.
- the electronic device 101 is a terminal or a server, which is not specifically limited in the embodiment of the present application.
- the types of terminals include but are not limited to: mobile terminals and fixed terminals.
- mobile terminals include, but are not limited to: smart phones, tablet computers, notebook computers, e-readers, MP3 players (Moving Picture Experts Group Audio Layer III, Motion Picture Experts compress standard audio layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, Motion Picture Experts Compresses Standard Audio Layer 4) Players, etc.; stationary terminals include, but are not limited to, desktop computers, which are not specifically limited in the embodiments of this application.
- a music application program with audio processing functions is usually installed on the terminal to execute the audio processing method provided in the embodiments of the present application.
- the terminal can also upload the audio signal to be processed to the server through a music application or a video application, and the server executes the audio processing method provided by the embodiment of the application, and The result is returned to the terminal, which is not specifically limited in the embodiment of the present application.
- the electronic device 101 in order to make the sound more full and beautiful, the electronic device 101 usually performs artificial reverberation processing on the collected human voice signal.
- the BGM audio signal is transformed from the time domain to the time-frequency domain through the short-time Fourier transform to obtain a BGM audio signal
- the amplitude information of each frame of accompaniment audio is obtained, and the frequency domain richness of the amplitude information of each frame of accompaniment audio is calculated based on this; in addition, the BGM audio signal can also be obtained for the specified duration (for example, every minute) Based on the number of beats, the rhythm speed of the BGM audio signal is calculated.
- the most suitable reverberation intensity parameter values can be dynamically or pre-calculated, and then the artificial reverberation algorithm can be guided to control the output person.
- the size of the reverberation of the sound part so as to achieve an adaptive K song sound effect.
- the embodiments of the present disclosure comprehensively consider the frequency domain richness of the song, the rhythm speed, the singer and other factors, and accordingly generate different reverberation intensity parameter values adaptively, thereby achieving an adaptive K song sound effects.
- Fig. 2 is a flowchart of an audio processing method according to an embodiment. As shown in Fig. 2, the audio processing method is used in an electronic device and includes the following steps.
- the accompaniment audio signal and the human voice signal of the currently to-be-processed music are collected.
- a target reverberation intensity parameter value of the collected accompaniment audio signal is determined, and the target reverberation intensity parameter value is used to indicate at least one of the rhythm speed, accompaniment type, and singer's singing score of the current music to be processed.
- reverberation processing is performed on the collected human voice signal based on the target reverberation intensity parameter value.
- the embodiment of the present disclosure will determine the target reverberation intensity parameter value of the collected accompaniment audio signal, where the target reverberation intensity
- the parameter value is used to indicate at least one of the rhythm speed, accompaniment type, and singer's singing score of the currently to-be-processed music; after that, the collected human voice signal is subjected to reverberation processing based on the target reverberation intensity parameter value.
- the embodiments of the present disclosure consider various factors such as the accompaniment type of the music, the rhythm speed, and the singer’s singing score, and accordingly, adaptively generate the parameter value of the reverberation intensity of the current music to be processed.
- the self-adaptive K song sound effect makes the sound output by the electronic device more full and beautiful.
- the determining the target reverberation intensity parameter value of the collected accompaniment audio signal includes:
- the target reverberation intensity parameter value is determined based on the first reverberation intensity parameter value, the second type reverberation intensity parameter value, and the third type reverberation intensity parameter value.
- the determining the first reverberation intensity parameter value of the collected accompaniment audio signal includes:
- the frequency domain richness coefficient is used to indicate the frequency domain richness of the amplitude information of each frame of accompaniment audio, and the frequency domain richness reflects the accompaniment type of the current music to be processed;
- the first reverberation intensity parameter value is determined based on the frequency domain rich coefficient of each frame of accompaniment audio.
- the determining the first reverberation intensity parameter value based on the frequency domain rich coefficient of each frame of accompaniment audio includes:
- a first ratio between the global frequency domain rich coefficient and the maximum frequency domain rich coefficient is obtained, and the smallest of the first ratio and the target value is determined as the first reverberation intensity parameter value.
- the determining the first reverberation intensity parameter value based on the frequency domain rich coefficient of each frame of accompaniment audio includes:
- the smallest of the second ratio and the target value is determined as the first reverberation intensity parameter value.
- the determining the second reverberation intensity parameter value of the collected accompaniment audio signal includes:
- the smallest of the third ratio and the target value is determined as the second reverberation intensity parameter value.
- the determining the third reverberation intensity parameter value of the collected accompaniment audio signal includes:
- the target reverberation intensity parameter value is determined based on the first reverberation intensity parameter value, the second type reverberation intensity parameter value, and the third type reverberation intensity parameter value ,include:
- the performing reverberation processing on the collected human voice signal based on the target reverberation intensity parameter value includes:
- the method further includes:
- Fig. 3 is a flowchart of an audio processing method according to an embodiment.
- the audio processing method is used in an electronic device.
- the audio processing method includes the following steps.
- the accompaniment audio signal and the vocal signal of the current music to be processed are collected.
- the currently to-be-processed music is the song currently being sung by the user.
- the accompaniment audio signal is also referred to herein as background music accompaniment or BGM audio signal.
- the electronic device collects the accompaniment audio signal and the human voice signal of the current music to be processed through its own configured or external microphone.
- a target reverberation intensity parameter value of the collected accompaniment audio signal is determined, where the target reverberation intensity parameter value is used to indicate at least one of the rhythm speed, accompaniment type, and singer's singing score of the current music to be processed. kind.
- reverb processing Under normal circumstances, a basic principle of reverb processing is: For songs with simple background music accompaniment (such as pure guitar accompaniment) and slow speed, small reverberation will be added to make the human voice more pure; for background music accompaniment components are diverse (Such as band song accompaniment), fast songs, will add a large reverberation, play a role in setting off the atmosphere and highlighting the human voice.
- simple background music accompaniment such as pure guitar accompaniment
- background music accompaniment components are diverse (Such as band song accompaniment)
- fast songs will add a large reverberation, play a role in setting off the atmosphere and highlighting the human voice.
- the target reverberation intensity parameter value is used to indicate at least one of the rhythm speed, accompaniment type, and singer’s singing score of the current music to be processed, including the following situations: the target reverberation intensity parameter value is used to indicate the current Process the rhythm and speed of the music; the target reverberation intensity parameter value is used to indicate the accompaniment type of the current music to be processed; the target reverberation intensity parameter value is used to indicate the singing score of the singer of the current music to be processed; the target reverberation intensity parameter value is used It is used to indicate the rhythm speed and accompaniment type of the current music to be processed; the target reverberation intensity parameter value is used to indicate the rhythm speed of the current music to be processed and the singer’s singing score; the target reverberation intensity parameter value is used to indicate the current music to be processed The type of accompaniment and the singer's singing score; the target reverb intensity parameter value is used to indicate the rhythm speed, accompaniment type, and singer's singing score of the current music to be processed.
- determining the target reverberation intensity parameter value of the collected accompaniment audio signal includes the following steps:
- the accompaniment type of the current music piece to be processed is characterized by the richness of the frequency domain.
- the richer the accompaniment of the song itself the higher the richness of the corresponding frequency domain; and vice versa.
- a song with a strong accompaniment has a higher frequency domain richness factor than a song with a simple accompaniment.
- the frequency domain richness coefficient is used to indicate the frequency domain richness of the amplitude information of each frame of accompaniment audio, that is, the frequency domain richness reflects the accompaniment type of the current music to be processed.
- determining the first reverberation intensity parameter value of the collected accompaniment audio signal includes but is not limited to the following steps:
- the embodiment of the present disclosure performs a short-time Fourier transform on the BCM audio signal of the currently to-be-processed music, so as to realize the transformation from the time domain to the time-frequency domain.
- an audio signal x of length T is x(t) in the time domain, where t represents time, and 0 ⁇ t ⁇ T.
- n refers to any frame in the obtained accompaniment audio frame sequence, 0 ⁇ n ⁇ N, N is the total number of frames, k refers to any frequency point in the center frequency sequence, 0 ⁇ k ⁇ K, K is Total frequency points.
- the frequency domain richness of each frame of accompaniment audio SpecRichness that is, the frequency domain richness coefficient is:
- Figure 6 shows the frequency domain richness of two songs. Because the accompaniment of song A is strong, and the accompaniment of song B is simpler than the former, the frequency domain richness of song A is higher than that of the former.
- Song B Figure 6 shows the original calculated SpecRichness for the two songs, and Figure 7 shows the smoothed SpecRichness. It can be seen from Figure 6 and Figure 7 that songs with strong accompaniment have higher SpecRichness than songs with simple accompaniment.
- the first reverberation intensity parameter value is determined based on the frequency domain rich coefficient of each frame of accompaniment audio.
- one implementation is to assign different degrees of reverberation to different songs through a pre-calculated global SpecRichness.
- the first reverberation intensity parameter value is determined based on the frequency domain rich coefficient of each frame of accompaniment audio, including but not limited to: based on the frequency domain rich coefficient of each frame of accompaniment audio, the global value of the current music to be processed is determined Frequency domain enrichment coefficient; obtain the first ratio between the global frequency domain enrichment coefficient and the maximum value of the frequency domain enrichment coefficient, and determine the smallest of the first ratio and the target value as the first reverberation intensity parameter value.
- the global frequency domain richness coefficient is the average value of the frequency domain richness coefficient of each frame of accompaniment audio, which is not specifically limited in the embodiment of the present disclosure.
- the target value is referred to as the value 1 herein.
- the formula for calculating the value of the first reverberation intensity parameter through the calculated SpecRichness is:
- G SpecRichness refers to the first reverberation intensity parameter value
- SpecRichness_max refers to the preset maximum allowable SpecRichness value
- another implementation manner is to assign different degrees of reverberation to different parts of each song through the smoothed SpecRichness. For example, the reverberation of the chorus will be stronger, as shown in the upper curve in Figure 7.
- the first reverberation intensity parameter value is determined based on the frequency domain rich coefficient of each frame of accompaniment audio, including but not limited to: based on the frequency domain rich coefficient of each frame of accompaniment audio, generated to indicate the frequency domain
- the richness waveform diagram is shown in Figure 7; the generated waveform diagram is smoothed, and the frequency domain rich coefficients of different parts of the current music to be processed are determined based on the smoothed waveform diagram; the frequency domain rich coefficients of different parts are obtained respectively
- the second ratio to the maximum value of the frequency domain rich coefficient; for each acquired second ratio, the smallest of the second ratio and the target value is determined as the first reverberation intensity parameter value.
- the frequency domain rich coefficients of different parts are the average values of the frequency domain rich coefficients of each frame of accompaniment audio of the corresponding parts, which are not specifically limited in the embodiment of the present disclosure.
- the above-mentioned different parts include at least the verse part and the chorus part.
- the rhythm speed of the current music piece to be processed is characterized by the number of beats. That is, in some embodiments, determining the second reverberation intensity parameter value of the collected accompaniment audio signal includes, but is not limited to: obtaining the number of beats of the collected accompaniment audio signal in a specified duration; determining the number of beats and The third ratio between the maximum number of beats; the smallest of the third ratio and the target value is determined as the second reverberation intensity parameter value.
- the number of beats within the prescribed time period is the number of beats per minute, which is not specifically limited in the embodiments of the present disclosure.
- BPM Beat Per Minute
- BPM the number of beats per minute, that is, the number of sound beats emitted between periods of one minute, and the unit of this number is BPM, also called the number of beats.
- the number of beats per minute of the current music to be processed is obtained through the beat number analysis algorithm.
- the calculation formula of the second reverberation intensity parameter value is:
- G bgm refers to the second reverberation intensity parameter value
- BGM refers to the calculated beats per minute
- BGM_max refers to the preset maximum allowable beats per minute.
- the third reverberation intensity parameter value of the collected accompaniment audio signal is determined, where the third reverberation intensity parameter value is used to indicate the singing score of the singer of the currently to-be-processed music.
- the embodiments of the present disclosure can also perform reverberation intensity control by extracting the singing score (audio singing score) of the singer of the currently to-be-processed music. That is, in some embodiments, determining the third reverberation strength parameter value of the collected accompaniment audio signal includes, but is not limited to: obtaining the audio singing score of the singer of the current music to be processed, and determining the first audio singing score based on the audio singing score. Three parameter values of reverberation intensity.
- the audio singing score refers to the historical song score or real-time song score of the singer, and the historical song score is the song score within the last month, the last 3 months, the last six months, or the last year.
- the implementation of the present disclosure The example does not specifically limit this.
- the full score of the song score is 100 points.
- the calculation formula of the third reverberation intensity parameter value is:
- G vocalGoodness refers to the third reverberation strength parameter value
- KTV_Score refers to the obtained audio singing score
- the target reverberation intensity parameter value is determined based on the first reverberation intensity parameter value, the second type reverberation intensity parameter value, and the third type reverberation intensity parameter value.
- the target reverberation intensity parameter value is determined based on the first reverberation intensity parameter value, the second type reverberation intensity parameter value, and the third type reverberation intensity parameter value, including but not limited to:
- the calculation formula of the target reverberation intensity parameter value is:
- G reverb min(1, G reverb_0 +w SpecRichness G SpecRichness +w bgm G bgm +w vocalGoodness G vocalGoodness )
- G reverb refers to the target reverberation intensity parameter value
- G reverb_0 refers to the preset basic reverberation intensity parameter value
- w SpecRichness refers to the first weight value corresponding to G SpecRichness
- w bgm refers to the value corresponding to G bgm
- the second weight value, w vocalGoodness refers to the third weight value corresponding to G vocalGoodness.
- the values of the above three weight values are set according to the magnitude of the influence on the reverberation intensity. For example, the first weight value has the largest value, and the second weight value has the smallest value. This is not specifically limited.
- reverberation processing is performed on the collected human voice signal based on the target reverberation intensity parameter value.
- the KTV reverberation algorithm includes two layers of parameters, one layer is the total reverberation gain, and the other layer is the internal parameters of the reverberation algorithm, which can then directly control the reverberation Part of the energy size achieves the purpose of controlling the intensity of the reverberation.
- reverberation processing is performed on the collected human voice signal based on the target reverberation intensity parameter value, including but not limited to:
- G reverb can be directly loaded as the total reverberation gain, and can also be loaded into one or more parameters within the reverberation algorithm, such as adjusting the echo gain, delay time, feedback network gain, etc.
- the embodiments of the present disclosure are This is not specifically limited.
- mixing processing is performed on the collected accompaniment audio signal and the human voice signal after the reverberation processing, and the audio signal after the mixing processing is output.
- the collected accompaniment audio signal and the vocal signal after the reverberation process will continue to be mixed, and after the mixing process After that, the audio signal can be directly output, for example, the audio signal after the mixing process is played through the speaker of the electronic device to realize the KTV sound effect.
- the embodiments of the present disclosure dynamically or pre-calculate the most suitable reverberation intensity parameter values for music of different rhythms and speeds, music of different accompaniment types, different parts of the same music, and music of different singers, thereby guiding the control of artificial reverberation algorithms
- the size of the reverberation of the output part of the human voice so as to achieve an adaptive K song sound effect.
- the embodiments of the present disclosure comprehensively consider the frequency domain richness, rhythm speed, and singer of the music.
- the frequency domain richness, rhythm speed, and singer of the music will be adaptively different.
- the embodiment of the present disclosure also provides a fusion method to finally obtain the total reverberation intensity parameter value, and the total reverberation intensity
- the parameter value can be loaded into the total gain of the reverberation, and it can also be loaded into one or more parameters inside the reverberation algorithm. Therefore, this kind of audio processing method achieves an adaptive K song sound effect, which makes the output sound of the electronic device more Full and graceful.
- Fig. 8 is a block diagram showing an audio processing device according to an embodiment.
- the device includes an acquisition module 801, a determination module 802, and a processing module 803.
- the collection module 801 is configured to collect accompaniment audio signals and vocal signals of the current music to be processed
- the determining module 802 is configured to determine a target reverberation intensity parameter value of the collected accompaniment audio signal, where the target reverberation intensity parameter value is used to indicate the rhythm speed, accompaniment type, and singer's singing score of the currently to-be-processed music At least one of
- the processing module 803 is configured to perform reverberation processing on the collected human voice signal based on the target reverberation intensity parameter value.
- the embodiment of the present disclosure after collecting the accompaniment audio signal and the vocal signal of the currently to-be-processed music, the embodiment of the present disclosure will determine the target reverberation intensity parameter value of the collected accompaniment audio signal, where the target reverberation intensity
- the parameter value is used to indicate at least one of the rhythm speed, accompaniment type, and singer's singing score of the currently to-be-processed music; after that, the collected human voice signal is subjected to reverberation processing based on the target reverberation intensity parameter value.
- the embodiments of the present disclosure consider various factors such as the accompaniment type of the music, the rhythm speed, and the singer’s singing score, and accordingly, adaptively generate the parameter value of the reverberation intensity of the current music to be processed.
- the self-adaptive K song sound effect makes the sound output by the electronic device more full and beautiful.
- the determining module 802 is further configured to determine a first reverberation strength parameter value of the collected accompaniment audio signal, where the first reverberation strength parameter value is used to indicate the accompaniment type of the currently to-be-processed music piece; Determine the second reverberation strength parameter value of the collected accompaniment audio signal, where the second reverberation strength parameter value is used to indicate the rhythm speed of the current music to be processed; determine the third reverberation strength parameter of the collected accompaniment audio signal Value, the third reverberation intensity parameter value is used to indicate the singing score of the singer of the currently to-be-processed music; based on the first reverberation intensity parameter value, the second type reverberation intensity parameter value, and the first The three types of reverberation intensity parameter values are used to determine the target reverberation intensity parameter value.
- the determining module 802 is further configured to transform the collected accompaniment audio signal from the time domain to the time-frequency domain to obtain the accompaniment audio frame sequence; obtain the amplitude information of each frame of the accompaniment audio; based on each frame of the accompaniment audio To determine the frequency domain richness coefficient of each frame of accompaniment audio; wherein, the frequency domain richness coefficient is used to indicate the frequency domain richness of the amplitude information of each frame of accompaniment audio, and the frequency domain richness reflects the current to-be-processed The accompaniment type of the music; the first reverberation intensity parameter value is determined based on the frequency domain rich coefficient of each frame of accompaniment audio.
- the determining module 802 is further configured to determine the global frequency domain rich coefficients of the current music piece to be processed based on the frequency domain rich coefficients of each frame of accompaniment audio; obtain the global frequency domain rich coefficients and frequency domain rich coefficients The first ratio between the maximum values, the smallest of the first ratio and the target value is determined as the first reverberation intensity parameter value.
- the determining module 802 is further configured to generate a waveform indicating the richness of the frequency domain based on the frequency domain rich coefficient of each frame of accompaniment audio; perform smoothing processing on the generated waveform graph based on the smoothed
- the waveform diagram determines the frequency domain rich coefficients of different parts of the current music to be processed; obtains the second ratio between the frequency domain rich coefficients of the different parts and the maximum value of the frequency domain rich coefficient; for each acquired second ratio , Determining the smallest of the second ratio and the target value as the first reverberation intensity parameter value.
- the determining module 802 is further configured to obtain the number of beats of the collected accompaniment audio signal in a prescribed time period; determine the third ratio between the obtained number of beats and the maximum value of the number of beats; The smallest of the three ratios and the target value is determined as the second reverberation intensity parameter value.
- the determining module 802 is further configured to obtain the audio singing score of the singer of the current music to be processed, and determine the third reverberation intensity parameter value based on the audio singing score.
- the determining module 802 is further configured to obtain a basic reverberation intensity parameter value, a first weight value, a second weight value, and a third weight value; determine the first weight value and the first mixing value.
- the processing module 803 is further configured to adjust the total reverberation gain of the collected human voice signal based on the target reverberation strength parameter value; or, based on the target reverberation strength parameter value , To adjust at least one reverberation algorithm parameter of the collected human voice signal.
- the processing module 803 is further configured to perform mixing processing on the collected accompaniment audio signal and the human voice signal after the reverberation processing after performing reverberation processing on the collected human voice signal, Output the audio signal after mixing.
- FIG. 9 shows a structural block diagram of an electronic device 900 according to an embodiment of the present disclosure.
- the device 900 is a portable mobile terminal, such as: smart phones, tablet computers, MP3 players (Moving Picture Experts Group Audio Layer III, Motion Picture Experts Compression Standard Audio Layer 3), MP4 (Moving Picture Experts Group Audio Layer IV, The dynamic image expert compresses the standard audio level 4) Player, laptop or desktop computer.
- the device 900 may also be called user equipment, portable terminal, laptop terminal, desktop terminal, and other names.
- the device 900 includes a processor 901 and a memory 902.
- the processor 901 includes one or more processing cores, such as a 4-core processor, an 8-core processor, and so on.
- the processor 901 adopts at least one hardware form among DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array, Programmable Logic Array).
- the processor 901 also includes a main processor and a coprocessor.
- the main processor is a processor used to process data in the wake-up state, also called a CPU (Central Processing Unit, central processing unit); It is a low-power processor for processing data in the standby state.
- the processor 901 is integrated with a GPU (Graphics Processing Unit, image processor), and the GPU is responsible for rendering and drawing content that needs to be displayed on the display screen.
- the processor 901 further includes an AI (Artificial Intelligence) processor, and the AI processor is used to process computing operations related to machine learning.
- AI Artificial Intelligence
- the memory 902 includes one or more computer-readable storage media, which are non-transitory.
- the memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more magnetic disk storage devices and flash memory storage devices.
- the device 900 further includes: a peripheral device interface 903 and at least one peripheral device.
- the processor 901, the memory 902, and the peripheral device interface 903 are connected by a bus or signal line.
- Each peripheral device is connected to the peripheral device interface 903 through a bus, a signal line or a circuit board.
- the peripheral device includes at least one of a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, a positioning assembly 908, and a power supply 909.
- the peripheral device interface 903 can be used to connect at least one peripheral device related to I/O (Input/Output) to the processor 901 and the memory 902.
- the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one of the processor 901, the memory 902, and the peripheral device interface 903 or The two are implemented on a separate chip or circuit board, which is not limited in this embodiment.
- the radio frequency circuit 904 is used for receiving and transmitting RF (Radio Frequency, radio frequency) signals, also called electromagnetic signals.
- the radio frequency circuit 904 communicates with a communication network and other communication devices through electromagnetic signals.
- the radio frequency circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals.
- the radio frequency circuit 904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on.
- the radio frequency circuit 904 communicates with other terminals through at least one wireless communication protocol.
- the wireless communication protocol includes but is not limited to: World Wide Web, Metropolitan Area Network, Intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network and/or WiFi (Wireless Fidelity, wireless fidelity) network.
- the radio frequency circuit 904 further includes a circuit related to NFC (Near Field Communication), which is not limited in the present disclosure.
- the display screen 905 is used to display a UI (User Interface, user interface).
- the UI includes graphics, text, icons, videos, and any combination of them.
- the display screen 905 also has the ability to collect touch signals on or above the surface of the display screen 905.
- the touch signal is input to the processor 901 as a control signal for processing.
- the display screen 905 is also used to provide virtual buttons and/or virtual keyboards, also called soft buttons and/or soft keyboards.
- the display screen 905 is a flexible display screen, which is disposed on the curved surface or the folding surface of the device 900. Furthermore, the display screen 905 is also configured as a non-rectangular irregular pattern, that is, a special-shaped screen.
- the display screen 905 is made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
- the camera assembly 906 is used to capture images or videos.
- the camera assembly 906 includes a front camera and a rear camera.
- the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal.
- the camera assembly 906 also includes a flash.
- the flash is a single-color temperature flash and also a dual-color temperature flash. Dual color temperature flash refers to a combination of warm light flash and cold light flash used for light compensation under different color temperatures.
- the audio circuit 907 includes a microphone and a speaker.
- the microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals and input them to the processor 901 for processing, or input to the radio frequency circuit 904 to implement voice communication.
- the microphone is also an array microphone or an omnidirectional acquisition microphone.
- the speaker is used to convert the electrical signal from the processor 901 or the radio frequency circuit 904 into sound waves.
- the speaker is a traditional thin-film speaker and also a piezoelectric ceramic speaker.
- the speaker When the speaker is a piezoelectric ceramic speaker, it not only converts electrical signals into sound waves that are audible to humans, but also converts electrical signals into sound waves that are inaudible to humans for distance measurement and other purposes.
- the audio circuit 907 also includes a headphone jack.
- the positioning component 908 is used to locate the current geographic location of the device 900 to implement navigation or LBS (Location Based Service, location-based service).
- the positioning component 908 is a positioning component based on the GPS (Global Positioning System, Global Positioning System) of the United States, the Beidou system of China, or the Galileo system of Russia.
- the power supply 909 is used to supply power to various components in the device 900.
- the power source 909 is alternating current, direct current, disposable batteries, or rechargeable batteries.
- the rechargeable battery is a wired rechargeable battery or a wireless rechargeable battery.
- a wired rechargeable battery is a battery charged through a wired line
- a wireless rechargeable battery is a battery charged through a wireless coil.
- the rechargeable battery is also used to support fast charging technology.
- the device 900 further includes one or more sensors 910.
- the one or more sensors 910 include, but are not limited to: an acceleration sensor 911, a gyroscope sensor 912, a pressure sensor 913, a fingerprint sensor 914, an optical sensor 915, and a proximity sensor 916.
- the acceleration sensor 911 detects the magnitude of acceleration on the three coordinate axes of the coordinate system established by the device 900.
- the acceleration sensor 911 is used to detect the components of gravitational acceleration on three coordinate axes.
- the processor 901 controls the display screen 905 to display the user interface in a horizontal view or a vertical view according to the gravity acceleration signal collected by the acceleration sensor 911.
- the acceleration sensor 911 is also used for the collection of game or user motion data.
- the gyroscope sensor 912 detects the body direction and rotation angle of the device 900, and the gyroscope sensor 912 and the acceleration sensor 911 cooperate to collect the user's 3D actions on the device 900.
- the processor 901 implements the following functions according to the data collected by the gyroscope sensor 912: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
- the pressure sensor 913 is disposed on the side frame of the device 900 and/or the lower layer of the display screen 905.
- the processor 901 performs left and right hand recognition or quick operation according to the holding signal collected by the pressure sensor 913.
- the processor 901 controls the operability controls on the UI interface according to the pressure operation of the user on the display screen 905.
- the operability control includes at least one of a button control, a scroll bar control, an icon control, and a menu control.
- the fingerprint sensor 914 is used to collect the user's fingerprint, and the processor 901 can identify the user's identity according to the fingerprint collected by the fingerprint sensor 914, or the fingerprint sensor 914 can identify the user's identity according to the collected fingerprint. When it is recognized that the user's identity is a trusted identity, the processor 901 authorizes the user to perform related sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, paying, and changing settings.
- the fingerprint sensor 914 is provided on the front, back or side of the device 900. When a physical button or manufacturer logo is provided on the device 900, the fingerprint sensor 914 is integrated with the physical button or manufacturer logo.
- the optical sensor 915 is used to collect the ambient light intensity.
- the processor 901 controls the display brightness of the display screen 905 according to the ambient light intensity collected by the optical sensor 915. Specifically, when the ambient light intensity is high, the display brightness of the display screen 905 is increased; when the ambient light intensity is low, the display brightness of the display screen 905 is decreased.
- the processor 901 also dynamically adjusts the shooting parameter values of the camera assembly 906 according to the ambient light intensity collected by the optical sensor 915.
- the proximity sensor 916 also called a distance sensor, is usually installed on the front panel of the device 900.
- the proximity sensor 916 is used to collect the distance between the user and the front of the device 900.
- the processor 901 controls the display screen 905 to switch from the on-screen state to the off-screen state; when the proximity sensor 916 detects When the distance between the user and the front of the device 900 gradually increases, the processor 901 controls the display screen 905 to switch from the rest screen state to the bright screen state.
- FIG. 10 is a structural block diagram of an electronic device 1000 provided by an embodiment of the present disclosure.
- the device 1000 behaves as a server.
- the server 1000 may have relatively large differences due to different configurations or performances, and includes one or more processors (central processing units, CPU) 1001 and one or more memories 1002.
- processors central processing units, CPU
- the server also has components such as a wired or wireless network interface, a keyboard, an input and output interface for input and output, and the server also includes other components for implementing device functions, which will not be repeated here.
- the embodiments of the present disclosure provide an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein the processor is configured to execute the instructions to implement
- the following steps are as follows: Collect the accompaniment audio signal and the vocal signal of the current music to be processed; determine the target reverberation intensity parameter value of the collected accompaniment audio signal, and the target reverberation intensity parameter value is used to indicate the rhythm speed of the current music to be processed , At least one of the accompaniment type and the singer's singing score; performing reverberation processing on the collected human voice signal based on the target reverberation intensity parameter value.
- the processor is configured to execute the instructions to implement the following steps: determine a first reverberation intensity parameter value of the collected accompaniment audio signal, and the first reverberation intensity parameter value is used for Indicate the accompaniment type of the current music to be processed; determine the second reverberation intensity parameter value of the collected accompaniment audio signal, the second reverberation intensity parameter value is used to indicate the rhythm speed of the current music to be processed; determine the collected accompaniment The third reverberation strength parameter value of the audio signal, where the third reverberation strength parameter value is used to indicate the singing score of the singer of the current music to be processed; based on the first reverberation strength parameter value, the second type The reverberation intensity parameter value and the third type of reverberation intensity parameter value determine the target reverberation intensity parameter value.
- the processor is configured to execute the instructions to implement the following steps: transform the collected accompaniment audio signal from the time domain to the time-frequency domain to obtain an accompaniment audio frame sequence; obtain each frame of accompaniment audio Based on the amplitude information of each frame of accompaniment audio, determine the frequency domain rich coefficient of each frame of accompaniment audio; wherein, the frequency domain rich coefficient is used to indicate the frequency domain richness of the amplitude information of each frame of accompaniment audio, the The frequency domain richness reflects the accompaniment type of the current music to be processed; the first reverberation intensity parameter value is determined based on the frequency domain richness coefficient of each frame of accompaniment audio.
- the processor is configured to execute the instructions to implement the following steps: based on the frequency domain rich coefficients of each frame of accompaniment audio, determine the global frequency domain rich coefficients of the current music to be processed; obtain the global The first ratio between the frequency domain rich coefficient and the maximum value of the frequency domain rich coefficient is determined as the first reverberation intensity parameter value, which is the smallest of the first ratio and the target value.
- the processor is configured to execute the instructions to implement the following steps: based on the frequency domain richness coefficient of each frame of accompaniment audio, generating a waveform diagram indicating the frequency domain richness; Perform smoothing processing on the graph, determine the frequency domain rich coefficients of different parts of the current music to be processed based on the smoothed waveform diagram; obtain the second ratio between the frequency domain rich coefficients of the different parts and the maximum value of the frequency domain rich coefficients; For each acquired second ratio, the smallest of the second ratio and the target value is determined as the first reverberation intensity parameter value.
- the processor is configured to execute the instructions to implement the following steps: obtain the number of beats of the collected accompaniment audio signal in a specified period of time; determine between the obtained number of beats and the maximum value of the number of beats The third ratio of the third ratio; the smallest of the third ratio and the target value is determined as the second reverberation intensity parameter value.
- the processor is configured to execute the instructions to implement the following steps: obtain the audio singing score of the singer of the currently to-be-processed music, and determine the third mix based on the audio singing score. Sound intensity parameter value.
- the processor is configured to execute the instructions to implement the following steps: obtain a basic reverberation intensity parameter value, a first weight value, a second weight value, and a third weight value; determine the first weight value; A first sum value between a weight value and the first reverberation intensity parameter value; determine a second sum value between the second weight value and the second reverberation intensity parameter value; determine the first The third sum value between the three-weight value and the third reverberation intensity parameter value; obtaining the basic reverberation intensity parameter value, the first sum value, the second sum value and the third sum value The fourth sum value between the values determines the smallest of the fourth ratio and the target value as the target reverberation intensity parameter value.
- the processor is configured to execute the instructions to implement the following steps: adjust the total reverberation gain of the collected human voice signal based on the target reverberation intensity parameter value; or, Based on the target reverberation intensity parameter value, at least one reverberation algorithm parameter of the collected human voice signal is adjusted.
- the processor is configured to execute the instructions to implement the following steps: perform mixing processing on the collected accompaniment audio signal and the human voice signal after reverberation processing, and output the mixing processing After the audio signal.
- the embodiment of the present disclosure also provides a storage medium including instructions, such as a memory including instructions, which may be executed by a processor of the electronic device 900 or the electronic device 1000 to complete the audio processing method described above.
- the storage medium is a non-transitory computer-readable storage medium, for example, the non-transitory computer-readable storage medium is ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage Equipment, etc.
- the embodiments of the present disclosure also provide a computer program product.
- the instructions in the computer program product are executed by the processor of the electronic device 900 or the electronic device 1000, the electronic device 900 or the electronic device 1000 can execute the method as described above.
- the audio processing method in.
Landscapes
- Physics & Mathematics (AREA)
- Engineering & Computer Science (AREA)
- Acoustics & Sound (AREA)
- Multimedia (AREA)
- Electrophonic Musical Instruments (AREA)
- Reverberation, Karaoke And Other Acoustics (AREA)
Abstract
一种音频处理方法及电子设备,涉及信号处理技术领域。该方法包括:采集当前待处理乐曲的伴奏音频信号和人声信号(201);确定采集到的伴奏音频信号的目标混响强度参数值,目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种(202);基于目标混响强度参数值对采集到的人声信号进行混响处理(203)。
Description
本公开要求于2020年01月22日提交的申请号为202010074552.2的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
本公开涉及信号处理技术领域,尤其涉及一种音频处理方法及电子设备。
长久以来,唱歌作为一项日常休闲娱乐活动一直广受用户追捧。时下随着诸如智能手机或平板电脑等电子设备不断地推陈出新,用户通过电子设备上安装的应用程序即可实现唱歌,甚至通过电子设备上安装的应用程序用户无需走进KTV即可实现K歌音效。
其中,K歌音效是指通过对采集到的人声和背景音乐进行音频处理,使得处理后的人声相较于处理前的人声更加悦耳,同时能够掩蔽掉人声一部分的音高不准等问题。
发明内容
本公开提供一种音频处理方法及电子设备,能够使得电子设备输出的声音更加饱满和优美。本公开的技术方案如下:
根据本公开实施例的一方面,提供一种音频处理方法,包括:
采集当前待处理乐曲的伴奏音频信号和人声信号;
确定采集到的伴奏音频信号的目标混响强度参数值,所述目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;
基于所述目标混响强度参数值对采集到的人声信号进行混响处理。
在一些实施例中,所述确定采集到的伴奏音频信号的目标混响强度参数值,包括:
确定采集到的伴奏音频信号的第一混响强度参数值,所述第一混响强度参数值用于指示当前待处理乐曲的伴奏类型;
确定采集到的伴奏音频信号的第二混响强度参数值,所述第二混响强度参数值用于指示当前待处理乐曲的节奏速度;
确定采集到的伴奏音频信号的第三混响强度参数值,所述第三混响强度参数值用于指示当前待处理乐曲的演唱者的演唱评分;
基于所述第一混响强度参数值、所述第二类混响强度参数值和所述第三类混响强度参数值,确定所述目标混响强度参数值。
在一些实施例中,所述确定采集到的伴奏音频信号的第一混响强度参数值,包括:
将采集到的伴奏音频信号由时域变换到时频域,得到伴奏音频帧序列;
获取每帧伴奏音频的幅度信息;
基于每帧伴奏音频的幅度信息,确定每帧伴奏音频的频域丰富系数;
其中,所述频域丰富系数用于指示每帧伴奏音频的幅度信息的频域丰富程度,所述频域丰富程度反映了当前待处理乐曲的伴奏类型;
基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值。
在一些实施例中,所述基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值,包括:
基于每帧伴奏音频的频域丰富系数,确定当前待处理乐曲的全局频域丰富系数;
获取所述全局频域丰富系数与频域丰富系数最大值之间的第一比值,将所述第一比值和目标数值中的最小者确定为所述第一混响强度参数值。
在一些实施例中,所述基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值,包括:
基于每帧伴奏音频的频域丰富系数,生成用于指示频域丰富程度的波形图;
对生成的波形图进行平滑处理,基于平滑后的波形图确定当前待处理乐曲的不同部分的频域丰富系数;
获取所述不同部分的频域丰富系数分别与频域丰富系数最大值之间的第二比值;
对于获取到的每个第二比值,将所述第二比值和目标数值中的最小者确定为所述第一混响强度参数值。
在一些实施例中,所述确定采集到的伴奏音频信号的第二混响强度参数值,包括:
获取采集到的伴奏音频信号在规定时长的节拍数;
确定获取到的节拍数与节拍数最大值之间的第三比值;
将所述第三比值和目标数值中的最小者,确定为所述第二混响强度参数值。
在一些实施例中,所述确定采集到的伴奏音频信号的第三混响强度参数值,包括:
获取当前待处理乐曲的演唱者的音频演唱分值,基于所述音频演唱分值确定所述第三混响强度参数值。
在一些实施例中,所述基于所述第一混响强度参数值、所述第二类混响强度参数值和所述第三类混响强度参数值,确定所述目标混响强度参数值,包括:
获取基础混响强度参数值、第一权重值、第二权重值以及第三权重值;
确定所述第一权重值与所述第一混响强度参数值之间的第一和值;
确定所述第二权重值与所述第二混响强度参数值之间的第二和值;
确定所述第三权重值与所述第三混响强度参数值之间的第三和值;
获取所述基础混响强度参数值、所述第一和值、所述第二和值与所述第三和值之间的第四和值,将所述第四比值和目标数值中的最小者,确定为所述目标混响强度参数值。
在一些实施例中,所述基于所述目标混响强度参数值对采集到的人声信号进行混响处理,包括:
基于所述目标混响强度参数值,对采集到的人声信号的混响总增益进行调整;
或,基于所述目标混响强度参数值,对采集到的人声信号的至少一项混响算法参数进行调整。
在一些实施例中,在对采集到的人声信号进行混响处理后,所述方法还包括:
对采集到的伴奏音频信号和经过混响处理后的人声信号进行混音处理,输出经过混音处理后的音频信号。
根据本公开实施例的另一方面,提供一种音频处理装置,包括:
采集模块,被配置为采集当前待处理乐曲的伴奏音频信号和人声信号;
确定模块,被配置为确定采集到的伴奏音频信号的目标混响强度参数值,所述目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;
处理模块,被配置为基于所述目标混响强度参数值对采集到的人声信号进行混响处理。
在一些实施例中,所述确定模块,还被配置为确定采集到的伴奏音频信号的第一混响强度参数值,所述第一混响强度参数值用于指示当前待处理乐曲的伴奏类型;确定采集到的伴奏音频信号的第二混响强度参数值,所述第二混响强度参数值用于指示当前待处理乐曲的节奏速度;确定采集到的伴奏音频信号的第三混响强度参数值,所述第三混响强度参数值用于指示当前待处理乐曲的演唱者的演唱评分;基于所述第一混响强度参数值、所述第二类混响强度参数值和所述第三类混响强度参数值,确定所述目标混响强度参数值。
在一些实施例中,所述确定模块,还被配置为将采集到的伴奏音频信号由时域变换到时频域,得到伴奏音频帧序列;获取每帧伴奏音频的幅度信息;基于每帧伴奏音频的幅度信息,确定每帧伴奏音频的频域丰富系数;其中,所述频域丰富系数用于指示每帧伴奏音频的幅度 信息的频域丰富程度,所述频域丰富程度反映了当前待处理乐曲的伴奏类型;基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值。
在一些实施例中,所述确定模块,还被配置为基于每帧伴奏音频的频域丰富系数,确定当前待处理乐曲的全局频域丰富系数;获取所述全局频域丰富系数与频域丰富系数最大值之间的第一比值,将所述第一比值和目标数值中的最小者确定为所述第一混响强度参数值。
在一些实施例中,所述确定模块,还被配置为基于每帧伴奏音频的频域丰富系数,生成用于指示频域丰富程度的波形图;对生成的波形图进行平滑处理,基于平滑后的波形图确定当前待处理乐曲的不同部分的频域丰富系数;获取所述不同部分的频域丰富系数分别与频域丰富系数最大值之间的第二比值;对于获取到的每个第二比值,将所述第二比值和目标数值中的最小者确定为所述第一混响强度参数值。
在一些实施例中,所述确定模块,还被配置为获取采集到的伴奏音频信号在规定时长的节拍数;确定获取到的节拍数与节拍数最大值之间的第三比值;将所述第三比值和目标数值中的最小者,确定为所述第二混响强度参数值。
在一些实施例中,所述确定模块,还被配置为获取当前待处理乐曲的演唱者的音频演唱分值,基于所述音频演唱分值确定所述第三混响强度参数值。
在一些实施例中,所述确定模块,还被配置为获取基础混响强度参数值、第一权重值、第二权重值以及第三权重值;确定所述第一权重值与所述第一混响强度参数值之间的第一和值;确定所述第二权重值与所述第二混响强度参数值之间的第二和值;确定所述第三权重值与所述第三混响强度参数值之间的第三和值;获取所述基础混响强度参数值、所述第一和值、所述第二和值与所述第三和值之间的第四和值,将所述第四比值和目标数值中的最小者,确定为所述目标混响强度参数值。
在一些实施例中,所述处理模块,还被配置为基于所述目标混响强度参数值,对采集到的人声信号的混响总增益进行调整;或,基于所述目标混响强度参数值,对采集到的人声信号的至少一项混响算法参数进行调整。
在一些实施例中,所述处理模块,还被配置为在对采集到的人声信号进行混响处理后,对采集到的伴奏音频信号和经过混响处理后的人声信号进行混音处理,输出经过混音处理后的音频信号。
根据本公开实施例的另一方面,提供一种电子设备,包括:
处理器;
用于存储所述处理器可执行指令的存储器;
其中,所述处理器被配置为执行所述指令,以实现上述的音频处理方法。
根据本公开实施例的另一方面,提供一种存储介质,所述存储介质中的指令由电子设备的处理器执行,使得电子设备能够执行上述的音频处理方法。
根据本公开实施例的另一方面,提供一种计算机程序产品,所述计算机程序产品中的指令由电子设备的处理器执行,使得电子设备能够执行如上述的音频处理方法。
图1是根据一实施例示出的一种音频处理方法涉及的实施环境的示意图。
图2是根据一实施例示出的一种音频处理方法的流程图。
图3是根据一实施例示出的一种音频处理方法的流程图。
图4是根据一实施例示出的一种音频处理方法的整体系统框图。
图5是根据一实施例示出的一种音频处理方法的流程图。
图6是根据一实施例示出的一种关于频域丰富程度的波形图。
图7是根据一实施例示出的一种关于频域丰富程度的平滑后的波形图。
图8是根据一实施例示出的一种音频处理装置的框图。
图9是根据一实施例示出的一种电子设备的框图。
图10是根据一实施例示出的另一种电子设备的框图。
本公开所涉及的用户信息为经用户授权或者经过各方充分授权的信息。其中,A、B、C中的至少一种,包括以下几种情况:单独的A,单独的B,单独的C,A和B,A和C,B和C,以及A、B、C。
在对本公开实施例进行详细地解释说明之前,先对本公开实施例涉及的一些名词术语或缩略语进行介绍。
K歌音效:是指通过对采集到的人声和背景音乐进行音频处理,使得处理后的人声相较于处理前的人声更加悦耳,同时能够掩蔽掉人声一部分的音高不准等问题。简言之,K歌音效用于修饰采集到的人声。
BGM(Background Music,伴奏音乐或背景音乐):简称为伴乐、配乐。广义上来讲,BGM通常是指在电视剧、电影、动画、电子游戏、网站中用于调节气氛的一种音乐,插入于对话之中,能够增强情感的表达,达到一种让观众身临其境的感受。另外,在一些公共场合(比如酒吧、咖啡厅或商场等)播放的音乐也称为背景音乐。在本公开实施例中,针对唱歌场景, BGM指代歌曲伴奏。
STFT(Short-Time Fourier Transform,短时傅里叶变换):是和傅里叶变换相关的一种数学变换,用以确定时变信号其局部区域正弦波的频率与相位。即,将长的非平稳信号看作是一系列短时平稳信号的叠加,短时平稳信号通过加窗函数实现,也即是,截取多段信号分别进行傅里叶变换。它的时频分析特性表现在:通过时间窗内的一段信号来表示某一时刻的特征。
混响(reverberation):声波在室内传播时,会被诸如墙壁、天花板或地板等障碍物反射,每反射一次都要被障碍物吸收一些。这样,当声源停止发声后,声波在室内要经过多次反射和吸收,最后才消失,人耳会感觉到声源停止发声后还有若干个声波混合持续一段时间,即在声源停止发声后仍然存在声音延续现象,这种现象即称为混响。在一些实施例中,混响主要用于唱卡拉OK,增加话筒声音的延时,产生适量的回声,使唱歌的声音更圆润更优美,歌声不那么干瘪。即,针对K歌的歌声来讲,为了使得效果更好声音不那么干瘪无力,一般都会在后期人工地加上混响,以使得声音更加饱满和优美。
下面对本公开实施例提供的一种音频处理方法涉及的实施环境进行介绍。
参见图1,该实施环境包括:用于音频处理的电子设备101。其中,电子设备101为终端或服务器,本申请实施例对此不进行具体限定。以终端为例,则终端的类型包括但不限于:移动式终端和固定式终端。
在一些实施例中,移动式终端包括但不限于:智能手机、平板电脑、笔记本电脑、电子阅读器、MP3播放器(Moving Picture Experts Group Audio Layer III,动态影像专家压缩标准音频层面3)、MP4(Moving Picture Experts Group Audio Layer IV,动态影像专家压缩标准音频层面4)播放器等;固定式终端包括但不限于台式电脑,本申请实施例对此不进行具体限定。
在一些实施例中,终端上通常安装有具有音频处理功能的音乐应用程序,以执行本申请实施例提供的音频处理方法。另外,除了可在终端上执行该方法以外,终端还可通过音乐类应用程序或视频类应用程序将待处理的音频信号上传至服务器,由服务器执行本申请实施例提供的音频处理方法,并将结果返回给终端,本申请实施例对此不进行具体限定。
基于上述的实施环境,为了使得声音更加饱满和优美,电子设备101通常会对采集到的人声信号进行人工混响处理。
简言之,在采集到伴奏音频信号(也称为BGM音频信号)和人声信号后,通过短时傅里叶变换将BGM音频信号从时域变换至时频域,得到一个关于BGM音频信号的帧序列;之后,获取每帧伴奏音频的幅度信息,据此计算每帧伴奏音频的幅度信息的频域丰富程度;除此之 外,还能够获取BGM音频信号在规定时长(比如每分钟)的节拍数,据此计算BGM音频信号的节奏速度。
通常情况下,对于背景音乐伴奏成分简单(比如纯吉他伴奏)、慢速的歌曲,会加入小混响,使得人声更纯净;而对于背景音乐伴奏成分多样(比如乐队歌曲伴奏)、快速的歌曲,会加入大混响,起到烘托气氛以及突出人声的作用。
在本公开实施例中,针对不同节奏和伴奏类型的歌曲、同一歌曲的不同部分、不同演唱者,能够动态或预先计算出最适合的混响强度参数值,进而指导人工混响算法控制输出人声部分混响大小,从而达到自适应的K歌音效。换一表达方式,本公开实施例综合考虑了歌曲的频域丰富程度、节奏速度以及演唱者等多方面因素,并据此自适应地生成不同的混响强度参数值,从而达到了自适应的K歌音效。
下面通过以下实施方式对本公开的实施例提供的音频处理方法进行详细地解释说明。
图2是根据一实施例示出的一种音频处理方法的流程图,如图2所示,该音频处理方法用于电子设备中,包括以下步骤。
在201中,采集当前待处理乐曲的伴奏音频信号和人声信号。
在202中,确定采集到的伴奏音频信号的目标混响强度参数值,目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种。
在203中,基于目标混响强度参数值对采集到的人声信号进行混响处理。
本公开实施例提供的方法,在采集到当前待处理乐曲的伴奏音频信号和人声信号后,本公开实施例会确定采集到的伴奏音频信号的目标混响强度参数值,其中,目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;之后,基于目标混响强度参数值对采集到的人声信号进行混响处理。基于以上描述可知,本公开实施例考虑了乐曲的伴奏类型、节奏速度以及演唱者的演唱评分等多方面的因素,并据此自适应地生成当前待处理乐曲的混响强度参数值,达到了自适应的K歌音效,使得电子设备输出的声音更加饱满和优美。
在一些实施例中,所述确定采集到的伴奏音频信号的目标混响强度参数值,包括:
确定采集到的伴奏音频信号的第一混响强度参数值,所述第一混响强度参数值用于指示当前待处理乐曲的伴奏类型;
确定采集到的伴奏音频信号的第二混响强度参数值,所述第二混响强度参数值用于指示当前待处理乐曲的节奏速度;
确定采集到的伴奏音频信号的第三混响强度参数值,所述第三混响强度参数值用于指示当前待处理乐曲的演唱者的演唱评分;
基于所述第一混响强度参数值、所述第二类混响强度参数值和所述第三类混响强度参数值,确定所述目标混响强度参数值。
在一些实施例中,所述确定采集到的伴奏音频信号的第一混响强度参数值,包括:
将采集到的伴奏音频信号由时域变换到时频域,得到伴奏音频帧序列;
获取每帧伴奏音频的幅度信息;
基于每帧伴奏音频的幅度信息,确定每帧伴奏音频的频域丰富系数;
其中,所述频域丰富系数用于指示每帧伴奏音频的幅度信息的频域丰富程度,所述频域丰富程度反映了当前待处理乐曲的伴奏类型;
基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值。
在一些实施例中,所述基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值,包括:
基于每帧伴奏音频的频域丰富系数,确定当前待处理乐曲的全局频域丰富系数;
获取所述全局频域丰富系数与频域丰富系数最大值之间的第一比值,将所述第一比值和目标数值中的最小者确定为所述第一混响强度参数值。
在一些实施例中,所述基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值,包括:
基于每帧伴奏音频的频域丰富系数,生成用于指示频域丰富程度的波形图;
对生成的波形图进行平滑处理,基于平滑后的波形图确定当前待处理乐曲的不同部分的频域丰富系数;
获取所述不同部分的频域丰富系数分别与频域丰富系数最大值之间的第二比值;
对于获取到的每个第二比值,将所述第二比值和目标数值中的最小者确定为所述第一混响强度参数值。
在一些实施例中,所述确定采集到的伴奏音频信号的第二混响强度参数值,包括:
获取采集到的伴奏音频信号在规定时长的节拍数;
确定获取到的节拍数与节拍数最大值之间的第三比值;
将所述第三比值和目标数值中的最小者,确定为所述第二混响强度参数值。
在一些实施例中,所述确定采集到的伴奏音频信号的第三混响强度参数值,包括:
获取当前待处理乐曲的演唱者的音频演唱分值,基于所述音频演唱分值确定所述第三混响强度参数值。
在一些实施例中,所述基于所述第一混响强度参数值、所述第二类混响强度参数值和所述第三类混响强度参数值,确定所述目标混响强度参数值,包括:
获取基础混响强度参数值、第一权重值、第二权重值以及第三权重值;
确定所述第一权重值与所述第一混响强度参数值之间的第一和值;
确定所述第二权重值与所述第二混响强度参数值之间的第二和值;
确定所述第三权重值与所述第三混响强度参数值之间的第三和值;
获取所述基础混响强度参数值、所述第一和值、所述第二和值与所述第三和值之间的第四和值,将所述第四比值和目标数值中的最小者,确定为所述目标混响强度参数值。
在一些实施例中,所述基于所述目标混响强度参数值对采集到的人声信号进行混响处理,包括:
基于所述目标混响强度参数值,对采集到的人声信号的混响总增益进行调整;
或,基于所述目标混响强度参数值,对采集到的人声信号的至少一项混响算法参数进行调整。
在一些实施例中,所述方法还包括:
对采集到的伴奏音频信号和经过混响处理后的人声信号进行混音处理,输出经过混音处理后的音频信号。
上述所有可选技术方案,能够采用任意结合形成本公开的可选实施例,在此不再一一赘述。
图3是根据一实施例示出的一种音频处理方法的流程图,该音频处理方法用于电子设备中,结合图4所示的整体系统框图,该音频处理方法包括以下步骤。
在301中,采集当前待处理乐曲的伴奏音频信号和人声信号。
其中,当前待处理乐曲为用户当前正在演唱的歌曲,相应地,伴奏音频信号在本文中也被称之为背景音乐伴奏或BGM音频信号。以电子设备为智能手机为例,则电子设备通过自身配置的或外置的麦克风,采集当前待处理乐曲的伴奏音频信号和人声信号。
在302中,确定采集到的伴奏音频信号的目标混响强度参数值,其中,目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种。
通常情况下,进行混响处理的一个基本原则是:对于背景音乐伴奏成分简单(比如纯吉他伴奏)、慢速的歌曲,会加入小混响,使得人声更纯净;对于背景音乐伴奏成分多样(比如乐队歌曲伴奏)、快速的歌曲,会加入大混响,起到烘托气氛以及突出人声的作用。
其中,目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种,包括如下几种情况:目标混响强度参数值用于指示当前待处理乐曲的节奏速度;目标混响强度参数值用于指示当前待处理乐曲的伴奏类型;目标混响强度参数值用于指示当前待处理乐曲的演唱者的演唱评分;目标混响强度参数值用于指示当前待处理 乐曲的节奏速度和伴奏类型;目标混响强度参数值用于指示当前待处理乐曲的节奏速度和演唱者的演唱评分;目标混响强度参数值用于指示当前待处理乐曲的伴奏类型和演唱者的演唱评分;目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分。
在一些实施例中,如图5所示,确定采集到的伴奏音频信号的目标混响强度参数值,包括如下步骤:
在3021中,确定采集到的伴奏音频信号的第一混响强度参数值,其中,第一混响强度参数值用于指示当前待处理乐曲的伴奏类型。
在本公开实施例中,当前待处理乐曲的伴奏类型通过频域丰富程度来表征。其中,歌曲本身的伴奏越丰富,相应的频域丰富程度越高;反之亦然。换一种表达方式,伴奏强烈的歌曲相较于伴奏简单的歌曲来说,具有更高的频域丰富系数。其中,频域丰富系数用于指示每帧伴奏音频的幅度信息的频域丰富程度,即频域丰富程度反映了当前待处理乐曲的伴奏类型。
在一些实施例中,确定采集到的伴奏音频信号的第一混响强度参数值,包括但不限于如下步骤:
将采集到的伴奏音频信号由时域变换到时频域,得到伴奏音频帧序列。
如图4所示,本公开实施例对当前待处理乐曲的BCM音频信号进行短时傅里叶变换,实现由时域变换至时频域。
比如,长度为T的音频信号x在时域上为x(t),其中t代表时间,0<t≤T,则经过短时傅里叶变换后,x(t)在频域上表示为:X(n,k)=STFT(x(t))。
其中,n指代得到的伴奏音频帧序列中的任意一帧,0<n≤N,N为总帧数,k指代中心频率序列中的任意一个频点,0<k≤K,K为总频点数。
获取每帧伴奏音频的幅度信息;基于每帧伴奏音频的幅度信息,确定每帧伴奏音频的频域丰富系数。
在通过短时傅里叶变换由时域变换至时频域后,会获得每帧音频信号的幅度信息和相位信息。在一些实施例中,通过如下公式来确定每帧伴奏音频的幅度Mag。即,在频域上BGM音频信号的幅度为:Mag(n,k)=abs(X(n,k))。
相应地,每帧伴奏音频的频域丰富程度SpecRichness,即频域丰富系数为:
需要说明的是,对于一首歌曲来讲,歌曲本身的伴奏越丰富,相应的频域丰富程度越高;反之亦然。在一些实施例中,图6示出了两首歌曲的频域丰富程度,由于歌曲A的伴奏强烈, 而歌曲B的伴奏相较于前者较为简单,因此歌曲A的频域丰富程度要高于歌曲B。图6展示的是关于两首歌曲的原始计算的SpecRichness,而图7展示的为平滑后的SpecRichness。由图6和图7能够看出,伴奏强烈的歌曲相较于伴奏简单的歌曲来说具有更高的SpecRichness。
基于每帧伴奏音频的频域丰富系数确定第一混响强度参数值。
在本公开实施例中,一种实现方式为通过预先计算好的全局SpecRichness对不同歌曲分配不同的混响程度。
即,在一些实施例中,基于每帧伴奏音频的频域丰富系数确定第一混响强度参数值,包括但不限于:基于每帧伴奏音频的频域丰富系数,确定当前待处理乐曲的全局频域丰富系数;获取全局频域丰富系数与频域丰富系数最大值之间的第一比值,将第一比值和目标数值中的最小者确定为第一混响强度参数值。
在一些实施例中,全局频域丰富系数为每帧伴奏音频的频域丰富系数的均值,本公开实施例对此不进行具体限定。另外,目标数值在本文中指代数值1。相应地,通过计算出的SpecRichness计算第一混响强度参数值的公式为:
其中,G
SpecRichness指代第一混响强度参数值,SpecRichness_max指代预设的最大允许的SpecRichness值。
在本公开实施例中,另一种实现方式为通过平滑后的SpecRichness对每首歌曲的不同部分分配不同的混响程度。比如,副歌部分混响程度会更强,如图7中上方曲线所示。
即,在另一些实施例中,基于每帧伴奏音频的频域丰富系数确定第一混响强度参数值,包括但不限于:基于每帧伴奏音频的频域丰富系数,生成用于指示频域丰富程度的波形图,如图7所示;对生成的波形图进行平滑处理,基于平滑后的波形图确定当前待处理乐曲的不同部分的频域丰富系数;获取不同部分的频域丰富系数分别与频域丰富系数最大值之间的第二比值;对于获取到的每个第二比值,将第二比值和目标数值中的最小者确定为第一混响强度参数值。
针对该种计算方式,对于一首歌曲来说,会通过计算出来的SpecRichness,计算出多个第一混响强度参数值。
在一些实施例中,不同部分的频域丰富系数为相应部分的各帧伴奏音频的频域丰富系数的均值,本公开实施例对此不进行具体限定。其中,上述不同部分至少包括主歌部分和副歌部分。
在3022中,确定采集到的伴奏音频信号的第二混响强度参数值,其中,第二混响强度参数值用于指示当前待处理乐曲的节奏速度。
在本公开实施例中,通过节拍数来表征当前待处理乐曲的节奏速度。即,在一些实施例中,确定采集到的伴奏音频信号的第二混响强度参数值,包括但不限于:获取采集到的伴奏音频信号在规定时长的节拍数;确定获取到的节拍数与节拍数最大值之间的第三比值;将第三比值和目标数值中的最小者,确定为第二混响强度参数值。
在一些实施例中,在规定时长内的节拍数,为每分钟的节拍数,本公开实施例对此不进行具体限定。其中,BPM(Beat Per Minute)释为每分钟的节拍数的单位,即是在一分钟的时间段落之间所发出的声音节拍的数量,这个数量的单位便是BPM,也叫做拍子数。
其中,通过节拍数分析算法来获取当前待处理乐曲每分钟的节拍数。相应地,第二混响强度参数值的计算公式为:
其中,G
bgm指代第二混响强度参数值,BGM指代计算出来的每分钟节拍数,BGM_max指代预设的最大允许的每分钟节拍数。
在3023中,确定采集到的伴奏音频信号的第三混响强度参数值,其中,第三混响强度参数值用于指示当前待处理乐曲的演唱者的演唱评分。
通常情况下,歌唱水平高(演唱评分相对较高)的演唱者偏好小混响;而歌唱水平差(演唱评分相对较低)的演唱者偏好大混响。在一些实施例中,本公开实施例还能够通过提取当前待处理乐曲的演唱者的演唱评分(音频演唱分值)来进行混响强度控制。即,在一些实施例中,确定采集到的伴奏音频信号的第三混响强度参数值,包括但不限于:获取当前待处理乐曲的演唱者的音频演唱分值,基于音频演唱分值确定第三混响强度参数值。
在一些实施例中,音频演唱分值指代演唱者的历史歌曲评分或实时歌曲评分,而历史歌曲评分为最近一个月、最近3个月、最近半年或最近1年内的歌曲评分,本公开实施例对此不进行具体限定。其中,歌曲评分的满分为100分。
相应地,第三混响强度参数值的计算公式为:
其中,G
vocalGoodness指代第三混响强度参数值,KTV_Score指代获取到的音频演唱分值。
在3024中,基于第一混响强度参数值、第二类混响强度参数值和第三类混响强度参数值, 确定目标混响强度参数值。
在一些实施例中,基于第一混响强度参数值、第二类混响强度参数值和第三类混响强度参数值,确定目标混响强度参数值,包括但不限于:
获取基础混响强度参数值、第一权重值、第二权重值以及第三权重值;确定第一权重值与第一混响强度参数值之间的第一和值;确定第二权重值与第二混响强度参数值之间的第二和值;确定第三权重值与第三混响强度参数值之间的第三和值;获取基础混响强度参数值、第一和值、第二和值与第三和值之间的第四和值,将第四比值和目标数值中的最小者,确定为目标混响强度参数值。
相应地,目标混响强度参数值的计算公式为:
G
reverb=min(1,G
reverb_0+w
SpecRichnessG
SpecRichness+w
bgmG
bgm+w
vocalGoodnessG
vocalGoodness)
其中,G
reverb指代目标混响强度参数值,G
reverb_0指代预设的基础混响强度参数值,w
SpecRichness指代与G
SpecRichness对应的第一权重值,w
bgm指代与G
bgm对应的第二权重值,w
vocalGoodness指代与G
vocalGoodness对应的第三权重值。
在一些实施例中,上述三个权重值的取值依据对混响强度的影响大小来设置,比如第一权重值的取值最大,而第二权重值的取值最小,本公开实施例对此不进行具体限定。
在303中,基于目标混响强度参数值对采集到的人声信号进行混响处理。
在本公开实施例中,如图4所示,KTV混响算法中包括两层参数,一层为混响总增益,另一层为该混响算法内部的参数,进而能够通过直接控制混响部分的能量大小实现控制混响强度的目的。在一些实施例中,基于目标混响强度参数值对采集到的人声信号进行混响处理,包括但不限于:
基于目标混响强度参数值,对采集到的人声信号的混响总增益进行调整;或,基于目标混响强度参数值,对采集到的人声信号的至少一项混响算法参数进行调整。即,G
reverb既能够被作为混响总增益直接加载,也能够加载至该混响算法内部的一个或多个参数中,比如调整回声增益、延迟时间、反馈网络增益等,本公开实施例对此不进行具体限定。
在304中,对采集到的伴奏音频信号和经过混响处理后的人声信号进行混音处理,输出经过混音处理后的音频信号。
如图4所示,在经过KTV混响算法对人声信号进行处理后,会继续对采集到的伴奏音频信号和经过混响处理后的人声信号进行混音处理,而在经过混音处理后,便可直接输出音频信号,比如通过电子设备的扬声器播放经过混音处理后的音频信号,实现KTV音效。
本公开实施例针对不同节奏速度的乐曲、不同伴奏类型的乐曲、同一乐曲的不同部分、不同演唱者的乐曲,动态或预先计算出最适合的混响强度参数值,进而指导人工混响算法控制输出人声部分混响大小,从而达到自适应的K歌音效。
换一表达方式,本公开实施例综合考虑了乐曲的频域丰富程度、节奏速度以及演唱者等多方面因素,比如针对乐曲的频域丰富程度、节奏速度、演唱者,会自适应地产生不同的混响强度参数值,而对于各种影响混响强度的混响强度参数值,本公开实施例还提供了一种融合方式,最终得到总的混响强度参数值,而总的混响强度参数值既能够加载到混响总增益上,也能够加载到混响算法内部的一个或多个参数中,因此该种音频处理方式达到了自适应的K歌音效,使得电子设备输出的声音更加饱满和优美。
图8是根据一实施例示出的一种音频处理装置的框图。参照图8,该装置包括采集模块801,确定模块802和处理模块803。
采集模块801,被配置为采集当前待处理乐曲的伴奏音频信号和人声信号;
确定模块802,被配置为确定采集到的伴奏音频信号的目标混响强度参数值,所述目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;
处理模块803,被配置为基于所述目标混响强度参数值对采集到的人声信号进行混响处理。
本公开实施例提供的装置,在采集到当前待处理乐曲的伴奏音频信号和人声信号后,本公开实施例会确定采集到的伴奏音频信号的目标混响强度参数值,其中,目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;之后,基于目标混响强度参数值对采集到的人声信号进行混响处理。基于以上描述可知,本公开实施例考虑了乐曲的伴奏类型、节奏速度以及演唱者的演唱评分等多方面的因素,并据此自适应地生成当前待处理乐曲的混响强度参数值,达到了自适应的K歌音效,使得电子设备输出的声音更加饱满和优美。
在一些实施例中,确定模块802,还被配置为确定采集到的伴奏音频信号的第一混响强度参数值,所述第一混响强度参数值用于指示当前待处理乐曲的伴奏类型;确定采集到的伴奏音频信号的第二混响强度参数值,所述第二混响强度参数值用于指示当前待处理乐曲的节奏速度;确定采集到的伴奏音频信号的第三混响强度参数值,所述第三混响强度参数值用于指示当前待处理乐曲的演唱者的演唱评分;基于所述第一混响强度参数值、所述第二类混响强度参数值和所述第三类混响强度参数值,确定所述目标混响强度参数值。
在一些实施例中,确定模块802,还被配置为将采集到的伴奏音频信号由时域变换到时 频域,得到伴奏音频帧序列;获取每帧伴奏音频的幅度信息;基于每帧伴奏音频的幅度信息,确定每帧伴奏音频的频域丰富系数;其中,所述频域丰富系数用于指示每帧伴奏音频的幅度信息的频域丰富程度,所述频域丰富程度反映了当前待处理乐曲的伴奏类型;基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值。
在一些实施例中,确定模块802,还被配置为基于每帧伴奏音频的频域丰富系数,确定当前待处理乐曲的全局频域丰富系数;获取所述全局频域丰富系数与频域丰富系数最大值之间的第一比值,将所述第一比值和目标数值中的最小者确定为所述第一混响强度参数值。
在一些实施例中,确定模块802,还被配置为基于每帧伴奏音频的频域丰富系数,生成用于指示频域丰富程度的波形图;对生成的波形图进行平滑处理,基于平滑后的波形图确定当前待处理乐曲的不同部分的频域丰富系数;获取所述不同部分的频域丰富系数分别与频域丰富系数最大值之间的第二比值;对于获取到的每个第二比值,将所述第二比值和目标数值中的最小者确定为所述第一混响强度参数值。
在一些实施例中,确定模块802,还被配置为获取采集到的伴奏音频信号在规定时长的节拍数;确定获取到的节拍数与节拍数最大值之间的第三比值;将所述第三比值和目标数值中的最小者,确定为所述第二混响强度参数值。
在一些实施例中,确定模块802,还被配置为获取当前待处理乐曲的演唱者的音频演唱分值,基于所述音频演唱分值确定所述第三混响强度参数值。
在一些实施例中,确定模块802,还被配置为获取基础混响强度参数值、第一权重值、第二权重值以及第三权重值;确定所述第一权重值与所述第一混响强度参数值之间的第一和值;确定所述第二权重值与所述第二混响强度参数值之间的第二和值;确定所述第三权重值与所述第三混响强度参数值之间的第三和值;获取所述基础混响强度参数值、所述第一和值、所述第二和值与所述第三和值之间的第四和值,将所述第四比值和目标数值中的最小者,确定为所述目标混响强度参数值。
在一些实施例中,处理模块803,还被配置为基于所述目标混响强度参数值,对采集到的人声信号的混响总增益进行调整;或,基于所述目标混响强度参数值,对采集到的人声信号的至少一项混响算法参数进行调整。
在一些实施例中,处理模块803,还被配置为在对采集到的人声信号进行混响处理后,对采集到的伴奏音频信号和经过混响处理后的人声信号进行混音处理,输出经过混音处理后的音频信号。
上述所有可选技术方案,能够采用任意结合形成本公开的可选实施例,在此不再一一赘述。
关于上述实施例中的装置,其中各个模块执行操作的具体方式已经在有关该方法的实施例中进行了详细描述,此处将不做详细阐述说明。
图9示出了本公开一个实施例提供的一种电子设备900的结构框图。其中,该设备900是便携式移动终端,比如:智能手机、平板电脑、MP3播放器(Moving Picture Experts Group Audio Layer III,动态影像专家压缩标准音频层面3)、MP4(Moving Picture Experts Group Audio Layer IV,动态影像专家压缩标准音频层面4)播放器、笔记本电脑或台式电脑。设备900还可能被称为用户设备、便携式终端、膝上型终端、台式终端等其他名称。
通常,设备900包括有:处理器901和存储器902。
处理器901包括一个或多个处理核心,比如4核心处理器、8核心处理器等。处理器901采用DSP(Digital Signal Processing,数字信号处理)、FPGA(Field-Programmable Gate Array,现场可编程门阵列)、PLA(Programmable Logic Array,可编程逻辑阵列)中的至少一种硬件形式来实现。处理器901也包括主处理器和协处理器,主处理器是用于对在唤醒状态下的数据进行处理的处理器,也称CPU(Central Processing Unit,中央处理器);协处理器是用于对在待机状态下的数据进行处理的低功耗处理器。在一些实施例中,处理器901集成有GPU(Graphics Processing Unit,图像处理器),GPU用于负责显示屏所需要显示的内容的渲染和绘制。一些实施例中,处理器901还包括AI(Artificial Intelligence,人工智能)处理器,该AI处理器用于处理有关机器学习的计算操作。
存储器902包括一个或多个计算机可读存储介质,该计算机可读存储介质是非暂态的。存储器902还可包括高速随机存取存储器,以及非易失性存储器,比如一个或多个磁盘存储设备、闪存存储设备。
在一些实施例中,设备900还包括有:外围设备接口903和至少一个外围设备。处理器901、存储器902和外围设备接口903之间通过总线或信号线相连。各个外围设备通过总线、信号线或电路板与外围设备接口903相连。在一些实施例中,外围设备包括:射频电路904、显示屏905、摄像头组件906、音频电路907、定位组件908和电源909中的至少一种。
外围设备接口903可被用于将I/O(Input/Output,输入/输出)相关的至少一个外围设备连接到处理器901和存储器902。在一些实施例中,处理器901、存储器902和外围设备接口903被集成在同一芯片或电路板上;在一些其他实施例中,处理器901、存储器902和外围设备接口903中的任意一个或两个在单独的芯片或电路板上实现,本实施例对此不加以限定。
射频电路904用于接收和发射RF(Radio Frequency,射频)信号,也称电磁信号。射频电路904通过电磁信号与通信网络以及其他通信设备进行通信。射频电路904将电信号转 换为电磁信号进行发送,或者,将接收到的电磁信号转换为电信号。在一些实施例中,射频电路904包括:天线系统、RF收发器、一个或多个放大器、调谐器、振荡器、数字信号处理器、编解码芯片组、用户身份模块卡等等。射频电路904通过至少一种无线通信协议来与其它终端进行通信。该无线通信协议包括但不限于:万维网、城域网、内联网、各代移动通信网络(2G、3G、4G及5G)、无线局域网和/或WiFi(Wireless Fidelity,无线保真)网络。在一些实施例中,射频电路904还包括NFC(Near Field Communication,近距离无线通信)有关的电路,本公开对此不加以限定。
显示屏905用于显示UI(User Interface,用户界面)。该UI包括图形、文本、图标、视频及其它们的任意组合。当显示屏905是触摸显示屏时,显示屏905还具有采集在显示屏905的表面或表面上方的触摸信号的能力。该触摸信号作为控制信号输入至处理器901进行处理。此时,显示屏905还用于提供虚拟按钮和/或虚拟键盘,也称软按钮和/或软键盘。在一些实施例中,显示屏905为一个,设置在设备900的前面板;在另一些实施例中,显示屏905为至少两个,分别设置在设备900的不同表面或呈折叠设计;在另一些实施例中,显示屏905是柔性显示屏,设置在设备900的弯曲表面上或折叠面上。甚至,显示屏905还设置成非矩形的不规则图形,也即异形屏。显示屏905采用LCD(Liquid Crystal Display,液晶显示屏)、OLED(Organic Light-Emitting Diode,有机发光二极管)等材质制备。
摄像头组件906用于采集图像或视频。在一些实施例中,摄像头组件906包括前置摄像头和后置摄像头。通常,前置摄像头设置在终端的前面板,后置摄像头设置在终端的背面。在一些实施例中,后置摄像头为至少两个,分别为主摄像头、景深摄像头、广角摄像头、长焦摄像头中的任意一种,以实现主摄像头和景深摄像头融合实现背景虚化功能、主摄像头和广角摄像头融合实现全景拍摄以及VR(Virtual Reality,虚拟现实)拍摄功能或者其它融合拍摄功能。在一些实施例中,摄像头组件906还包括闪光灯。闪光灯是单色温闪光灯,也是双色温闪光灯。双色温闪光灯是指暖光闪光灯和冷光闪光灯的组合,用于不同色温下的光线补偿。
音频电路907包括麦克风和扬声器。麦克风用于采集用户及环境的声波,并将声波转换为电信号输入至处理器901进行处理,或者输入至射频电路904以实现语音通信。出于立体声采集或降噪的目的,麦克风为多个,分别设置在设备900的不同部位。麦克风还是阵列麦克风或全向采集型麦克风。扬声器则用于将来自处理器901或射频电路904的电信号转换为声波。扬声器是传统的薄膜扬声器,也是压电陶瓷扬声器。当扬声器是压电陶瓷扬声器时,不仅将电信号转换为人类可听见的声波,也将电信号转换为人类听不见的声波以进行测距等用途。在一些实施例中,音频电路907还包括耳机插孔。
定位组件908用于定位设备900的当前地理位置,以实现导航或LBS(Location Based Service,基于位置的服务)。定位组件908是基于美国的GPS(Global Positioning System,全球定位系统)、中国的北斗系统或俄罗斯的伽利略系统的定位组件。
电源909用于为设备900中的各个组件进行供电。电源909是交流电、直流电、一次性电池或可充电电池。当电源909包括可充电电池时,该可充电电池是有线充电电池或无线充电电池。有线充电电池是通过有线线路充电的电池,无线充电电池是通过无线线圈充电的电池。该可充电电池还用于支持快充技术。
在一些实施例中,设备900还包括有一个或多个传感器910。该一个或多个传感器910包括但不限于:加速度传感器911、陀螺仪传感器912、压力传感器913、指纹传感器914、光学传感器915以及接近传感器916。
加速度传感器911检测以设备900建立的坐标系的三个坐标轴上的加速度大小。比如,加速度传感器911用于检测重力加速度在三个坐标轴上的分量。处理器901根据加速度传感器911采集的重力加速度信号,控制显示屏905以横向视图或纵向视图进行用户界面的显示。加速度传感器911还用于游戏或者用户的运动数据的采集。
陀螺仪传感器912检测设备900的机体方向及转动角度,陀螺仪传感器912与加速度传感器911协同采集用户对设备900的3D动作。处理器901根据陀螺仪传感器912采集的数据,实现如下功能:动作感应(比如根据用户的倾斜操作来改变UI)、拍摄时的图像稳定、游戏控制以及惯性导航。
压力传感器913设置在设备900的侧边框和/或显示屏905的下层。当压力传感器913设置在设备900的侧边框时,检测用户对设备900的握持信号,由处理器901根据压力传感器913采集的握持信号进行左右手识别或快捷操作。当压力传感器913设置在显示屏905的下层时,由处理器901根据用户对显示屏905的压力操作,实现对UI界面上的可操作性控件进行控制。可操作性控件包括按钮控件、滚动条控件、图标控件、菜单控件中的至少一种。
指纹传感器914用于采集用户的指纹,由处理器901根据指纹传感器914采集到的指纹识别用户的身份,或者,由指纹传感器914根据采集到的指纹识别用户的身份。在识别出用户的身份为可信身份时,由处理器901授权该用户执行相关的敏感操作,该敏感操作包括解锁屏幕、查看加密信息、下载软件、支付及更改设置等。指纹传感器914被设置设备900的正面、背面或侧面。当设备900上设置有物理按键或厂商Logo时,指纹传感器914与物理按键或厂商Logo集成在一起。
光学传感器915用于采集环境光强度。在一个实施例中,处理器901根据光学传感器915采集的环境光强度,控制显示屏905的显示亮度。具体地,当环境光强度较高时,调高显示 屏905的显示亮度;当环境光强度较低时,调低显示屏905的显示亮度。在另一个实施例中,处理器901还根据光学传感器915采集的环境光强度,动态调整摄像头组件906的拍摄参数值。
接近传感器916,也称距离传感器,通常设置在设备900的前面板。接近传感器916用于采集用户与设备900的正面之间的距离。在一个实施例中,当接近传感器916检测到用户与设备900的正面之间的距离逐渐变小时,由处理器901控制显示屏905从亮屏状态切换为息屏状态;当接近传感器916检测到用户与设备900的正面之间的距离逐渐变大时,由处理器901控制显示屏905从息屏状态切换为亮屏状态。
图10是本公开的实施例提供的一种电子设备1000的结构框图。该设备1000表现为服务器。该服务器1000可因配置或性能不同而产生比较大的差异,包括一个或一个以上处理器(central processing units,CPU)1001和一个或一个以上的存储器1002。当然,该服务器还具有有线或无线网络接口、键盘以及输入输出接口等部件,以便进行输入输出,该服务器还包括其他用于实现设备功能的部件,在此不做赘述。
综上所述,本公开实施例提供了一种电子设备,包括:处理器;用于存储所述处理器可执行指令的存储器;其中,所述处理器被配置为执行所述指令,以实现如下步骤:采集当前待处理乐曲的伴奏音频信号和人声信号;确定采集到的伴奏音频信号的目标混响强度参数值,所述目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;基于所述目标混响强度参数值对采集到的人声信号进行混响处理。
在一些实施例中,所述处理器被配置为执行所述指令,以实现如下步骤:确定采集到的伴奏音频信号的第一混响强度参数值,所述第一混响强度参数值用于指示当前待处理乐曲的伴奏类型;确定采集到的伴奏音频信号的第二混响强度参数值,所述第二混响强度参数值用于指示当前待处理乐曲的节奏速度;确定采集到的伴奏音频信号的第三混响强度参数值,所述第三混响强度参数值用于指示当前待处理乐曲的演唱者的演唱评分;基于所述第一混响强度参数值、所述第二类混响强度参数值和所述第三类混响强度参数值,确定所述目标混响强度参数值。
在一些实施例中,所述处理器被配置为执行所述指令,以实现如下步骤:将采集到的伴奏音频信号由时域变换到时频域,得到伴奏音频帧序列;获取每帧伴奏音频的幅度信息;基于每帧伴奏音频的幅度信息,确定每帧伴奏音频的频域丰富系数;其中,所述频域丰富系数用于指示每帧伴奏音频的幅度信息的频域丰富程度,所述频域丰富程度反映了当前待处理乐曲的伴奏类型;基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值。
在一些实施例中,所述处理器被配置为执行所述指令,以实现如下步骤:基于每帧伴奏 音频的频域丰富系数,确定当前待处理乐曲的全局频域丰富系数;获取所述全局频域丰富系数与频域丰富系数最大值之间的第一比值,将所述第一比值和目标数值中的最小者确定为所述第一混响强度参数值。
在一些实施例中,所述处理器被配置为执行所述指令,以实现如下步骤:基于每帧伴奏音频的频域丰富系数,生成用于指示频域丰富程度的波形图;对生成的波形图进行平滑处理,基于平滑后的波形图确定当前待处理乐曲的不同部分的频域丰富系数;获取所述不同部分的频域丰富系数分别与频域丰富系数最大值之间的第二比值;对于获取到的每个第二比值,将所述第二比值和目标数值中的最小者确定为所述第一混响强度参数值。
在一些实施例中,所述处理器被配置为执行所述指令,以实现如下步骤:获取采集到的伴奏音频信号在规定时长的节拍数;确定获取到的节拍数与节拍数最大值之间的第三比值;将所述第三比值和目标数值中的最小者,确定为所述第二混响强度参数值。
在一些实施例中,所述处理器被配置为执行所述指令,以实现如下步骤:获取当前待处理乐曲的演唱者的音频演唱分值,基于所述音频演唱分值确定所述第三混响强度参数值。
在一些实施例中,所述处理器被配置为执行所述指令,以实现如下步骤:获取基础混响强度参数值、第一权重值、第二权重值以及第三权重值;确定所述第一权重值与所述第一混响强度参数值之间的第一和值;确定所述第二权重值与所述第二混响强度参数值之间的第二和值;确定所述第三权重值与所述第三混响强度参数值之间的第三和值;获取所述基础混响强度参数值、所述第一和值、所述第二和值与所述第三和值之间的第四和值,将所述第四比值和目标数值中的最小者,确定为所述目标混响强度参数值。
在一些实施例中,所述处理器被配置为执行所述指令,以实现如下步骤:基于所述目标混响强度参数值,对采集到的人声信号的混响总增益进行调整;或,基于所述目标混响强度参数值,对采集到的人声信号的至少一项混响算法参数进行调整。
在一些实施例中,所述处理器被配置为执行所述指令,以实现如下步骤:对采集到的伴奏音频信号和经过混响处理后的人声信号进行混音处理,输出经过混音处理后的音频信号。
本公开实施例还提供了一种包括指令的存储介质,例如包括指令的存储器,上述指令可由电子设备900或电子设备1000的处理器执行以完成上述音频处理方法。在一些实施例中,存储介质是非临时性计算机可读存储介质,例如,所述非临时性计算机可读存储介质是ROM、随机存取存储器(RAM)、CD-ROM、磁带、软盘和光数据存储设备等。
本公开实施例还提供了一种计算机程序产品,所述计算机程序产品中的指令由电子设备900或电子设备1000的处理器执行时,使得电子设备900或电子设备1000能够执行如上述方法实施例中的音频处理方法。
Claims (20)
- 一种音频处理方法,包括:采集当前待处理乐曲的伴奏音频信号和人声信号;确定采集到的伴奏音频信号的目标混响强度参数值,所述目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;基于所述目标混响强度参数值对采集到的人声信号进行混响处理。
- 根据权利要求1所述的音频处理方法,其中,所述确定采集到的伴奏音频信号的目标混响强度参数值,包括:确定采集到的伴奏音频信号的第一混响强度参数值,所述第一混响强度参数值用于指示当前待处理乐曲的伴奏类型;确定采集到的伴奏音频信号的第二混响强度参数值,所述第二混响强度参数值用于指示当前待处理乐曲的节奏速度;确定采集到的伴奏音频信号的第三混响强度参数值,所述第三混响强度参数值用于指示当前待处理乐曲的演唱者的演唱评分;基于所述第一混响强度参数值、所述第二类混响强度参数值和所述第三类混响强度参数值,确定所述目标混响强度参数值。
- 根据权利要求2所述的音频处理方法,其中,所述确定采集到的伴奏音频信号的第一混响强度参数值,包括:将采集到的伴奏音频信号由时域变换到时频域,得到伴奏音频帧序列;获取每帧伴奏音频的幅度信息;基于每帧伴奏音频的幅度信息,确定每帧伴奏音频的频域丰富系数;其中,所述频域丰富系数用于指示每帧伴奏音频的幅度信息的频域丰富程度,所述频域丰富程度反映了当前待处理乐曲的伴奏类型;基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值。
- 根据权利要求3所述的音频处理方法,其中,所述基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值,包括:基于每帧伴奏音频的频域丰富系数,确定当前待处理乐曲的全局频域丰富系数;获取所述全局频域丰富系数与频域丰富系数最大值之间的第一比值,将所述第一比值和目标数值中的最小者确定为所述第一混响强度参数值。
- 根据权利要求3所述的音频处理方法,其中,所述基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值,包括:基于每帧伴奏音频的频域丰富系数,生成用于指示频域丰富程度的波形图;对生成的波形图进行平滑处理,基于平滑后的波形图确定当前待处理乐曲的不同部分的频域丰富系数;获取所述不同部分的频域丰富系数分别与频域丰富系数最大值之间的第二比值;对于获取到的每个第二比值,将所述第二比值和目标数值中的最小者确定为所述第一混响强度参数值。
- 根据权利要求2所述的音频处理方法,其中,所述确定采集到的伴奏音频信号的第二混响强度参数值,包括:获取采集到的伴奏音频信号在规定时长的节拍数;确定获取到的节拍数与节拍数最大值之间的第三比值;将所述第三比值和目标数值中的最小者,确定为所述第二混响强度参数值。
- 根据权利要求2所述的音频处理方法,其中,所述确定采集到的伴奏音频信号的第三混响强度参数值,包括:获取当前待处理乐曲的演唱者的音频演唱分值,基于所述音频演唱分值确定所述第三混响强度参数值。
- 根据权利要求2所述的音频处理方法,其中,所述基于所述第一混响强度参数值、所述第二类混响强度参数值和所述第三类混响强度参数值,确定所述目标混响强度参数值,包括:获取基础混响强度参数值、第一权重值、第二权重值以及第三权重值;确定所述第一权重值与所述第一混响强度参数值之间的第一和值;确定所述第二权重值与所述第二混响强度参数值之间的第二和值;确定所述第三权重值与所述第三混响强度参数值之间的第三和值;获取所述基础混响强度参数值、所述第一和值、所述第二和值与所述第三和值之间的第四和值,将所述第四比值和目标数值中的最小者,确定为所述目标混响强度参数值。
- 根据权利要求1所述的音频处理方法,其中,所述基于所述目标混响强度参数值对采集到的人声信号进行混响处理,包括:基于所述目标混响强度参数值,对采集到的人声信号的混响总增益进行调整;或,基于所述目标混响强度参数值,对采集到的人声信号的至少一项混响算法参数进行调整。
- 根据权利要求1至9中任一项权利要求所述的音频处理方法,其中,所述方法还包括:对采集到的伴奏音频信号和经过混响处理后的人声信号进行混音处理,输出经过混音处理后的音频信号。
- 一种音频处理装置,包括:采集模块,被配置为采集当前待处理乐曲的伴奏音频信号和人声信号;确定模块,被配置为确定采集到的伴奏音频信号的目标混响强度参数值,所述目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;处理模块,被配置为基于所述目标混响强度参数值对采集到的人声信号进行混响处理。
- 一种电子设备,包括:处理器;用于存储所述处理器可执行指令的存储器;其中,所述处理器被配置为执行所述指令,以实现如下步骤:采集当前待处理乐曲的伴奏音频信号和人声信号;确定采集到的伴奏音频信号的目标混响强度参数值,所述目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;基于所述目标混响强度参数值对采集到的人声信号进行混响处理。
- 根据权利要求12所述的电子设备,其中,所述处理器被配置为执行所述指令,以实现如下步骤:确定采集到的伴奏音频信号的第一混响强度参数值,所述第一混响强度参数值用于指示当前待处理乐曲的伴奏类型;确定采集到的伴奏音频信号的第二混响强度参数值,所述第二混响强度参数值用于指示当前待处理乐曲的节奏速度;确定采集到的伴奏音频信号的第三混响强度参数值,所述第三混响强度参数值用于指示当前待处理乐曲的演唱者的演唱评分;基于所述第一混响强度参数值、所述第二类混响强度参数值和所述第三类混响强度参数值,确定所述目标混响强度参数值。
- 根据权利要求13所述的电子设备,其中,所述处理器被配置为执行所述指令,以实现如下步骤:将采集到的伴奏音频信号由时域变换到时频域,得到伴奏音频帧序列;获取每帧伴奏音频的幅度信息;基于每帧伴奏音频的幅度信息,确定每帧伴奏音频的频域丰富系数;其中,所述频域丰富系数用于指示每帧伴奏音频的幅度信息的频域丰富程度,所述频域丰富程度反映了当前待处理乐曲的伴奏类型;基于每帧伴奏音频的频域丰富系数确定所述第一混响强度参数值。
- 根据权利要求14所述的电子设备,其中,所述处理器被配置为执行所述指令,以实现如下步骤:基于每帧伴奏音频的频域丰富系数,确定当前待处理乐曲的全局频域丰富系数;获取所述全局频域丰富系数与频域丰富系数最大值之间的第一比值,将所述第一比值和目标数值中的最小者确定为所述第一混响强度参数值。
- 根据权利要求14所述的电子设备,其中,所述处理器被配置为执行所述指令,以实现如下步骤:基于每帧伴奏音频的频域丰富系数,生成用于指示频域丰富程度的波形图;对生成的波形图进行平滑处理,基于平滑后的波形图确定当前待处理乐曲的不同部分的频域丰富系数;获取所述不同部分的频域丰富系数分别与频域丰富系数最大值之间的第二比值;对于获取到的每个第二比值,将所述第二比值和目标数值中的最小者确定为所述第一混响强度参数值。
- 根据权利要求13所述的电子设备,其中,所述处理器被配置为执行所述指令,以实现如下步骤:获取采集到的伴奏音频信号在规定时长的节拍数;确定获取到的节拍数与节拍数最大值之间的第三比值;将所述第三比值和目标数值中的最小者,确定为所述第二混响强度参数值。
- 根据权利要求13所述的电子设备,其中,所述处理器被配置为执行所述指令,以实现如下步骤:获取当前待处理乐曲的演唱者的音频演唱分值,基于所述音频演唱分值确定所述第三混响强度参数值。
- 一种存储介质,所述存储介质中的指令由电子设备的处理器执行,使得电子设备能够执行如下步骤:采集当前待处理乐曲的伴奏音频信号和人声信号;确定采集到的伴奏音频信号的目标混响强度参数值,所述目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;基于所述目标混响强度参数值对采集到的人声信号进行混响处理。
- 一种计算机程序产品,所述计算机程序产品中的指令由电子设备的处理器执行,使得电子设备能够执行如下步骤:采集当前待处理乐曲的伴奏音频信号和人声信号;确定采集到的伴奏音频信号的目标混响强度参数值,所述目标混响强度参数值用于指示当前待处理乐曲的节奏速度、伴奏类型和演唱者的演唱评分中的至少一种;基于所述目标混响强度参数值对采集到的人声信号进行混响处理。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP21743735.9A EP4006897A4 (en) | 2020-01-22 | 2021-01-22 | AUDIO PROCESSING METHOD AND ELECTRONIC DEVICE |
| US17/702,416 US11636836B2 (en) | 2020-01-22 | 2022-03-23 | Method for processing audio and electronic device |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202010074552.2A CN111326132B (zh) | 2020-01-22 | 2020-01-22 | 音频处理方法、装置、存储介质及电子设备 |
| CN202010074552.2 | 2020-01-22 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/702,416 Continuation US11636836B2 (en) | 2020-01-22 | 2022-03-23 | Method for processing audio and electronic device |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021148009A1 true WO2021148009A1 (zh) | 2021-07-29 |
Family
ID=71172108
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/073380 Ceased WO2021148009A1 (zh) | 2020-01-22 | 2021-01-22 | 音频处理方法及电子设备 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US11636836B2 (zh) |
| EP (1) | EP4006897A4 (zh) |
| CN (1) | CN111326132B (zh) |
| WO (1) | WO2021148009A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20220215821A1 (en) * | 2020-01-22 | 2022-07-07 | Beijing Dajia Internet Information Technology Co., Ltd. | Method for processing audio and electronic device |
Families Citing this family (17)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN110047514B (zh) * | 2019-05-30 | 2021-05-28 | 腾讯音乐娱乐科技(深圳)有限公司 | 一种伴奏纯净度评估方法以及相关设备 |
| US12262082B1 (en) * | 2020-07-16 | 2025-03-25 | Apple Inc. | Audience reactive media |
| CN112216294B (zh) * | 2020-08-31 | 2024-03-19 | 北京达佳互联信息技术有限公司 | 音频处理方法、装置、电子设备及存储介质 |
| CN116437256A (zh) * | 2020-09-23 | 2023-07-14 | 华为技术有限公司 | 音频处理方法、计算机可读存储介质、及电子设备 |
| CN112365868B (zh) * | 2020-11-17 | 2024-05-28 | 北京达佳互联信息技术有限公司 | 声音处理方法、装置、电子设备及存储介质 |
| CN112435643B (zh) * | 2020-11-20 | 2024-07-19 | 腾讯音乐娱乐科技(深圳)有限公司 | 生成电音风格歌曲音频的方法、装置、设备及存储介质 |
| CN112669811B (zh) * | 2020-12-23 | 2024-02-23 | 腾讯音乐娱乐科技(深圳)有限公司 | 一种歌曲处理方法、装置、电子设备及可读存储介质 |
| CN112866732B (zh) * | 2020-12-30 | 2023-04-25 | 广州方硅信息技术有限公司 | 音乐广播方法及其装置、设备与介质 |
| CN112669797B (zh) * | 2020-12-30 | 2023-11-14 | 北京达佳互联信息技术有限公司 | 音频处理方法、装置、电子设备及存储介质 |
| CN112951265B (zh) * | 2021-01-27 | 2022-07-19 | 杭州网易云音乐科技有限公司 | 音频处理方法、装置、电子设备和存储介质 |
| CN112967705B (zh) * | 2021-02-24 | 2023-11-28 | 腾讯音乐娱乐科技(深圳)有限公司 | 一种混音歌曲生成方法、装置、设备及存储介质 |
| CN115942224A (zh) * | 2021-08-17 | 2023-04-07 | 上海艾为电子技术股份有限公司 | 声场扩展方法和系统、电子设备 |
| CN114449339B (zh) * | 2022-02-16 | 2024-04-12 | 深圳万兴软件有限公司 | 背景音效的转换方法、装置、计算机设备及存储介质 |
| CN114743527B (zh) * | 2022-04-21 | 2025-03-21 | 上海炉石信息科技有限公司 | 一种美声滤镜匹配方法 |
| CN114842820A (zh) * | 2022-05-18 | 2022-08-02 | 北京地平线信息技术有限公司 | K歌音频处理方法、装置及计算机可读存储介质 |
| CN115240709B (zh) * | 2022-07-25 | 2023-09-19 | 镁佳(北京)科技有限公司 | 一种音频文件的声场分析方法及装置 |
| CN115910098B (zh) * | 2022-11-02 | 2025-06-24 | 未鲲(上海)科技服务有限公司 | 基于变声识别的反诈预警方法、装置、电子设备及介质 |
Citations (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5502768A (en) * | 1992-09-28 | 1996-03-26 | Kabushiki Kaisha Kawai Gakki Seisakusho | Reverberator |
| US6091824A (en) * | 1997-09-26 | 2000-07-18 | Crystal Semiconductor Corporation | Reduced-memory early reflection and reverberation simulator and method |
| CN101454825A (zh) * | 2006-09-20 | 2009-06-10 | 哈曼国际工业有限公司 | 用于提取和改变输入信号的混响内容的方法和装置 |
| CN101609667A (zh) * | 2009-07-22 | 2009-12-23 | 福州瑞芯微电子有限公司 | Pmp播放器中实现卡拉ok功能的方法 |
| CN105654932A (zh) * | 2014-11-10 | 2016-06-08 | 乐视致新电子科技(天津)有限公司 | 实现卡拉ok应用的系统和方法 |
| CN108282712A (zh) * | 2018-02-06 | 2018-07-13 | 北京唱吧科技股份有限公司 | 一种麦克风 |
| CN109830244A (zh) * | 2019-01-21 | 2019-05-31 | 北京小唱科技有限公司 | 用于音频的动态混响处理方法及装置 |
| CN110211556A (zh) * | 2019-05-10 | 2019-09-06 | 北京字节跳动网络技术有限公司 | 音乐文件的处理方法、装置、终端及存储介质 |
| CN111326132A (zh) * | 2020-01-22 | 2020-06-23 | 北京达佳互联信息技术有限公司 | 音频处理方法、装置、存储介质及电子设备 |
Family Cites Families (19)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| KR100717324B1 (ko) * | 2005-11-01 | 2007-05-15 | 테크온팜 주식회사 | 휴대용 디지털음원 재생기를 이용한 노래방 시스템 |
| US9601127B2 (en) * | 2010-04-12 | 2017-03-21 | Smule, Inc. | Social music system and method with continuous, real-time pitch correction of vocal performance and dry vocal capture for subsequent re-rendering based on selectively applicable vocal effect(s) schedule(s) |
| US10930256B2 (en) * | 2010-04-12 | 2021-02-23 | Smule, Inc. | Social music system and method with continuous, real-time pitch correction of vocal performance and dry vocal capture for subsequent re-rendering based on selectively applicable vocal effect(s) schedule(s) |
| KR102246623B1 (ko) * | 2012-08-07 | 2021-04-29 | 스뮬, 인코포레이티드 | 선택적으로 적용가능한 보컬 효과 스케줄에 기초한 후속적 리렌더링을 위한 보컬 연주 및 드라이 보컬 캡쳐의 연속적인 실시간 피치 보정에 의한 소셜 음악 시스템 및 방법 |
| CN103295568B (zh) * | 2013-05-30 | 2015-10-14 | 小米科技有限责任公司 | 一种异步合唱方法和装置 |
| US9847078B2 (en) * | 2014-07-07 | 2017-12-19 | Sensibol Audio Technologies Pvt. Ltd. | Music performance system and method thereof |
| US10032443B2 (en) * | 2014-07-10 | 2018-07-24 | Rensselaer Polytechnic Institute | Interactive, expressive music accompaniment system |
| CN108040497B (zh) * | 2015-06-03 | 2022-03-04 | 思妙公司 | 用于自动产生协调的视听作品的方法和系统 |
| CN105161081B (zh) * | 2015-08-06 | 2019-06-04 | 蔡雨声 | 一种app哼唱作曲系统及其方法 |
| US9721551B2 (en) | 2015-09-29 | 2017-08-01 | Amper Music, Inc. | Machines, systems, processes for automated music composition and generation employing linguistic and/or graphical icon based musical experience descriptions |
| US9812105B2 (en) * | 2016-03-29 | 2017-11-07 | Mixed In Key Llc | Apparatus, method, and computer-readable storage medium for compensating for latency in musical collaboration |
| CN108305603B (zh) * | 2017-10-20 | 2021-07-27 | 腾讯科技(深圳)有限公司 | 音效处理方法及其设备、存储介质、服务器、音响终端 |
| CN108008930B (zh) * | 2017-11-30 | 2020-06-30 | 广州酷狗计算机科技有限公司 | 确定k歌分值的方法和装置 |
| CN108922506A (zh) * | 2018-06-29 | 2018-11-30 | 广州酷狗计算机科技有限公司 | 歌曲音频生成方法、装置和计算机可读存储介质 |
| CN108986842B (zh) * | 2018-08-14 | 2019-10-18 | 百度在线网络技术(北京)有限公司 | 音乐风格识别处理方法及终端 |
| CN109741723A (zh) * | 2018-12-29 | 2019-05-10 | 广州小鹏汽车科技有限公司 | 一种卡拉ok音效优化方法及卡拉ok装置 |
| CN109785820B (zh) * | 2019-03-01 | 2022-12-27 | 腾讯音乐娱乐科技(深圳)有限公司 | 一种处理方法、装置及设备 |
| CN109872710B (zh) * | 2019-03-13 | 2021-01-08 | 腾讯音乐娱乐科技(深圳)有限公司 | 音效调制方法、装置及存储介质 |
| CN110688082B (zh) * | 2019-10-10 | 2021-08-03 | 腾讯音乐娱乐科技(深圳)有限公司 | 确定音量的调节比例信息的方法、装置、设备及存储介质 |
-
2020
- 2020-01-22 CN CN202010074552.2A patent/CN111326132B/zh active Active
-
2021
- 2021-01-22 EP EP21743735.9A patent/EP4006897A4/en not_active Withdrawn
- 2021-01-22 WO PCT/CN2021/073380 patent/WO2021148009A1/zh not_active Ceased
-
2022
- 2022-03-23 US US17/702,416 patent/US11636836B2/en active Active
Patent Citations (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5502768A (en) * | 1992-09-28 | 1996-03-26 | Kabushiki Kaisha Kawai Gakki Seisakusho | Reverberator |
| US6091824A (en) * | 1997-09-26 | 2000-07-18 | Crystal Semiconductor Corporation | Reduced-memory early reflection and reverberation simulator and method |
| CN101454825A (zh) * | 2006-09-20 | 2009-06-10 | 哈曼国际工业有限公司 | 用于提取和改变输入信号的混响内容的方法和装置 |
| CN101609667A (zh) * | 2009-07-22 | 2009-12-23 | 福州瑞芯微电子有限公司 | Pmp播放器中实现卡拉ok功能的方法 |
| CN105654932A (zh) * | 2014-11-10 | 2016-06-08 | 乐视致新电子科技(天津)有限公司 | 实现卡拉ok应用的系统和方法 |
| CN108282712A (zh) * | 2018-02-06 | 2018-07-13 | 北京唱吧科技股份有限公司 | 一种麦克风 |
| CN109830244A (zh) * | 2019-01-21 | 2019-05-31 | 北京小唱科技有限公司 | 用于音频的动态混响处理方法及装置 |
| CN110211556A (zh) * | 2019-05-10 | 2019-09-06 | 北京字节跳动网络技术有限公司 | 音乐文件的处理方法、装置、终端及存储介质 |
| CN111326132A (zh) * | 2020-01-22 | 2020-06-23 | 北京达佳互联信息技术有限公司 | 音频处理方法、装置、存储介质及电子设备 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4006897A4 * |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20220215821A1 (en) * | 2020-01-22 | 2022-07-07 | Beijing Dajia Internet Information Technology Co., Ltd. | Method for processing audio and electronic device |
| US11636836B2 (en) * | 2020-01-22 | 2023-04-25 | Beijing Dajia Internet Information Technology Co., Ltd. | Method for processing audio and electronic device |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4006897A1 (en) | 2022-06-01 |
| US20220215821A1 (en) | 2022-07-07 |
| US11636836B2 (en) | 2023-04-25 |
| EP4006897A4 (en) | 2022-12-21 |
| CN111326132B (zh) | 2021-10-22 |
| CN111326132A (zh) | 2020-06-23 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11636836B2 (en) | Method for processing audio and electronic device | |
| CN110688082B (zh) | 确定音量的调节比例信息的方法、装置、设备及存储介质 | |
| CN108008930B (zh) | 确定k歌分值的方法和装置 | |
| CN110956971B (zh) | 音频处理方法、装置、终端及存储介质 | |
| WO2020103550A1 (zh) | 音频信号的评分方法、装置、终端设备及计算机存储介质 | |
| CN111128232B (zh) | 音乐的小节信息确定方法、装置、存储介质及设备 | |
| CN110867194B (zh) | 音频的评分方法、装置、设备及存储介质 | |
| CN111753125A (zh) | 歌曲音频显示的方法和装置 | |
| CN112435643B (zh) | 生成电音风格歌曲音频的方法、装置、设备及存储介质 | |
| WO2022111168A1 (zh) | 视频的分类方法和装置 | |
| CN113963707B (zh) | 音频处理方法、装置、设备和存储介质 | |
| CN112086102B (zh) | 扩展音频频带的方法、装置、设备以及存储介质 | |
| WO2021139535A1 (zh) | 播放音频的方法、装置、系统、设备及存储介质 | |
| CN111984222B (zh) | 调节音量的方法、装置、电子设备及可读存储介质 | |
| CN109192223B (zh) | 音频对齐的方法和装置 | |
| CN112992107B (zh) | 训练声学转换模型的方法、终端及存储介质 | |
| CN113257222B (zh) | 合成歌曲音频的方法、终端及存储介质 | |
| WO2023061330A1 (zh) | 音频合成方法、装置、设备及计算机可读存储介质 | |
| CN112597331B (zh) | 显示音域匹配信息的方法、装置、设备和存储介质 | |
| CN111063364B (zh) | 生成音频的方法、装置、计算机设备和存储介质 | |
| CN113450823A (zh) | 基于音频的场景识别方法、装置、设备及存储介质 | |
| CN115767117B (zh) | 直播互动操作的方法、设备和存储介质 | |
| CN117496923A (zh) | 歌曲生成方法、装置、设备及存储介质 | |
| CN113192531B (zh) | 检测音频是否是纯音乐音频方法、终端及存储介质 | |
| CN114329001B (zh) | 动态图片的显示方法、装置、电子设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21743735 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2021743735 Country of ref document: EP Effective date: 20220228 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
