WO2024152802A1 - 视频的编辑方法、装置、电子设备和存储介质 - Google Patents

视频的编辑方法、装置、电子设备和存储介质 Download PDF

Info

Publication number
WO2024152802A1
WO2024152802A1 PCT/CN2023/138239 CN2023138239W WO2024152802A1 WO 2024152802 A1 WO2024152802 A1 WO 2024152802A1 CN 2023138239 W CN2023138239 W CN 2023138239W WO 2024152802 A1 WO2024152802 A1 WO 2024152802A1
Authority
WO
WIPO (PCT)
Prior art keywords
text
target video
invalid
video
timeline
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2023/138239
Other languages
English (en)
French (fr)
Inventor
曾祥瑞
冀佳玉
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Zitiao Network Technology Co Ltd
Original Assignee
Beijing Zitiao Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Zitiao Network Technology Co Ltd filed Critical Beijing Zitiao Network Technology Co Ltd
Priority to JP2023578854A priority Critical patent/JP7680578B2/ja
Priority to EP23821101.5A priority patent/EP4429256A4/en
Priority to KR1020257018467A priority patent/KR20250105651A/ko
Priority to US18/542,524 priority patent/US12051445B1/en
Publication of WO2024152802A1 publication Critical patent/WO2024152802A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/04Segmentation; Word boundary detection
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/08Speech classification or search
    • G10L15/18Speech classification or search using natural language modelling
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/26Speech to text systems
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/439Processing of audio elementary streams
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/845Structuring of content, e.g. decomposing content into time segments
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/845Structuring of content, e.g. decomposing content into time segments
    • H04N21/8456Structuring of content, e.g. decomposing content into time segments by decomposing the content in the time domain, e.g. in time segments

Definitions

  • the embodiments of the present disclosure relate to the field of computer technology, and in particular, to a video editing method, device, electronic device, and storage medium.
  • voice-over videos For example, users can shoot and upload voice-over videos to video applications.
  • the voice-over or audio carries the core content of the voice-over video. Therefore, how to better edit the voice-over video is an urgent problem that needs to be solved.
  • the editing of spoken video mainly needs to deal with the pauses, stuttering, slurred words and other unnecessary segments in the video.
  • the existing editing method generally identifies invalid segments through the waveform of the audio in the video, but this method takes a lot of time and has low editing efficiency.
  • the embodiments of the present disclosure provide a video editing method, device, electronic device and storage medium to improve the efficiency of video editing.
  • an embodiment of the present disclosure provides a method for editing a video, comprising:
  • the video segment on the editing track segment that is within the timeline interval of the invalid text is deleted from the target video.
  • the present disclosure also provides a video editing device, including:
  • a text determination module used to determine invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video, wherein the timeline position of the invalid text is used to indicate the appearance time of the speech audio of the invalid text in the target video;
  • An interval identification module is used to display the editing track segment of the target video on the editing interface of the target video, and identify the timeline interval of the invalid text on the editing track segment based on the timeline position of the invalid text, wherein the timeline interval of the invalid text is the time interval in which the voice audio of the invalid text appears in the target video;
  • An interval adjustment module configured to adjust the timeline interval of the invalid text on the editing track segment of the target video in response to an adjustment operation of the invalid text
  • the segment deletion module is used to delete the video segment on the editing track segment within the timeline interval of the invalid text from the target video in response to the video segment deletion operation of the invalid text.
  • an embodiment of the present disclosure further provides an electronic device, including:
  • processors one or more processors
  • a memory for storing one or more programs
  • the one or more processors When the one or more programs are executed by the one or more processors, the one or more processors implement the video editing method as described in the embodiment of the present disclosure.
  • the embodiments of the present disclosure further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video editing method as described in the embodiments of the present disclosure.
  • the embodiments of the present disclosure provide a method, device, electronic device and storage medium for editing a video, the method comprising: determining invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video, the timeline position of the invalid text being used to indicate the appearance time of the speech audio of the invalid text in the target video; displaying an editing track segment of the target video on an editing interface of the target video, and marking a timeline interval of the invalid text on the editing track segment based on the timeline position of the invalid text, the timeline interval of the invalid text being the time interval when the speech audio of the invalid text appears in the target video; in response to an adjustment operation of the invalid text, adjusting the timeline interval of the invalid text on the editing track segment of the target video; and in response to a video segment deletion operation of the invalid text, deleting the video segment on the editing track segment that is within the timeline interval of the invalid text from the target video.
  • FIG1 is a schematic flow chart of a video editing method provided by an embodiment of the present disclosure
  • FIG2 is a schematic diagram of triggering speech recognition provided by an embodiment of the present disclosure
  • FIG3 is a schematic diagram of an editing interface provided by an embodiment of the present disclosure.
  • FIG4 is a schematic flow chart of another video editing method provided by an embodiment of the present disclosure.
  • FIG5 is a schematic diagram of another editing interface provided by an embodiment of the present disclosure.
  • FIG6 is a schematic diagram of updating a target text provided by an embodiment of the present disclosure.
  • FIG7 is a schematic diagram of another method for updating a target text provided by an embodiment of the present disclosure.
  • FIG8 is a schematic diagram of another method for updating a target text provided by an embodiment of the present disclosure.
  • FIG9 is a schematic structural diagram of a video editing device provided by an embodiment of the present disclosure.
  • FIG. 10 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure.
  • FIG1 is a flow chart of a method for editing a video provided in an embodiment of the present disclosure.
  • the method is applicable to video editing.
  • the method can be executed by a video editing device, wherein the device can be implemented by software and/or hardware and is generally integrated on an electronic device.
  • the electronic device includes but is not limited to: mobile phones, computers and other devices.
  • a video editing method provided by an embodiment of the present disclosure includes the following steps:
  • S110 Determine invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video.
  • the timeline position of the invalid text is used to indicate the appearance time of the speech audio of the invalid text in the target video.
  • the target video may refer to a video that needs to be edited, such as a video uploaded by a user or a video shot online by a user.
  • the target video contains audio
  • the voice text may be considered as text information corresponding to the audio of the target video, which may be obtained by performing voice recognition on the audio of the target video.
  • the voice text includes invalid text
  • the invalid text may include text corresponding to the pauses, repetitions, and unnecessary words in the target video that the user does not need; the timeline position of the invalid text may be used to indicate the appearance time of the voice audio of the invalid text in the target video.
  • the target video is a continuous video
  • the target video may be a continuous and complete video, or may be a segment of a complete video.
  • the invalid text and the timeline position of the invalid text in the speech text of the target video can be determined by performing speech recognition on the audio in the target video.
  • This step does not limit the specific method of voice recognition and the timing of triggering voice recognition.
  • a certain recognition control can be triggered in the interface to trigger voice recognition of the audio in the target video, so that invalid text in the target video and the timeline position of the invalid text can be automatically identified.
  • the position and style of the recognition control are not limited and can be set according to the actual page situation.
  • the identification conditions of invalid text are: a segment with a default silence duration greater than a preset duration (such as 120ms, etc.) is a pause segment; invalid words (such as repetitions and modal particles) are identified by reusing a preset algorithm logic, and the preset algorithm can be configured in the current application or in the server.
  • a preset duration such as 120ms, etc.
  • FIG2 is a schematic diagram of triggering voice recognition provided by an embodiment of the present disclosure. As shown in FIG2, after selecting a video (such as a target video), you can trigger control 1 in the pop-up window by right-clicking to trigger voice recognition of the audio in the target video. You can also click the shortcut control 2 in the current interface to realize voice recognition of the audio in the target video.
  • a video such as a target video
  • the editing interface can be used to edit clips in the target video.
  • the editing track clip of the target video can be understood as the editing track clip corresponding to the target video, such as an editing track clip whose starting point corresponds to the starting point of the target video and whose end point corresponds to the end point of the target video.
  • the editing track clip can be the editing track of the continuous and complete video; when the target video is a clip of a complete video, the editing track clip can be the track clip in the editing track of the complete video that corresponds to the target video.
  • the editing track clip is the audio track clip of the target video, and the audio track clip can refer to the audio track clip corresponding to the target video.
  • the timeline interval of invalid text can be considered as the time interval in which the voice audio of the invalid text appears in the target video.
  • the invalid text can be displayed in the target video.
  • the editing track segment of the target video is displayed on the editing interface, and the timeline interval of the invalid text is marked on the editing track segment based on the timeline position of the invalid text.
  • a preset mark can be used to mark the timeline interval of the invalid text in the editing track segment of the target video.
  • the preset mark can be specifically set according to, for example, the preset mark can be a rectangular frame.
  • a rectangular frame can be used to frame the timeline interval of the invalid text in the editing track segment of the target video; or, the timeline interval of the invalid text can be marked as a first display style in the editing track segment, and the timeline interval of other texts can be marked as a second display style in the editing track segment, and the first display style is different from the second display style, such as the first display style and the second display style can be distinguished by different colors, or can be distinguished by different lines, etc.
  • the timeline intervals of all invalid texts identified can be marked on the editing track segment by default to distinguish the timeline intervals of invalid texts from the timeline intervals of other texts; or a set number of timeline intervals of invalid texts can be marked on the editing track segment by default, and the set number can be pre-set by relevant personnel. Users can adjust the timeline intervals of invalid texts as needed.
  • Figure 3 is a schematic diagram of an editing interface provided by an embodiment of the present disclosure.
  • the editing track segment 3 of the target video can be displayed on the editing interface of the target video, and based on the timeline position of the invalid text identified in the previous step, the timeline interval of the invalid text is marked on the editing track segment 3, as shown in Figure 3, the timeline interval 4 of the invalid text is marked with a black rectangular frame.
  • the invalid text adjustment operation may refer to an operation for triggering the adjustment of the timeline interval of the invalid text.
  • the specific operation method of the invalid text adjustment operation in this embodiment is not limited.
  • the invalid text adjustment operation may be an operation of clicking a certain control in the editing interface or a preset triggering operation acting on the editing track segment, or an operation of performing a preset gesture in the editing interface.
  • the preset gesture may be a pre-set gesture, such as a preset gesture. For a right swipe gesture, etc.
  • This embodiment does not limit the triggering position of the adjustment operation. For example, it can act on the timeline interval of invalid text, or it can act on other positions, as long as it can trigger the adjustment of the timeline interval of invalid text.
  • the present embodiment can adjust the timeline interval of the invalid text on the editing track segment of the target video in response to the adjustment operation of the invalid text.
  • the specific method of adjusting the timeline interval may vary according to the different adjustment operations. For example, when the adjustment operation of the invalid text is to trigger a control to change the state of the invalid text from a selected state to an unselected state, the number of timeline intervals of the invalid text can be adjusted on the editing track segment of the target video, such as adding a timeline interval marking a certain text on the editing track segment to adjust the text to invalid text, or canceling the time interval marking a certain text on the editing track segment to adjust the text to non-invalid text.
  • the invalid text video clip deletion operation may refer to an operation for triggering deletion of the video clip corresponding to the invalid text on the editing track segment.
  • the invalid text video clip deletion operation may be an operation of clicking a deletion control in the editing interface.
  • This step can respond to the video segment deletion operation of the invalid text, and delete the video segment in the timeline interval of the invalid text on the editing track segment from the target video.
  • the specific deletion method is not limited.
  • the trace corresponding to the timeline interval of the invalid text on the editing track segment can be retained or removed according to the user's needs.
  • the user triggering control 5 can be considered as a video segment deletion operation.
  • the electronic device can respond to the video segment deletion operation of the invalid text and delete the video segment corresponding to the timeline interval 6 of the invalid text from the target video.
  • deleting the video segment on the editing track segment within the timeline interval of the invalid text from the target video includes:
  • the video segment on the editing track segment that is located in the timeline interval of the invalid text is deleted from the target video, and the deleted video segment is used as the segmentation point, Splitting the target video into multiple independent video segments; or
  • the video segment on the editing track segment that is located in the timeline interval of the invalid text is deleted from the target video, and the target video after the video segment is deleted is integrated into one video.
  • the video segment located in the timeline interval of the invalid text on the editing track segment can be deleted from the target video, and the target video can be divided into multiple independent video segments with the deleted video segment as the division point. On this basis, the user can operate each independent video segment as needed.
  • the video segments in the timeline interval of the invalid text on the editing track segment can be deleted from the target video, and the target video after the video segments are deleted can be integrated into one video.
  • the operation of deleting the invalid segments in the target video is realized, and the quality of the target video is improved.
  • the invalid segment is the segment corresponding to the timeline interval of the invalid text, in other words, the text corresponding to the timeline interval marked on the editing track segment in the speech text of the target video.
  • the embodiment of the present disclosure provides a video editing method, which determines invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video, wherein the timeline position of the invalid text is used to indicate the appearance time of the speech audio of the invalid text in the target video; displays the editing track segment of the target video on the editing interface of the target video, and identifies the timeline interval of the invalid text on the editing track segment based on the timeline position of the invalid text, wherein the timeline interval of the invalid text is the time interval when the speech audio of the invalid text appears in the target video; in response to the adjustment operation of the invalid text, adjusts the timeline interval of the invalid text on the editing track segment of the target video; in response to the video segment deletion operation of the invalid text, deletes the video segment on the editing track segment that is located in the timeline interval of the invalid text from the target video.
  • the invalid text and the position of the invalid text in the target video can be determined by performing speech recognition on the audio in the target video, and then by identifying the timeline interval of the invalid text on the editing track segment of the target video, the invalid text can be deleted in response to the user's adjustment operation on the invalid text and the video segment deletion operation.
  • the invalid segments are deleted through the deletion operation, thereby completing the video editing and improving the efficiency of video editing.
  • the method further comprises:
  • the playback progress of the target video is adjusted to the first time node, and the target video continues to be played from the first time node, wherein the second timeline interval is the timeline interval of the invalid text, and the first time node is the time node corresponding to the end point of the second timeline interval in the target video.
  • the video area can be considered as a certain area in the editing interface, which is used to play the target video.
  • the specific position and size of the video area can be set by relevant personnel according to the situation of the editing interface.
  • the second timeline interval can be any timeline interval of invalid text, and the first time node can be understood as the time node corresponding to the end position of the second timeline interval in the target video.
  • the target video can be played in the video area of the editing interface.
  • the timing of playing the target video is not limited.
  • the target video can be played automatically when the editing interface is displayed; or after the editing interface is displayed, when the user triggers a certain playback control in the editing interface, the target video can be played in the video area of the editing interface; or the target video can be automatically played after the editing interface is displayed for a preset length of time, etc. This embodiment does not limit this.
  • the present embodiment can automatically adjust the play progress of the target video to the first time node, and continue to play the target video from the first time node. On this basis, by automatically skipping invalid segments when playing the target video, the user can watch the edited video in advance to better edit the target video.
  • the method after playing the target video in the video area of the editing interface, the method further includes:
  • the fifth text is displayed in the text area of the editing interface, the fifth text is invalid text or non-invalid text, and the second time node is the time node corresponding to the starting position of the fifth text in the target video.
  • the fifth text may be a text triggered by a user, which may be an invalid text or a non-invalid text.
  • Non-invalid text is text other than invalid text, such as valid text in the voice text of the target video.
  • the fifth text may be displayed in a text area of the editing interface, and the text area may be considered as a certain area in the editing interface for displaying the voice text of the target video, such as the fifth text in the voice text for displaying the target video.
  • the specific position and size of the text area may be set by relevant personnel according to the situation of the editing interface.
  • the invalid text in the text area is formed into a paragraph by itself, for example, the "pause" text can be formed into a separate line, and the pause duration is displayed at the same time.
  • the trigger operation on the fifth text can be used to trigger the adjustment of the playback progress of the target video to the second time node.
  • the trigger operation on the fifth text can be an operation of clicking the fifth text; the second time node is the time node corresponding to the starting position of the fifth text in the target video.
  • the target video's playback progress can be adjusted to the second time node in response to the triggering operation of the fifth text, and the target video can be played continuously from the second time node.
  • the playback progress of the target video is adjusted, thereby locating the corresponding text in the target video, so that the user can watch the video clip corresponding to the corresponding text.
  • a time mark is also displayed on the editing track segment, and after the target video is played in the video area of the editing interface, the method further includes:
  • the playback progress of the target video is adjusted to the third time node, and the target video is played from the third time node.
  • the target video continues to be played at the node, wherein the third time node is the time node corresponding to the starting position or the ending position of the third timeline interval in the target video, or the time node to which the target video was last played before the position progress adjustment operation is received;
  • the playback progress of the target video is adjusted to a fourth time node, and the target video continues to be played from the fourth time node, wherein the fourth time node is the time node corresponding to the target position in the target video.
  • the time identifier is used to represent the time node to which the target video is currently played in the editing track segment of the target video. In other words, this time identifier can be used to indicate the time node corresponding to the current playback progress of the target video.
  • the position of the target identifier in the editing track segment of the target video can be updated in real time according to the playback progress of the target video.
  • the position adjustment operation of the target identifier can be understood as an operation to adjust the position of the target identifier, such as the position adjustment operation of the target identifier can be an operation of clicking on a certain position a in the editing track segment to adjust the target identifier to the position a, and the position adjustment operation of the target identifier can also be an operation of dragging the target identifier to adjust the target identifier to the position a, etc.
  • the target position can be considered as the position of the target identifier on the editing track segment when the position adjustment operation is completed, and the third timeline interval is the timeline interval where the target position of the target identifier is located.
  • the third time node can be the time node corresponding to the starting position of the third timeline interval in the target video, or the time node corresponding to the end position of the third timeline interval in the target video, or it can be the time node to which the target video was last played before receiving the position progress adjustment operation; the fourth time node can be considered as the time node corresponding to the target position in the target video.
  • the position of the target marker may be adjusted in response to the position adjustment operation of the target marker.
  • the playback progress of the target video can be adjusted according to the position of the time marker in the editing track segment.
  • the third timeline interval where the target position is located is determined, and the playback progress of the target video is adjusted differently according to different third timeline intervals.
  • the playback progress of the target video can be adjusted to the fourth time node, and the target video can continue to be played from the fourth time node.
  • this embodiment automatically adjusts the playback progress of the target video to the end position of the timeline interval A to continue playing the target video from the end position of the timeline interval A, or restores the position of the target identifier to the last position of the target identifier before receiving the position adjustment operation, and continues to play the target video from this position.
  • this embodiment automatically adjusts the playback progress of the target video to the starting position of the timeline interval A, and continues to play the target video from the starting position of the timeline interval A.
  • the play progress of the target video can be adjusted to the time node corresponding to the target position in the target video, and the target video can be continued to be played from this time node.
  • the target position b to which the target identifier is dragged is not located in the timeline interval of invalid text, or when the target position b to which the target identifier is moved by clicking is not located in the timeline interval of invalid text, this embodiment can adjust the play progress of the target video to the time node c corresponding to the target position b in the target video, and continue to play the target video from time node c.
  • the video segment corresponding to the invalid text can be automatically skipped or the video segment corresponding to the invalid text can be played completely from the starting position; and when the position of the target identifier is adjusted not to an invalid segment, the function of playing the target video from the time node corresponding to this position.
  • FIG4 is a flow chart of another video editing method provided by an embodiment of the present disclosure.
  • the solution in this embodiment can be combined with one or more optional solutions in the above embodiments.
  • the target The timeline interval of the invalid text is adjusted on the editing track segment of the video, including at least one of the following: in response to a first adjustment operation of the invalid text, the length of the timeline interval of the invalid text is adjusted on the editing track segment of the target video; in response to a second adjustment operation of the invalid text, the number of timeline intervals marked on the editing track segment of the target video is adjusted.
  • the method includes:
  • S210 Determine invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video.
  • the timeline position of the invalid text is used to indicate the appearance time of the speech audio of the invalid text in the target video.
  • the first adjustment operation can be used to adjust the length of the timeline interval of the invalid text.
  • the first adjustment operation can be performed on the editing track segment of the target video.
  • the first adjustment operation can be an operation of stretching or compressing the timeline interval B of a certain invalid text on the editing track segment of the target video, such as dragging the starting position of the timeline interval B to the left.
  • This step can respond to the first adjustment operation of the invalid text and adjust the length of the timeline interval of the invalid text on the editing track segment of the target video. For example, when the left border of a timeline interval is dragged to the left or the right border of a timeline interval is dragged to the right, the interval length of the timeline interval can be extended; when the left border of a timeline interval is dragged to the right or the right border of a timeline interval is dragged to the right, the interval length of the timeline interval can be shortened.
  • the second adjustment operation can be used to adjust the number of timeline intervals marked on the editing track segment.
  • the second adjustment operation can act on a certain adjustment control.
  • Each invalid text can correspond to an adjustment control, and the adjustment control is used to control whether the timeline interval corresponding to the invalid text is marked on the editing track segment.
  • the second adjustment operation can also act on the timeline interval of the editing track segment.
  • This step can adjust the number of timeline intervals marked on the editing track segment of the target video in response to the second adjustment operation of the invalid text.
  • This embodiment does not limit the specific process of the adjustment, as long as the number of timeline intervals marked on the editing track segment can be adjusted.
  • a video editing method provided by an embodiment of the present disclosure can adjust the length and number of timeline intervals of invalid text on an editing track segment of a target video in response to a first adjustment operation and a second adjustment operation of the invalid text, thereby enabling accurate editing of the target video according to user needs.
  • displaying the editing track segment of the target video on the editing interface of the target video includes:
  • the editing track segment of the target video is displayed in the track area of the editing interface, and the invalid text and non-invalid text in the voice text are displayed in the text area of the editing interface, wherein the invalid text is displayed as a selected state on the editing interface, and the non-invalid text is displayed as a non-selected state on the editing interface.
  • the track area can be an area in the editing interface, which is used to display the editing track segment of the target video.
  • the selected state can be understood as a text being in a selected state, thereby indicating that the text will be deleted; the unselected state can be understood as a text being in an unselected state, thereby indicating that the text will not be deleted.
  • the specific process of displaying the editing track segment of the target video on the editing interface of the target video may be: displaying the editing track segment of the target video in the track area of the editing interface, and displaying the non-text in the voice text in the text area of the editing interface.
  • the invalid text and non-invalid text are displayed, among which the invalid text can be displayed as selected on the editing interface, and the non-invalid text can be displayed as unselected on the editing interface.
  • the invalid text and non-invalid text in the voice text are displayed in the text area of the editing interface while displaying the editing track fragment; in addition, the invalid text is displayed as selected by default, which can prompt the user to delete the invalid text and avoid the user from performing the corresponding selection operation to select the invalid text such as pause and repeat before deleting it, further simplifying the operation required for the user to delete the invalid text in the target video.
  • Figure 5 is a schematic diagram of another editing interface provided by an embodiment of the present disclosure.
  • the editing track segment of the target video is displayed in the track area 7 of the editing interface, and the invalid text (such as "pause 0.12s" or “repeat") and non-invalid text (such as "mine”) in the voice text are displayed in the text area 8 of the editing interface, wherein the invalid text is displayed as selected on the editing interface, and the non-invalid text is displayed as unselected on the editing interface.
  • the invalid text such as "pause 0.12s" or "repeat
  • non-invalid text such as "mine
  • the adjusting the length of the timeline interval of the invalid text on the editing track segment of the target video includes:
  • the length of a first timeline interval of a first text is adjusted on an editing track segment of the target video, and the target text displayed in the text area is updated, wherein the target text includes the first text and/or a second text associated with the first text.
  • the first text may be an invalid text, such as the text belonging to the timeline interval targeted by the first trigger operation.
  • the first timeline interval is the timeline interval of the first text.
  • the target text may refer to the text that needs to be adjusted, such as the target text may include the first text and/or the second text associated with the first text, the second text may be the text corresponding to the timeline interval adjacent to the timeline interval of the first text, or the text corresponding to the new timeline interval after the timeline interval adjacent to the timeline interval of the first text is reduced.
  • the target text displayed in the text area needs to be updated, and the specific content of the update can be determined by the specific operation of the first adjustment operation.
  • the text corresponding to the adjusted part can be updated in real time in the text area to accurately adjust the length of the timeline interval corresponding to the invalid text.
  • the first adjustment operation includes a length extension operation
  • the updating of the target text displayed in the text area includes:
  • the text content in the second text corresponding to the extended portion of the first timeline interval is moved to the first text.
  • the length extension operation may be considered as an operation of extending the length of the first timeline interval.
  • the length extension operation may be an operation of dragging the start position of the first timeline interval to the left or dragging the end position of the first timeline interval to the right.
  • the text content in the second text corresponding to the extended portion of the first timeline interval can be synchronously moved to the first text in the text area.
  • the first adjustment operation is an operation of dragging the starting position of the first timeline interval to the left
  • the second text is the text corresponding to the timeline interval located to the left of the first timeline interval. At this time, the text content in the second text corresponding to the extended portion of the first timeline interval needs to be moved to the first text.
  • the length extension operation may support moving part of the adjacent non-invalid text into the first text.
  • the length extension operation may be supported or not supported.
  • Figure 6 is a schematic diagram of updating a target text provided by an embodiment of the present disclosure.
  • the end position of the first timeline interval corresponding to the first text 9 is dragged rightward from position A to position B, the length of the first timeline interval of the first text 9 can be extended on the editing track segment of the target video, and the text content (such as "sun") in the second text 10 corresponding to the extended part of the first timeline interval can be moved to the first text 9.
  • the first adjustment operation includes a length shortening operation
  • the updating of the target text displayed in the text area includes:
  • a second text is added to the text area, and text content in the first text corresponding to the shortened portion of the first timeline interval is moved to the second text.
  • the length shortening operation may be considered as an operation of shortening the length of the first timeline interval.
  • the length shortening operation may be an operation of dragging the start position of the first timeline interval to the right or dragging the end position of the first timeline interval to the left.
  • the update method of the target text displayed in the text area can be determined according to the different contents of the first text. For example, when the first text content is a pause, the text content in the first text corresponding to the shortened portion of the first timeline interval can be moved to the second text; and when the first text content is repeated text, invalid words, or normal text other than pauses, repetitions and invalid words, a second text can be added to the text area, and the text content in the first text corresponding to the shortened portion of the first timeline interval can be moved to the newly added second text.
  • Figure 7 is a schematic diagram of another method of updating the target text provided by an embodiment of the present disclosure.
  • the end position of the first timeline interval corresponding to the first text 9 is dragged leftward from position A to position C, the length of the first timeline interval of the first text 9 can be shortened on the editing track segment of the target video, and the text content in the first text 9 corresponding to the shortened part of the first timeline interval (such as "pause 0.02s") is moved to the second text 10.
  • Figure 8 is a schematic diagram of another method of updating the target text provided by an embodiment of the present disclosure.
  • the end position of the first timeline interval corresponding to the first text 11 is dragged rightward from position D to position E, the length of the first timeline interval of the first text 11 can be shortened on the editing track segment of the target video.
  • a second text 12 can be added to the text area, and the text content in the first text 11 corresponding to the shortened part of the first timeline interval (i.e., "repeat") can be moved to the second text 12.
  • the invalid text is text in the track area that is in a selected state
  • the second adjustment operation in response to the invalid text is performed on the target
  • the number of timeline intervals identified on the editing track segment of the video is adjusted, including at least one of the following:
  • the fourth text is switched from a selected state to an unselected state, and the timeline interval marking the fourth text is deselected on the editing track segment.
  • the third text may be any text in the text area, such as a text in an unselected state in the text area.
  • the selection operation may refer to an operation of switching the text from an unselected state to a selected state, such as the selection operation may be an operation of clicking a control, the control may have a first display style and a second display style, and the selected state of the text may change with the change of the display style of the control, and the specific contents of the first display style and the second display style are not limited, as long as different display styles can be distinguished.
  • the fourth text may be any text in the text area, such as a text in a selected state in the text area.
  • the fourth text may be the same as or different from the third text.
  • the deselect operation may refer to an operation of switching the text from a selected state to an unselected state, such as the deselect operation may correspond to the selection operation, which is an operation of clicking a certain control.
  • this embodiment can respond to the selection operation of the third text in the text area, switch the third text from an unselected state to a selected state, and add a timeline interval identifying the third text on the editing track segment. On this basis, the user can add the third text to be deleted and the timeline interval of the third text as needed.
  • the fourth text in response to the deselection operation on the fourth text in the text area, the fourth text may be switched from a selected state to an unselected state, and the timeline interval marking the fourth text may be deselected on the editing track segment.
  • the user may deselect the fourth text to be deleted and deselect the timeline interval marking the fourth text as needed.
  • a selection operation on the third text corresponding to the control is triggered.
  • the third text can be switched from an unselected state to a selected state, and an edit track segment can be added.
  • a timeline interval for marking a third text is added; when a control is clicked to display the control in the second display style, it can be considered that a deselection operation of a fourth text corresponding to the control is triggered.
  • the fourth text can be switched from a selected state to an unselected state, and the timeline interval for marking the fourth text can be deselected on the editing track segment.
  • FIG9 is a schematic diagram of the structure of a video editing device provided in an embodiment of the present disclosure.
  • the device may be applicable to video editing, wherein the device may be implemented by software and/or hardware and is generally integrated on an electronic device.
  • the device includes:
  • a text determination module 310 is used to determine invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video, wherein the timeline position of the invalid text is used to indicate the appearance time of the speech audio of the invalid text in the target video;
  • An interval identification module 320 is used to display the editing track segment of the target video on the editing interface of the target video, and identify the timeline interval of the invalid text on the editing track segment based on the timeline position of the invalid text, wherein the timeline interval of the invalid text is the time interval in which the voice audio of the invalid text appears in the target video;
  • An interval adjustment module 330 configured to adjust the timeline interval of the invalid text on the editing track segment of the target video in response to the adjustment operation of the invalid text;
  • the segment deletion module 340 is used to delete the video segment on the editing track segment within the timeline interval of the invalid text from the target video in response to the video segment deletion operation of the invalid text.
  • the embodiment of the present disclosure provides a video editing device, which uses a text determination module to perform voice recognition on the audio in the target video to determine the invalid text in the voice text of the target video and the timeline position of the invalid text, wherein the timeline position of the invalid text is used to indicate the appearance time of the voice audio of the invalid text in the target video; and uses an interval identification module to display the invalid text on the editing interface of the target video.
  • the editing track segment of the target video is identified, and based on the timeline position of the invalid text, the timeline interval of the invalid text is marked on the editing track segment, and the timeline interval of the invalid text is the time interval in which the voice and audio of the invalid text appear in the target video; the timeline interval of the invalid text is adjusted on the editing track segment of the target video in response to the adjustment operation of the invalid text by the interval adjustment module; the video segment located in the timeline interval of the invalid text on the editing track segment is deleted from the target video in response to the video segment deletion operation of the invalid text by the segment deletion module.
  • the invalid text and the position of the invalid text in the target video can be determined, and then by marking the timeline interval of the invalid text on the editing track segment of the target video, the invalid segment can be deleted in response to the user's adjustment operation on the invalid text and the video segment deletion operation, thereby completing the video editing and improving the efficiency of video editing.
  • the interval adjustment module 330 specifically includes at least one of the following:
  • a first response unit configured to adjust the length of the timeline section of the invalid text on the editing track segment of the target video in response to a first adjustment operation of the invalid text
  • the second response unit is used to adjust the number of timeline intervals marked on the editing track segment of the target video in response to the second adjustment operation of the invalid text.
  • the interval identification module 320 is specifically used for:
  • the editing track segment of the target video is displayed in the track area of the editing interface, and the invalid text and non-invalid text in the voice text are displayed in the text area of the editing interface, wherein the invalid text is displayed as a selected state on the editing interface, and the non-invalid text is displayed as a non-selected state on the editing interface.
  • the first response unit includes:
  • an adjusting subunit configured to adjust the length of a first timeline interval of a first text on the editing track segment of the target video, and update a target text displayed in the text area, wherein the target text includes the first text and/or a timeline interval related to the first text; The associated second text.
  • the first adjustment operation includes a length extension operation
  • the adjustment subunit is specifically configured to:
  • the text content in the second text corresponding to the extended portion of the first timeline interval is moved to the first text.
  • the first adjustment operation includes a length shortening operation
  • the adjustment subunit is specifically configured to:
  • a second text is added to the text area, and text content in the first text corresponding to the shortened portion of the first timeline interval is moved to the second text.
  • the invalid text is text in the track area that is in a selected state
  • the second response unit is specifically used for at least one of the following:
  • the fourth text is switched from a selected state to an unselected state, and the timeline interval marking the fourth text is deselected on the editing track segment.
  • a video editing device provided by an embodiment of the present disclosure further includes:
  • a playing module used for playing the target video in the video area of the editing interface
  • the first progress adjustment module is used to adjust the playback progress of the target video to the first time node when the target video is played to the starting point of the second timeline interval, and continue to play the target video from the first time node, wherein the second timeline interval is the timeline interval of the invalid text, and the first time node is the time node corresponding to the end position of the second timeline interval in the target video.
  • a video editing device provided by an embodiment of the present disclosure further includes:
  • the first response module is used to play the target in the video area of the editing interface.
  • the playback progress of the target video is adjusted to the second time node, and the target video continues to be played from the second time node, the fifth text is displayed in the text area of the editing interface, the fifth text is invalid text or non-invalid text, and the second time node is the time node corresponding to the starting position of the fifth text in the target video.
  • a time mark is also displayed on the editing track segment.
  • the video editing device provided by the embodiment of the present disclosure further includes:
  • a second response module configured to determine, after playing the target video in the video area of the editing interface, in response to a position adjustment operation on the target identifier, a third timeline interval where a target position of the target identifier is located, wherein the target position is a position of the target identifier on the editing track segment when the position adjustment operation is completed;
  • a second progress adjustment module is used to adjust the playback progress of the target video to a third time node if the third timeline interval is the timeline interval of the invalid text, and continue to play the target video from the third time node, wherein the third time node is the time node corresponding to the starting position or the ending position of the third timeline interval in the target video, or the time node to which the target video was last played before receiving the position progress adjustment operation;
  • the third progress adjustment module is used to adjust the playback progress of the target video to a fourth time node if the third timeline interval is not the timeline interval of the invalid text, and continue to play the target video from the fourth time node, wherein the fourth time node is the time node corresponding to the target position in the target video.
  • segment deletion module 340 is specifically used for:
  • the video clip on the editing track segment that is located in the timeline interval of the invalid text is deleted from the target video, and the target video after deleting the video clip is The videos are combined into one video.
  • the editing track segment is an audio track segment of the target video.
  • the above-mentioned video editing device can execute the video editing method provided by any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution method.
  • the terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
  • mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
  • PDAs personal digital assistants
  • PADs tablet computers
  • PMPs portable multimedia players
  • vehicle-mounted terminals such as vehicle-mounted navigation terminals
  • fixed terminals such as digital TVs, desktop computers, etc.
  • the electronic device shown in FIG10 is only an example and should not bring any limitation to the functions and scope of use of
  • the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 to a random access memory (RAM) 403.
  • a processing device e.g., a central processing unit, a graphics processing unit, etc.
  • RAM random access memory
  • various programs and data required for the operation of the electronic device 400 are also stored.
  • the processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404.
  • An input/output (I/O) interface 405 is also connected to the bus 404.
  • the following devices may be connected to the I/O interface 405: input devices 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 408 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 409.
  • the communication device 409 may allow the electronic device 400 to communicate wirelessly or wired with other devices to exchange data.
  • FIG. 10 shows an electronic device 400 with various devices, it should be understood that it is not required to implement or have all the devices shown. More or fewer devices may be implemented or have alternatively.
  • an embodiment of the present disclosure includes a computer program.
  • a computer program product includes a computer program carried on a non-transitory computer readable medium, the computer program including program code for executing the method shown in the flowchart.
  • the computer program can be downloaded and installed from a network through a communication device 409, or installed from a storage device 408, or installed from a ROM 402.
  • the processing device 401 executes the above functions defined in the method of the embodiment of the present disclosure.
  • the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two.
  • the computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above.
  • Computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
  • a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
  • a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried.
  • This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above.
  • the computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device.
  • the program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
  • the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may communicate with any form or medium.
  • Digital data communications e.g., communication networks
  • Examples of communication networks include a local area network ("LAN”), a wide area network (“WAN”), an internetwork (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
  • LAN local area network
  • WAN wide area network
  • Internet internetwork
  • peer-to-peer network e.g., an ad hoc peer-to-peer network
  • the computer-readable medium may be included in the electronic device, or may exist independently without being installed in the electronic device.
  • the computer-readable medium carries one or more programs.
  • the electronic device When the one or more programs are executed by the electronic device, the electronic device:
  • the video segment on the editing track segment that is within the timeline interval of the invalid text is deleted from the target video.
  • Computer program code for performing operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages.
  • the program code may execute entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server.
  • the remote computer may be connected via any type of network, including a local area network.
  • a computer system is a network (LAN) or wide area network (WAN)—connected to a user's computer, or it can be connected to an external computer (for example, through the Internet using an Internet service provider).
  • each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function.
  • the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
  • each square box in the block diagram and/or flow chart, and the combination of the square boxes in the block diagram and/or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
  • the units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a module does not, in some cases, limit the unit itself.
  • exemplary types of hardware logic components include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
  • FPGAs field programmable gate arrays
  • ASICs application specific integrated circuits
  • ASSPs application specific standard products
  • SOCs systems on chip
  • CPLDs complex programmable logic devices
  • a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
  • a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
  • a machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
  • machine-readable storage media would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable Read only memory (EPROM or flash memory), optical fiber, portable compact disk read only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.
  • RAM random access memory
  • ROM read-only memory
  • EPROM or flash memory erasable programmable Read only memory
  • CD-ROM portable compact disk read only memory
  • magnetic storage device or any suitable combination of the above.
  • Example 1 provides a video editing method, including:
  • the video segment on the editing track segment that is within the timeline interval of the invalid text is deleted from the target video.
  • Example 2 is the method according to Example 1, wherein in response to the adjustment operation of the invalid text, the timeline interval of the invalid text is adjusted on the editing track segment of the target video, including at least one of the following:
  • the number of timeline intervals identified on the editing track segment of the target video is adjusted.
  • Example 3 is the method according to Example 2, wherein the editing track of the target video is displayed on the editing interface of the target video.
  • the snippet includes:
  • the editing track segment of the target video is displayed in the track area of the editing interface, and the invalid text and non-invalid text in the voice text are displayed in the text area of the editing interface, wherein the invalid text is displayed as a selected state on the editing interface, and the non-invalid text is displayed as a non-selected state on the editing interface.
  • Example 4 is the method according to Example 3, wherein adjusting the length of the timeline interval of the invalid text on the editing track segment of the target video includes:
  • the length of a first timeline interval of a first text is adjusted on an editing track segment of the target video, and the target text displayed in the text area is updated, wherein the target text includes the first text and/or a second text associated with the first text.
  • Example 5 is the method according to Example 4, wherein the first adjustment operation includes a length extension operation, and the updating of the target text displayed in the text area includes:
  • the text content in the second text corresponding to the extended portion of the first timeline interval is moved to the first text.
  • Example 6 is the method according to Example 4, wherein the first adjustment operation includes a length shortening operation, and the updating of the target text displayed in the text area includes:
  • a second text is added to the text area, and text content in the first text corresponding to the shortened portion of the first timeline interval is moved to the second text.
  • Example 7 is the method according to Example 3, wherein the invalid text is text in the track area that is in a selected state, and the second adjustment operation in response to the invalid text adjusts the number of timeline intervals identified on the editing track segment of the target video, including at least one of the following:
  • the fourth text is switched from a selected state to an unselected state, and the timeline interval marking the fourth text is deselected on the editing track segment.
  • Example 8 is a method according to any one of Examples 1-7, further comprising:
  • the playback progress of the target video is adjusted to the first time node, and the target video continues to be played from the first time node, wherein the second timeline interval is the timeline interval of the invalid text, and the first time node is the time node corresponding to the end point of the second timeline interval in the target video.
  • Example 9 is the method according to Example 8, after playing the target video in the video area of the editing interface, further comprising:
  • the playback progress of the target video is adjusted to a second time node, and the target video continues to be played from the second time node, the fifth text is displayed in the text area of the editing interface, the fifth text is an invalid text or a non-invalid text, and the second time node is the time node corresponding to the starting position of the fifth text in the target video.
  • Example 10 is the method according to Example 8, wherein a time mark is also displayed on the editing track segment, and after the target video is played in the video area of the editing interface, the method further includes:
  • the playback progress of the target video is adjusted to the third time node, and the target video is played from the third time node.
  • the target video continues to be played at the node, wherein the third time node is the time node corresponding to the starting position or the ending position of the third timeline interval in the target video, or the time node to which the target video was last played before the position progress adjustment operation is received;
  • the playback progress of the target video is adjusted to a fourth time node, and the target video continues to be played from the fourth time node, wherein the fourth time node is the time node corresponding to the target position in the target video.
  • Example 11 is a method according to any one of Examples 1-7, wherein the step of deleting the video segment on the editing track segment within the timeline interval of the invalid text from the target video includes:
  • the video segment on the editing track segment that is located in the timeline interval of the invalid text is deleted from the target video, and the target video after the video segment is deleted is integrated into one video.
  • Example 12 is a method according to any one of Examples 1-7, wherein the edited track segment is an audio track segment of the target video.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Computational Linguistics (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Artificial Intelligence (AREA)
  • Television Signal Processing For Recording (AREA)
  • Management Or Editing Of Information On Record Carriers (AREA)

Abstract

本公开实施例提供了一种视频的编辑方法、装置、电子设备和存储介质,方法包括:通过对目标视频中的音频进行语音识别,确定目标视频的语音文本中的无效文本及无效文本的时间线位置;在目标视频的编辑界面上展示目标视频的编辑轨道片段,并基于无效文本的时间线位置,在编辑轨道片段上标识无效文本的时间线区间;响应于无效文本的调整操作,在目标视频的编辑轨道片段上对无效文本的时间线区间进行调整;响应于无效文本的视频片段删除操作,将编辑轨道片段上位于无效文本的时间线区间内的视频片段从目标视频中删除。该方法能够响应于用户对无效文本的调整操作和视频片段删除操作,对无效片段进行删除,从而提升了视频剪辑的效率。

Description

视频的编辑方法、装置、电子设备和存储介质
本申请要求于2023年01月19日提交中国专利局、申请号为202310107566.3、发明名称为“视频的编辑方法、装置、电子设备和存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本公开实施例涉及计算机技术领域,尤其涉及一种视频的编辑方法、装置、电子设备和存储介质。
背景技术
随着科技的快速发展,视频类应用程序应运而生,如用户可以拍摄并上传口播视频至视频类应用程序,其中,出镜口播或音频承载了口播视频的核心内容,故如何对口播视频进行更好地剪辑是当前亟需解决的问题。
口播视频的剪辑主要需要处理掉视频中的停顿、磕巴、口水词等用户不需要的片段,现有的剪辑方法一般通过视频中音频的波形来识别无效片段,但该方法需要耗费大量的时间,剪辑效率较低。
发明内容
本公开实施例提供一种视频的编辑方法、装置、电子设备和存储介质,以提升视频的剪辑效率。
第一方面,本公开实施例提供了一种视频的编辑方法,包括:
通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;
在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片 段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;
响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;
响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
第二方面,本公开实施例还提供了一种视频的编辑装置,包括:
文本确定模块,用于通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;
区间标识模块,用于在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;
区间调整模块,用于响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;
片段删除模块,用于响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
第三方面,本公开实施例还提供了一种电子设备,包括:
一个或多个处理器;
存储器,用于存储一个或多个程序,
当所述一个或多个程序被所述一个或多个处理器执行,使得所述一个或多个处理器实现如本公开实施例所述的视频的编辑方法。
第四方面,本公开实施例还提供了一种计算机可读存储介质,其上存储有计算机程序,该程序被处理器执行时实现如本公开实施例所述的视频的编辑方法。
本公开实施例提供了一种视频的编辑方法、装置、电子设备和存储介质,所述方法包括:通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。利用上述技术方案,通过对目标视频中的音频进行语音识别,能够对目标视频中的无效文本以及无效文本的位置进行确定,进而通过在目标视频的编辑轨道片段上标识无效文本的时间线区间,能够响应于用户对无效文本的调整操作和视频片段删除操作,对无效片段进行删除,从而完成对视频的剪辑,提升了视频剪辑的效率。
附图说明
结合附图并参考以下具体实施方式,本公开各实施例的上述和其他特征、优点及方面将变得更加明显。贯穿附图中,相同或相似的附图标记表示相同或相似的元素。应当理解附图是示意性的,原件和元素不一定按照比例绘制。
图1为本公开实施例提供的一种视频的编辑方法的流程示意图;
图2为本公开实施例提供的一种触发语音识别的示意图;
图3为本公开实施例提供的一种编辑界面的示意图;
图4为本公开实施例提供的另一种视频的编辑方法的流程示意图;
图5为本公开实施例提供的另一种编辑界面的示意图;
图6为本公开实施例提供的一种更新目标文本的示意图;
图7为本公开实施例提供的另一种更新目标文本的示意图;
图8为本公开实施例提供的另一种更新目标文本的示意图;
图9为本公开实施例提供的一种视频的编辑装置的结构示意图;
图10为本公开实施例提供的一种电子设备的结构示意图。
具体实施方式
下面将参照附图更详细地描述本公开的实施例。虽然附图中显示了本公开的某些实施例,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实施例,相反提供这些实施例是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实施例仅用于示例性作用,并非用于限制本公开的保护范围。
应当理解,本公开的方法实施方式中记载的各个步骤可以按照不同的顺序执行,和/或并行执行。此外,方法实施方式可以包括附加的步骤和/或省略执行示出的步骤。本公开的范围在此方面不受限制。
本文使用的术语“包括”及其变形是开放性包括,即“包括但不限于”。术语“基于”是“至少部分地基于”。术语“一个实施例”表示“至少一个实施例”;术语“另一实施例”表示“至少一个另外的实施例”;术语“一些实施例”表示“至少一些实施例”。其他术语的相关定义将在下文描述中给出。
需要注意,本公开中提及的“第一”、“第二”等概念仅用于对不同的装置、模块或单元进行区分,并非用于限定这些装置、模块或单元所执行的功能的顺序或者相互依存关系。
需要注意,本公开中提及的“一个”、“多个”的修饰是示意 性而非限制性的,本领域技术人员应当理解,除非在上下文另有明确指出,否则应该理解为“一个或多个”。
本公开实施方式中的多个装置之间所交互的消息或者信息的名称仅用于说明性的目的,而并不是用于对这些消息或信息的范围进行限制。
图1为本公开实施例提供的一种视频的编辑方法的流程示意图,该方法可适用于对视频进行编辑的情况,该方法可以由视频的编辑装置来执行,其中该装置可由软件和/或硬件实现,并一般集成在电子设备上,在本实施例中电子设备包括但不限于:手机或电脑等设备。
如图1所示,本公开实施例提供的一种视频的编辑方法,包括如下步骤:
S110、通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间。
目标视频可以是指需要进行编辑的视频,如目标视频可以为用户上传的视频,也可以为用户在线拍摄的视频。在本实施例中,目标视频中包含音频,语音文本可以认为是与目标视频的音频对应的文本信息,其可以通过对目标视频的音频进行语音识别得到。示例性的,语音文本中包括无效文本,无效文本可以包括目标视频中停顿、重复、口水词等用户不需要的片段对应的文本;无效文本的时间线位置可以用于表示无效文本的语音音频在目标视频中的出现时间。
在一个实施例中,目标视频为连续视频,如目标视频可以为某连续且完整视频,也可以为某完整视频的某一片段。
在本实施例中,可以通过对目标视频中的音频进行语音识别,来确定目标视频的语音文本中的无效文本及无效文本的时间线位 置,以进行后续的编辑步骤。其中,本步骤不对语音识别的具体方式以及触发语音识别的时机进行限定,如可以在界面中触发某识别控件来触发对目标视频中音频的语音识别,从而能够自动识别出目标视频中无效文本及无效文本的时间线位置,识别控件的位置及样式不限,可以根据实际页面的情况进行设置。
在一个实施方式中,无效文本的识别条件为:默认静音时长大于预设时长(如120ms等)的片段是停顿片段;通过复用预设算法逻辑来识别无效词(如重复与语气词),该预设算法可以配置于当前应用程序内,也可以配置于服务器中。
图2为本公开实施例提供的一种触发语音识别的示意图,如图2所示,在选中某视频(如目标视频)后,可以在点击右键的弹出窗口中触发控件1,即可触发对目标视频中音频的语音识别。还可以点击当前界面中的快捷控件2实现对目标视频中音频的语音识别。
S120、在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间。
编辑界面可以用于对目标视频中的片段进行编辑。目标视频的编辑轨道片段可以理解为与目标视频对应的编辑轨道片段,如起点与目标视频的起点对应、终点与目标视频的终点对应的编辑轨道片段。示例性的,当目标视频为连续且完整视频时,该编辑轨道片段可以为该连续且完整视频的编辑轨道;当目标视频为某完整视频的一个片段时,该编辑轨道片段可以为该完整视频的编辑轨道中与目标视频对应的轨道片段。可选的,所述编辑轨道片段为所述目标视频的音频轨道片段,音频轨道片段可以是指与目标视频对应的音频轨道片段。无效文本的时间线区间可以认为是无效文本的语音音频在目标视频中出现的时间区间。
在一个实施例中,当对目标视频中的音频进行语音识别确定目标视频的无效文本及无效文本的时间线位置后,可以在目标视频的 编辑界面上对目标视频的编辑轨道片段进行展示,并基于无效文本的时间线位置,在编辑轨道片段上标识无效文本的时间线区间。其中,展示编辑轨道片段以及标识无效文本的时间线区间的方式不限,可以由相关人员预先根据编辑界面的实际情况进行设置,如可以采用预设标识在目标视频的编辑轨道片段上标识无效文本中无效文本的时间线区间,该预设标识具体可以根据进行设置,示例性的,该预设标识可以为矩形框,此时,可以采用矩形框在目标视频的编辑轨道片段上框选目标视频中无效文本的时间线区间;或者,无效文本的时间线区间在编辑轨道片段中可以标识为第一展示样式,而其他文本的时间线区间在编辑轨道片段中标识为第二展示样式,第一展示样式区别于第二展示样式,如第一展示样式和第二展示样式可以以不同颜色进行区分,也可以通过不同的线条进行区分等。
在一个实施方式中,可以默认在编辑轨道片段上标识识别得到的所有无效文本的时间线区间,以区分无效文本的时间线区间与其他文本的时间线区间;也可以默认在编辑轨道片段上标识设定个数的无效文本的时间线区间,设定个数可以由相关人员预先进行设置。用户可以根据需要对无效文本的时间线区间进行调整。
图3为本公开实施例提供的一种编辑界面的示意图,如图3所示,目标视频的编辑界面上可以展示有目标视频的编辑轨道片段3,并基于上步骤识别出的无效文本的时间线位置,在编辑轨道片段3上对无效文本的时间线区间进行标识,如图3中以黑色的矩形框对无效文本的时间线区间4进行了标识。
S130、响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整。
无效文本的调整操作可以是指用于触发调整无效文本的时间线区间的操作,本实施例中无效文本的调整操作的具体操作方式不限,如无效文本的调整操作可以为点击编辑界面中某一控件的操作或者作用于编辑轨道片段上的预设出发操作,也可以为在编辑界面中进行预设手势的操作,预设手势可为预先设定的手势,如预设手势可 为向右滑动的手势等。
本实施例不对调整操作的触发位置进行限定,如可以作用在无效文本的时间线区间,也可以作用在其他位置,只要能触发调整无效文本的时间线区间即可。
具体的,本实施例可以响应于无效文本的调整操作,在目标视频的编辑轨道片段上对无效文本的时间线区间进行调整,具体调整时间线区间的方式可以根据调整操作的不同而有所区别,如当无效文本的调整操作为触发某控件以使无效文本的状态由选中状态变为未选中状态时,可以在目标视频的编辑轨道片段上对无效文本的时间线区间的数量进行调整,如在编辑轨道片段上增加标识某文本的时间线区间,以将此文本调整为无效文本,或者,在编辑轨道片段上取消标识某文本的时间区间,以将此文本调整为非无效文本。
S140、响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
无效文本的视频片段删除操作可以是指用于触发删除编辑轨道片段上无效文本对应的视频片段的操作,例如无效文本的视频片段删除操作可以为点击编辑界面中某删除控件的操作等。
本步骤可以响应于无效文本的视频片段删除操作,将编辑轨道片段上位于无效文本的时间线区间内的视频片段从目标视频中删除,具体删除的方式不作限定,如可以根据用户需要保留或去除编辑轨道片段上无效文本的时间线区间对应的痕迹。如图3所示,用户触发控件5即可以认为是视频片段删除操作,在用户触发控件5后,电子设备可以响应于无效文本的视频片段删除操作,将无效文本的时间线区间6对应的视频片段从目标视频中删除。
在一个实施例中,所述将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,包括:
将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,并以删除的所述视频片段为分割点, 将所述目标视频分割为多个独立的视频片段;或者
将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,并将删除所述视频片段后的目标视频整合为一个视频。
在一个实施方式中,可以将编辑轨道片段上位于无效文本的时间线区间内的视频片段从目标视频中删除,并以删除的视频片段为分割点,将目标视频分割为多个独立的视频片段,在此基础上,能够便于用户根据需要对各独立的视频片段进行操作。
在一个实施方式中,可以将编辑轨道片段上位于无效文本的时间线区间内的视频片段从目标视频中删除,并将删除视频片段后的目标视频整合为一个视频。在此基础上,实现了删除目标视频中无效片段的操作,提高了目标视频的质量。无效片段即为无效文本的时间线区间对应的片段,换言之,编辑轨道片段上所标识的时间线区间在目标视频的语音文本中所对应的文本。
本公开实施例提供的一种视频的编辑方法,通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。利用该方法,通过对目标视频中的音频进行语音识别,能够对目标视频中的无效文本以及无效文本的位置进行确定,进而通过在目标视频的编辑轨道片段上标识无效文本的时间线区间,能够响应于用户对无效文本的调整操作和视频片段删 除操作,对无效片段进行删除,从而完成对视频的剪辑,提升了视频剪辑的效率。
在一个实施例中,所述方法还包括:
在所述编辑界面的视频区域播放所述目标视频;
当所述目标视频播放至第二时间线区间的起点位置时,将所述目标视频的播放进度调整至第一时间节点,并从所述第一时间节点继续播放所述目标视频,其中,所述第二时间线区间为所述无效文本的时间线区间,所述第一时间节点为所述第二时间线区间的终点位置在所述目标视频中所对应的时间节点。
其中,视频区域可以认为是编辑界面中的某个区域,用于播放目标视频,视频区域的具体位置及大小可以由相关人员根据编辑界面的情况进行设置。第二时间线区间可以是任一无效文本的时间线区间,第一时间节点可以理解为第二时间线区间的终点位置在目标视频中所对应的时间节点。
在一个实施方式中,可以在编辑界面的视频区域播放目标视频,播放目标视频的时机不限,如在显示编辑界面时可以自动对目标视频进行播放;也可以在显示编辑界面后,当用户触发编辑界面中的某播放控件时,在编辑界面的视频区域播放目标视频;还可以在显示编辑界面的预设时长后,自动对目标视频进行播放等,本实施例对此不作限定。
当目标视频播放至第二时间线区间的起点位置时,本实施例可以自动将目标视频的播放进度调整至第一时间节点,并从第一时间节点继续播放目标视频。在此基础上,通过播放目标视频时自动跳过无效片段,能够使用户预先观看到编辑后的视频,以对目标视频进行更好地编辑。
在一个实施例中,在所述编辑界面的视频区域播放所述目标视频之后,还包括:
响应于对第五文本的触发操作,将所述目标视频的播放进度调整至第二时间节点,并从所述第二时间节点继续播放所述目标视频, 所述第五文本显示于所述编辑界面的文本区域内,所述第五文本为无效文本或非无效文本,所述第二时间节点为所述第五文本的起点位置在所述目标视频中所对应的时间节点。
在本实施例中,第五文本可以为用户所触发的文本,其可以为无效文本或非无效文本。非无效文本即为除无效文本之外的文本,如目标视频的语音文本中的有效文本。第五文本可以显示于编辑界面的文本区域内,文本区域可以认为是编辑界面中的某个区域,用于显示目标视频的语音文本,如用于显示目标视频的语音文本中的第五文本,文本区域的具体位置及大小可以由相关人员根据编辑界面的情况进行设置。
在一个实施方式中,文本区域中的无效文本自成一段,如可以将“停顿”文本单独成行,同时展示停顿时长。
对第五文本的触发操作可以用于触发将目标视频的播放进度调整第二时间节点,如对第五文本的触发操作可以为点击第五文本的操作;第二时间节点即为第五文本的起点位置在目标视频中所对应的时间节点。
具体的,在编辑界面的视频区域播放目标视频之后,还可以响应于对第五文本的触发操作,将目标视频的播放进度调整至第二时间节点,并从第二时间节点继续播放目标视频。在此基础上,通过响应于第五文本的触发操作,实现了对目标视频播放进度的调整,以此实现在目标视频中对相应文本的定位,便于用户观看相应文本所对应的视频片段。
在一个实施例中,所述编辑轨道片段上还展示有时间标识,在所述编辑界面的视频区域播放所述目标视频之后,还包括:
响应于对所述目标标识的位置调整操作,确定所述目标标识的目标位置所位于的第三时间线区间,所述目标位置为所述位置调整操作执行完毕时所述目标标识在所述编辑轨道片段上的位置;
如果所述第三时间线区间为所述无效文本的时间线区间,则将所述目标视频的播放进度调整至第三时间节点,并从所述第三时间 节点继续播放所述目标视频,其中,所述第三时间节点为所述第三时间线区间的起点位置或终点位置在所述目标视频中所对应的时间节点,或者,接收到所述位置进度调整操作之前所述目标视频最后播放至的时间节点;
如果所述第三时间线区间不为所述无效文本的时间线区间,则将所述目标视频的播放进度调整至第四时间节点,并从所述第四时间节点继续播放所述目标视频,其中,所述第四时间节点为所述目标位置在所述目标视频中所对应的时间节点。
时间标识用于在目标视频的编辑轨道片段中表征目标视频当前所播放至的时间节点,换言之,此时间标识可用于指示目标视频当前的播放进度对应的时间节点。当前应用程序在编辑界面的视频区域播放目标视频时,可以依据目标视频的播放进度,实时更新目标标识在目标视频的编辑轨道片段中的位置。目标标识的位置调整操作则可以理解为调整目标标识位置的操作,如目标标识的位置调整操作可以为点击编辑轨道片段中的某一位置a以将目标标识调整至该位置a的操作,目标标识的位置调整操作也可以为拖拽目标标识以将目标标识调整至位置a的操作等。目标位置可以认为是位置调整操作执行完毕时目标标识在编辑轨道片段上的位置,第三时间线区间即为目标标识的目标位置所位于的时间线区间。
第三时间节点可以是第三时间线区间的起点位置在目标视频中所对应的时间节点,或第三时间线区间的终点位置在目标视频中所对应的时间节点,也可以是接收到位置进度调整操作之前目标视频最后播放至的时间节点;第四时间节点可以认为是目标位置在目标视频中所对应的时间节点。
本步骤中,在编辑界面的视频区域播放目标视频之后,还可以响应于对目标标识的位置调整操作,对目标标识的位置进行调整。
示例性的,在对目标标识的位置调整操作的执行过程中,可以依据时间标识在所述编辑轨道片段中的位置,调整所述目标视频的播放进度。当该位置调整操作执行结束时,确定目标标识的目标位 置,对目标位置所位于的第三时间线区间进行确定,并根据第三时间线区间的不同分别对目标视频的播放进度进行不同的调整。
在一个实施方式中,当第三时间线区间为无效文本的时间线区间时,可以将目标视频的播放进度调整至第四时间节点,并从第四时间节点继续播放目标视频。示例性的,当拖拽目标标识至某无效文本的时间线区间A内时,本实施例自动将目标视频的播放进度调整至时间线区间A的终点位置,以从时间线区间A的终点位置继续播放目标视频,或者,将目标标识的位置还原为接收到位置调整操作之前目标标识最后位于的位置,并从该位置继续播放目标视频。当点击目标视频的编辑轨道片段上所标识的时间线区间A内的某位置时,本实施例自动将目标视频的播放进度调整至时间线区间A的起点位置,病从时间线区间A的起点位置继续播放目标视频。
在一个实施方式中,当第三时间线区间不为无效文本的时间线区间时,则可以将目标视频的播放进度调整至目标位置在目标视频中所对应的时间节点,并从该时间节点继续播放目标视频。示例性的,当拖拽目标标识至的目标位置b不位于无效文本的时间线区间内时,或者,当通过点击的方式将目标标识移动至的目标位置b不位于无效文本的时间线区间内时,本实施例则可以将目标视频的播放进度调整至目标位置b在目标视频中所对应的时间节点c,并从时间节点c继续播放目标视频。在此基础上,通过响应于对目标标识的位置调整操作,能够使得当调整目标标识的位置位于无效片段时,能够自动跳过无效文本所对应的视频片段或者从无效文本所对应视频片段的起点位置完整播放该视频片段;而当调整目标标识的位置不位于无效片段时,从此位置所对应时间节点处播放目标视频的功能。
图4为本公开实施例提供的另一种视频的编辑方法的流程示意图。本实施例中的方案可以与上述实施例中的一个或多个可选方案组合。可选的,所述响应于所述无效文本的调整操作,在所述目标 视频的编辑轨道片段上对所述无效文本的时间线区间进行调整,包括下述至少之一:响应于所述无效文本的第一调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间的长度进行调整;响应于所述无效文本的第二调整操作,对所述目标视频的编辑轨道片段上所标识的时间线区间的数量进行调整。如图4所示,该方法包括:
S210、通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间。
S220、在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间。
S230、响应于所述无效文本的第一调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间的长度进行调整。
第一调整操作可以用于调整无效文本的时间线区间的长度,第一调整操作可以作用在目标视频的编辑轨道片段上,示例性的,第一调整操作可以为在目标视频的编辑轨道片段上拉伸或压缩某无效文本的时间线区间B的操作,如向左拖拽时间线区间B的起始位置的操作等。
本步骤可以响应于无效文本的第一调整操作,在目标视频的编辑轨道片段上对无效文本的时间线区间的长度进行调整,如当向左拖拽某时间线区间的左边界或向右拖拽某时间线区间的右边界时可以延长该时间线区间的区间长度,当向右拖拽某时间线区间的左边界或者向右拖拽某时间线区间右边界时可以缩短该时间线区间的区间长度。
S240、响应于所述无效文本的第二调整操作,对所述目标视频的编辑轨道片段上所标识的时间线区间的数量进行调整。
第二调整操作可以用于调整编辑轨道片段上所标识的时间线区间的数量,如第二调整操作可以作用在某调整控件上,每个无效文本可以对应一个调整控件,调整控件用于控制编辑轨道片段上是否标识此无效文本对应的时间线区间;第二调整操作也可以作用在编辑轨道片段的时间线区间上。
本步骤可以响应于无效文本的第二调整操作,对目标视频的编辑轨道片段上所标识的时间线区间的数量进行调整,本实施例不对调整的具体过程进行限定,只要能调整编辑轨道片段上所标识的时间线区间的数量即可。
S250、响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
本公开实施例提供的一种视频的编辑方法,通过响应于无效文本的第一调整操作和第二调整操作,能够在目标视频的编辑轨道片段上对无效文本的时间线区间的长度和数量进行调整,从而能够根据用户需求实现对目标视频的精确剪辑。
在一个实施例中,所述在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,包括:
在所述编辑界面的轨道区域展示所述目标视频的编辑轨道片段,并在所述编辑界面的文本区域展示所述语音文本中的无效文本和非无效文本,其中,所述无效文本在所述编辑界面上展示为选中状态,所述非无效文本在所述编辑界面上展示为非选中状态。
轨道区域可以为编辑界面内的某个区域,用于展示目标视频的编辑轨道片段。选中状态可以理解为某文本处于被选中的状态,以此表征此文本将要被删除;未选中状态可以理解为某文本处于未被选中的状态,以此表征此文本不会被删除。
本实施例中,在目标视频的编辑界面上展示目标视频的编辑轨道片段的具体过程可以为:在编辑界面的轨道区域对目标视频的编辑轨道片段进行展示,并在编辑界面的文本区域对语音文本中的无 效文本和非无效文本进行展示,其中,无效文本在编辑界面上可以展示为选中状态,非无效文本在编辑界面上可以展示为非选中状态。在此基础上,实现了展示编辑轨道片段的同时,在编辑界面的文本区域展示语音文本中的无效文本和非无效文本;此外,默认无效文本展示为选中状态,能够提示用户可以对无效文本进行删除,并避免用户需要执行相应的选中操作选中停顿、重复等无效文本才能对其进行删除,进一步简化用户删除目标视频中的无效文本所需的操作。
图5为本公开实施例提供的另一种编辑界面的示意图,如图5所示,在编辑界面的轨道区域7展示有目标视频的编辑轨道片段,并在编辑界面的文本区域8展示有语音文本中的无效文本(如“停顿0.12s”或“重复”)和非无效文本(如“我的”),其中,无效文本在编辑界面上展示为选中状态,非无效文本在编辑界面上展示为非选中状态。
在一个实施例中,所述在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间的长度进行调整,包括:
在所述目标视频的编辑轨道片段上对第一文本的第一时间线区间的长度进行调整,并更新所述文本区域中所展示的目标文本,所述目标文本包括所述第一文本和/或与所述第一文本相关联的第二文本。
第一文本可以是无效文本,如第一触发操作所针对时间线区间所属的文本。第一时间线区间即为第一文本的时间线区间。目标文本可以是指对应需要进行调整的文本,如目标文本可以包括第一文本和/或与第一文本相关联的第二文本,第二文本可以为与第一文本的时间线区间相邻时间线区间对应的文本,也可以为与第一文本的时间线区间相邻的时间线区间缩减后的新时间线区间对应的文本。
具体的,在目标视频的编辑轨道片段上对第一文本的第一时间线区间的长度进行调整的同时,需要对文本区域中所展示的目标文本进行更新,更新的具体内容可以由第一调整操作的具体操作来确 定。在此基础上,能够在调整第一时间线区间长度的同时,在文本区域将调整部分对应的文本进行实时更新,以精确化调整无效文本的时间线区间对应的时长。
在一个实施例中,所述第一调整操作包括长度延长操作,所述更新所述文本区域中所展示的目标文本,包括:
将所述第二文本中与所述第一时间线区间的延长部分对应的文本内容移动至所述第一文本中。
长度延长操作可以认为是延长第一时间线区间长度的操作,如长度延长操作可以为向左拖拽第一时间线区间的起始位置或者向右拖拽第一时间线区间的终止位置的操作。
在一个实施方式中,当在编辑轨道片段上延长第一文本的第一时间线区间的长度时,可以在文本区域同步将第二文本中与第一时间线区间的延长部分对应的文本内容移动至第一文本中,示例性的,当第一调整操作为向左拖拽第一时间线区间的起始位置的操作时,第二文本为位于第一时间线区间左侧的时间线区间对应的文本,此时,需要将第二文本中与第一时间线区间的延长部分对应的文本内容移动至第一文本中。
在一个实施方式中,长度延长操作可以支持将相邻非无效文本中的部分文本移动至第一文本中,当第二文本为无效文本时,则可以支持或不支持进行长度延长操作。
图6为本公开实施例提供的一种更新目标文本的示意图,如图6所示,当由位置A向右拖拽第一文本9对应的第一时间线区间的终点位置至位置B时,可以在目标视频的编辑轨道片段上对第一文本9的第一时间线区间的长度进行延长,同时将第二文本10中与第一时间线区间的延长部分对应的文本内容(如“太阳”)移动至第一文本9中。
在一个实施例中,所述第一调整操作包括长度缩短操作,所述更新所述文本区域中所展示的目标文本,包括:
将所述第一文本中与所述第一时间线区间的缩短部分对应的文 本内容移动至所述第二文本中;或者
在所述文本区域中新增第二文本,并将所述第一文本中与所述第一时间线区间的缩短部分对应的文本内容移动至所述第二文本中。
长度缩短操作可以认为是缩短第一时间线区间长度的操作,如长度缩短操作可以为向右拖拽第一时间线区间的起始位置或向左拖拽第一时间区间的终止位置的操作。
在一个实施方式中,当在编辑轨道片段上缩短第一文本的第一时间线区间的长度时,可以根据第一文本内容的不同来确定文本区域中所展示目标文本的更新方式,示例性的,当第一文本内容为停顿时,可以将第一文本中与第一时间线区间的缩短部分对应的文本内容移动至第二文本中;而当第一文本内容为重复文本、无效词或除停顿、重复和无效词之外的正常文本时,则可以在文本区域中新增第二文本,并将第一文本中与第一时间线区间的缩短部分对应的文本内容移动至新增的第二文本中。
图7为本公开实施例提供的另一种更新目标文本的示意图,如图7所示,当由位置A向左拖拽第一文本9对应的第一时间线区间的终点位置至位置C时,可以在目标视频的编辑轨道片段上对第一文本9的第一时间线区间的长度进行缩短,同时将第一文本9中与第一时间线区间的缩短部分对应的文本内容(如“停顿0.02s”)移动至第二文本10中。
图8为本公开实施例提供的另一种更新目标文本的示意图,如图8所示,当由位置D向右拖拽第一文本11对应的第一时间线区间的终点位置至位置E时,可以在目标视频的编辑轨道片段上对第一文本11的第一时间线区间的长度进行缩短,同时可以在文本区域中新增第二文本12,并将第一文本11中与第一时间线区间的缩短部分对应的文本内容(即“重复”)移动至第二文本12中。
在一个实施例中,所述无效文本为所述轨道区域中处于选中状态的文本,所述响应于所述无效文本的第二调整操作,对所述目标 视频的编辑轨道片段上所标识的时间线区间的数量进行调整,包括下述至少之一:
响应于对所述文本区域中的第三文本的选中操作,将所述第三文本由未选中状态切换为选中状态,并在所述编辑轨道片段上增加标识所述第三文本的时间线区间;
响应于对所述文本区域中的第四文本的取消选中操作,将所述第四文本由选中状态切换为未选中状态,并在所述编辑轨道片段上取消标识所述第四文本的时间线区间。
第三文本可以是文本区域中的任一文本,如文本区域中处于未选中状态的某文本。选中操作可以是指将文本由未选中状态切换为选中状态的操作,如选中操作可以为点击某控件的操作,该控件可以具有第一显示样式和第二显示样式,且文本的选中状态可以跟随该控件显示样式的变化而变化,第一显示样式和第二显示样式的具体内容不限,只要能区分不同的显示样式即可。
第四文本可以是文本区域中的任一文本,如文本区域中处于选中状态的某文本。第四文本可以与第三文本相同,也可以不同。取消选中操作可以是指将文本由选中状态切换为未选中状态的操作,如取消选中操作可以与选中操作相对应,为点击某控件的操作。
具体的,本实施例可以响应于对文本区域中的第三文本的选中操作,将第三文本由未选中状态切换为选中状态,并在编辑轨道片段上增加标识第三文本的时间线区间,在此基础上,用户可以根据需要增加将要删除的第三文本以及第三文本的时间线区间。
本实施例也可以响应于对文本区域中的第四文本的取消选中操作,将第四文本由选中状态切换为未选中状态,并在编辑轨道片段上取消标识第四文本的时间线区间。在此基础上,用户可以根据需要取消将要删除的第四文本以及取消标识第四文本的时间线区间。
示例性的,当点击某控件使该控件显示为第一显示样式时,可以认为是触发了对该控件所对应第三文本的选中操作,此时,可以将第三文本由未选中状态切换为选中状态,并在编辑轨道片段上增 加标识第三文本的时间线区间;当点击某控件使该控件显示为第二显示样式时,可以认为是触发了对该控件所对应第四文本的取消选中操作,此时,可以将第四文本由选中状态切换为未选中状态,并在编辑轨道片段上取消标识第四文本的时间线区间。
图9为本公开实施例提供的一种视频的编辑装置的结构示意图,该装置可适用于对视频进行编辑的情况,其中该装置可由软件和/或硬件实现,并一般集成在电子设备上。
如图9所示,该装置包括:
文本确定模块310,用于通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;
区间标识模块320,用于在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;
区间调整模块330,用于响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;
片段删除模块340,用于响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
本公开实施例提供的一种视频的编辑装置,通过文本确定模块通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;通过区间标识模块在所述目标视频的编辑界面上展示所 述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;通过区间调整模块响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;通过片段删除模块响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。利用该装置,通过对目标视频中的音频进行语音识别,能够对目标视频中的无效文本以及无效文本的位置进行确定,进而通过在目标视频的编辑轨道片段上标识无效文本的时间线区间,能够响应于用户对无效文本的调整操作和视频片段删除操作,对无效片段进行删除,从而完成对视频的剪辑,提升了视频剪辑的效率。
可选的,所述区间调整模块330具体包括下述至少之一:
第一响应单元,用于响应于所述无效文本的第一调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间的长度进行调整;
第二响应单元,用于响应于所述无效文本的第二调整操作,对所述目标视频的编辑轨道片段上所标识的时间线区间的数量进行调整。
可选的,所述区间标识模块320具体用于:
在所述编辑界面的轨道区域展示所述目标视频的编辑轨道片段,并在所述编辑界面的文本区域展示所述语音文本中的无效文本和非无效文本,其中,所述无效文本在所述编辑界面上展示为选中状态,所述非无效文本在所述编辑界面上展示为非选中状态。
可选的,所述第一响应单元包括:
调整子单元,用于在所述目标视频的编辑轨道片段上对第一文本的第一时间线区间的长度进行调整,并更新所述文本区域中所展示的目标文本,所述目标文本包括所述第一文本和/或与所述第一文 本相关联的第二文本。
可选的,所述第一调整操作包括长度延长操作,所述调整子单元具体用于:
将所述第二文本中与所述第一时间线区间的延长部分对应的文本内容移动至所述第一文本中。
可选的,所述第一调整操作包括长度缩短操作,所述调整子单元具体用于:
将所述第一文本中与所述第一时间线区间的缩短部分对应的文本内容移动至所述第二文本中;或者
在所述文本区域中新增第二文本,并将所述第一文本中与所述第一时间线区间的缩短部分对应的文本内容移动至所述第二文本中。
可选的,所述无效文本为所述轨道区域中处于选中状态的文本,所述第二响应单元具体用于下述至少之一:
响应于对所述文本区域中的第三文本的选中操作,将所述第三文本由未选中状态切换为选中状态,并在所述编辑轨道片段上增加标识所述第三文本的时间线区间;
响应于对所述文本区域中的第四文本的取消选中操作,将所述第四文本由选中状态切换为未选中状态,并在所述编辑轨道片段上取消标识所述第四文本的时间线区间。
可选的,本公开实施例提供的一种视频的编辑装置,还包括:
播放模块,用于在所述编辑界面的视频区域播放所述目标视频;
第一进度调整模块,用于当所述目标视频播放至第二时间线区间的起点位置时,将所述目标视频的播放进度调整至第一时间节点,并从所述第一时间节点继续播放所述目标视频,其中,所述第二时间线区间为所述无效文本的时间线区间,所述第一时间节点为所述第二时间线区间的终点位置在所述目标视频中所对应的时间节点。
可选的,本公开实施例提供的一种视频的编辑装置,还包括:
第一响应模块,用于在所述编辑界面的视频区域播放所述目标 视频之后,响应于对第五文本的触发操作,将所述目标视频的播放进度调整至第二时间节点,并从所述第二时间节点继续播放所述目标视频,所述第五文本显示于所述编辑界面的文本区域内,所述第五文本为无效文本或非无效文本,所述第二时间节点为所述第五文本的起点位置在所述目标视频中所对应的时间节点。
可选的,所述编辑轨道片段上还展示有时间标识,本公开实施例提供的一种视频的编辑装置,还包括:
第二响应模块,用于在所述编辑界面的视频区域播放所述目标视频之后,响应于对所述目标标识的位置调整操作,确定所述目标标识的目标位置所位于的第三时间线区间,所述目标位置为所述位置调整操作执行完毕时所述目标标识在所述编辑轨道片段上的位置;
第二进度调整模块,用于如果所述第三时间线区间为所述无效文本的时间线区间,则将所述目标视频的播放进度调整至第三时间节点,并从所述第三时间节点继续播放所述目标视频,其中,所述第三时间节点为所述第三时间线区间的起点位置或终点位置在所述目标视频中所对应的时间节点,或者,接收到所述位置进度调整操作之前所述目标视频最后播放至的时间节点;
第三进度调整模块,用于如果所述第三时间线区间不为所述无效文本的时间线区间,则将所述目标视频的播放进度调整至第四时间节点,并从所述第四时间节点继续播放所述目标视频,其中,所述第四时间节点为所述目标位置在所述目标视频中所对应的时间节点。
可选的,所述片段删除模块340具体用于:
将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,并以删除的所述视频片段为分割点,将所述目标视频分割为多个独立的视频片段;或者
将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,并将删除所述视频片段后的目标视 频整合为一个视频。
可选的,所述编辑轨道片段为所述目标视频的音频轨道片段。
上述视频的编辑装置可执行本公开任意实施例所提供的视频的编辑方法,具备执行方法相应的功能模块和有益效果。
下面参考图10,其示出了适于用来实现本公开实施例的电子设备400的结构示意图。本公开实施例中的终端设备可以包括但不限于诸如移动电话、笔记本电脑、数字广播接收器、PDA(个人数字助理)、PAD(平板电脑)、PMP(便携式多媒体播放器)、车载终端(例如车载导航终端)等等的移动终端以及诸如数字TV、台式计算机等等的固定终端。图10示出的电子设备仅仅是一个示例,不应对本公开实施例的功能和使用范围带来任何限制。
如图10所示,电子设备400可以包括处理装置(例如中央处理器、图形处理器等)401,其可以根据存储在只读存储器(ROM)402中的程序或者从存储装置408加载到随机访问存储器(RAM)403中的程序而执行各种适当的动作和处理。在RAM 403中,还存储有电子设备400操作所需的各种程序和数据。处理装置401、ROM 402以及RAM 403通过总线404彼此相连。输入/输出(I/O)接口405也连接至总线404。
通常,以下装置可以连接至I/O接口405:包括例如触摸屏、触摸板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置406;包括例如液晶显示器(LCD)、扬声器、振动器等的输出装置407;包括例如磁带、硬盘等的存储装置408;以及通信装置409。通信装置409可以允许电子设备400与其他设备进行无线或有线通信以交换数据。虽然图10示出了具有各种装置的电子设备400,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。
特别地,根据本公开的实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的实施例包括一种计算 机程序产品,其包括承载在非暂态计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的实施例中,该计算机程序可以通过通信装置409从网络上被下载和安装,或者从存储装置408被安装,或者从ROM 402被安装。在该计算机程序被处理装置401执行时,执行本公开实施例的方法中限定的上述功能。
需要说明的是,本公开上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频)等等,或者上述的任意合适的组合。
在一些实施方式中,客户端、服务器可以利用诸如HTTP(HyperText Transfer Protocol,超文本传输协议)之类的任何当前已知或未来研发的网络协议进行通信,并且可以与任意形式或介质的 数字数据通信(例如,通信网络)互连。通信网络的示例包括局域网(“LAN”),广域网(“WAN”),网际网(例如,互联网)以及端对端网络(例如,ad hoc端对端网络),以及任何当前已知或未来研发的网络。
上述计算机可读介质可以是上述电子设备中所包含的;也可以是单独存在,而未装配入该电子设备中。
上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备:
通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;
在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;
响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;
响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的操作的计算机程序代码,上述程序设计语言包括但不限于面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域 网(LAN)或广域网(WAN)—连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开实施例中所涉及到的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。其中,模块的名称在某种情况下并不构成对该单元本身的限定。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准产品(ASSP)、片上系统(SOC)、复杂可编程逻辑设备(CPLD)等等。
在本公开的上下文中,机器可读介质可以是有形的介质,其可以包含或存储以供指令执行系统、装置或设备使用或与指令执行系统、装置或设备结合地使用的程序。机器可读介质可以是机器可读信号介质或机器可读储存介质。机器可读介质可以包括但不限于电子的、磁性的、光学的、电磁的、红外的、或半导体系统、装置或设备,或者上述内容的任何合适组合。机器可读存储介质的更具体示例会包括基于一个或多个线的电气连接、便携式计算机盘、硬盘、随机存取存储器(RAM)、只读存储器(ROM)、可擦除可编程只 读存储器(EPROM或快闪存储器)、光纤、便捷式紧凑盘只读存储器(CD-ROM)、光学储存设备、磁储存设备、或上述内容的任何合适组合。
根据本公开的一个或多个实施例,示例1提供了一种视频的编辑方法,包括:
通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;
在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;
响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;
响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
根据本公开的一个或多个实施例,示例2根据示例1所述的方法,所述响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整,包括下述至少之一:
响应于所述无效文本的第一调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间的长度进行调整;
响应于所述无效文本的第二调整操作,对所述目标视频的编辑轨道片段上所标识的时间线区间的数量进行调整。
根据本公开的一个或多个实施例,示例3根据示例2所述的方法,所述在所述目标视频的编辑界面上展示所述目标视频的编辑轨 道片段,包括:
在所述编辑界面的轨道区域展示所述目标视频的编辑轨道片段,并在所述编辑界面的文本区域展示所述语音文本中的无效文本和非无效文本,其中,所述无效文本在所述编辑界面上展示为选中状态,所述非无效文本在所述编辑界面上展示为非选中状态。
根据本公开的一个或多个实施例,示例4根据示例3所述的方法,所述在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间的长度进行调整,包括:
在所述目标视频的编辑轨道片段上对第一文本的第一时间线区间的长度进行调整,并更新所述文本区域中所展示的目标文本,所述目标文本包括所述第一文本和/或与所述第一文本相关联的第二文本。
根据本公开的一个或多个实施例,示例5根据示例4所述的方法,所述第一调整操作包括长度延长操作,所述更新所述文本区域中所展示的目标文本,包括:
将所述第二文本中与所述第一时间线区间的延长部分对应的文本内容移动至所述第一文本中。
根据本公开的一个或多个实施例,示例6根据示例4所述的方法,所述第一调整操作包括长度缩短操作,所述更新所述文本区域中所展示的目标文本,包括:
将所述第一文本中与所述第一时间线区间的缩短部分对应的文本内容移动至所述第二文本中;或者
在所述文本区域中新增第二文本,并将所述第一文本中与所述第一时间线区间的缩短部分对应的文本内容移动至所述第二文本中。
根据本公开的一个或多个实施例,示例7根据示例3所述的方法,所述无效文本为所述轨道区域中处于选中状态的文本,所述响应于所述无效文本的第二调整操作,对所述目标视频的编辑轨道片段上所标识的时间线区间的数量进行调整,包括下述至少之一:
响应于对所述文本区域中的第三文本的选中操作,将所述第三文本由未选中状态切换为选中状态,并在所述编辑轨道片段上增加标识所述第三文本的时间线区间;
响应于对所述文本区域中的第四文本的取消选中操作,将所述第四文本由选中状态切换为未选中状态,并在所述编辑轨道片段上取消标识所述第四文本的时间线区间。
根据本公开的一个或多个实施例,示例8根据示例1-7任一所述的方法,还包括:
在所述编辑界面的视频区域播放所述目标视频;
当所述目标视频播放至第二时间线区间的起点位置时,将所述目标视频的播放进度调整至第一时间节点,并从所述第一时间节点继续播放所述目标视频,其中,所述第二时间线区间为所述无效文本的时间线区间,所述第一时间节点为所述第二时间线区间的终点位置在所述目标视频中所对应的时间节点。
根据本公开的一个或多个实施例,示例9根据示例8所述的方法,在所述编辑界面的视频区域播放所述目标视频之后,还包括:
响应于对第五文本的触发操作,将所述目标视频的播放进度调整至第二时间节点,并从所述第二时间节点继续播放所述目标视频,所述第五文本显示于所述编辑界面的文本区域内,所述第五文本为无效文本或非无效文本,所述第二时间节点为所述第五文本的起点位置在所述目标视频中所对应的时间节点。
根据本公开的一个或多个实施例,示例10根据示例8所述的方法,所述编辑轨道片段上还展示有时间标识,在所述编辑界面的视频区域播放所述目标视频之后,还包括:
响应于对所述目标标识的位置调整操作,确定所述目标标识的目标位置所位于的第三时间线区间,所述目标位置为所述位置调整操作执行完毕时所述目标标识在所述编辑轨道片段上的位置;
如果所述第三时间线区间为所述无效文本的时间线区间,则将所述目标视频的播放进度调整至第三时间节点,并从所述第三时间 节点继续播放所述目标视频,其中,所述第三时间节点为所述第三时间线区间的起点位置或终点位置在所述目标视频中所对应的时间节点,或者,接收到所述位置进度调整操作之前所述目标视频最后播放至的时间节点;
如果所述第三时间线区间不为所述无效文本的时间线区间,则将所述目标视频的播放进度调整至第四时间节点,并从所述第四时间节点继续播放所述目标视频,其中,所述第四时间节点为所述目标位置在所述目标视频中所对应的时间节点。
根据本公开的一个或多个实施例,示例11根据示例1-7任一所述的方法,所述将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,包括:
将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,并以删除的所述视频片段为分割点,将所述目标视频分割为多个独立的视频片段;或者
将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,并将删除所述视频片段后的目标视频整合为一个视频。
根据本公开的一个或多个实施例,示例12根据示例1-7任一所述的方法,所述编辑轨道片段为所述目标视频的音频轨道片段。
以上描述仅为本公开的较佳实施例以及对所运用技术原理的说明。本领域技术人员应当理解,本公开中所涉及的公开范围,并不限于上述技术特征的特定组合而成的技术方案,同时也应涵盖在不脱离上述公开构思的情况下,由上述技术特征或其等同特征进行任意组合而形成的其它技术方案。例如上述特征与本公开中公开的(但不限于)具有类似功能的技术特征进行互相替换而形成的技术方案。
此外,虽然采用特定次序描绘了各操作,但是这不应当理解为要求这些操作以所示出的特定次序或以顺序次序执行来执行。在一定环境下,多任务和并行处理可能是有利的。同样地,虽然在上面论述中包含了若干具体实现细节,但是这些不应当被解释为对本公 开的范围的限制。在单独的实施例的上下文中描述的某些特征还可以组合地实现在单个实施例中。相反地,在单个实施例的上下文中描述的各种特征也可以单独地或以任何合适的子组合的方式实现在多个实施例中。
尽管已经采用特定于结构特征和/或方法逻辑动作的语言描述了本主题,但是应当理解所附权利要求书中所限定的主题未必局限于上面描述的特定特征或动作。相反,上面所描述的特定特征和动作仅仅是实现权利要求书的示例形式。

Claims (15)

  1. 一种视频的编辑方法,包括:
    通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;
    在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;
    响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;
    响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
  2. 根据权利要求1所述的方法,其中所述响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整,包括下述至少之一:
    响应于所述无效文本的第一调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间的长度进行调整;
    响应于所述无效文本的第二调整操作,对所述目标视频的编辑轨道片段上所标识的时间线区间的数量进行调整。
  3. 根据权利要求2所述的方法,其中所述在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,包括:
    在所述编辑界面的轨道区域展示所述目标视频的编辑轨道片段,并在所述编辑界面的文本区域展示所述语音文本中的无效文本和非无效文本,其中,所述无效文本在所述编辑界面上展示为选中状态,所述非无效文本在所述编辑界面上展示为非选中状态。
  4. 根据权利要求3所述的方法,其中所述在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间的长度进行调整,包括:
    在所述目标视频的编辑轨道片段上对第一文本的第一时间线区间的长度进行调整,并更新所述文本区域中所展示的目标文本,所述目标文本包括所述第一文本和/或与所述第一文本相关联的第二文本。
  5. 根据权利要求4所述的方法,其中所述第一调整操作包括长度延长操作,所述更新所述文本区域中所展示的目标文本,包括:
    将所述第二文本中与所述第一时间线区间的延长部分对应的文本内容移动至所述第一文本中。
  6. 根据权利要求4所述的方法,其中所述第一调整操作包括长度缩短操作,所述更新所述文本区域中所展示的目标文本,包括:
    将所述第一文本中与所述第一时间线区间的缩短部分对应的文本内容移动至所述第二文本中;或者
    在所述文本区域中新增第二文本,并将所述第一文本中与所述第一时间线区间的缩短部分对应的文本内容移动至所述第二文本中。
  7. 根据权利要求3所述的方法,其中所述无效文本为所述轨道区域中处于选中状态的文本,所述响应于所述无效文本的第二调整操作,对所述目标视频的编辑轨道片段上所标识的时间线区间的数量进行调整,包括下述至少之一:
    响应于对所述文本区域中的第三文本的选中操作,将所述第三文本由未选中状态切换为选中状态,并在所述编辑轨道片段上增加标识所述第三文本的时间线区间;
    响应于对所述文本区域中的第四文本的取消选中操作,将所述第四文本由选中状态切换为未选中状态,并在所述编辑轨道片段上取消标识所述第四文本的时间线区间。
  8. 根据权利要求1-7任一所述的方法,还包括:
    在所述编辑界面的视频区域播放所述目标视频;
    当所述目标视频播放至第二时间线区间的起点位置时,将所述目标视频的播放进度调整至第一时间节点,并从所述第一时间节点继续播放所述目标视频,其中,所述第二时间线区间为所述无效文本的时间线区间,所述第一时间节点为所述第二时间线区间的终点位置在所述目标视频中所对应的时间节点。
  9. 根据权利要求8所述的方法,其中在所述编辑界面的视频区域播放所述目标视频之后,还包括:
    响应于对第五文本的触发操作,将所述目标视频的播放进度调整至第二时间节点,并从所述第二时间节点继续播放所述目标视频,所述第五文本显示于所述编辑界面的文本区域内,所述第五文本为无效文本或非无效文本,所述第二时间节点为所述第五文本的起点位置在所述目标视频中所对应的时间节点。
  10. 根据权利要求8所述的方法,其中所述编辑轨道片段上还展示有时间标识,在所述编辑界面的视频区域播放所述目标视频之后,还包括:
    响应于对所述目标标识的位置调整操作,确定所述目标标识的目标位置所位于的第三时间线区间,所述目标位置为所述位置调整操作执行完毕时所述目标标识在所述编辑轨道片段上的位置;
    如果所述第三时间线区间为所述无效文本的时间线区间,则将所述目标视频的播放进度调整至第三时间节点,并从所述第三时间节点继续播放所述目标视频,其中,所述第三时间节点为所述第三时间线区间的起点位置或终点位置在所述目标视频中所对应的时间节点,或者,接收到所述位置进度调整操作之前所述目标视频最后播放至的时间节点;
    如果所述第三时间线区间不为所述无效文本的时间线区间,则将所述目标视频的播放进度调整至第四时间节点,并从所述第四时间节点继续播放所述目标视频,其中,所述第四时间节点为所述目标位置在所述目标视频中所对应的时间节点。
  11. 根据权利要求1-7任一所述的方法,其中所述将所述编辑轨 道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,包括:
    将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,并以删除的所述视频片段为分割点,将所述目标视频分割为多个独立的视频片段;或者
    将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,并将删除所述视频片段后的目标视频整合为一个视频。
  12. 根据权利要求1-7任一所述的方法,其中所述编辑轨道片段为所述目标视频的音频轨道片段。
  13. 一种视频的编辑装置,包括:
    文本确定模块,用于通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;
    区间标识模块,用于在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;
    区间调整模块,用于响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;
    片段删除模块,用于响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
  14. 一种电子设备,包括:
    至少一个处理器;以及
    与所述至少一个处理器通信连接的存储器;其中,
    所述存储器存储有可被所述至少一个处理器执行的计算机程序,所述计算机程序被所述至少一个处理器执行,以使所述至少一个处理器能够执行权利要求1-12中任一项所述的视频的编辑方法。
  15. 一种计算机可读存储介质,所述计算机可读存储介质存储有计算机指令,所述计算机指令用于使处理器执行时实现权利要求1-12中任一项所述的视频的编辑方法。
PCT/CN2023/138239 2023-01-19 2023-12-12 视频的编辑方法、装置、电子设备和存储介质 Ceased WO2024152802A1 (zh)

Priority Applications (4)

Application Number Priority Date Filing Date Title
JP2023578854A JP7680578B2 (ja) 2023-01-19 2023-12-12 ビデオの編集方法、装置、電子機器及び記憶媒体
EP23821101.5A EP4429256A4 (en) 2023-01-19 2023-12-12 VIDEO EDITING METHOD AND APPARATUS, ELECTRONIC DEVICE AND RECORDING MEDIUM
KR1020257018467A KR20250105651A (ko) 2023-01-19 2023-12-12 비디오 편집 방법 및 장치, 그리고 전자 디바이스 및 저장 매체
US18/542,524 US12051445B1 (en) 2023-01-19 2023-12-15 Method and apparatus of video editing, and electronic device and storage medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202310107566.3A CN118368478A (zh) 2023-01-19 2023-01-19 视频的编辑方法、装置、电子设备和存储介质
CN202310107566.3 2023-01-19

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US18/542,524 Continuation US12051445B1 (en) 2023-01-19 2023-12-15 Method and apparatus of video editing, and electronic device and storage medium

Publications (1)

Publication Number Publication Date
WO2024152802A1 true WO2024152802A1 (zh) 2024-07-25

Family

ID=89430577

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2023/138239 Ceased WO2024152802A1 (zh) 2023-01-19 2023-12-12 视频的编辑方法、装置、电子设备和存储介质

Country Status (2)

Country Link
CN (1) CN118368478A (zh)
WO (1) WO2024152802A1 (zh)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20120323897A1 (en) * 2011-06-14 2012-12-20 Microsoft Corporation Query-dependent audio/video clip search result previews
CN109040773A (zh) * 2018-07-10 2018-12-18 武汉斗鱼网络科技有限公司 一种视频改进方法、装置、设备及介质
CN113613068A (zh) * 2021-08-03 2021-11-05 北京字跳网络技术有限公司 视频的处理方法、装置、电子设备和存储介质
CN114999530A (zh) * 2022-05-18 2022-09-02 北京飞象星球科技有限公司 音视频剪辑方法及装置
CN115623279A (zh) * 2021-07-15 2023-01-17 脸萌有限公司 多媒体处理方法、装置、电子设备及存储介质

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20120323897A1 (en) * 2011-06-14 2012-12-20 Microsoft Corporation Query-dependent audio/video clip search result previews
CN109040773A (zh) * 2018-07-10 2018-12-18 武汉斗鱼网络科技有限公司 一种视频改进方法、装置、设备及介质
CN115623279A (zh) * 2021-07-15 2023-01-17 脸萌有限公司 多媒体处理方法、装置、电子设备及存储介质
CN113613068A (zh) * 2021-08-03 2021-11-05 北京字跳网络技术有限公司 视频的处理方法、装置、电子设备和存储介质
CN114999530A (zh) * 2022-05-18 2022-09-02 北京飞象星球科技有限公司 音视频剪辑方法及装置

Also Published As

Publication number Publication date
CN118368478A (zh) 2024-07-19

Similar Documents

Publication Publication Date Title
JP7572108B2 (ja) 議事録のインタラクション方法、装置、機器及び媒体
US12573429B2 (en) Video processing method and apparatus, electronic device and storage medium
CN111970577A (zh) 字幕编辑方法、装置和电子设备
CN114679628B (zh) 一种弹幕添加方法、装置、电子设备和存储介质
JP7334362B2 (ja) 文書内テーブル閲覧方法、装置、電子機器及び記憶媒体
US20240289398A1 (en) Method, apparatus, device and storage medium for content display
CN113395568B (zh) 视频互动方法、装置、电子设备和存储介质
CN115061602B (zh) 页面显示方法、装置、设备、计算机可读存储介质及产品
US20250106486A1 (en) Multimedia data processing method and apparatus, device and medium
JP2024525192A (ja) オーディオ処理方法、装置、電子機器及び記憶媒体
US20250173045A1 (en) Book information display method and apparatus, device, and storage medium
WO2024140239A1 (zh) 页面显示方法、装置、设备、计算机可读存储介质及产品
US20250363293A1 (en) Method for generating table field content, storage medium, and electronic device
WO2025185040A1 (zh) 媒体内容的生成方法、装置、电子设备、存储介质和程序产品
WO2024099376A1 (zh) 视频编辑方法、装置、设备及介质
CN117857829A (zh) 内容生成方法、装置、可读介质及电子设备
US20240103802A1 (en) Method, apparatus, device and medium for multimedia processing
CN118170297A (zh) 特效编辑方法、装置、电子设备、存储介质及程序产品
US12051445B1 (en) Method and apparatus of video editing, and electronic device and storage medium
WO2026067488A1 (zh) 互动信息的发布方法、装置、电子设备、存储介质和程序产品
WO2025130695A1 (zh) 特效创作方法、装置、设备、计算机可读存储介质及产品
WO2026056410A1 (zh) 信息显示方法、装置、电子设备、存储介质和程序产品
WO2025252119A1 (zh) 作品展示方法、装置、电子设备、存储介质和程序产品
CN118368478A (zh) 视频的编辑方法、装置、电子设备和存储介质
WO2025093007A1 (zh) 信息回复方法、装置、电子设备和存储介质

Legal Events

Date Code Title Description
WWE Wipo information: entry into national phase

Ref document number: 2023578854

Country of ref document: JP

ENP Entry into the national phase

Ref document number: 2023821101

Country of ref document: EP

Effective date: 20231218

121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23821101

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 202527050779

Country of ref document: IN

WWE Wipo information: entry into national phase

Ref document number: 1020257018467

Country of ref document: KR

WWP Wipo information: published in national office

Ref document number: 202527050779

Country of ref document: IN

WWE Wipo information: entry into national phase

Ref document number: 11202503525X

Country of ref document: SG

WWP Wipo information: published in national office

Ref document number: 11202503525X

Country of ref document: SG

WWP Wipo information: published in national office

Ref document number: 1020257018467

Country of ref document: KR

REG Reference to national code

Ref country code: BR

Ref legal event code: B01A

Ref document number: 112025011769

Country of ref document: BR

NENP Non-entry into the national phase

Ref country code: DE

ENP Entry into the national phase

Ref document number: 112025011769

Country of ref document: BR

Kind code of ref document: A2

Effective date: 20250610