WO2024152802A1 - 视频的编辑方法、装置、电子设备和存储介质 - Google Patents
视频的编辑方法、装置、电子设备和存储介质 Download PDFInfo
- Publication number
- WO2024152802A1 WO2024152802A1 PCT/CN2023/138239 CN2023138239W WO2024152802A1 WO 2024152802 A1 WO2024152802 A1 WO 2024152802A1 CN 2023138239 W CN2023138239 W CN 2023138239W WO 2024152802 A1 WO2024152802 A1 WO 2024152802A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- text
- target video
- invalid
- video
- timeline
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/04—Segmentation; Word boundary detection
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/18—Speech classification or search using natural language modelling
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/26—Speech to text systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/83—Generation or processing of protective or descriptive data associated with content; Content structuring
- H04N21/845—Structuring of content, e.g. decomposing content into time segments
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/83—Generation or processing of protective or descriptive data associated with content; Content structuring
- H04N21/845—Structuring of content, e.g. decomposing content into time segments
- H04N21/8456—Structuring of content, e.g. decomposing content into time segments by decomposing the content in the time domain, e.g. in time segments
Definitions
- the embodiments of the present disclosure relate to the field of computer technology, and in particular, to a video editing method, device, electronic device, and storage medium.
- voice-over videos For example, users can shoot and upload voice-over videos to video applications.
- the voice-over or audio carries the core content of the voice-over video. Therefore, how to better edit the voice-over video is an urgent problem that needs to be solved.
- the editing of spoken video mainly needs to deal with the pauses, stuttering, slurred words and other unnecessary segments in the video.
- the existing editing method generally identifies invalid segments through the waveform of the audio in the video, but this method takes a lot of time and has low editing efficiency.
- the embodiments of the present disclosure provide a video editing method, device, electronic device and storage medium to improve the efficiency of video editing.
- an embodiment of the present disclosure provides a method for editing a video, comprising:
- the video segment on the editing track segment that is within the timeline interval of the invalid text is deleted from the target video.
- the present disclosure also provides a video editing device, including:
- a text determination module used to determine invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video, wherein the timeline position of the invalid text is used to indicate the appearance time of the speech audio of the invalid text in the target video;
- An interval identification module is used to display the editing track segment of the target video on the editing interface of the target video, and identify the timeline interval of the invalid text on the editing track segment based on the timeline position of the invalid text, wherein the timeline interval of the invalid text is the time interval in which the voice audio of the invalid text appears in the target video;
- An interval adjustment module configured to adjust the timeline interval of the invalid text on the editing track segment of the target video in response to an adjustment operation of the invalid text
- the segment deletion module is used to delete the video segment on the editing track segment within the timeline interval of the invalid text from the target video in response to the video segment deletion operation of the invalid text.
- an embodiment of the present disclosure further provides an electronic device, including:
- processors one or more processors
- a memory for storing one or more programs
- the one or more processors When the one or more programs are executed by the one or more processors, the one or more processors implement the video editing method as described in the embodiment of the present disclosure.
- the embodiments of the present disclosure further provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video editing method as described in the embodiments of the present disclosure.
- the embodiments of the present disclosure provide a method, device, electronic device and storage medium for editing a video, the method comprising: determining invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video, the timeline position of the invalid text being used to indicate the appearance time of the speech audio of the invalid text in the target video; displaying an editing track segment of the target video on an editing interface of the target video, and marking a timeline interval of the invalid text on the editing track segment based on the timeline position of the invalid text, the timeline interval of the invalid text being the time interval when the speech audio of the invalid text appears in the target video; in response to an adjustment operation of the invalid text, adjusting the timeline interval of the invalid text on the editing track segment of the target video; and in response to a video segment deletion operation of the invalid text, deleting the video segment on the editing track segment that is within the timeline interval of the invalid text from the target video.
- FIG1 is a schematic flow chart of a video editing method provided by an embodiment of the present disclosure
- FIG2 is a schematic diagram of triggering speech recognition provided by an embodiment of the present disclosure
- FIG3 is a schematic diagram of an editing interface provided by an embodiment of the present disclosure.
- FIG4 is a schematic flow chart of another video editing method provided by an embodiment of the present disclosure.
- FIG5 is a schematic diagram of another editing interface provided by an embodiment of the present disclosure.
- FIG6 is a schematic diagram of updating a target text provided by an embodiment of the present disclosure.
- FIG7 is a schematic diagram of another method for updating a target text provided by an embodiment of the present disclosure.
- FIG8 is a schematic diagram of another method for updating a target text provided by an embodiment of the present disclosure.
- FIG9 is a schematic structural diagram of a video editing device provided by an embodiment of the present disclosure.
- FIG. 10 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure.
- FIG1 is a flow chart of a method for editing a video provided in an embodiment of the present disclosure.
- the method is applicable to video editing.
- the method can be executed by a video editing device, wherein the device can be implemented by software and/or hardware and is generally integrated on an electronic device.
- the electronic device includes but is not limited to: mobile phones, computers and other devices.
- a video editing method provided by an embodiment of the present disclosure includes the following steps:
- S110 Determine invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video.
- the timeline position of the invalid text is used to indicate the appearance time of the speech audio of the invalid text in the target video.
- the target video may refer to a video that needs to be edited, such as a video uploaded by a user or a video shot online by a user.
- the target video contains audio
- the voice text may be considered as text information corresponding to the audio of the target video, which may be obtained by performing voice recognition on the audio of the target video.
- the voice text includes invalid text
- the invalid text may include text corresponding to the pauses, repetitions, and unnecessary words in the target video that the user does not need; the timeline position of the invalid text may be used to indicate the appearance time of the voice audio of the invalid text in the target video.
- the target video is a continuous video
- the target video may be a continuous and complete video, or may be a segment of a complete video.
- the invalid text and the timeline position of the invalid text in the speech text of the target video can be determined by performing speech recognition on the audio in the target video.
- This step does not limit the specific method of voice recognition and the timing of triggering voice recognition.
- a certain recognition control can be triggered in the interface to trigger voice recognition of the audio in the target video, so that invalid text in the target video and the timeline position of the invalid text can be automatically identified.
- the position and style of the recognition control are not limited and can be set according to the actual page situation.
- the identification conditions of invalid text are: a segment with a default silence duration greater than a preset duration (such as 120ms, etc.) is a pause segment; invalid words (such as repetitions and modal particles) are identified by reusing a preset algorithm logic, and the preset algorithm can be configured in the current application or in the server.
- a preset duration such as 120ms, etc.
- FIG2 is a schematic diagram of triggering voice recognition provided by an embodiment of the present disclosure. As shown in FIG2, after selecting a video (such as a target video), you can trigger control 1 in the pop-up window by right-clicking to trigger voice recognition of the audio in the target video. You can also click the shortcut control 2 in the current interface to realize voice recognition of the audio in the target video.
- a video such as a target video
- the editing interface can be used to edit clips in the target video.
- the editing track clip of the target video can be understood as the editing track clip corresponding to the target video, such as an editing track clip whose starting point corresponds to the starting point of the target video and whose end point corresponds to the end point of the target video.
- the editing track clip can be the editing track of the continuous and complete video; when the target video is a clip of a complete video, the editing track clip can be the track clip in the editing track of the complete video that corresponds to the target video.
- the editing track clip is the audio track clip of the target video, and the audio track clip can refer to the audio track clip corresponding to the target video.
- the timeline interval of invalid text can be considered as the time interval in which the voice audio of the invalid text appears in the target video.
- the invalid text can be displayed in the target video.
- the editing track segment of the target video is displayed on the editing interface, and the timeline interval of the invalid text is marked on the editing track segment based on the timeline position of the invalid text.
- a preset mark can be used to mark the timeline interval of the invalid text in the editing track segment of the target video.
- the preset mark can be specifically set according to, for example, the preset mark can be a rectangular frame.
- a rectangular frame can be used to frame the timeline interval of the invalid text in the editing track segment of the target video; or, the timeline interval of the invalid text can be marked as a first display style in the editing track segment, and the timeline interval of other texts can be marked as a second display style in the editing track segment, and the first display style is different from the second display style, such as the first display style and the second display style can be distinguished by different colors, or can be distinguished by different lines, etc.
- the timeline intervals of all invalid texts identified can be marked on the editing track segment by default to distinguish the timeline intervals of invalid texts from the timeline intervals of other texts; or a set number of timeline intervals of invalid texts can be marked on the editing track segment by default, and the set number can be pre-set by relevant personnel. Users can adjust the timeline intervals of invalid texts as needed.
- Figure 3 is a schematic diagram of an editing interface provided by an embodiment of the present disclosure.
- the editing track segment 3 of the target video can be displayed on the editing interface of the target video, and based on the timeline position of the invalid text identified in the previous step, the timeline interval of the invalid text is marked on the editing track segment 3, as shown in Figure 3, the timeline interval 4 of the invalid text is marked with a black rectangular frame.
- the invalid text adjustment operation may refer to an operation for triggering the adjustment of the timeline interval of the invalid text.
- the specific operation method of the invalid text adjustment operation in this embodiment is not limited.
- the invalid text adjustment operation may be an operation of clicking a certain control in the editing interface or a preset triggering operation acting on the editing track segment, or an operation of performing a preset gesture in the editing interface.
- the preset gesture may be a pre-set gesture, such as a preset gesture. For a right swipe gesture, etc.
- This embodiment does not limit the triggering position of the adjustment operation. For example, it can act on the timeline interval of invalid text, or it can act on other positions, as long as it can trigger the adjustment of the timeline interval of invalid text.
- the present embodiment can adjust the timeline interval of the invalid text on the editing track segment of the target video in response to the adjustment operation of the invalid text.
- the specific method of adjusting the timeline interval may vary according to the different adjustment operations. For example, when the adjustment operation of the invalid text is to trigger a control to change the state of the invalid text from a selected state to an unselected state, the number of timeline intervals of the invalid text can be adjusted on the editing track segment of the target video, such as adding a timeline interval marking a certain text on the editing track segment to adjust the text to invalid text, or canceling the time interval marking a certain text on the editing track segment to adjust the text to non-invalid text.
- the invalid text video clip deletion operation may refer to an operation for triggering deletion of the video clip corresponding to the invalid text on the editing track segment.
- the invalid text video clip deletion operation may be an operation of clicking a deletion control in the editing interface.
- This step can respond to the video segment deletion operation of the invalid text, and delete the video segment in the timeline interval of the invalid text on the editing track segment from the target video.
- the specific deletion method is not limited.
- the trace corresponding to the timeline interval of the invalid text on the editing track segment can be retained or removed according to the user's needs.
- the user triggering control 5 can be considered as a video segment deletion operation.
- the electronic device can respond to the video segment deletion operation of the invalid text and delete the video segment corresponding to the timeline interval 6 of the invalid text from the target video.
- deleting the video segment on the editing track segment within the timeline interval of the invalid text from the target video includes:
- the video segment on the editing track segment that is located in the timeline interval of the invalid text is deleted from the target video, and the deleted video segment is used as the segmentation point, Splitting the target video into multiple independent video segments; or
- the video segment on the editing track segment that is located in the timeline interval of the invalid text is deleted from the target video, and the target video after the video segment is deleted is integrated into one video.
- the video segment located in the timeline interval of the invalid text on the editing track segment can be deleted from the target video, and the target video can be divided into multiple independent video segments with the deleted video segment as the division point. On this basis, the user can operate each independent video segment as needed.
- the video segments in the timeline interval of the invalid text on the editing track segment can be deleted from the target video, and the target video after the video segments are deleted can be integrated into one video.
- the operation of deleting the invalid segments in the target video is realized, and the quality of the target video is improved.
- the invalid segment is the segment corresponding to the timeline interval of the invalid text, in other words, the text corresponding to the timeline interval marked on the editing track segment in the speech text of the target video.
- the embodiment of the present disclosure provides a video editing method, which determines invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video, wherein the timeline position of the invalid text is used to indicate the appearance time of the speech audio of the invalid text in the target video; displays the editing track segment of the target video on the editing interface of the target video, and identifies the timeline interval of the invalid text on the editing track segment based on the timeline position of the invalid text, wherein the timeline interval of the invalid text is the time interval when the speech audio of the invalid text appears in the target video; in response to the adjustment operation of the invalid text, adjusts the timeline interval of the invalid text on the editing track segment of the target video; in response to the video segment deletion operation of the invalid text, deletes the video segment on the editing track segment that is located in the timeline interval of the invalid text from the target video.
- the invalid text and the position of the invalid text in the target video can be determined by performing speech recognition on the audio in the target video, and then by identifying the timeline interval of the invalid text on the editing track segment of the target video, the invalid text can be deleted in response to the user's adjustment operation on the invalid text and the video segment deletion operation.
- the invalid segments are deleted through the deletion operation, thereby completing the video editing and improving the efficiency of video editing.
- the method further comprises:
- the playback progress of the target video is adjusted to the first time node, and the target video continues to be played from the first time node, wherein the second timeline interval is the timeline interval of the invalid text, and the first time node is the time node corresponding to the end point of the second timeline interval in the target video.
- the video area can be considered as a certain area in the editing interface, which is used to play the target video.
- the specific position and size of the video area can be set by relevant personnel according to the situation of the editing interface.
- the second timeline interval can be any timeline interval of invalid text, and the first time node can be understood as the time node corresponding to the end position of the second timeline interval in the target video.
- the target video can be played in the video area of the editing interface.
- the timing of playing the target video is not limited.
- the target video can be played automatically when the editing interface is displayed; or after the editing interface is displayed, when the user triggers a certain playback control in the editing interface, the target video can be played in the video area of the editing interface; or the target video can be automatically played after the editing interface is displayed for a preset length of time, etc. This embodiment does not limit this.
- the present embodiment can automatically adjust the play progress of the target video to the first time node, and continue to play the target video from the first time node. On this basis, by automatically skipping invalid segments when playing the target video, the user can watch the edited video in advance to better edit the target video.
- the method after playing the target video in the video area of the editing interface, the method further includes:
- the fifth text is displayed in the text area of the editing interface, the fifth text is invalid text or non-invalid text, and the second time node is the time node corresponding to the starting position of the fifth text in the target video.
- the fifth text may be a text triggered by a user, which may be an invalid text or a non-invalid text.
- Non-invalid text is text other than invalid text, such as valid text in the voice text of the target video.
- the fifth text may be displayed in a text area of the editing interface, and the text area may be considered as a certain area in the editing interface for displaying the voice text of the target video, such as the fifth text in the voice text for displaying the target video.
- the specific position and size of the text area may be set by relevant personnel according to the situation of the editing interface.
- the invalid text in the text area is formed into a paragraph by itself, for example, the "pause" text can be formed into a separate line, and the pause duration is displayed at the same time.
- the trigger operation on the fifth text can be used to trigger the adjustment of the playback progress of the target video to the second time node.
- the trigger operation on the fifth text can be an operation of clicking the fifth text; the second time node is the time node corresponding to the starting position of the fifth text in the target video.
- the target video's playback progress can be adjusted to the second time node in response to the triggering operation of the fifth text, and the target video can be played continuously from the second time node.
- the playback progress of the target video is adjusted, thereby locating the corresponding text in the target video, so that the user can watch the video clip corresponding to the corresponding text.
- a time mark is also displayed on the editing track segment, and after the target video is played in the video area of the editing interface, the method further includes:
- the playback progress of the target video is adjusted to the third time node, and the target video is played from the third time node.
- the target video continues to be played at the node, wherein the third time node is the time node corresponding to the starting position or the ending position of the third timeline interval in the target video, or the time node to which the target video was last played before the position progress adjustment operation is received;
- the playback progress of the target video is adjusted to a fourth time node, and the target video continues to be played from the fourth time node, wherein the fourth time node is the time node corresponding to the target position in the target video.
- the time identifier is used to represent the time node to which the target video is currently played in the editing track segment of the target video. In other words, this time identifier can be used to indicate the time node corresponding to the current playback progress of the target video.
- the position of the target identifier in the editing track segment of the target video can be updated in real time according to the playback progress of the target video.
- the position adjustment operation of the target identifier can be understood as an operation to adjust the position of the target identifier, such as the position adjustment operation of the target identifier can be an operation of clicking on a certain position a in the editing track segment to adjust the target identifier to the position a, and the position adjustment operation of the target identifier can also be an operation of dragging the target identifier to adjust the target identifier to the position a, etc.
- the target position can be considered as the position of the target identifier on the editing track segment when the position adjustment operation is completed, and the third timeline interval is the timeline interval where the target position of the target identifier is located.
- the third time node can be the time node corresponding to the starting position of the third timeline interval in the target video, or the time node corresponding to the end position of the third timeline interval in the target video, or it can be the time node to which the target video was last played before receiving the position progress adjustment operation; the fourth time node can be considered as the time node corresponding to the target position in the target video.
- the position of the target marker may be adjusted in response to the position adjustment operation of the target marker.
- the playback progress of the target video can be adjusted according to the position of the time marker in the editing track segment.
- the third timeline interval where the target position is located is determined, and the playback progress of the target video is adjusted differently according to different third timeline intervals.
- the playback progress of the target video can be adjusted to the fourth time node, and the target video can continue to be played from the fourth time node.
- this embodiment automatically adjusts the playback progress of the target video to the end position of the timeline interval A to continue playing the target video from the end position of the timeline interval A, or restores the position of the target identifier to the last position of the target identifier before receiving the position adjustment operation, and continues to play the target video from this position.
- this embodiment automatically adjusts the playback progress of the target video to the starting position of the timeline interval A, and continues to play the target video from the starting position of the timeline interval A.
- the play progress of the target video can be adjusted to the time node corresponding to the target position in the target video, and the target video can be continued to be played from this time node.
- the target position b to which the target identifier is dragged is not located in the timeline interval of invalid text, or when the target position b to which the target identifier is moved by clicking is not located in the timeline interval of invalid text, this embodiment can adjust the play progress of the target video to the time node c corresponding to the target position b in the target video, and continue to play the target video from time node c.
- the video segment corresponding to the invalid text can be automatically skipped or the video segment corresponding to the invalid text can be played completely from the starting position; and when the position of the target identifier is adjusted not to an invalid segment, the function of playing the target video from the time node corresponding to this position.
- FIG4 is a flow chart of another video editing method provided by an embodiment of the present disclosure.
- the solution in this embodiment can be combined with one or more optional solutions in the above embodiments.
- the target The timeline interval of the invalid text is adjusted on the editing track segment of the video, including at least one of the following: in response to a first adjustment operation of the invalid text, the length of the timeline interval of the invalid text is adjusted on the editing track segment of the target video; in response to a second adjustment operation of the invalid text, the number of timeline intervals marked on the editing track segment of the target video is adjusted.
- the method includes:
- S210 Determine invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video.
- the timeline position of the invalid text is used to indicate the appearance time of the speech audio of the invalid text in the target video.
- the first adjustment operation can be used to adjust the length of the timeline interval of the invalid text.
- the first adjustment operation can be performed on the editing track segment of the target video.
- the first adjustment operation can be an operation of stretching or compressing the timeline interval B of a certain invalid text on the editing track segment of the target video, such as dragging the starting position of the timeline interval B to the left.
- This step can respond to the first adjustment operation of the invalid text and adjust the length of the timeline interval of the invalid text on the editing track segment of the target video. For example, when the left border of a timeline interval is dragged to the left or the right border of a timeline interval is dragged to the right, the interval length of the timeline interval can be extended; when the left border of a timeline interval is dragged to the right or the right border of a timeline interval is dragged to the right, the interval length of the timeline interval can be shortened.
- the second adjustment operation can be used to adjust the number of timeline intervals marked on the editing track segment.
- the second adjustment operation can act on a certain adjustment control.
- Each invalid text can correspond to an adjustment control, and the adjustment control is used to control whether the timeline interval corresponding to the invalid text is marked on the editing track segment.
- the second adjustment operation can also act on the timeline interval of the editing track segment.
- This step can adjust the number of timeline intervals marked on the editing track segment of the target video in response to the second adjustment operation of the invalid text.
- This embodiment does not limit the specific process of the adjustment, as long as the number of timeline intervals marked on the editing track segment can be adjusted.
- a video editing method provided by an embodiment of the present disclosure can adjust the length and number of timeline intervals of invalid text on an editing track segment of a target video in response to a first adjustment operation and a second adjustment operation of the invalid text, thereby enabling accurate editing of the target video according to user needs.
- displaying the editing track segment of the target video on the editing interface of the target video includes:
- the editing track segment of the target video is displayed in the track area of the editing interface, and the invalid text and non-invalid text in the voice text are displayed in the text area of the editing interface, wherein the invalid text is displayed as a selected state on the editing interface, and the non-invalid text is displayed as a non-selected state on the editing interface.
- the track area can be an area in the editing interface, which is used to display the editing track segment of the target video.
- the selected state can be understood as a text being in a selected state, thereby indicating that the text will be deleted; the unselected state can be understood as a text being in an unselected state, thereby indicating that the text will not be deleted.
- the specific process of displaying the editing track segment of the target video on the editing interface of the target video may be: displaying the editing track segment of the target video in the track area of the editing interface, and displaying the non-text in the voice text in the text area of the editing interface.
- the invalid text and non-invalid text are displayed, among which the invalid text can be displayed as selected on the editing interface, and the non-invalid text can be displayed as unselected on the editing interface.
- the invalid text and non-invalid text in the voice text are displayed in the text area of the editing interface while displaying the editing track fragment; in addition, the invalid text is displayed as selected by default, which can prompt the user to delete the invalid text and avoid the user from performing the corresponding selection operation to select the invalid text such as pause and repeat before deleting it, further simplifying the operation required for the user to delete the invalid text in the target video.
- Figure 5 is a schematic diagram of another editing interface provided by an embodiment of the present disclosure.
- the editing track segment of the target video is displayed in the track area 7 of the editing interface, and the invalid text (such as "pause 0.12s" or “repeat") and non-invalid text (such as "mine”) in the voice text are displayed in the text area 8 of the editing interface, wherein the invalid text is displayed as selected on the editing interface, and the non-invalid text is displayed as unselected on the editing interface.
- the invalid text such as "pause 0.12s" or "repeat
- non-invalid text such as "mine
- the adjusting the length of the timeline interval of the invalid text on the editing track segment of the target video includes:
- the length of a first timeline interval of a first text is adjusted on an editing track segment of the target video, and the target text displayed in the text area is updated, wherein the target text includes the first text and/or a second text associated with the first text.
- the first text may be an invalid text, such as the text belonging to the timeline interval targeted by the first trigger operation.
- the first timeline interval is the timeline interval of the first text.
- the target text may refer to the text that needs to be adjusted, such as the target text may include the first text and/or the second text associated with the first text, the second text may be the text corresponding to the timeline interval adjacent to the timeline interval of the first text, or the text corresponding to the new timeline interval after the timeline interval adjacent to the timeline interval of the first text is reduced.
- the target text displayed in the text area needs to be updated, and the specific content of the update can be determined by the specific operation of the first adjustment operation.
- the text corresponding to the adjusted part can be updated in real time in the text area to accurately adjust the length of the timeline interval corresponding to the invalid text.
- the first adjustment operation includes a length extension operation
- the updating of the target text displayed in the text area includes:
- the text content in the second text corresponding to the extended portion of the first timeline interval is moved to the first text.
- the length extension operation may be considered as an operation of extending the length of the first timeline interval.
- the length extension operation may be an operation of dragging the start position of the first timeline interval to the left or dragging the end position of the first timeline interval to the right.
- the text content in the second text corresponding to the extended portion of the first timeline interval can be synchronously moved to the first text in the text area.
- the first adjustment operation is an operation of dragging the starting position of the first timeline interval to the left
- the second text is the text corresponding to the timeline interval located to the left of the first timeline interval. At this time, the text content in the second text corresponding to the extended portion of the first timeline interval needs to be moved to the first text.
- the length extension operation may support moving part of the adjacent non-invalid text into the first text.
- the length extension operation may be supported or not supported.
- Figure 6 is a schematic diagram of updating a target text provided by an embodiment of the present disclosure.
- the end position of the first timeline interval corresponding to the first text 9 is dragged rightward from position A to position B, the length of the first timeline interval of the first text 9 can be extended on the editing track segment of the target video, and the text content (such as "sun") in the second text 10 corresponding to the extended part of the first timeline interval can be moved to the first text 9.
- the first adjustment operation includes a length shortening operation
- the updating of the target text displayed in the text area includes:
- a second text is added to the text area, and text content in the first text corresponding to the shortened portion of the first timeline interval is moved to the second text.
- the length shortening operation may be considered as an operation of shortening the length of the first timeline interval.
- the length shortening operation may be an operation of dragging the start position of the first timeline interval to the right or dragging the end position of the first timeline interval to the left.
- the update method of the target text displayed in the text area can be determined according to the different contents of the first text. For example, when the first text content is a pause, the text content in the first text corresponding to the shortened portion of the first timeline interval can be moved to the second text; and when the first text content is repeated text, invalid words, or normal text other than pauses, repetitions and invalid words, a second text can be added to the text area, and the text content in the first text corresponding to the shortened portion of the first timeline interval can be moved to the newly added second text.
- Figure 7 is a schematic diagram of another method of updating the target text provided by an embodiment of the present disclosure.
- the end position of the first timeline interval corresponding to the first text 9 is dragged leftward from position A to position C, the length of the first timeline interval of the first text 9 can be shortened on the editing track segment of the target video, and the text content in the first text 9 corresponding to the shortened part of the first timeline interval (such as "pause 0.02s") is moved to the second text 10.
- Figure 8 is a schematic diagram of another method of updating the target text provided by an embodiment of the present disclosure.
- the end position of the first timeline interval corresponding to the first text 11 is dragged rightward from position D to position E, the length of the first timeline interval of the first text 11 can be shortened on the editing track segment of the target video.
- a second text 12 can be added to the text area, and the text content in the first text 11 corresponding to the shortened part of the first timeline interval (i.e., "repeat") can be moved to the second text 12.
- the invalid text is text in the track area that is in a selected state
- the second adjustment operation in response to the invalid text is performed on the target
- the number of timeline intervals identified on the editing track segment of the video is adjusted, including at least one of the following:
- the fourth text is switched from a selected state to an unselected state, and the timeline interval marking the fourth text is deselected on the editing track segment.
- the third text may be any text in the text area, such as a text in an unselected state in the text area.
- the selection operation may refer to an operation of switching the text from an unselected state to a selected state, such as the selection operation may be an operation of clicking a control, the control may have a first display style and a second display style, and the selected state of the text may change with the change of the display style of the control, and the specific contents of the first display style and the second display style are not limited, as long as different display styles can be distinguished.
- the fourth text may be any text in the text area, such as a text in a selected state in the text area.
- the fourth text may be the same as or different from the third text.
- the deselect operation may refer to an operation of switching the text from a selected state to an unselected state, such as the deselect operation may correspond to the selection operation, which is an operation of clicking a certain control.
- this embodiment can respond to the selection operation of the third text in the text area, switch the third text from an unselected state to a selected state, and add a timeline interval identifying the third text on the editing track segment. On this basis, the user can add the third text to be deleted and the timeline interval of the third text as needed.
- the fourth text in response to the deselection operation on the fourth text in the text area, the fourth text may be switched from a selected state to an unselected state, and the timeline interval marking the fourth text may be deselected on the editing track segment.
- the user may deselect the fourth text to be deleted and deselect the timeline interval marking the fourth text as needed.
- a selection operation on the third text corresponding to the control is triggered.
- the third text can be switched from an unselected state to a selected state, and an edit track segment can be added.
- a timeline interval for marking a third text is added; when a control is clicked to display the control in the second display style, it can be considered that a deselection operation of a fourth text corresponding to the control is triggered.
- the fourth text can be switched from a selected state to an unselected state, and the timeline interval for marking the fourth text can be deselected on the editing track segment.
- FIG9 is a schematic diagram of the structure of a video editing device provided in an embodiment of the present disclosure.
- the device may be applicable to video editing, wherein the device may be implemented by software and/or hardware and is generally integrated on an electronic device.
- the device includes:
- a text determination module 310 is used to determine invalid text in the speech text of the target video and the timeline position of the invalid text by performing speech recognition on the audio in the target video, wherein the timeline position of the invalid text is used to indicate the appearance time of the speech audio of the invalid text in the target video;
- An interval identification module 320 is used to display the editing track segment of the target video on the editing interface of the target video, and identify the timeline interval of the invalid text on the editing track segment based on the timeline position of the invalid text, wherein the timeline interval of the invalid text is the time interval in which the voice audio of the invalid text appears in the target video;
- An interval adjustment module 330 configured to adjust the timeline interval of the invalid text on the editing track segment of the target video in response to the adjustment operation of the invalid text;
- the segment deletion module 340 is used to delete the video segment on the editing track segment within the timeline interval of the invalid text from the target video in response to the video segment deletion operation of the invalid text.
- the embodiment of the present disclosure provides a video editing device, which uses a text determination module to perform voice recognition on the audio in the target video to determine the invalid text in the voice text of the target video and the timeline position of the invalid text, wherein the timeline position of the invalid text is used to indicate the appearance time of the voice audio of the invalid text in the target video; and uses an interval identification module to display the invalid text on the editing interface of the target video.
- the editing track segment of the target video is identified, and based on the timeline position of the invalid text, the timeline interval of the invalid text is marked on the editing track segment, and the timeline interval of the invalid text is the time interval in which the voice and audio of the invalid text appear in the target video; the timeline interval of the invalid text is adjusted on the editing track segment of the target video in response to the adjustment operation of the invalid text by the interval adjustment module; the video segment located in the timeline interval of the invalid text on the editing track segment is deleted from the target video in response to the video segment deletion operation of the invalid text by the segment deletion module.
- the invalid text and the position of the invalid text in the target video can be determined, and then by marking the timeline interval of the invalid text on the editing track segment of the target video, the invalid segment can be deleted in response to the user's adjustment operation on the invalid text and the video segment deletion operation, thereby completing the video editing and improving the efficiency of video editing.
- the interval adjustment module 330 specifically includes at least one of the following:
- a first response unit configured to adjust the length of the timeline section of the invalid text on the editing track segment of the target video in response to a first adjustment operation of the invalid text
- the second response unit is used to adjust the number of timeline intervals marked on the editing track segment of the target video in response to the second adjustment operation of the invalid text.
- the interval identification module 320 is specifically used for:
- the editing track segment of the target video is displayed in the track area of the editing interface, and the invalid text and non-invalid text in the voice text are displayed in the text area of the editing interface, wherein the invalid text is displayed as a selected state on the editing interface, and the non-invalid text is displayed as a non-selected state on the editing interface.
- the first response unit includes:
- an adjusting subunit configured to adjust the length of a first timeline interval of a first text on the editing track segment of the target video, and update a target text displayed in the text area, wherein the target text includes the first text and/or a timeline interval related to the first text; The associated second text.
- the first adjustment operation includes a length extension operation
- the adjustment subunit is specifically configured to:
- the text content in the second text corresponding to the extended portion of the first timeline interval is moved to the first text.
- the first adjustment operation includes a length shortening operation
- the adjustment subunit is specifically configured to:
- a second text is added to the text area, and text content in the first text corresponding to the shortened portion of the first timeline interval is moved to the second text.
- the invalid text is text in the track area that is in a selected state
- the second response unit is specifically used for at least one of the following:
- the fourth text is switched from a selected state to an unselected state, and the timeline interval marking the fourth text is deselected on the editing track segment.
- a video editing device provided by an embodiment of the present disclosure further includes:
- a playing module used for playing the target video in the video area of the editing interface
- the first progress adjustment module is used to adjust the playback progress of the target video to the first time node when the target video is played to the starting point of the second timeline interval, and continue to play the target video from the first time node, wherein the second timeline interval is the timeline interval of the invalid text, and the first time node is the time node corresponding to the end position of the second timeline interval in the target video.
- a video editing device provided by an embodiment of the present disclosure further includes:
- the first response module is used to play the target in the video area of the editing interface.
- the playback progress of the target video is adjusted to the second time node, and the target video continues to be played from the second time node, the fifth text is displayed in the text area of the editing interface, the fifth text is invalid text or non-invalid text, and the second time node is the time node corresponding to the starting position of the fifth text in the target video.
- a time mark is also displayed on the editing track segment.
- the video editing device provided by the embodiment of the present disclosure further includes:
- a second response module configured to determine, after playing the target video in the video area of the editing interface, in response to a position adjustment operation on the target identifier, a third timeline interval where a target position of the target identifier is located, wherein the target position is a position of the target identifier on the editing track segment when the position adjustment operation is completed;
- a second progress adjustment module is used to adjust the playback progress of the target video to a third time node if the third timeline interval is the timeline interval of the invalid text, and continue to play the target video from the third time node, wherein the third time node is the time node corresponding to the starting position or the ending position of the third timeline interval in the target video, or the time node to which the target video was last played before receiving the position progress adjustment operation;
- the third progress adjustment module is used to adjust the playback progress of the target video to a fourth time node if the third timeline interval is not the timeline interval of the invalid text, and continue to play the target video from the fourth time node, wherein the fourth time node is the time node corresponding to the target position in the target video.
- segment deletion module 340 is specifically used for:
- the video clip on the editing track segment that is located in the timeline interval of the invalid text is deleted from the target video, and the target video after deleting the video clip is The videos are combined into one video.
- the editing track segment is an audio track segment of the target video.
- the above-mentioned video editing device can execute the video editing method provided by any embodiment of the present disclosure, and has functional modules and beneficial effects corresponding to the execution method.
- the terminal device in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
- mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc.
- PDAs personal digital assistants
- PADs tablet computers
- PMPs portable multimedia players
- vehicle-mounted terminals such as vehicle-mounted navigation terminals
- fixed terminals such as digital TVs, desktop computers, etc.
- the electronic device shown in FIG10 is only an example and should not bring any limitation to the functions and scope of use of
- the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 to a random access memory (RAM) 403.
- a processing device e.g., a central processing unit, a graphics processing unit, etc.
- RAM random access memory
- various programs and data required for the operation of the electronic device 400 are also stored.
- the processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404.
- An input/output (I/O) interface 405 is also connected to the bus 404.
- the following devices may be connected to the I/O interface 405: input devices 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 408 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 409.
- the communication device 409 may allow the electronic device 400 to communicate wirelessly or wired with other devices to exchange data.
- FIG. 10 shows an electronic device 400 with various devices, it should be understood that it is not required to implement or have all the devices shown. More or fewer devices may be implemented or have alternatively.
- an embodiment of the present disclosure includes a computer program.
- a computer program product includes a computer program carried on a non-transitory computer readable medium, the computer program including program code for executing the method shown in the flowchart.
- the computer program can be downloaded and installed from a network through a communication device 409, or installed from a storage device 408, or installed from a ROM 402.
- the processing device 401 executes the above functions defined in the method of the embodiment of the present disclosure.
- the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two.
- the computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above.
- Computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
- a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device.
- a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried.
- This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above.
- the computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device.
- the program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
- the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may communicate with any form or medium.
- Digital data communications e.g., communication networks
- Examples of communication networks include a local area network ("LAN”), a wide area network (“WAN”), an internetwork (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
- LAN local area network
- WAN wide area network
- Internet internetwork
- peer-to-peer network e.g., an ad hoc peer-to-peer network
- the computer-readable medium may be included in the electronic device, or may exist independently without being installed in the electronic device.
- the computer-readable medium carries one or more programs.
- the electronic device When the one or more programs are executed by the electronic device, the electronic device:
- the video segment on the editing track segment that is within the timeline interval of the invalid text is deleted from the target video.
- Computer program code for performing operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages.
- the program code may execute entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server.
- the remote computer may be connected via any type of network, including a local area network.
- a computer system is a network (LAN) or wide area network (WAN)—connected to a user's computer, or it can be connected to an external computer (for example, through the Internet using an Internet service provider).
- each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function.
- the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved.
- each square box in the block diagram and/or flow chart, and the combination of the square boxes in the block diagram and/or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
- the units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a module does not, in some cases, limit the unit itself.
- exemplary types of hardware logic components include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
- FPGAs field programmable gate arrays
- ASICs application specific integrated circuits
- ASSPs application specific standard products
- SOCs systems on chip
- CPLDs complex programmable logic devices
- a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device.
- a machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- a machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.
- machine-readable storage media would include electrical connections based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable Read only memory (EPROM or flash memory), optical fiber, portable compact disk read only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the above.
- RAM random access memory
- ROM read-only memory
- EPROM or flash memory erasable programmable Read only memory
- CD-ROM portable compact disk read only memory
- magnetic storage device or any suitable combination of the above.
- Example 1 provides a video editing method, including:
- the video segment on the editing track segment that is within the timeline interval of the invalid text is deleted from the target video.
- Example 2 is the method according to Example 1, wherein in response to the adjustment operation of the invalid text, the timeline interval of the invalid text is adjusted on the editing track segment of the target video, including at least one of the following:
- the number of timeline intervals identified on the editing track segment of the target video is adjusted.
- Example 3 is the method according to Example 2, wherein the editing track of the target video is displayed on the editing interface of the target video.
- the snippet includes:
- the editing track segment of the target video is displayed in the track area of the editing interface, and the invalid text and non-invalid text in the voice text are displayed in the text area of the editing interface, wherein the invalid text is displayed as a selected state on the editing interface, and the non-invalid text is displayed as a non-selected state on the editing interface.
- Example 4 is the method according to Example 3, wherein adjusting the length of the timeline interval of the invalid text on the editing track segment of the target video includes:
- the length of a first timeline interval of a first text is adjusted on an editing track segment of the target video, and the target text displayed in the text area is updated, wherein the target text includes the first text and/or a second text associated with the first text.
- Example 5 is the method according to Example 4, wherein the first adjustment operation includes a length extension operation, and the updating of the target text displayed in the text area includes:
- the text content in the second text corresponding to the extended portion of the first timeline interval is moved to the first text.
- Example 6 is the method according to Example 4, wherein the first adjustment operation includes a length shortening operation, and the updating of the target text displayed in the text area includes:
- a second text is added to the text area, and text content in the first text corresponding to the shortened portion of the first timeline interval is moved to the second text.
- Example 7 is the method according to Example 3, wherein the invalid text is text in the track area that is in a selected state, and the second adjustment operation in response to the invalid text adjusts the number of timeline intervals identified on the editing track segment of the target video, including at least one of the following:
- the fourth text is switched from a selected state to an unselected state, and the timeline interval marking the fourth text is deselected on the editing track segment.
- Example 8 is a method according to any one of Examples 1-7, further comprising:
- the playback progress of the target video is adjusted to the first time node, and the target video continues to be played from the first time node, wherein the second timeline interval is the timeline interval of the invalid text, and the first time node is the time node corresponding to the end point of the second timeline interval in the target video.
- Example 9 is the method according to Example 8, after playing the target video in the video area of the editing interface, further comprising:
- the playback progress of the target video is adjusted to a second time node, and the target video continues to be played from the second time node, the fifth text is displayed in the text area of the editing interface, the fifth text is an invalid text or a non-invalid text, and the second time node is the time node corresponding to the starting position of the fifth text in the target video.
- Example 10 is the method according to Example 8, wherein a time mark is also displayed on the editing track segment, and after the target video is played in the video area of the editing interface, the method further includes:
- the playback progress of the target video is adjusted to the third time node, and the target video is played from the third time node.
- the target video continues to be played at the node, wherein the third time node is the time node corresponding to the starting position or the ending position of the third timeline interval in the target video, or the time node to which the target video was last played before the position progress adjustment operation is received;
- the playback progress of the target video is adjusted to a fourth time node, and the target video continues to be played from the fourth time node, wherein the fourth time node is the time node corresponding to the target position in the target video.
- Example 11 is a method according to any one of Examples 1-7, wherein the step of deleting the video segment on the editing track segment within the timeline interval of the invalid text from the target video includes:
- the video segment on the editing track segment that is located in the timeline interval of the invalid text is deleted from the target video, and the target video after the video segment is deleted is integrated into one video.
- Example 12 is a method according to any one of Examples 1-7, wherein the edited track segment is an audio track segment of the target video.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Signal Processing (AREA)
- Artificial Intelligence (AREA)
- Television Signal Processing For Recording (AREA)
- Management Or Editing Of Information On Record Carriers (AREA)
Abstract
Description
Claims (15)
- 一种视频的编辑方法,包括:通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
- 根据权利要求1所述的方法,其中所述响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整,包括下述至少之一:响应于所述无效文本的第一调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间的长度进行调整;响应于所述无效文本的第二调整操作,对所述目标视频的编辑轨道片段上所标识的时间线区间的数量进行调整。
- 根据权利要求2所述的方法,其中所述在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,包括:在所述编辑界面的轨道区域展示所述目标视频的编辑轨道片段,并在所述编辑界面的文本区域展示所述语音文本中的无效文本和非无效文本,其中,所述无效文本在所述编辑界面上展示为选中状态,所述非无效文本在所述编辑界面上展示为非选中状态。
- 根据权利要求3所述的方法,其中所述在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间的长度进行调整,包括:在所述目标视频的编辑轨道片段上对第一文本的第一时间线区间的长度进行调整,并更新所述文本区域中所展示的目标文本,所述目标文本包括所述第一文本和/或与所述第一文本相关联的第二文本。
- 根据权利要求4所述的方法,其中所述第一调整操作包括长度延长操作,所述更新所述文本区域中所展示的目标文本,包括:将所述第二文本中与所述第一时间线区间的延长部分对应的文本内容移动至所述第一文本中。
- 根据权利要求4所述的方法,其中所述第一调整操作包括长度缩短操作,所述更新所述文本区域中所展示的目标文本,包括:将所述第一文本中与所述第一时间线区间的缩短部分对应的文本内容移动至所述第二文本中;或者在所述文本区域中新增第二文本,并将所述第一文本中与所述第一时间线区间的缩短部分对应的文本内容移动至所述第二文本中。
- 根据权利要求3所述的方法,其中所述无效文本为所述轨道区域中处于选中状态的文本,所述响应于所述无效文本的第二调整操作,对所述目标视频的编辑轨道片段上所标识的时间线区间的数量进行调整,包括下述至少之一:响应于对所述文本区域中的第三文本的选中操作,将所述第三文本由未选中状态切换为选中状态,并在所述编辑轨道片段上增加标识所述第三文本的时间线区间;响应于对所述文本区域中的第四文本的取消选中操作,将所述第四文本由选中状态切换为未选中状态,并在所述编辑轨道片段上取消标识所述第四文本的时间线区间。
- 根据权利要求1-7任一所述的方法,还包括:在所述编辑界面的视频区域播放所述目标视频;当所述目标视频播放至第二时间线区间的起点位置时,将所述目标视频的播放进度调整至第一时间节点,并从所述第一时间节点继续播放所述目标视频,其中,所述第二时间线区间为所述无效文本的时间线区间,所述第一时间节点为所述第二时间线区间的终点位置在所述目标视频中所对应的时间节点。
- 根据权利要求8所述的方法,其中在所述编辑界面的视频区域播放所述目标视频之后,还包括:响应于对第五文本的触发操作,将所述目标视频的播放进度调整至第二时间节点,并从所述第二时间节点继续播放所述目标视频,所述第五文本显示于所述编辑界面的文本区域内,所述第五文本为无效文本或非无效文本,所述第二时间节点为所述第五文本的起点位置在所述目标视频中所对应的时间节点。
- 根据权利要求8所述的方法,其中所述编辑轨道片段上还展示有时间标识,在所述编辑界面的视频区域播放所述目标视频之后,还包括:响应于对所述目标标识的位置调整操作,确定所述目标标识的目标位置所位于的第三时间线区间,所述目标位置为所述位置调整操作执行完毕时所述目标标识在所述编辑轨道片段上的位置;如果所述第三时间线区间为所述无效文本的时间线区间,则将所述目标视频的播放进度调整至第三时间节点,并从所述第三时间节点继续播放所述目标视频,其中,所述第三时间节点为所述第三时间线区间的起点位置或终点位置在所述目标视频中所对应的时间节点,或者,接收到所述位置进度调整操作之前所述目标视频最后播放至的时间节点;如果所述第三时间线区间不为所述无效文本的时间线区间,则将所述目标视频的播放进度调整至第四时间节点,并从所述第四时间节点继续播放所述目标视频,其中,所述第四时间节点为所述目标位置在所述目标视频中所对应的时间节点。
- 根据权利要求1-7任一所述的方法,其中所述将所述编辑轨 道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,包括:将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,并以删除的所述视频片段为分割点,将所述目标视频分割为多个独立的视频片段;或者将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除,并将删除所述视频片段后的目标视频整合为一个视频。
- 根据权利要求1-7任一所述的方法,其中所述编辑轨道片段为所述目标视频的音频轨道片段。
- 一种视频的编辑装置,包括:文本确定模块,用于通过对目标视频中的音频进行语音识别,确定所述目标视频的语音文本中的无效文本及所述无效文本的时间线位置,所述无效文本的时间线位置用于表示所述无效文本的语音音频在所述目标视频中的出现时间;区间标识模块,用于在所述目标视频的编辑界面上展示所述目标视频的编辑轨道片段,并基于所述无效文本的时间线位置,在所述编辑轨道片段上标识所述无效文本的时间线区间,所述无效文本的时间线区间为所述无效文本的语音音频在所述目标视频中出现的时间区间;区间调整模块,用于响应于所述无效文本的调整操作,在所述目标视频的编辑轨道片段上对所述无效文本的时间线区间进行调整;片段删除模块,用于响应于所述无效文本的视频片段删除操作,将所述编辑轨道片段上位于所述无效文本的时间线区间内的视频片段从所述目标视频中删除。
- 一种电子设备,包括:至少一个处理器;以及与所述至少一个处理器通信连接的存储器;其中,所述存储器存储有可被所述至少一个处理器执行的计算机程序,所述计算机程序被所述至少一个处理器执行,以使所述至少一个处理器能够执行权利要求1-12中任一项所述的视频的编辑方法。
- 一种计算机可读存储介质,所述计算机可读存储介质存储有计算机指令,所述计算机指令用于使处理器执行时实现权利要求1-12中任一项所述的视频的编辑方法。
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2023578854A JP7680578B2 (ja) | 2023-01-19 | 2023-12-12 | ビデオの編集方法、装置、電子機器及び記憶媒体 |
| EP23821101.5A EP4429256A4 (en) | 2023-01-19 | 2023-12-12 | VIDEO EDITING METHOD AND APPARATUS, ELECTRONIC DEVICE AND RECORDING MEDIUM |
| KR1020257018467A KR20250105651A (ko) | 2023-01-19 | 2023-12-12 | 비디오 편집 방법 및 장치, 그리고 전자 디바이스 및 저장 매체 |
| US18/542,524 US12051445B1 (en) | 2023-01-19 | 2023-12-15 | Method and apparatus of video editing, and electronic device and storage medium |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202310107566.3A CN118368478A (zh) | 2023-01-19 | 2023-01-19 | 视频的编辑方法、装置、电子设备和存储介质 |
| CN202310107566.3 | 2023-01-19 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US18/542,524 Continuation US12051445B1 (en) | 2023-01-19 | 2023-12-15 | Method and apparatus of video editing, and electronic device and storage medium |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024152802A1 true WO2024152802A1 (zh) | 2024-07-25 |
Family
ID=89430577
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/138239 Ceased WO2024152802A1 (zh) | 2023-01-19 | 2023-12-12 | 视频的编辑方法、装置、电子设备和存储介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN118368478A (zh) |
| WO (1) | WO2024152802A1 (zh) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20120323897A1 (en) * | 2011-06-14 | 2012-12-20 | Microsoft Corporation | Query-dependent audio/video clip search result previews |
| CN109040773A (zh) * | 2018-07-10 | 2018-12-18 | 武汉斗鱼网络科技有限公司 | 一种视频改进方法、装置、设备及介质 |
| CN113613068A (zh) * | 2021-08-03 | 2021-11-05 | 北京字跳网络技术有限公司 | 视频的处理方法、装置、电子设备和存储介质 |
| CN114999530A (zh) * | 2022-05-18 | 2022-09-02 | 北京飞象星球科技有限公司 | 音视频剪辑方法及装置 |
| CN115623279A (zh) * | 2021-07-15 | 2023-01-17 | 脸萌有限公司 | 多媒体处理方法、装置、电子设备及存储介质 |
-
2023
- 2023-01-19 CN CN202310107566.3A patent/CN118368478A/zh active Pending
- 2023-12-12 WO PCT/CN2023/138239 patent/WO2024152802A1/zh not_active Ceased
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20120323897A1 (en) * | 2011-06-14 | 2012-12-20 | Microsoft Corporation | Query-dependent audio/video clip search result previews |
| CN109040773A (zh) * | 2018-07-10 | 2018-12-18 | 武汉斗鱼网络科技有限公司 | 一种视频改进方法、装置、设备及介质 |
| CN115623279A (zh) * | 2021-07-15 | 2023-01-17 | 脸萌有限公司 | 多媒体处理方法、装置、电子设备及存储介质 |
| CN113613068A (zh) * | 2021-08-03 | 2021-11-05 | 北京字跳网络技术有限公司 | 视频的处理方法、装置、电子设备和存储介质 |
| CN114999530A (zh) * | 2022-05-18 | 2022-09-02 | 北京飞象星球科技有限公司 | 音视频剪辑方法及装置 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN118368478A (zh) | 2024-07-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7572108B2 (ja) | 議事録のインタラクション方法、装置、機器及び媒体 | |
| US12573429B2 (en) | Video processing method and apparatus, electronic device and storage medium | |
| CN111970577A (zh) | 字幕编辑方法、装置和电子设备 | |
| CN114679628B (zh) | 一种弹幕添加方法、装置、电子设备和存储介质 | |
| JP7334362B2 (ja) | 文書内テーブル閲覧方法、装置、電子機器及び記憶媒体 | |
| US20240289398A1 (en) | Method, apparatus, device and storage medium for content display | |
| CN113395568B (zh) | 视频互动方法、装置、电子设备和存储介质 | |
| CN115061602B (zh) | 页面显示方法、装置、设备、计算机可读存储介质及产品 | |
| US20250106486A1 (en) | Multimedia data processing method and apparatus, device and medium | |
| JP2024525192A (ja) | オーディオ処理方法、装置、電子機器及び記憶媒体 | |
| US20250173045A1 (en) | Book information display method and apparatus, device, and storage medium | |
| WO2024140239A1 (zh) | 页面显示方法、装置、设备、计算机可读存储介质及产品 | |
| US20250363293A1 (en) | Method for generating table field content, storage medium, and electronic device | |
| WO2025185040A1 (zh) | 媒体内容的生成方法、装置、电子设备、存储介质和程序产品 | |
| WO2024099376A1 (zh) | 视频编辑方法、装置、设备及介质 | |
| CN117857829A (zh) | 内容生成方法、装置、可读介质及电子设备 | |
| US20240103802A1 (en) | Method, apparatus, device and medium for multimedia processing | |
| CN118170297A (zh) | 特效编辑方法、装置、电子设备、存储介质及程序产品 | |
| US12051445B1 (en) | Method and apparatus of video editing, and electronic device and storage medium | |
| WO2026067488A1 (zh) | 互动信息的发布方法、装置、电子设备、存储介质和程序产品 | |
| WO2025130695A1 (zh) | 特效创作方法、装置、设备、计算机可读存储介质及产品 | |
| WO2026056410A1 (zh) | 信息显示方法、装置、电子设备、存储介质和程序产品 | |
| WO2025252119A1 (zh) | 作品展示方法、装置、电子设备、存储介质和程序产品 | |
| CN118368478A (zh) | 视频的编辑方法、装置、电子设备和存储介质 | |
| WO2025093007A1 (zh) | 信息回复方法、装置、电子设备和存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| WWE | Wipo information: entry into national phase |
Ref document number: 2023578854 Country of ref document: JP |
|
| ENP | Entry into the national phase |
Ref document number: 2023821101 Country of ref document: EP Effective date: 20231218 |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23821101 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202527050779 Country of ref document: IN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 1020257018467 Country of ref document: KR |
|
| WWP | Wipo information: published in national office |
Ref document number: 202527050779 Country of ref document: IN |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 11202503525X Country of ref document: SG |
|
| WWP | Wipo information: published in national office |
Ref document number: 11202503525X Country of ref document: SG |
|
| WWP | Wipo information: published in national office |
Ref document number: 1020257018467 Country of ref document: KR |
|
| REG | Reference to national code |
Ref country code: BR Ref legal event code: B01A Ref document number: 112025011769 Country of ref document: BR |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 112025011769 Country of ref document: BR Kind code of ref document: A2 Effective date: 20250610 |