WO2021259221A1 - 视频翻译方法和装置、存储介质和电子设备 - Google Patents
视频翻译方法和装置、存储介质和电子设备 Download PDFInfo
- Publication number
- WO2021259221A1 WO2021259221A1 PCT/CN2021/101388 CN2021101388W WO2021259221A1 WO 2021259221 A1 WO2021259221 A1 WO 2021259221A1 CN 2021101388 W CN2021101388 W CN 2021101388W WO 2021259221 A1 WO2021259221 A1 WO 2021259221A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- translation
- text
- user
- suggestion
- time information
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/40—Processing or translation of natural language
- G06F40/58—Use of machine translation, e.g. for multi-lingual retrieval, for server-side translation for client devices or for real-time translation
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/16—Sound input; Sound output
- G06F3/165—Management of the audio stream, e.g. setting of volume, audio stream path
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/166—Editing, e.g. inserting or deleting
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/40—Processing or translation of natural language
- G06F40/51—Translation evaluation
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/26—Speech to text systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
- G10L25/57—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for processing of video signals
-
- G—PHYSICS
- G11—INFORMATION STORAGE
- G11B—INFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
- G11B27/00—Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
- G11B27/02—Editing, e.g. varying the order of information signals recorded on, or reproduced from, record carriers
-
- G—PHYSICS
- G11—INFORMATION STORAGE
- G11B—INFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
- G11B27/00—Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
- G11B27/02—Editing, e.g. varying the order of information signals recorded on, or reproduced from, record carriers
- G11B27/031—Electronic editing of digitised analogue information signals, e.g. audio or video signals
-
- G—PHYSICS
- G11—INFORMATION STORAGE
- G11B—INFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
- G11B27/00—Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
- G11B27/10—Indexing; Addressing; Timing or synchronising; Measuring tape travel
- G11B27/19—Indexing; Addressing; Timing or synchronising; Measuring tape travel by using information detectable on the record carrier
- G11B27/28—Indexing; Addressing; Timing or synchronising; Measuring tape travel by using information detectable on the record carrier by using information signals recorded by the same method as the main recording
-
- G—PHYSICS
- G11—INFORMATION STORAGE
- G11B—INFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
- G11B27/00—Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
- G11B27/10—Indexing; Addressing; Timing or synchronising; Measuring tape travel
- G11B27/34—Indicating arrangements
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/233—Processing of audio elementary streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/2343—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements
- H04N21/234336—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving reformatting operations of video signals for distribution or compliance with end-user requests or end-user device requirements by media transcoding, e.g. video is transformed into a slideshow of still pictures or audio is converted into text
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/431—Generation of visual interfaces for content selection or interaction; Content or additional data rendering
- H04N21/4312—Generation of visual interfaces for content selection or interaction; Content or additional data rendering involving specific graphical features, e.g. screen layout, special fonts or colors, blinking icons, highlights or animations
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/4402—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display
- H04N21/440236—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving reformatting operations of video signals for household redistribution, storage or real-time display by media transcoding, e.g. video is transformed into a slideshow of still pictures, audio is converted into text
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/47—End-user applications
- H04N21/488—Data services, e.g. news ticker
- H04N21/4884—Data services, e.g. news ticker for displaying subtitles
Definitions
- the present disclosure relates to the field of machine translation, and in particular to a video translation method and device, storage medium and electronic equipment.
- the content of the invention is provided to introduce concepts in a brief form, and these concepts will be described in detail in the following specific embodiments.
- the content of the invention is not intended to identify the key features or essential features of the technical solution that is required to be protected, nor is it intended to be used to limit the scope of the technical solution that is required to be protected.
- the present disclosure provides a video translation method, including:
- the first time information is the start time of the text in the video
- the second time information is the The ending time of the text in the video
- an editing area In response to a user's operation on the text or the reference translation, an editing area is displayed, the editing area supports the user to input the translation;
- the present disclosure provides a video translation device, including:
- Conversion module used to convert the voice of the video to be translated into text
- the display module is used to display the text and the first time information, the second time information and the reference translation of the text, the first time information is the start time of the text in the video, and the first time information is the start time of the text in the video. Second, the time information is the end time of the text in the video;
- the display module is further configured to display an editing area in response to a user's operation on the text or the reference translation, and the editing area supports the user to input the translation;
- the suggestion module is used to follow the user's input in the editing area and provide translation suggestions from the reference translation;
- the display module is further configured to display the translation suggestion as a translation result in the editing area in the case of detecting the user's confirmation operation for the translation suggestion; and, after detecting that the user is targeting In the case of the non-confirmation operation of the translation suggestion, a translation input by the user that is different from the translation suggestion is received, and the translation input by the user is displayed as the translation result in the editing area, according to the The reference translation in the translation area is updated with the translation input by the user.
- the present disclosure provides a computer-readable medium on which a computer program is stored, and when the program is executed by a processing device, the steps of the method described in the first aspect of the present disclosure are implemented.
- an electronic device including:
- a storage device on which a computer program is stored
- the processing device is configured to execute the computer program in the storage device to implement the steps of the method in the first aspect of the present disclosure.
- the voice of the video to be translated can be converted into text
- the first time information, second time information and reference translation of the text can be provided, and follow the user's input in the editing area
- Fig. 1 is a flowchart showing a video translation method according to an exemplary disclosed embodiment.
- Fig. 2 is a schematic diagram showing a translation interface according to an exemplary disclosed embodiment.
- Fig. 3 is a schematic diagram showing a text splitting manner according to an exemplary disclosed embodiment.
- Fig. 4 is a block diagram showing a video translation device according to an exemplary disclosed embodiment.
- Fig. 5 is a block diagram showing an electronic device according to an exemplary disclosed embodiment.
- Figure 1 is a flow chart showing a video translation method according to an exemplary disclosed embodiment.
- This method can be used in terminals, servers and other independent electronic devices, and can also be applied to a translation system.
- the method in the method Each step can be completed by multiple devices in the translation system.
- S12 and S14 shown in FIG. 1 can be executed by the terminal, and S11 and S13 can be executed by the server.
- the video translation method includes the following steps:
- the voice content of the video can be translated, such as audio tracks, and convert the voice content into text content through voice recognition technology.
- the text content can be divided into multiple sentences according to the clauses in the voice content, and the text content of each sentence can correspond to the time information extracted from the speech content of the clause , Use it as the timeline information of the text content of the sentence.
- the voice content of the video to be translated is recognized as multiple sentences.
- the first sentence is "First introduce what is a hot spot”
- the sentence is located between the second to the fifth second of the video
- the sentence text The timeline information corresponding to the content is "00:00:02-00:00:05”
- the second round is "You can see from the right side of ppt”
- this sentence is located between the 5th and 7th seconds of the video
- the timeline information corresponding to the text content of the sentence is "00:00:05-00:00:07”.
- the text may be segmented according to the time information and/or picture frames corresponding to the text in the video to obtain the multiple segmented texts, For example, consider the recognized text of the voice within multiple consecutive seconds as a clause, or use the recognized text of the voice appearing in multiple consecutive frames as a clause; it is also possible to divide the sentence according to the pause in the speech content, for example , You can set a pause threshold. When the human voice content is not recognized within the pause threshold, the sentence can be segmented at any position where the human voice content is not recognized; it can also be segmented according to the semantics of the voice content.
- the recognized text content can be segmented through the sentence model to obtain the text content after the sentence.
- the first time information is the start time of the text in the video
- the second time information is the end time of the text in the video
- the text may be a text that has been divided into sentences
- the first time information is the start time of the current sentence of the sentenced text in the video
- the second time information is the current division of the sentenced text End time in the video.
- the editing area can be displayed above the reference translation of the text.
- the editing area supports the user to input the translation, and the user can perform editing operations in the editing area to obtain the translation result of the text. Among them, the editing area can be displayed above the reference translation for the user to compare and modify.
- the text may be a text that has been divided into sentences, and each sentence text is displayed in a different area, and for each sentence text, the first time information, the second time information, and the reference translation of the sentence text are displayed.
- the text includes a plurality of clause texts, each of the clause texts is displayed in a different area, and a split that provides the user to split the clause text can also be displayed.
- the function bar and in response to the user's splitting operation for any one of the sentenced texts, splits the sentenced text into at least two sentenced sentenced texts, and associated display for each of the sentenced texts.
- the split function bar may be provided in response to a user's operation on the clause text or reference translation, and the split function bar may be hidden before the user selects the clause text or reference text.
- the timeline information of this piece of text is "00:00:15-00:00:18”
- the first The first time information is 00:00:15
- the second time information is 00:00:18.
- the user divides it into two clauses: "What I want to introduce to you today" and "The three cities that are about to rise in our country”, then You can set the time axis for each clause according to the length of the text before editing and the length of the text of each clause after editing.
- Figure 3 shows a schematic diagram of a possible text splitting method.
- the user can use the cursor to select the position where the sentence needs to be split, and click the sentence button; the text before the sentence will be split into two subsections.
- the first time information and the second time information of each clause are all obtained by splitting the first time information and the second time information before the clause.
- a section of text in the dashed box before splitting is split into two subsections in the dashed box.
- the text includes a plurality of clause texts, each of the clause texts is displayed in a different area, and a merge function column that provides the user to merge the clauses may also be displayed, And in response to the user's merging operation for any two adjacent clause texts, the two adjacent clause texts are merged into a new clause text, and the new clause text is displayed in association The new clause text, the first time information, the second time information of the new clause text, and the reference translation of the new clause text.
- the merging function bar may be provided in response to the user's operation on the clause text or the reference translation, and the merging function bar may be hidden before the user selects the clause text or the reference text.
- the text includes multiple sentence texts, each of the sentence texts is displayed in a different area, and a play function bar that provides the user to play the sentence texts may also be displayed. , And in response to the user's operation on the play function bar, play the voice corresponding to the clause text.
- the play function bar may be provided in response to the user's operation on the sentence text or the reference translation, and the play function bar may be hidden before the user selects the sentence text or the reference text.
- the reference translation or the translation result may be used as a subtitle to play the video corresponding to the sentence text, so that the user can view the effect of the translated subtitle.
- FIG 2 is a schematic diagram of a possible translation interface.
- the inside of the dashed box is the translation interface of a section of text content that has been selected by the user.
- the selected text content will display the editing area and the play, merge, and split function bars.
- the text content of the video to be translated is displayed above the reference translation, and different clauses have different display areas. Each display area can be translated independently and will not be updated due to the modification of other areas.
- the user can enter characters in the editing area or modify the characters of the text to be translated.
- the translation interface may also include time axis information, including first time information representing the start time and second time information representing the end time.
- the reference translation is gray, and the translation is suggested to be black.
- the reference translation can be moved down one line to align with the function bar, while the original reference translation is located
- the area becomes the editing area, which is used to display translation suggestions and receive user modifications.
- the method provided by the embodiment of the present disclosure includes displaying the translation suggestion as the translation result in the editing area in the case of detecting the user's confirmation operation for the translation suggestion, and when it is detected that the user is targeting the translation suggestion In the case of a non-confirmation operation of the translation suggestion, a translation that is different from the translation suggestion input by the user is received, and the reference translation in the translation area is updated according to the translation input by the user.
- the above confirmation operation may be the user's operation on a preset shortcut key.
- the translation suggestion is displayed in the editing area as the translation result.
- the action of displaying the translation suggestion as the translation result in the editing area will be the user's input in the editing area as described in step S14, that is, in this case, step S14 is It shows that the method provided by the embodiment of the present disclosure can provide the next translation suggestion from the reference translation in response to displaying the translation suggestion provided this time as the translation result in the editing area (the next translation suggestion may be the provided translation It is recommended to refer to the subsequent translation in the reference translation).
- the detection of the non-confirmation operation of the user for the translation suggestion may be the detection of the inconsistency between the translation input by the user and the translation suggestion provided this time.
- the method provided by the embodiment of the present disclosure It can receive a translation that is different from the suggested translation input by the user, and update the reference translation in the translation area according to the translation input by the user.
- step S14 indicates that this The method provided by the disclosed embodiment can provide the next translation suggestion from the reference translation updated according to the translation input by the user in response to the user inputting a translation in the editing area that is different from the translation suggestion.
- the translation suggestion provided this time is "my”. If it is detected that the translation input by the user is a translation "I” that is different from the translation suggestion "my”, the reference translation is updated according to the translation "I”, and the updated Suggestions for the next translation of the translation "I" are provided in the reference translation of.
- the translation suggestions from the reference translation can be provided according to the user's input, and the user can directly use the translation suggestion as the translation result through the confirmation operation, reducing the user's input time.
- the present disclosure combines manual accuracy and machine efficiency , Can improve the efficiency of translation and the quality of translation.
- the providing translation suggestions in step S14 may include: highlighting the translation suggestions from the reference translation in the translation area.
- the highlighting of the translation suggestion in the translation area can be cancelled.
- the highlighting can be bold fonts, highlighted fonts, heterochromatic characters, heterochromatic backgrounds, shading effects, etc., which can highlight the translation suggestions.
- the highlighting may be a display mode different from the display mode of the input translation, for example, the input translation may be a bold font, the translation suggestion is a normal font, or the input translation It can be black words, the translation suggestion is gray words, etc.
- the display mode of the translation suggestion can be adjusted to be the same as the display method of the input translation.
- the input translation may be in bold font, and the translation suggestion is in normal font.
- the translation suggestion is adjusted to be displayed in bold font.
- the confirmation operation may be a user's input operation on a shortcut key of the electronic device.
- the electronic device may be a mobile phone, and the shortcut key may be a virtual key on the display area of the mobile phone or an entity of the mobile phone.
- Key for example: volume key
- the user can operate the above shortcut key to adopt the translation suggestion, then in the case of detecting the user's input operation on the above shortcut key, the translation suggestion can be displayed in the editing area as a question result ;
- the electronic device can also be a computer, and the shortcut key can be a designated or custom key on the computer keyboard or mouse (for example: keyboard alt key, mouse side key, etc.).
- the confirmation operation may also be a gesture confirmation operation recognized after being captured by the camera, such as nodding, blinking, making a preset gesture, etc.; it may also be a voice operation recognized after being captured by a microphone.
- the translation suggestion from the reference translation includes at least one of a word, a phrase, and a sentence.
- the user When the user translates the text content, he can refer to the reference translation displayed in the translation area and input in the editing area (it is worth noting that the input here includes the input of characters, such as typing letters, words, etc., as well as Key operation input, such as clicking the editing area, etc.), can provide translation suggestions from the reference translation.
- the translation suggestion can be a translation suggestion for the whole sentence of a clause, or a more fine-grained translation suggestion provided word by word or phrase by phrase.
- the user can adopt the translation suggestion through the confirmation operation, and use the confirmation operation as an input operation in the editing area to continue to provide translation suggestions from the reference translation, for example, in the case of detecting the user's confirmation operation for "Some” , Display "Some” as the translation result in the editing area, and provide users with the next translation suggestion "cities".
- the non-confirmation operation can be a preset operation that represents non-confirmation (clicking a preset button, making a preset action, etc.), or it can refer to other conditions besides the aforementioned confirmation operation, for example, The confirmation operation is not performed within the set time, or the operation to continue input is performed.
- the reference translation of the text "Some cities continue to rise with the advantage of a complete high-speed rail network” is "Some cities continue to rise with the advantage of the perfect high-speed rail network.”
- the click input operation provide the translation suggestion "Some” from the reference translation, and based on the user's confirmation operation, display the translation suggestion "Some” as the translation result in the editing area, and continue to provide the user with the next translation suggestion "cities” .
- the translation suggestion "with” if you receive the user's input "b” that is different from the translation suggestion, you can update the reference translation to "Some cities continue to rise because of the advantage of the perfect high-speed” based on the translation input by the user rail network.” and provide users with translation suggestions "because”.
- the user can directly edit the translation suggestion in the editing area, for example, insert a word in the translation suggestion, delete a word in the translation suggestion, and change the translation Suggested words etc.
- the translation suggestion for the text "Some cities continue to rise by virtue of the perfect high-speed rail network” is the same as the reference translation, which is “Some cities continue to rise with the advantage of the perfect high-speed rail network.”
- users can Directly modify "with” to "because of” in the translation suggestion, and update the reference translation to "ome city continue to rise because of the advantage of the perfect high-speed rail network.” according to the user's modification, and inform the user Provide a translation suggestion from the reference translation, and the user can confirm the translation suggestion as the translation result.
- the reference translation and translation suggestions can be provided by machine translation (for example, a deep learning translation model, etc.). It is worth noting that when a reference translation that conforms to the text content cannot be generated based on the translation entered by the user in the editing area, the user can correct the translation characters entered by the user based on the pre-stored dictionary content, and update the translation based on the corrected translation.
- the reference translation can be provided by machine translation (for example, a deep learning translation model, etc.).
- the translation language is English and the original text is Chinese as examples in this disclosure, this disclosure does not limit the language of the translation and the language of the original text.
- the original text in this disclosure may also be Chinese classical Chinese or translated text. It can be in the Chinese vernacular, or in various combinations such as the original text in Japanese and the translated text in English.
- the original text display area is an editable area, and in response to a user's operation to modify the text content in the original text display area, the reference translation in the translation area can be updated.
- the user Before or after the user enters the translation in the translation area, the user can edit the content of the text, that is, the original translation, and the entered translation will not be overwritten due to the modification of the original text, but will be based on the modified text by the user
- the content and the input translation characters update the translation result.
- the text content before editing is "Some cities continue to rise by virtue of the perfect Qualcomm network”
- the corresponding translation suggestion is "Some cities continue to rise with the advantage of a perfect Qualcomm network.”
- the user is editing
- the translation result of the regional input is "Some cities continue to rise b”. If the translation that is different from the translation suggestion is "b”, you can update the reference translation to "Some cities continue to rise because of the advantage of a perfect Qualcomm network.” ".
- the text content of the sentence may be misrecognized text due to noise, the accent of the voice narrator and other factors.
- the length of the text content after editing is greater than the length of the text content before editing. According to the time axis information of the text content before editing, the time axis information of the edited text content is obtained through interpolation processing. .
- the timeline information of each text will be reset to the original 9/11, and when subsequent users perform operations such as clauses, merging, etc., the clause or merged sub-segment is determined based on the timeline information of each text Timeline information.
- the translation result may be added as a subtitle to the frame of the video to be translated.
- the timeline of the translation result of the first sentence of the video to be translated is "00:00:00-00:00:02" (the first time information is 00:00:00, and the second time information is 00:00:02 ), the timeline of the second translation result is "00:00:03-00:00:07” (the first time information is 00:00:03, the second time information is 00:00:07), you can Insert the translation result with the timeline of "00:00:00-00:00:02" from the 0th second to the 2nd second of the video to be translated, and insert the translation result from the 3rd to the 7th second of the video to be translated
- the timeline is the translation result of "00:00:03-00:00:07", which can be inserted into the video to be translated in the form of subtitles.
- the translated video can be generated in a format specified by the user and provided to the user for download.
- the voice of the video to be translated can be converted into text
- the first time information, second time information and reference translation of the text can be provided, and follow the user's input in the editing area
- Fig. 4 is a block diagram showing a video translation device according to an exemplary disclosed embodiment. As shown in FIG. 4, the video translation device 400 includes:
- the conversion module 410 is used to convert the voice of the video to be translated into text.
- the display module 420 is used to display the text and the first time information, the second time information and the reference translation of the text, the first time information is the start time of the text in the video, the The second time information is the end time of the text in the video.
- the display module 420 is further configured to display an editing area in response to a user's operation on the text or the reference translation, and the editing area supports the user to input a translation.
- the suggestion module 430 is configured to follow the user's input in the editing area and provide translation suggestions from the reference translation;
- the display module 420 is further configured to display the translation suggestion as a translation result in the editing area when the user's confirmation operation for the translation suggestion is detected; and, when the user is detected In the case of a non-confirmation operation for the translation suggestion, receiving a translation input by the user that is different from the translation suggestion, displaying the translation input by the user as the translation result in the editing area, according to The translation input by the user updates the reference translation in the translation area.
- the display module 420 is further configured to segment the text according to the time information and/or picture frames corresponding to the text in the video to obtain the multiple segmented texts; The clause text displays the clause text and the first time information, the second time information and the reference translation of the clause text.
- the text includes a plurality of sectioned texts, each of the sectioned texts is displayed in a different area
- the device further includes a splitting module for displaying a splitting function bar, and the splitting function bar supports The user splits the sentence text; in response to the user's split operation for any one of the sentence texts, the sentence text is split into at least two sentence sentence texts, and each 1.
- the text includes a plurality of clause texts, each of the clause texts is displayed in a different area, the device further includes a merging module for displaying a merged function column, and the merged function column supports the user Merging the clause text; in response to the user's merging operation for any two adjacent clause texts, merging the two adjacent clause texts into a new clause text, and targeting the
- the new clause text is associated and displayed with the new clause text, the first time information, the second time information of the new clause text, and the reference translation of the new clause text.
- the text includes a plurality of sectioned texts, each of the sectioned texts is displayed in a different area, and the device further includes a play module for displaying a play function bar, and the play function bar supports the user Playing the voice corresponding to the sentence text; in response to the user's operation on the playback function bar, playing the voice corresponding to the sentence text.
- the suggestion module 430 is configured to display the translation suggestion in a display mode different from the input translation in the editing area; and in response to the user's confirmation operation on the translation suggestion, the The display of the translation suggestion as the translation result to the editing area includes: in response to the user confirming the translation suggestion, displaying it in the editing area in the same manner as the input translation as the translation result The said translation suggestion.
- the suggestion module 430 is further configured to display the translation suggestion as a translation result in the editing area in response to a user's triggering operation of the shortcut key.
- the voice of the video to be translated can be converted into text
- the first time information, second time information and reference translation of the text can be provided, and follow the user's input in the editing area
- FIG. 5 shows a schematic structural diagram of an electronic device (for example, the terminal device or the server in FIG. 1) 500 suitable for implementing the embodiments of the present disclosure.
- the terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile phones, notebook computers, digital broadcast receivers, personal digital assistants (personal digital assistants, PDAs), tablet computers (portable android devices, PAD), and portable multimedia players.
- Mobile terminals such as portable media player (PMP), in-vehicle terminals (for example, in-vehicle navigation terminals), and fixed terminals such as digital television (DTV) and desktop computers.
- PMP portable media player
- in-vehicle terminals for example, in-vehicle navigation terminals
- DTV digital television
- the electronic device shown in FIG. 5 is only an example, and should not bring any limitation to the function and scope of use of the embodiments of the present disclosure.
- the electronic device 500 may include a processing device (such as a central processing unit, a graphics processor, etc.) 501, which may be based on a program stored in a read-only memory (read-only memory, ROM) 502 or from a storage device 508 is loaded into a program in a random access memory (RAM) 503 to execute various appropriate actions and processing.
- a processing device such as a central processing unit, a graphics processor, etc.
- ROM read-only memory
- RAM random access memory
- various programs and data required for the operation of the electronic device 500 are also stored.
- the processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504.
- An input/output (I/O) interface 505 is also connected to the bus 504.
- the following devices can be connected to the I/O interface 505: including input devices 506 such as touch screen, touch panel, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; including, for example, liquid crystal display (LCD) Output devices 507 such as speakers, vibrators, etc.; storage devices 508 such as magnetic tapes, hard disks, etc.; and communication devices 509.
- the communication device 509 may allow the electronic device 500 to perform wireless or wired communication with other devices to exchange data.
- FIG. 5 shows an electronic device 500 having various devices, it should be understood that it is not required to implement or have all of the illustrated devices. It may be implemented alternatively or provided with more or fewer devices.
- an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer readable medium, and the computer program contains program code for executing the method shown in the flowchart.
- the computer program may be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502.
- the processing device 501 the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
- the above-mentioned computer-readable medium in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two.
- the computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or a combination of any of the above.
- Computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable Erasable programmable read-only memory (EPROM) or flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any of the above The right combination.
- a computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
- a computer-readable signal medium may include a data signal propagated in a baseband or as a part of a carrier wave, and a computer-readable program code is carried therein.
- This propagated data signal can take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing.
- the computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium.
- the computer-readable signal medium may send, propagate or transmit the program for use by or in combination with the instruction execution system, apparatus, or device .
- the program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wire, optical cable, radio frequency (RF), etc., or any suitable combination of the foregoing.
- the client and server can communicate with any currently known or future-developed network protocol, such as hypertext transfer protocol (HTTP), and can communicate with digital data in any form or medium.
- Communication e.g., communication network
- Examples of communication networks include local area networks (LAN), wide area networks (WAN), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any current Know or develop a network in the future.
- LAN local area networks
- WAN wide area networks
- the Internet e.g., the Internet
- end-to-end networks e.g., ad hoc end-to-end networks
- the above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist alone without being assembled into the electronic device.
- the above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device: obtains at least two Internet protocol addresses; A node evaluation request for an Internet Protocol address, wherein the node evaluation device selects an Internet Protocol address from the at least two Internet Protocol addresses and returns it; receives the Internet Protocol address returned by the node evaluation device; wherein, the obtained The Internet Protocol address indicates the edge node in the content distribution network.
- the aforementioned computer-readable medium carries one or more programs, and when the aforementioned one or more programs are executed by the electronic device, the electronic device: receives a node evaluation request including at least two Internet Protocol addresses; Among at least two Internet Protocol addresses, select an Internet Protocol address; return the selected Internet Protocol address; wherein, the received Internet Protocol address indicates an edge node in the content distribution network.
- the computer program code used to perform the operations of the present disclosure can be written in one or more programming languages or a combination thereof.
- the above-mentioned programming languages include, but are not limited to, object-oriented programming languages—such as Java, Smalltalk, C++, and Including conventional procedural programming languages-such as "C" language or similar programming languages.
- the program code can be executed entirely on the user's computer, partly on the user's computer, executed as an independent software package, partly on the user's computer and partly executed on a remote computer, or entirely executed on the remote computer or server.
- the remote computer can be connected to the user's computer through any kind of network including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, using an Internet service provider to Connect via the Internet).
- LAN local area network
- WAN wide area network
- each block in the flowchart or block diagram can represent a module, program segment, or part of code, and the module, program segment, or part of code contains one or more for realizing the specified logic function.
- Executable instructions can also occur in a different order from the order marked in the drawings. For example, two blocks shown one after the other can actually be executed substantially in parallel, or they can sometimes be executed in the reverse order, depending on the functions involved.
- each block in the block diagram and/or flowchart, and the combination of the blocks in the block diagram and/or flowchart can be implemented by a dedicated hardware-based system that performs the specified functions or operations Or it can be realized by a combination of dedicated hardware and computer instructions.
- the modules involved in the embodiments described in the present disclosure can be implemented in software or hardware. Wherein, the name of the module does not constitute a limitation on the module itself under certain circumstances.
- the first obtaining module can also be described as "a module for obtaining at least two Internet Protocol addresses.”
- exemplary types of hardware logic components include: field programmable gate array (FPGA), application specific integrated circuit (ASIC), application specific standard product parts, ASSP), system on a chip (SOC), complex programmable logic device (CPLD), etc.
- FPGA field programmable gate array
- ASIC application specific integrated circuit
- ASSP application specific standard product parts
- SOC system on a chip
- CPLD complex programmable logic device
- a machine-readable medium may be a tangible medium, which may contain or store a program for use by the instruction execution system, apparatus, or device or in combination with the instruction execution system, apparatus, or device.
- the machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium.
- the machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing.
- machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing.
- RAM random access memory
- ROM read-only memory
- EPROM or flash memory erasable programmable read-only memory
- CD-ROM compact disk read only memory
- magnetic storage device or any suitable combination of the foregoing.
- Example 1 provides a video translation method, including converting the voice of the video to be translated into text; displaying the text and the first time information and second time information of the text Information and a reference translation, the first time information is the start time of the text in the video, and the second time information is the end time of the text in the video; in response to the user’s response to the The operation of the text or the reference translation shows the editing area, the editing area supports the user to input the translation; following the user’s input in the editing area, the translation suggestions from the reference translation are provided; wherein, in the detection In the case of the user's confirmation operation for the translation suggestion, the translation suggestion is displayed in the editing area as the translation result; and, in the case of detecting the user's non-confirmation operation for the translation suggestion Next, receiving a translation input by the user that is different from the translation suggestion, displaying the translation input by the user as the translation result in the editing area, and updating the translation according to the translation input by the user The reference translation in the translation
- Example 2 provides the method of Example 1.
- the text is segmented according to the time information and/or picture frames corresponding to the text in the video to obtain the multiple A clause text; for each clause text, the clause text and the first time information, the second time information and the reference translation of the clause text are displayed.
- Example 3 provides the method of Example 1.
- the text includes a plurality of clause texts, each of the clause texts is displayed in a different area, and the method further includes: A sub-function column, where the split function column supports the user to split the clause text; in response to the user's split operation for any one of the clause texts, split the clause text At least two sentence-divided texts are formed, and for each of the sentence-divided texts, the sentence-divided text, the first time information, the second time information of the sentence-divided text, and the reference translation of the sentence-divided text are displayed in association with each other .
- Example 4 provides the method of Example 1.
- the text includes a plurality of clause texts, each of the clause texts is displayed in a different area, and the method further includes: display merge A function bar, the merging function bar supports the user to merge the clause text; in response to the user's merging operation for any two adjacent clause texts, the two adjacent clause texts are merged Into a new clause text, and for the new clause text, the new clause text, the first time information, the second time information of the new clause text, and the new clause text are displayed in association with each other The reference translation of the clause text.
- Example 5 provides the method of Examples 1-4, the text includes a plurality of clause texts, each of the clause texts is displayed in a different area, and the method further includes: A play function bar is displayed, the play function bar supports the user to play the voice corresponding to the sentence text; in response to the user's operation on the play function bar, the voice corresponding to the sentence text is played.
- Example 6 provides the method of Examples 1-4, and the providing the translation suggestions from the reference translation includes: displaying in the editing area in a manner different from the input translation Displaying the translation suggestion; the displaying the translation suggestion in the editing area as a translation result in response to the user's confirmation operation on the translation suggestion includes: responding to the user's confirmation of the translation suggestion Operation to display the translation suggestion as the translation result in the editing area in the same manner as the display manner of the input translation.
- Example 7 provides the method of Examples 1-4.
- the translation suggestion is displayed to the translation result as a translation result.
- the editing area includes: in response to a user's input operation on the shortcut key, displaying the translation suggestion as a translation result in the editing area.
- Example 8 provides a video translation device.
- a conversion module is used to convert the voice of a video to be translated into text; and a display module is used to display the text and the content of the text.
- the first time information, the second time information and the reference translation the first time information is the start time of the text in the video, and the second time information is the end of the text in the video Time;
- the display module is also used to display the editing area in response to the user's operation on the text or the reference translation, and the editing area supports the user to input the translation;
- the suggestion module is used to follow the user in the editing area To provide translation suggestions from the reference translation;
- the display module is also used to display the translation suggestions as a translation result in the case of detecting the user's confirmation operation for the translation suggestions The editing area; and, in the case of detecting a non-confirmation operation of the user for the translation suggestion, receiving a translation input by the user that is different from the translation suggestion, and using the translation input by the user as The translation result is displayed in
- Example 9 provides the device of Example 8, and the display module is further configured to perform a display on the text according to the time information and/or picture frames corresponding to the text in the video. Clause to obtain the multiple clause texts; for each clause text, display the clause text and the first time information, the second time information and the reference translation of the clause text.
- Example 10 provides the device of Example 8.
- the text includes a plurality of clause texts, each of the clause texts is displayed in a different area, and the device further includes a splitting module , Used to display the split function bar, the split function bar supports the user to split the sentence text; in response to the user's split operation for any of the sentence text, the The sentenced text is split into at least two sentenced sentence texts, and for each sentenced sentence text, the sentenced sentence text, the first time information, the second time information of the sentenced text, and the divided sentence text are displayed in association with each other.
- the reference translation of the sentence text is used to display the split function bar, the split function bar supports the user to split the sentence text; in response to the user's split operation for any of the sentence text, the The sentenced text is split into at least two sentenced sentence texts, and for each sentenced sentence text, the sentenced sentence text, the first time information, the second time information of the sentenced text, and the divided sentence text are displayed in association with each other.
- the reference translation of the sentence text
- Example 11 provides the device of Example 8.
- the text includes a plurality of clause texts, each of the clause texts is displayed in a different area, and the device further includes a merging module, Used to display the merge function column, the merge function column supports the user to merge the clause text; in response to the user's merging operation for any two adjacent clause texts, the two adjacent The clause text is merged into a new clause text, and for the new clause text, the new clause text, the first time information, the second time information, and the new clause text are displayed in association with each other.
- the reference translation of the new clause text Used to display the merge function column, the merge function column supports the user to merge the clause text; in response to the user's merging operation for any two adjacent clause texts, the two adjacent The clause text is merged into a new clause text, and for the new clause text, the new clause text, the first time information, the second time information, and the new clause text are displayed in association with each other.
- the reference translation of the new clause text is used to display the merge function column, the merge
- Example 12 provides the device of Examples 8-11, where the text includes a plurality of sectioned texts, each of the sectioned texts is displayed in a different area, and the apparatus further includes playing Module, used to display the playback function bar, the playback function bar supports the user to play the voice corresponding to the sentence text; in response to the user's operation on the playback function bar, play the voice corresponding to the sentence text voice.
- Example 13 provides the device of Examples 8-11, and the suggestion module is used to display the translation suggestion in a display manner different from the input translation in the editing area; Said that in response to the user's confirmation operation on the translation suggestion, displaying the translation suggestion as a translation result in the editing area includes: responding to the user's confirmation operation on the translation suggestion to compare with the input translation suggestion.
- the translation suggestion as the translation result is displayed in the editing area in the same manner as the display method.
- Example 14 provides the device of Examples 8-11.
- the suggestion module is further configured to display the translation suggestion as a translation result in response to a user's triggering operation on the shortcut key.
- the editing area is further configured to display the translation suggestion as a translation result in response to a user's triggering operation on the shortcut key.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Multimedia (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Physics & Mathematics (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Human Computer Interaction (AREA)
- Acoustics & Sound (AREA)
- Machine Translation (AREA)
- Electrically Operated Instructional Devices (AREA)
Abstract
Description
Claims (10)
- 一种视频翻译方法,其特征在于,所述方法包括:将待翻译的视频的语音转换为文本;展示所述文本和所述文本的第一时间信息、第二时间信息及参考译文,所述第一时间信息为所述文本在所述视频中的起始时间,所述第二时间信息为所述文本在所述视频中的结束时间;响应于用户对所述文本或所述参考译文的操作,展示编辑区域,所述编辑区域支持所述用户输入译文;跟随所述用户在所述编辑区域的输入,提供来自所述参考译文的译文建议;其中,在检测到所述用户针对所述译文建议的确认操作的情况下,将所述译文建议作为译文结果显示到所述编辑区域;以及,在检测到所述用户针对所述译文建议的非确认操作的情况下,接收所述用户输入的不同于所述译文建议的译文,将所述用户输入的所述译文作为所述译文结果显示到所述编辑区域,根据所述用户输入的所述译文更新所述译文区域中的参考译文。
- 根据权利要求1所述的方法,其特征在于,所述展示所述文本和所述文本的第一时间信息、第二时间信息及参考译文,包括:根据所述文本在所述视频中对应的时刻信息和/或画面帧对所述文本进行分句,得到所述多个分句文本;针对每一所述分句文本,展示所述分句文本和所述分句文本的第一时间信息、第二时间信息及参考译文。
- 根据权利要求1所述的方法,其特征在于,所述文本包括多个分句文本,每一所述分句文本在不同区域展示,所述方法还包括:展示拆分功能栏,所述拆分功能栏支持所述用户对所述分句文本进行拆分;响应于所述用户针对任一所述分句文本的拆分操作,将所述分句文本拆分成至少两句分句子文本,并针对每一所述分句子文本,关联显示所述分句子文本、所述分句子文本的第一时间信息、第二时间信息以及所述分句子文本的参考译文。
- 根据权利要求1所述的方法,其特征在于,所述文本包括多个分句文本,每一所述分句文本在不同区域展示,所述方法还包括:展示合并功能栏,所述合并功能栏支持所述用户对所述分句文本进行合并;响应于所述用户针对任意相邻两个分句文本的合并操作,将所述相邻两个分句文本合并成一段新的分句文本,并针对所述新的分句文本,关联显示所述新的分句文本、所述新的分句文本的第一时间信息、第二时间信息以及所述新的分句文本的参考译文。
- 根据权利要求1-4任一项所述的方法,其特征在于,所述文本包括多个分句文本,每一所述分句文本在不同区域展示,所述方法还包括:展示播放功能栏,所述播放功能栏支持所述用户对所述分句文本对应的语音进行播放;响应于所述用户针对所述播放功能栏的操作,播放所述分句文本对应的语音。
- 根据权利要求1-4任一项所述的方法,其特征在于,所述提供来自所述参考译文的译文建议包括:在所述编辑区域以不同于已输入译文的显示方式来显示所述译文建议;所述响应于所述用户对所述译文建议的确认操作,将所述译文建议作为译文结果显示到所述编辑区域包括:响应于所述用户对所述译文建议的确认操作,以与已输入译文的显示方式相同的方式来在所述编辑区域内显示作为译文结果的所述译文建议。
- 根据权利要求1-4任一项所述的方法,其特征在于,所述响应于所述用户对所述译文建议的确认操作,将所述译文建议作为译文结果显示到所述编辑区域,包括:响应于用户对快捷键的触发操作,将所述译文建议作为译文结果显示到所述编辑区域。
- 一种视频翻译装置,其特征在于,所述装置包括:转换模块,用于将待翻译视频的语音转换为文本;展示模块,用于展示所述文本和所述文本的第一时间信息、第二时间信息及参考译文,所述第一时间信息为所述文本在所述视频中的起始时间,所述第二时间信息为所述文本在所述视频中的结束时间;展示模块,还用于响应于用户对所述文本或所述参考译文的操作,展示编辑区域,所述编辑区域支持所述用户输入译文;建议模块,用于跟随用户在所述编辑区域的输入,提供来自所述参考译文的译文建议;所述展示模块,还用于在检测到所述用户针对所述译文建议的确认操作的情况下,将所述译文建议作为译文结果显示到所述编辑区域;以及,在检测到所述用户针对所述译文建议的非确认操作的情况下,接收所述用户输入的不同于所述译文建议的译文,将所述用户输入的所述译文作为所述译文结果显示到所述编辑区域,根据所述用户输入的所述译文更新所述译文区域中的参考译文。
- 一种计算机可读介质,其上存储有计算机程序,其特征在于,该程序被处理装置执行时实现权利要求1-7中任一项所述方法的步骤。
- 一种电子设备,其特征在于,包括:存储装置,其上存储有计算机程序;处理装置,用于执行所述存储装置中的所述计算机程序,以实现权利要求1-7中任一项所述方法的步骤。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP21830302.2A EP4170543A4 (en) | 2020-06-23 | 2021-06-22 | VIDEO TRANSLATION METHOD AND APPARATUS, STORAGE MEDIUM AND ELECTRONIC DEVICE |
| JP2022564506A JP7548602B2 (ja) | 2020-06-23 | 2021-06-22 | ビデオ翻訳方法、装置、記憶媒体及び電子機器 |
| KR1020227030540A KR20220127361A (ko) | 2020-06-23 | 2021-06-22 | 비디오 번역 방법 및 장치, 저장 매체 및 전자 디바이스 |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202010583177.4 | 2020-06-23 | ||
| CN202010583177.4A CN111753558B (zh) | 2020-06-23 | 2020-06-23 | 视频翻译方法和装置、存储介质和电子设备 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US17/818,969 Continuation US11763103B2 (en) | 2020-06-23 | 2022-08-10 | Video translation method and apparatus, storage medium, and electronic device |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2021259221A1 true WO2021259221A1 (zh) | 2021-12-30 |
Family
ID=72676904
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2021/101388 Ceased WO2021259221A1 (zh) | 2020-06-23 | 2021-06-22 | 视频翻译方法和装置、存储介质和电子设备 |
Country Status (6)
| Country | Link |
|---|---|
| US (1) | US11763103B2 (zh) |
| EP (1) | EP4170543A4 (zh) |
| JP (1) | JP7548602B2 (zh) |
| KR (1) | KR20220127361A (zh) |
| CN (1) | CN111753558B (zh) |
| WO (1) | WO2021259221A1 (zh) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114596882A (zh) * | 2022-03-09 | 2022-06-07 | 云学堂信息科技(江苏)有限公司 | 一种可实现对课程内容快速定位的剪辑方法 |
| CN116095421A (zh) * | 2023-01-04 | 2023-05-09 | 北京达佳互联信息技术有限公司 | 字幕编辑方法、装置、电子设备及存储介质 |
| WO2023212920A1 (zh) * | 2022-05-06 | 2023-11-09 | 湖南师范大学 | 一种基于自建模板的多模态快速转写及标注系统 |
Families Citing this family (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111753558B (zh) * | 2020-06-23 | 2022-03-04 | 北京字节跳动网络技术有限公司 | 视频翻译方法和装置、存储介质和电子设备 |
| KR102812342B1 (ko) * | 2022-02-18 | 2025-05-27 | 에이아이링고 주식회사 | 번역된 콘텐츠의 편집 인터페이스 제공 방법 및 컴퓨터 프로그램 |
| JP2025048994A (ja) * | 2023-09-21 | 2025-04-03 | ソフトバンクグループ株式会社 | システム |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20150042771A1 (en) * | 2013-08-07 | 2015-02-12 | United Video Properties, Inc. | Methods and systems for presenting supplemental content in media assets |
| CN105828101A (zh) * | 2016-03-29 | 2016-08-03 | 北京小米移动软件有限公司 | 生成字幕文件的方法及装置 |
| CN107885729A (zh) * | 2017-09-25 | 2018-04-06 | 沈阳航空航天大学 | 基于双语片段的交互式机器翻译方法 |
| CN111753558A (zh) * | 2020-06-23 | 2020-10-09 | 北京字节跳动网络技术有限公司 | 视频翻译方法和装置、存储介质和电子设备 |
Family Cites Families (37)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US6549911B2 (en) * | 1998-11-02 | 2003-04-15 | Survivors Of The Shoah Visual History Foundation | Method and apparatus for cataloguing multimedia data |
| US6782384B2 (en) * | 2000-10-04 | 2004-08-24 | Idiom Merger Sub, Inc. | Method of and system for splitting and/or merging content to facilitate content processing |
| US7035804B2 (en) * | 2001-04-26 | 2006-04-25 | Stenograph, L.L.C. | Systems and methods for automated audio transcription, translation, and transfer |
| JP2005129971A (ja) * | 2002-01-28 | 2005-05-19 | Telecommunication Advancement Organization Of Japan | 半自動型字幕番組制作システム |
| US7111044B2 (en) * | 2002-07-17 | 2006-09-19 | Fastmobile, Inc. | Method and system for displaying group chat sessions on wireless mobile terminals |
| JP3999771B2 (ja) * | 2004-07-06 | 2007-10-31 | 株式会社東芝 | 翻訳支援プログラム、翻訳支援装置、翻訳支援方法 |
| JP2006166407A (ja) * | 2004-11-09 | 2006-06-22 | Canon Inc | 撮像装置及びその制御方法 |
| JP2007035056A (ja) * | 2006-08-29 | 2007-02-08 | Ebook Initiative Japan Co Ltd | 翻訳情報生成装置、翻訳情報生成方法並びにコンピュータプログラム |
| WO2009038209A1 (ja) * | 2007-09-20 | 2009-03-26 | Nec Corporation | 機械翻訳システム、機械翻訳方法及び機械翻訳プログラム |
| JP2010074482A (ja) * | 2008-09-18 | 2010-04-02 | Toshiba Corp | 外国語放送編集システム、翻訳サーバおよび翻訳支援方法 |
| US8843359B2 (en) * | 2009-02-27 | 2014-09-23 | Andrew Nelthropp Lauder | Language translation employing a combination of machine and human translations |
| US20100332214A1 (en) * | 2009-06-30 | 2010-12-30 | Shpalter Shahar | System and method for network transmision of subtitles |
| US20110246172A1 (en) * | 2010-03-30 | 2011-10-06 | Polycom, Inc. | Method and System for Adding Translation in a Videoconference |
| GB2502944A (en) * | 2012-03-30 | 2013-12-18 | Jpal Ltd | Segmentation and transcription of speech |
| EP2946279B1 (en) * | 2013-01-15 | 2019-10-16 | Viki, Inc. | System and method for captioning media |
| US9183198B2 (en) * | 2013-03-19 | 2015-11-10 | International Business Machines Corporation | Customizable and low-latency interactive computer-aided translation |
| CN103226947B (zh) * | 2013-03-27 | 2016-08-17 | 广东欧珀移动通信有限公司 | 一种基于移动终端的音频处理方法及装置 |
| WO2014198035A1 (en) * | 2013-06-13 | 2014-12-18 | Google Inc. | Techniques for user identification of and translation of media |
| JP6327848B2 (ja) * | 2013-12-20 | 2018-05-23 | 株式会社東芝 | コミュニケーション支援装置、コミュニケーション支援方法およびプログラム |
| US10169313B2 (en) * | 2014-12-04 | 2019-01-01 | Sap Se | In-context editing of text for elements of a graphical user interface |
| US9772816B1 (en) * | 2014-12-22 | 2017-09-26 | Google Inc. | Transcription and tagging system |
| CN104731776B (zh) * | 2015-03-27 | 2017-12-26 | 百度在线网络技术(北京)有限公司 | 翻译信息的提供方法及系统 |
| JP6470097B2 (ja) * | 2015-04-22 | 2019-02-13 | 株式会社東芝 | 通訳装置、方法およびプログラム |
| JP6471074B2 (ja) * | 2015-09-30 | 2019-02-13 | 株式会社東芝 | 機械翻訳装置、方法及びプログラム |
| US9558182B1 (en) * | 2016-01-08 | 2017-01-31 | International Business Machines Corporation | Smart terminology marker system for a language translation system |
| KR102495517B1 (ko) * | 2016-01-26 | 2023-02-03 | 삼성전자 주식회사 | 전자 장치, 전자 장치의 음성 인식 방법 |
| JP2017151768A (ja) * | 2016-02-25 | 2017-08-31 | 富士ゼロックス株式会社 | 翻訳プログラム及び情報処理装置 |
| CN108664201B (zh) * | 2017-03-29 | 2021-12-28 | 北京搜狗科技发展有限公司 | 一种文本编辑方法、装置及电子设备 |
| CN107943797A (zh) * | 2017-11-22 | 2018-04-20 | 语联网(武汉)信息技术有限公司 | 一种全原文参考的在线翻译系统 |
| CN108259965B (zh) * | 2018-03-31 | 2020-05-12 | 湖南广播电视台广播传媒中心 | 一种视频剪辑方法和剪辑系统 |
| KR102085908B1 (ko) * | 2018-05-10 | 2020-03-09 | 네이버 주식회사 | 컨텐츠 제공 서버, 컨텐츠 제공 단말 및 컨텐츠 제공 방법 |
| KR20210041007A (ko) * | 2018-08-29 | 2021-04-14 | 유장현 | 특허 문서 작성 장치, 방법, 컴퓨터 프로그램, 컴퓨터로 판독 가능한 기록매체, 서버 및 시스템 |
| US11636273B2 (en) * | 2019-06-14 | 2023-04-25 | Netflix, Inc. | Machine-assisted translation for subtitle localization |
| CN110489763B (zh) * | 2019-07-18 | 2023-03-10 | 深圳市轱辘车联数据技术有限公司 | 一种视频翻译方法及装置 |
| US11301644B2 (en) * | 2019-12-03 | 2022-04-12 | Trint Limited | Generating and editing media |
| US11580312B2 (en) * | 2020-03-16 | 2023-02-14 | Servicenow, Inc. | Machine translation of chat sessions |
| US11545156B2 (en) * | 2020-05-27 | 2023-01-03 | Microsoft Technology Licensing, Llc | Automated meeting minutes generation service |
-
2020
- 2020-06-23 CN CN202010583177.4A patent/CN111753558B/zh active Active
-
2021
- 2021-06-22 KR KR1020227030540A patent/KR20220127361A/ko not_active Ceased
- 2021-06-22 JP JP2022564506A patent/JP7548602B2/ja active Active
- 2021-06-22 EP EP21830302.2A patent/EP4170543A4/en active Pending
- 2021-06-22 WO PCT/CN2021/101388 patent/WO2021259221A1/zh not_active Ceased
-
2022
- 2022-08-10 US US17/818,969 patent/US11763103B2/en active Active
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20150042771A1 (en) * | 2013-08-07 | 2015-02-12 | United Video Properties, Inc. | Methods and systems for presenting supplemental content in media assets |
| CN105828101A (zh) * | 2016-03-29 | 2016-08-03 | 北京小米移动软件有限公司 | 生成字幕文件的方法及装置 |
| CN107885729A (zh) * | 2017-09-25 | 2018-04-06 | 沈阳航空航天大学 | 基于双语片段的交互式机器翻译方法 |
| CN111753558A (zh) * | 2020-06-23 | 2020-10-09 | 北京字节跳动网络技术有限公司 | 视频翻译方法和装置、存储介质和电子设备 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP4170543A4 |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN114596882A (zh) * | 2022-03-09 | 2022-06-07 | 云学堂信息科技(江苏)有限公司 | 一种可实现对课程内容快速定位的剪辑方法 |
| CN114596882B (zh) * | 2022-03-09 | 2024-02-02 | 云学堂信息科技(江苏)有限公司 | 一种可实现对课程内容快速定位的剪辑方法 |
| WO2023212920A1 (zh) * | 2022-05-06 | 2023-11-09 | 湖南师范大学 | 一种基于自建模板的多模态快速转写及标注系统 |
| CN116095421A (zh) * | 2023-01-04 | 2023-05-09 | 北京达佳互联信息技术有限公司 | 字幕编辑方法、装置、电子设备及存储介质 |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4170543A4 (en) | 2023-10-25 |
| CN111753558A (zh) | 2020-10-09 |
| JP2023522469A (ja) | 2023-05-30 |
| KR20220127361A (ko) | 2022-09-19 |
| EP4170543A1 (en) | 2023-04-26 |
| US11763103B2 (en) | 2023-09-19 |
| US20220383000A1 (en) | 2022-12-01 |
| JP7548602B2 (ja) | 2024-09-10 |
| CN111753558B (zh) | 2022-03-04 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12483683B2 (en) | Interactive information processing method, device and medium | |
| WO2021259221A1 (zh) | 视频翻译方法和装置、存储介质和电子设备 | |
| CN111898388B (zh) | 视频字幕翻译编辑方法、装置、电子设备及存储介质 | |
| CN113010698B (zh) | 多媒体的交互方法、信息交互方法、装置、设备及介质 | |
| US11954455B2 (en) | Method for translating words in a picture, electronic device, and storage medium | |
| CN113778419B (zh) | 多媒体数据的生成方法、装置、可读介质及电子设备 | |
| US11580314B2 (en) | Document translation method and apparatus, storage medium, and electronic device | |
| JP7548678B2 (ja) | オーディオとテキストとの同期方法、装置、読取可能な媒体及び電子機器 | |
| US20250106486A1 (en) | Multimedia data processing method and apparatus, device and medium | |
| CN112163433B (zh) | 关键词汇的匹配方法、装置、电子设备及存储介质 | |
| CN115906875A (zh) | 文档翻译方法、设备、存储介质及程序产品 | |
| US12174880B2 (en) | Method for searching target content, and electronic device and storage medium | |
| CN111523330A (zh) | 用于生成文本的方法、装置、电子设备和介质 | |
| CN120596699A (zh) | 有声书籍的播放方法、装置、电子设备、介质和产品 | |
| WO2021161908A1 (ja) | 情報処理装置及び情報処理方法 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 21830302 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2022564506 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2021830302 Country of ref document: EP Effective date: 20230123 |
|
| WWR | Wipo information: refused in national office |
Ref document number: 1020227030540 Country of ref document: KR |