JP2021069117A - ローカライズされたコンテキストのビデオ注釈を生成するためのシステム及び方法 - Google Patents
ローカライズされたコンテキストのビデオ注釈を生成するためのシステム及び方法 Download PDFInfo
- Publication number
- JP2021069117A JP2021069117A JP2020177288A JP2020177288A JP2021069117A JP 2021069117 A JP2021069117 A JP 2021069117A JP 2020177288 A JP2020177288 A JP 2020177288A JP 2020177288 A JP2020177288 A JP 2020177288A JP 2021069117 A JP2021069117 A JP 2021069117A
- Authority
- JP
- Japan
- Prior art keywords
- segment
- video
- segments
- classification
- annotation
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Granted
Links
Images
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/83—Generation or processing of protective or descriptive data associated with content; Content structuring
- H04N21/845—Structuring of content, e.g. decomposing content into time segments
- H04N21/8456—Structuring of content, e.g. decomposing content into time segments by decomposing the content in the time domain, e.g. in time segments
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/25—Fusion techniques
- G06F18/253—Fusion techniques of extracted features
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/25—Fusion techniques
- G06F18/254—Fusion techniques of classification results, e.g. of results related to same input data
- G06F18/256—Fusion techniques of classification results, e.g. of results related to same input data of results relating to different input data, e.g. multimodal recognition
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/103—Formatting, i.e. changing of presentation of documents
- G06F40/117—Tagging; Marking up; Designating a block; Setting of attributes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/41—Higher-level, semantic clustering, classification or understanding of video scenes, e.g. detection, labelling or Markovian modelling of sport events or news items
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
- G06V20/49—Segmenting video sequences, i.e. computational techniques such as parsing or cutting the sequence, low-level clustering or determining units such as shots or scenes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/70—Labelling scene content, e.g. deriving syntactic or semantic representations
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/04—Segmentation; Word boundary detection
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/08—Speech classification or search
- G10L15/18—Speech classification or search using natural language modelling
- G10L15/1815—Semantic context, e.g. disambiguation of the recognition hypotheses based on word meaning
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0272—Voice signal separating
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/48—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
- G10L25/51—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination
- G10L25/63—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for comparison or discrimination for estimating an emotional state
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/233—Processing of audio elementary streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/23—Processing of content or additional data; Elementary server operations; Server middleware
- H04N21/234—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs
- H04N21/23418—Processing of video elementary streams, e.g. splicing of video streams or manipulating encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/25—Management operations performed by the server for facilitating the content distribution or administrating data related to end-users or client devices, e.g. end-user or client device authentication, learning user preferences for recommending movies
- H04N21/251—Learning process for intelligent management, e.g. learning user preferences for recommending movies
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/20—Servers specifically adapted for the distribution of content, e.g. VOD servers; Operations thereof
- H04N21/25—Management operations performed by the server for facilitating the content distribution or administrating data related to end-users or client devices, e.g. end-user or client device authentication, learning user preferences for recommending movies
- H04N21/266—Channel or content management, e.g. generation and management of keys and entitlement messages in a conditional access system, merging a VOD unicast channel into a multicast channel
- H04N21/26603—Channel or content management, e.g. generation and management of keys and entitlement messages in a conditional access system, merging a VOD unicast channel into a multicast channel for automatically generating descriptors from content, e.g. when it is not made available by its provider, using content analysis techniques
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
- H04N21/4394—Processing of audio elementary streams involving operations for analysing the audio stream, e.g. detecting features or characteristics in audio streams
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/439—Processing of audio elementary streams
- H04N21/4398—Processing of audio elementary streams involving reformatting operations of audio signals
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44008—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics in the video stream
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/40—Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
- H04N21/43—Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
- H04N21/44—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
- H04N21/44016—Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving splicing one content stream with another content stream, e.g. for substituting a video clip
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N21/00—Selective content distribution, e.g. interactive television or video on demand [VOD]
- H04N21/80—Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
- H04N21/81—Monomedia components thereof
- H04N21/812—Monomedia components thereof involving advertisement data
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Physics & Mathematics (AREA)
- Theoretical Computer Science (AREA)
- Signal Processing (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- General Physics & Mathematics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Acoustics & Sound (AREA)
- Human Computer Interaction (AREA)
- Artificial Intelligence (AREA)
- General Engineering & Computer Science (AREA)
- Data Mining & Analysis (AREA)
- General Health & Medical Sciences (AREA)
- Computing Systems (AREA)
- Databases & Information Systems (AREA)
- Evolutionary Biology (AREA)
- Evolutionary Computation (AREA)
- Bioinformatics & Computational Biology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Software Systems (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Quality & Reliability (AREA)
- Child & Adolescent Psychology (AREA)
- Hospice & Palliative Care (AREA)
- Psychiatry (AREA)
- Business, Economics & Management (AREA)
- Marketing (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
- Television Signal Processing For Recording (AREA)
Abstract
Description
本出願は、2019年10月22日に出願され、発明者Suresh Kumar及びRaja Balaによる「Personalized Viewing Experience and Contextual Advertising in Streaming Media based on Scene Level Annotation of Videos」と題する米国仮出願第62/924,586号、代理人整理番号PARC−20190487US01の利益を主張するものであり、その開示は、参照により本明細書に組み込まれる。
本開示は、一般的に人工知能(artificial intelligence、AI)の分野に関連する。より具体的には、本開示は、ビデオからローカライズされたコンテキスト情報を取得し、ビデオの個々のセグメントに注釈付けするためにローカライズされたコンテキスト情報を利用するためのシステム及び方法に関する。
概要
例示的なシステム
セグメントレベルの注釈
システムアーキテクチャ
広告システム
操作
例示的なコンピュータシステム及び装置
Claims (26)
- ローカライズされたコンテキストビデオ注釈のための方法であって、
セグメント化ユニットに基づいて、ビデオを複数のセグメントに分割することと、
セグメントの複数の入力モダリティを決定するために、それぞれの前記セグメントを解析することであって、それぞれの入力モダリティが、前記セグメントのコンテンツの形態を示す、解析することと、
前記入力モダリティに基づいて、前記セグメントを一組の意味クラスに分類することと、
前記一組の意味クラスに基づいて、前記セグメントの注釈を決定することであって、前記注釈が、前記セグメントと関連付けられた意味コンテキスト情報を示す、決定することと、を含む、方法。 - 前記セグメントを分類することが、
入力モダリティの分類を決定するために、対応する分類子をそれぞれの前記入力モダリティに適用することと、
前記複数の入力モダリティの前記分類に基づいて、前記セグメントの統一分類を決定することと、
前記統一分類に基づいて、前記セグメントの前記注釈を決定することと、を更に含む、請求項1に記載の方法。 - 前記統一分類を決定することが、前記複数の入力モダリティの前記分類を互いに融合させて、前記統一分類を生成することを更に含む、請求項2に記載の方法。
- 前記複数の入力モダリティが、前記セグメントのオーディオ信号から分離されたビデオフレームを含み、
前記セグメントを分類することが、ディープビジュアル分類子を前記ビデオフレームに適用して、前記セグメントのビジュアル分類を生成することを更に含む、請求項1に記載の方法。 - 前記複数の入力モダリティが、前記セグメントのビデオフレームから分離されたオーディオ信号を含み、
前記セグメントを分類することが、
前記オーディオ信号をバックグラウンド信号及び音声信号に分解することと、
オーディオ分類子を前記バックグラウンド信号に適用して、前記セグメントのバックグラウンドオーディオ分類を生成することと、
感情分類子を前記音声信号に適用して、前記セグメントの感情分類を生成することと、を更に含む、請求項1に記載の方法。 - 前記複数の入力モダリティが、前記セグメントのオーディオビジュアル信号から分離されたテキスト情報を含み、
前記セグメントを分類することが、
前記セグメントの言語的音声を表す音声テキストを取得することと、
前記テキスト情報を前記音声テキストと整列させることと、
テキストベースの分類子を前記整列させたテキストに適用して、前記セグメントのテキスト分類を生成することと、を更に含む、請求項1に記載の方法。 - 前記ムービーのスクリプトを取得し、更に、前記スクリプトを前記音声テキストと整列させることを更に含む、請求項6に記載の方法。
- 前記セグメントを分類することが、
前記複数の入力モダリティのそれぞれの特徴埋め込みを取得することと、
前記特徴埋め込みを組み合わせて、統一埋め込みを生成することと、
意味分類子を前記統一埋め込みに適用して、前記統一分類を決定することと、を更に含む、請求項1に記載の方法。 - 前記特徴埋め込みを組み合わせることが、特徴連結を前記特徴埋め込みに適用することを更に含む、請求項8に記載の方法。
- 前記注釈が、前記セグメントの前記意味コンテキスト情報を示す一組のキーを含み、それぞれのキーが、値と、強度と、を含み、前記値が、前記セグメントの特徴を示し、前記強度は、前記値が前記セグメントと関連付けられる可能性を示す、請求項1に記載の方法。
- 前記セグメント化ユニットが、前記ビデオのアクト、シーン、ビート、及びショットのうちの1つ以上である、請求項1に記載の方法。
- それぞれの意味クラスが、アクション、危険、ロマンス、フレンドシップ、及びアウトドアアドベンチャーのうちの1つに対応する、請求項1に記載の方法。
- コンピュータによって実行されると、前記コンピュータに、ローカライズされたコンテキストビデオ注釈のための方法を実行させる命令を記憶している、非一時的コンピュータ可読記憶媒体であって、前記方法が、
セグメント化ユニットに基づいて、ビデオを複数のセグメントに分割することと、
前記セグメントの複数の入力モダリティを生成するために、それぞれのセグメントを解析することであって、それぞれの入力モダリティが、前記セグメントのコンテンツの形態を示す、解析することと、
前記入力モダリティに基づいて、前記セグメントを一組の意味クラスに分類することと、
前記一組の意味クラスに基づいて、前記セグメントの注釈を決定することであって、前記注釈が、前記セグメントと関連付けられた意味コンテキスト情報を示す、決定することと、を含む、非一時的コンピュータ可読記憶媒体。 - 前記セグメントを分類することが、
前記入力モダリティの分類を決定するために、対応する分類子をそれぞれの入力モダリティに適用することと、
前記複数の入力モダリティの前記分類に基づいて、前記セグメントの統一分類を決定することと、
前記統一分類に基づいて、前記セグメントの前記注釈を決定することと、を更に含む、請求項13に記載のコンピュータ可読記憶媒体。 - 前記統一分類を決定することが、前記複数の入力モダリティの前記分類を互いに融合させて、前記統一分類を生成することを更に含む、請求項14に記載のコンピュータ可読記憶媒体。
- 前記複数の入力モダリティが、前記セグメントのオーディオ信号から分離されたビデオフレームを含み、
前記セグメントを分類することが、ディープビジュアル分類子を前記ビデオフレームに適用して、前記セグメントのビジュアル分類を生成することを更に含む、請求項13に記載のコンピュータ可読記憶媒体。 - 前記複数の入力モダリティが、前記セグメントのビデオフレームから分離されたオーディオ信号を含み、
前記セグメントを分類することが、
前記オーディオ信号をバックグラウンド信号及び音声信号に分解することと、
オーディオ分類子を前記バックグラウンド信号に適用して、前記セグメントのバックグラウンドオーディオ分類を生成することと、
感情分類子を前記音声信号に適用して、前記セグメントの感情分類を生成することと、を更に含む、請求項13に記載のコンピュータ可読記憶媒体。 - 前記複数の入力モダリティが、前記セグメントのオーディオビジュアル信号から分離されたテキスト情報を含み、
前記セグメントを分類することが、
前記セグメントの言語的音声を表す音声テキストを取得することと、
前記テキスト情報を前記音声テキストと整列させることと、
テキストベースの分類子を前記整列させたテキストに適用して、前記セグメントのテキスト分類を生成することと、を更に含む、請求項13に記載のコンピュータ可読記憶媒体。 - 前記方法が、前記ムービーのスクリプトを取得し、更に、前記スクリプトを前記音声テキストと整列させることを更に含む、請求項18に記載のコンピュータ可読記憶媒体。
- 前記セグメントを分類することが、
前記複数の入力モダリティの前記分類からそれぞれの特徴埋め込みを取得することと、
前記特徴埋め込みを組み合わせて、統一埋め込みを生成することと、
意味分類子を前記統一埋め込みに適用して、前記統一分類を決定することと、を更に含む、請求項13に記載のコンピュータ可読記憶媒体。 - 前記特徴埋め込みを組み合わせることが、特徴連結を前記特徴埋め込みに適用することを更に含む、請求項20に記載のコンピュータ可読記憶媒体。
- 前記注釈が、前記セグメントの前記意味コンテキスト情報を示す一組のキーを含み、それぞれのキーが、値と、強度と、を含み、前記値が、前記セグメントの特徴を示し、前記強度は、前記値が前記セグメントと関連付けられる可能性を示す、請求項13に記載のコンピュータ可読記憶媒体。
- 前記セグメント化ユニットが、前記ビデオのアクト、シーン、ビート、及びショットのうちの1つ以上である、請求項13に記載のコンピュータ可読記憶媒体。
- それぞれの意味クラスが、アクション、危険、ロマンス、フレンドシップ、及びアウトドアアドベンチャーのうちの1つに対応する、請求項13に記載のコンピュータ可読記憶媒体。
- ローカライズされたコンテキストビデオ注釈に基づいて広告を配置するための方法であって、
セグメント化ユニットに基づいて、ビデオを複数のセグメントに分割することと、
セグメントの複数の入力モダリティを決定するために、それぞれの前記セグメントを解析することであって、それぞれの入力モダリティが、前記セグメントのコンテンツの形態を示す、解析することと、
前記入力モダリティに基づいて、前記セグメントを一組の意味クラスに分類することと、
前記一組の意味クラスに基づいて、前記セグメントの注釈を決定することであって、前記注釈が、前記セグメントと関連付けられた意味コンテキスト情報を示す、決定することと、
広告を配置するための標的位置として、前記ビデオファイルのセグメント間のセグメント間可用性(ISA)を識別することと、
前記ISAと関連付けられた一組のセグメントの注釈を広告システムに送信することであって、前記一組のセグメントが、前記ISAの先行セグメント及び前記ISAの後続セグメントのうちの1つ以上を含む、送信することと、を含む、方法。 - ローカライズされたコンテキストビデオ注釈に基づいて自由裁量の視聴を容易にするための方法であって、
セグメント化ユニットに基づいて、ビデオを複数のセグメントに分割することと、
セグメントの複数の入力モダリティを決定するために、それぞれの前記セグメントを解析することであって、それぞれの入力モダリティが、前記セグメントのコンテンツの形態を示す、解析することと、
前記入力モダリティに基づいて、前記セグメントを一組の意味クラスに分類することと、
前記一組の意味クラスに基づいて、前記セグメントの注釈を決定することであって、前記注釈が、前記セグメントと関連付けられた意味コンテキスト情報を示す、決定することと、
前記ビデオの視聴者から視聴選好を取得することと、
前記複数のセグメントの前記注釈に基づいて、前記複数のセグメントから一組の視聴セグメントを決定することであって、前記一組の視聴セグメントが、前記視聴選好に従う、決定することと、を含む、方法。
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US201962924586P | 2019-10-22 | 2019-10-22 | |
| US62/924,586 | 2019-10-22 | ||
| US17/026,897 | 2020-09-21 | ||
| US17/026,897 US11270123B2 (en) | 2019-10-22 | 2020-09-21 | System and method for generating localized contextual video annotation |
Publications (3)
| Publication Number | Publication Date |
|---|---|
| JP2021069117A true JP2021069117A (ja) | 2021-04-30 |
| JP2021069117A5 JP2021069117A5 (ja) | 2023-10-20 |
| JP7498640B2 JP7498640B2 (ja) | 2024-06-12 |
Family
ID=73005486
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| JP2020177288A Active JP7498640B2 (ja) | 2019-10-22 | 2020-10-22 | ローカライズされたコンテキストのビデオ注釈を生成するためのシステム及び方法 |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US11270123B2 (ja) |
| EP (1) | EP3813376A1 (ja) |
| JP (1) | JP7498640B2 (ja) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20230029242A1 (en) * | 2021-07-16 | 2023-01-26 | Sintokogio, Ltd. | Screen image generation method, screen image generation device, and storage medium |
| JP2024528440A (ja) * | 2022-05-10 | 2024-07-30 | 北京字跳▲網▼絡技▲術▼有限公司 | ビデオ生成方法、装置、デバイス、記憶媒体およびプログラム製品 |
Families Citing this family (29)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US11392791B2 (en) * | 2018-08-31 | 2022-07-19 | Writer, Inc. | Generating training data for natural language processing |
| AU2021231754A1 (en) * | 2020-03-02 | 2022-09-15 | Visual Supply Company | Systems and methods for automating video editing |
| US11458409B2 (en) * | 2020-05-27 | 2022-10-04 | Nvidia Corporation | Automatic classification and reporting of inappropriate language in online applications |
| US11973993B2 (en) * | 2020-10-28 | 2024-04-30 | Nagravision S.A. | Machine learning based media content annotation |
| CN112584062B (zh) * | 2020-12-10 | 2023-08-08 | 上海幻电信息科技有限公司 | 背景音频构建方法及装置 |
| US11636677B2 (en) * | 2021-01-08 | 2023-04-25 | Huawei Technologies Co., Ltd. | Systems, devices and methods for distributed hierarchical video analysis |
| US11418821B1 (en) * | 2021-02-09 | 2022-08-16 | Gracenote, Inc. | Classifying segments of media content using closed captioning |
| US12132953B2 (en) | 2021-02-16 | 2024-10-29 | Gracenote, Inc. | Identifying and labeling segments within video content |
| US12249147B2 (en) * | 2021-03-11 | 2025-03-11 | International Business Machines Corporation | Adaptive selection of data modalities for efficient video recognition |
| US11361421B1 (en) | 2021-05-21 | 2022-06-14 | Airbnb, Inc. | Visual attractiveness scoring system |
| US11200449B1 (en) | 2021-05-21 | 2021-12-14 | Airbnb, Inc. | Image ranking system |
| US12333794B2 (en) | 2021-11-12 | 2025-06-17 | Sony Group Corporation | Emotion recognition in multimedia videos using multi-modal fusion-based deep neural network |
| WO2023084348A1 (en) * | 2021-11-12 | 2023-05-19 | Sony Group Corporation | Emotion recognition in multimedia videos using multi-modal fusion-based deep neural network |
| CN114125566B (zh) * | 2021-12-29 | 2024-03-08 | 阿里巴巴(中国)有限公司 | 互动方法、系统及电子设备 |
| US11910061B2 (en) | 2022-05-23 | 2024-02-20 | Rovi Guides, Inc. | Leveraging emotional transitions in media to modulate emotional impact of secondary content |
| US11871081B2 (en) | 2022-05-23 | 2024-01-09 | Rovi Guides, Inc. | Leveraging emotional transitions in media to modulate emotional impact of secondary content |
| US12041278B1 (en) * | 2022-06-29 | 2024-07-16 | Amazon Technologies, Inc. | Computer-implemented methods of an automated framework for virtual product placement in video frames |
| US12573146B2 (en) * | 2023-04-20 | 2026-03-10 | Adeia Guides Inc. | Product placement systems and methods for 3D productions |
| US20240354785A1 (en) * | 2023-04-21 | 2024-10-24 | Disney Enterprises, Inc. | Feature addition analysis for an item |
| US12549815B2 (en) * | 2023-06-30 | 2026-02-10 | Roku, Inc. | Context classification of streaming content using machine learning |
| US20250037734A1 (en) * | 2023-07-28 | 2025-01-30 | Qualcomm Incorporated | Selective processing of segments of time-series data based on segment classification |
| US12541949B2 (en) | 2023-10-31 | 2026-02-03 | Roku, Inc. | Contextual understanding of media content to generate targeted media content |
| US20250139969A1 (en) * | 2023-10-31 | 2025-05-01 | Roku, Inc. | Processing and contextual understanding of video segments |
| CN117692676B (zh) * | 2023-12-08 | 2024-11-15 | 广东创意热店互联网科技有限公司 | 一种基于人工智能技术的视频快速剪辑方法 |
| CN118035565B (zh) * | 2024-04-11 | 2024-06-11 | 烟台大学 | 基于多模态情绪感知的主动服务推荐方法、系统和设备 |
| JP2025167140A (ja) * | 2024-04-25 | 2025-11-07 | 株式会社東芝 | プログラム、映像処理装置および映像処理方法 |
| CN118296186B (zh) * | 2024-06-05 | 2024-10-11 | 上海蜜度科技股份有限公司 | 视频广告检测方法、系统、存储介质及电子设备 |
| CN118741176B (zh) * | 2024-09-03 | 2024-12-31 | 腾讯科技(深圳)有限公司 | 广告植入信息处理方法、相关装置和介质 |
| US20260075295A1 (en) * | 2024-09-10 | 2026-03-12 | Adobe Inc. | Determining topic chapters for digital videos utilizing video segmentation machine learning models |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2011155329A (ja) * | 2010-01-26 | 2011-08-11 | Nippon Telegr & Teleph Corp <Ntt> | 映像コンテンツ編集装置,映像コンテンツ編集方法および映像コンテンツ編集プログラム |
| WO2018124309A1 (en) * | 2016-12-30 | 2018-07-05 | Mitsubishi Electric Corporation | Method and system for multi-modal fusion model |
Family Cites Families (43)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US4570232A (en) * | 1981-12-21 | 1986-02-11 | Nippon Telegraph & Telephone Public Corporation | Speech recognition apparatus |
| US5598557A (en) * | 1992-09-22 | 1997-01-28 | Caere Corporation | Apparatus and method for retrieving and grouping images representing text files based on the relevance of key words extracted from a selected file to the text files |
| US6112186A (en) * | 1995-06-30 | 2000-08-29 | Microsoft Corporation | Distributed system for facilitating exchange of user information and opinion using automated collaborative filtering |
| US6098082A (en) * | 1996-07-15 | 2000-08-01 | At&T Corp | Method for automatically providing a compressed rendition of a video program in a format suitable for electronic searching and retrieval |
| US6085160A (en) * | 1998-07-10 | 2000-07-04 | Lernout & Hauspie Speech Products N.V. | Language independent speech recognition |
| US6243676B1 (en) * | 1998-12-23 | 2001-06-05 | Openwave Systems Inc. | Searching and retrieving multimedia information |
| US6473778B1 (en) * | 1998-12-24 | 2002-10-29 | At&T Corporation | Generating hypermedia documents from transcriptions of television programs using parallel text alignment |
| CA2366057C (en) * | 1999-03-05 | 2009-03-24 | Canon Kabushiki Kaisha | Database annotation and retrieval |
| US6442518B1 (en) * | 1999-07-14 | 2002-08-27 | Compaq Information Technologies Group, L.P. | Method for refining time alignments of closed captions |
| US7047191B2 (en) * | 2000-03-06 | 2006-05-16 | Rochester Institute Of Technology | Method and system for providing automated captioning for AV signals |
| US6925455B2 (en) * | 2000-12-12 | 2005-08-02 | Nec Corporation | Creating audio-centric, image-centric, and integrated audio-visual summaries |
| US7065524B1 (en) * | 2001-03-30 | 2006-06-20 | Pharsight Corporation | Identification and correction of confounders in a statistical analysis |
| US7110664B2 (en) * | 2001-04-20 | 2006-09-19 | Front Porch Digital, Inc. | Methods and apparatus for indexing and archiving encoded audio-video data |
| US7035468B2 (en) * | 2001-04-20 | 2006-04-25 | Front Porch Digital Inc. | Methods and apparatus for archiving, indexing and accessing audio and video data |
| US7908628B2 (en) * | 2001-08-03 | 2011-03-15 | Comcast Ip Holdings I, Llc | Video and digital multimedia aggregator content coding and formatting |
| US20030061028A1 (en) * | 2001-09-21 | 2003-03-27 | Knumi Inc. | Tool for automatically mapping multimedia annotations to ontologies |
| US7092888B1 (en) * | 2001-10-26 | 2006-08-15 | Verizon Corporate Services Group Inc. | Unsupervised training in natural language call routing |
| AU2002363907A1 (en) * | 2001-12-24 | 2003-07-30 | Scientific Generics Limited | Captioning system |
| US8522267B2 (en) * | 2002-03-08 | 2013-08-27 | Caption Colorado Llc | Method and apparatus for control of closed captioning |
| US7440895B1 (en) * | 2003-12-01 | 2008-10-21 | Lumenvox, Llc. | System and method for tuning and testing in a speech recognition system |
| US20070124788A1 (en) * | 2004-11-25 | 2007-05-31 | Erland Wittkoter | Appliance and method for client-sided synchronization of audio/video content and external data |
| US7739253B1 (en) * | 2005-04-21 | 2010-06-15 | Sonicwall, Inc. | Link-based content ratings of pages |
| US7382933B2 (en) * | 2005-08-24 | 2008-06-03 | International Business Machines Corporation | System and method for semantic video segmentation based on joint audiovisual and text analysis |
| US7801910B2 (en) * | 2005-11-09 | 2010-09-21 | Ramp Holdings, Inc. | Method and apparatus for timed tagging of media content |
| US20070124147A1 (en) * | 2005-11-30 | 2007-05-31 | International Business Machines Corporation | Methods and apparatus for use in speech recognition systems for identifying unknown words and for adding previously unknown words to vocabularies and grammars of speech recognition systems |
| US8209724B2 (en) * | 2007-04-25 | 2012-06-26 | Samsung Electronics Co., Ltd. | Method and system for providing access to information of potential interest to a user |
| JP4158937B2 (ja) * | 2006-03-24 | 2008-10-01 | インターナショナル・ビジネス・マシーンズ・コーポレーション | 字幕修正装置 |
| US8045054B2 (en) * | 2006-09-13 | 2011-10-25 | Nortel Networks Limited | Closed captioning language translation |
| US20080120646A1 (en) * | 2006-11-20 | 2008-05-22 | Stern Benjamin J | Automatically associating relevant advertising with video content |
| US8583416B2 (en) * | 2007-12-27 | 2013-11-12 | Fluential, Llc | Robust information extraction from utterances |
| US7509385B1 (en) * | 2008-05-29 | 2009-03-24 | International Business Machines Corporation | Method of system for creating an electronic message |
| US8131545B1 (en) * | 2008-09-25 | 2012-03-06 | Google Inc. | Aligning a transcript to audio data |
| US8423363B2 (en) * | 2009-01-13 | 2013-04-16 | CRIM (Centre de Recherche Informatique de Montréal) | Identifying keyword occurrences in audio data |
| US8843368B2 (en) * | 2009-08-17 | 2014-09-23 | At&T Intellectual Property I, L.P. | Systems, computer-implemented methods, and tangible computer-readable storage media for transcription alignment |
| US8707381B2 (en) * | 2009-09-22 | 2014-04-22 | Caption Colorado L.L.C. | Caption and/or metadata synchronization for replay of previously or simultaneously recorded live programs |
| US8572488B2 (en) * | 2010-03-29 | 2013-10-29 | Avid Technology, Inc. | Spot dialog editor |
| US8867891B2 (en) * | 2011-10-10 | 2014-10-21 | Intellectual Ventures Fund 83 Llc | Video concept classification using audio-visual grouplets |
| EP2642484A1 (en) * | 2012-03-23 | 2013-09-25 | Thomson Licensing | Method for setting a watching level for an audiovisual content |
| US9100721B2 (en) * | 2012-12-06 | 2015-08-04 | Cable Television Laboratories, Inc. | Advertisement insertion |
| US20160014482A1 (en) * | 2014-07-14 | 2016-01-14 | The Board Of Trustees Of The Leland Stanford Junior University | Systems and Methods for Generating Video Summary Sequences From One or More Video Segments |
| US9955218B2 (en) * | 2015-04-28 | 2018-04-24 | Rovi Guides, Inc. | Smart mechanism for blocking media responsive to user environment |
| KR102450840B1 (ko) * | 2015-11-19 | 2022-10-05 | 엘지전자 주식회사 | 전자 기기 및 전자 기기의 제어 방법 |
| US10262239B2 (en) * | 2016-07-26 | 2019-04-16 | Viisights Solutions Ltd. | Video content contextual classification |
-
2020
- 2020-09-21 US US17/026,897 patent/US11270123B2/en active Active
- 2020-10-21 EP EP20203180.3A patent/EP3813376A1/en not_active Ceased
- 2020-10-22 JP JP2020177288A patent/JP7498640B2/ja active Active
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2011155329A (ja) * | 2010-01-26 | 2011-08-11 | Nippon Telegr & Teleph Corp <Ntt> | 映像コンテンツ編集装置,映像コンテンツ編集方法および映像コンテンツ編集プログラム |
| WO2018124309A1 (en) * | 2016-12-30 | 2018-07-05 | Mitsubishi Electric Corporation | Method and system for multi-modal fusion model |
Cited By (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20230029242A1 (en) * | 2021-07-16 | 2023-01-26 | Sintokogio, Ltd. | Screen image generation method, screen image generation device, and storage medium |
| JP2024528440A (ja) * | 2022-05-10 | 2024-07-30 | 北京字跳▲網▼絡技▲術▼有限公司 | ビデオ生成方法、装置、デバイス、記憶媒体およびプログラム製品 |
| JP7732004B2 (ja) | 2022-05-10 | 2025-09-01 | 北京字跳▲網▼絡技▲術▼有限公司 | ビデオ生成方法、装置、デバイス、記憶媒体およびプログラム製品 |
| US12586610B2 (en) | 2022-05-10 | 2026-03-24 | Beijing Zitiao Network Technology Co., Ltd. | Method, apparatus, device, storage medium and program product for video generation |
Also Published As
| Publication number | Publication date |
|---|---|
| US11270123B2 (en) | 2022-03-08 |
| JP7498640B2 (ja) | 2024-06-12 |
| US20210117685A1 (en) | 2021-04-22 |
| EP3813376A1 (en) | 2021-04-28 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP7498640B2 (ja) | ローカライズされたコンテキストのビデオ注釈を生成するためのシステム及び方法 | |
| US20210065249A1 (en) | System and method for delivering personalized content based on deep learning of user emotions | |
| CN102244807B (zh) | 自适应视频变焦 | |
| US10970334B2 (en) | Navigating video scenes using cognitive insights | |
| US8804999B2 (en) | Video recommendation system and method thereof | |
| US9813779B2 (en) | Method and apparatus for increasing user engagement with video advertisements and content by summarization | |
| US9888279B2 (en) | Content based video content segmentation | |
| US9100701B2 (en) | Enhanced video systems and methods | |
| JP2021069117A5 (ja) | ||
| US20140143218A1 (en) | Method for Crowd Sourced Multimedia Captioning for Video Content | |
| Tran et al. | Exploiting character networks for movie summarization | |
| CN102207954A (zh) | 电子设备、内容推荐方法及其程序 | |
| US20240312488A1 (en) | System and method for generating personalized video trailers | |
| CN114845149A (zh) | 视频片段的剪辑方法、视频推荐方法、装置、设备及介质 | |
| US20250246176A1 (en) | Video scene describer | |
| US20190313154A1 (en) | Customizing digital content based on consumer context data | |
| CN118803350A (zh) | 视频播放方法、装置、设备、介质及产品 | |
| KR100768074B1 (ko) | 광고 동영상을 제공하는 시스템 및 그 서비스 방법 | |
| US20250008189A1 (en) | Personalized Bandwidth Preservation | |
| CN119383401B (zh) | 一种用于提取和插入图片或视频广告信息的方法 | |
| HK40073412A (en) | Video clip editing method, video recommendation method, device, equipment and medium | |
| Yamauchi et al. | Searching emotional scenes in TV programs based on twitter emotion analysis | |
| CN119316648A (zh) | 弹幕信息显示方法、装置、计算机设备和存储介质 | |
| CN120751203A (zh) | 一种点播内容推荐方法、装置、存储介质及电子设备 | |
| CN121585848A (zh) | 视频处理方法、装置、设备及存储介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| RD03 | Notification of appointment of power of attorney |
Free format text: JAPANESE INTERMEDIATE CODE: A7423 Effective date: 20201102 |
|
| RD04 | Notification of resignation of power of attorney |
Free format text: JAPANESE INTERMEDIATE CODE: A7424 Effective date: 20210219 |
|
| A521 | Request for written amendment filed |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20231012 |
|
| A621 | Written request for application examination |
Free format text: JAPANESE INTERMEDIATE CODE: A621 Effective date: 20231012 |
|
| A871 | Explanation of circumstances concerning accelerated examination |
Free format text: JAPANESE INTERMEDIATE CODE: A871 Effective date: 20231012 |
|
| A131 | Notification of reasons for refusal |
Free format text: JAPANESE INTERMEDIATE CODE: A131 Effective date: 20231214 |
|
| A521 | Request for written amendment filed |
Free format text: JAPANESE INTERMEDIATE CODE: A523 Effective date: 20240314 |
|
| TRDD | Decision of grant or rejection written | ||
| A01 | Written decision to grant a patent or to grant a registration (utility model) |
Free format text: JAPANESE INTERMEDIATE CODE: A01 Effective date: 20240501 |
|
| A61 | First payment of annual fees (during grant procedure) |
Free format text: JAPANESE INTERMEDIATE CODE: A61 Effective date: 20240531 |
|
| R150 | Certificate of patent or registration of utility model |
Ref document number: 7498640 Country of ref document: JP Free format text: JAPANESE INTERMEDIATE CODE: R150 |