WO2013107184A1 - 记录会议的方法和会议系统 - Google Patents
记录会议的方法和会议系统 Download PDFInfo
- Publication number
- WO2013107184A1 WO2013107184A1 PCT/CN2012/081220 CN2012081220W WO2013107184A1 WO 2013107184 A1 WO2013107184 A1 WO 2013107184A1 CN 2012081220 W CN2012081220 W CN 2012081220W WO 2013107184 A1 WO2013107184 A1 WO 2013107184A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- conference
- key information
- information
- key
- meeting
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N7/00—Television systems
- H04N7/14—Systems for two-way working
- H04N7/15—Conference systems
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N7/00—Television systems
- H04N7/14—Systems for two-way working
- H04N7/15—Conference systems
- H04N7/155—Conference systems involving storage of or access to video conference sessions
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F16/00—Information retrieval; Database structures therefor; File system structures therefor
- G06F16/70—Information retrieval; Database structures therefor; File system structures therefor of video data
- G06F16/73—Querying
- G06F16/738—Presentation of query results
- G06F16/739—Presentation of query results in form of a video summary, e.g. the video summary being a video sequence, a composite still image or having synthesized frames
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/20—Movements or behaviour, e.g. gesture recognition
-
- G—PHYSICS
- G11—INFORMATION STORAGE
- G11B—INFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
- G11B27/00—Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
- G11B27/02—Editing, e.g. varying the order of information signals recorded on, or reproduced from, record carriers
- G11B27/031—Electronic editing of digitised analogue information signals, e.g. audio or video signals
-
- G—PHYSICS
- G11—INFORMATION STORAGE
- G11B—INFORMATION STORAGE BASED ON RELATIVE MOVEMENT BETWEEN RECORD CARRIER AND TRANSDUCER
- G11B27/00—Editing; Indexing; Addressing; Timing or synchronising; Monitoring; Measuring tape travel
- G11B27/10—Indexing; Addressing; Timing or synchronising; Measuring tape travel
- G11B27/19—Indexing; Addressing; Timing or synchronising; Measuring tape travel by using information detectable on the record carrier
- G11B27/28—Indexing; Addressing; Timing or synchronising; Measuring tape travel by using information detectable on the record carrier by using information signals recorded by the same method as the main recording
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04N—PICTORIAL COMMUNICATION, e.g. TELEVISION
- H04N7/00—Television systems
- H04N7/14—Systems for two-way working
- H04N7/141—Systems for two-way working between two video terminals, e.g. videophone
- H04N7/147—Communication arrangements, e.g. identifying the communication as a video-communication, intermediate storage of the signals
Definitions
- Embodiments of the present invention relate to the field of multimedia communications and, more particularly, to a method of recording a conference and a conferencing system.
- Multimedia communication is a new way of communication that is different from traditional telephone communication.
- Multimedia communication is the addition of video and computer data interaction based on sound interaction.
- An important feature of multimedia communication is that the communicating parties can see each other's active video and environmental video. Since the use of visually conveyed information is more straightforward, video interaction can greatly improve the quality of communication.
- a large number of multimedia communication technologies are used. Participants can quickly integrate into the conference through multimedia communication tools used in the conference. For example, a conference recording function that records the entire process of multimedia communication that occurs in real time for playback can meet the needs of conference memorandum, post-session training, and the like.
- a conference recording function that records the entire process of multimedia communication that occurs in real time for playback can meet the needs of conference memorandum, post-session training, and the like.
- people have been looking for effective ways to save storage resources and save network bandwidth to record the ongoing multi-party conference in real time to meet or Deepen the need for understanding of the meeting, especially the need to quickly grasp the content of the meeting.
- a viable option is to generate a summary of the meetings, and the summary of the meetings provides key information about the meeting.
- a technique for automatic face extraction of a recorded conference timeline has been proposed.
- the face of the speaker is automatically detected, and the face image corresponding to each speaker is stored in the face database; creating a timeline to graphically identify the speaking time of the speaker in the playback of the meeting record;
- the face image is identified to identify each speaker associated with the timeline, instead of identifying each user in the timeline in general.
- using only face information as the key information of the meeting record is not The theme of the meeting record can be fully provided by the summary of the meeting.
- a technique for performance enhancement of a video conference is also proposed, wherein the conference server requests a key frame from the conference participant in response to determining that the conference participant should be the most active participant, the conference server responding to participate from the conference The keyframe is received such that the conference participant becomes the most active participant.
- the method of performance enhancement of the video conference is also incapable of fully grasping the subject matter of the conference record when viewing the record summary.
- the embodiment of the invention provides a method for recording a conference and a conference system, which can make the subject matter of the conference record as a whole when viewing the conference summary.
- a method for recording a conference including: extracting, at each of a plurality of time points on a conference timeline, key information of each site based on a configuration file, where the conference timeline and the conference Time-related, the configuration file is used to define the key information of the conference and the format of the conference summary; the key information of each site is combined into a key index point, and the key index point is used to interact or edit with the conference summary.
- Index point combines multiple key index points corresponding to multiple points in time into a meeting summary.
- a conference system including: an extracting unit, configured to extract key information of each site based on a configuration file, at each of a plurality of time points on the conference timeline, where the conference The time line is associated with the meeting time, and the configuration file is used to define the key information of the meeting and the format of the meeting summary.
- the combining unit is configured to combine the key information of each meeting site into a key index point, where the key index point is used. An index point for interacting or editing with the meeting summary; a combining unit for combining a plurality of key index points corresponding to the plurality of time points into the meeting summary.
- the method for recording a conference and the conference system in the embodiment of the present invention may automatically generate a conference summary based on a customized configuration file that can be modified at any time, and present the conference summary, and then pass the conference summary.
- the key index points get more detailed meeting information.
- FIG. 1 is a flow chart of a method of recording a conference in accordance with an embodiment of the present invention.
- FIG. 2 is a schematic flow chart of a method of recording a conference according to an embodiment of the present invention.
- FIG. 3 is a flow diagram of generating key index points in accordance with an embodiment of the present invention.
- FIG. 4 is a schematic structural diagram of a conference system according to an embodiment of the present invention.
- FIG. 5 is a schematic structural diagram of a conference system according to an embodiment of the present invention.
- FIG. 6 is a schematic structural diagram of a conference system according to an embodiment of the present invention.
- FIG. 7 is a schematic structural diagram of a conference system according to an embodiment of the present invention.
- a conference summary can be generated based on the key index points in a multi-point video conference. In addition, you can interact with the information in the meeting summary at any time.
- a method of recording a conference according to an embodiment of the present invention will be specifically described below with reference to FIGS. 1 through 3.
- the method for recording a conference as shown in FIG. 1 includes the following steps.
- the conference system may extract key information of each site based on the configuration file, where the conference timeline is associated with the conference time, and the configuration file is used at each of the multiple time points on the conference timeline. Define key information for the meeting and the format of the meeting summary.
- the conference system also needs to generate a configuration file, which may be a configuration file generated in a human-computer interaction mode, or may be customized in advance. When the configuration file is generated, the conference system will automatically save the configuration file. In this way, after a configuration file is generated in a site, other sites can retrieve the configuration file in the conference system.
- the configuration file can include various voice, video detection and recognition modules, key information extraction modules, event determination and analysis modules, and the like.
- these modules can detect and identify faces, specify the voice detection and recognition of participants, detect and identify the actions or behaviors of designated people, detect and identify information of participants in multi-point conferences, and specific products. Demonstration, special needs of people with disabilities, amplification of specific speaker voices, partial enlargement of specific product information, etc.
- the configuration file defines the key information for the meeting.
- the key information may include one or more of the following information: face, voice, body movements, key frames, speeches, specific events.
- the so-called specific events are special events during the meeting, such as raising hands, asking questions, dozing, no spirits, bowing, laughing, crying, seatless, etc., and can also include other custom things.
- the configuration file also defines the format of the conference summary to be generated, such as a text file, an audio file, a video file, a Flash file, a PPT file, and the like.
- the configuration file also defines how to generate a meeting summary in the above format.
- the conference system combines the key information of each site into a key index point, and the key index point is used as an index point for interacting or editing with the conference summary.
- the conference system combines the key information of the respective venues into key index points corresponding to all the information contained in the key information.
- the key information defined by the profile includes face and voice
- the face key information and voice key information corresponding to the time point on the conference timeline in each site are extracted, and then the face key information and voice key are selected.
- the information is combined into a key index point.
- a method of organizing and arranging key information to form key index points according to a certain pattern For example, when sorting by character, you can sort according to the order of the characters in the meeting. The important characters are from the middle to the two sides, or from left to right. After the characters are arranged, the characters are simultaneously arranged and the sounds are synchronized to ensure the lip sounds are synchronized. For example, when sorting according to the specific attention events in the obtained key information, a specific event of interest can be obtained, and the event can be made significant; the successor is accompanied by the associated person; and the specific event commentary is used to explain the specific event.
- the conference system combines multiple key index points corresponding to multiple time points into a conference summary. Since there are multiple points in time on the conference timeline, the conference system combines multiple key index points corresponding to each point in time to form a conference summary.
- the format of the meeting summary can be determined based on the configuration file.
- the generated conference summary can be presented in the form of a screen or an icon in the conference display screen of each venue. Participants can view the topic of the previous meeting by watching the summary of the meeting.
- the method for recording a conference in the embodiment of the present invention can automatically generate a conference summary based on a customized configuration file that can be modified at any time, and present the conference summary, and then obtain more detailed conference information through key index points in the conference summary.
- the following describes the process of generating a conference summary of a video conference according to the method for recording a conference according to the embodiment of the present invention.
- the two conference sites in the video conference are taken as an example for description.
- the two venues for video conferencing are called “Conference A” and “Conference Site B.” 201.
- site B can also invoke the configuration file generated above.
- the profile can define the following information: face, gestures (eg gestures, gestures), voices sent, PPT play status, conference scenes, events (emergency, such as sudden someone running, demo product landing , the meeting was suddenly interrupted, most people left the game, most people are playing mobile phones or sleeping, etc.), the introduction of the participants (gender, age, knowledge background, hobbies, etc.).
- the key information of the site A and/or the site B can be detected or extracted in the conference system of the site A. For example, through a face detection module, a face recognition module, a gesture recognition module, a gesture recognition module, a voice recognition module, a PPT determination module, an event detection module, Specific scene modeling detection modules, etc. generate key information.
- 204 generating a key index point of the video conference based on the key information, where the key index point is a set of multiple key information.
- 205 based on the key index points, combine the key index points generated by the multiple time points/segments to generate a conference summary of the video conference.
- Figure 3 illustrates a specific embodiment illustrating how to generate key index points.
- the information of the conference dispute point and the PPT important display point is defined.
- the microphone array is detected and identified. If the detected voice is not intense, it indicates that the topic of the discussion is not an issue of the conference, so that the voice information is discarded as the key information of the conference; otherwise, the voice information is recorded.
- 303 detecting whether there is a large limb movement between the speakers by the depth acquiring device, and if not, discarding the motion information as the key information of the conference, otherwise, recording the motion information.
- the key information of the conference dispute point is generated based on the recorded voice information and action information.
- 304 generates a key index point for the video conference based on the generated key information for the PPT important presentation point and key information for the conference issue.
- these key information can be organized and arranged according to a certain pattern to generate key index points. For example, you can generate key index points according to the order of the characters. First, sort the characters according to the order of the participants. The important characters are from the middle to the two sides, or from left to right. Then, according to the characters, the sounds are arranged in synchronization with each character. Lip sync.
- the specific attention events in the key information obtained may be sorted, firstly, the specific attention event is obtained, the event is marked, and then associated with the associated person, and finally with a specific event commentary to explain the specific event.
- a conference summary of the video conference is generated. For example, multiple key index points are concatenated according to a certain motion pattern. That is to solve how to switch between two consecutive key index points. Can follow similar
- the situation of PPT playback adds a custom animation that links two consecutive key index points. Alternatively, it may be defined according to the character association mode of the conference scene.
- the steps are as follows: Firstly, the character information in the two key index points is obtained; then the character is judged and associated, if the character of the two consecutive key index points appears, then The character is defined as the main body of the action, so that the coherent action of the character is defined as the play action of two consecutive key index points here; if there is no character, the voice of the speaker is used as the key action, and the corresponding voice is defined according to the priority of the voice.
- the priority of the key index point action if there is no character, no sound, the PPT switching mode in the video conference is used as the playback action of two consecutive key index points; if there is no character, no sound, no PPT, the user uses the default mode Define the playback action of the key index points.
- participants can interact or edit with the conference summary based on key information in the conference summary, such as extending the information associated with the key information.
- the participant clicks on the face in the meeting summary the person's brief information is displayed in real time, or a further reference index is provided.
- the participant clicks on the key information the original video conference video related to the key information needs to be obtained, and then the key information is used to jump to the video segment of the original video conference.
- the user can obtain the PPT speech or automatically send the PPT speech to a pre-defined user's e-mail box.
- the voice information segment or the video segment corresponding to the limb motion may be associated.
- the participant previews a lot of face key information
- the basic information of the person can be obtained, and the action information, key voice information, and even the person's speech information in the multi-point meeting can be further obtained.
- the interaction of the conference summary according to the key information in the conference summary includes: if the key information includes a face, obtaining information of the participant corresponding to the face, and understanding the participation The meeting of the person speaks or communicates with the participant; if the key information includes the speech, the file information of the speech is obtained; if the key information includes a voice or a physical action, the information of the object targeted by the voice or the body motion is obtained. ; If the key information includes product information, obtain additional information about the product.
- Participant A in the first meeting wants to communicate with Participant B in the second meeting
- the face of Participant B can be selected from the meeting summary, if the face image of Participant B is The face information database of the secondary meeting can be matched, and the basic information of the participant B is imported in the meeting summary, and the participant A has further information about the speaking information of the participant B in the meeting.
- the basic information provided by Participant B includes Instant Messenger (IM) information
- the participant can chat with Participant B through the IM.
- the basic information provided by Participant B includes email information
- Participant A can contact Participant B via this email. Even if the basic information of the participant B is included to include the body motion information, the participant A can quickly understand the speech style of the participant B.
- IM Instant Messenger
- real-time interactions can be made between the audience and the speaker.
- the audience is interested in a part of the PPT speech, pointing to the part of the PPT speech, the camera or sensor gets the corresponding position recognized by the PPT speech and draws a circle in the corresponding position of the PPT speech through the conference system.
- the speaker knows that the audience is interested in a part of the PPT speech and explains the part.
- the speaker will explain the important parts of some important topics or PPT speeches over and over again, perhaps with some habitual actions, which can be distinguished by the speech recognition module, gesture gesture recognition module, and feedback to the audience.
- the conference system prompts to display the character. information.
- the conference system prompts whether further interaction is required? For example, if it is an online real-time conference summary, the chat prompt is sent online in real time, and the chat program is started, including text chat, voice private chat, video private chat, and the like; If it is an offline meeting summary, the user is prompted whether or not to send an email.
- the meeting information is included in the key information
- the user is provided with further information by means of a link.
- the conference system prompts how to obtain the file and how to apply for the file permission.
- the user information is further provided to the user in a link manner, for example, providing the source of the product, the manufacturer, the stereoscopic three-dimensional display model, and the like.
- participants can edit the meeting summary based on key information in the meeting summary.
- the editing of a character includes associating more information, whether to start a chat program, whether to start an email (email) program, whether to automatically send an email, whether to be a specific person or leader, whether to display a brief introduction of a character in time, etc.;
- the editor also includes instructions for adding characters, such as physical movement instructions, background knowledge descriptions, and statement status.
- the editing of a character also includes analyzing the behavioral habits of the main speaker by a computer visual algorithm based on a physical motion of the main speaker, and then using these action habits as a characteristic of the character.
- the editing of the characters also includes the classification of the participants.
- the editing of the meeting summary can be edited based on the automated learning result of one of the main speaker's body movements. For example, after the entire video conference is finished, the computer vision algorithm analyzes the main speaker's behavioral habits. Add these habits as a characteristic of the character to the basic information of the character.
- the conference system 40 includes an extracting unit 41, a combining unit 42, and a combining unit 43.
- the extracting unit 41 is configured to extract key information of each site based on the configuration file, where the meeting timeline is associated with the meeting time, at each of the plurality of time points on the meeting timeline, where the configuration is
- the file is used to define the key information of the meeting and the format of the meeting summary.
- the combining unit 42 is configured to combine key information of the respective sites into key index points, where the key index points are used as index points for interacting or editing with the meeting summary, that is, the key information of the respective sites is combined to correspond to the A key index point for all the information contained in the key information.
- the combining unit 43 is configured to combine a plurality of key index points corresponding to a plurality of time points into a conference summary.
- profiles include voice, video detection and recognition modules, key information extraction modules, and event determination and analysis modules.
- Key information includes one or more of the following information: Faces, Voices, Limbs, Keyframes, Presentations, Custom Events.
- the format of the meeting summary is a text file, an audio file, a video file, a flash file, or a PPT file.
- the conference system 50 further includes a generating and holding unit 44 for generating and saving a configuration file before extracting key information of each site based on the configuration file. , so that the configuration file is retrieved by other sites.
- the conferencing system 60 further includes a display unit 45 for presenting the meeting summary in the form of a picture or an icon.
- the conferencing system 70 further includes an interaction and editing unit 46 for selecting keys according to the meeting summary
- the information interacts or is edited with the summary of the meeting.
- the interaction and editing unit 46 is specifically configured to: if the key information includes a face, obtain information of the participant corresponding to the face, and understand the conference speech of the participant or communicate with the participant; The information includes a speech, obtaining file information of the speech; if the key information includes a voice or a body motion, acquiring information about the object targeted by the voice or the body motion.
- the interaction and editing unit 46 is specifically configured to: according to the key information, including the face to edit the conference summary, including adding a physical motion to the character Ming, background description, statement of speech status, or classification of participants; or editing of the summary of the meeting based on the action including the action, including editing based on the automated learning result of a physical action of the primary speaker; or
- the key information is associated and arranged, and the switching action is defined for the key index points generated by the key information.
- the disclosed systems, devices, and methods may be implemented in other ways.
- the device embodiments described above are merely illustrative.
- the division of the unit is only a logical function division.
- there may be another division manner for example, multiple units or components may be combined or Can be integrated into another system, or some features can be ignored, or not executed.
- the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interface, device or unit, and may be electrical, mechanical or otherwise.
- the components displayed for the unit may or may not be physical units, ie may be located in one place, or may be distributed over multiple network units. Some or all of the units may be selected according to actual needs to achieve the objectives of the solution of the embodiment.
- each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
- the functions may be stored in a computer readable storage medium if implemented in the form of a software functional unit and sold or used as a standalone product.
- the technical solution of the present invention which is essential or contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product, which is stored in a storage medium, including
- the instructions are used to cause a computer device (which may be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention.
- the foregoing storage medium includes: a U disk, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and the like, which can store program codes. .
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Theoretical Computer Science (AREA)
- Signal Processing (AREA)
- General Physics & Mathematics (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Databases & Information Systems (AREA)
- Data Mining & Analysis (AREA)
- Computational Linguistics (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Human Computer Interaction (AREA)
- Social Psychology (AREA)
- Psychiatry (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- Two-Way Televisions, Distribution Of Moving Picture Or The Like (AREA)
- Telephonic Communication Services (AREA)
Abstract
本发明实施例提供了记录会议的方法和会议系统。其中,该记录会议的方法包括:在会议时间线上的多个时间点中的每个时间点上,基于配置文件提取各个会场的关键信息,其中所述会议时间线与会议时间相关联,所述配置文件用于定义会议的关键信息以及会议摘要的格式;将所述各个会场的关键信息组合成关键索引点,所述关键索引点用作与会议摘要进行互动或编辑的索引点;将对应于多个时间点的多个关键索引点结合为会议摘要。从而,基于自定义可随时修改的配置文件自动生成会议摘要、并呈现所述会议摘要,进而通过会议摘要中的关键索引点获得更详尽的会议信息。
Description
记录会议的方法和会议系统 本申请要求于 2012 年 1 月 16 日提交中国专利局、 申请号为 201210012814.8,发明名称为"记录会议的方法和会议系统,,的中国专利申请的 优先权, 在先申请文件的内容通过引用结合在本申请中。
技术领域
本发明实施例涉及多媒体通信领域, 并且更具体地, 涉及记录会议的方 法和会议系统。
背景技术
多媒体通信是一种不同于传统的电话通信的新的通信方式。 多媒体通信 是在声音交互的基础上又增加了视频与计算机数据的交互。 多媒体通信的一 个重要特征是通信双方可以相互看到对方的活动视像以及环境视像。 由于使 用视觉方式传递信息更为直接, 因此视频的交互可以大大提升交流的质量。
在目前视频会议中, 大量釆用多媒体通信技术。 参会人员通过会议釆用 的多媒体通信工具可以快速融入会议中。 例如, 将一个实时发生的多媒体通 信的整个过程均录制下来以便回放的会议录播功能, 可以满足会议备忘、 会 后培训等所需。 考虑到现有的会议终端系统的处理能力、 存储能力、 网络带 宽等问题, 人们一直在寻找有效的、 节省存储资源以及节约网络带宽的方法 来对正在进行的多方会议进行实时录制, 以便满足或加深对会议的理解等需 求, 特别是快速把握会议内容的需求。 为此, 一种可行的方案是生成一些会 议摘要, 并由这些会议摘要提供会议的关键信息。
例如, 已经提出了用于记录的会议时间线的自动脸部提取的技术。 其中, 发言者的脸部被自动检测, 对应于每个发言者的脸部图像被存储在脸部数据 库中; 创建一条时间线在会议记录的回放中图形化地标识发言者的发言时间; 示出脸部图像以标识与时间线相关联的每个发言者, 取代了一般地在时间线 中识别每个用户。 但是, 仅仅利用人脸信息作为会议记录的关键信息, 并不
能由会议摘要完备地提供整体把握会议记录的主题内容。
例如, 还提出了视频会议的性能增强的技术, 其中会议服务器响应于确 定会议参与者应成为最活跃的参与者向所述会议参与者请求关键帧, 所述会 议服务器响应于从所述会议参与者接收到所述关键帧使得所述会议参与者成 为最活跃的参与者。 同样的, 该视频会议的性能增强的方法同样无法完备地 当观看该记录摘要时能够整体把握会议记录的主题内容。
另外, 已经提出了在线的会议记录系统, 可以编辑或打开会议记录, 并 且参会者可以在在线会议时听到和看到该会议记录。 该在线的会议记录系统 仍然无法完备地当观看该记录摘要时能够整体把握会议记录的主题内容。 发明内容
本发明实施例提供一种记录会议的方法及会议系统, 能够使得当观看会 议摘要时整体把握会议记录的主题内容。
一方面, 提供了一种记录会议的方法, 包括: 在会议时间线上的多个时 间点中的每个时间点上, 基于配置文件提取各个会场的关键信息, 其中所述 会议时间线与会议时间相关联, 所述配置文件用于定义会议的关键信息以及 会议摘要的格式; 将所述各个会场的关键信息组合成关键索引点, 所述关键 索引点用作与会议摘要进行互动或编辑的索引点; 将对应于多个时间点的多 个关键索引点结合为会议摘要。
另一方面, 提供了一种会议系统, 包括: 提取单元, 用于在会议时间线 上的多个时间点中的每个时间点上, 基于配置文件提取各个会场的关键信息 , 其中所述会议时间线与会议时间相关联, 所述配置文件用于定义会议的关键 信息以及会议摘要的格式; 组合单元, 用于将所述各个会场的关键信息组合 成关键索引点, 所述关键索引点用作与会议摘要进行互动或编辑的索引点; 结合单元, 用于将对应于多个时间点的多个关键索引点结合为会议摘要。
本发明实施例的记录会议的方法及会议系统可以基于自定义可随时修改 的配置文件自动生成会议摘要、 并呈现所述会议摘要, 进而通过会议摘要中
的关键索引点获得更详尽的会议信息。
附图说明
为了更清楚地说明本发明实施例的技术方案, 下面将对实施例或现有技 术描述中所需要使用的附图作简单地介绍, 显而易见地, 下面描述中的附图 仅仅是本发明的一些实施例, 对于本领域普通技术人员来讲, 在不付出创造 性劳动的前提下, 还可以根据这些附图获得其他的附图。
图 1是根据本发明实施例的记录会议的方法的流程图。
图 2是根据本发明实施例的记录会议的方法的示意性流程图。
图 3是根据本发明实施例的生成关键索引点的流程图。
图 4是根据本发明实施例的会议系统的结构示意图。
图 5是根据本发明实施例的会议系统的结构示意图。
图 6是根据本发明实施例的会议系统的结构示意图。
图 7是根据本发明实施例的会议系统的结构示意图。
具体实施方式
下面将结合本发明实施例中的附图, 对本发明实施例中的技术方案进行 清楚、 完整地描述, 显然, 所描述的实施例是本发明一部分实施例, 而不是 全部的实施例。 基于本发明中的实施例, 本领域普通技术人员在没有作出创 造性劳动前提下所获得的所有其他实施例, 都属于本发明保护的范围。
为了方便观看者从整体上把握会议情况, 可以在一个多点视频会议中基 于关键索引点生成会议摘要。 此外, 可以实现与会议摘要中的信息随时进行 互动。
下面将结合图 1至图 3具体描述根据本发明实施例的记录会议的方法。 如图 1所示的记录会议的方法, 包括以下步骤。
11 , 在会议时间线上的多个时间点中的每个时间点上, 会议系统可以基 于配置文件提取各个会场的关键信息, 其中所述会议时间线与会议时间相关 联, 所述配置文件用于定义会议的关键信息以及会议摘要的格式。
可选地, 在基于配置文件提取各个会场的关键信息之前, 会议系统 还需要生成配置文件, 可以是以人机交互模式生成配置文件, 也可以预先自 定义好配置文件。 当配置文件生成后, 会议系统将自动保存该配置文件。 这 样, 在一个会场中生成配置文件后, 其他会场可以在会议系统中调取该配置 文件。
另外, 配置文件可包括各种语音、 视频检测与识别模块, 关键信息提取 模块, 事件判定与分析模块等。 例如: 这些模块可以实现人脸的检测与识别、 指定参会者语音的检测与识别、 指定人的动作或行为的检测与识别、 多点会 议参会者共同关注信息的检测与识别、 特定产品的演示、 残障人员的特别需 求、 特定发言者声音的放大、 特定产品信息的局部放大表示等。
如上所述, 配置文件定义了会议的关键信息。 其中, 关键信息可以包括 以下信息中的一个或多个: 人脸、 语音、 肢体动作、 关键帧、 演讲稿、 特定 事件。 其中所谓特定事件是开会过程中的一些特殊事件, 例如包括举手问问 题、 打瞌睡、 无精神、 低头、 大笑、 哭丧、 座位无人等场景, 也可以包括其 他的自定义的事情。
可选地, 配置文件还定义了将生成的会议摘要的格式, 例如文本文件、 音频文件、 视频文件、 Flash文件、 PPT文件等。
可选地, 配置文件还定义如何生成上述格式的会议摘要。
12, 会议系统将所述各个会场的关键信息组合成关键索引点, 所述关键 索引点用作与会议摘要进行互动或编辑的索引点。
从而, 会议系统将所述各个会场的关键信息组合为对应于所述关键信息 所含的全部信息的关键索引点。
例如, 配置文件定义的关键信息中包括人脸、 语音, 那么提取各个会场 中对应于会议时间线上的一个时间点处的人脸关键信息以及语音关键信息, 然后将人脸关键信息和语音关键信息组合成一个关键索引点。
按照一定的模式把关键信息组织、 排列起来形成关键索引点的方法中,
例如, 按照人物排序时, 可以按照参会的人物顺序排序, 重要的人物从中间 往两边, 或者从左至右; 按照人物排列后, 同时给各个人物同步排列声音, 保证唇音同步。 例如, 按照所获得的关键信息中特定关注事件排序时, 可以 获得特定关注事件, 把该事件显著化; 后继配以相关联人物陪衬; 配以特定 事件解说词, 解释特定事件。
13 , 会议系统将对应于多个时间点的多个关键索引点结合为会议摘要。 由于在会议时间线上有多个时间点, 会议系统将多个对应于每个时间点 的关键索引点结合在一起就形成了会议摘要。 该会议摘要的格式可以依据配 置文件确定。
生成的会议摘要可以以画面形式或者图标形式呈现在每个会场的会议显 示屏中。 参会者可以通过观看该会议摘要来整体把握之前会议的主题内容。
由此可见, 本发明实施例的记录会议的方法可以基于自定义可随时修改 的配置文件自动生成会议摘要、 并呈现该会议摘要, 进而通过会议摘要中的 关键索引点获得更详尽的会议信息。
下面参见图 2具体描述根据本发明实施例的记录会议的方法生成视频会 议的会议摘要的过程, 以视频会议中的 2个会场为例说明。
例如, 这 2个进行视频会议的会场称为 "会场 A" 和 "会场 B"。 201 , 在 会场 A中通过手动配置或自动配置生成配置文件, 并在 202中将生成的配置 文件进行保存。 这样, 会场 B也可以调用上述生成的配置文件。 其中, 配置 文件可以定义以下信息: 人脸、 发出的肢体动作 (如: 手势、 姿势)、 发出的 语音、 PPT 播放状态、 会议场景、 事件 (突发事件, 如突然有人快跑、 演示 产品落地、 会议突然中断、 大部分人离场、 大部分人在玩手机或瞌睡等)、 参 会人员的简介(性别、 年龄、 知识背景、 爱好等)等。
203 , 在配置文件的基础上, 可以在会场 A的会议系统中检测或提取会场 A和 /或会场 B的关键信息。 例如, 通过人脸检测模块、 人脸识别模块、 手势 识别模块、 姿势识别模块、 语音识别模块、 PPT 判定模块、 事件检测模块、
特定场景建模检测模块等生成关键信息。
之后, 204, 基于关键信息生成视频会议的关键索引点, 该关键索引点是 多个关键信息的集合。 最后, 205 , 在关键索引点的基础上, 将多个时间点 / 段生成的关键索引点结合在一起就生成了该视频会议的会议摘要。
图 3示出了一个具体实施例说明如何生成关键索引点。 例如在会场 A的 会议系统的配置文件中定义了会议争论点和 PPT重要展示点的信息。 首先,
301 , 判断会场 A中播放的 PPT是否长时间不变。 如果不是, 则丟弃该信息 不作为关键信息; 否则, 如果播放的 PPT长时间不变, 表明该 PPT的内容可 能是会议中重要的关注点, 由此生成针对 PPT重要展示点的关键信息。 然后,
302, 再对麦克风阵列进行检测和识别, 如果检测到的语音不激烈, 表示该讨 论的主题不是会议的争论点, 从而丟弃该语音信息不作为会议的关键信息; 否则, 记录该语音信息。 同时, 303 , 通过深度获取装置检测发言者之间是否 有大的肢体动作, 如果没有, 也丟弃该动作信息不作为会议的关键信息, 否 则, 记录该动作信息。 基于记录的语音信息和动作信息生成会议争论点的关 键信息。 最后, 304, 基于生成的针对 PPT重要展示点的关键信息以及针对会 议争论点的关键信息生成该视频会议的关键索引点。
也就是说, 当获得会议的关键信息之后, 可以按照一定的模式把这些关 键信息组织、 排列起来以生成关键索引点。 例如, 可以按照人物排序来生成 关键索引点, 首先按照参会的人物顺序排序, 重要的人物从中间往两边, 或 者从左至右, 然后按照人物排列后, 同时给各个人物同步排列声音, 保证唇 音同步。 或者, 可以按照所获得的关键信息中特定关注事件排序, 首先获得 特定关注事件, 把该事件显著化, 再后继配以相关联人物陪衬, 最后配以特 定事件解说词, 解释特定事件。
最后, 将各个时间点的关键索引点结合在一起并与时间线相关联, 就生 成了该视频会议的会议摘要。 例如, 按照一定的运动模式把多个关键索引点 串联起来。 也就是解决连续的两个关键索引点之间如何切换。 可以按照类似
PPT 播放的形势加入自定义的动画, 把连续的两个关键索引点关联起来。 或 者, 可以按照会议场景的人物关联模式定义, 步骤如下: 首先获得前后两个 关键索引点中的人物信息; 然后判断人物并进行关联, 若是连续两个关键索 引点中的那个人物都出现, 则把该人物作为动作定义的主体, 以便给人物连 贯的动作定义为这里的连续两个关键索引点的播放动作; 若是无人物, 则以 发言者的声音作为关键动作, 按照声音的轻重緩急定义对应的关键索引点动 作的轻重緩急; 若无人物、 无声音, 以视频会议中 PPT切换的方式作为连续 两个关键索引点的播放动作; 若是无人物、 无声音、 无 PPT, 则用户使用默 认方式定义关键索引点的播放动作。
为了获得更佳的会议摘要, 参会者可以根据所述会议摘要中的关键 信息与所述会议摘要进行互动或编辑, 例如延伸加入与关键信息相关联的信 息。
可选地, 当参会者点击会议摘要中的人脸时, 实时地显示出该人的简要 信息, 或提供更进一步的参考索引。
可选地, 当参会者点击关键信息, 需要获得该关键信息相关的原始视频 会议视频, 再通过该关键信息跳转到原始视频会议的视频片段中。
可选地, 当参会者预览到某个会场讲解的 PPT演讲稿时, 用户可以获取 到该 PPT演讲稿, 或者将该 PPT演讲稿自动发送到预先定义的用户的电子邮 箱中。
可选地, 当参会者预览到肢体动作丰富的关键信息时, 可关联出该肢体 动作对应的语音信息片段或视频片段。
或者, 当参会者预览到非常多的人脸关键信息时, 可获取到所有人的相 关基本信息, 更进一步可获得某个人在此次多点会议中的发言动作信息、 关 键语音信息、 乃至于总结出来的发言特色信息等。 甚至, 可以呼叫对应的人。
由此可见, 根据会议摘要中的关键信息对所述会议摘要进行互动包括: 如果关键信息包括人脸, 获取该人脸对应的参会者的信息, 并了解所述参会
者的会议发言或与所述参会者进行通信; 如果关键信息包括演讲稿, 获取该 演讲稿的文件信息; 如果关键信息包括语音或肢体动作, 获取所述语音或肢 体动作针对的对象的信息; 如果关键信息包括产品信息, 获取该产品的其他 信息。
以最简单的 2个会场为例说明与会议摘要的互动方式。
第一会场中的参会者 A如果希望与第二会场中的参会者 B交流, 那么就 可以从会议摘要中选中参会者 B的脸部, 如果参会者 B的脸部图像与本次会 议的人脸信息库中能够匹配, 会议摘要中就导入该参会者 B的基本信息, 参 会者 A还有进一步了解参会者 B在本次会议中的有关发言信息。 假设提供的 参会者 B的基本信息中包括即时通讯( Instant Messenger, IM )信息, 则参会 者 Α可以通过该 IM与参会者 B聊天。 假设提供的参会者 B的基本信息中包 括电子邮件( email )信息,则参会者 A可以通过该电子邮件与参会者 B联系。 甚至, 假设提供的参会者 B的基本信息中包括肢体动作信息, 则参会者 A可 以快速了解参会者 B的演讲风格。
此外, 听众与演讲者之间也可以进行实时互动。 例如, 听众对 PPT演讲 稿的某个部分感兴趣, 手指着 PPT演讲稿的该部分, 摄像机或传感器获取识 别到 PPT演讲稿的对应位置并且通过会议系统在 PPT演讲稿的对应位置画上 圈圈, 演讲者知道了听众对 PPT演讲稿的某个部分感兴趣, 就对该部分进行 详解。 或者, 演讲者对某些重要的议题或 PPT演讲稿的重要部分会大声反复 解释, 或许带些习惯性的动作, 这些可以通过语音识别模块、 手势姿势识别 模块区分出来, 并反馈到听众所在会场的会议系统的显示屏或所在会议系统。
此外, 若是关键信息中有人物, 则把人物相关简介信息加入进去, 可以 以链接方式提供人物的相关信息, 即当会议系统的指示装置移到人物脸部时, 会议系统提示要显示该人物哪些信息。 在显示了人物信息之后, 会议系统提 示是否需要进行进一步交互? 比如, 若是在线实时会议摘要, 则在线实时发 出聊天提示, 聊天程序启动, 包括文本聊天, 语音私聊、 视频私聊等方式;
若是离线会议摘要, 则提示用户是否打算发送 email。
可选地, 若是关键信息中包括会议文件, 以链接方式提供用户更进一步 的信息。 例如用户想获取该文件时, 则会议系统提示如何获取该文件以及如 何申请该文件权限等信息。
可选地, 若是关键信息中包括产品信息, 以链接方式提供给用户更进一 步关于该产品的产品信息, 例如提供该产品的来源、 厂家、 立体三维展示模 型等。
此外, 参会者还可以根据会议摘要中的关键信息对会议摘要进行编辑。 例如, 对人物的编辑包括关联更多信息、 是否启动聊天程序、 是否启动电子 邮件 (email )程序、 是否自动发送邮件、 是否作为特定的人物或领导、 是否 及时显示人物的简要介绍等; 对人物的编辑还包括对人物添加说明, 如肢体 动作说明、 背景知识说明、 发言状况说明等。 例如, 对人物的编辑还包括基 于主要演讲者的一个肢体动作由计算机视觉算法分析出该主要演讲者的行为 动作习惯, 再把这些动作习惯作为人物的一个特性。 对人物的编辑还包括对 参会人物的分类。
根据关键信息中包括人脸对会议摘要的编辑可以关联更多信息, 例如对 人物添加说明, 如肢体动作说明、 背景知识说明、 发言状况说明等, 或者对 参会人员进行分类
根据关键信息中包括动作对会议摘要的编辑可以基于主要演讲者的一个 肢体动作的自动化学习结果进行编辑操作, 例如, 在整个视频会议开完后, 计算机视觉算法分析出主要演讲者的行为动作习惯, 把这些习惯作为人物的 一个特性加入该人物的基本信息中。
或者是对这些关键信息的关联与排列, 以及对这些关键信息生成的关键 索引点定义切换的动作, 如是緩慢变化还是快速变化等, 时间轴上的连贯性 等。
下面结合图 4至图 7说明根据本发明实施例的会议系统的结构示意图。
如图 4所示,会议系统 40包括提取单元 41、组合单元 42和结合单元 43。 其中, 提取单元 41用于在会议时间线上的多个时间点中的每个时间点上, 基 于配置文件提取各个会场的关键信息, 其中所述会议时间线与会议时间相关 联, 所述配置文件用于定义会议的关键信息以及会议摘要的格式。 组合单元 42用于将所述各个会场的关键信息组合成关键索引点, 所述关键索引点用作 与会议摘要进行互动或编辑的索引点, 即将所述各个会场的关键信息组合为 对应于所述关键信息所含的全部信息的关键索引点。 结合单元 43用于将对应 于多个时间点的多个关键索引点结合为会议摘要。
一般而言, 配置文件包括语音、 视频检测与识别模块、 关键信息提取模 块、 事件判定与分析模块。 关键信息包括以下信息中的一个或多个: 人脸、 语音、 肢体动作、 关键帧、 演讲稿、 自定义事件。 会议摘要的格式为文本文 件、 音频文件、 视频文件、 flash文件或 PPT文件。
在图 5中, 除了提取单元 41、 组合单元 42和结合单元 43外, 会议系统 50还包括生成与保存单元 44 , 其用于在基于配置文件提取各个会场的关键信 息之前, 生成并保存配置文件, 以便所述配置文件被其他会场调取。
在图 6中, 除了提取单元 41、 组合单元 42、 结合单元 43和生成与保存 单元 44外, 会议系统 60还包括显示单元 45 , 用于以画面或图标形式呈现所 述会议摘要。
在图 7中, 除了提取单元 41、 组合单元 42、 结合单元 43、 生成与保存单 元 44和显示单元 45外,会议系统 70还包括互动与编辑单元 46 , 用于根据所 述会议摘要中的关键信息与所述会议摘要进行互动或编辑。 该互动与编辑单 元 46具体用于如果关键信息包括人脸, 获取该人脸对应的参会者的信息, 并 了解所述参会者的会议发言或与所述参会者进行通信; 如果关键信息包括演 讲稿, 获取该演讲稿的文件信息; 如果关键信息包括语音或肢体动作, 获取 所述语音或肢体动作针对的对象的相关信息。 或者该互动与编辑单元 46具体 用于根据关键信息中包括人脸对会议摘要的编辑包括对人物添加肢体动作说
明、 背景知识说明、 发言状况说明, 或者对参会人员进行分类; 或者根据关 键信息中包括动作对会议摘要的编辑包括基于主要演讲者的一个肢体动作的 自动化学习结果进行编辑操作; 或者对所述关键信息进行关联与排列, 再对 这些关键信息生成的关键索引点定义切换动作。
本领域普通技术人员可以意识到, 结合本文中所公开的实施例描述的各 示例的单元及算法步骤, 能够以电子硬件、 或者计算机软件和电子硬件的结 合来实现。 这些功能究竟以硬件还是软件方式来执行, 取决于技术方案的特 定应用和设计约束条件。 专业技术人员可以对每个特定的应用来使用不同方 法来实现所描述的功能, 但是这种实现不应认为超出本发明的范围。
所属领域的技术人员可以清楚地了解到, 为描述的方便和简洁, 上述描 述的系统、 装置和单元的具体工作过程, 可以参考前述方法实施例中的对应 过程, 在此不再赘述。
在本申请所提供的几个实施例中, 应该理解到, 所揭露的系统、 装置和 方法, 可以通过其它的方式实现。 例如, 以上所描述的装置实施例仅仅是示 意性的, 例如, 所述单元的划分, 仅仅为一种逻辑功能划分, 实际实现时可 以有另外的划分方式, 例如多个单元或组件可以结合或者可以集成到另一个 系统, 或一些特征可以忽略, 或不执行。 另一点, 所显示或讨论的相互之间 的耦合或直接耦合或通信连接可以是通过一些接口, 装置或单元的间接耦合 或通信连接, 可以是电性, 机械或其它的形式。 为单元显示的部件可以是或者也可以不是物理单元, 即可以位于一个地方, 或者也可以分布到多个网络单元上。 可以根据实际的需要选择其中的部分或 者全部单元来实现本实施例方案的目的。
另外, 在本发明各个实施例中的各功能单元可以集成在一个处理单元中, 也可以是各个单元单独物理存在, 也可以两个或两个以上单元集成在一个单 元中。
所述功能如果以软件功能单元的形式实现并作为独立的产品销售或使用 时, 可以存储在一个计算机可读取存储介质中。 基于这样的理解, 本发明的 技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的部分可 以以软件产品的形式体现出来, 该计算机软件产品存储在一个存储介质中, 包括若干指令用以使得一台计算机设备(可以是个人计算机, 服务器, 或者 网络设备等)执行本发明各个实施例所述方法的全部或部分步骤。 而前述的 存储介质包括: U盘、 移动硬盘、 只读存储器(ROM, Read-Only Memory ), 随机存取存储器(RAM, Random Access Memory ),磁碟或者光盘等各种可以 存储程序代码的介质。
以上所述, 仅为本发明的具体实施方式, 但本发明的保护范围并不局限 于此, 任何熟悉本技术领域的技术人员在本发明揭露的技术范围内, 可轻易 想到变化或替换, 都应涵盖在本发明的保护范围之内。 因此, 本发明的保护 范围应所述以权利要求的保护范围为准。
Claims
1、 一种记录会议的方法, 其特征在于, 包括:
在会议时间线上的多个时间点中的每个时间点上, 基于配置文件提取各个 会场的关键信息, 其中所述会议时间线与会议时间相关联, 所述配置文件用于 定义会议的关键信息以及会议摘要的格式;
将所述各个会场的关键信息组合成关键索引点, 所述关键索引点用作与会 议摘要进行互动或编辑的索引点; 将对应于多个时间点的多个关键索引点结合为会议摘要。
2、 根据权利要求 1所述的方法, 其特征在于, 在基于配置文件提取各个会 场的关键信息之前, 还包括生成并保存配置文件, 以便所述配置文件被其他会 场调取。
3、根据权利要求 1或 2所述的方法, 其特征在于, 所述配置文件包括语音、 视频检测与识别模块, 关键信息提取模块, 或者事件判定与分析模块。
4、 根据权利要求 1或 2所述的方法, 其特征在于, 所述关键信息包括以下 信息中的一个或多个: 人脸、 语音、 肢体动作、 关键帧、 演讲稿、 自定义事件。
5、 根据权利要求 1所述的方法, 其特征在于, 所述将所述各个会场的关键 信息组合成关键索引点包括:
将所述各个会场的关键信息组合为对应于所述关键信息所含的全部信息的 关键索引点。
6、 根据权利要求 1所述的方法, 其特征在于, 所述会议摘要的格式为文本 文件、 音频文件、 视频文件、 f lash文件或 PPT文件。
7、 根据权利要求 1所述的方法, 其特征在于, 所述会议摘要以画面或图标 形式呈现。
8、 根据权利要求 1所述的方法, 其特征在于, 还包括: 根据所述会议摘要 中的关键信息与所述会议摘要进行互动或编辑。
9、 根据权利要求 8所述的方法, 其特征在于, 所述根据所述会议摘要中的 关键信息对所述会议摘要进行互动包括:
当关键信息包括人脸, 获取该人脸对应的参会者的信息, 并了解所述参会 者的会议发言或与所述参会者进行通信;
当关键信息包括演讲稿, 获取所述演讲稿的文件信息;
当关键信息包括语音或肢体动作, 获取所述语音或肢体动作针对的对象的 相关信息。
10、 根据权利要求 8 所述的方法, 其特征在于, 所述根据所述会议摘要中 的关键信息对所述会议摘要进行编辑包括:
根据关键信息中包括人脸对会议摘要的编辑包括对人物添加肢体动作说 明、 背景知识说明、 发言状况说明, 或者对参会人员进行分类; 或者
根据关键信息中包括动作对会议摘要的编辑包括基于主要演讲者的一个肢 体动作的自动化学习结果进行编辑操作; 或者
对所述关键信息进行关联与排列, 再对这些关键信息生成的关键索引点定 义切换动作。
11、 一种会议系统, 其特征在于, 包括: 提取单元, 用于在会议时间线上的多个时间点中的每个时间点上, 基于配 置文件提取各个会场的关键信息, 其中所述会议时间线与会议时间相关联, 所 述配置文件用于定义会议的关键信息以及会议摘要的格式; 组合单元, 用于将所述各个会场的关键信息组合成关键索引点, 所述关键 索引点用作与会议摘要进行互动或编辑的索引点;
结合单元, 用于将对应于多个时间点的多个关键索引点结合为会议摘要。
12、 根据权利要求 11所述的会议系统, 其特征在于, 还包括生成与保存单 元, 用于在基于配置文件提取各个会场的关键信息之前, 生成并保存配置文件, 以便所述配置文件被其他会场调取。
13、 根据权利要求 11或 12所述的会议系统, 其特征在于, 所述配置文件 包括语音、 视频检测与识别模块、 关键信息提取模块、 事件判定与分析模块。
14、 根据权利要求 11或 12所述的会议系统, 其特征在于, 所述关键信息 包括以下信息中的一个或多个: 人脸、 语音、 肢体动作、 关键帧、 演讲稿、 自 定义事件。
15、 根据权利要求 11所述的会议系统, 其特征在于, 所述组合单元进一步 用于:
将所述各个会场的关键信息组合为对应于所述关键信息所含的全部信息的 关键索引点。
16、 根据权利要求 11所述的会议系统, 其特征在于, 所述会议摘要的格式 为文本文件、 音频文件、 视频文件、 f lash文件或 PPT文件。
17、 根据权利要求 11所述的会议系统, 其特征在于, 还包括显示单元, 用 于以画面或图标形式呈现所述会议摘要。
18、 根据权利要求 11所述的会议系统, 其特征在于, 还包括互动与编辑单 元, 用于根据所述会议摘要中的关键信息与所述会议摘要进行互动或编辑。
19、 根据权利要求 18所述的会议系统, 其特征在于, 所述互动与编辑单元 具体用于:
如果关键信息包括人脸, 获取该人脸对应的参会者的信息, 并了解所述参 会者的会议发言或与所述参会者进行通信;
如果关键信息包括演讲稿, 获取该演讲稿的文件信息; 如果关键信息包括语音或肢体动作, 获取所述语音或肢体动作针对的对象 的相关信息。
20、 根据权利要求 18所述的会议系统, 其特征在于, 所述互动与编辑单元 具体用于:
根据关键信息中包括人脸对会议摘要的编辑包括对人物添加肢体动作说 明、 背景知识说明、 发言状况说明, 或者对参会人员进行分类; 或者
根据关键信息中包括动作对会议摘要的编辑包括基于主要演讲者的一个肢 体动作的自动化学习结果进行编辑操作; 或者
对所述关键信息进行关联与排列, 再对这些关键信息生成的关键索引点定 义切换动作。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP12866313.5A EP2709357B1 (en) | 2012-01-16 | 2012-09-11 | Conference recording method and conference system |
| US14/103,542 US8750678B2 (en) | 2012-01-16 | 2013-12-11 | Conference recording method and conference system |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN201210012814.8A CN102572356B (zh) | 2012-01-16 | 2012-01-16 | 记录会议的方法和会议系统 |
| CN201210012814.8 | 2012-01-16 |
Related Child Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| US14/103,542 Continuation US8750678B2 (en) | 2012-01-16 | 2013-12-11 | Conference recording method and conference system |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2013107184A1 true WO2013107184A1 (zh) | 2013-07-25 |
Family
ID=46416682
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2012/081220 Ceased WO2013107184A1 (zh) | 2012-01-16 | 2012-09-11 | 记录会议的方法和会议系统 |
Country Status (4)
| Country | Link |
|---|---|
| US (1) | US8750678B2 (zh) |
| EP (1) | EP2709357B1 (zh) |
| CN (1) | CN102572356B (zh) |
| WO (1) | WO2013107184A1 (zh) |
Families Citing this family (44)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN102572356B (zh) * | 2012-01-16 | 2014-09-03 | 华为技术有限公司 | 记录会议的方法和会议系统 |
| KR101862128B1 (ko) * | 2012-02-23 | 2018-05-29 | 삼성전자 주식회사 | 얼굴을 포함하는 영상 처리 방법 및 장치 |
| US8976226B2 (en) * | 2012-10-15 | 2015-03-10 | Google Inc. | Generating an animated preview of a multi-party video communication session |
| CN103841358B (zh) * | 2012-11-23 | 2017-12-26 | 中兴通讯股份有限公司 | 低码流的视频会议系统及方法、发送端设备、接收端设备 |
| CN103903074B (zh) * | 2012-12-24 | 2018-10-30 | 华为技术有限公司 | 一种视频交流的信息处理方法及装置 |
| WO2015001492A1 (en) * | 2013-07-02 | 2015-01-08 | Family Systems, Limited | Systems and methods for improving audio conferencing services |
| CN104423936B (zh) * | 2013-08-23 | 2017-12-26 | 联想(北京)有限公司 | 一种获取数据方法及电子设备 |
| US9660824B2 (en) * | 2013-09-25 | 2017-05-23 | Cisco Technology, Inc. | Renewing an in-process meeting without interruption in a network environment |
| CN105323531A (zh) * | 2014-06-30 | 2016-02-10 | 三亚中兴软件有限责任公司 | 视频会议热点场景的检测方法和装置 |
| US9672829B2 (en) * | 2015-03-23 | 2017-06-06 | International Business Machines Corporation | Extracting and displaying key points of a video conference |
| CN104954151A (zh) * | 2015-04-24 | 2015-09-30 | 成都腾悦科技有限公司 | 一种基于网络会议的会议纪要提取与推送方法 |
| EP3101838A1 (en) * | 2015-06-03 | 2016-12-07 | Thomson Licensing | Method and apparatus for isolating an active participant in a group of participants |
| US9530426B1 (en) * | 2015-06-24 | 2016-12-27 | Microsoft Technology Licensing, Llc | Filtering sounds for conferencing applications |
| CN106897304B (zh) * | 2015-12-18 | 2021-01-29 | 北京奇虎科技有限公司 | 一种多媒体数据的处理方法和装置 |
| CN106982344B (zh) * | 2016-01-15 | 2020-02-21 | 阿里巴巴集团控股有限公司 | 视频信息处理方法及装置 |
| CN107295291A (zh) * | 2016-04-11 | 2017-10-24 | 中兴通讯股份有限公司 | 一种会议记录方法及系统 |
| CN107123060A (zh) * | 2016-12-02 | 2017-09-01 | 国网四川省电力公司信息通信公司 | 一种用于支持地区供电公司早会系统 |
| US10171256B2 (en) | 2017-02-07 | 2019-01-01 | Microsoft Technology Licensing, Llc | Interactive timeline for a teleconference session |
| US10193940B2 (en) | 2017-02-07 | 2019-01-29 | Microsoft Technology Licensing, Llc | Adding recorded content to an interactive timeline of a teleconference session |
| US10070093B1 (en) | 2017-02-24 | 2018-09-04 | Microsoft Technology Licensing, Llc | Concurrent viewing of live content and recorded content |
| CN107181927A (zh) * | 2017-04-17 | 2017-09-19 | 努比亚技术有限公司 | 一种信息共享方法和设备 |
| CN108305632B (zh) * | 2018-02-02 | 2020-03-27 | 深圳市鹰硕技术有限公司 | 一种会议的语音摘要形成方法及系统 |
| CN108346034B (zh) * | 2018-02-02 | 2021-10-15 | 深圳市鹰硕技术有限公司 | 一种会议智能管理方法及系统 |
| TWI684159B (zh) * | 2018-03-20 | 2020-02-01 | 麥奇數位股份有限公司 | 用於互動式線上教學的即時監控方法 |
| CN108597518A (zh) * | 2018-03-21 | 2018-09-28 | 安徽咪鼠科技有限公司 | 一种基于语音识别的会议记录智能麦克风系统 |
| CN108537508A (zh) * | 2018-03-30 | 2018-09-14 | 上海爱优威软件开发有限公司 | 会议记录方法及系统 |
| CN108595645B (zh) * | 2018-04-26 | 2020-10-30 | 深圳市鹰硕技术有限公司 | 会议发言管理方法以及装置 |
| CN110473545A (zh) * | 2018-05-11 | 2019-11-19 | 视联动力信息技术股份有限公司 | 一种基于会议室的会议处理方法和装置 |
| CN108810446A (zh) * | 2018-06-07 | 2018-11-13 | 北京智能管家科技有限公司 | 一种视频会议的标签生成方法、装置、设备和介质 |
| CN115469785B (zh) * | 2018-10-19 | 2023-11-28 | 华为技术有限公司 | 时间轴用户界面 |
| CN109561273A (zh) * | 2018-10-23 | 2019-04-02 | 视联动力信息技术股份有限公司 | 识别视频会议发言人的方法和装置 |
| CN109274922A (zh) * | 2018-11-19 | 2019-01-25 | 国网山东省电力公司信息通信公司 | 一种基于语音识别的视频会议控制系统 |
| CN111601125A (zh) * | 2019-02-21 | 2020-08-28 | 杭州海康威视数字技术股份有限公司 | 录播视频摘要生成方法、装置、电子设备及可读存储介质 |
| CN111246150A (zh) * | 2019-12-30 | 2020-06-05 | 咪咕视讯科技有限公司 | 视频会议的控制方法、系统、服务器和可读存储介质 |
| JP7218802B2 (ja) * | 2020-02-27 | 2023-02-07 | 日本電気株式会社 | サーバ装置、会議支援方法及びプログラム |
| CN113014854B (zh) * | 2020-04-30 | 2022-11-11 | 北京字节跳动网络技术有限公司 | 互动记录的生成方法、装置、设备及介质 |
| CN113556501A (zh) * | 2020-08-26 | 2021-10-26 | 华为技术有限公司 | 音频处理方法及电子设备 |
| CN112685534B (zh) * | 2020-12-23 | 2022-12-30 | 上海掌门科技有限公司 | 在创作过程中生成已创作内容的脉络信息的方法与设备 |
| CN115567670B (zh) * | 2021-07-02 | 2024-10-01 | 信骅科技股份有限公司 | 会议检视方法及装置 |
| CN113779234B (zh) * | 2021-09-09 | 2024-07-05 | 京东方科技集团股份有限公司 | 会议发言人的讲话纪要生成方法、装置、设备及介质 |
| CN115002396B (zh) * | 2022-06-01 | 2023-03-21 | 北京美迪康信息咨询有限公司 | 一种多信号源统一处理系统及方法 |
| US12531061B2 (en) * | 2023-04-03 | 2026-01-20 | Comcast Cable Communications, Llc | Methods and systems for enhanced conferencing |
| TWI890350B (zh) * | 2023-12-21 | 2025-07-11 | 英濟股份有限公司 | 智能會議輔助系統及生成會議紀錄的方法 |
| CN119003759B (zh) * | 2024-08-22 | 2025-05-30 | 航天物联网技术有限公司 | 一种基于大型语言模型的会议纪要生成方法 |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN1601531A (zh) * | 2003-09-26 | 2005-03-30 | 精工爱普生株式会社 | 用于为视听演示内容制作摘要和索引的方法与设备 |
| CN1655609A (zh) * | 2004-02-13 | 2005-08-17 | 精工爱普生株式会社 | 记录视频会议数据的方法和系统 |
| CN1662057A (zh) * | 2004-02-25 | 2005-08-31 | 日本先锋公司 | 备忘录文件创建和管理方法、会议服务器及网络会议系统 |
| CN1674672A (zh) * | 2004-03-22 | 2005-09-28 | 富士施乐株式会社 | 会议信息处理装置和方法以及计算机可读存储介质 |
| CN102572356A (zh) * | 2012-01-16 | 2012-07-11 | 华为技术有限公司 | 记录会议的方法和会议系统 |
Family Cites Families (16)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH11191882A (ja) * | 1997-12-25 | 1999-07-13 | Nec Corp | テレビ会議予約システムおよびテレビ会議予約プログラムを記録した記録媒体 |
| GB2342802B (en) * | 1998-10-14 | 2003-04-16 | Picturetel Corp | Method and apparatus for indexing conference content |
| US7213051B2 (en) | 2002-03-28 | 2007-05-01 | Webex Communications, Inc. | On-line conference recording system |
| US7598975B2 (en) | 2002-06-21 | 2009-10-06 | Microsoft Corporation | Automatic face extraction for use in recorded meetings timelines |
| JP4430929B2 (ja) | 2003-12-18 | 2010-03-10 | 株式会社日立製作所 | 自動録画システム |
| CN1801726A (zh) | 2004-12-31 | 2006-07-12 | 研华股份有限公司 | 一种会议记录装置及内建有该装置的影音撷取与输出设备 |
| US8036263B2 (en) | 2005-12-23 | 2011-10-11 | Qualcomm Incorporated | Selecting key frames from video frames |
| US7822811B2 (en) | 2006-06-16 | 2010-10-26 | Microsoft Corporation | Performance enhancements for video conferencing |
| US7847815B2 (en) | 2006-10-11 | 2010-12-07 | Cisco Technology, Inc. | Interaction based on facial recognition of conference participants |
| US7987423B2 (en) | 2006-10-11 | 2011-07-26 | Hewlett-Packard Development Company, L.P. | Personalized slide show generation |
| US8121277B2 (en) | 2006-12-12 | 2012-02-21 | Cisco Technology, Inc. | Catch-up playback in a conferencing system |
| US8887067B2 (en) | 2008-05-30 | 2014-11-11 | Microsoft Corporation | Techniques to manage recordings for multimedia conference events |
| US8290124B2 (en) | 2008-12-19 | 2012-10-16 | At&T Mobility Ii Llc | Conference call replay |
| US8373743B2 (en) | 2009-03-13 | 2013-02-12 | Avaya Inc. | System and method for playing back individual conference callers |
| US8907984B2 (en) | 2009-07-08 | 2014-12-09 | Apple Inc. | Generating slideshows using facial detection information |
| US8560309B2 (en) | 2009-12-29 | 2013-10-15 | Apple Inc. | Remote conferencing center |
-
2012
- 2012-01-16 CN CN201210012814.8A patent/CN102572356B/zh not_active Expired - Fee Related
- 2012-09-11 WO PCT/CN2012/081220 patent/WO2013107184A1/zh not_active Ceased
- 2012-09-11 EP EP12866313.5A patent/EP2709357B1/en active Active
-
2013
- 2013-12-11 US US14/103,542 patent/US8750678B2/en active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN1601531A (zh) * | 2003-09-26 | 2005-03-30 | 精工爱普生株式会社 | 用于为视听演示内容制作摘要和索引的方法与设备 |
| CN1655609A (zh) * | 2004-02-13 | 2005-08-17 | 精工爱普生株式会社 | 记录视频会议数据的方法和系统 |
| CN1662057A (zh) * | 2004-02-25 | 2005-08-31 | 日本先锋公司 | 备忘录文件创建和管理方法、会议服务器及网络会议系统 |
| CN1674672A (zh) * | 2004-03-22 | 2005-09-28 | 富士施乐株式会社 | 会议信息处理装置和方法以及计算机可读存储介质 |
| CN102572356A (zh) * | 2012-01-16 | 2012-07-11 | 华为技术有限公司 | 记录会议的方法和会议系统 |
Non-Patent Citations (1)
| Title |
|---|
| See also references of EP2709357A4 * |
Also Published As
| Publication number | Publication date |
|---|---|
| CN102572356A (zh) | 2012-07-11 |
| EP2709357A1 (en) | 2014-03-19 |
| CN102572356B (zh) | 2014-09-03 |
| EP2709357A4 (en) | 2014-11-12 |
| EP2709357B1 (en) | 2017-05-17 |
| US8750678B2 (en) | 2014-06-10 |
| US20140099075A1 (en) | 2014-04-10 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN102572356B (zh) | 记录会议的方法和会议系统 | |
| US20250124637A1 (en) | Integrated input/output (i/o) for a three-dimensional (3d) environment | |
| US11209956B2 (en) | Collaborative media sharing | |
| US8630854B2 (en) | System and method for generating videoconference transcriptions | |
| US9247205B2 (en) | System and method for editing recorded videoconference data | |
| US10163077B2 (en) | Proxy for asynchronous meeting participation | |
| CN112653902B (zh) | 说话人识别方法、装置及电子设备 | |
| US8243116B2 (en) | Method and system for modifying non-verbal behavior for social appropriateness in video conferencing and other computer mediated communications | |
| US8391455B2 (en) | Method and system for live collaborative tagging of audio conferences | |
| US11714595B1 (en) | Adaptive audio for immersive individual conference spaces | |
| CN101860713A (zh) | 为不支持视频的视频电话参与者提供非言辞通信的描述 | |
| WO2016039835A1 (en) | Real-time video transformations in video conferences | |
| US8693842B2 (en) | Systems and methods for enriching audio/video recordings | |
| CN111428450A (zh) | 基于社交应用的会议纪要处理方法及电子设备 | |
| CN117897930A (zh) | 用于混合在线会议的流式数据处理 | |
| US11894938B2 (en) | Executing scripting for events of an online conferencing service | |
| CN108737903B (zh) | 一种多媒体处理系统及多媒体处理方法 | |
| JPWO2013061389A1 (ja) | 会議通話システム、コンテンツ表示システム、要約コンテンツ再生方法およびプログラム | |
| JP2012156965A (ja) | ダイジェスト生成方法、ダイジェスト生成装置及びプログラム | |
| CN114629868B (zh) | 适用于远程工作的多媒体群聊室通信方法和系统及智能终端 | |
| US20240121280A1 (en) | Simulated choral audio chatter | |
| WO2024032111A1 (zh) | 在线会议的数据处理方法、装置、设备、介质及产品 | |
| JP2023552119A (ja) | 撮影中のパフォーマに対する聴衆反応のシミュレーション | |
| CN117750119A (zh) | 用于播放音频数据的方法、装置、设备和存储介质 | |
| Carey | Expressive communication in multimedia environments |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 12866313 Country of ref document: EP Kind code of ref document: A1 |
|
| REEP | Request for entry into the european phase |
Ref document number: 2012866313 Country of ref document: EP |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2012866313 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |