WO2006016461A1 - 撮像装置 - Google Patents

撮像装置 Download PDF

Info

Publication number
WO2006016461A1
WO2006016461A1 PCT/JP2005/012801 JP2005012801W WO2006016461A1 WO 2006016461 A1 WO2006016461 A1 WO 2006016461A1 JP 2005012801 W JP2005012801 W JP 2005012801W WO 2006016461 A1 WO2006016461 A1 WO 2006016461A1
Authority
WO
WIPO (PCT)
Prior art keywords
unit
keyword
image
imaging
subject
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2005/012801
Other languages
English (en)
French (fr)
Inventor
Naoaki Yorita
Tetsuo In
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Nikon Corp
Original Assignee
Nikon Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nikon Corp filed Critical Nikon Corp
Priority to EP05765650A priority Critical patent/EP1784006A4/en
Priority to JP2006531347A priority patent/JP4636024B2/ja
Priority to US11/659,100 priority patent/US20080095449A1/en
Publication of WO2006016461A1 publication Critical patent/WO2006016461A1/ja
Anticipated expiration legal-status Critical
Priority to US12/929,427 priority patent/US20110122292A1/en
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/41Structure of client; Structure of client peripherals
    • H04N21/422Input-only peripherals, i.e. input devices connected to specially adapted client devices, e.g. global positioning system [GPS]
    • H04N21/4223Cameras
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/40Client devices specifically adapted for the reception of or interaction with content, e.g. set-top-box [STB]; Operations thereof
    • H04N21/43Processing of content or additional data, e.g. demultiplexing additional data from a digital video stream; Elementary client operations, e.g. monitoring of home network or synchronising decoder's clock; Client middleware
    • H04N21/44Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs
    • H04N21/44008Processing of video elementary streams, e.g. splicing a video clip retrieved from local storage with an incoming video stream or rendering scenes according to encoded video stream scene graphs involving operations for analysing video streams, e.g. detecting features or characteristics in the video stream
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N21/00Selective content distribution, e.g. interactive television or video on demand [VOD]
    • H04N21/80Generation or processing of content or additional data by content creator independently of the distribution process; Content per se
    • H04N21/83Generation or processing of protective or descriptive data associated with content; Content structuring
    • H04N21/84Generation or processing of descriptive data, e.g. content descriptors
    • H04N21/8405Generation or processing of descriptive data, e.g. content descriptors represented by keywords
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04NPICTORIAL COMMUNICATION, e.g. TELEVISION
    • H04N5/00Details of television systems
    • H04N5/76Television signal recording
    • H04N5/765Interface circuits between an apparatus for recording and another apparatus
    • H04N5/77Interface circuits between an apparatus for recording and another apparatus between a recording apparatus and a television camera
    • H04N5/772Interface circuits between an apparatus for recording and another apparatus between a recording apparatus and a television camera the recording apparatus and the television camera being placed in the same enclosure

Definitions

  • the present invention relates to an imaging apparatus having a moving image shooting function composed of a plurality of images.
  • Patent Document 1 discloses a moving image processing device that generates and assigns at least one keyword to each scene of a moving image based on a user operation, and stores the generated moving image together with the moving image in a storage device. Yes.
  • Patent Document 1 Japanese Patent Application Laid-Open No. 6-309381
  • Patent Document 2 JP-A-5-204990
  • Patent Document 2 the user only has to specify one image in which a subject exists, but this one image is one of a series of moving images. Therefore, in order to specify this one image, the user plays a moving image, There is a problem that it is necessary to wait until a desired portion is reproduced. Even if only one image is specified, in order to assign multiple types of keywords, it is necessary to specify one image with the subject for each keyword. It takes much more time in proportion to the number of types of cards.
  • An object of the present invention is to provide an imaging device capable of creating metadata quickly and easily when shooting a moving image.
  • the imaging apparatus of the present invention repeatedly captures an image of a scene and a storage unit that stores in advance a plurality of keywords and a detection method for detecting a subject corresponding to the keyword for each keyword.
  • An image capturing unit that generates a moving image composed of a plurality of frames of images, and the keyword for a plurality of frames of the moving image obtained by the image capturing in parallel with the image capturing by the image capturing unit.
  • a detection unit that detects the subject based on the detection method for each, a keyword corresponding to the detected subject according to a detection result by the detection unit, and the subject corresponding to the keyword
  • An image information creation unit that creates image information including information indicating time existing in the scene, the moving image generated by the imaging unit, and the image information creation unit. And and an image recording unit for recording in association with the image information.
  • the image information creation unit as information indicating the time, for each of the subjects detected by the detection unit, the start time and end of the time when the subject was present in the object scene Information including the time may be created.
  • the image information creation unit detects, as the image information, the detection unit in the frame image for each of the images of the plurality of frames in which the subject is detected by the detection unit.
  • the key corresponding to the subject You can also create information that includes a single word.
  • the imaging unit includes a temporary recording unit that temporarily records an image obtained by imaging, and the detection unit parallels the writing of the image to the temporary recording unit. And the generation of the image information may be performed.
  • a keyword designating unit that designates the keyword to be detected by the detection unit among the plurality of keywords stored in advance in the storage unit is provided, and the detection unit is configured by the keyword designating unit. Only the specified keyword may be detected based on the detection method for each keyword.
  • a shooting scene recording unit that records shooting conditions in association with each of a plurality of predetermined shooting scenes, and selects one shooting scene from the plurality of shooting scenes.
  • a selection unit wherein the imaging unit generates the moving image according to the shooting condition associated with the one shooting scene selected by the selection unit, and the keyword designating unit is The keyword to be detected by the detection unit may be designated based on the one photographing scene selected by the above.
  • a keyword-specific moving image generating unit that generates a moving image for each of the keywords specified by the keyword specifying unit, wherein the image recording unit includes the moving image generated by the imaging unit, It is also possible to record a moving image for each of the keywords generated by the keyword-specific moving image generation unit!
  • the recording unit detects, by the detection unit, the subject corresponding to the keyword specified by the keyword specifying unit among the images of the plurality of frames generated by the imaging unit. It is also possible to record the moving image excluding the frame image and the image information in association with each other.
  • the storage unit has a registration mode for newly registering the keyword and a detection method for detecting the subject corresponding to the keyword, and the keyword to be newly registered is input in the registration mode.
  • Keyword input unit and the keyword A specification unit that specifies a subject corresponding to the keyword input by the input unit based on an image obtained by the imaging unit, and a feature amount of the subject specified by the specification unit are extracted, and the feature is extracted.
  • a control unit for setting a detection method for detecting the subject based on a quantity, the keyword input by the keyword input unit, and the detection method set by the setting unit in the storage unit May be provided.
  • FIG. 1 is a functional block diagram of an electronic camera 1.
  • FIG. 2 is a flowchart regarding image recognition.
  • FIG. 3 is a flowchart regarding creation of metadata.
  • FIG. 4 is a diagram for explaining an example of metadata.
  • FIG. 5 is a diagram for explaining an example of metadata.
  • FIG. 6 is a diagram for explaining an example of metadata.
  • FIG. 7 is a diagram for explaining an example of metadata.
  • FIG. 1 is a functional block diagram of the electronic camera 1.
  • the electronic camera 1 includes a photographic optical system 2 that is a zoom lens, a CCD (Charge
  • Image sensor 3 which is a photoelectric conversion element such as Coupled Device
  • a / D conversion unit (circuit) 4 image determination unit (circuit) 5, temporary recording unit (circuit) 6, image processing unit (circuit) 7, compression unit (Circuit) 8, Image recording unit 9, AE, AF judgment unit (Circuit) 10, AE'AF unit 11, Image recognition unit (Circuit) 12, Metadata creation unit (Circuit) 13 and control each part Equipped with control unit 14 Brown.
  • the control unit 14 forms a subject image on the image sensor 3 via the photographing optical system 2 when the user instructs to shoot a moving image.
  • the image sensor 3 photoelectrically converts the subject image formed on the imaging surface and outputs analog data.
  • the AZD conversion unit 4 converts analog data output from the image sensor 3 into digital data, outputs the digital data to the image determination unit 5, and temporarily records it in the temporary recording unit 6.
  • the image determination unit 5 determines the exposure state and the focus adjustment state based on the brightness information and contrast information of the input image data.
  • the temporary recording unit 6 is a nota memory and outputs data input by the AZD conversion unit to an image processing unit 7, an image recognition unit 12 and a control unit 14 which will be described later.
  • the image processing unit 7 performs ⁇ processing and image processing described later, and outputs the image data after the image processing to the compression unit 8.
  • the compression unit 8 performs a predetermined compression process on the image data after image processing, and the image recording unit 9 records the compressed image data.
  • the AE'AF determination unit 10 determines the exposure control value and the focus adjustment control value based on the exposure state and the focus adjustment state determined by the image determination unit 5, and supplies them to the AE'AF unit 11. At the same time, distance information to the subject and luminance information in the shooting screen are output to the image recognition unit 12 as necessary.
  • the AE / AF unit 1 1 controls the photographing optical system 2 and the image sensor 3 based on the exposure control value and the focus adjustment control value.
  • the image recognition unit 12 determines conditions for image processing in the image processing unit 7 and supplies the conditions to the image processing unit 7.
  • the image processing unit 7 outputs the image recognition unit 12 (for example, conditions for specifying the “blue sky” area in the shooting screen (for example, the hue range that can be regarded as the blue sky, the luminance range, the position of the blue area in the shooting screen, etc.) ))
  • To identify the area of the specific subject eg “blue sky”
  • the specified area will be more appropriate color (for example, a blue sky-like color)
  • Perform image processing color conversion).
  • the image recognition unit 12 stores in advance a detection method for detecting a subject corresponding to the keyword for each of the plurality of keywords, and the AE'AF determination unit described above as necessary. Considering the distance information to the subject acquired from 10 and the luminance information in the shooting screen, image recognition based on the feature amount of the image data is performed. Further, the metadata creation unit 13 stores a plurality of keywords similar to those stored in the image recognition unit 12 in advance. And then, according to the recognition result by the image recognition unit 12, the creation of metadata which is a feature of the present invention is performed, and the created metadata is supplied to the image recording unit 9.
  • the electronic camera 1 includes a control unit 14 that controls each unit.
  • the control unit 14 records each operation program in an internal memory (not shown) in advance, and controls each unit according to the operation program. Further, the control unit 14 outputs image data to a display unit 15 to be described later and detects the state of the operation member 16 to be described later.
  • FIG. 1 only the characteristic portions of the present invention are illustrated with arrows indicating the connection between the control unit 14 and each unit.
  • the electronic camera 1 includes a display unit 15 that displays an image being captured, a menu display at the time of a user operation described later, and an operation member 16 such as a button that receives a user operation.
  • the electronic camera 1 is an imaging device having a moving image shooting function, and generates a moving image composed of a plurality of images in response to a user operation.
  • the control unit 14 detects this and controls the imaging optical system 2 and the image sensor 3 to start capturing a moving image.
  • Moving image capturing is realized by repeatedly capturing images at predetermined time intervals.
  • the control unit 14 performs image recognition in the image recognition unit 12 on the image obtained by the imaging in parallel with the imaging. Specifically, the control unit 14 recognizes the image one frame before (or before) the image in parallel with the writing of the image to the temporary recording unit 6 through the image recognition unit 12. I do. Further, the metadata creation unit 13 creates metadata according to the image recognition result.
  • the image recognition unit 12 stores “person” and “blue sky” as a plurality of keywords, and detects a subject corresponding to the keyword (specifically, “face part” and “blue sky part”).
  • the method for recognizing “)” and the metadata creation unit 13 store “person” and “blue sky” as a plurality of keywords.
  • step S1 whether the control unit 14 recognizes the face portion via the image recognition unit 12 or not. Determine whether or not. Note that the image recognition unit 12 recognizes a face part based on information obtained when the image processing condition is determined in the image processing unit 7 described above.
  • the image recognition unit 12 sets a hue range that can be recognized as a skin color and performs recognition. At this time, it is considered that even areas recognized as the same skin color are faces, hands, and feet. Therefore, when recognizing a face portion, the skin color region of the hand or foot can be excluded by a method such as determining the contour of the skin color region or the region where the hair color exists in the adjacent region as the face portion. At this time, the lower limit value and the upper limit value of the size of the skin color area that can be regarded as a face may be set based on the shooting distance information input from the AE'AF determination unit 10.
  • control unit 14 recognizes "person” as the subject through the image recognition unit 12 and the keyword ("person") corresponding to the recognized subject in step S2. ) Is output to the metadata creation unit 13.
  • step S1 if the face portion is not recognized, or if a keyword (“person”) corresponding to the recognized subject is output to the metadata creation unit 13, the control unit 14 in step S3 It is determined whether or not the blue sky portion has been recognized through the recognition unit 12. Note that the image recognition unit 12 recognizes the blue portion based on the information obtained when the image processing condition is determined in the image processing unit 7 described above.
  • the image recognition unit 12 recognizes a blue portion and determines whether or not the blue portion exists in the upper part of the image.
  • the image recognition unit 12 performs a differentiation process in the vertical direction of the image on the blue portion recognized based on a predetermined hue range, and if there is an edge in the horizontal direction of the image, The region above is determined as a blue sky region. By performing such processing, it is possible to extract a horizontal line even for an image having the blue sky and the sea as subjects, and as a result, it is possible to accurately grasp the blue sky region. Then, the image recognition unit 12 outputs information indicating the blue sky area to the image processing unit 7 as an image processing condition.
  • the control unit 14 recognizes "blue sky” as a subject through the image recognition unit 12 in step S4, and a keyword ("blue sky") corresponding to the recognized subject. ) Is output to the metadata creation unit 13. The control unit 14 then passes through the image recognition unit 12. The image recognition process ends.
  • meta data is “data related to data” and is image information including a keyword corresponding to the subject and information indicating the time when the subject corresponding to the keyword exists in the scene.
  • step S 10 the control unit 14 determines whether there is a keyword output from the image recognition unit 12 via the metadata creation unit 13. If there is a keyword, in step S11, the control unit 14 causes the subject corresponding to the keyword to be recognized in the previous image in time through the metadata creation unit 13. It is determined whether or not. If the subject corresponding to the keyword is not the subject recognized in the previous image! (If the subject is not recognized in the previous image but recognized in the current image), step S12 Then, the control unit 14 reads out the keyword corresponding to the object via the metadata creation unit 13 and gives information indicating the start time to the keyword data. Note that the control unit 14 performs the processing of step S11 and step S12 on all subjects recognized by the image recognition unit 12.
  • step S 13 the control unit 14 determines whether or not there is a missing subject via the metadata creation unit 13. That is, when there is a subject that is recognized in the previous image in time but not recognized in the current image, in step S14, the control unit 14 passes the metadata creation unit 13 through Information indicating the end time is added to the key data corresponding to the disappeared subject. Note that the metadata creation unit 13 determines whether or not there is a power of an erased subject for all subjects recognized by the image recognition unit 12 for the previous image. Then, the control unit 14 outputs the created metadata from the metadata creation unit 13 to the image recording unit 9, and ends the metadata creation process.
  • the image recognition unit 12 stores a “person”, “X person (number of people)”, and “ball” as a plurality of keywords, and detects a subject corresponding to the key word (specifically, , ⁇ Face part '', ⁇ Number of face parts '', and ⁇ Circle part '' recognition method), and the metadata creation unit 13 uses ⁇ people '', ⁇ X people (number of people) as keywords ) "And” Ball "memorize and show examples [0031]
  • the horizontal axis indicates the shooting time, and shows each image and a keyword corresponding to the recognized subject in the image.
  • the image recognition unit 12 recognizes “one” “face part”
  • the metadata creation unit 13 creates the keyword “person” and the keyword “one person”, and sets the start time as T1. .
  • the image recognition unit 12 recognizes the “face part” as “two” and “circular part”, and the metadata creation unit 13 uses the keywords “person”, “two people”, and “ball”. As well as creating it, the start time is T2 for “2 people” and “ball”, and the end time is T2 for the keyword “1 person”.
  • the image recognition unit 1 2 recognizes the “face part” as “three” and “circular part”, and the metadata creation unit 13 sets the keywords “people”, “three people”, and “ball”.
  • the start time is T3 for “3 people”
  • the end time is T3 for the keywords “2 people”.
  • the image recognition unit 12 recognizes the “face part” as “two” and “circular part”, and the metadata creation unit 13 creates the keywords “person”, “two people”, and “ball”.
  • the start time for T2 is T4
  • the end time is T4.
  • the metadata creating unit 13 determines the start time and end time of the time when the subject exists in the scene for each detected subject, as shown in FIG. Create metadata that includes it.
  • control unit 14 records the metadata supplied from the metadata creation unit 13 in the image recording unit 9 in association with the image supplied from the compression unit 8.
  • the electronic camera 1 performs image recognition described using the flowchart of FIG. 2 and creation of metadata described using the flowchart of FIG. 3 in parallel with the imaging of the moving image.
  • a plurality of keywords and a detection method for detecting a subject corresponding to the keyword for each keyword are stored in advance, and when generating a moving image,
  • a subject is detected for a plurality of frames of moving images obtained by imaging, and according to the detection result, a keyword, Image information including information indicating the time when the subject corresponding to the keyword exists in the scene is created and recorded in association with the moving image. Therefore, it is possible to create metadata quickly and easily at the time of moving image shooting without requiring playback of moving images after shooting and designation based on user operations. That is, since the creation of metadata is completed almost simultaneously with the end of moving image shooting, it is possible to eliminate the hassle of playing back a moving image and transferring an image for adding a keyword.
  • information indicating the time information including a start time and an end time of the time when the subject exists in the object scene is created for each detected subject. . Accordingly, it is possible to easily indicate the time when the subject corresponding to a certain keyword exists in the object scene.
  • the subject is detected in parallel with the writing of the image obtained by the imaging to the temporary recording unit 6. Therefore, shortening of processing time can be expected.
  • the electronic camera of the second embodiment has the same configuration as the electronic camera 1 of the first embodiment. Hereinafter, description will be made using the same reference numerals as in the first embodiment.
  • the image recognition unit 12 stores “person A”, “person B”, “person C”, and “ball” as a plurality of keywords, and uses a detection method for detecting a subject corresponding to the keyword.
  • An example is described in which the metadata creation unit 13 stores “person A”, “person B”, “person C”, and “ball” as a plurality of keywords.
  • a feature amount is stored such that the face contour is a base type, the minor axis / major axis is 0.5, the hair color is brown, and the eye interval Z minor axis is 0.5.
  • a feature amount that has a circular outline and that includes both white and black at the same time is stored.
  • the control unit 14 recognizes “person A”, “person B”, “person C”, and “ball” by performing the same processing as in the flowchart shown in FIG. Then, instead of the flow chart process shown in FIG. 3, metadata is created for each image of each frame. For example, as shown in FIG. 6, when creating metadata for images of frame numbers 1 to 7, as shown in FIG. 7, for each frame, the time when the image of the frame was shot and the image Create metadata that contains information (" ⁇ " and "X” in Figures 6 and 7) that indicates the recognition results when "person A”, “person B”, “person C”, and "ball” are recognized .
  • the control unit 14 records metadata from the metadata creating unit 13 to the image recording unit 9 independently of the image recording from the compression unit 8 to the image recording unit 9. Since the creation of metadata requires time for image recognition by the image recognition unit 12, it may take longer than the time required for image processing by the image processing unit 7 and compression processing by the compression unit 8. Therefore, the control unit 14 starts recording an image from the compression unit 8 to the image recording unit 9 without waiting for the metadata creation unit 13 to finish creating the metadata. Record metadata in association with the recorded (or recording) image. When the creation and recording of metadata is complete, use the display unit 15 etc. It is preferable to notify the user that the creation and recording have been completed.
  • a keyword corresponding to the subject detected in the image of the frame is used as image information. Create information to include. Therefore, it is possible to easily indicate the time during which the subject corresponding to a certain keyword is present in the scene.
  • the compression format of the moving image was not described in particular, but the moving image is compressed in any format such as MPEG, motion JPEG, JPEG2000. May be. Of course, non-compressed recording may be used to improve image quality.
  • the subject detection and the creation of image information which are features of the present invention, do not have to be performed for all the images constituting the moving image. For example, in MPE G, subject detection and image information may be generated only for the I picture that is the reference image, or subject detection and image information may be thinned out at an appropriate time interval. You can also make it.
  • the image and the image information are recorded in association with each other.
  • the image and the image information are stored as separate files. It may be recorded as the same file. That is, any format such as tag information and header information may be used.
  • other elements may be arranged or combined.
  • the chromaticity value of the skin color, the ratio between the eye interval and the eye-nose (mouth) interval, or the like may be used as the feature amount.
  • a keyword is selected according to the selected shooting scene. May be specified.
  • a keyword based on the shooting conditions associated with the shooting scene and detecting the subject only for that keyword, for example, when shooting in "Portrait” mode .
  • the user intends to photograph “people”. Therefore, images retrieved with the intention of “people” in a later search can be searched for “blue sky”. The search is not performed, the accuracy of the search can be improved, and the processing time for adding metadata to each image is shortened.
  • each process may be performed by specifying a keyword. good.
  • a keyword is specified according to the focus adjustment condition (for example, when the shooting distance between the electronic camera 1 and the focused subject is short, “people” is specified as the keyword and the background “ You can exclude keywords such as “blue sky” and “mountain”, and if the shooting distance is long, exclude “people” and specify keywords such as “blue sky” and “mountain”.
  • the user may perform an operation via the operation member 16 and directly specify the keypad on the display unit 15.
  • a keyword to be detected from among a plurality of keywords and detecting a subject only for the keyword it is possible to limit processing for keywords that are not necessary in advance. Therefore, it is possible to avoid the creation of wrong (unnecessary) keyword metadata due to misrecognition.
  • a keyword it is preferable to have a configuration in which not only one keyword but also a plurality of keywords can be specified.
  • multiple keywords it should be possible to specify them using logical expressions. For example, it can be specified by the following logical expression that “the subject of keyword A does not exist and the subject of keyword B exists”.
  • a moving image for each keyword is generated and recorded when shooting of a moving image is completed or recorded in the image recording unit 9. May be.
  • the user can perform playback.
  • it is possible to easily reproduce only the moving image of the necessary part.
  • it is possible to continuously reproduce only the images having the keyword specified by the user as metadata, without generating separate images.
  • any keyword may be provided.
  • various keywords can be provided without imposing a load.
  • a through image is displayed on the display unit 15 and a frame including a subject corresponding to a specified keyword among the already captured images.
  • These images may be reduced and displayed as a list on the display unit 15 at the same time as the through image.
  • the representative image of the scene among the images of the scene consisting of frame 4 to frame 6 is the high frequency component. It may be determined based on the above, or the image of the first or last frame of the scene may be used as the representative image.
  • images constituting a scene may be displayed in a superimposed manner. With such a representative image, it is possible to grasp the image of the entire scene from one image.
  • a display similar to that during shooting of the moving image described above may be performed. Further, the moving image recorded in the image recording unit 9 may be subjected to a decoding process to detect a subject and create image information.
  • a registration mode for newly registering a keyword and a detection method for detecting a subject corresponding to the keyword may be provided. In other words, in the registration mode, new registration is performed via the operation member 16 or the like. Enter a keyword and specify the subject corresponding to the keyword. The subject may be specified based on an image already obtained by imaging, or a detection method (feature amount) may be specifically specified.
  • the feature amount of the subject may be extracted from the image, and the detection method may be determined based on the feature amount, or the image itself may be used. Then, detection may be performed by pattern matching. In either case, the input keyword and the detection method (or image for no-turn matching) are stored in the image recognition unit 12, so that the keyword can be detected during subsequent imaging.
  • an external database having keywords and a subject detection method may be used to input via a communication / recording medium.
  • the operation performed by the imaging apparatus may be realized by a computer.
  • the moving image generated by the imaging device is captured in a computer, and the captured image is subjected to processing such as subject detection and image information creation described in the first and second embodiments. It's okay. Further, the detection of the subject and creation of the image information may be performed while capturing the image.
  • the invention described in the first embodiment and the second embodiment may be executed by appropriately replacing or combining them.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Signal Processing (AREA)
  • Television Signal Processing For Recording (AREA)
  • Studio Devices (AREA)

Abstract

 複数のキーワードと、キーワードごとに、そのキーワードに対応する被写体を検出する検出方法とを予め記憶する記憶部と、被写界を繰り返し撮像して、複数フレームの画像から構成される動画像を生成する撮像部と、撮像部による撮像と並行して、撮像により得られた動画像のうち、複数フレームの画像に対して、キーワードごとの検出方法に基づいて被写体の検出を行う検出部と、検出部による検出結果に応じて、検出された被写体に対応するキーワードと、そのキーワードに対応する被写体が被写界に存在した時間を示す情報とを含む画像情報を作成する画像情報作成部と、撮像部により生成された動画像と、画像情報作成部により作成された画像情報とを関連づけて記録する画像記録部とを備えることで、動画撮影の際に、迅速かつ手軽にメタデータを作成可能な撮像装置を提供する。                                                                             

Description

明 細 書
撮像装置
技術分野
[0001] 本発明は、複数の画像から構成された動画撮影機能を有する撮像装置に関する。
背景技術
[0002] 近年、記録媒体の記録容量が増大し、一般家庭にお!、てもディジタルィ匕された動 画像を扱う環境が整ってきている。動画像の中から所望のシーンを探すために、動 画像を構成する画像ごとにキーワードを付与し、このキーワードを基に所望のシーン を検索して表示する技術が知られている (例えば、特許文献 1,特許文献 2参照)。 特許文献 1に記載の発明では、ユーザ操作に基づいて、動画像の各シーンに少な くとも 1つのキーワードを生成して付与し、動画像とともに記憶装置に記憶する動画像 処理装置が開示されている。
[0003] また、特許文献 2に記載の発明では、ユーザ操作に基づいて、ある被写体が存在 する画像を 1枚指定すると、動画像中の同一の被写体が現れている部分区間を探し 出して共通のキーワードおよび関連情報を自動的に付与する技術が開示されている 特許文献 1:特開平 6 - 309381号公報
特許文献 2:特開平 5 - 204990号公報
発明の開示
発明が解決しょうとする課題
[0004] し力しながら、特許文献 1に開示された動画処理装置においては、ユーザは、動画 像を再生させた上で、画像 1枚ずつに対して逐一キーワードを入力しなければならな い。したがって、特に長時間にわたる動画像の場合には、キーワードの入力に多大 の時間を要するという問題があった。
また、特許文献 2に開示された技術においては、ユーザは、被写体が存在する画 像を 1枚指定すれば良いが、この 1枚の画像は、一連の動画像の中の 1枚である。し たがって、ユーザは、この 1枚の画像を指定するために、動画像を再生させた上で、 所望の部分が再生されるまで待たなければならないという問題があった。また、指定 する画像は 1枚であっても、複数種類のキーワードを付与するためには、各キーヮー ドごとに被写体が存在する画像を 1枚指定しなければならな 、ので、付与するキーヮ ードの種類数に比例して、さらに多大な時間を要する。
[0005] さらに、撮像装置にこれらの動画像処理装置を適用する場合には、撮像により生成 された動画像を再生させ、前述したキーワードの指定などを行うことになる。したがつ て、このような動作を行っている間は、撮影を行うことができない。さらに、撮像装置が 備える表示装置 (液晶モニタなど)は、サイズが小さぐ細部にわたる画像の確認が困 難である場合もある。
[0006] 本発明は、動画撮影の際に、迅速かつ手軽にメタデータを作成可能な撮像装置を 提供することを目的とする。
課題を解決するための手段
[0007] 本発明の撮像装置は、複数のキーワードと、前記キーワードごとに、そのキーワード に対応する被写体を検出する検出方法とを予め記憶する記憶部と、被写界を繰り返 し撮像して、複数フレームの画像から構成される動画像を生成する撮像部と、前記撮 像部による撮像と並行して、撮像により得られた前記動画像のうち、複数フレームの 画像に対して、前記キーワードごとの前記検出方法に基づいて前記被写体の検出を 行う検出部と、前記検出部による検出結果に応じて、検出された前記被写体に対応 する前記キーワードと、そのキーワードに対応する前記被写体が前記被写界に存在 した時間を示す情報とを含む画像情報を作成する画像情報作成部と、前記撮像部 により生成された前記動画像と、前記画像情報作成部により作成された前記画像情 報とを関連づけて記録する画像記録部とを備える。
[0008] なお、好ましくは、前記画像情報作成部は、前記時間を示す情報として、前記検出 部により検出された前記被写体ごとに、その被写体が前記被写界に存在した時間の 開始時刻と終了時刻とを含む情報を作成するようにしても良い。
また、好ましくは、前記画像情報作成部は、前記画像情報として、前記検出部によ り前記被写体の検出を行った前記複数フレームの画像のそれぞれについて、そのフ レームの画像において前記検出部により検出された前記被写体に対応する前記キ 一ワードを含む情報を作成するようにしても良 、。
[0009] また、好ましくは、前記撮像部は、撮像により得られた画像を一時記録する一時記 録部を備え、前記検出部は、前記一時記録部に対する画像の書き込みに並行して、 前記被写体の検出、および前記画像情報の生成を行うようにしても良い。
また、好ましくは、前記記憶部に予め記憶された前記複数のキーワードのうち、前記 検出部による検出の対象となる前記キーワードを指定するキーワード指定部を備え、 前記検出部は、前記キーワード指定部により指定された前記キーワードについての み、前記キーワードごとの前記検出方法に基づいて前記被写体の検出を行うようにし ても良い。
[0010] また、好ましくは、予め定められた複数種類の撮影シーンごとに撮影条件を対応付 けて記録する撮影シーン記録部と、前記複数種類の撮影シーンのうち、 1つの撮影 シーンを選択する選択部とをさらに備え、前記撮像部は、前記選択部により選択され た前記 1つの撮影シーンに対応付けられた前記撮影条件にしたがって前記動画像を 生成し、前記キーワード指定部は、前記選択部により選択された前記 1つの撮影シー ンに基づいて、前記検出部による検出の対象となる前記キーワードを指定するように しても良い。
[0011] また、好ましくは、前記画像情報作成部により作成された前記画像情報に基づいて
、前記キーワード指定部により指定された前記キーワードごとの動画像を生成するキ 一ワード別動画像生成部を備え、前記画像記録部は、前記撮像部により生成された 前記動画像に加えて、前記キーワード別動画像生成部により生成された前記キーヮ ードごとの動画像を記録するようにしても良!、。
[0012] また、好ましくは、前記記録部は、前記撮像部により生成された前記複数フレーム の画像のうち、前記キーワード指定部により指定された前記キーワードに対応する前 記被写体が前記検出部により検出されたフレームの画像を除いた動画像と、前記画 像情報とを関連づけて記録するようにしても良 、。
また、好ましくは、前記記憶部に、前記キーワードと、そのキーワードに対応する前 記被写体を検出する検出方法とを新規登録する登録モードを有し、前記登録モード において、新規登録するキーワードを入力するキーワード入力部と、前記キーワード 入力部により入力された前記キーワードに対応する被写体を、前記撮像部により得ら れた画像に基づいて指定する指定部と、前記指定部により指定された前記被写体の 特徴量を抽出し、その特徴量に基づいて前記被写体を検出する検出方法を設定す る設定部と、前記キーワード入力部により入力された前記キーワードと、前記設定部 により設定された前記検出方法とを前記記憶部に記憶させる制御部とを備えるように しても良い。
発明の効果
[0013] 本発明によれば、動画撮影の際に、迅速かつ手軽にメタデータを作成可能な撮像 装置を提供することができる。
図面の簡単な説明
[0014] [図 1]電子カメラ 1の機能ブロック図である。
[図 2]画像認識に関するフローチャートである。
[図 3]メタデータの作成に関するフローチャートである。
[図 4]メタデータの例について説明する図である。
[図 5]メタデータの例について説明する図である。
[図 6]メタデータの例について説明する図である。
[図 7]メタデータの例について説明する図である。
発明を実施するための最良の形態
[0015] <第 1実施形態 >
以下、図面に基づいて、本発明の第 1実施形態について詳細に説明する。 なお、第 1実施形態では、本発明の撮像装置の一例として、動画撮影機能を有す る電子カメラを例に挙げて説明を行う。
図 1は電子カメラ 1の機能ブロック図である。
[0016] 電子カメラ 1は、図 1に示すようにズームレンズである撮影光学系 2、 CCD (Charge
Coupled Device)などの光電変換素子である撮像素子 3、 A/D変換部(回路) 4 、画像判定部(回路) 5、一時記録部(回路) 6、画像処理部(回路) 7、圧縮部(回路) 8、画像記録部 9、 AE,AF判定部(回路) 10、 AE'AFユニット 11、画像認識部(回 路) 12、メタデータ作成部(回路) 13を備えるとともに、各部を制御する制御部 14を備 える。
[0017] 制御部 14は、ユーザにより動画像の撮影が指示されると撮影光学系 2を介して被 写体像を撮像素子 3に結像する。撮像素子 3は、撮像面上に結像された被写体像を 光電変換して、アナログデータを出力する。 AZD変換部 4は、撮像素子 3より出力さ れるアナログデータをディジタルデータに変換して、画像判定部 5に出力するとともに 、一時記録部 6に一時記録する。画像判定部 5は、入力された画像データの明るさ情 報やコントラスト情報などに基づき、露出状態や焦点調節状態を判定する。一時記録 部 6は、ノ ッファメモリであり、 AZD変換部により入力されたデータを後述する画像 処理部 7、画像認識部 12および制御部 14に出力する。画像処理部 7は、 γ処理や、 後述する画像処理を行い、圧縮部 8に画像処理後の画像データを出力する。圧縮部 8は、画像処理後の画像データに対して所定の圧縮処理を行い、画像記録部 9は圧 縮後の画像データを記録する。なお、 AE'AF判定部 10は、画像判定部 5において 判定された露出状態や焦点調節状態に基づ 、て、露出制御値や焦点調節制御値 を決定し、 AE'AFユニット 11に供給するとともに、必要に応じて、被写体までの距離 情報や撮影画面内の輝度情報などを画像認識部 12に出力する。 AE · AFユニット 1 1は、露出制御値や焦点調節制御値に基づいて、撮影光学系 2や撮像素子 3を制御 する。
[0018] また、画像認識部 12は、画像処理部 7における画像処理の条件を決定して、画像 処理部 7に供給する。画像処理部 7は、画像認識部 12の出力(例えば、撮影画面中 の「青空」の領域を特定する条件 (例えば、青空と見なせる色相範囲、輝度範囲、撮 影画面内の青色領域の位置等))に基づいて、特定被写体 (例えば「青空」)の領域 を特定し、特定された領域について、より適切な色となるように(例えば、青空らしい 色となるように)、必要に応じて画像処理 (色変換)を施す。
[0019] また、画像認識部 12は、複数のキーワードごとに、そのキーワードに対応する被写 体を検出する検出方法を、予め記憶しており、必要に応じて、前述した AE'AF判定 部 10から取得した被写体までの距離情報や撮影画面内の輝度情報などを加味して 、画像データの特徴量に基づいた画像認識を行う。また、メタデータ作成部 13は、画 像認識部 12が記憶しているものと同様の複数のキーワードを予め記憶している。そし て、画像認識部 12による認識結果に応じて、本発明の特徴であるメタデータの作成 を行 ヽ、作成したメタデータを画像記録部 9に供給する。
[0020] また、電子カメラ 1は、各部を制御する制御部 14を備える。制御部 14は、内部の不 図示のメモリに各動作プログラムを予め記録し、この動作プログラムにしたがって各部 を制御する。また、制御部 14は、後述する表示部 15に対する画像データの出力を行 うとともに、後述する操作部材 16の状態を検知する。なお、図 1においては、本発明 の特徴部分についてのみ、制御部 14と各部の接続を示す矢印を図示した。
[0021] また、電子カメラ 1は、撮影中の画像の表示や後述するユーザ操作時のメニュー表 示などを表示する表示部 15、ユーザ操作を受け付けるボタンなどの操作部材 16を 備える。
なお、電子カメラ 1は、動画像撮影機能を有する撮像装置であり、ユーザ操作に応 じて、複数の画像から構成される動画像を生成する。
以上説明した構成の電子カメラ 1において、本発明の特徴である動画像撮影時の 動作について説明する。
[0022] ユーザにより操作部 16を介して動画像の撮影が指示されると制御部 14はこれを検 知し、撮影光学系 2、撮像素子 3を制御して動画像の撮影を開始する。動画像の撮 影は、所定の時間間隔で繰り返し撮像を行うことにより実現する。そして、制御部 14 は、撮像と並行して、撮像により得られた画像に対して、画像認識部 12において画 像認識を行う。具体的には、制御部 14は、画像認識部 12を介して、 AZD変換部 4 力も一時記録部 6への画像の書き込みと並行して、 1フレーム前 (またはさらに前)の 画像に対する画像認識を行う。さらに、画像認識結果に応じて、メタデータ作成部 13 において、メタデータの作成を行う。
[0023] まず、画像認識について、図 2のフローチャートを用いて説明する。なお、ここでは 、画像認識部 12は、複数のキーワードとして「人」および「青空」を記憶し、キーワード に対応する被写体を検出する検出方法 (具体的には、「顔部分」および「青空部分」 を認識する方法)をそれぞれ記憶していて、かつ、メタデータ作成部 13は、複数のキ 一ワードとして「人」および「青空」を記憶して 、る例を説明する。
[0024] ステップ S1にお ヽて、制御部 14は、画像認識部 12を介して、顔部分を認識したか 否かを判定する。なお、画像認識部 12は、前述した画像処理部 7における画像処理 の条件の決定の際に得た情報に基づいて顔部分の認識を行う。
画像認識部 12は、顔部分を認識するために、肌色として認識できる色相範囲を設 定して、認識を行う。この際、同じ肌色として認識される領域でも、顔、手、足であるこ とが考えられる。したがって、顔部分を認識する場合には、肌色領域の輪郭や隣接 領域に髪の毛の色が存在する領域を顔部分と判断するなどの手法によって、手や足 の肌色領域を除外することができる。また、この際 AE'AF判定部 10から入力される 撮影距離情報をもとに、顔と見なせる肌色領域の大きさの下限値や上限値を設定す るようにしても良い。
[0025] 顔部分を認識した場合、ステップ S2にお ヽて、制御部 14は、画像認識部 12を介し て、「人」を被写体として認識し、認識した被写体に対応するキーワード(「人」)をメタ データ作成部 13に出力する。
ステップ S1において、顔部分を認識しな力つた場合、もしくは、認識した被写体に 対応するキーワード(「人」)をメタデータ作成部 13に出力すると、ステップ S3にお ヽ て制御部 14は、画像認識部 12を介して、青空部分を認識したか否かを判定する。な お、画像認識部 12は、前述した画像処理部 7における画像処理の条件の決定の際 に得た情報に基づ 、て青色部分の認識を行う。
[0026] さらに、画像認識部 12は、青色部分を認識し、その青色部分が画像上部に存在す るカゝ否かを判定する。
画像認識部 12は、予め定められた色相範囲に基づ 、て認識した青色部分に対し て、画像の垂直方向における微分処理を行い、画像の水平方向にエッジが存在する 場合には、このエッジより上方の領域を青空の領域と判定する。このような処理を行う ことで、青空と海とを被写体とする画像に対しても、水平線を抽出することが可能であ り、結果として青空の領域を正確に把握することができる。そして、画像認識部 12は 、青空の領域を示す情報を画像処理条件として画像処理部 7に出力する。
[0027] 青空部分を認識した場合、ステップ S4にお ヽて、制御部 14は、画像認識部 12を 介して、「青空」を被写体として認識し、認識した被写体に対応するキーワード(「青空 」)をメタデータ作成部 13に出力する。そして、制御部 14は、画像認識部 12を介した 画像認識処理を終了する。
次に、メタデータの作成について、図 3のフローチャートを用いて説明する。なお、メ タデータとは、「データに関するデータ」のことであり、被写体に対応するキーワードと 、そのキーワードに対応する被写体が被写界に存在した時間を示す情報とを含む画 像情報である。
[0028] まず、ステップ S10において、制御部 14は、メタデータ作成部 13を介して、画像認 識部 12から出力されたキーワードがあるか否かを判定する。そして、キーワードがあ る場合、ステップ S 11において、制御部 14は、メタデータ作成部 13を介して、そのキ 一ワードに対応する被写体が、時間的に 1つ前の画像でも認識された被写体である か否かを判定する。キーワードに対応する被写体が、 1つ前の画像でも認識された被 写体でな!、場合(1つ前の画像では認識されず、今回の画像で認識された被写体で ある場合)、ステップ S12において、制御部 14は、メタデータ作成部 13を介して、被 写体に対応するキーワードを読み出し、キーワードのデータに、開始時刻を示す情 報を付与する。なお、制御部 14は、ステップ S11およびステップ S12の処理を、画像 認識部 12にお 、て認識されたすベての被写体に対して行う。
[0029] そして、ステップ S 13において、制御部 14は、メタデータ作成部 13を介して、消失 した被写体がある力否かを判定する。すなわち、時間的に 1つ前の画像では認識さ れ、今回の画像では認識されなカゝつた被写体が存在する場合、ステップ S14におい て、制御部 14は、メタデータ作成部 13を介して、消失した被写体に対応するキーヮ ードのデータに、終了時刻を示す情報を付与する。なお、メタデータ作成部 13は、消 失した被写体がある力否かの判定を、 1つ前の画像について画像認識部 12が認識 したすベての被写体に対して行う。そして、制御部 14は、作成したメタデータをメタデ ータ作成部 13から画像記録部 9に出力して、メタデータの作成処理を終了する。
[0030] 図 4を用いて作成されるメタデータの例について説明する。図 4では、画像認識部 1 2は、複数のキーワードとして「人」、「X人(人数)」および「ボール」を記憶し、キーヮ ードに対応する被写体をする検出方法 (具体的には、「顔部分」、「顔部分の数」、「円 形部分」を認識する方法)をそれぞれ記憶していて、かつ、メタデータ作成部 13は、 キーワードとして「人」、「X人 (人数)」および「ボール」を記憶して 、る場合の例を示す [0031] 横軸は撮影時刻を示し、各画像と、その画像にぉ 、て認識された被写体に対応す るキーワードを示す。
時間 TOより撮影を開始し、時間 T1において被写界内に 1人現れたとする。ここで、 画像認識部 12は「顔部分」を「1つ」認識し、メタデータ作成部 13は、「人」というキー ワードと「1人」というキーワードを作成し、開始時刻を T1とする。
[0032] 時間 T2になったときにもう 1人被写界内に現れ、同時にボールが現れたとする。こ こで画像認識部 12は、「顔部分」を「2つ」と「円形部分」とを認識し、メタデータ作成 部 13は、「人」、「2人」、「ボール」というキーワードを作成するとともに、「2人」および「 ボール」については開始時刻を T2とし、「1人」というキーワードについては終了時刻 を T2とする。
[0033] 時間 T3となったときにさらにもう 1人被写界内に現れたとする。ここで画像認識部 1 2は、「顔部分」を「3つ」と「円形部分」とを認識し、メタデータ作成部 13は、「人」、「3 人」、「ボール」というキーワードを作成するとともに、「3人」については開始時刻を T3 とし、「2人」と 、うキーワードにつ 、ては終了時刻を T3とする。
時間 T4となったときに 1人が被写界内から消えたとする。ここで画像認識部 12は、「 顔部分」を「2つ」と「円形部分」とを認識し、メタデータ作成部 13は、「人」、「2人」、「 ボール」というキーワードを作成するとともに、「2人」については開始時刻を T4とし、「 3人」と 、うキーワードにつ 、ては終了時刻を T4とする。
[0034] 時間 T5となったときにボールと 2人が被写界内から消えたとする。ここで画像認識 部 12は、「顔部分」を「1つ」認識し、メタデータ作成部 13は、「人」、「1人」というキー ワードを作成するとともに、「1人」については開始時刻を T5とし、「2人」および「ボー ル」というキーワードについては終了時刻を T5とする。その後時間 T6で撮影を終了 したとする。
[0035] 以上説明した処理により、メタデータ作成部 13は、図 5に示すように、検出された被 写体ごとに、その被写体が被写界に存在した時間の開始時刻と終了時刻とを含むメ タデータを作成する。
上記したようにメタデータ作成部 13を介してメタデータの作成を終了すると、制御部 14は、メタデータ作成部 13から供給されたメタデータを圧縮部 8から供給された画像 と関連づけて、画像記録部 9に記録する。
[0036] このように、電子カメラ 1は、動画像の撮像と並行して、図 2のフローチャートを用い て説明した画像認識および図 3のフローチャートを用いて説明したメタデータの作成 を行う。
以上説明したように、第 1実施形態によれば、複数のキーワードと、キーワードごと に、そのキーワードに対応する被写体を検出する検出方法とを予め記憶しておき、動 画像を生成する際に、撮像動作、および一時記録部 6への書き込み動作と並行して 、撮像により得られた動画像のうち、複数フレームの画像に対して、被写体の検出を 行い、検出結果に応じて、キーワードと、そのキーワードに対応する被写体が被写界 に存在した時間を示す情報とを含む画像情報を作成して動画像と関連づけて記録 する。したがって、撮影後の動画像の再生およびユーザ操作に基づく指定などを必 要としないで、動画撮影の際に、迅速かつ手軽にメタデータを作成することができる。 すなわち、動画像の撮影終了と略同時にメタデータの作成が終了するので、動画像 の再生の手間や、キーワード付与のための画像の転送などの手間をなくすことができ る。
[0037] また、第 1実施形態によれば、時間を示す情報として、検出された被写体ごとに、そ の被写体が被写界に存在した時間の開始時刻と終了時刻とを含む情報を作成する 。したがって、あるキーワードに対応する被写体が被写界に存在した時間を分力り易 く示すことができる。
また、第 1実施形態によれば、撮像により得られた画像の一時記録部 6に対する書 き込みに並行して、被写体の検出を行う。したがって、処理時間の短縮が期待できる
[0038] <第 2実施形態 >
以下、図面に基づいて、本発明の第 2実施形態について詳細に説明する。 なお、第 2実施形態では、本発明の撮像装置の一例として、第 1実施形態と同様に
、動画撮影機能を有する電子カメラを例に挙げて説明を行う。以下では、第 1実施形 態と異なる部分についてのみ説明を行う。 [0039] 第 2実施形態の電子カメラは、第 1実施形態の電子カメラ 1と同様の構成を有する。 以下では、第 1実施形態と同様の符号を用いて説明する。
第 2実施形態では、画像認識部 12は、複数のキーワードとして「人物 A」、「人物 B」 、「人物 C」および「ボール」を記憶し、キーワードに対応する被写体を検出する検出 方法をそれぞれ記憶していて、かつ、メタデータ作成部 13は、複数のキーワードとし て「人物 A」、「人物 B」、「人物 C」および「ボール」を記憶している例を説明する。なお 、検出方法としては、「人物 A」については、顔輪郭が丸型、かつ、短径 Z長径 =0. 8であり、髪色が肌色 (スキンヘッド)で、目間隔 Z短径 =0. 6という特徴量を記憶す る。また、「人物 B」については、顔輪郭がベース型、かつ、短径/長径 =0. 5であり 、髪色が茶色で、目間隔 Z短径 =0. 5という特徴量を記憶する。また、「人物 C」につ いては、顔輪郭が楕円、かつ、短径 Z長径 =0. 6であり、髪色が黒色で、目間隔 Z 短径 =0. 5という特徴量を記憶する。また、「ボール」については、輪郭が円形であり 、色が白と黒とを同時に含むという特徴量を記憶する。
[0040] そして、制御部 14は、図 2で示したフローチャートと同様の処理を行って、「人物 A」 、「人物 B」、「人物 C」および「ボール」を認識する。そして、図 3で示したフローチヤ一 トの処理に代えて、各フレームの画像ごとにメタデータを作成する。例えば、図 6に示 すように、フレーム番号 1〜7の画像についてメタデータを作成する場合、図 7に示す ように、フレームごとに、そのフレームの画像が撮影された時刻と、その画像において 「人物 A」、「人物 B」、「人物 C」および「ボール」を認識した際の認識結果を示す情報 (図 6および図 7では「〇」および「 X」)を含むメタデータを作成する。
[0041] そして、制御部 14は、圧縮部 8から画像記録部 9への画像の記録と独立に、メタデ ータ作成部 13から画像記録部 9へのメタデータの記録を行う。メタデータの作成には 、画像認識部 12による画像認識に時間を要するため、画像処理部 7による画像処理 および圧縮部 8による圧縮処理に要する時間よりも時間を要する場合がある。そこで 、制御部 14は、メタデータ作成部 13によるメタデータの作成終了を待たずに、圧縮 部 8から画像記録部 9への画像の記録を開始し、メタデータの作成が終了すると、既 に記録した (または記録中の)画像に関連づけてメタデータの記録を行う。なお、メタ データの作成および記録が終了した際には、表示部 15などを用いて、メタデータの 作成および記録が終了したことをユーザに報知するようにすると良い。
[0042] 以上説明したように、第 2実施形態によれば、画像情報として、被写体の検出を行 つた複数フレームの画像のそれぞれについて、そのフレームの画像において検出さ れた被写体に対応するキーワードを含む情報を作成する。したがって、あるキーヮー ドに対応する被写体が被写界に存在した時間を分力り易く示すことができる。
なお、第 1実施形態および第 2実施形態では、動画像の圧縮形式について特に説 明しなかったが、動画像は MPEG、モーション JPEG、JPEG2000など、どのような形 式で圧縮されるものであっても良い。勿論、高画質化のために非圧縮で記録するも のであっても構わない。また、本発明の特徴である被写体の検出および画像情報の 作成は、動画像を構成するすべての画像に対して行わなくても良い。例えば、 MPE Gにお 、て、基準画像である Iピクチャのみに対して被写体の検出および画像情報の 作成を行うようにしても良いし、適当な時間間隔で間引いて被写体の検出および画 像情報の作成を行うようにしても良 、。
[0043] また、第 1実施形態および第 2実施形態では、画像と画像情報とを関連づけて記録 することを説明したが、画像と画像情報とが関連づけ付けられていれば、別ファイルと して記録しても同じファイルとして記録しても良い。すなわち、タグ情報、ヘッダ情報 など、どのような形式であっても良い。
また、第 1実施形態および第 2実施形態で説明した被写体の検出方法以外に、他 の要素をカ卩えたり組み合わせたりするようにしても良い。例えば、被写体が人物の場 合に、特徴量として肌の色の色度値、目の間隔と目と鼻(口)の間隔との比などを利 用するようにしても良い。また、画像データとして、輪郭データなどを利用し、パターン マッチングを行う構成としても良 、。
[0044] また、第 1実施形態および第 2実施形態において、電子カメラ 1が、撮影条件が対 応付けられた 、わゆるシーンモードを備えて 、る場合、選択された撮影シーンに応じ てキーワードを指定しても良い。このように、撮影シーンに対応付けられた撮影条件 に基づいて、キーワードを指定し、そのキーワードについてのみ、被写体の検出を行 うことにより、例えば、「ポートレート」モードにて撮影している場合、キーワードとして「 人」を指定し、「青空」を除外することで、被写界内に青空と認識してしまうような青い 壁の部分が存在して ヽても「青空」を誤認識してしまうのを回避することができる。また 、たとえ実際に青空が写っていた場合でもユーザは、「人」の撮影を意図しているの で、後の検索時に「人」を意図して撮影された画像が「青空」の検索で検索されること がなくなり、検索時の精度の向上も期待することができるとともに、各画像に対するメ タデータの付与処理時間が短縮化される。
[0045] また、第 1実施形態および第 2実施形態では、予め定められた複数のキーワードの すべてについて、画像認識を行う例を示したが、キーワードを指定して各処理を行う ようにしても良い。例えば、焦点調節の条件に応じてキーワードを指定する(例えば、 電子カメラ 1と合焦被写体との間の撮影距離が短い場合には、「人」をキーワードとし て指定して、背景である「青空」、「山」などのキーワードは除外し、逆に撮影距離が長 い場合には、「人」を除外して、「青空」、「山」などのキーワードを指定する)ようにして も良いし、操作部材 16を介して、ユーザが操作を行って、表示部 15上で直接キーヮ ードを指定するようにしても良い。このように、複数のキーワードのうち、検出の対象と なるキーワードを指定し、そのキーワードについてのみ、被写体の検出を行うことによ り、予め必要のないキーワードにおける処理を制限することができる。そのため、誤認 識により間違った (不用な)キーワードのメタデータが作成されるのを回避することが できる。なお、キーワードを指定する際には、 1つのキーワードを指定するだけでなく 、複数のキーワードを指定可能な構成にするのが好ましい。複数のキーワードを指定 する場合には、論理式を用いて指定可能にすると良い。例えば、以下の論理式によ り、「キーワード Aの被写体は存在せず、かつ、キーワード Bの被写体は存在する」と いう指定が可能である。
[0046] [数 1]
[0047] また、第 1実施形態および第 2実施形態において、動画像の撮影が終了した際また は、画像記録部 9に記録する際に、キーワードごとの動画像を生成して記録するよう にしても良い。このように、作成した画像情報に基づいて、キーワード指定部により指 定されたキーワードごとの動画像を生成して記録することにより、ユーザは、再生の際 に、必要な部分の動画像のみを簡単に再生することができる。また、別に画像を生成 しなくても、再生の際に、ユーザが指定されたキーワードをメタデータとして持ってい る画像のみを連続的に再生するようにしても良い。
[0048] また、第 1実施形態および第 2実施形態において、動画像の撮影が終了した際また は、画像記録部 9に記録する際に、あるキーワードを除く動画像を生成して記録する ようにしても良い。このように、作成した画像情報に基づいて、キーワード指定部によ り指定されたキーワードを除く動画像を生成して記録することにより、ユーザは、ある キーワードに対応する被写体を除外した動画像を得ることができる。
[0049] また、第 1実施形態および第 2実施形態で例を挙げたキーワード以外にどのような キーワードを備えるようにしても良い。ここで、キーワードに対応する被写体の検出に 、画質調整や画像処理などに使用している情報を流用することにより、負荷をかけず に様々なキーワードを備えることができる。
また、第 1実施形態および第 2実施形態において、動画像の撮影中に、表示部 15 にスルー画像を表示するとともに、既に撮像された画像のうち、指定されたキーワード に対応する被写体を含むフレームの画像を縮小して、スルー画像と同時に表示部 1 5に一覧表示するようにしても良い。例えば、図 6のフレーム 4〜フレーム 6に示すよう に、「人物 B」が連続して検出された場合には、フレーム 4〜フレーム 6からなるシーン の画像のうち、シーンの代表画像を高周波成分などに基づいて決定するようにしても 良いし、シーンの最初または最後のフレームの画像を代表画像としても良い。また、 一覧表示する代わりにシーンを構成する画像を重ね合わせて表示するようにしても 良い。このような代表画像とすれば、 1つの画像によりシーン全体の画像について把 握することができる。
[0050] また、画像記録部 9に記録した動画像の再生時にも、上述した動画像の撮影中と 同様の表示を行うようにしても良い。また、画像記録部 9に記録した動画像に対して、 復号化処理を施して、被写体の検出および画像情報の作成を行うようにしても良 、。 また、第 1実施形態および第 2実施形態において、キーワードと、そのキーワードに 対応する被写体を検出する検出方法とを新規登録する登録モードを備えるようにし ても良い。すなわち、登録モードにおいて、操作部材 16などを介して新規登録する キーワードを入力し、そのキーワードに対応する被写体を指定する。被写体の指定は 、撮像により既に得られた画像に基づいて行うようにしても良いし、検出方法 (特徴量 )を具体的に指定するようにしても良い。なお、撮像により既に得られた画像に基づい て指定する場合には、画像から被写体の特徴量を抽出し、その特徴量に基づいて検 出方法を決めるようにしても良いし、画像そのものを利用してパターンマッチングによ り検出を行うようにしても良い。いずれの場合も、入力されたキーワードと、検出方法( または、ノターンマッチング用の画像)とを画像認識部 12に記憶させることにより、以 降の撮像時にキーワード検出の対象とすることができる。また、キーワードと被写体の 検出方法とを持つ外部データベースから、通信'記録媒体を介して入力するようにし ても良い。
[0051] また、第 1実施形態および第 2実施形態において、撮像装置が行った動作を、コン ピュータで実現するようにしても良い。すなわち、撮像装置によって生成された動画 像をコンピュータに取り込み、取り込んだ画像に対して、第 1実施形態および第 2実 施形態で説明した被写体の検出および画像情報の作成などの処理を行うようにして も良い。また、画像の取り込みを行いながら、被写体の検出および画像情報の作成 を行うようにしても良い。
[0052] また、第 1実施形態および第 2実施形態で説明した発明を、適宜入れ替えまたは組 み合わせて実行するようにしても良!、。

Claims

請求の範囲
[1] 複数のキーワードと、前記キーワードごとに、そのキーワードに対応する被写体を検 出する検出方法とを予め記憶する記憶部と、
被写界を繰り返し撮像して、複数フレームの画像から構成される動画像を生成する 撮像部と、
前記撮像部による撮像と並行して、撮像により得られた前記動画像のうち、複数フ レームの画像に対して、前記キーワードごとの前記検出方法に基づいて前記被写体 の検出を行う検出部と、
前記検出部による検出結果に応じて、検出された前記被写体に対応する前記キー ワードと、そのキーワードに対応する前記被写体が前記被写界に存在した時間を示 す情報とを含む画像情報を作成する画像情報作成部と、
前記撮像部により生成された前記動画像と、前記画像情報作成部により作成され た前記画像情報とを関連づけて記録する画像記録部と
を備えたことを特徴とする撮像装置。
[2] 請求項 1に記載の撮像装置にぉ 、て、
前記画像情報作成部は、前記時間を示す情報として、前記検出部により検出され た前記被写体ごとに、その被写体が前記被写界に存在した時間の開始時刻と終了 時刻とを含む情報を作成する
ことを特徴とする撮像装置。
[3] 請求項 1に記載の撮像装置にぉ 、て、
前記画像情報作成部は、前記画像情報として、前記検出部により前記被写体の検 出を行った前記複数フレームの画像のそれぞれにつ 、て、そのフレームの画像にお いて前記検出部により検出された前記被写体に対応する前記キーワードを含む情報 を作成する
ことを特徴とする撮像装置。
[4] 請求項 1に記載の撮像装置にぉ 、て、
前記撮像部は、撮像により得られた画像を一時記録する一時記録部を備え、 前記検出部は、前記一時記録部に対する画像の書き込みに並行して、前記被写 体の検出、および前記画像情報の生成を行う
ことを特徴とする撮像装置。
[5] 請求項 1に記載の撮像装置にぉ 、て、
前記記憶部に予め記憶された前記複数のキーワードのうち、前記検出部による検 出の対象となる前記キーワードを指定するキーワード指定部を備え、
前記検出部は、前記キーワード指定部により指定された前記キーワードについての み、前記キーワードごとの前記検出方法に基づいて前記被写体の検出を行う ことを特徴とする撮像装置。
[6] 請求項 5に記載の撮像装置において、
予め定められた複数種類の撮影シーンごとに撮影条件を対応付けて記録する撮影 シーン記録部と、
前記複数種類の撮影シーンのうち、 1つの撮影シーンを選択する選択部とをさらに 備え、
前記撮像部は、前記選択部により選択された前記 1つの撮影シーンに対応付けら れた前記撮影条件にしたがって前記動画像を生成し、
前記キーワード指定部は、前記選択部により選択された前記 1つの撮影シーンに 基づいて、前記検出部による検出の対象となる前記キーワードを指定する
ことを特徴とする撮像装置。
[7] 請求項 5または請求項 6に記載の撮像装置にぉ 、て、
前記画像情報作成部により作成された前記画像情報に基づいて、前記キーワード 指定部により指定された前記キーワードごとの動画像を生成するキーワード別動画 像生成部を備え、
前記画像記録部は、前記撮像部により生成された前記動画像に加えて、前記キー ワード別動画像生成部により生成された前記キーワードごとの動画像を記録する ことを特徴とする撮像装置。
[8] 請求項 5または請求項 6に記載の撮像装置にぉ 、て、
前記記録部は、前記撮像部により生成された前記複数フレームの画像のうち、前記 キーワード指定部により指定された前記キーワードに対応する前記被写体が前記検 出部により検出されたフレームの画像を除いた動画像と、前記画像情報とを関連づ けて記録する
ことを特徴とする撮像装置。
請求項 1に記載の撮像装置にぉ 、て、
前記記憶部に、前記キーワードと、そのキーワードに対応する前記被写体を検出す る検出方法とを新規登録する登録モードを有し、
前記登録モードにぉ 、て、新規登録するキーワードを入力するキーワード入力部と 前記キーワード入力部により入力された前記キーワードに対応する被写体を、前記 撮像部により得られた画像に基づいて指定する指定部と、
前記指定部により指定された前記被写体の特徴量を抽出し、その特徴量に基づい て前記被写体を検出する検出方法を設定する設定部と、
前記キーワード入力部により入力された前記キーワードと、前記設定部により設定 された前記検出方法とを前記記憶部に記憶させる制御部と
を備えたことを特徴とする撮像装置。
PCT/JP2005/012801 2004-08-09 2005-07-12 撮像装置 Ceased WO2006016461A1 (ja)

Priority Applications (4)

Application Number Priority Date Filing Date Title
EP05765650A EP1784006A4 (en) 2004-08-09 2005-07-12 IMAGE FORMING DEVICE
JP2006531347A JP4636024B2 (ja) 2004-08-09 2005-07-12 撮像装置
US11/659,100 US20080095449A1 (en) 2004-08-09 2005-07-12 Imaging Device
US12/929,427 US20110122292A1 (en) 2004-08-09 2011-01-24 Imaging device and metadata preparing apparatus

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2004232761 2004-08-09
JP2004-232761 2004-08-09

Related Child Applications (1)

Application Number Title Priority Date Filing Date
US12/929,427 Continuation US20110122292A1 (en) 2004-08-09 2011-01-24 Imaging device and metadata preparing apparatus

Publications (1)

Publication Number Publication Date
WO2006016461A1 true WO2006016461A1 (ja) 2006-02-16

Family

ID=35839237

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2005/012801 Ceased WO2006016461A1 (ja) 2004-08-09 2005-07-12 撮像装置

Country Status (4)

Country Link
US (2) US20080095449A1 (ja)
EP (1) EP1784006A4 (ja)
JP (1) JP4636024B2 (ja)
WO (1) WO2006016461A1 (ja)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008193563A (ja) * 2007-02-07 2008-08-21 Nec Design Ltd 撮像装置、再生装置、撮像方法、再生方法及びプログラム
WO2023188652A1 (ja) * 2022-03-30 2023-10-05 富士フイルム株式会社 記録方法、記録装置、及びプログラム

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP6327816B2 (ja) * 2013-09-13 2018-05-23 キヤノン株式会社 送信装置、受信装置、送受信システム、送信装置の制御方法、受信装置の制御方法、送受信システムの制御方法、及びプログラム

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH05204990A (ja) 1992-01-29 1993-08-13 Hitachi Ltd 動画像情報のキーワード付与方法
JPH06309381A (ja) 1993-04-20 1994-11-04 Ibm Japan Ltd 動画像処理装置
JPH07322125A (ja) * 1994-05-28 1995-12-08 Sony Corp 目標追尾装置
EP0805405A2 (en) 1996-02-05 1997-11-05 Texas Instruments Incorporated Motion event detection for video indexing
JPH10326278A (ja) * 1997-03-27 1998-12-08 Minolta Co Ltd 情報処理装置及び方法並びに情報処理プログラムを記録した記録媒体
JP2001057660A (ja) * 1999-08-17 2001-02-27 Hitachi Kokusai Electric Inc 動画像編集装置
JP2004023656A (ja) * 2002-06-19 2004-01-22 Canon Inc 画像処理装置、画像処理方法、及び、プログラム

Family Cites Families (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP3780623B2 (ja) * 1997-05-16 2006-05-31 株式会社日立製作所 動画像の記述方法
US7023469B1 (en) * 1998-04-30 2006-04-04 Texas Instruments Incorporated Automatic video monitoring system which selectively saves information
EP1081960B1 (en) * 1999-01-29 2007-12-19 Sony Corporation Signal processing method and video/voice processing device
JP3176893B2 (ja) * 1999-03-05 2001-06-18 株式会社次世代情報放送システム研究所 ダイジェスト作成装置,ダイジェスト作成方法およびその方法の各工程をコンピュータに実行させるためのプログラムを記録したコンピュータ読み取り可能な記録媒体
JP4683253B2 (ja) * 2000-07-14 2011-05-18 ソニー株式会社 Av信号処理装置および方法、プログラム、並びに記録媒体
JP3984029B2 (ja) * 2001-11-12 2007-09-26 オリンパス株式会社 画像処理装置およびプログラム
JP2003169312A (ja) * 2001-11-30 2003-06-13 Ricoh Co Ltd 電子番組表提供システム、電子番組表提供方法、そのプログラム、及びそのプログラムを記録した記録媒体
GB2395853A (en) * 2002-11-29 2004-06-02 Sony Uk Ltd Association of metadata derived from facial images
JP4081680B2 (ja) * 2003-11-10 2008-04-30 ソニー株式会社 記録装置、記録方法、記録媒体、再生装置、再生方法およびコンテンツの伝送方法
JP2005269604A (ja) * 2004-02-20 2005-09-29 Fuji Photo Film Co Ltd 撮像装置、撮像方法、及び撮像プログラム
US7586517B2 (en) * 2004-10-27 2009-09-08 Panasonic Corporation Image pickup apparatus

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH05204990A (ja) 1992-01-29 1993-08-13 Hitachi Ltd 動画像情報のキーワード付与方法
JPH06309381A (ja) 1993-04-20 1994-11-04 Ibm Japan Ltd 動画像処理装置
JPH07322125A (ja) * 1994-05-28 1995-12-08 Sony Corp 目標追尾装置
EP0805405A2 (en) 1996-02-05 1997-11-05 Texas Instruments Incorporated Motion event detection for video indexing
JPH10326278A (ja) * 1997-03-27 1998-12-08 Minolta Co Ltd 情報処理装置及び方法並びに情報処理プログラムを記録した記録媒体
JP2001057660A (ja) * 1999-08-17 2001-02-27 Hitachi Kokusai Electric Inc 動画像編集装置
JP2004023656A (ja) * 2002-06-19 2004-01-22 Canon Inc 画像処理装置、画像処理方法、及び、プログラム

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP1784006A4

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008193563A (ja) * 2007-02-07 2008-08-21 Nec Design Ltd 撮像装置、再生装置、撮像方法、再生方法及びプログラム
WO2023188652A1 (ja) * 2022-03-30 2023-10-05 富士フイルム株式会社 記録方法、記録装置、及びプログラム

Also Published As

Publication number Publication date
JPWO2006016461A1 (ja) 2008-05-01
US20110122292A1 (en) 2011-05-26
US20080095449A1 (en) 2008-04-24
JP4636024B2 (ja) 2011-02-23
EP1784006A1 (en) 2007-05-09
EP1784006A4 (en) 2012-04-18

Similar Documents

Publication Publication Date Title
JP4645685B2 (ja) カメラ、カメラ制御プログラム及び撮影方法
KR101532294B1 (ko) 자동 태깅 장치 및 방법
CN101796814B (zh) 摄像设备和摄像方法
US20090322896A1 (en) Image recording apparatus, image recording method, image processing apparatus, image processing method, and program
JP4474885B2 (ja) 画像分類装置及び画像分類プログラム
JP2010165012A (ja) 撮像装置、画像検索方法及びプログラム
JP4574459B2 (ja) 撮影装置及びその制御方法及びプログラム及び記憶媒体
JP4577252B2 (ja) カメラ、ベストショット撮影方法、プログラム
CN101611347A (zh) 成像装置
JP5375846B2 (ja) 撮像装置、自動焦点調整方法、およびプログラム
JP4839908B2 (ja) 撮像装置、自動焦点調整方法、およびプログラム
JP5181841B2 (ja) 撮影装置、撮影制御プログラム、並びに画像再生装置、画像再生プログラム
JP2007310813A (ja) 画像検索装置およびカメラ
JP2013121097A (ja) 撮像装置及び方法、並びに画像生成装置及び方法、プログラム
WO2007142237A1 (ja) 画像再生システム、デジタルカメラ、および画像再生装置
JP5168320B2 (ja) カメラ、ベストショット撮影方法、プログラム
JP5332369B2 (ja) 画像処理装置及び画像処理方法、並びにコンピュータ・プログラム
JP2010161547A (ja) 構図選択装置及びプログラム
JP5266701B2 (ja) 撮像装置、被写体分離方法、およびプログラム
JP5267645B2 (ja) 撮像装置、撮影制御方法、およびプログラム
JP2011119936A (ja) 撮影装置及び再生方法
JP5369776B2 (ja) 撮像装置、撮像方法、及び撮像プログラム
JP4636024B2 (ja) 撮像装置
JP2010237911A (ja) 電子機器
JP4842232B2 (ja) 撮影装置及び画像再生装置

Legal Events

Date Code Title Description
AK Designated states

Kind code of ref document: A1

Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS JP KE KG KM KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NG NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SM SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW

AL Designated countries for regional patents

Kind code of ref document: A1

Designated state(s): GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IS IT LT LU LV MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG

121 Ep: the epo has been informed by wipo that ep was designated in this application
WWE Wipo information: entry into national phase

Ref document number: 2006531347

Country of ref document: JP

WWE Wipo information: entry into national phase

Ref document number: 11659100

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

WWE Wipo information: entry into national phase

Ref document number: 2005765650

Country of ref document: EP

WWP Wipo information: published in national office

Ref document number: 2005765650

Country of ref document: EP

WWP Wipo information: published in national office

Ref document number: 11659100

Country of ref document: US