WO2017120681A1 - Procédé et système pour déterminer automatiquement une sortie tridimensionnelle de position d'informations audio sur la base de l'orientation d'un utilisateur dans un environnement immersif artificiel. - Google Patents

Procédé et système pour déterminer automatiquement une sortie tridimensionnelle de position d'informations audio sur la base de l'orientation d'un utilisateur dans un environnement immersif artificiel. Download PDF

Info

Publication number
WO2017120681A1
WO2017120681A1 PCT/CA2017/050050 CA2017050050W WO2017120681A1 WO 2017120681 A1 WO2017120681 A1 WO 2017120681A1 CA 2017050050 W CA2017050050 W CA 2017050050W WO 2017120681 A1 WO2017120681 A1 WO 2017120681A1
Authority
WO
WIPO (PCT)
Prior art keywords
audio
user
subset
input audio
channels
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CA2017/050050
Other languages
English (en)
Inventor
Michael Godfrey
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Individual
Original Assignee
Individual
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Individual filed Critical Individual
Publication of WO2017120681A1 publication Critical patent/WO2017120681A1/fr
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/302Electronic adaptation of stereophonic sound system to listener position or orientation
    • H04S7/303Tracking of listener position or orientation
    • H04S7/304For headphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R2460/00Details of hearing devices, i.e. of ear- or headphones covered by H04R1/10 or H04R5/033 but not provided for in any of their subgroups, or of hearing aids covered by H04R25/00 but not provided for in any of its subgroups
    • H04R2460/07Use of position data from wide-area or local-area positioning systems in hearing devices, e.g. program or information selection
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S2400/00Details of stereophonic systems covered by H04S but not provided for in its groups
    • H04S2400/03Aspects of down-mixing multi-channel audio to configurations with lower numbers of playback channels, e.g. 7.1 -> 5.1
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S2420/00Techniques used stereophonic systems covered by H04S but not provided for in its groups
    • H04S2420/01Enhancing the perception of the sound image or of the spatial distribution using head related transfer functions [HRTF's] or equivalents thereof, e.g. interaural time difference [ITD] or interaural level difference [ILD]

Definitions

  • the present specification relates generally to audio processing, and more specifically to a method and system for automatically determining a positional three dimensional output of audio information based on a user's orientation within an artificial immersive environment.
  • virtual reality otherwise known as artificial reality, provides users with an immersive and interactive audiovisual experience and may be applied to any form of entertainment, broadcasting or communication.
  • the user is immersed into a virtual reality environment when real-world stimuli are replaced with electronically generated stimuli.
  • This usually involves the use of audiovisual devices that provide audiovisual content, which may be based on any one or a combination of pre-recorded content, real-time content, or digitally simulated content.
  • audiovisual devices typically provide "spherical" three-dimensional visual simulations surrounding the user, which is synchronized to three-dimensional audio simulations.
  • current three-dimensional audio simulation is very limited and often lacks synchronicity with the three-dimensional video simulation.
  • a method for automatically determining a positional three dimensional output of audio information based on a user's orientation within an artificial immersive environment comprising: based on a fixed geometric relationship between a plurality of input audio channels, defining a plurality of unique geometric audio zones in the artificial immersive environment, each geometric audio zone uniquely characterized by a subset of one or more of the input audio channels; receiving data corresponding to the user's orientation within the artificial immersive environment; automatically, for each of a plurality of output channels: based on the user's orientation with respect to the unique geometric audio zones, identifying a respective subset and producing a mix profile for the one or more input audio channels in the subset; and outputting data defining the subset and its corresponding mix profile.
  • a system for automatically determining a positional three dimensional output of audio information based on a user's orientation within an artificial immersive environment comprising: an input interface receiving data corresponding to the user' s orientation within the artificial immersive environment; processing structure defining a plurality of unique geometric audio zones in the artificial immersive environment based on a fixed geometric relationship between a plurality of input audio channels, each geometric audio zone uniquely characterized by a subset of one or more of the input audio channels, the processing structure automatically, for each of a plurality of output channels, identifying a respective subset and establishing a mix profile for the one or more input audio channels in the subset based on the user's orientation with respect to the unique geometric audio zones; and an output interface outputting data defining the subset and its corresponding mix profile.
  • a processor readable medium embodying a computer program for automatically determining a positional three dimensional output of audio information based on a user's orientation within an artificial immersive environment
  • the computer program comprising: program code for, based on a fixed geometric relationship between a plurality of input audio channels, defining a plurality of unique geometric audio zones in the artificial immersive environment, each geometric audio zone uniquely characterized by a subset of one or more of the input audio channels; program code for receiving data corresponding to the user's orientation within the artificial immersive environment; and program code for automatically, for each of a plurality of output channels: based on the user's orientation with respect to the unique geometric audio zones, identifying a respective subset and producing a mix profile for the one or more input audio channels in the subset; and outputting data defining the subset and its corresponding mix profile.
  • Figure 1 is a flowchart depicting steps in a method, according to an embodiment
  • Figure 2 is a schematic diagram of a computing system according to an embodiment
  • Figures 3A, 3B and 3C show top views of different user orientations resulting in different audio outputs for multiple audio output channels, according to an embodiment of the invention
  • Figure 4 shows an example of a fixed geometric relationship between a plurality of input audio channels, defined unique geometric audio zones, and a user orientation corresponding to a particular geometric audio zone;
  • Figures 5A, 5B and 5C show examples of calculating weights for a mix profile based on a degree of geometrical alignment of the user orientation with input audio channels for the corresponding geometric audio zone;
  • Figure 6 is a schematic diagram of a system for automatically determining a positional three dimensional output of audio information based on a user's orientation within an artificial immersive environment, according to an embodiment.
  • the present invention relates to the provision of spatially accurate three-dimensional binaural audio information, which is coordinated with the user's orientation and, in embodiments, location in an artificial immersive environment such as a virtual reality or artificial reality environment.
  • an artificial immersive environment such as a virtual reality or artificial reality environment.
  • the audio output orientation will be matched similarly to the video output orientation.
  • Figure 1 is a flowchart depicting steps in a process 90 for automatically determining a positional three dimensional output of audio information based on a user's orientation within an artificial immersive environment, according to an embodiment.
  • process 90 is executed on one or more systems such as special purpose computing system 1000 shown in Figure 2.
  • Computing system 1000 may also be specially configured with software applications and hardware components to enable a user to author, edit and/or play media such as digital video and audio, as well as to encode, decode and/or transcode the digital video and corresponding audio from and into various formats such as MP4, AVI, MOV, WEBM and using a selected compression algorithm such as H.264 or H. 265 and according to various selected parameters, thereby to compress, decompress, view and/or manipulate the digital video and audio as desired for a particular application, media player, or platform.
  • a selected compression algorithm such as H.264 or H. 265
  • Computing system 1000 may also be configured to enable an author or editor to form multiple copies of a particular digital video, each encoded with a respective bitrate, to facilitate streaming of the same digital video to various downstream users who may have different or time-varying capacities to stream it through adaptive bitrate streaming.
  • Computing system 1000 includes a bus 1010 or other communication mechanism for communicating information, and a processor 1018 coupled with the bus 1010 for processing the information.
  • the computing system 1000 also includes a main memory 1004, such as a random access memory (RAM) or other dynamic storage device (e.g., dynamic RAM (DRAM), static RAM (SRAM), and synchronous DRAM (SDRAM)), coupled to the bus 1010 for storing information and instructions to be executed by processor 1018.
  • main memory 1004 may be used for storing temporary variables or other intermediate information during the execution of instructions by the processor 1018.
  • Processor 1018 may include memory structures such as registers for storing such temporary variables or other intermediate information during execution of instructions.
  • the computing system 1000 further includes a read only memory (ROM) 1006 or other static storage device (e.g., programmable ROM (PROM), erasable PROM (EPROM), and electrically erasable PROM (EEPROM)) coupled to the bus 1010 for storing static information and instructions for the processor 1018.
  • ROM read only memory
  • PROM programmable ROM
  • EPROM erasable PROM
  • EEPROM electrically erasable PROM
  • the computing system 1000 also includes a disk controller 1008 coupled to the bus 1010 to control one or more storage devices for storing information and instructions, such as a magnetic hard disk 1022 and/or a solid state drive (SSD) and/or a flash drive, and a removable media drive 1024 (e.g., solid state drive such as USB key or external hard drive, floppy disk drive, read-only compact disc drive, read/write compact disc drive, compact disc jukebox, tape drive, and removable magneto-optical drive).
  • SSD solid state drive
  • removable media drive 1024 e.g., solid state drive such as USB key or external hard drive, floppy disk drive, read-only compact disc drive, read/write compact disc drive, compact disc jukebox, tape drive, and removable magneto-optical drive.
  • the storage devices may be added to the computing system 1000 using an appropriate device interface (e.g., Serial ATA (SATA), peripheral component interconnect (PCI), small computing system interface (SCSI), integrated device electronics (IDE), enhanced-IDE (E-IDE), direct memory access (DMA), ultra- DMA, as well as cloud-based device interfaces).
  • SATA Serial ATA
  • PCI peripheral component interconnect
  • SCSI small computing system interface
  • IDE integrated device electronics
  • E-IDE enhanced-IDE
  • DMA direct memory access
  • ultra- DMA ultra- DMA
  • the computing system 1000 may also include special purpose logic devices (e.g., application specific integrated circuits (ASICs)) or configurable logic devices (e.g., simple programmable logic devices (SPLDs), complex programmable logic devices (CPLDs), and field programmable gate arrays (FPGAs)).
  • ASICs application specific integrated circuits
  • SPLDs simple programmable logic devices
  • CPLDs complex programmable logic devices
  • FPGAs field programmable gate arrays
  • the computing system 1000 also includes a display controller 1002 coupled to the bus 1010 to control a display 1012, such as an LED (light emitting diode) screen, organic LED (OLED) screen, liquid crystal display (LCD) screen or some other device suitable for displaying information to a computer user, or for controlling data to an external display system.
  • display controller 1002 incorporates a dedicated graphics processing unit (GPU) for processing mainly graphics-intensive or other highly-parallel operations. Such operations may include rendering by applying texturing, shading and the like to wireframe objects including polygons such as spheres and cubes thereby to relieve processor 1018 of having to undertake such intensive operations at the expense of overall performance of computing system 1000.
  • GPU graphics processing unit
  • the GPU may incorporate dedicated graphics memory for storing data generated during its operations, and includes a frame buffer RAM memory for storing processing results as bitmaps to be used to activate pixels of display 1012.
  • the GPU may be instructed to undertake various operations by applications running on computing system 1000 using a graphics-directed application programming interface (API) such as OpenGL, Direct3D and the like.
  • API application programming interface
  • the computing system 1000 includes input/output devices, such as a head mounted display (HMD) having an external display system and associated audio headphones 1016, for interacting with a computer user and providing information such as orientation and/or position information to the processor 1018.
  • input/output devices may be employed, such as those that provide data to the computing system via wires or wirelessly, such as gesture detectors including infrared detectors, gyroscopes, accelerometers, radar/sonar and the like.
  • the computing system 1000 performs a portion or all of the processing steps discussed herein in response to the processor 1018 and/or GPU of display controller 1002 executing one or more sequences of one or more instructions contained in a memory, such as the main memory 1004. Such instructions may be read into the main memory 1004 from another processor readable medium, such as a hard disk 1022 or a removable media drive 1024.
  • processors in a multi -processing arrangement such as computing system 1000 having both a central processing unit and one or more graphics processing unit may also be employed to execute the sequences of instructions contained in main memory 1004 or in dedicated graphics memory of the GPU.
  • hard-wired circuitry may be used in place of or in combination with software instructions.
  • the computing system 1000 includes at least one processor readable medium or memory for holding instructions programmed according to the teachings of the invention and for containing data structures, tables, records, or other data described herein.
  • processor readable media are solid state devices (SSD), flash-based drives, compact discs, hard disks, floppy disks, tape, magneto-optical disks, PROMs (EPROM, EEPROM, flash EPROM), DRAM, SRAM, SDRAM, or any other magnetic medium, compact discs (e.g., CD-ROM), or any other optical medium, punch cards, paper tape, or other physical medium with patterns of holes, a carrier wave (described below), or any other medium from which a computer can read.
  • processor readable media Stored on any one or on a combination of processor readable media, is software for controlling the computing system 1000, for driving a device or devices to perform the functions discussed herein, and for enabling the computing system 1000 to interact with a human user (e.g., digital video author/editor/user).
  • software may include, but is not limited to, device drivers, operating systems, development tools, and applications software.
  • processor readable media further includes the computer program product for performing all or a portion (if processing is distributed) of the processing performed discussed herein.
  • the computer code devices discussed herein may be any interpretable or executable code mechanism, including but not limited to scripts, interpretable programs, dynamic link libraries (DLLs), Java classes, and complete executable programs. Moreover, parts of the processing of the present invention may be distributed for better performance, reliability, and/or cost.
  • a processor readable medium providing instructions to a processor 1018 may take many forms, including but not limited to, non-volatile media, volatile media, and transmission media.
  • Nonvolatile media includes, for example, optical, magnetic disks, and magneto-optical disks, such as the hard disk 1022 or the removable media drive 1024.
  • Volatile media includes dynamic memory, such as the main memory 1004.
  • Transmission media includes coaxial cables, copper wire and fiber optics, including the wires that make up the bus 1010. Transmission media also may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications using various communications protocols.
  • processor readable media may be involved in carrying out one or more sequences of one or more instructions to processor 1018 for execution.
  • the instructions may initially be carried on a magnetic disk of a remote computer or in combination with a processor such as a processor resident on HMD 1014.
  • the remote computer can load the instructions for implementing all or a portion of the present invention remotely and send the instructions or processing results over a wired or wireless connection.
  • a modem local to the computing system 1000 may receive data via wired Ethernet or wirelessly via WiFi and place data on the bus 1010.
  • the bus 1010 carries the data to the main memory 1004, from which the processor 1018 retrieves and executes the instructions.
  • the instructions received by the main memory 1004 may optionally be stored on storage device 1022 or 1024 either before or after execution by processor 1018.
  • the computing system 1000 also includes a communication interface 1020 coupled to the bus 1010.
  • the communication interface 1020 provides a two-way data communication coupling to a network link that is connected to, for example, a local area network (LAN) 1500, or to another communications network 2000 such as the Internet.
  • the communication interface 1020 may be a network interface card to attach to any packet switched LAN.
  • the communication interface 1020 may be an asymmetrical digital subscriber line (ADSL) card, an integrated services digital network (ISDN) card or a modem to provide a data communication connection to a corresponding type of communications line.
  • Wireless links may also be implemented.
  • the communication interface 1020 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.
  • the network link typically provides data communication through one or more networks to other data devices, including without limitation to enable the flow of electronic information.
  • the network link may provide a connection to another computer through a local network 1500 (e.g., a LAN) or through equipment operated by a service provider, which provides communication services through a communications network 2000.
  • the local network 1500 and the communications network 2000 use, for example, electrical, electromagnetic, or optical signals that carry digital data streams, and the associated physical layer (e.g., CAT 5 cable, coaxial cable, optical fiber, etc).
  • the signals through the various networks and the signals on the network link and through the communication interface 1020, which carry the digital data to and from the computing system 1000 may be implemented in baseband signals, or carrier wave based signals.
  • the baseband signals convey the digital data as unmodulated electrical pulses that are descriptive of a stream of digital data bits, where the term "bits" is to be construed broadly to mean symbol, where each symbol conveys at least one or more information bits.
  • the digital data may also be used to modulate a carrier wave, such as with amplitude, phase and/or frequency shift keyed signals that are propagated over a conductive media, or transmitted as electromagnetic waves through a propagation medium.
  • the digital data may be sent as unmodulated baseband data through a "wired" communication channel and/or sent within a predetermined frequency band, different than baseband, by modulating a carrier wave.
  • the computing system 1000 can transmit and receive data, including program code, through the network(s) 1500 and 2000, the network link and the communication interface 1020. Moreover, the network link may provide a connection through a LAN 1500 to a mobile device 1300 such as a personal digital assistant (PDA) laptop computer, or cellular telephone.
  • PDA personal digital assistant
  • Computing system 1000 may be provisioned with or be in communication with live broadcast/streaming equipment that receives and transmits, in near real-time, a stream of digital video and audio content captured in near real-time from a particular live event.
  • the electronic data store implemented in the database or other data structures described herein may be one or more of a table, an array, a database, a structured data file, an XML file, or some other functional data store, such as hard disk 1022 or removable media 1024.
  • each geometric audio zone is characterized by a subset of one or more of the input audio channels.
  • the process 90 proceeds with receiving data corresponding to the user's orientation within the artificial immersive environment (step 200). Automatically, for each of a plurality of output channels, a respective subset is identified (step 300) and a mix profile for the one or more input audio channels in the subset is produced (step 400). Pursuant to this, data defining the subset and its corresponding mix profile are outputted (step 500) for use downstream as will be described.
  • the receiving, identifying, producing and outputting are conducted continuously during use of the artificial immersive environment. That is, as the user is continuously navigating the artificial immersive environment, the data defining the subset and its corresponding mix profile are continuously being updated and outputted according to the user's orientation.
  • processing resources can be efficiently managed such that, only if the user has changed his or her orientation in the artificial immersive environment are the identifying, producing and outputting conducted.
  • the fixed geometric relationship corresponds to the fixed geometric relationship between positions and orientations of microphones in a particular array configuration with respect to an origin.
  • An example configuration is described in United States Patent No. 5,778,083 to Godfrey, the contents of which are incorporated herein by reference.
  • a generally football-shaped frame supports eight (8) microphones distributed laterally around the generally spherical cross-sectioned frame (center front, front left, left, rear left, center rear, rear right, right, and front right) as well as both top and bottom microphones at the top and bottom, respectively, of the frame.
  • the microphones are each oriented such that their diaphragms are oriented outwardly from the frame.
  • the lateral microphones have a hypercardiod pickup pattern while the top and bottom microphones each have a hemispherical pickup pattern, but the positions and orientations of the microphones on the frame form a fixed geometric relationship.
  • the fixed geometric relationship be preserved in some way from capture to playback, or that the fixed geometric relationship at least is synthesized very carefully in a manner that provides spatial integrity. Maintaining or synthesizing the integrity of the spatial positioning of the audio sources provides a very visceral auditory "core" to the listener. When the user re -orients himself or herself in the artificial immersive environment, it is this auditory core that should re -orient accordingly thereby to re -orient all of the corresponding audio information together as one. If no re -orientation of audio information is done, then the orientation of the audio information presented to the user can be perceptually misaligned - severely so in some cases - with the orientation of the video information.
  • a user could see an airplane approaching from the front right, but hear it as though it was approaching from the rear left. Still further, if only a portion of the audio information is re-oriented so as to be inconsistent with the fixed geometric relationship in connection with other audio information, the user may still feel disoriented or, at the least, will not be receiving the visceral benefits of the auditory information being oriented in unison to preserve the auditory core.
  • Data defining the fixed geometric relationship may be encoded along with the input audio information thereby to enable downstream processing to recreate the fixed geometric relationship after decoding.
  • object-based audio produced post-capture
  • producing each mix profile includes calculating a contribution of each of the one or more input audio channels in the subset to the output channel.
  • the mix profile specifies simply that the one input audio channel contributes 100% to the corresponding output audio channel.
  • Table 1 shows an embodiment of a data structure, in this embodiment a two-dimensional matrix, which is appropriate for embodiments where a mix profile is to specify simply that one input audio channel contributes 100% to its corresponding output channel.
  • the two- dimensional matrix stores a sequence of input audio channels according to multiple different user orientations. The sequence itself, constant through the various possible rotational positions available in the matrix, represents the fixed geometric relationship between audio channels distributed laterally such as those described above in the '083 patent.
  • the front left (FL), left (L), left rear (LR), center rear (CR), right rear (RR), right (R) and front right (FR) input audio channels are spatially positioned and oriented in sequence around an ellipse in a counterclockwise direction.
  • the geometric audio zones are defined as the FC zone, the FL zone, the L zone, the LR zone, the CR zone, the RR zone, the R zone and the FR zone.
  • the subset characterizing the FC zone includes only the FC input audio channel
  • the subset characterizing the FL zone includes only the FL audio input channel
  • the subset characterizing the L zone includes only the L audio input channel
  • the subset characterizing the LR zone includes only the LR input audio channel
  • the subset characterizing the CR zone includes only the CR input audio channel
  • the subset characterizing the RR zone includes only the RR input audio channel
  • the subset characterizing the R zone includes only the R input audio channel
  • the subset characterizing the FR zone includes only the FR input audio channel.
  • the user has re -oriented to face in a direction corresponding to the FL zone.
  • This orientation information is employed to, for each output audio channel, identify a respective subset and produce a mix profile.
  • the FL zone corresponds to the user's orientation so the FL input audio channel is now to contribute 100% to the FC output audio channel, the L input audio channel is now to contribute 100% to the FL output audio channel, the LR input audio channel is now to contribute 100% to the L output audio channel, and so forth.
  • the embodiment described above in connection with Table 1 is illustrative of principles of the invention, where one input audio channel contributes fully (weighted to 100%) to its corresponding output audio channel.
  • the example embodiment provides for one-to-one re-assignment (or routing) of audio input channels to particular output audio channels for a discrete number of positions (8 positions) corresponding to spatial locations of audio channels, based on the orientation of a user in an artificial immersive environment. Due to the preserved sequence, as described above, the audio core is preserved as the user re -orients in the artificial immersive environment.
  • any Top or Bottom input audio channel is to be simply routed to the Top or Bottom output audio channels accordingly. That is, in the straightforward example above, only certain changes in yaw of a user's orientation, and not changes in pitch or roll, result in re -orientation of the audio information. However, in a practical system, it would be valuable to provide audio orientation changes when a user changed pitch or roll. For this reason, a third dimension added to the two- dimensional matrix depicted above would enable encoding of the corresponding fixed geometrical structure that account for top and bottom channels in various ways could be used. As would be understood however, unless the matrix is very large, there is likely to be orientations that would not have exact predefined counterparts in the matrix.
  • additional processing may be provided to derive output audio information from multiple input audio channels rather than simply allocate a full input channel to a full output channel, and to smooth changes between zones specified in the matrix so that the resultant audio information does not present clicks or sharp changes across zone boundaries.
  • Figure 4 shows a configuration for audio capture that preserves a fixed geometric relationship between audio input channels T (Top), FC, FL, L, RL and CR as well as those not shown in Figure 4: audio input channels RR, R, FR and B (Bottom).
  • a plurality of unique geometric zones is each defined to be characterized by a subset of three (3) input audio channels.
  • the zones are defined as shown in Table 2, below.
  • an Orientation is shown that represents the orientation of the user in the artificial immersive environment. It will be noted that, despite the appearance as drawn in Figure 4, the origin for the Orientation is not the same location as the L input audio channel. Rather, in this embodiment, the origin corresponds with the geometric center about which the geometric audio zones are positioned. [0054] As shown, the Orientation in Figure 4 does not align with any of spatial locations of the input channels T, FC, FL, L, RL, CR, RR, R, FR or B. As such, in this embodiment, the mix profile for each output audio channel will include non-zero values corresponding to each of the input audio channels for the corresponding subset, with the values summing to 100%.
  • the Orientation corresponds with Zone A.
  • the subset of input audio channels that are to be combined to produce the output audio channel for FC is T-FC-FL.
  • the corresponding mix profile could simply be an even mix of all three input audio channels (33%, 33%, 33%), it may be advantageous to further refine the mix profile for improved realism according to the proximity of Orientation to each of the three input audio channels.
  • Orientation while there is no alignment between the Orientation and any one of input audio channels T, FC and FL, Orientation differs least in alignment in Zone A from input audio channel FC, and differs most in alignment in Zone A from input audio channel T.
  • the technique for weighting the various contributions to an output signal for an output channel may vary based on audio principles such as differences in pickup patterns of microphones (such as hypercardiod versus hemispherical), differences in perceptions of different frequencies, and differences in human perception of sound approaching the ears laterally versus from above or below, as well as other factors. However, these factors could also be codified or otherwise processed and accounted for in the relative weights in the mix profiles as the user navigates in the artificial geometric environment.
  • the audio channel outputs from the from the mixer/blender, containing audio information produced as described above, may then be directed to a respective one of a plurality of corresponding audio output devices, where each of the audio output devices is positioned and oriented with respect to a user position according to the fixed geometric relationship.
  • the audio channel outputs may be directed to an audio virtualizer.
  • An audio virtualizer generally "spreads" a plurality of output channels across a software- or hardware-implemented geometric structure such as the inner walls of a sphere or the inner walls of a sound booth or the inner walls of a concert hall (as desired), such that each of a plurality of locations (typically many more than there are actual individual output audio channels) on the surface of the geometric structure can be notionally considered a point source of audio.
  • the virtualized audio can more readily be presented for output on downstream audio output devices of various configurations instead of the particular configuration used to capture the audio information in the first place.
  • the virtualized audio information may be directed to the inputs of both left and right head related transfer functions (HRTFs), either implemented in software or hardware or a combination thereof, the outputs of which are conveyed to respective earpieces of headphones or earbuds or the like, or other combinations of left and right audio output devices directed at the user's ears.
  • HRTFs head related transfer functions
  • additional information about user's navigation in the artificial immersive environment may be available.
  • an HMD system for virtual reality gaming and the like may offer information in six degrees of freedom (6-DOF), including information about translation in X, Y, and Z dimensions in addition to the information about yaw, pitch and roll described above.
  • 6-DOF degrees of freedom
  • data is received corresponding to the user's position within the artificial immersive environment.
  • beamforming of one or more of the output audio channels is applied or modified in order to enable a user to receive more sound in the direction of change of the user's position.
  • amplitude of one or more output audio channels corresponding to a direction of change in the user's position is increased, and amplitude of one or more output audio channels corresponding to the opposite of the direction of chance in the user's position is decreased.
  • This feature may be enabled in an implementation of an artificial immersive environment where 6-DOF is tracked by user equipment (such as an HMD or accompanying device) but may simply be disabled in an implementation of an artificial immersive environment where only 3-DOF is tracked by the user equipment.
  • the methods described herein may be implemented where input audio channels have never been encoded, but are more likely to be implemented where input audio channels have been decoded from one or more sources of either pre-recorded audio information or streaming audio information. Furthermore, the methods described herein may be implemented where one or more of the input audio channels have been synthesized rather than captured using an audio sensor, and the synthesized one or more channels have been oriented according to a fixed geometric relationship with each other and with any other input channels.
  • FIG. 6 is schematic diagram of a system 2000 for automatically determining a positional three dimensional output of audio information based on a user' s orientation within an artificial immersive environment.
  • System 2000 may be implemented in software, hardware or a combination of both and sharing or using resources provided by system 1000 in Figure 2.
  • the system 2000 includes an input interface 2002 receiving data corresponding to the user's orientation within the artificial immersive environment.
  • Processing structure is configured to defining a plurality of unique geometric audio zones 2004 in the artificial immersive environment based on a fixed geometric relationship between a plurality of input audio channels 2500. As described above, each geometric audio zone is uniquely characterized by a subset of one or more of the input audio channels 2500.
  • the processing structure is also configured to automatically, for each of a plurality of output audio channels 2600, identify a respective subset 2006 and establish a mix profile 2008 for the one or more input audio channels 2500 in the subset 2006 based on the user's orientation with respect to the unique geometric audio zones.
  • An output interface 2010 outputs data 2012 defining the subset and its corresponding mix profile for each output audio channel.
  • the data 2012 may be routed to a mixer/blender 2014 that also receives input audio information via the audio input channels 2500.
  • the mixer/blender produces output audio information 2600 with the input audio information and the data 2012.
  • the output audio information can be conveyed to one or more downstream systems, such as a speaker system 2700 or a virtualizer/spatializer 2800.
  • outputs of the virtualizer/spatializer 2800 may be fed to two head related transfer functions 2900R and 2900L, the outputs of each of which are, in turn, conveyed to respective speakers of headphones.
  • the input interface 2002 also optionally receives data corresponding to the user's position within the artificial immersive environment; and provides that position data to a beamformer 2009 to apply or modifying beamforming for one or more of the output audio channels in response to changes in the user's position.

Landscapes

  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Stereophonic System (AREA)

Abstract

La présente invention concerne un procédé, un système et un produit de programme informatique permettant de déterminer automatiquement une sortie tridimensionnelle de position d'informations audio sur la base de l'orientation d'un utilisateur dans un environnement immersif artificiel. Le procédé consiste, sur la base d'une relation géométrique fixée entre une pluralité de canaux audio d'entrée, à définir une pluralité de zones audios géométriques uniques dans l'environnement immersif artificiel, chaque zone audio géométrique unique caractérisée par un sous-ensemble d'un ou plusieurs des canaux audios d'entrée ; à recevoir des données correspondantes à l'orientation de l'utilisateur dans l'environnement immersif artificiel ; automatiquement, pour chacun de la pluralité des canaux de sortie : sur la base de l'orientation de l'utilisateur par rapport aux zones géométriques audios uniques, à identifier un sous-ensemble respectif et à produire un profil de mixage pour lesdits canaux audios d'entrées dans le sous-ensemble ; et à produire des données de sortie définissant le sous-ensemble et son profil de mixage correspondant.
PCT/CA2017/050050 2016-01-15 2017-01-16 Procédé et système pour déterminer automatiquement une sortie tridimensionnelle de position d'informations audio sur la base de l'orientation d'un utilisateur dans un environnement immersif artificiel. Ceased WO2017120681A1 (fr)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201662279140P 2016-01-15 2016-01-15
US62/279,140 2016-01-15

Publications (1)

Publication Number Publication Date
WO2017120681A1 true WO2017120681A1 (fr) 2017-07-20

Family

ID=59310500

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CA2017/050050 Ceased WO2017120681A1 (fr) 2016-01-15 2017-01-16 Procédé et système pour déterminer automatiquement une sortie tridimensionnelle de position d'informations audio sur la base de l'orientation d'un utilisateur dans un environnement immersif artificiel.

Country Status (1)

Country Link
WO (1) WO2017120681A1 (fr)

Cited By (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10165386B2 (en) 2017-05-16 2018-12-25 Nokia Technologies Oy VR audio superzoom
WO2019072984A1 (fr) * 2017-10-12 2019-04-18 Ffraunhofer-Gesellschaft Zur Förderung Der Angewandten Forschung E.V. Optimisation de diffusion audio pour applications de réalité virtuelle
US20190306651A1 (en) 2018-03-27 2019-10-03 Nokia Technologies Oy Audio Content Modification for Playback Audio
US10531219B2 (en) 2017-03-20 2020-01-07 Nokia Technologies Oy Smooth rendering of overlapping audio-object interactions
US11074036B2 (en) 2017-05-05 2021-07-27 Nokia Technologies Oy Metadata-free audio-object interactions
US11096004B2 (en) 2017-01-23 2021-08-17 Nokia Technologies Oy Spatial audio rendering point extension
US11109178B2 (en) 2017-12-18 2021-08-31 Dolby International Ab Method and system for handling local transitions between listening positions in a virtual reality environment
US11395087B2 (en) 2017-09-29 2022-07-19 Nokia Technologies Oy Level-based audio-object interactions
CN115499772A (zh) * 2022-08-08 2022-12-20 深圳感臻智能股份有限公司 一种声道变换方法及装置
RU2801698C2 (ru) * 2017-10-12 2023-08-14 Фраунхофер-Гезелльшафт Цур Фердерунг Дер Ангевандтен Форшунг Е.Ф. Оптимизация доставки звука для приложений виртуальной реальности

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1296155B1 (fr) * 2001-09-25 2006-11-22 Symbol Technologies, Inc. Système de localisation d'un objet avec une balise acoustique et méthode correspondante
WO2010140088A1 (fr) * 2009-06-03 2010-12-09 Koninklijke Philips Electronics N.V. Estimation de positions de haut-parleur
US8767968B2 (en) * 2010-10-13 2014-07-01 Microsoft Corporation System and method for high-precision 3-dimensional audio for augmented reality

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP1296155B1 (fr) * 2001-09-25 2006-11-22 Symbol Technologies, Inc. Système de localisation d'un objet avec une balise acoustique et méthode correspondante
WO2010140088A1 (fr) * 2009-06-03 2010-12-09 Koninklijke Philips Electronics N.V. Estimation de positions de haut-parleur
US8767968B2 (en) * 2010-10-13 2014-07-01 Microsoft Corporation System and method for high-precision 3-dimensional audio for augmented reality

Cited By (28)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US12538089B2 (en) 2017-01-23 2026-01-27 Nokia Technologies Oy Spatial audio rendering point extension
US11096004B2 (en) 2017-01-23 2021-08-17 Nokia Technologies Oy Spatial audio rendering point extension
US10531219B2 (en) 2017-03-20 2020-01-07 Nokia Technologies Oy Smooth rendering of overlapping audio-object interactions
US11044570B2 (en) 2017-03-20 2021-06-22 Nokia Technologies Oy Overlapping audio-object interactions
US11074036B2 (en) 2017-05-05 2021-07-27 Nokia Technologies Oy Metadata-free audio-object interactions
US11604624B2 (en) 2017-05-05 2023-03-14 Nokia Technologies Oy Metadata-free audio-object interactions
US11442693B2 (en) 2017-05-05 2022-09-13 Nokia Technologies Oy Metadata-free audio-object interactions
US10165386B2 (en) 2017-05-16 2018-12-25 Nokia Technologies Oy VR audio superzoom
US11395087B2 (en) 2017-09-29 2022-07-19 Nokia Technologies Oy Level-based audio-object interactions
EP4329319A3 (fr) * 2017-10-12 2024-04-24 Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. Optimisation de distribution audio pour applications de réalité virtuelle
KR102568373B1 (ko) * 2017-10-12 2023-08-18 프라운 호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. 가상 현실 애플리케이션들에 대한 오디오 전달의 최적화
WO2019072984A1 (fr) * 2017-10-12 2019-04-18 Ffraunhofer-Gesellschaft Zur Förderung Der Angewandten Forschung E.V. Optimisation de diffusion audio pour applications de réalité virtuelle
RU2801698C2 (ru) * 2017-10-12 2023-08-14 Фраунхофер-Гезелльшафт Цур Фердерунг Дер Ангевандтен Форшунг Е.Ф. Оптимизация доставки звука для приложений виртуальной реальности
US11354084B2 (en) 2017-10-12 2022-06-07 Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. Optimizing audio delivery for virtual reality applications
KR20200078537A (ko) * 2017-10-12 2020-07-01 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. 가상 현실 애플리케이션들에 대한 오디오 전달의 최적화
US12411648B2 (en) 2017-10-12 2025-09-09 Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V. Optimizing audio delivery for virtual reality applications
KR102774081B1 (ko) 2017-10-12 2025-02-26 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. 가상 현실 애플리케이션들에 대한 오디오 전달의 최적화
TWI713911B (zh) * 2017-10-12 2020-12-21 弗勞恩霍夫爾協會 用於虛擬實境應用之音訊遞送最佳化技術
KR20240137132A (ko) * 2017-10-12 2024-09-19 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. 가상 현실 애플리케이션들에 대한 오디오 전달의 최적화
RU2765569C1 (ru) * 2017-10-12 2022-02-01 Фраунхофер-Гезелльшафт Цур Фердерунг Дер Ангевандтен Форшунг Е.Ф. Оптимизация доставки звука для приложений виртуальной реальности
KR20230130729A (ko) * 2017-10-12 2023-09-12 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. 가상 현실 애플리케이션들에 대한 오디오 전달의 최적화
RU2750505C1 (ru) * 2017-10-12 2021-06-29 Фраунхофер-Гезелльшафт Цур Фердерунг Дер Ангевандтен Форшунг Е.Ф. Оптимизация доставки звука для приложений виртуальной реальности
KR102707356B1 (ko) * 2017-10-12 2024-09-13 프라운호퍼 게젤샤프트 쭈르 푀르데룽 데어 안겐반텐 포르슝 에. 베. 가상 현실 애플리케이션들에 대한 오디오 전달의 최적화
US12238506B2 (en) 2017-12-18 2025-02-25 Dolby International Ab Method and system for handling local transitions between listening positions in a virtual reality environment
US11109178B2 (en) 2017-12-18 2021-08-31 Dolby International Ab Method and system for handling local transitions between listening positions in a virtual reality environment
US20190306651A1 (en) 2018-03-27 2019-10-03 Nokia Technologies Oy Audio Content Modification for Playback Audio
US10542368B2 (en) 2018-03-27 2020-01-21 Nokia Technologies Oy Audio content modification for playback audio
CN115499772A (zh) * 2022-08-08 2022-12-20 深圳感臻智能股份有限公司 一种声道变换方法及装置

Similar Documents

Publication Publication Date Title
US11832086B2 (en) Spatial audio downmixing
US10952009B2 (en) Audio parallax for virtual reality, augmented reality, and mixed reality
CN106993249B (zh) 一种声场的音频数据的处理方法及装置
US11055057B2 (en) Apparatus and associated methods in the field of virtual reality
EP2954703B1 (fr) Détermination de dispositifs de restitution pour des coefficients d'harmoniques sphériques
US10993067B2 (en) Apparatus and associated methods
JP2021535632A (ja) オーディオ信号の処理用の方法及び装置
CN114424587A (zh) 控制音频数据的呈现
US12204815B2 (en) Adaptive audio delivery and rendering
CN105264914A (zh) 音频再生装置以及方法
EP3506080B1 (fr) Traitement de scène audio
WO2021003351A1 (fr) Adaptation de flux audio pour obtenir un rendu
WO2018121524A1 (fr) Procédé et dispositif de traitement de données, dispositif d'acquisition et support de stockage
CN114072792A (zh) 用于音频渲染的基于密码的授权
WO2019057530A1 (fr) Appareil et procédés associés pour la présentation d'audio sous la forme d'audio spatial
Villegas Locating virtual sound sources at arbitrary distances in real-time binaural reproduction
US20200267492A1 (en) An Apparatus and Associated Methods for Presentation of a Bird's Eye View
Geronazzo et al. HOBA-VR: HRTF On Demand for Binaural Audio in immersive virtual reality environments
US20230059715A1 (en) Immersive media interoperability
US20240406669A1 (en) Metadata for Spatial Audio Rendering
CN114128312A (zh) 用于低频效果的音频渲染
US12137336B2 (en) Immersive media compatibility
US12531077B2 (en) Method and apparatus in audio processing
US20240406660A1 (en) Spatial Audio Rendering with Listener Motion Compensation using Metadata
US20240406661A1 (en) Spatial Audio Rendering with Listener Motion Compensation using Metadata

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 17738069

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 17738069

Country of ref document: EP

Kind code of ref document: A1