WO2024253828A1 - Techniques pour environnements tridimensionnels - Google Patents
Techniques pour environnements tridimensionnels Download PDFInfo
- Publication number
- WO2024253828A1 WO2024253828A1 PCT/US2024/030181 US2024030181W WO2024253828A1 WO 2024253828 A1 WO2024253828 A1 WO 2024253828A1 US 2024030181 W US2024030181 W US 2024030181W WO 2024253828 A1 WO2024253828 A1 WO 2024253828A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- entity
- input
- environment
- representation
- type
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T19/00—Manipulating three-dimensional [3D] models or images for computer graphics
- G06T19/20—Editing of three-dimensional [3D] images, e.g. changing shapes or colours, aligning objects or positioning parts
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
- G06F3/013—Eye tracking input arrangements
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/017—Gesture based interaction, e.g. based on a set of recognized hand gestures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/03—Arrangements for converting the position or the displacement of a member into a coded form
- G06F3/0304—Detection arrangements using opto-electronic means
Definitions
- Some techniques are described herein for integrating a 2D framework with a 3D framework. Such techniques use a concept referred to as a hidden entity to link the two frameworks together. Other techniques are described herein for translating gestures from a first type to a second type in certain situations. [0005] In some examples, a method that is performed by a computer system is described.
- the method comprises: receiving, from an application, a request to add a two-dimensional entity at a first location in a three-dimensional environment, wherein the 3D environment includes one or more 3D entities; and in response to receiving the request to add the 2D entity at the first location in the 3D environment: adding a first 3D entity to the first location in the 3D environment; rendering, via a 2D framework, a representation of the 2D entity; and rendering, via a 3D framework, a representation of a second 3D entity by: performing one or more operations on the representation of the 2D entity; and placing the representation of the second 3D entity at the first location.
- a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system.
- the one or more programs includes instructions for: receiving, from an application, a request to add a two-dimensional entity at a first location in a three-dimensional environment, wherein the 3D environment includes one or more 3D entities; and in response to receiving the request to add the 2D entity at the first location in the 3D environment: adding a first 3D entity to the first location in the 3D environment; rendering, via a 2D framework, a representation of the 2D entity; and rendering, via a 3D framework, a representation of a second 3D entity by: performing one or more operations on the representation of the 2D entity; and placing the representation of the second 3D entity at the first location.
- a transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system.
- the one or more programs includes instructions for: receiving, from an application, a request to add a two-dimensional entity at a first location in a three- dimensional environment, wherein the 3D environment includes one or more 3D entities; and in response to receiving the request to add the 2D entity at the first location in the 3D environment: adding a first 3D entity to the first location in the 3D environment; rendering, via a 2D framework, a representation of the 2D entity; and rendering, via a 3D framework, a representation of a second 3D entity by: performing one or more operations on the representation of the 2D entity; and placing the representation of the second 3D entity at the first location.
- a computer system comprising one or more processors and memory storing one or more programs configured to be executed by the one or more processors.
- the one or more programs includes instructions for: receiving, from an application, a request to add a two-dimensional entity at a first location in a three-dimensional environment, wherein the 3D environment includes one or more 3D entities; and in response to receiving the request to add the 2D entity at the first location in the 3D environment: adding a first 3D entity to the first location in the 3D environment; rendering, via a 2D framework, a representation of the 2D entity; and rendering, via a 3D framework, a representation of a second 3D entity by: performing one or more operations on the representation of the 2D entity; and placing the representation of the second 3D entity at the first location.
- a computer system is comprising means for performing each of the following steps: receiving, from an application, a request to add a two-dimensional entity at a first location in a three-dimensional environment, wherein the 3D environment includes one or more 3D entities; and in response to receiving the request to add the 2D entity at the first location in the 3D environment: adding a first 3D entity to the first location in the 3D environment; rendering, via a 2D framework, a representation of the 2D entity; and rendering, via a 3D framework, a representation of a second 3D entity by: performing one or more operations on the representation of the 2D entity; and placing the representation of the second 3D entity at the first location.
- a computer program product comprises one or more programs configured to be executed by one or more processors of a computer system.
- the one or more programs include instructions for: receiving, from an application, a request to add a two-dimensional entity at a first location in a three-dimensional environment, wherein the 3D environment includes one or more 3D entities; and in response to receiving the request to add the 2D entity at the first location in the 3D environment: adding a first 3D entity to the first location in the 3D environment; rendering, via a 2D framework, a representation of the 2D entity; and rendering, via a 3D framework, a representation of a second 3D entity by: performing one or more operations on the representation of the 2D entity; and placing the representation of the second 3D entity at the first location.
- a method that is performed at a computer system in communication with one or more input devices comprises: detecting, via the one or more input devices, a first input corresponding to a three- dimensional environment; in response to detecting the first input corresponding to the 3D environment: in accordance with a determination that a first set of one or more criteria is satisfied, wherein the first set of one or more criteria includes a criterion that is satisfied when the input is directed to a first entity of a first type in the 3D environment: translating the first input to a second input different from the first input; and sending, to a first application, an indication of the second input; and in accordance with a determination that a second set of one or more criteria is satisfied, wherein the second set of one or more criteria includes a criterion that is satisfied when the input is directed to a second entity of a second type in the 3D environment, sending, to a second application, an indication of the first input, wherein the second type of
- a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more input devices.
- the one or more programs includes instructions for: detecting, via the one or more input devices, a first input corresponding to a three-dimensional environment; in response to detecting the first input corresponding to the 3D environment: in accordance with a determination that a first set of one or more criteria is satisfied, wherein the first set of one or more criteria includes a criterion that is satisfied when the input is directed to a first entity of a first type in the 3D environment: translating the first input to a second input different from the first input; and sending, to a first application, an indication of the second input; and in accordance with a determination that a second set of one or more criteria is satisfied, wherein the second set of one or more criteria includes a criterion that is satisfied when the input is directed to a second entity of a second type in
- a transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computer system in communication with one or more input devices.
- the one or more programs includes instructions for: detecting, via the one or more input devices, a first input corresponding to a three-dimensional environment; in response to detecting the first input corresponding to the 3D environment: in accordance with a determination that a first set of one or more criteria is satisfied, wherein the first set of one or more criteria includes a criterion that is satisfied when the input is directed to a first entity of a first type in the 3D environment: translating the first input to a second input different from the first input; and sending, to a first application, an indication of the second input; and in accordance with a determination that a second set of one or more criteria is satisfied, wherein the second set of one or more criteria includes a criterion that is satisfied when the input is directed to a second entity of a second type in the 3
- a computer system in communication with one or more input devices comprises one or more processors and memory storing one or more programs configured to be executed by the one or more processors.
- the one or more programs includes instructions for: detecting, via the one or more input devices, a first input corresponding to a three-dimensional environment; in response to detecting the first input corresponding to the 3D environment: in accordance with a determination that a first set of one or more criteria is satisfied, wherein the first set of one or more criteria includes a criterion that is satisfied when the input is directed to a first entity of a first type in the 3D environment: translating the first input to a second input different from the first input; and sending, to a first application, an indication of the second input; and in accordance with a determination that a second set of one or more criteria is satisfied, wherein the second set of one or more criteria includes a criterion that is satisfied when the input is directed to a second entity of a second type in the 3D environment, sending, to a second application, an indication of the first input, wherein the second type of entity is different from the first type of entity, and wherein the second set of one or more criteria is different from the first
- a computer system in communication with one or more input devices comprises means for performing each of the following steps: detecting, via the one or more input devices, a first input corresponding to a three-dimensional environment; in response to detecting the first input corresponding to the 3D environment: in accordance with a determination that a first set of one or more criteria is satisfied, wherein the first set of one or more criteria includes a criterion that is satisfied when the input is directed to a first entity of a first type in the 3D environment: translating the first input to a second input different from the first input; and sending, to a first application, an indication of the second input; and in accordance with a determination that a second set of one or more criteria is satisfied, wherein the second set of one or more criteria includes a criterion that is satisfied when the input is directed to a second entity of a second type in the 3D environment, sending, to a second application, an indication of the
- a computer program product comprises one or more programs configured to be executed by one or more processors of a computer system in communication with one or more input devices.
- the one or more programs include instructions for: detecting, via the one or more input devices, a first input corresponding to a three-dimensional environment; in response to detecting the first input corresponding to the 3D environment: in accordance with a determination that a first set of one or more criteria is satisfied, wherein the first set of one or more criteria includes a criterion that is satisfied when the input is directed to a first entity of a first type in the 3D environment: translating the first input to a second input different from the first input; and sending, to a first application, an indication of the second input; and in accordance with a determination that a second set of one or more criteria is satisfied, wherein the second set of one or more criteria includes a criterion that is satisfied when the input is directed to a second entity of a second type
- Executable instructions for performing these functions are, optionally, included in a non-transitory computer-readable storage medium or other computer program product configured for execution by one or more processors. Executable instructions for performing these functions are, optionally, included in a transitory computer-readable storage medium or other computer program product configured for execution by one or more processors.
- FIG. 1 illustrates an example system architecture including various electronic devices that may implement the subject system in accordance with some examples.
- FIG. 2 illustrates a block diagram of example features of an electronic device in accordance with some examples.
- FIG. 3 is a block diagram illustrating a computer system in accordance with some examples.
- FIG. 4 is a block diagram of a scene graph in accordance with some examples.
- FIG. 5 is a block diagram illustrating a process for processing gestures in a 3D environment according to some examples.
- FIG. 6 is a flow diagram illustrating a method for integrating a 2D framework with a 3D framework in accordance with some examples.
- FIG. 7 is a flow diagram illustrating a method for translating between gestures in accordance with some examples.
- FIG. 8 illustrates an electronic system with which some examples of the subject technology may be implemented.
- Methods and/or processes described herein can include one or more steps that are contingent upon one or more conditions being satisfied. It should be understood that a method can occur over multiple iterations of the same process with different steps of the method being satisfied in different iterations. For example, if a method requires performing a first step upon a determination that a set of one or more criteria is met and a second step upon a determination that the set of one or more criteria is not met, a person of ordinary skill in the art would appreciate that the steps of the method are repeated until both conditions, in no particular order, are satisfied. Thus, a method described with steps that are contingent upon a condition being satisfied can be rewritten as a method that is repeated until each of the conditions described in the method are satisfied.
- system or computer readable medium claims include instructions for performing one or more steps that are contingent upon one or more conditions being satisfied. Because the instructions for the system or computer readable medium claims are stored in one or more processors and/or at one or more memory locations, the system or computer readable medium claims include logic that can determine whether the one or more conditions have been satisfied without explicitly repeating steps of a method until all of the conditions upon which steps in the method are contingent have been satisfied. A person having ordinary skill in the art would also understand that, similar to a method with contingent steps, a system or computer readable storage medium can repeat the steps of a method as many times as needed to ensure that all of the contingent steps have been performed.
- first used to distinguish one element from another.
- a first subsystem could be termed a second subsystem, and, similarly, a subsystem device could be termed a subsystem device, without departing from the scope of the various described examples.
- the first subsystem and the second subsystem are two separate references to the same subsystem.
- the first subsystem and the second subsystem are both subsystems, but they are not the same subsystem or the same type of subsystem.
- the term “if’ is, optionally, construed to mean “when,” “upon,” “in response to determining,” “in response to detecting,” or “in accordance with a determination that” depending on the context.
- the phrase “if it is determined” or “if [a stated condition or event] is detected” is, optionally, construed to mean “upon determining,” “in response to determining,” “upon detecting [the stated condition or event],” “in response to detecting [the stated condition or event],” or “in accordance with a determination that [the stated condition or event]” depending on the context.
- a physical environment refers to a physical world that people can sense and/or interact with without aid of electronic devices.
- the physical environment may include physical features such as a physical surface or a physical object.
- the physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People can directly sense and/or interact with the physical environment such as through sight, touch, hearing, taste, and smell.
- an extended reality (XR) environment refers to a wholly or partially simulated environment that people sense and/or interact with via an electronic device.
- the XR environment may include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, and/or the like.
- an XR system With an XR system, a subset of a person’s physical motions, or representations thereof, are tracked, and, in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that comports with at least one law of physics.
- the XR system may detect head movement and, in response, adjust graphical content and an acoustic field presented to the person in a manner similar to how such views and sounds would change in a physical environment.
- the XR system may detect movement of the electronic device presenting the XR environment (e.g., a mobile phone, a tablet, a laptop, or the like) and, in response, adjust graphical content and an acoustic field presented to the person in a manner similar to how such views and sounds would change in a physical environment.
- the XR system may adjust character! stic(s) of graphical content in the XR environment in response to representations of physical motions (e.g., vocal commands).
- a head mountable system may have one or more speaker(s) and an integrated opaque display.
- a head mountable system may be configured to accept an external opaque display (e.g., a smartphone).
- the head mountable system may incorporate one or more imaging sensors to capture images or video of the physical environment, and/or one or more microphones to capture audio of the physical environment.
- a head mountable system may have a transparent or translucent display.
- the transparent or translucent display may have a medium through which light representative of images is directed to a person’s eyes.
- the display may utilize digital light projection, OLEDs, LEDs, uLEDs, liquid crystal on silicon, laser scanning light source, or any combination of these technologies.
- the medium may be an optical waveguide, a hologram medium, an optical combiner, an optical reflector, or any combination thereof.
- the transparent or translucent display may be configured to become opaque selectively.
- Projection-based systems may employ retinal projection technology that projects graphical images onto a person’s retina. Projection systems also may be configured to project virtual objects into the physical environment, for example, as a hologram or on a physical surface.
- FIG. 1 illustrates an example system architecture 100 including various electronic devices that may implement the subject system in accordance with some examples. Not all of the depicted components may be used in all examples, however, and some examples may include additional or different components than those shown in the figure. Variations in the arrangement and type of the components may be made without departing from the spirit or scope of the claims as set forth herein. Additional components, different components, or fewer components may be provided.
- the system architecture 100 includes an electronic device 105, a handheld electronic device 104, an electronic device 110, an electronic device 115, and a server 120.
- the system architecture 100 is illustrated in FIG. 1 as including the electronic device 105, the handheld electronic device 104, the electronic device 110, the electronic device 115, and the server 120; however, the system architecture 100 may include any number of electronic devices, and any number of servers or a data center including multiple servers.
- the electronic device 105 may be implemented, for example, as a tablet device, a smartphone, or as a head mountable portable system (e.g., worn by a user 101).
- the electronic device 105 includes a display system capable of presenting a visualization of an extended reality environment to the user.
- the electronic device 105 may be powered with a battery and/or another power supply.
- the display system of the electronic device 105 provides a stereoscopic presentation of the extended reality environment, enabling a three-dimensional visual display of a rendering of a particular scene, to the user.
- the user may use a handheld electronic device 104, such as a tablet, watch, mobile device, and the like.
- the electronic device 105 may include one or more cameras such as camera(s) 150 (e.g., visible light cameras, infrared cameras, etc.).
- the electronic device 105 may include multiple cameras 150.
- the multiple cameras 150 may include a left facing camera, a front facing camera, a right facing camera, a down facing camera, a leftdown facing camera, a right-down facing camera, an up facing camera, one or more eye- facing cameras, and/or other cameras.
- Each of the cameras 150 may include one or more image sensors (e.g., charged coupled device (CCD) image sensors, complementary metal oxide semiconductor (CMOS) image sensors, or the like).
- CCD charged coupled device
- CMOS complementary metal oxide semiconductor
- the electronic device 105 may include various sensors 152 including, but not limited to, other cameras, other image sensors, touch sensors, microphones, inertial measurement units (IMU), heart rate sensors, temperature sensors, depth sensors (e.g., Lidar sensors, radar sensors, sonar sensors, time-of-flight sensors, etc.), GPS sensors, Wi-Fi sensors, near-field communications sensors, radio frequency sensors, etc.
- the electronic device 105 may include hardware elements that can receive user input such as hardware buttons or switches. User inputs detected by such cameras, sensors, and/or hardware elements may correspond to, for example, various input modalities.
- such input modalities may include, but are not limited to, facial tracking, eye tracking (e.g., gaze direction), hand tracking, gesture tracking, biometric readings (e.g., heart rate, pulse, pupil dilation, breath, temperature, electroencephalogram, olfactory), recognizing speech or audio (e.g., particular hotwords), and activating buttons or switches, etc.
- facial tracking, gaze tracking, hand tracking, gesture tracking, object tracking, and/or physical environment mapping processes may utilize images (e.g., image frames) captured by one or more image sensors of the cameras 150 and/or the sensors 152.
- the electronic device 105 may be communicatively coupled to a base device such as the electronic device 110 and/or the electronic device 115.
- a base device may, in general, include more computing resources and/or available power in comparison with the electronic device 105.
- the electronic device 105 may operate in various modes. For instance, the electronic device 105 can operate in a standalone mode independent of any base device. When the electronic device 105 operates in the standalone mode, the number of input modalities may be constrained by power and/or processing limitations of the electronic device 105 such as available battery power of the device. In response to power limitations, the electronic device 105 may deactivate certain sensors within the device itself to preserve battery power and/or to free processing resources.
- the electronic device 105 may also operate in a wireless tethered mode (e.g., connected via a wireless connection with a base device), working in conjunction with a given base device.
- the electronic device 105 may also work in a connected mode where the electronic device 105 is physically connected to a base device (e.g., via a cable or some other physical connector) and may utilize power resources provided by the base device (e.g., where the base device is charging the electronic device 105 and/or providing power to the electronic device 105 while physically connected).
- the electronic device 105 When the electronic device 105 operates in the wireless tethered mode or the connected mode, a least a portion of processing user inputs and/or rendering the extended reality environment may be offloaded to the base device thereby reducing processing burdens on the electronic device 105.
- the electronic device 105 works in conjunction with the electronic device 110 or the electronic device 115 to generate an extended reality environment including physical and/or virtual objects that enables different forms of interaction (e.g., visual, auditory, and/or physical or tactile interaction) between the user and the generated extended reality environment in a real-time manner.
- the electronic device 105 provides a rendering of a scene corresponding to the extended reality environment that can be perceived by the user and interacted with in a real-time manner, such as a host environment for a group session with another user. Additionally, as part of presenting the rendered scene, the electronic device 105 may provide sound, and/or haptic or tactile feedback to the user.
- the content of a given rendered scene may be dependent on available processing capability, network availability and capacity, available battery power, and current system workload.
- the electronic device 105 may be, and/or may include all or part of, the electronic system discussed below with respect to FIG. 8.
- the network 106 may communicatively (directly or indirectly) couple, for example, the electronic device 105, the electronic device 110, and/or the electronic device 115 with each other device and/or the server 120.
- the network 106 may be an interconnected network of devices that may include, or may be communicatively coupled to, the Internet.
- the handheld electronic device 104 may be, for example, a smartphone, a portable computing device such as a laptop computer, a companion device (e.g., a digital camera, headphones), a tablet device, a wearable device such as a watch, a band, and the like, or any other appropriate device that includes, for example, one or more speakers, communications circuitry, processing circuitry, memory, a touchscreen, and/or a touchpad.
- the handheld electronic device 104 may not include a touchscreen but may support touchscreen-like gestures, such as in an extended reality environment.
- the handheld electronic device 104 may include a touchpad. In FIG. 1, by way of example, the handheld electronic device 104 is depicted as a tablet device.
- the electronic device 110 may be, for example, a smartphone, a portable computing device such as a laptop computer, a companion device (e.g., a digital camera, headphones), a tablet device, a wearable device such as a watch, a band, and the like, or any other appropriate device that includes, for example, one or more speakers, communications circuitry, processing circuitry, memory, a touchscreen, and/or a touchpad.
- the electronic device 110 may not include a touchscreen but may support touchscreen-like gestures, such as in an extended reality environment.
- the electronic device 110 may include a touchpad. In FIG. 1, by way of example, the electronic device 110 is depicted as a tablet device.
- the electronic device 110, the handheld electronic device 104, and/or the electronic device 105 may be, and/or may include all or part of, the electronic system discussed below with respect to FIG. 7.
- the electronic device 110 may be another device such as an Internet Protocol (IP) camera, a tablet, or a companion device such as an electronic stylus, etc.
- IP Internet Protocol
- the electronic device 115 may be, for example, desktop computer, a portable computing device such as a laptop computer, a smartphone, a companion device (e.g., a digital camera, headphones), a tablet device, a wearable device such as a watch, a band, and the like.
- the electronic device 115 is depicted as a desktop computer having one or more cameras 150 (e.g., multiple cameras 150).
- the electronic device 115 may be, and/or may include all or part of, the electronic system discussed below with respect to FIG. 7.
- the server 120 may form all or part of a network of computers or a group of servers 130, such as in a cloud computing or data center implementation.
- the server 120 stores data and software, and includes specific hardware (e.g., processors, graphics processors and other specialized or custom processors) for rendering and generating content such as graphics, images, video, audio and multi-media files for extended reality environments.
- the server 120 may function as a cloud storage server that stores any of the aforementioned extended reality content generated by the above-discussed devices and/or the server 120.
- FIG. 2 illustrates a block diagram of various components that may be included in electronic device 105, in accordance with aspects of the disclosure. As shown in FIG.
- electronic device 105 may include one or more cameras such as camera(s) 150 (e.g., multiple cameras 150, each including one or more image sensors 215) that capture images and/or video of the physical environment around the electronic device, one or more sensors 152 that obtain environment information (e.g., depth information) associated with the physical environment around the electronic device 105.
- Sensors 152 may include depth sensors (e.g., time-of-flight sensors, infrared sensors, radar, sonar, lidar, etc.), one or more microphones, and/or other types of sensors for sensing the physical environment.
- one or more microphones included in the sensor(s) 152 may be operable to capture audio input from a user of the electronic device 105, such as a voice input corresponding to the user speaking into the microphones.
- electronic device 105 also includes communications circuitry 208 for communication with electronic device 110, electronic device 115, servers 120, and/or other devices and/or systems in some examples.
- Communications circuitry 208 may include radio frequency (RF) communications circuitry for detecting radio frequency identification (RFID) tags, Bluetooth Low Energy (BLE) communications circuitry, other near-field communications (NFC) circuitry, WiFi communications circuitry, cellular communications circuitry, and/or other wired and/or wireless communications circuitry.
- RFID radio frequency identification
- BLE Bluetooth Low Energy
- NFC near-field communications
- electronic device 105 includes processing circuitry 204 (e.g., one or more processors and/or integrated circuits) and memory 206.
- Memory 206 may store (e.g., temporarily or permanently) content generated by and/or otherwise obtained by electronic device 105.
- memory 206 may temporarily store images of a physical environment captured by camera(s) 150, depth information corresponding to the images generated, for example, using a depth sensor of sensors 152, meshes and/or textures corresponding to the physical environment, virtual objects such as virtual objects generated by processing circuitry 204 to include virtual content, and/or virtual depth information for the virtual objects.
- Memory 206 may store (e.g., temporarily or permanently) intermediate images and/or information generated by processing circuitry 204 for combining the image(s) of the physical environment and the virtual objects and/or virtual image(s) to form, e.g., composite images for display by display 200, such as by compositing one or more virtual objects onto a pass-through video stream obtained from one or more of the cameras 150.
- the electronic device 105 may include one or more speakers 211. The speakers may be operable to output audio content, including audio content stored and/or generated at the electronic device 105, and/or audio content received from a remote device or server via the communications circuitry 208.
- Memory 206 may store instructions or code for execution by processing circuitry 204, such as, for example operating system code corresponding to an operating system installed on the electronic device 105, and application code corresponding to one or more applications installed on the electronic device 105.
- the operating system code and/or the application code when executed, may correspond to one or more operating system level processes and/or application level processes, such as processes that support capture of images, obtaining and/or processing environmental condition information, and/or determination of inputs to the electronic device 105 and/or outputs (e.g., display content on display 200) from the electronic device 105.
- one or more input devices include one or more camera sensors (e.g., one or more optical sensors and/or one or more depth camera sensors such as for tracking a user’s gestures (e.g., hand gestures and/or air gestures) as input.
- the one or more input devices are integrated with the computer system. In some examples, the one or more input devices are separate from the computer system.
- an air gesture is a gesture that is detected without the user touching an input element that is part of the device (or independently of an input element that is a part of the device) and is based on detected motion of a portion of the user’s body through the air including motion of the user’s body relative to an absolute reference (e.g., an angle of the user’s arm relative to the ground or a distance of the user’s hand relative to the ground), relative to another portion of the user’s body (e.g., movement of a hand of the user relative to a shoulder of the user, movement of one hand of the user relative to another hand of the user, and/or movement of a finger of the user relative to another finger or portion of a hand of the user), and/or absolute motion of a portion of the user’s body (e.g., a tap gesture that includes movement of a hand in a predetermined pose by a predetermined amount and/or speed, or a shake gesture that includes a predetermined speed or amount of rotation of a portion of the user’
- FIG. 3 is a block diagram illustrating a computer system (e.g., computer system 300) in accordance with some examples. Not all of the illustrated components are used in all examples; however, one or more examples can include additional and/or different components than those shown in FIG. 3.
- computer system 300 includes one or more components described above with respect to electronic device 105, handheld electronic device 104, electronic device 110, electronic device 115, and/or server 120 as shown in FIG. 1. Variations in the arrangement and type of the components can be made without departing from the spirit or scope of the claims as set forth herein. Additional components, different components, and/or fewer components can be used as well.
- computer system 300 loads, renders, manages, and/or displays computer-generated content in a 3D environment.
- the 3D environment can be either virtual or physical, with the computer-generated content either completely covering a field of view of a user or supplementing the field of view.
- computer system 300 can cause a virtual environment to be rendered and displayed to a user such that the user is provided content that is reactive to movements of the user.
- computer system 300 detects and processes the actions (e.g., movements and/or gestures of the user) to provide tailored information to applications executing on computer system 300.
- computer system 300 includes 3D environment process 310, 3D framework 320 (e.g., a 3D UI framework and/or other type of 3D framework), 2D framework 330 (e.g., a 2D UI framework and/or other type of 2D framework), display process 340, first user application 350, and second user application 360. While FIG. 3 illustrates that each of these components are on a single computer system, it should be recognized that one or more components can be on another computer system in communication (e.g., wired and/or wireless communication) with computer system 300. In addition, while each component will be discussed separately, in some examples, the functionality of one or more components are combined together or separated further. In some examples, one or more components of computer system 300 communicate with other components via application programming interfaces (APIs), inter-process communications (IPCs), and/or serial peripheral interfaces (SPIs).
- APIs application programming interfaces
- IPCs inter-process communications
- SPIs serial peripheral interfaces
- 3D environment process 310 executes as a background process (e.g., a daemon, a service, a system process, an application process, and/or one or more instructions) to manage a 3D environment on behalf of one or more applications (e.g., first user application 350 and/or second user application 360).
- 3D environment process 310 can create the 3D environment, manage a state of the 3D environment, receive requests from the one or more applications to render content in the 3D environment, communicate with 3D framework 320 and/or 2D framework 330 to service the requests, cause display process 340 to display the 3D environment, and/or detect and process inputs from a number of different sources.
- 3D environment process 310 provides one or more APIs to be used by the one or more applications for setting up the 3D environment.
- the APIs can work in a declarative form that allows for developers to create views, animations, and/or other user-interface elements without needing to configure the 3D environment imperatively.
- 3D environment process 310 creates a scene via a scene graph, adds one or more entities to the scene, and/or causes the scene to be rendered.
- 3D environment process 310 combines functionality of 3D framework 320 and 2D framework 330 such that user-interface elements and/or functionality provided by 3D framework 320 and/or 2D framework 330 can be used with each other rather than requiring one or the other to be used at a time.
- 3D environment process 310 acts as a bridge between 3D framework 320 and 2D framework 330, providing each the ability to render objects together in a single scene.
- 3D framework 320 renders 3D objects (e.g., via a first render server) and manages interactions with respect to the 3D objects and/or other objects.
- 2D framework renders 2D objects (e.g., via a second render server different from the first render server) (e.g., and not 3D objects) and manages interactions with respect to the 2D objects and/or other objects.
- 2D framework Rather than requiring each framework to work independently, such as providing a separate space for each to own, techniques described herein provide a single space that combines functionality of 3D framework 320 and 2D framework 330 to create the 3D environment.
- 2D environment can render objects to be used by 3D framework 320 when rendering the 3D environment.
- 3D environment process 310 creates a view (e.g., sometimes referred to as a world view) of a 3D environment and adds one or more 3D objects to the view.
- an object of the one or more objects can be hidden, as described further below.
- the object can be used by 3D framework 320 to maintain a place for 2D content from 2D framework 330.
- one technique for implementing such is via a scene graph.
- the scene graph can include multiple 3D entities that are managed by environment process 310 and/or 3D framework 320.
- Such 3D entities can include both visible entities and hidden entities.
- a hidden entity e.g., sometimes referred to as an invisible and/or nondisplayed entity
- the hidden entity is connected to a 2D entity such that 3D framework 320 communicates with 2D framework via the hidden entity and/or vice versa.
- FIG. 4 is a block diagram of a scene graph (e.g., scene graph 400) in accordance with some examples. It should be recognized that the block diagram is not meant to be limiting and that such is used for discussion purposes.
- scene graph 400 is a topological representation of a scene, with logical entities as nodes.
- scene graph 400 can encode entities, their relationships, and operations required to render content.
- world 410 of scene graph 400 can be a root node of scene graph 400 and include one or more render operations for the 3D environment.
- scene graph 400 can include one or more branch and/or leaf nodes that correspond to entities in the 3D environment and their respective rendering operations.
- first entity 420 can correspond to visible content in the 3D environment and include operations to be performed to render the visible content. Such operations can include a position and/or orientation of the visible content along with textures to be used for the rendering process of the visible content.
- the rendering process for the visible content is performed by 3D framework 320 via one or more APIs between 3D environment process 310 and 3D framework 320. Such operations can be performed without the use of 2D framework 330 as the visible content does not include any content rendered and/or managed by 2D framework 330.
- scene graph 400 also includes first group 430 as a branch node of world 410.
- first group 430 represents multiple entities that are related to each other.
- second entity 440 can be a leaf node of first group 430 and represent more visible content that is rendered via 3D framework 320 without the use of 2D framework 330.
- third entity 450 is another leaf node of first group 430.
- Third entity 450 can be a hidden entity that does not directly correspond to visible content.
- third entity 450 might not by itself include any visible user-interface elements. Instead, third entity 450 can be a proxy for a 2D entity that is managed by 2D framework 330.
- the 2D entity includes information related to content rendered via 2D framework 330 (sometimes referred to as 2D content and/or 2D rendered content herein) and communicates such information and/or rendered texture to third entity 450 for use by 3D framework 320 when rendering.
- a respective application e.g., first user application 350 and/or second user application 360
- 3D environment process 310 can add a 3D object to a scene graph for the 3D environment, the 3D object is not visible in the 3D environment and is intended to hold a location and/or orientation for 2D content.
- the 3D object can have a size corresponding to a size of the 2D content and be located at a position and/or an orientation requested by the respective application.
- the 3D object can also include a mesh (e.g., a 3D mesh) that is used by 3D framework 320 when adding 2D rendered content.
- a mesh e.g., a 3D mesh
- 2D content is rendered by 2D framework 330 and placed on the 3D mesh by 3D framework 320 so that rendering the 3D environment by 3D framework 320 includes rendering the 2D rendered content within the 3D environment.
- third entity 450 is used to inform 3D framework 320 when to update.
- 2D framework 330 and/or another process can detect an event that requires an update and, in response, send a request to 3D framework 320 via third entity 450 to update the 3D environment based on information associated with the 2D entity.
- the combination of third entity 450 and the 2D entity allows 2D framework 330 to communicate with 3D framework 320, whereas typically the two frameworks would be completely independent.
- third entity 450 is attached to and/or configured relative to another 3D entity (e.g., second entity 440).
- third entity 450 can ensure that a position and/or orientation of third entity 450 stays consistent with and/or maintains relative positioning and/or orientation with second entity 440 so that rendered content received by the 2D entity can move with second entity 440 without the 2D entity and/or 2D framework 330 needing to track where second entity 440 is located.
- third entity 450 includes information such as transformations (e.g., rotating, shrinking, enlarging, stretching, modifying a shape, and/or modifying a color characteristic) to be performed to content received from the 2D entity, collision dynamics to indicate how third entity 450 (and, by consequence to the 2D content tracking location and/or orientation of third entity 450, the 2D content) reacts to collisions with other 3D objects in the 3D environment, and/or physics dynamics to indicate how physics affects position and/or orientation of third entity 450 (and, by consequence to the 2D content tracking location and/or orientation of third entity 450, the 2D content) within the 3D environment.
- transformations e.g., rotating, shrinking, enlarging, stretching, modifying a shape, and/or modifying a color characteristic
- 3D environment process 310 can manage multiple different environments (e.g., serially and/or simultaneously) and that different environments can include overlapping visible and/or hidden entities.
- FIG. 5 is a block diagram illustrating a process (e.g., process 500) for processing gestures in a 3D environment according to some examples. Some operations in process 500 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted.
- process 500 is performed by a process (e.g., a system process or an application process) executing on a computer system, such as 3D environment process 310 described above.
- the process can detect a 3D gesture and conditionally translate the 3D gestures into a 2D gesture.
- Process 500 begins at 510, where the process detects a first 3D gesture.
- the first 3D gesture includes 6 degrees of freedom, such as an x, y, and z value.
- the first 3D gesture can correspond to movement of a user (e.g., a hand, an eye, and/or a leg) and be captured via one or more cameras in communication with the computer system.
- Other examples of input devices include a camera, a motion sensor, a depth sensor, a remote control, a gyroscope, an accelerometer, a touch-sensitive surface, and/or a physical input mechanism (e.g., a keyboard, a mouse, a rotatable input mechanism, and/or a physical button).
- the process determines which entity in a 3D environment (e.g., a virtual or a physical 3D environment) that the first 3D gesture is directed to.
- a 3D environment e.g., a virtual or a physical 3D environment
- the entity is determined using a scene graph (e.g., scene graph 400 of FIG. 4).
- the process can determine that the first 3D gesture is directed to a particular location within the 3D environment and, based on the particular location and the scene graph, determines that the first 3D gesture is directed to a particular entity in the scene graph.
- the particular entity is a visible entity (e.g., content that corresponds to a 3D framework (e.g., 3D framework 320) and not a 2D framework (e.g., 2D framework 330)).
- the particular entity is a hidden entity (e.g., content that corresponds to the 2D framework and not the 3D framework).
- the first process in response to determining that the first 3D gesture is directed to a visible entity in the scene graph, the first process sends an indication of the first 3D gesture to the 3D framework and/or an application (e.g., a user application executing on the computer system, such as first user application 350 and/or second user application 360) corresponding to the visible entity.
- an application e.g., a user application executing on the computer system, such as first user application 350 and/or second user application 360
- the indication of the first 3D gesture is sent to the application corresponding to the visible entity via the 3D framework.
- the indication of the first 3D gesture includes and/or is an indication of a type of gesture and does not specify one or more locations and/or orientations of an input that is detected (e.g., the indication is not at the level of joints and positions of fingers but rather at the level of communicating a category).
- the first process determines whether the 2D framework is aware of a type of gesture corresponding to the first 3D gesture.
- the 2D framework can be configured to recognize some types of 3D gestures and/or some 3D gestures are not specific to a 3D coordinate space and therefore do not need to be translated to a type of 2D gesture.
- the first process in response to determining that the 2D framework is aware of a type of gesture corresponding to the first 3D gesture, the first process sends the indication of the first 3D gesture to the 2D framework and/or an application (e.g., a user application executing on the computer system, such as first user application 350 and/or second user application 360) corresponding to the hidden entity.
- an application e.g., a user application executing on the computer system, such as first user application 350 and/or second user application 360
- the indication of the first 3D gesture is sent to the application corresponding to the hidden entity via the 2D framework.
- the indication of the first 3D gesture includes and/or is an indication of a type of gesture and does not specify one or more locations and/or orientations of an input that is detected (e.g., the indication is not at the level of joints and positions of fingers but rather at the level of communicating a category).
- the first process in response to determining that the 2D framework is not aware of a type of gesture corresponding to the first 3D gesture, the first process translates the first 3D gesture into a first 2D gesture (e.g., a gesture that is a type of 2D gesture) and sends an indication of the first 2D gesture to the 2D framework and/or the application corresponding to the hidden entity.
- a first 2D gesture e.g., a gesture that is a type of 2D gesture
- translating the first 3D gesture includes identifying a 2D gesture corresponding to the first 3D gesture.
- translating the first 3D gesture includes removing data from the first 3D gesture to create the first 2D gesture.
- translating the first 3D gesture includes removing an axis (e.g., removing z axis from x, y, and z) and/or reducing from 6 degrees of freedom to 2 degrees of freedom. It should be recognized that the examples for translating the first 3D gesture described above are not an exhaustive list and that some examples can be combined with others.
- the indication of the first 2D gesture is sent to the application corresponding to the hidden entity via the 2D framework.
- the indication of the first 2D gesture includes and/or is an indication of a type of gesture and does not specify one or more locations and/or orientations of an input that is detected (e.g., the indication is not at the level of joints and positions of fingers but rather at the level of communicating a category).
- an indication that is sent can correspond to a first type of coordinate space (e.g., a world space (e.g., a specific location in the 3D environment) and/or a user space (e.g., a location corresponding to a field of view of the computer system)).
- a first type of coordinate space e.g., a world space (e.g., a specific location in the 3D environment) and/or a user space (e.g., a location corresponding to a field of view of the computer system)
- the process receives, from an application, a request for a second type of coordinate space different from the first type of coordinate space.
- first process can convert the indication that was sent to a different coordinate space and send the converted indication to the application.
- the application defines a type of coordinate space for the process to use when communicating with the computer system.
- an indication that is sent can correspond to a first type of unit (e.g., a type of measurement).
- the process receives, from an application, a request for a second type of unit different from the first type of unit.
- first process can convert the indication that was sent to the second type of unit and send the converted indication to the application.
- the application defines a type of unit for the process to use when communicating with the computer system.
- a 2D gesture can be detected and, in accordance with a determination that the 2D gesture is being sent to a component requesting and/or needing a 3D gesture, the first process can translate the 2D gesture to a first 3D gesture using the opposite of one or more of the techniques described above with respect to translating from a 3D gesture to a 2D gesture.
- translation from a 2D gesture to a 3D gesture can be used in the context of a 2D window and/or application providing the first process a first 2D gesture and needing a 3D gesture corresponding to the first 2D gesture so that the application can send the 3D gesture to a 3D framework.
- FIG. 6 is a flow diagram illustrating a method (e.g., method 600) for integrating a 2D framework with a 3D framework in accordance with some examples. Some operations in method 600 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted. In some examples, method 600 is performed by a computer system (e.g., 300).
- a computer system e.g., 300
- the computer system receives, from an application (e.g., 340 and/or 350) (e.g., a user application and/or an application installed on the computer system (e.g., by a user and/or another computer system)) (e.g., of the computer system (e.g., a device, a personal device, a user device, and/or a head-mounted display (HMD))), a request to add (e.g., display, configure, and/or place) a two-dimensional (2D) entity (e.g., the 2D entity corresponding to 450) (e.g., a 2D object, a 2D model, a 2D user interface, a 2D window, and/or a 2D user interface element) at a first location in a three-dimensional (3D) environment (e.g., a virtual reality environment, a mixed reality environment, and/or an augmented reality environment), wherein the 3D environment includes one or more 3D entities (
- an application e
- the computer system is a phone, a watch, a tablet, a fitness tracking device, a wearable device, a television, a multi-media device, an accessory, a speaker, and/or a personal computing device.
- the computer system is in communication with input/output devices, such as one or more cameras (e.g., a telephoto camera, a wide-angle camera, and/or an ultra-wide-angle camera), speakers, microphones, sensors (e.g., heart rate sensor, monitors, antennas (e.g., using Bluetooth and/or Wi-Fi), fitness tracking devices (e.g., a smart watch and/or a smart ring), and/or near-field communication sensors).
- cameras e.g., a telephoto camera, a wide-angle camera, and/or an ultra-wide-angle camera
- speakers e.g., speakers, microphones, sensors (e.g., heart rate sensor, monitors, antennas (e.g., using Bluetooth and/or Wi-Fi), fitness tracking
- the computer system is in communication with a display generation component (e.g., a projector, a display, a display screen, a touch-sensitive display, and/or a transparent display).
- receiving the request to add the 2D entity at the first location in the 3D environment includes detecting, via one or more input devices in communication with the computer system, an input (e.g., a tap input and/or a nontap input, such as an air input (e.g., a pointing air gesture, a tapping air gesture, a swiping air gesture, and/or a moving air gesture), a gaze input, a gaze-and-hold input, a mouse click, a mouse click-and-drag, a key input of a keyboard, a voice command, a selection input, and/or an input that moves the computer system in a particular direction and/or to a particular location).
- one or more operations of method 600 are performed via a process (e.g., 310) of the computer system
- the computer system adds a first 3D entity (e.g., 450) (e.g., a hidden or proxy entity) to the first location in the 3D environment.
- adding the first 3D entity does not include displaying the first 3D entity.
- the first 3D entity is not visible in the 3D environment.
- the first 3D entity is different and/or separate from the 2D entity.
- the computer system renders (e.g., synthesizes, generates, and/or creates), via a 2D framework (e.g., 330) (e.g., of the computer system or in communication with the computer system) (e.g., a 2D user interface framework) (e.g., and not a 3D framework (e.g., 320)) (e.g., software that includes one or more predefined userinterface elements and/or one or more operations to build, generate, render, enable interactions with, and/or display a user interface and/or user interface element in two dimensions (e.g., and not three dimensions)), a representation (e.g., a graphical and/or visual representation) of the 2D entity.
- a 2D framework e.g., 330
- a 2D user interface framework e.g., and not a 3D framework (e.g., 320)
- a representation e.g., a graphical and/or visual representation
- the computer system renders, via a 3D framework (e.g., 320) (e.g., of the computer system or in communication with the computer system) (e.g., software that includes one or more predefined user-interface elements and/or one or more operations to build, generate, render, enable interactions with, and/or display a user interface and/or user interface element in three dimensions (e.g., and not two dimensions)), a representation (e.g., a graphical and/or visual representation) of a second 3D entity (e.g., 450) (e.g., different and/or separate from the first 3D entity) by (e.g., rendering the representation of the second 3D entity includes) performing (e.g., via the 3D framework) one or more operations on the representation of the 2D entity.
- the second 3D entity is a visible representation of the first 3D entity.
- the computer system renders, via the 3D framework, the representation of the second 3D entity by placing (e.g., via the 3D framework) the representation of the second 3D entity at the first location (e.g., at a location corresponding to the first 3D entity).
- the computer system displays, via a display generation component in communication with the computer system, the 3D environment including display of at least a portion of the representation of the second 3D entity at the first location.
- an observable (e.g., perceivable, visible, and/or audible) representation e.g., a visual and/or an auditory representation
- the first 3D entity is used for tracking purposes for the application, the 3D framework, and/or the 2D framework, such as to track position, orientation, and/or pose for the second 3D entity and/or information related to the 2D entity).
- performing the one or more operations include rotating the representation of the 2D entity, shirking the representation of the 2D entity, enlarging the representation of the 2D entity, stretching the representation of the 2D entity, modifying a shape of the representation of the 2D entity, modifying a color characteristic (e.g., hue, saturation, and/or brightness) of the representation of the 2D entity, or one or more combinations thereof (e.g., based on information (e.g., position, orientation, and/or pose) included in the first 3D entity) (e.g., based on information (e.g., a size and/or a color characteristic) included in the second 2D entity) (e.g., based on information (e.g., a color characteristic and/or a physics setting to augment an appearance of the representation of the 2D entity) corresponding to the 3D environment).
- information e.g., position, orientation, and/or pose
- information e.g., a size and/or a color characteristic
- the computer system sends, to a display process (e.g., 340) (e.g., of the computer system and/or of a display generation component (e.g., of a headmounted display (HMD) device) in communication with the computer system) (e.g., a system process or another application different from the application), a request to display a representation of the one or more 3D entities (e.g., the computer system causes the display process to display the representation of the one or more 3D entities).
- the display process is different from the process of the computer system.
- the computer system detects movement of a visible representation (e.g., visible within the 3D environment) of a third 3D entity of the one or more 3D entities from a second location to a third location.
- the third location is different from the second location.
- the second location is a first distance from the first location (e.g., zero or more units).
- the third location is different from the first location.
- the visible representation of the third 3D entity is not currently visible as a result of a current orientation of a view area but is visible as a result of another orientation different from the current orientation of the view area.
- the computer system in response to detecting the movement of the visible representation of the third 3D entity, places (e.g., via the 3D framework) the representation of the second 3D entity at a fourth location different from the first location.
- the fourth location is the first distance from the third location (e.g., the representation of the second 3D entity maintains a relative position to the visible representation of the third 3D entity).
- the computer system receives (e.g., via the 2D framework and/or the application) a change to the 2D entity.
- the change is caused by the application.
- the change is caused by an interaction of the representation of the 2D entity with a representation of a 3D entity in the 3D environment.
- the change is caused by an input (e.g., a tap input and/or a non-tap input, such as an air input (e.g., a pointing air gesture, a tapping air gesture, a swiping air gesture, and/or a moving air gesture), a gaze input, a gaze-and-hold input, a mouse click, a mouse click-and-drag, a key input of a keyboard, a voice command, a selection input, and/or an input that moves the computer system in a particular direction and/or to a particular location) detected via one or more input devices (e.g., a camera (e.g., a telephoto camera, a wide-angle camera, and/or an ultra-wide-angle camera), a microphone, a sensor (such as a heart rate sensor), a touch-sensitive surface, a mouse, a keyboard, a touch pad, and/or an input mechanism (e.g., a physical input mechanism, such as a rotatable
- the computer system in response to receiving the change to the 2D entity, renders, via the 3D framework, a second representation of the 2D entity, wherein the second representation is different from (e.g., a different appearance, location, orientation, and/or pose) the representation of the 2D entity.
- the computer system detects, via one or more input devices (e.g., a camera (e.g., a telephoto camera, a wide-angle camera, and/or an ultra-wide-angle camera), a microphone, a sensor (such as a heart rate sensor), a touch-sensitive surface, a mouse, a keyboard, a touch pad, and/or an input mechanism (e.g., a physical input mechanism, such as a rotatable input mechanism and/or a button)) (e.g., in communication with the computer system), an input (e.g., a tap input and/or a non-tap input, such as an air input (e.g., a pointing air gesture, a tapping air gesture, a swiping air gesture, and/or a moving air gesture), a gaze input, a gaze-and-hold input, a mouse click, a mouse click-and-drag, a key input of a keyboard, a voice command, a
- a camera e
- the computer system in response to detecting the input directed to the representation of the second 3D entity and in accordance with a determination that a first set of one or more criteria (e.g., as described below with respect to method 700) is satisfied, the computer system translates the input to a second input (e.g., generating the second input) (e.g., a 2D representation of the input) different from the input (e.g., as described below with respect to method 700).
- the first set of one or more criteria includes a criterion that is satisfied when the input is a first type of input (e.g., as described below with respect to method 700).
- the computer system in response to detecting the input directed to the representation of the second 3D entity and in accordance with a determination that a first set of one or more criteria (e.g., as described below with respect to method 700) is satisfied, the computer system sends, to the application (e.g., via the 2D framework), an indication of the second input (e.g., without sending an indication of the input).
- the application e.g., via the 2D framework
- the computer system in response to detecting the input directed to the representation of the second 3D entity and in accordance with a determination that a second set of one or more criteria is satisfied (e.g., as described below with respect to method 700), sends, to the application (e.g., via the 2D framework), an indication of the input (e.g., without sending the indication of the second input).
- the second set of one or more criteria is different from the first set of one or more criteria.
- the second set of one or more criteria includes a criterion that is satisfied when the first set of one or more criteria is not satisfied.
- the computer system in response to detecting the input directed to the representation of the second 3D entity and in accordance with a determination that a third set of one or more criteria (e.g., as described below with respect to method 700) is satisfied, the computer system (and/or the process of the computer system) sends, to the application (e.g., via the 2D framework), the indication of the second input and the indication of the input.
- the third set of one or more criteria is different from the first set of one or more criteria and/or the second set of one or more criteria.
- the third set of one or more criteria includes a criterion that is satisfied when the first set of one or more criteria and/or the second set of one or more criteria is not satisfied.
- the computer system detects an interaction (e.g., a collision and/or an effect of a proximity (e.g., gravity, a pull effect, and/or a push effect)) between the representation of the second 3D entity and a representation of a fourth 3D entity of the one or more 3D entities.
- the computer system renders (e.g., synthesizes, generates, and/or creates), via the 2D framework, an updated representation (e.g., a graphical and/or visual representation) of the 2D entity.
- the updated representation of the 2D entity is different from the representation of the 2D entity.
- the computer system in response to detecting the interaction, renders, via the 3D framework, an updated representation (e.g., a graphical and/or visual representation) of the second 3D entity.
- an updated representation e.g., a graphical and/or visual representation
- the updated representation of the second 3D entity is different from the representation of the second 3D entity.
- the computer system receives, from the application, information (e.g., an effect of the 2D entity on one or more other representations in the 3D environment, such as a glow or other outward effect on an area outside of a representation of the 2D entity) corresponding to the 2D entity.
- information e.g., an effect of the 2D entity on one or more other representations in the 3D environment, such as a glow or other outward effect on an area outside of a representation of the 2D entity.
- the information is included in the request to add the 2D entity at the first location in the 3D environment.
- the computer system after receiving the request to add the 2D entity at the first location in the 3D environment, the computer system renders (e.g., via the 3D framework) the 3D environment (e.g., renders one or more representations of one or more visible entities in the 3D environment) (e.g., based on the information corresponding to the 2D entity), wherein the information affects rendering a visual representation of a fifth 3D entity in the 3D environment, and wherein the fifth 3D entity does not correspond to the 2D entity.
- the information affects rendering the visual representation of the fifth 3D entity in the 3D environment by causing a color of the fifth 3D entity to be based on a color of the 2D entity.
- the first 3D entity is added to a graph (e.g., 400) (e.g., a scene graph) of the 3D environment.
- a graph e.g., 400
- a scene graph e.g., a scene graph
- the first 3D entity is included in a second 3D environment different from the 3D environment.
- the second 3D environment is not displayed concurrently with the 3D environment.
- the second 3D environment is displayed in response to a request to display the second 3D environment (e.g., and not the 3D environment).
- updates to the first 3D entity affects the 3D environment and the second 3D environment.
- the first 3D entity includes a size (e.g., corresponding to the 2D entity, the first 3D entity, the second 3D entity, and/or the representation of the second 3D entity), a location (e.g., corresponding to the 2D entity, the first 3D entity, the second 3D entity, and/or the representation of the second 3D entity) (e.g., within the 3D environment), an orientation (e.g., corresponding to the 2D entity, the first 3D entity, the second 3D entity, and/or the representation of the second 3D entity) (e.g., within the 3D environment), a pose (e.g., corresponding to the 2D entity, the first 3D entity, the second 3D entity, and/or the representation of the second 3D entity) (e.g., within the 3D environment), or any combination thereof.
- a size e.g., corresponding to the 2D entity, the first 3D entity, the second 3D entity, and/or the representation of the second 3D entity
- the computer system receives, from a second application (e.g., 350 and/or 360) (e.g., a user application and/or an application installed on the computer system (e.g., by a user and/or another computer system)) (e.g., of a computer system (e.g., a device, a personal device, a user device, and/or a head-mounted display (HMD))) different from the application, a request to add a sixth 3D entity (e.g., corresponding to another 2D entity and/or not corresponding to another 2D entity) to the 3D environment.
- the second application is not associated with the application.
- the second application is a different type of application than the application.
- the computer system after receiving the request to add the sixth 3D entity to the 3D environment, the computer system renders, via the 3D framework, a representation of the sixth 3D entity.
- the computer system causes concurrent display, in the 3D environment, of the representation of the sixth 3D entity and the representation of the third 3D entity.
- the application causes the one or more 3D entities to be rendered for (and/or displayed in) the 3D environment.
- method 700 optionally includes one or more of the characteristics of the various methods described above with reference to method 600.
- the 3D environment of method 700 is the 3D environment of method 600. For brevity, these details are not repeated below.
- FIG. 7 is a flow diagram illustrating a method (e.g., method 700) for translating between gestures in accordance with some examples. Some operations in method 700 are, optionally, combined, the orders of some operations are, optionally, changed, and some operations are, optionally, omitted.
- method 700 is performed at a computer system (e.g., 300) (e.g., a device, a personal device, a user device, and/or a head-mounted display (HMD)) in communication with one or more input devices (e.g., a camera, a touch-sensitive surface, a motion detector, and/or a microphone).
- the computer system is a phone, a watch, a tablet, a fitness tracking device, a wearable device, a television, a multi-media device, an accessory, a speaker, and/or a personal computing device.
- method 700 is performed by a daemon (e.g., 310), a system process (e.g., 310), and/or a user process (e.g., 310, 350, and/or 360).
- the computer system detects, via the one or more input devices, a first input (e.g., 510) (e.g., a tap input and/or a non-tap input, such as an air input (e.g., a pointing air gesture, a tapping air gesture, a swiping air gesture, and/or a moving air gesture), a gaze input, a gaze-and-hold input, a mouse click, a mouse click-and-drag, a key input of a keyboard, a voice command, a selection input, and/or an input that moves the computer system in a particular direction and/or to a particular location) corresponding to a three- dimensional (3D) environment (e.g., a virtual reality environment, a mixed reality environment, and/or an augmented reality environment).
- a first input e.g., 510)
- a tap input and/or a non-tap input such as an air input (e.g., a pointing air gesture, a tapping air gesture,
- the computer system in response to detecting the first input corresponding to the 3D environment and in accordance with a determination (e.g., 520) that a first set of one or more criteria is satisfied, wherein the first set of one or more criteria includes a criterion that is satisfied when the input is directed to a first entity (e.g., 450) (e.g., an object, a model, a user interface, a window, an element in a scene graph, and/or a user-interface element) of a first type (e.g., a hidden or proxy entity) in the 3D environment, the computer system translates (e.g., 560) the first input to a second input different from the first input.
- the second input is a different format than the first input.
- the second input corresponds to the first input.
- the computer system sends (e.g., 560), to a first application (e.g., 350 and/or 360) (e.g., a user application and/or an application installed on the computer system (e.g., by a user)) (e.g., of the computer system (e.g., corresponding to the first entity)) (e.g., via a first framework (e.g., 330) (e.g., of the computer system or in communication with the computer system) (e.g., a two-dimensional (2D) framework and/or a 2D user interface framework) (e.g., and not a 3D framework) (e.g., software that includes one or more predefined userinterface elements and/or one or more operations to build, generate, render, enable interactions with, and/or display a user interface and/or user
- a first framework e.g., 330
- 2D framework and/or a 2D user interface framework e.g., and not a 3D
- the computer system sends (e.g., 430), to a second application (e.g., of the computer system) (e.g., via a second framework (e.g., 320) different from the first framework (e.g., of the computer system or in communication with the computer system) (e.g., a three-dimensional (3D) framework and/or a 3D user interface framework) (e.g., and not
- a method includes: displaying 2D content; while displaying the 2D content, detecting, via one or more input devices, a 3D gesture directed to the 2D content; in response to detecting the 3D gesture directed to the 2D content, translating the 3D gesture to a 2D gesture and providing the 2D gesture to an application (e.g., via a 2D framework as described above) corresponding to the 2D content (e.g., the application caused display of and/or includes the 2D content).
- an application e.g., via a 2D framework as described above
- an observable (e.g., perceivable, visible, and/or audible) representation e.g., a visual and/or an auditory representation
- the first entity of the first type e.g., a hidden entity corresponding to a 2D entity, as described above with respect to method 600
- the first entity of the first type is not present (e.g., visible, displayed, and/or output) in the 3D environment (e.g., the first entity of the first type is used for tracking purposes for the first application, the 3D framework, and/or the 2D framework, such as to track position, orientation, and/or pose for a representation of an entity related to, corresponding to, and/or associated with the first entity of the first type).
- an observable (e.g., perceivable, visible, and/or audible) representation e.g., a visual and/or an auditory representation
- the first entity of the second type e.g., a visible entity, as described above with respect to method 600
- an auditory representation e.g., a visual and/or an auditory representation of the first entity of the second type (e.g., a visible entity, as described above with respect to method 600) is present (e.g., visible, displayed, and/or output) in the 3D environment.
- the computer system in response to detecting the first input corresponding to the 3D environment and in accordance with a determination that a third set of one or more criteria is satisfied, wherein the third set of one or more criteria includes a criterion that is satisfied when the input is a first type of input (e.g., a type of gesture known by the first application and/or the 2D framework) directed to the first entity of the first type in the 3D environment, the computer system sends (e.g., 550), to the first application (e.g., via the first framework), the indication of the first input (e.g., without translating the first input to the second input), wherein the third set of one or more criteria is different from the first set of one or more criteria and the second set of one or more criteria, and wherein the first set of one or more criteria includes a criterion that is satisfied when the input is a second type of input (e.g., a type of gesture not known by the first application and/or the 2D framework) different from the first
- the first input includes a 3D gesture (e.g., a hand pose in three dimensions and/or movement in three dimensions).
- the first input is the 3D gesture.
- the first input includes an air gesture. In some examples, the first input is the air gesture.
- the first input includes a gaze of a user directed to a location (e.g., a location corresponding to the first entity and/or the second entity) in the 3D environment.
- the first input is the gaze of the user directed to the location in the 3D environment.
- the computer system in response to detecting the first input corresponding to the 3D environment and in accordance with a determination that a fourth set of one or more criteria is satisfied, wherein the fourth set of one or more criteria includes a criterion that is satisfied when the input is directed to a third entity (e.g., the first entity, the second entity, or another entity different from the first entity and the second entity) of a third type in the 3D environment, the computer system sends, to a fourth application (e.g., via the first framework and/or the second framework), an indication of a location corresponding to a first type of coordinate space for the first input, an orientation corresponding to the first type of coordinate space for the first input, a pose corresponding to the first type of coordinate space for the first input, a magnitude corresponding to the first type of coordinate space for the first input, or any combination thereof, wherein the first type of coordinate space is defined by the fourth application.
- a fourth application e.g., via the first framework and/or the second framework
- the fourth application is the first application or the second application.
- the indication corresponding to the first type of coordinate space is sent with an indication of an input (e.g., the indication of the first input and/or the indication of the second input).
- the first type of coordinate space is a coordinate space with respect to a window of the fourth application (e.g., application centric, such as a distance from the window) (e.g., and not with respect to a location in the 3D environment other than a location corresponding to the window).
- the indication corresponding to the first type of coordinate space includes an identification of a type of gesture.
- the computer system sends, to the fourth application (e.g., via the first framework and/or the second framework), an indication of a location corresponding to a second type of coordinate space for the first input, an orientation corresponding to the second type of coordinate space for the first input, a pose corresponding to the second type of coordinate space for the first input, a magnitude corresponding to the second type of coordinate space for the first input, or any combination thereof, wherein the second type of coordinate space is defined by the fourth application, wherein the fifth set of one or more criteria is different from the fourth set of one or more criteria, wherein the fourth type of entity is different from the third type of entity, and wherein the second type of coordinate space is different from the first type of coordinate space (e.g., via the first framework and/or the second framework), an indication of a location corresponding to a second type of coordinate space for the first input, an orientation corresponding to the second type of coordinate space for the first input, a pose corresponding to the second type of coordinate space for the first input, a magnitude corresponding to the second type
- the fourth entity is different from the third entity.
- the third type of entity is different from the first type of entity and the second type of entity.
- the fourth type of entity is different from the first type of entity and the second type of entity.
- the indication corresponding to the second type of coordinate space is sent with an indication of an input (e.g., the indication of the first input and/or the indication of the second input).
- the second type of coordinate space is a coordinate space with respect to the 3D environment (e.g., a location within the 3D environment, such as a location of a user, a location of a view point, and/or a location of an object other than the window) (e.g., world centric, such as a distance from a user, and/or world centric, such as a stance from a location in the world) (e.g., and not with respect to a location corresponding to the window).
- a coordinate space with respect to the 3D environment e.g., a location within the 3D environment, such as a location of a user, a location of a view point, and/or a location of an object other than the window
- world centric such as a distance from a user
- world centric such as a stance from a location in the world
- the indication corresponding to the first type of coordinate space and the indication corresponding to the second type of coordinate space is sent to the fourth application (e.g., the fourth application is defined to receive indications corresponding to the first type of coordinate space and the second type of coordinate space).
- the computer system sends, to the fourth application, the indication corresponding to the first type of coordinate space and the indication corresponding to the second type of coordinate space.
- the computer system receives, from the fourth application, a request for an indication for an input corresponding to a particular type of coordinate space.
- the computer system in response to receiving the request for the indication for the input corresponding to the particular type of coordinate space, the computer system sends, to the fourth application, the indication for the input corresponding to the particular type of coordinate space.
- the indication corresponding to the second type of coordinate space includes an identification of a type of gesture.
- the second application is the first application. In some examples, the second application is different from the first application.
- the second input includes a two-dimensional representation of the first input (e.g., a location, orientation, and/or pose with two or less dimensions).
- translating the first input to the second input includes modifying a representation of a respective input from having six degrees of freedom to two degrees of freedom.
- method 600 optionally includes one or more of the characteristics of the various methods described above with reference to method 700.
- the first entity of the first type of method 700 is the first 3D entity of method 600.
- a view for hosting a 3D framework simulation within a 2D framework view hierarchy is provided.
- the view is defined in a crossimport overlay of a 3D framework (e.g., a 3D UI framework) and a 2D framework (e.g., a 2D UI framework).
- the view can be implemented as a wrapper for a view of the 3D framework.
- the view can be implemented as an integration of the 3D framework with the 2D framework via one or more private serial peripheral interfaces (SPIs).
- SPIs serial peripheral interfaces
- the view uses inline closure syntax to bridge between the declarative world of the 2D framework and the imperative world of the 3D framework.
- the view provides a mutable 'Content' struct. This struct serves as a container for entities in the view’s hierarchy, as well as an entry point to one or more top-level APIs corresponding to the 3D framework.
- var modelEntity ModelEntity(mesh: ,generateSphere(radius: 0.1))
- Handling state updates for the view can be done by adding an 'update' closure, evaluated any time the containing view's body is re-evaluated:
- the view omits the use of a context or coordinator, as developers can instead use standard constructs corresponding to the 2D framework like ' ⁇ State' and ' ⁇ Environment' as needed.
- a placeholder view is shown instead.
- a developer can customize this placeholder view using the optional 'placeholder' ' ViewBuilder', such as to show a 'Progress View' spinner.
- a developer can use another instance of the view to display a 3D placeholder model:
- the 'Content' struct provided to the view’s closures is a generic type conforming to a protocol.
- the protocol represents the core interface common to some permutations of the view. Concrete types conforming to this protocol can add additional functionality specific to their configuration. For example, an isolated view could provide additional functionality in its 'Content' struct that a shared view would not have access to (more discussion on isolation later).
- the protocol itself is defined like so:
- componentType Component. Type?
- processes add and/or remove entities using convenience APIs via the protocol:
- clients can use the 'subscribe(to:on:componentType:)' method on the protocol, or one of the convenience methods provided in an extension:
- EventSource? nil
- componentType Component.Type?
- 'EventSubscription' is defined in the 3D framework.
- subscription content.subscribe(to: CollisionEvents.Began.self) ⁇
- the view provides content in a shared scene on mixed reality operating system, and content isolated to a 2D camera projection (AR or non- AR) on a phone operating system and/or a computer operating system.
- AR 2D camera projection
- additional APIs can be provided on each type separately.
- To convert coordinates between the 2D framework space and the 3D framework entity space we introduce new methods on 'Content' for a mixed reality operating system:
- Entity? nil
- the "to / from" 'CoordinateSpace' represents the 2D framework reference space, while the other is implicitly the local space of the Content.
- the 2D framework space is represented in points, with Y pointing down and a top-left-back origin, while the 3D framework space is represented in meters, with Y pointing up and a center origin.
- We also use the currency types for each framework (' SIMD3 ⁇ Float>', 'BoundingBox', and 'Transform' as used in the 3D framework, as opposed to the Spatial types 'Point3D', 'Rect3D', and ' AffineTransform3D' used in the 2D framework).
- a flexible container view of the view with a frame can be explicitly set in the same way as any other view.
- clients that wish to size a view to match the size of a containing entity can do so by manually calculating the bounding box of that entity and setting that as the view’s frame, like so:
- modelSize content. convert(bounds, to: .local). size
- componentType Component. Type?
- componentType Component.Type?
Landscapes
- Engineering & Computer Science (AREA)
- General Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Human Computer Interaction (AREA)
- Architecture (AREA)
- Computer Graphics (AREA)
- Computer Hardware Design (AREA)
- Software Systems (AREA)
- Processing Or Creating Images (AREA)
Abstract
Certaines techniques sont décrites dans la description en vue d'intégrer un cadre 2D à un cadre 3D. De telles techniques utilisent un concept appelé entité masquée pour relier les deux cadres ensemble. D'autres techniques sont décrites dans la description en vue de traduire des gestes d'un premier type à un second type dans certaines situations.
Applications Claiming Priority (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363471209P | 2023-06-05 | 2023-06-05 | |
| US63/471,209 | 2023-06-05 | ||
| US18/622,447 US20240402872A1 (en) | 2023-06-05 | 2024-03-29 | Techniques for three-dimensional environments |
| US18/622,447 | 2024-03-29 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024253828A1 true WO2024253828A1 (fr) | 2024-12-12 |
Family
ID=91585567
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2024/030181 Ceased WO2024253828A1 (fr) | 2023-06-05 | 2024-05-20 | Techniques pour environnements tridimensionnels |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2024253828A1 (fr) |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090123046A1 (en) * | 2006-05-11 | 2009-05-14 | Koninklijke Philips Electronics N.V. | System and method for generating intraoperative 3-dimensional images using non-contrast image data |
| US20190247130A1 (en) * | 2009-02-17 | 2019-08-15 | Inneroptic Technology, Inc. | Systems, methods, apparatuses, and computer-readable media for image management in image-guided medical procedures |
| US20230165639A1 (en) * | 2021-12-01 | 2023-06-01 | Globus Medical, Inc. | Extended reality systems with three-dimensional visualizations of medical image scan slices |
-
2024
- 2024-05-20 WO PCT/US2024/030181 patent/WO2024253828A1/fr not_active Ceased
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20090123046A1 (en) * | 2006-05-11 | 2009-05-14 | Koninklijke Philips Electronics N.V. | System and method for generating intraoperative 3-dimensional images using non-contrast image data |
| US20190247130A1 (en) * | 2009-02-17 | 2019-08-15 | Inneroptic Technology, Inc. | Systems, methods, apparatuses, and computer-readable media for image management in image-guided medical procedures |
| US20230165639A1 (en) * | 2021-12-01 | 2023-06-01 | Globus Medical, Inc. | Extended reality systems with three-dimensional visualizations of medical image scan slices |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US12079626B2 (en) | Methods and systems for creating applications using scene trees | |
| CN112136101B (zh) | 为用户界面和动画提供应用程序编程接口的框架 | |
| AU2017200358B2 (en) | Multiplatform based experience generation | |
| EP4462253A2 (fr) | Plate-forme de réalité générée par ordinateur | |
| US11195323B2 (en) | Managing multi-modal rendering of application content | |
| WO2023102139A1 (fr) | Modes d'interface utilisateur pour affichage tridimensionnel | |
| Alshaal et al. | Enhancing virtual reality systems with smart wearable devices | |
| Kim et al. | The augmented reality internet of things: Opportunities of embodied interactions in transreality | |
| Colombo et al. | Mixed reality to design lower limb prosthesis | |
| US20240402872A1 (en) | Techniques for three-dimensional environments | |
| WO2024253828A1 (fr) | Techniques pour environnements tridimensionnels | |
| Sanna et al. | Developing touch-less interfaces to interact with 3D contents in public exhibitions | |
| Placitelli et al. | Toward a framework for rapid prototyping of touchless user interfaces | |
| US20240404228A1 (en) | Techniques for managing computer-generated experiences | |
| US20240402891A1 (en) | Techniques for placing user interface objects | |
| WO2024253835A1 (fr) | Techniques de gestion d'expériences générées par ordinateur | |
| US20250209720A1 (en) | Techniques for rendering content | |
| US20250378651A1 (en) | Techniques for managing three-dimensional content | |
| US20240404168A1 (en) | Techniques for rendering content | |
| WO2024253823A1 (fr) | Techniques de mise en place d'objets d'interface utilisateur | |
| Siandri et al. | Computer Vision in Extended Reality | |
| Pavlopoulou et al. | A Mixed Reality application for Object detection with audiovisual feedback through MS HoloLenses | |
| Balado et al. | 3D as‐built environments in extended reality applications: a systematic review | |
| Elvezio | XR Development with the Relay and Responder Pattern | |
| Santos | MARC. MD-Multimedia Augmented Reality Collaboration in Mobile Devices Exploring Collaborative ar Interactions |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24734467 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |