WO2026016140A1 - 用于执行用户任务的方法、装置、设备和介质 - Google Patents

用于执行用户任务的方法、装置、设备和介质

Info

Publication number
WO2026016140A1
WO2026016140A1 PCT/CN2024/106255 CN2024106255W WO2026016140A1 WO 2026016140 A1 WO2026016140 A1 WO 2026016140A1 CN 2024106255 W CN2024106255 W CN 2024106255W WO 2026016140 A1 WO2026016140 A1 WO 2026016140A1
Authority
WO
WIPO (PCT)
Prior art keywords
image
objects
location
physical space
user
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2024/106255
Other languages
English (en)
French (fr)
Inventor
吴弘涛
郑子琳
孔涛
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Youzhuju Network Technology Co Ltd
Original Assignee
Beijing Youzhuju Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Youzhuju Network Technology Co Ltd filed Critical Beijing Youzhuju Network Technology Co Ltd
Priority to EP24841099.5A priority Critical patent/EP4706904A4/en
Priority to PCT/CN2024/106255 priority patent/WO2026016140A1/zh
Priority to CN202480003521.4A priority patent/CN121712620A/zh
Publication of WO2026016140A1 publication Critical patent/WO2026016140A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J9/00Program-controlled manipulators
    • B25J9/16Program controls
    • B25J9/1694Program controls characterised by use of sensors other than normal servo-feedback from position, speed or acceleration sensors, perception control, multi-sensor controlled systems, sensor fusion
    • B25J9/1697Vision controlled systems
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J9/00Program-controlled manipulators
    • B25J9/16Program controls
    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J9/00Program-controlled manipulators
    • B25J9/16Program controls
    • B25J9/1656Program controls characterised by programming, planning systems for manipulators
    • B25J9/1664Program controls characterised by programming, planning systems for manipulators characterised by motion, path, trajectory planning
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05BCONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
    • G05B2219/00Program-control systems
    • G05B2219/30Nc systems
    • G05B2219/37Measurements
    • G05B2219/37555Camera detects orientation, position workpiece, points of workpiece
    • GPHYSICS
    • G05CONTROLLING; REGULATING
    • G05BCONTROL OR REGULATING SYSTEMS IN GENERAL; FUNCTIONAL ELEMENTS OF SUCH SYSTEMS; MONITORING OR TESTING ARRANGEMENTS FOR SUCH SYSTEMS OR ELEMENTS
    • G05B2219/00Program-control systems
    • G05B2219/30Nc systems
    • G05B2219/45Nc applications
    • G05B2219/45084Service robot

Definitions

  • Robotics technology has developed rapidly and is widely used in many technological fields.
  • Various specialized robotic devices have been developed; for example, in industrial environments, robots can perform a variety of tasks such as processing, grasping, sorting, and packaging.
  • robots In home environments, for instance, robotic vacuum cleaners and window cleaning robots have been developed.
  • robots typically can only perform pre-set, fixed tasks and cannot perform different user-defined tasks according to user needs.
  • a method for performing a user task is provided.
  • a user task is received from a user, the user task instructing a robotic device to classify multiple objects within a first range in a physical space; an image including the multiple objects is acquired; for a first object among the multiple objects, a first destination location in the physical space is determined based on the image of the first object; and the robotic device moves the first object to the first destination location.
  • an apparatus for performing a user task includes: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to classify a plurality of objects within a first range in a physical space; an acquiring module configured to acquire an image including the plurality of objects; and a determining module configured to, for a first object among the plurality of objects, determine, based on the image, a first object in the physical space.
  • the destination location; and the execution module are configured to cause the robotic device to move the first object to the first destination location.
  • an electronic device in a third aspect of this disclosure, includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.
  • a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.
  • a computer program product comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.
  • Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure
  • Figure 2 shows a block diagram of some implementations of the present disclosure for performing user tasks
  • Figure 3 shows a block diagram of an image acquisition process according to some implementations of this disclosure
  • Figure 4 shows a block diagram of the process of moving an object according to some implementations of this disclosure
  • Figure 5 shows a block diagram of the calling model process according to some implementations of this disclosure
  • Figure 6 illustrates the process of invoking the action model according to some implementations of this disclosure.
  • Figure 7 shows a flowchart of a method for performing user tasks according to some implementations of this disclosure
  • Figure 8 shows a block diagram of an apparatus for performing user tasks according to some implementations of the present disclosure.
  • Figure 9 shows a block diagram of a device capable of implementing various implementations of the present disclosure.
  • the term “comprising” and similar terms should be understood as open inclusion, i.e., “including but not limited to”.
  • the term “based on” should be understood as “at least partially based on”.
  • the term “one implementation” or “the implementation” should be understood as “at least one implementation”.
  • the term “some implementations” should be understood as “at least some implementations”.
  • Other explicit and implicit definitions may also be included below.
  • the term “model” can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and/or future-developed technical solutions.
  • a notification message is sent to the user.
  • a prompt message in response to a user's active request, can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format.
  • the pop-up window can also include a selection control allowing the user to choose whether to "agree” or "disagree” to provide personal information to the electronic device.
  • the term "in response to” as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
  • robots and machine learning technologies have been widely applied in various scenarios.
  • robots typically can only perform pre-set, fixed tasks and cannot execute different user tasks according to user needs.
  • robotic devices struggle to determine user requirements and thus perform corresponding tasks.
  • Simple robotic devices have been developed to perform specific tasks. However, these devices cannot understand complex user instructions, nor can they execute the desired tasks according to user commands in complex physical spaces. Therefore, it is desirable to control the robot's operation in an effective way to perform the desired tasks.
  • Figure 1 shows a block diagram 100 of the application environment according to an exemplary implementation of this disclosure.
  • a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform various tasks.
  • the physical space 160 can include, but is not limited to, one or more rooms.
  • the physical space 160 can include, but is not limited to, a living room, bedroom, study, kitchen, toilet, etc., or a combination of one or more of the above.
  • the robot device 110 may include multiple parts.
  • the control unit 111 can serve as the control center of the robot device 110, and an application can be loaded into the control unit 111 to control the various parts of the robot device.
  • the user 120 can use the interaction unit 112 to interact with the robot device 110, for example, by inputting control commands to the robot device 110 to perform desired tasks.
  • the robot device 110 may include an arm 113 for performing actions such as grasping and releasing.
  • the arm 113 can grasp an object and move it to a desired position, and so on.
  • the robot device 110 may also include a data acquisition unit 114.
  • the data acquisition unit 114 may include various types, such as an image acquisition unit, a sound acquisition unit, etc.
  • the robot device 110 may further include a sensing unit for detecting surrounding objects, for example, detecting the distance between the robot and surrounding objects based on laser light, etc.
  • the robot device 110 may also include a drive unit 115; for example, the robot device 110 may be deployed on a movable base, and the drive unit 115 may drive the wheels of the base to move along a desired path.
  • Physical environment 160 may include one or more acquisition units 130, ..., and 132.
  • one or more image acquisition devices may be deployed in a room to acquire images of the room from various angles.
  • Physical environment 160 may include control device 140, which can control one or more acquisition units 130, ..., and 132, etc., via a network (not shown).
  • control device 140 can control various electrical devices in physical space 160.
  • a machine learning model (e.g., model 150) may be provided to manage physical space 160.
  • model 150 may be located inside physical space 160, alternatively and/or additionally, model 150 may be located at a remote device outside physical space 160, and control device 140, robotic device 110, or other device may access the remote model 150 via a network.
  • Model 150 may include one or more models. If model 150 includes multiple models, these multiple models may include multiple types of models. Model 150 may, for example, include at least a language model (LM) and an action model. The language model, by learning from a large corpus, is capable of question answering. The action model can control the robotic device 110 to perform various actions. Model 150 may also include, for example, an image recognition model, a text recognition model, and so on.
  • LM language model
  • the action model can control the robotic device 110 to perform various actions.
  • Model 150 may also include, for example, an image recognition model, a text recognition model, and so on.
  • user 120 can instruct robot device 110 to manipulate various objects in physical space 110.
  • objects can be various items in the home environment.
  • user 120 can instruct robot device 110 to find a certain object in physical space 160; or user 120 can instruct robot device 110 to place the found object in a designated location, and so on.
  • Figure 2 illustrates a block diagram 200 for performing user tasks according to some implementations of this disclosure.
  • the robot device 110 in the physical space 160 can receive a user task 210 from the user 120.
  • the user task 210 can instruct the robot device 110 to categorize multiple objects within a specified area (e.g., area 230, also referred to as the first area) in the physical space 160.
  • a specified area e.g., area 230, also referred to as the first area
  • the user 120 can say in natural language, "Organize the items on the desktop.”
  • the first area is the "desktop”
  • the multiple objects are the multiple objects placed on the desktop, including object 222 and object 224.
  • the robot device 110 can acquire images including the multiple objects. For example, this can be done via acquisition units 114, 130, ... And at least one of 132 to obtain the image.
  • a first destination location in physical space is determined based on an image.
  • the first object can be any suitable object among the multiple objects.
  • the first destination location can be the place where the first object should be placed.
  • a table includes a bottle of water (i.e., object 222) and a book (i.e., object 224).
  • Object 222 should be placed in a refrigerator, kitchen, or other location; therefore, if object 222 is the first object, the first destination location could be a refrigerator (e.g., destination location 250) and/or a kitchen.
  • the book should be placed on a bookshelf, in a study, or other location; therefore, if the book is the first object, the first destination location could be a bookshelf and/or a study.
  • the robot device 110 can be instructed to move the first object to the first destination location. For example, the robot device 110 can be instructed to move object 222 to destination location 250.
  • the methods described above can be executed at any computing device with computing capabilities.
  • the methods described above can be executed using an application deployed at robot device 110.
  • an application can be deployed at control device 140 to execute the methods described above.
  • the powerful processing capabilities of model 150 can be invoked to find a first range that may include multiple objects and the destination location in physical space for each of the multiple objects.
  • robot device 110 can move along path 242 to range 230 and along path 244 move object 222 of the multiple objects within range 230 to its destination location 250 in physical space 160.
  • a robotic device can be controlled to move multiple objects within a first range to their corresponding destination locations. This can improve the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby completing the intended user task.
  • the image can come from at least one of the following: an acquisition device at the robot device, an acquisition device in a first physical space, or an acquisition device in a second physical space.
  • Figure 3 shows a block diagram 300 of the image acquisition process according to some implementations of this disclosure.
  • images e.g., one or more images 310) of the physical space 160 can be acquired from the acquisition unit 114 at the robot device 110. Since the robot device 110 can move freely within the physical space 160, the acquisition unit 114 can acquire images from various locations within the physical space, thereby facilitating the location of a first range and multiple objects.
  • images of the physical space 160 can be acquired from acquisition units 130, ..., and 132.
  • acquisition units 130, ..., and 132 can be pre-deployed at designated locations within the physical space 160, such as a corner of the ceiling, etc.
  • images of the physical space 160 taken from a top-down angle can be obtained, facilitating an overall understanding of the layout of the physical space 160. This can help determine a first range and the location of a first destination corresponding to a first object within the first range.
  • an image of the physical space 160 where the robot device 110 is located can be acquired, and a first range can be located based on the image. For example, if user 120 says “tidy up the items on the desktop" in natural language, then the user task 210 instructs the robot device 110 to sort multiple objects on the desktop.
  • the range 230 where the desktop is located can be determined based on the image of the physical space 160. For example, a prompt for locating the first range can be acquired: "Please determine the location of the desktop in the following image," and the range 230 can be located using a model. Then, the robot device 110 can be instructed to move to the range 230.
  • images of multiple objects within a first range can be acquired simultaneously with the image of the physical space 160 where the robot device 110 is located.
  • the images of multiple objects within the first range can be determined based on the image of the physical space 160.
  • images of multiple objects within the first range can be cropped from the image of the physical space 160.
  • images of multiple objects within the first range can also be acquired in response to the robot device 110 moving to the first range.
  • Images of multiple objects within a first range are acquired using the acquisition unit 114 at the robot device 110.
  • the images of multiple objects may include one or more images. If the images of multiple objects include multiple images, each image may correspond to one object. For example, the images of multiple objects may include multiple images corresponding to each of the multiple objects.
  • the position of the first object can be adjusted to acquire the image of the first object.
  • the position of the acquisition device used to acquire the image can be adjusted (e.g., an angle that avoids occlusion can be found) to acquire the image of the first object.
  • the position of object 222 and/or the position of the mobile robot device 110 can be adjusted so that the robot device 110 can acquire an image containing only object 222 using the acquisition unit 114.
  • the image corresponding to each of the multiple objects can be acquired separately, which can improve the accuracy of the image corresponding to each object, and thus improve the accuracy of subsequent image-based task execution.
  • the images of multiple objects consist of only one image, it can be determined whether an occlusion relationship exists between the multiple objects.
  • Any suitable method can be used to determine whether an occlusion relationship exists between the multiple objects.
  • model 150 can be used to determine whether an occlusion relationship exists between the multiple objects.
  • it can be based on any suitable rule or algorithm.
  • prompts can be provided to user 120 to allow user 120 to determine whether an occlusion relationship exists between the multiple objects.
  • the determination result provided by user 120 can be received, and the existence of an occlusion relationship between the multiple objects can be determined based on that determination result.
  • the user can be instructed to eliminate the occlusion relationship.
  • an image including multiple objects can be directly acquired. If there is an occlusion relationship between multiple objects, in response to determining that an occlusion relationship exists between multiple objects, the robot device 110 can be instructed to move at least one of the multiple objects, thereby acquiring an image including multiple objects. See Figure 4 for further details, which shows a block diagram 400 of the process of moving objects according to some implementations of this disclosure. As shown in Figure 4, object 402 is located in front of object 401 and occludes object 401. At this time, the robot device can be instructed to move object 402 from position 430 to a position that does not obstruct object 401 (such as position 430' in image 420).
  • a target position can be determined, and the robotic device can be instructed to move object 402 to the target position.
  • actions can be generated using model 150 to control the robotic device to move object 402 from position 430 to position 430'.
  • the robotic device can be supported in handling complex problems in complex environments, thereby performing user tasks in a more accurate manner.
  • the first destination location of the first object in physical space can be determined based on the images.
  • the first type of the first object can be determined based on the images.
  • any suitable method can be used.
  • the first type of the first object can be determined based on a pre-stored database.
  • the pre-stored database can store multiple objects and their respective types.
  • the first object can be retrieved from this pre-stored database to determine its first type.
  • multiple text items (e.g., labels on objects) associated with multiple objects can be identified from an image (e.g., using model 150 to identify the image).
  • a first type can be determined based on the first text item associated with the first object.
  • the text item associated with object 222 could be, for example, the text item on the bottled water packaging.
  • Object 222 can be determined to be bottled water based on the text item on the bottled water packaging.
  • the location of a second object of the first type in physical space can also be determined so as to determine the location of a first destination based on the location of the second object.
  • any suitable method can be used to determine the location of the second object.
  • the location of the second object can be determined based on an image of the physical space 160 where the robot device 110 is located. Taking bottled water as an example, an image of the physical space 160 where the robot device 110 is located can be acquired, and the locations of other bottled water items in the physical space 160 can be determined based on this image. Taking the example that other bottled water items are placed in the kitchen, the location of the second object can be determined... The location is determined to be the kitchen.
  • the location of the second object can be determined based on predetermined rules or algorithms. Continuing with the example of bottled water as the first object, if predetermined rules indicate that the bottled water should be placed in the refrigerator, then the location of the second object can be determined to be the refrigerator based on these rules. It is understandable that the location of the second object can also be determined using a model (e.g., model 150), or manually by a user (120), and so on. For example, robot device 110 can ask the user, "Where should I put the bottled water?" and place the bottled water in the location specified by the user.
  • model 150 e.g., model 150
  • robot device 110 can ask the user, "Where should I put the bottled water?" and place the bottled water in the location specified by the user.
  • the cue word can be provided to a machine learning model (e.g., model 150) to determine the location of a second object of the first type.
  • the machine learning model's response to the cue word i.e., the model output of the machine learning model to the cue word, can indicate the location of the second object (e.g., a refrigerator).
  • the location of the second object i.e., the first destination location of the first object, can be determined based on the machine learning model's response to the cue word.
  • the machine learning model can process images, and if the image includes a second object, the model can output the location of the second object (e.g., the region coordinates of the second object in the image, and/or directly output the image of the region where the second object is located, etc.). If the image does not include the second object, the model can output a response such as "not found".
  • the presence or absence of a second object in an image can be detected in multiple ways, thereby improving the performance of robotic devices in detecting the location of a second object in an image.
  • another physical space associated with the physical space 160 may be determined.
  • the physical space 160 in which the robot device 110 is located may be referred to as the first physical space
  • the physical space associated with the robot device 110 may be referred to as the first physical space
  • the other physical space associated with space 160 is called the second physical space.
  • the second physical space here is a potential physical space that may include the first object.
  • the second physical space can be identified as the refrigerator.
  • cue words for locating the second object can be obtained, and the response of the machine learning model to the cue words can be received to determine the second physical space.
  • the location of the second physical space in the first physical space can be determined as the location of the second object, that is, the first destination location of the first object.
  • Figure 5 illustrates a block diagram 500 of the process of invoking the model according to some implementations of this disclosure.
  • multiple objects on the desktop e.g., bottled water and books
  • a corresponding prompt 510 can be obtained based on image 310.
  • Prompt 510 can be expressed, for example, as: "Please determine the physical space that may include 'bottled water' from the following images," or, for example, as: "Where might 'bottled water' be placed in the following images,” and so on.
  • Prompt 510 and image 310 can be input to language model 520 so that language model 520 can find the second physical space that may include bottled water from image 310.
  • Language model 520 can be, for example, a model included in model 150.
  • a message associated with the first object can be provided to the user 120.
  • This message can be used to prompt the user 120 to confirm whether to move the first object to the first destination location.
  • it can be provided to the user 120 via the display screen of the robot device 110, a voice playback device, etc.
  • the decision to move the first object to the first destination location can be determined based on the reply. If the reply indicates that the first object should be moved to the first destination location, the robot device 110 can be instructed to move the first object to the first destination location. Therefore, the object can be moved only upon obtaining user confirmation, and the robot device's execution can be improved by following user instructions. The accuracy of the task.
  • the motion trajectory from the position of the first object to the first destination position can also be determined based on an image of the physical space 160. It is understood that the motion trajectory can also be determined by any suitable method, such as using model 150, manual intervention, or based on predetermined rules or algorithms. This motion trajectory can, for example, instruct the robotic device 110 to avoid obstacles in the physical space 160 and move from the position of the first object to the first destination position along a shorter path. After determining the motion trajectory, the robotic device 110 can be instructed to move the first object to the first destination position according to the motion trajectory.
  • the actions to be performed by the robot device 110 can be determined using model 150, and the robot device 110 can be controlled to perform these actions using model 150.
  • the specific method of moving the first object to the refrigerator e.g., opening the refrigerator, placing the first object, etc.
  • a corresponding prompt can be constructed to ask model 150 how to open the refrigerator.
  • the prompt could be, for example, "Determine how to open the refrigerator from the following images," and the prompt and the corresponding image (e.g., an image of the physical environment 160 including the refrigerator) can be sent to model 150.
  • Model 150 could, for example, return: pull the handle to open the refrigerator door. Then, it can instruct robot device 110 to pull the handle to open the refrigerator.
  • the powerful processing capabilities of the model can be invoked to solve unknown problems in complex environments, thereby determining the actions that the robot device needs to perform. In this way, the robot device's ability to handle complex tasks can be improved, thereby executing user tasks in a more accurate manner.
  • a motion model can be used to determine the specific actions to be performed by the robot device. See Figure 6 for further details, which shows a block diagram 600 illustrating the process of invoking a motion model according to some implementations of this disclosure.
  • a motion model 630 can be provided, which can determine the specific actions to be performed by the robot device based on the current state and instructions of the robot device.
  • This motion model 630 can be a pre-trained and fine-tuned model.
  • the motion model 630 can also be one of the models included in model 150.
  • the current state 620 may include data from various aspects, such as an image of the robot device 110, an image of the robot device 110's environment, pose data of the robot arm (e.g., the positions of the robot arm's joints (POS1, ...)), and the state of the tool (e.g., a gripper, a cutting tool, etc.) fixed to the end of the robot arm.
  • pose data of the robot arm e.g., the positions of the robot arm's joints (POS1, 7)
  • the state of the tool e.g., a gripper, a cutting tool, etc.
  • Instructions and the current state can be input into the motion model 630, which then uses the motion model to determine the action to be performed by the robot device 110 based on the instructions and the current state.
  • the action can represent the difference between the robot device 110's current pose and the next pose, and the difference between the tool's current state and the next state, etc.
  • An instruction 610 (e.g., "open the refrigerator") can be input to the motion model 630.
  • the instruction 610 can be expressed in natural language, and the instruction 610 can be determined from the response of the language model 630.
  • the current state of the robot device 110 can be obtained, and the motion model 630 can determine the corresponding action 640 based on the input data. For example, the orientation, position, speed, acceleration, etc., of each joint in the arm, and/or the wheels and/or other movable devices of the robot device at the next time point can be determined.
  • the determined action 640 can be used to control the state of the robot device 110 at the next time point.
  • the robot's movements can be controlled by using a model.
  • the robot can be instructed to move a first object based on the model's output, and the robot's movements can be precisely controlled, thereby executing user tasks more efficiently.
  • multiple objects within a first range can be identified as having multiple destination locations.
  • a similar approach can be used to control the robot device 110 to move the multiple objects sequentially to their respective destination locations.
  • the robot device 110 can be controlled to move bottled water to the refrigerator, then move a book to the bookshelf, and so on.
  • the robot device 110 can be instructed to move this group of objects at once. Therefore, by moving a group of objects at a time based on their different types, the efficiency of the robot device 110 in moving objects can be improved.
  • the robot device 110 may be instructed to acquire a third object.
  • the threshold condition may, for example, indicate a threshold number of objects in a group.
  • the robot device 110 may be instructed to acquire a third object in response to the number of objects included in the group of objects of the first type reaching 4.
  • the third object may, for example, be another object used to move the group of objects.
  • the third object may be a pre-determined object.
  • user 120 can pre-configure robot device 110 to move multiple objects using specified objects.
  • the third object could be any object that helps move this group of bottled water, such as a tray, basket, bag, or trailer.
  • robot device 110 can be instructed to move a group of objects to a first destination location via the third object. For example, if the third object is a tray, robot device 110 can be instructed to move a group of objects (e.g., a group of bottled water) to the first destination location using the tray.
  • a group of objects e.g., a group of bottled water
  • model 150 can be used to determine the way to leave the first destination location, the movement trajectory to return to the first range, the movement trajectory to move from the first range to the second destination, and so on.
  • New images can be acquired and new prompts can be constructed to query model 150 (e.g., a language model) for the next instruction. Prompts can be represented, for example, as: "Please determine the next instruction based on the following image," "What to do next,” etc.
  • the language model can return "Close the refrigerator” based on the currently received image.
  • a corresponding action can be generated to instruct the robot device to close the refrigerator.
  • the priority of multiple objects within a first range can also be determined, and the multiple objects can be moved sequentially based on their priority. Any method can be used to...
  • the priority of multiple objects can be determined, for example, using model 150 or predetermined rules. For instance, objects that need to be frozen (e.g., ice cream, frozen meat) have a higher priority than objects that can be stored at room temperature. In this way, higher-priority objects can be moved first, improving the quality of task execution by the robotic device 110.
  • users can interact with the robot device through language, actions, gestures, etc.
  • users can state the user task they wish to perform, predefine a certain action to specify the user task, and so on.
  • users can make the action of placing items, and this action can be used as a trigger for the robot device to perform the user task of sorting and placing items.
  • the robot device can automatically ask the user whether items need to be sorted and place and ask the user about the scope of the task to be processed. If an affirmative answer is received, the robot device can perform the user task.
  • a user can interact with the robot device via the interaction unit 112, for example, the user inputs a task represented by text and/or images, and controls the robot device to perform the task.
  • the user can specify the execution conditions of the task, for example, to execute the task immediately, to execute the task after a predetermined time, or to execute the task when predetermined conditions are determined to be met (e.g., after the user has eaten), etc.
  • the robotic device can provide users with various messages. For example, regarding a first object (e.g., bottled water), if the physical space includes multiple first destination locations (e.g., a refrigerator and a kitchen), it can ask the user if they need to move the bottled water. To the refrigerator or to the kitchen. Or, for example, if the robot cannot find the initial destination, it can ask the user where to move the first object, and so on.
  • a first object e.g., bottled water
  • first destination locations e.g., a refrigerator and a kitchen
  • Alternate and/or additional locations can be determined using a visual positioning system to pinpoint the location of the robotic device and/or individual objects.
  • a map of the physical space can be pre-acquired, and the locations of each object can be marked on this map.
  • the robotic device can utilize echo detection units to detect distances to surrounding objects and, by combining the acquired images with the physical space map, determine the precise location of each object.
  • CAD computer-aided design
  • GIS geographic information systems
  • tracking units can be added to remote controls for household appliances (e.g., television remotes, air conditioner remotes) so that the robotic device can promptly acquire the precise location of important objects, and so on.
  • household appliances e.g., television remotes, air conditioner remotes
  • the robot's initial position and desired destination can be determined based on the methods described above.
  • the robot can determine a path from its initial position to its destination. For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, continuously update the path, enabling the robot to move along the path to its destination.
  • the robotic device can perform a specified task. For example, it can acquire a specified object and move it to the appropriate location.
  • Constraints i.e., the constraints that should be followed during task execution, can be determined using a language model and/or a knowledge base. For example, an image and corresponding prompts can be acquired, and the image and prompts can be input into the language model, thereby receiving the constraints from the language model.
  • prompts can be determined as: "Based on the following image, determine the constraints that should be followed during the movement of object XXX,” or "Please determine the precautions during the movement of object XXX,” etc.
  • an object e.g., bottled water, plate, bowl, etc.
  • the object's original posture should be maintained (e.g., remaining vertical and not tilted).
  • constraints can be input into the motion model, at which point the series of actions output by the motion model will perform the corresponding tasks while ensuring the constraints are met.
  • safety during the operation of robotic devices can be ensured, thereby preventing accidental damage to an object, and so on.
  • a robotic device can perform user tasks in complex physical spaces.
  • the robotic device can autonomously move multiple objects within a first range to their corresponding destination locations. This improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby completing the intended user task.
  • Figure 7 illustrates a flowchart of a method 700 for performing a user task according to some implementations of this disclosure.
  • a user task is received from a user, instructing a robotic device to classify multiple objects within a first range in physical space.
  • an image including the multiple objects is acquired.
  • a first destination location in physical space is determined based on the image.
  • the robotic device moves the first object to the first destination location.
  • obtaining an image includes: in response to determining that there is an occlusion relationship between multiple objects, the robotic device moves at least one of the multiple objects; And to obtain an image that includes multiple objects.
  • the image of the first object is determined based on: adjusting the position of the first object to acquire an image of the first object; and adjusting the position of the acquisition device used to acquire the image to acquire an image of the first object.
  • method 700 further includes: acquiring an image of the physical space where the robot device is located; locating a first range based on the image; and moving the robot device to the first range.
  • determining the location of the first destination includes: determining a first type of a first object based on an image; determining the location of a second object of the first type in physical space; and determining the location of the first destination based on the location of the second object.
  • determining the first type of the first object further includes: identifying multiple text items from the image that are associated with multiple objects respectively; and, for the first object among the multiple objects, determining the first type based on the first text item associated with the first object among the multiple text items.
  • determining the location of the second object includes at least one of the following: identifying the second object in an image in physical space to determine the location of the second object.
  • determining the location of the second object includes at least one of the following: constructing a cue word in an image in physical space and a first type of the first object, the cue word being used to determine the location of the first type of object in the image in physical space; and determining the location of the second object based on the response of a machine learning model to the cue word.
  • the robotic device moves a first object to a first destination location by: determining a motion trajectory from the location of the first object to the first destination location based on an image of the physical space; and moving the first object to the first destination location according to the motion trajectory.
  • moving a first object to a first destination location by a robotic device includes: providing a user with a message associated with the first object; and moving the first object to the first destination location in response to receiving a response from the user to the message.
  • the robot device moves a first object to a first destination location by: in response to determining that the number of a group of objects of a first type among a plurality of objects meets a threshold condition, the robot device acquires a third object; the robot device moves a group of objects to the first destination location via the third object.
  • FIG. 8 shows a block diagram of an apparatus 800 for performing a user task according to some implementations of the present disclosure.
  • the apparatus 800 includes: a receiving module 810 configured to receive a user task from a user, the user task instructing a robotic device to classify a plurality of objects within a first range in physical space; an acquiring module 820 configured to acquire an image including the plurality of objects; a determining module 830 configured to determine a first destination location in physical space for a first object among the plurality of objects based on the image; and an executing module 840 configured to cause the robotic device to move the first object to the first destination location.
  • a receiving module 810 configured to receive a user task from a user, the user task instructing a robotic device to classify a plurality of objects within a first range in physical space
  • an acquiring module 820 configured to acquire an image including the plurality of objects
  • a determining module 830 configured to determine a first destination location in physical space for a first object among the plurality of objects
  • the acquisition module 820 is further configured to: in response to determining that there is an occlusion relationship between multiple objects, cause the robot device to move at least one of the multiple objects; and acquire an image including the multiple objects.
  • the image of the first object is determined based on: adjusting the position of the first object to acquire an image of the first object; and adjusting the position of the acquisition device used to acquire the image to acquire an image of the first object.
  • the acquisition module 820 is further configured to: acquire an image of the physical space where the robot device is located; locate a first range based on the image; and move the robot device to the first range.
  • the determining module 830 is further configured to: determine a first type of a first object based on an image; determine the location of a second object of the first type in physical space; and determine the location of a first destination based on the location of the second object.
  • the determining module 830 is further configured to: identify from an image a plurality of text items associated with a plurality of objects respectively; and, for a first object among the plurality of objects, based on a first text item associated with the first object among the plurality of text items, Identify the first type.
  • the determining module 830 is further configured to: identify a second object in an image in physical space to determine the location of the second object.
  • the determining module 830 is further configured to: construct a prompt word in the image of the first object in the physical space and the first type of the first object, the prompt word being used to determine the location of the object of the first type in the image of the physical space; and determine the location of the second object based on the response of the prompt word by a machine learning model.
  • the execution module 840 is further configured to: determine a motion trajectory from the position of the first object to the first destination position based on an image of the physical space; and cause the robot device to move the first object to the first destination position according to the motion trajectory.
  • the execution module 840 is further configured to: provide a message associated with the first object to a user; and in response to receiving a response from the user to the message, cause the robotic device to move the first object to a first destination location.
  • the execution module 840 is further configured to: in response to determining that the number of a group of objects of a first type among a plurality of objects meets a threshold condition, cause the robot device to acquire a third object; cause the robot device to move a group of objects to a first destination location via the third object.
  • Figure 9 shows a block diagram of a device 900 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 900 shown in Figure 9 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 900 shown in Figure 9 can be used to implement the methods described above.
  • the computing device 900 is in the form of a general-purpose computing device.
  • Components of the computing device 900 may include, but are not limited to, one or more processors or processing units 910, memory 920, storage devices 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960.
  • the processing unit 910 may be a physical or virtual processor and is capable of performing various processes according to programs stored in the memory 920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve computing efficiency.
  • the computing device has a parallel processing capability of 900.
  • Computing device 900 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 900, including but not limited to volatile and non-volatile media, removable and non-removable media.
  • Memory 920 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof.
  • Storage device 930 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media capable of storing information and/or data (e.g., training data for training) and accessible within computing device 900.
  • the computing device 900 may further include additional removable/non-removable, volatile/non-volatile storage media.
  • disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided.
  • each drive may be connected to a bus (not shown) via one or more data media interfaces.
  • the memory 920 may include a computer program product 925 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
  • the communication unit 940 enables communication with other computing devices via a communication medium. Additionally, the components of the computing device 900 can function as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
  • PCs network personal computers
  • Input device 950 can be one or more input devices, such as a mouse, keyboard, trackball, etc.
  • Output device 960 can be one or more output devices, such as a monitor, speaker, printer, etc.
  • Computing device 900 can also communicate with one or more external devices (not shown) via communication unit 940 as needed.
  • External devices include storage devices, display devices, etc., and one or more devices that enable user interaction with computing device 900.
  • Communication, or communication with any device that enables computing device 900 to communicate with one or more other computing devices e.g., a network interface card, modem, etc.
  • Such communication may be performed via an input/output (I/O) interface (not shown).
  • a computer-readable storage medium that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above.
  • a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
  • a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
  • These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
  • These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and/or other device to operate in a particular manner.
  • the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
  • Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby enabling the instructions that execute on the computer, other programmable data processing apparatus, or other device to be implemented.
  • each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function.
  • the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.
  • each block in the block diagrams and/or flowcharts, and combinations of blocks in the block diagrams and/or flowcharts may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Landscapes

  • Engineering & Computer Science (AREA)
  • Robotics (AREA)
  • Mechanical Engineering (AREA)
  • Manipulator (AREA)

Abstract

提供了用于执行用户任务的方法、装置、设备和介质。在一种方法中,接接收来自用户的用户任务,用户任务指示机器人设备来分类物理空间中的第一范围内的多个对象。获取包括多个对象的图像。针对多个对象中的第一对象,基于图像来确定第一对象在物理空间中的第一目的地位置。机器人设备将第一对象移动至第一目的地位置。利用本公开的示例性实现方式,机器人设备可以在复杂的物理空间中执行用户任务,可以提高机器人设备在复杂环境下执行任务的灵活度和精确度,进而完成预期的用户任务。

Description

用于执行用户任务的方法、装置、设备和介质 技术领域
本公开的示例性实现方式总体涉及机器人领域,特别地涉及利用机器人来执行用户任务的方法、装置、设备和计算机可读存储介质。
背景技术
机器人技术已经得到了迅速发展,并且已经被广泛地用于多个技术领域。目前已经开发出了多种专用机器人设备,例如,在工业环境中,可以使用机器人来执行加工、抓取、分类、包装等多种任务。又例如,在家居环境下,已经开发出了扫地机器人、擦玻璃机器人,等等。然而,机器人通常仅能执行预先设置的固定任务,并不能按照用户需求来执行不同的用户任务。
发明内容
在本公开的第一方面,提供了一种用于执行用户任务的方法。在该方法中,接收来自用户的用户任务,用户任务指示机器人设备来分类物理空间中的第一范围内的多个对象;获取包括多个对象的图像;针对多个对象中的第一对象,基于图像来确定第一对象在物理空间中的第一目的地位置;以及机器人设备将第一对象移动至第一目的地位置。
在本公开的第二方面,提供了一种用于执行用户任务的装置。该装置包括:接收模块,被配置为接收来自用户的用户任务,用户任务指示机器人设备来分类物理空间中的第一范围内的多个对象;获取模块,被配置为获取包括多个对象的图像;确定模块,被配置为针对多个对象中的第一对象,基于图像来确定第一对象在物理空间中的第一 目的地位置;以及执行模块,被配置为使得机器人设备将第一对象移动至第一目的地位置。
在本公开的第三方面,提供了一种电子设备。该电子设备包括:至少一个处理单元;以及至少一个存储器,至少一个存储器被耦合到至少一个处理单元并且存储用于由至少一个处理单元执行的指令,指令在由至少一个处理单元执行时使电子设备执行根据本公开第一方面的方法。
在本公开的第四方面,提供了一种计算机可读存储介质,其上存储有计算机程序,计算机程序在被处理器执行时使处理器实现根据本公开第一方面的方法。
在本公开的第五方面,提供了一种计算机程序产品,包括计算机程序,其中所述计算机程序在被处理器执行时实现根据本公开第一方面的方法。
应当理解,本内容部分中所描述的内容并非旨在限定本公开的实现方式的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。
附图说明
在下文中,结合附图并参考以下详细说明,本公开各实现方式的上述和其他特征、优点及方面将变得更加明显。在附图中,相同或相似的附图标注表示相同或相似的元素,其中:
图1示出了根据本公开的一个示例性实现方式的应用环境的框图;
图2示出了根据本公开的一些实现方式的用于执行用户任务的框图;
图3示出了根据本公开的一些实现方式的图像采集过程的框图;
图4示出了根据本公开的一些实现方式的移动对象的过程的框图;
图5示出了根据本公开的一些实现方式的调用模型的过程的框图;
图6示出了根据本公开的一些实现方式的调用动作模型的过程的 框图;
图7示出了根据本公开的一些实现方式的用于执行用户任务的方法的流程图;
图8示出了根据本公开的一些实现方式的用于执行用户任务的装置的框图;以及
图9示出了能够实施本公开的多个实现方式的设备的框图。
具体实施方式
下面将参照附图更详细地描述本公开的实现方式。虽然附图中示出了本公开的某些实现方式,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实现方式,相反,提供这些实现方式是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实现方式仅用于示例性作用,并非用于限制本公开的保护范围。
在本公开的实现方式的描述中,术语“包括”及其类似用语应当理解为开放性包含,即“包括但不限于”。术语“基于”应当理解为“至少部分地基于”。术语“一个实现方式”或“该实现方式”应当理解为“至少一个实现方式”。术语“一些实现方式”应当理解为“至少一些实现方式”。下文还可能包括其他明确的和隐含的定义。如本文中所使用的,术语“模型”可以表示各个数据之间的关联关系。例如,可以基于目前已知的和/或将在未来开发的多种技术方案来获取上述关联关系。
可以理解的是,本技术方案所涉及的数据(包括但不限于数据本身、数据的获取或使用)应当遵循相应法律法规及相关规定的要求。
可以理解的是,在使用本公开各实现方式公开的技术方案之前,均应当根据相关法律法规通过适当的方式对本公开所涉及个人信息的类型、使用范围、使用场景等告知用户并获得用户的授权。
例如,在响应于接收到用户的主动请求时,向用户发送提示信息, 以明确地提示用户,其请求执行的操作将需要获取和使用到用户的个人信息。从而,使得用户可以根据提示信息来自主地选择是否向执行本公开技术方案的操作的电子设备、应用程序、服务器或存储介质等软件或硬件提供个人信息。
作为一种可选的但非限制性的实现方式,响应于接收到用户的主动请求,向用户发送提示信息的方式,例如可以是弹出窗口的方式,弹出窗口中可以以文字的方式呈现提示信息。此外,弹出窗口中还可以承载供用户选择“同意”或“不同意”向电子设备提供个人信息的选择控件。
可以理解的是,上述通知和获取用户授权过程仅是示意性的,不对本公开的实现方式构成限定,其他满足相关法律法规的方式也可应用于本公开的实现方式中。
在此使用的术语“响应于”表示相应的事件发生或者条件得以满足的状态。将会理解,响应于该事件或者条件而被执行的后续动作的执行时机,与该事件发生或者条件成立的时间,二者之间未必是强关联的。例如,在某些情况下,后续动作可在事件发生或者条件成立时立即被执行;而在另一些情况下,后续动作可在事件发生或者条件成立后经过一段时间才被执行。
示例环境
近年来,机器人技术和机器学习技术已经被广泛应用于多个应用场景。然而,机器人通常仅能执行预先设置的固定任务,并不能按照用户需求来执行不同的用户任务。尤其是,在复杂应用环境下,机器人设备难以确定用户需求进而执行相应的任务。
目前已经开发了执行特定任务的简单机器人设备,然而,此类简单机器人设备并不能理解复杂的用户指令,也不能在复杂的物理空间中按照用户指令来执行期望的任务。此时,期望可以以有效的方式控制机器人的操作,进而执行期望的任务。
根据本公开的一个示例性实现方式,提出了一种用于执行用户任务的方法。参见图1描述根据本公开的一个示例实现方式的应用环境,图1示出了根据本公开的一个示例性实现方式的应用环境的框图100。如图1所示,机器人设备110和用户120可以位于物理空间160中,并且用户120可以控制机器人设备110来执行多种任务。物理空间160可以包括但不限于一个或者多个房间。例如,在家居环境下,物理空间160可以包括但不限于客厅、卧室、书房、厨房、厕所,等等,或者包括以上一个或者多个的组合。
如图1所示,机器人设备110可以包括多个部分。例如,控制单元111可以作为机器人设备110的控制中心,可以向控制单元111中加载应用程序,以便控制机器人设备中的各个部分。用户120可以使用交互单元112来与机器人设备110交互,例如,向机器人设备110输入控制指令,以便利用机器人设备110执行期望的任务。机器人设备110可以包括手臂113,用于执行抓取、释放等动作。例如,手臂113可以抓取某个对象,并且将该对象移动至期望的位置,等等。
备选地和/或附加地,机器人设备110还可以包括采集单元114。在此,采集单元114可以包括多种类型,例如,图像采集单元、声音采集单元,等等。备选地和/或附加地,机器人设备110可以进一步包括用于检测周围物体的感测单元,例如,可以基于激光来检测机器人与周围物体的距离,等等。机器人设备110还可以包括驱动单元115,例如,机器人设备110可以被部署在可移动的基座之上,并且驱动单元115可以驱动基座的轮子来按照期望的路径移动。
物理环境160可以包括一个或者多个采集单元130、…、以及132,例如,可以在房间中部署一个或者多个图像采集设备以便从各个角度采集房间的图像。物理环境160可以包括控制设备140,该控制设备140可以经由网络(未示出)控制一个或者多个采集单元130、…、以及132,等等。备选地和/或附加地,在智能家居环境下,控制设备140可以控制物理空间160中的各种电器设备。
备选地和/或附加地,可以提供机器学习模型(例如,模型150)来管理物理空间160。应当理解,尽管图1示出了模型150位于物理空间160内部,备选地和/或附加地,该模型150可以位于物理空间160之外的远程设备处,并且控制设备140、机器人设备110或者其他设备可以经由网络来访问远程的模型150。
模型150可以包括一个或多个模型。如果模型150包括多个模型,这多个模型可以包括多个类型的模型。模型150例如可以至少包括语言模型(LM)和动作模型。语言模型通过从大量语料中学习,能够具备问答能力。动作模型可以控制机器人设备110来执行各种动作。模型150例如还可以包括图像识别模型、文本识别模型等等。
如图1所示,用户120可以指示机器人设备110来操作物理空间110中的各种对象。在此,对象可以是家居环境中的各种物品,例如,用户120可以指令机器人设备110来在物理空间160中寻找某个对象;又例如,用户120可以指令机器人设备110来将找到对象放置到指定位置,等等。
执行任务的概要
为了至少部分地解决现有技术中的不足,根据本公开的一个示例性实现方式,提出了一种用于执行用户任务的方法。参见图2描述根据本公开的一个示例性实现方式的概要,该图2示出了根据本公开的一些实现方式的用于执行用户任务的框图200。
如图2所示,物理空间160中的机器人设备110可以接收来自用户120的用户任务210。此时,用户任务210可以指示机器人设备110来分类物理空间160中的指定范围(例如范围230,也可以称之为第一范围)内的多个对象。例如,用户120可以以自然语言说出“整理桌面上的物品”。在图2的示例中第一范围为“桌面”,多个对象也即桌面上放置的多个对象,包括对象222和对象224。机器人设备110可以获取包括多个对象的图像。例如,可以经由采集单元114、130、…、 以及132中的至少任一项来获取图像。
针对多个对象中的第一对象,基于图像来确定第一对象在物理空间中的第一目的地位置。第一对象可以是多个对象中任一适当的对象。第一目的地位置可以是第一对象应该被放置的位置。例如,在家居环境下,桌面上包括一个瓶装水(也即对象222)和一本书(也即对象224)。对象222应当被放置在冰箱、厨房或者其他位置,因此,如果对象222为第一对象,第一目的地位置可以为冰箱(例如目的地位置250)和/或厨房。书应当被放置在书架、书房或者其他位置,因此,如果书为第一对象,第一目的地位置可以为书架和/或书房。此时,可以指示机器人设备110来将第一对象移动至第一目的地位置。例如,可以指示机器人设备110将对象222移动至目的地位置250。
根据本公开的一些实现方式,可以在具有计算能力的任何计算设备处执行上文描述的方法。例如,可以利用部署在机器人设备110处的应用程序来执行上文描述的方法。备选地和/或附加地,可以在控制设备140处部署应用程序,以便执行上述方法。具体地,可以调用模型150的强大处理能力,以便找到可能包括多个对象的第一范围以及多个对象中每个对象在物理空间中的目的地位置。继而,机器人设备110可以沿路径242移动至范围230,并沿路径244将范围230内的多个对象中的对象222移动至其在物理空间160中的目的地位置250。
利用本公开的示例性实现方式,可以控制机器人设备将第一范围内的多个对象移动至对应的目的地位置。可以提高机器人设备在复杂环境下执行任务的灵活度和精确度,进而完成预期的用户任务。
执行任务的详细过程
已经描述了根据本公开的一些实现方式的概要,在下文中,将描述有关执行用户任务的更多细节。为了便于描述,在下文中仅以控制机器人设备110整理桌面上的物品作为示例,来描述执行用户任务的更多细节。
根据本公开的一些实现方式,图像可以来自以下至少任一项:机器人设备处的采集设备、第一物理空间中的采集设备、第二物理空间中的采集设备。参见图3描述图像采集的更多细节,该图3示出了根据本公开的一些实现方式的图像采集过程的框图300。如图3所示,可以从机器人设备110处的采集单元114获取物理空间160的图像(例如,一个或者多个图像310)。由于机器人设备110可以在物理空间160中自由移动,因而采集单元114可以采集物理空间中的各个位置的图像,进而便于寻找第一范围以及多个对象。
备选地和/或附加地,可以从采集单元130、…、以及132来获取物理空间160的图像。在此,采集单元130、…、以及132可以被预先部署在物理空间160内的指定位置,例如,天花板的拐角位置,等等。以此方式,可以获得以俯视角度拍摄的物理空间160的图像,进而便于从整体上了解物理空间160的布局。这可以帮助确定第一范围以及第一范围内的第一对象对应的第一目的地位置。
根据本公开的一些实现方式,可以获取机器人设备110所在的物理空间160的图像,并基于图像来定位第一范围。例如,若用户120以自然语言说出“整理桌面上的物品”,则该用户任务210指示机器人设备110分类桌面上的多个对象。可以基于物理空间160的图像,来确定桌面所处的范围230。例如,可以获取用于定位第一范围的提示词:“请在如下图像中确定桌面的位置”,并且利用模型来定位范围230。进而可以指示机器人设备110移动至范围230。
根据本公开的一些实现方式,可以在获取机器人设备110所在的物理空间160的图像的同时,获取第一范围内的多个对象的图像。示例性地,可以基于物理空间160的图像来确定第一范围内的多个对象的图像。例如,可从物理空间160的图像中裁剪出第一范围内的多个对象的图像。备选地或者附加地,根据本公开的一些实现方式,还可以响应于机器人设备110移动至第一范围,获取第一范围内的多个对象的图像。示例性地,可以响应于机器人设备110移动至第一范围, 借助机器人设备110处的采集单元114来获取第一范围内的多个对象的图像。
根据本公开的一些实现方式,多个对象的图像可以包括一个或多个图像。如果多个对象的图像包括多个图像,每个图像可以对应一个对象。例如,多个对象的图像可以包括多个对象各自对应的多个图像,针对第一对象,可以调整第一对象的位置,以便获取第一对象的图像。此时,可以调整用于采集图像的采集设备的位置(例如,可以寻找避免遮挡的角度),以便获取第一对象的图像。例如,针对范围230内的多个对象,可以调整对象222的位置,和/或移动机器人设备110的位置,以便机器人设备110借助采集单元114来获取仅包括对象222的图像。由此,可以分别获取多个对象中每个对象对应的图像,可以提高每个对象对应的图像的准确性,进而提高后续基于图像来执行任务的准确性。
如果多个对象的图像仅包括一个图像,可以确定多个对象之间是否存在遮挡关系。可以采用任意适当的方式来确定多个对象之间是否存在遮挡关系。例如,可以借助模型150来确定多个对象之间是否存在遮挡关系。又例如,还可以基于任意适当的规则或算法来确定多个对象之间是否存在遮挡关系。备选地或者附加地,还可以向用户120提供提示信息以由用户120来确定多个对象之间是否存在遮挡关系。可以接收用户120提供的确定结果,并基于该确定结果来确定多个对象之间是否存在遮挡关系。备选地和/或附加地,如果存在遮挡关系,可以指示用户消除该遮挡关系。
如果多个对象之间不存在遮挡关系,可以直接获取包括多个对象的图像。如果多个对象之间存在遮挡关系,可以响应于确定多个对象之间存在遮挡关系,指示机器人设备110移动多个对象中的至少一个对象,进而获取包括多个对象的图像。参见图4描述更多细节,该图4示出了根据本公开的一些实现方式的移动对象的过程的框图400。如图4所示,对象402位于对象401前侧并且遮挡住了对象401。此 时,可以指示机器人设备来将对象402从位置430移动至不遮挡对象401的位置(如图像420中的位置430’)。
根据本公开的一些实现方式,可以确定目标位置,并且指示机器人设备来将对象402移动至目标位置。例如,可以借助模型150来生成动作,以便控制机器人设备来将对象402从位置430移动至位置430’。以此方式,可以支持机器人设备在复杂环境中处理复杂问题,从而以更为准确的方式执行用户任务。
获取到多个对象对应的图像后,可以基于图像来确定第一对象在物理空间中的第一目的地位置。根据本公开的一些实现方式,可以基于图像,确定第一对象的第一类型。关于确定第一对象的第一类型的具体方式,可以采用任意适当的方式来确定第一对象的第一类型。例如,可以基于预先存储的数据库来确定第一对象的第一类型。预先存储的数据库中可以存储有多个对象以及这多个对象各自对应的类型。可以在这个预先存储的数据库中检索第一对象,以确定第一对象的第一类型。
又例如,还可以从图像中识别(例如借助模型150来对图像进行识别)分别与多个对象相关联的多个文本项(例如,对象上的标签)。针对第一对象,进而可以基于第一对象相关联的第一文本项,确定第一类型。以对象222为例,对象222为瓶装水,与对象222相关联的文本项例如可以为瓶装水包装上的文本项。可以基于瓶装水包装上的文本项确定对象222为瓶装水。
根据本公开的一些实现方式,还可以确定物理空间中的第一类型的第二对象的位置,以便基于第二对象的位置来确定第一目的地位置。可以理解,可以采用任意适当的方式来确定第二对象的位置。例如,可以基于机器人设备110所处物理空间160的图像来确定第二对象的位置。以第一对象为瓶装水为例,可以获取机器人设备110所处物理空间160的图像,并基于这个图像来确定物理空间160中其他瓶装水的位置。以其他瓶装水都被放置在厨房为例,可以将第二对象的位置 确定为厨房。又例如,还可以基于预定的规则或算法来确定第二对象的位置。依旧以第一对象为瓶装水为例,若预定的规则指示瓶装水被放置在冰箱中,则可以基于预定规则确定第二对象的位置为冰箱。可以理解,还可以借助模型(例如模型150)来确定第二对象的位置,借助用户120来人工确定第二对象的位置,等等。例如,机器人设备110可以询问用户:“把瓶装水放在哪里”,并且将瓶装水放置到用户指定的位置。
如果要借助模型来确定第二对象的位置,根据本公开的一些实现方式,可以基于物理空间160的图像和第一对象的第一类型,构建提示词,这个提示词可以用于在物理空间160的图像中确定第一类型的对象的位置。以确定了第一对象的第一类型为瓶装水为例,提示词例如可以表示为:“请从如下图像中识别出‘瓶装水’被放置的位置”。
提示词可以被提供至机器学习模型(例如模型150)以利用该机器学习模型来确定第一类型的第二对象的位置。机器学习模型针对提示词的应答也即机器学习模型输出的针对提示词的模型输出,这个模型输出可以指示第二对象的位置(例如,冰箱)。由此,可以基于机器学习模型针对提示词的应答,确定第二对象的位置,也即第一对象的第一目的地位置。
具体地,机器学习模型可以处理图像,并且在图像包括第二对象的情况下,模型可以输出第二对象所在的位置(例如,第二对象在图像中的区域坐标,和/或直接输出第二对象所在区域的图像,等等)。如果图像不包括第二对象,模型可以输出“未找到”等应答。利用本公开的一些实现方式,可以基于多种方式检测图像是否包括第二对象,从而提高机器人设备检测第二图像的位置的性能。
备选地或者附加地,根据本公开的一些实现方式,还可以响应于确定图像表示机器人设备110所处的物理空间160中不包括第二对象,确定与物理空间160相关联的另一物理空间。在这种情况下,可以将机器人设备110所处的物理空间160称之为第一物理空间,将与物理 空间160相关联的另一物理空间称之为第二物理空间。
应当理解,在此的第二物理空间是可能包括第一对象的潜在物理空间。例如在家居环境下,由于瓶装水可能会被放置在冰箱中,因而可以确定第二物理空间为冰箱。具体地,在确定与第一物理空间相关联的第二物理空间的过程中,可以基于物理空间160的图像和第一类型,获取用于定位第二对象的提示词,以及接收机器学习模型针对提示词的应答,以便确定第二物理空间。确定可能包括第二对象的第二物理空间后,可以将第二物理空间在第一物理空间中所处的位置确定为第二对象的位置,也即第一对象的第一目的地位置。
参见图5描述有关确定第二物理空间的更多细节,该图5示出了根据本公开的一些实现方式的调用模型的过程的框图500。如图5所示,可以基于用户任务210,确定要将桌面上的多个对象(例如瓶装水和书)移动至对应的位置。针对瓶装水,可以基于图像310,获取相应的提示词510。提示词510例如可以表示为:“请从以下图像中确定可能包括‘瓶装水’的物理空间”,又例如,提示词510可以表示为:“在以下图像中,‘瓶装水’可能放在哪里”,等等。可以向语言模型520输入提示词510和图像310,以便语言模型520从图像310中找到可能包括瓶装水的第二物理空间。语言模型520例如可以为模型150所包括的一个模型。
根据本公开的一些实现方式,确定第一目的地位置后,可以向用户120提供与第一对象相关联的消息。这个消息可以用于提示用户120确认是否将第一对象移动至第一目的地位置。关于提供消息的具体方式,例如可以经由机器人设备110的显示屏、语音播放设备等来向用户120提供消息。还可以响应于接收到来自用户120的针对消息的应答,基于应答来确定是否将第一对象移动至第一目的地位置。如果应答指示确定将第一对象移动至第一目的地位置,则可以指示机器人设备110将第一对象移动至第一目的地位置。由此,可以在获取用户确认的情况下才移动对象,可以按照用户指示来提高机器人设备执 行任务的准确性。
根据本公开的一些实现方式,还可以基于物理空间160的图像,确定从第一对象的位置移动至第一目的地位置的运动轨迹。可以理解,同样可以采用借助模型150、借助人工、基于预定规则或算法等任意适当方式来确定运动轨迹。这个运动轨迹可以例如可以指示机器人设备110避开物理空间160中的障碍,以较短的路径从第一对象的位置移动至第一目的地位置。确定运动轨迹后,可以指示机器人设备110按照运动轨迹,将第一对象移动至第一目的地位置。
根据本公开的一些实现方式,可以借助模型150来确定机器人设备110要执行的动作以及借助模型150来控制机器人设备110执行动作。例如,若第一目的地位置为冰箱,可以借助模型150来确定将第一对象移动至冰箱的具体方式(例如打开冰箱、放入第一对象等)。示例性地,可以构建相应的提示词以询问模型150如何打开冰箱。提示词例如可以表示为:“从以下图像中确定打开冰箱的方式”,可以向模型150发送提示词和相应的图像(例如物理环境160的包括冰箱的图像)。
模型150例如可以返回:拉动手柄以打开冰箱门。继而,可以指示机器人设备110拉动手柄从而打开冰箱。由此,利用本公开的一些实现方式,可以调用模型的强大处理能力来解决复杂环境下的未知问题,进而确定需要机器人设备执行的动作。以此方式,可以提高机器人设备处理复杂任务的能力,从而以更为准确的方式来执行用户任务。
根据本公开的一些实现方式,可以利用动作模型来确定机器人设备执行的具体动作。参见图6描述更多细节,该图6示出了根据本公开的一些实现方式的调用动作模型的过程的框图600。如图6所示,可以提供动作模型630,该动作模型630可以基于机器人设备的当前状态和指令来确定将要由机器人设备执行的具体动作,该动作模型630可以是经过预训练和微调的模型。动作模型630同样可以为模型150所包括的模型。
当前状态620可以包括多个方面的数据,例如,机器人设备110的图像、机器人设备110的环境图像、机器人手臂的姿态数据(如,机器人手臂的各个关节的位置(POS1,…))、以及被固定在机器人手臂的末端的工具(例如,夹具,刀具,等等)的状态。例如,可以使用0来表示夹具的关闭状态,并且使用1来表示夹具的开放状态。可以将指令和当前状态输入至动作模型630,进而利用动作模型,基于指令和当前状态来确定由机器人设备110将要执行的动作。在此,动作可以表示机器人设备110的当前姿态与下一姿态之间的差异,以及工具的当前状态与下一状态之间的差异,等等。
可以向动作模型630输入指令610(例如,“打开冰箱”),在此,指令610可以是以自然语言表示的,并且可以从语言模型630的应答中确定该指令610。进一步,可以获取机器人设备110的当前状态,动作模型630可以基于输入的数据来确定相应的动作640。例如,可以确定手臂中的各个关节、和/或机器人设备的轮子和/或其他可移动装置在下一时间点的朝向、位置、速度、加速度,等等。进一步,可以利用确定的动作640来控制机器人设备110在下一时间点的状态。
以此方式,可以借助模型来控制机器人设备的动作,可以基于模型输出来指示机器人设备移动第一对象,可以精确地控制机器人设备的动作,进而以更高的效率执行用户任务。
可以理解,可以确定第一范围内的多个对象各自对应的多个目的地位置。可以采用上述类似方式来控制机器人设备110依次将多个对象分别移动至对应的目的地位置。例如,可以控制机器人设备110依次将瓶装水移动至冰箱,再将书移动至书架,等等。需要注意的是,如果多个对象中存在至少两个对象为同一类型的对象,可以将这至少两个对象一同移动。例如,若多个对象包括第一类型的一组对象,一组对象包括至少两个对象,则可以指示机器人设备110一次性移动这一组对象。由此,可以根据多个对象中不同对象的类型,一次移动一组对象,可以提高机器人设备110移动对象的效率。
备选地或者附加地,根据本公开的一些实现方式,如果确定多个对象中的第一类型的一组对象的数量满足阈值条件,可以指示机器人设备110获取第三对象。阈值条件例如可以指示一组对象的阈值数目。例如,若阈值数目为4,则可以响应于第一类型的一组对象所包括的对象的数目达到4个,指示机器人设备110获取第三对象。第三对象例如可以是用于移动一组对象的其他对象。第三对象可以是预先确定好的对象。
例如,用户120可以预先设置机器人设备110可以借助指定的对象来移动多个对象。例如,以一组对象为一组瓶装水为例,第三对象可以是托盘、筐、袋子、拖车等任意可以帮助移动这一组瓶装水的对象。进一步地,可以指示机器人设备110经由第三对象来将一组对象移动至第一目的地位置。例如,若第三对象为托盘,可以指示机器人设备110借助托盘将一组对象(例如一组瓶装水)移动至第一目的地位置。
根据本公开的一些实现方式,将第一类型的至少一个对象移动至第一目的地位置后,可以确定第二类型的至少一个对象的目的地位置并将第二类型的至少一个对象移动至第二类型的对象在物理空间中的第二目的地位置。例如可以借助模型150来确定离开第一目的地位置的方式、重新前往第一范围的运动轨迹、从第一范围移动至第二目的地的运动轨迹,等等。可以采集新的图像并且构造新的提示词,以便询问模型150(例如语言模型)下一指令。提示词例如可以表示为:“请基于如下图像确定下一指令”、“下一步怎么办”,等等。示例性地,当机器人设备已经将瓶装水放入冰箱后,语言模型可以基于当前接收到的图像返回“关闭冰箱”。此时可以基于指令“关闭冰箱”和机器人设备的当前状态,生成相应的动作以便指示机器人设备关闭冰箱。
根据本公开的一些实现方式,还可以确定第一范围内的多个对象的优先级,并基于优先级来依次移动多个对象。可以采用任意方式来 确定多个对象的优先级,例如,可以借助模型150或预定的规则来确定多个对象的优先级。示例性地,需要冷冻的对象(例如冰糕、冻肉等)的优先级会高于可以常温存储的对象的优先级。以此方式,可以先移动对应优先级较高的对象,提高机器人设备110执行任务的质量。
应当理解,尽管上文以中文语言环境为示例描述根据本公开的一个示例实现方式。备选地和/或附加地,可以在多种语言环境中执行根据本公开的一个示例实现方式的技术方案。例如,可以在中文、英文、日文、法文等环境中控制机器人。具体地,可以基于机器学习技术所提供的多语言能力,来在不同语言的应用环境中控制机器人。进一步,尽管上文以取回瓶装水作为示例描述了利用机器人设备执行用户任务的过程,备选地和/或附加地,可以控制机器人设备来执行其他用户任务,例如,在房间中寻找其他物品,将某个物品放置到指定位置,等等。
根据本公开的一些实现方式,用户可以经由语言、动作、手势等来与机器人设备交互。例如,用户可以说出期望执行的用户任务,预先定义某个动作来指定用户任务,等等。具体地,用户可以做出摆放物品的动作,可以以此动作来作为触发机器人设备执行分类放置物品的用户任务。当从采集到的图像序列中识别出该动作时,机器人设备可以自动询问用户是否需要分类放置物品并且询问用户待处理的范围,在获得肯定答复的情况下,机器人设备可以执行该用户任务。
备选地和/或附加地,用户可以经由交互单元112来与机器人设备交互,例如,用户输入以文字和/或图像表示的任务,并且控制机器人设备执行该任务。备选地和/或附加地,用户可以指定任务的执行条件,例如,立即执行任务,在预定时间之后执行任务,或者在确定满足预定条件(例如,在用户吃饭后)时执行任务,等等。
根据本公开的一些实现方式,机器人设备可以向用户提供多种消息,例如,针对第一对象(例如瓶装水),若物理空间中包括多个第一目的地位置(例如冰箱和厨房),可以询问用户需要将瓶装水移动 至冰箱还是移动至厨房。又例如,假设机器人设备没有找到第一目的地位置,机器人设备可以询问用户将第一对象移动至哪里,等等。
根据本公开的一些实现方式,可以利用多种定位算法来确定机器人设备、以及各个对象在物理环境中的位置。例如,可以在机器人设备处部署全球定位系统(GPS),并且使用卫星信号来确定机器人设备的精确位置。备选地和/或附加地,可以在机器人设备处部署通信单元,借助于该通信单元与基站之间的信号,并利用通信网络来确定机器人设备的位置。备选地和/或附加地,在物理空间中可以部署Wi-Fi接入点,机器人设备处的通信单元可以与Wi-Fi热点交互,以便经由Wi-Fi信号强度和已知Wi-Fi接入点的位置来确定位置。备选地和/或附加地,机器人设备处的通信单元可以支持蓝牙功能,此时可以使用蓝牙信号和已知的蓝牙设备位置来确定附近设备的位置。可以在机器人设备处部署惯性导航系统,并且使用加速度计和陀螺仪来测量和计算设备在空间中的移动和方向,进而确定机器人设备的位置。
备选地和/或附加地,可以使用视觉定位系统来确定机器人设备和/或各个对象的位置。可以预先获取物理空间的地图,并且在该地图中标注各个对象的位置。机器人设备可以利用回波检测单元来检测与周围对象之间的距离,并且结合采集的图像和物理空间的地图,来确定各个对象的具体位置。具体地,可以使用计算机辅助设计(CAD)和地理信息系统(GIS),并且利用定位算法来确定位置。备选地和/或附加地,可以在物理空间中的重要对象处部署跟踪单元,例如,可以在家用电器的遥控器(例如,电视遥控器、空调遥控器)处添加跟踪单元,以便机器人设备可以及时获取重要对象的精确位置,等等。
根据本公开的一些实现方式,可以基于上文描述的方法来确定机器人设备自身的原始位置和期望去往的目的地位置。机器人设备可以确定从原始位置到目的地位置的路径。例如,可以不断获取周围的环境图像,并且在确保躲避障碍的情况下,不断更新该路径,并且使得机器人设备沿着该路径移动至目的地位置。
根据本公开的一些实现方式,在到达目的地位置之后,机器人设备可以执行指定的任务。例如,可以获取指定的对象并且将该对象移动至相应位置。可以利用语言模型和/或知识库来确定约束条件,也即在执行任务期间应当遵循的约束条件。例如,可以获取图像和相应提示词,并且向语言模型输入该图像和提示词,进而从语言模型接收约束条件。例如,可以确定提示词:“请基于如下图像,确定在移动XXX对象期间应当遵循的约束条件”,或者“请确定移动XXX对象期间的注意事项”,等等。
此时,可以确定在移动对象(例如,瓶装水、盘子、碗等)期间,应当保持对象的原始姿态(例如,保持竖直方向,不会被倾斜)。进一步,可以向动作模型输入约束条件,此时,动作模型输出的一系列动作将会在确保约束条件的情况下,执行相应的任务。利用本公开的一些实现方式,可以确保机器人设备操作期间的安全性,从而避免造成意外损坏某个对象,等等。
利用本公开的示例性实现方式,机器人设备可以在复杂的物理空间中执行用户任务。以此方式,机器人设备可以自行将第一范围内的多个对象移动至对应的目的地位置。可以提高机器人设备在复杂环境下执行任务的灵活度和精确度,进而完成预期的用户任务。
示例过程
图7示出了根据本公开的一些实现方式的用于执行用户任务的方法700的流程图。在框710处,接收来自用户的用户任务,用户任务指示机器人设备来分类物理空间中的第一范围内的多个对象。在框720处,获取包括多个对象的图像。在框730处,针对多个对象中的第一对象,基于图像来确定第一对象在物理空间中的第一目的地位置。在框740处,机器人设备将第一对象移动至第一目的地位置。
根据本公开的一些实现方式,获取图像包括:响应于确定多个对象之间存在遮挡关系,机器人设备移动多个对象中的至少一个对象; 以及获取包括多个对象的图像。
根据本公开的一些实现方式,第一对象的图像是基于以下来确定的:调整第一对象的位置,以便获取第一对象的图像;以及调整用于采集图像的采集设备的位置,以便获取第一对象的图像。
根据本公开的一些实现方式,方法700进一步包括:获取机器人设备所在的物理空间的图像;基于图像来定位第一范围;以及机器人设备移动至第一范围。
根据本公开的一些实现方式,确定第一目的地位置包括:基于图像,确定第一对象的第一类型;确定物理空间中的第一类型的第二对象的位置;以及基于第二对象的位置,确定第一目的地位置。
根据本公开的一些实现方式,确定第一对象的第一类型进一步包括:从图像中识别分别与多个对象相关联的多个文本项;针对多个对象中的第一对象,基于多个文本项中的与第一对象相关联的第一文本项,确定第一类型。
根据本公开的一些实现方式,确定第二对象的位置包括以下至少任一项:在物理空间的图像中识别第二对象,以确定第二对象的位置。
根据本公开的一些实现方式,确定第二对象的位置包括以下至少任一项:在物理空间的图像和第一对象的第一类型,构建提示词,提示词用于在物理空间的图像中确定第一类型的对象的位置;以及基于机器学习模型针对提示词的应答,确定第二对象的位置。
根据本公开的一些实现方式,机器人设备将第一对象移动至第一目的地位置包括:基于物理空间的图像,确定从第一对象的位置移动至第一目的地位置的运动轨迹;以及机器人设备按照运动轨迹,将第一对象移动至第一目的地位置。
根据本公开的一些实现方式,机器人设备将第一对象移动至第一目的地位置包括:向用户提供与第一对象相关联的消息;以及响应于接收到来自用户的针对消息的应答,机器人设备将第一对象移动至第一目的地位置。
根据本公开的一些实现方式,机器人设备将第一对象移动至第一目的地位置包括:响应于确定多个对象中的第一类型的一组对象的数量满足阈值条件,机器人设备获取第三对象;机器人设备经由第三对象来将一组对象移动至第一目的地位置。
示例装置和设备
图8示出了根据本公开的一些实现方式的用于执行用户任务的装置800的框图。该装置800包括:接收模块810,被配置为接收来自用户的用户任务,用户任务指示机器人设备来分类物理空间中的第一范围内的多个对象;获取模块820,被配置为获取包括多个对象的图像;确定模块830,被配置为针对多个对象中的第一对象,基于图像来确定第一对象在物理空间中的第一目的地位置;以及执行模块840,被配置为使得机器人设备将第一对象移动至第一目的地位置。
根据本公开的一些实现方式,获取模块820进一步被配置用于:响应于确定多个对象之间存在遮挡关系,使得机器人设备移动多个对象中的至少一个对象;以及获取包括多个对象的图像。
根据本公开的一些实现方式,第一对象的图像是基于以下来确定的:调整第一对象的位置,以便获取第一对象的图像;以及调整用于采集图像的采集设备的位置,以便获取第一对象的图像。
根据本公开的一些实现方式,获取模块820进一步被配置用于:获取机器人设备所在的物理空间的图像;基于图像来定位第一范围;以及使得机器人设备移动至第一范围。
根据本公开的一些实现方式,确定模块830进一步被配置用于:基于图像,确定第一对象的第一类型;确定物理空间中的第一类型的第二对象的位置;以及基于第二对象的位置,确定第一目的地位置。
根据本公开的一些实现方式,确定模块830进一步被配置用于:从图像中识别分别与多个对象相关联的多个文本项;针对多个对象中的第一对象,基于多个文本项中的与第一对象相关联的第一文本项, 确定第一类型。
根据本公开的一些实现方式,确定模块830进一步被配置用于:在物理空间的图像中识别第二对象,以确定第二对象的位置。
根据本公开的一些实现方式,确定模块830进一步被配置用于:在物理空间的图像和第一对象的第一类型,构建提示词,提示词用于在物理空间的图像中确定第一类型的对象的位置;以及基于机器学习模型针对提示词的应答,确定第二对象的位置。
根据本公开的一些实现方式,执行模块840进一步被配置用于:基于物理空间的图像,确定从第一对象的位置移动至第一目的地位置的运动轨迹;以及使得机器人设备按照运动轨迹,将第一对象移动至第一目的地位置。
根据本公开的一些实现方式,执行模块840进一步被配置用于:向用户提供与第一对象相关联的消息;以及响应于接收到来自用户的针对消息的应答,使得机器人设备将第一对象移动至第一目的地位置。
根据本公开的一些实现方式,执行模块840进一步被配置用于:响应于确定多个对象中的第一类型的一组对象的数量满足阈值条件,使得机器人设备获取第三对象;使得机器人设备经由第三对象来将一组对象移动至第一目的地位置。
图9示出了能够实施本公开的多个实现方式的设备900的框图。应当理解,图9所示出的计算设备900仅仅是示例性的,而不应当构成对本文所描述的实现方式的功能和范围的任何限制。图9所示出的计算设备900可以用于实现上文描述的方法。
如图9所示,计算设备900是通用计算设备的形式。计算设备900的组件可以包括但不限于一个或多个处理器或处理单元910、存储器920、存储设备930、一个或多个通信单元940、一个或多个输入设备950以及一个或多个输出设备960。处理单元910可以是实际或虚拟处理器并且能够根据存储器920中存储的程序来执行各种处理。在多处理器系统中,多个处理单元并行执行计算机可执行指令,以提高计 算设备900的并行处理能力。
计算设备900通常包括多个计算机存储介质。这样的介质可以是计算设备900可访问的任何可以获得的介质,包括但不限于易失性和非易失性介质、可拆卸和不可拆卸介质。存储器920可以是易失性存储器(例如寄存器、高速缓存、随机访问存储器(RAM))、非易失性存储器(例如,只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、闪存)或它们的某种组合。存储设备930可以是可拆卸或不可拆卸的介质,并且可以包括机器可读介质,诸如闪存驱动、磁盘或者任何其他介质,其可以能够用于存储信息和/或数据(例如用于训练的训练数据)并且可以在计算设备900内被访问。
计算设备900可以进一步包括另外的可拆卸/不可拆卸、易失性/非易失性存储介质。尽管未在图9中示出,可以提供用于从可拆卸、非易失性磁盘(例如“软盘”)进行读取或写入的磁盘驱动和用于从可拆卸、非易失性光盘进行读取或写入的光盘驱动。在这些情况中,每个驱动可以由一个或多个数据介质接口被连接至总线(未示出)。存储器920可以包括计算机程序产品925,其具有一个或多个程序模块,这些程序模块被配置为执行本公开的各种实现方式的各种方法或动作。
通信单元940实现通过通信介质与其他计算设备进行通信。附加地,计算设备900的组件的功能可以以单个计算集群或多个计算机器来实现,这些计算机器能够通过通信连接进行通信。因此,计算设备900可以使用与一个或多个其他服务器、网络个人计算机(PC)或者另一个网络节点的逻辑连接来在联网环境中进行操作。
输入设备950可以是一个或多个输入设备,例如鼠标、键盘、追踪球等。输出设备960可以是一个或多个输出设备,例如显示器、扬声器、打印机等。计算设备900还可以根据需要通过通信单元940与一个或多个外部设备(未示出)进行通信,外部设备诸如存储设备、显示设备等,与一个或多个使得用户与计算设备900交互的设备进行 通信,或者与使得计算设备900与一个或多个其他计算设备通信的任何设备(例如,网卡、调制解调器等)进行通信。这样的通信可以经由输入/输出(I/O)接口(未示出)来执行。
根据本公开的示例性实现方式,提供了一种计算机可读存储介质,其上存储有计算机可执行指令,其中计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,还提供了一种计算机程序产品,计算机程序产品被有形地存储在非瞬态计算机可读介质上并且包括计算机可执行指令,而计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,提供了一种计算机程序产品,其上存储有计算机程序,程序被处理器执行时实现上文描述的方法。
这里参照根据本公开实现的方法、装置、设备和计算机程序产品的流程图和/或框图描述了本公开的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理单元,从而生产出一种机器,使得这些指令在通过计算机或其他可编程数据处理装置的处理单元执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
可以把计算机可读程序指令加载到计算机、其他可编程数据处理装置、或其他设备上,使得在计算机、其他可编程数据处理装置或其他设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其他可编程数据处理装置、或其他设备上执行的指令实现 流程图和/或框图中的一个或多个方框中规定的功能/动作。
附图中的流程图和框图显示了根据本公开的多个实现的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段或指令的一部分,模块、程序段或指令的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
以上已经描述了本公开的各实现,上述说明是示例性的,并非穷尽性的,并且也不限于所公开的各实现。在不偏离所说明的各实现的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实现的原理、实际应用或对市场中的技术的改进,或者使本技术领域的其他普通技术人员能理解本文公开的各个实现方式。

Claims (15)

  1. 一种用于执行用户任务的方法,包括:
    接收来自用户的用户任务,所述用户任务指示机器人设备来分类物理空间中的第一范围内的多个对象;
    获取所述包括所述多个对象的图像;
    针对所述多个对象中的第一对象,基于所述图像来确定所述第一对象在所述物理空间中的第一目的地位置;以及
    所述机器人设备将所述第一对象移动至所述第一目的地位置。
  2. 根据权利要求1所述的方法,其中获取所述图像包括:
    响应于确定所述多个对象之间存在遮挡关系,所述机器人设备移动所述多个对象中的至少一个对象;以及
    获取包括所述多个对象的所述图像。
  3. 根据权利要求1所述的方法,其中所述第一对象的图像是基于以下来确定的:
    调整所述第一对象的位置,以便获取所述第一对象的所述图像;以及
    调整用于采集图像的采集设备的位置,以便获取所述第一对象的所述图像。
  4. 根据权利要求1所述的方法,进一步包括:
    获取所述机器人设备所在的物理空间的图像;
    基于所述图像来定位所述第一范围;以及
    所述机器人设备移动至所述第一范围。
  5. 根据权利要求1所述的方法,其中确定所述第一目的地位置包括:
    基于所述图像,确定所述第一对象的第一类型;
    确定所述物理空间中的所述第一类型的第二对象的位置;以及
    基于所述第二对象的位置,确定所述第一目的地位置。
  6. 根据权利要求5所述的方法,其中确定所述第一对象的所述第一类型进一步包括:
    从所述图像中识别分别与所述多个对象相关联的多个文本项;
    针对所述多个对象中的第一对象,基于所述多个文本项中的与所述第一对象相关联的第一文本项,确定所述第一类型。
  7. 根据权利要求5所述的方法,其中确定所述第二对象的位置包括以下至少任一项:在所述物理空间的图像中识别所述第二对象,以确定所述第二对象的位置。
  8. 根据权利要求5所述的方法,其中确定所述第二对象的位置包括以下至少任一项:
    在所述物理空间的图像和所述第一对象的第一类型,构建提示词,所述提示词用于在所述物理空间的所述图像中确定所述第一类型的对象的位置;以及
    基于机器学习模型针对所述提示词的应答,确定所述第二对象的所述位置。
  9. 根据权利要求1所述的方法,其中所述机器人设备将所述第一对象移动至所述第一目的地位置包括:
    基于所述物理空间的图像,确定从所述第一对象的位置移动至所述第一目的地位置的运动轨迹;以及
    所述机器人设备按照所述运动轨迹,将所述第一对象移动至所述第一目的地位置。
  10. 根据权利要求1所述的方法,其中所述机器人设备将所述第一对象移动至所述第一目的地位置包括:
    向所述用户提供与所述第一对象相关联的消息;以及
    响应于接收到来自所述用户的针对所述消息的应答,所述机器人设备将所述第一对象移动至所述第一目的地位置。
  11. 根据权利要求5所述的方法,其中所述机器人设备将所述第一对象移动至所述第一目的地位置包括:
    响应于确定所述多个对象中的所述第一类型的一组对象的数量满足阈值条件,所述机器人设备获取第三对象;
    所述机器人设备经由所述第三对象来将所述一组对象移动至所述第一目的地位置。
  12. 一种用于执行用户任务的装置,包括:
    接收模块,被配置为接收来自用户的用户任务,所述用户任务指示机器人设备来分类物理空间中的第一范围内的多个对象;
    获取模块,被配置为获取所述包括所述多个对象的图像;
    确定模块,被配置为针对所述多个对象中的第一对象,基于所述图像来确定所述第一对象在所述物理空间中的第一目的地位置;以及
    执行模块,被配置为使得所述机器人设备将所述第一对象移动至所述第一目的地位置。
  13. 一种电子设备,包括:
    至少一个处理单元;以及
    至少一个存储器,所述至少一个存储器被耦合到所述至少一个处理单元并且存储用于由所述至少一个处理单元执行的指令,所述指令在由所述至少一个处理单元执行时使所述电子设备执行根据权利要求1至11中任一项所述的方法。
  14. 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序在被处理器执行时使所述处理器实现根据权利要求1至11中任一项所述的方法。
  15. 一种计算机程序产品,包括计算机程序,其中所述计算机程序在被处理器执行时实现根据权利要求1至11中任一项所述的方法。
PCT/CN2024/106255 2024-07-18 2024-07-18 用于执行用户任务的方法、装置、设备和介质 Pending WO2026016140A1 (zh)

Priority Applications (3)

Application Number Priority Date Filing Date Title
EP24841099.5A EP4706904A4 (en) 2024-07-18 2024-07-18 METHOD AND APPARATUS FOR PERFORMING A USER TASK, AND DEVICE AND SUPPORT
PCT/CN2024/106255 WO2026016140A1 (zh) 2024-07-18 2024-07-18 用于执行用户任务的方法、装置、设备和介质
CN202480003521.4A CN121712620A (zh) 2024-07-18 2024-07-18 用于执行用户任务的方法、装置、设备和介质

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/106255 WO2026016140A1 (zh) 2024-07-18 2024-07-18 用于执行用户任务的方法、装置、设备和介质

Publications (1)

Publication Number Publication Date
WO2026016140A1 true WO2026016140A1 (zh) 2026-01-22

Family

ID=98436610

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/106255 Pending WO2026016140A1 (zh) 2024-07-18 2024-07-18 用于执行用户任务的方法、装置、设备和介质

Country Status (3)

Country Link
EP (1) EP4706904A4 (zh)
CN (1) CN121712620A (zh)
WO (1) WO2026016140A1 (zh)

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170076222A1 (en) * 2015-09-14 2017-03-16 International Business Machines Corporation System and method to cognitively process and answer questions regarding content in images
CN209142945U (zh) * 2018-11-28 2019-07-23 中国人民解放军陆军工程大学 一种智能化搬运码垛实训机器人装置
CN112476443A (zh) * 2020-11-16 2021-03-12 刘玲玲 自动食材分类调度机器人及其控制方法
CN115500755A (zh) * 2022-08-30 2022-12-23 广州江水青智能科技有限公司 一种餐厅餐具收拾-清洗智能系统及其控制方法
CN116709962A (zh) * 2020-11-30 2023-09-05 克拉特博特股份有限公司 杂物清理机器人系统
CN117140476A (zh) * 2023-01-17 2023-12-01 浙江水利水电学院 一种家庭智能整理机器人
CN117992629A (zh) * 2024-02-09 2024-05-07 北京有竹居网络技术有限公司 用于从图像中识别对象的方法、装置、设备和介质
CN118163096A (zh) * 2024-03-06 2024-06-11 京东科技控股股份有限公司 用于机器人控制的方法、装置、设备和存储介质

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10572775B2 (en) * 2017-12-05 2020-02-25 X Development Llc Learning and applying empirical knowledge of environments by robots
US20240227190A9 (en) * 2021-03-04 2024-07-11 Tutor Intelligence, Inc. Robotic system

Patent Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170076222A1 (en) * 2015-09-14 2017-03-16 International Business Machines Corporation System and method to cognitively process and answer questions regarding content in images
CN209142945U (zh) * 2018-11-28 2019-07-23 中国人民解放军陆军工程大学 一种智能化搬运码垛实训机器人装置
CN112476443A (zh) * 2020-11-16 2021-03-12 刘玲玲 自动食材分类调度机器人及其控制方法
CN116709962A (zh) * 2020-11-30 2023-09-05 克拉特博特股份有限公司 杂物清理机器人系统
CN115500755A (zh) * 2022-08-30 2022-12-23 广州江水青智能科技有限公司 一种餐厅餐具收拾-清洗智能系统及其控制方法
CN117140476A (zh) * 2023-01-17 2023-12-01 浙江水利水电学院 一种家庭智能整理机器人
CN117992629A (zh) * 2024-02-09 2024-05-07 北京有竹居网络技术有限公司 用于从图像中识别对象的方法、装置、设备和介质
CN118163096A (zh) * 2024-03-06 2024-06-11 京东科技控股股份有限公司 用于机器人控制的方法、装置、设备和存储介质

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See also references of EP4706904A4 *

Also Published As

Publication number Publication date
CN121712620A (zh) 2026-03-20
EP4706904A1 (en) 2026-03-11
EP4706904A4 (en) 2026-03-11

Similar Documents

Publication Publication Date Title
TWI662388B (zh) 機器人之避障控制系統及方法
US11618164B2 (en) Robot and method of controlling same
EP4547451A1 (en) Robot control based on natural language instructions and on descriptors of objects that are present in the environment of the robot
US20220314432A1 (en) Information processing system, information processing method, and nonvolatile storage medium capable of being read by computer that stores information processing program
WO2022259600A1 (ja) 情報処理装置、情報処理システム、および情報処理方法、並びにプログラム
US20230245476A1 (en) Location discovery
Memmesheimer et al. RoboCup@ Home 2024 OPL winner NimbRo: Anthropomorphic service robots using foundation models for perception and planning
WO2026016139A1 (zh) 用于执行用户任务的方法、装置、设备和介质
Militaru et al. Object handling in cluttered indoor environment with a mobile manipulator
EP4706904A1 (en) Method and apparatus for executing user task, and device and medium
US11164002B2 (en) Method for human-machine interaction and apparatus for the same
WO2026016133A1 (zh) 用于执行用户任务的方法、装置、设备和介质
Deguchi et al. Enhanced robot navigation with human geometric instruction
WO2026016142A1 (zh) 用于执行用户任务的方法、装置、设备和介质
WO2026055826A1 (zh) 获取物品方法、装置、设备和介质
WO2026016134A1 (zh) 用于执行用户任务的方法、装置、设备和介质
WO2026016144A1 (zh) 用于执行用户任务的方法、装置、设备和介质
CN113176775B (zh) 对移动的机器人进行控制的方法以及机器人
WO2026016137A1 (zh) 用于执行用户任务的方法、装置、设备和介质
WO2026055824A1 (zh) 用于确定机器人设备的动作的方法、装置、设备和介质
CN119897856B (zh) 机器人的安全导航与操作方法、装置和电子设备
WO2026016136A1 (zh) 用于执行用户任务的方法、装置、设备和介质
US20240153230A1 (en) Generalized three dimensional multi-object search
Nabeuchi et al. Experiments on picking system using multimodal instructions for mobile assisted robot using Mixed Reality and voice activation
WO2026007183A1 (zh) 障碍物识别方法、装置以及自移动机器人的控制方法

Legal Events

Date Code Title Description
ENP Entry into the national phase

Ref document number: 2024841099

Country of ref document: EP

Effective date: 20250128

WWP Wipo information: published in national office

Ref document number: 2024841099

Country of ref document: EP