WO2026038902A1 - Procédé, appareil et système de mise en œuvre d'un outil de visualisation de données - Google Patents

Procédé, appareil et système de mise en œuvre d'un outil de visualisation de données

Info

Publication number
WO2026038902A1
WO2026038902A1 PCT/KR2025/012348 KR2025012348W WO2026038902A1 WO 2026038902 A1 WO2026038902 A1 WO 2026038902A1 KR 2025012348 W KR2025012348 W KR 2025012348W WO 2026038902 A1 WO2026038902 A1 WO 2026038902A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
computing device
processor
data set
information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/KR2025/012348
Other languages
English (en)
Korean (ko)
Inventor
이주행
이정원
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Pebblous Inc
Original Assignee
Pebblous Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from KR1020250019947A external-priority patent/KR102925436B1/ko
Application filed by Pebblous Inc filed Critical Pebblous Inc
Publication of WO2026038902A1 publication Critical patent/WO2026038902A1/fr
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/904Browsing; Visualisation therefor
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/21Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
    • G06F18/213Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/44Arrangements for executing specific programs
    • G06F9/451Execution arrangements for user interfaces
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/042Knowledge-based neural networks; Logical representations of neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/10Interfaces, programming languages or software development kits, e.g. for simulating neural networks

Definitions

  • the present disclosure relates to data processing technology for diagnosing and visualizing data. More specifically, it relates to technology for diagnosing data through data imaging and providing interactive functions by visualizing the data.
  • One task of the present disclosure is to diagnose the quality of a large data set.
  • one task of the present disclosure is to visualize a data set so that it can be effectively expressed while maintaining the inherent structure of the data through a vectorization method for precisely identifying the inherent distribution of a large data set.
  • a computing device comprising a memory and at least one processor electronically connected to the memory, wherein the at least one processor is configured to execute at least one instruction stored in the memory, and is configured to perform the following operations: acquiring a first data set; selecting a first level from among a plurality of levels classified according to a data processing method based on a user input; determining at least one property of a first data processing model corresponding to the first level, the first data processing model including at least one artificial intelligence model; providing, through a first GUI (Graphical User Interface), prior information related to a task of processing the first data set using the first data processing model; and constructing the first data processing model based on a user input to the first GUI.
  • GUI Graphic User Interface
  • a data processing method may be provided that is set to perform, by at least one processor executing at least one instruction stored in a memory, an operation of acquiring a first data set, an operation of selecting a first level from among a plurality of levels classified according to a data processing method based on a user input, an operation of determining at least one property of a first data processing model corresponding to the first level, the first data processing model including at least one artificial intelligence model, an operation of providing, through a first GUI (Graphical User Interface), prior information related to a task of processing the first data set using the first data processing model, and an operation of constructing the first data processing model based on a user input to the first GUI.
  • GUI Graphic User Interface
  • an electronic device including a display configured to display at least one GUI (Graphic User Interface), a memory, and at least one processor electronically connected to the memory, wherein the at least one processor configured to execute at least one instruction stored in the memory is configured to perform an operation of receiving a user input for selecting a first level among a plurality of levels classified according to a data processing method, an operation of displaying a first GUI (Graphic User Interface) through the display that indicates prior information related to a task of processing the first data set using a first data processing model corresponding to the first level, the first data processing model including at least one artificial intelligence model, an operation of obtaining a diagnosis result for the first data set based on a first vector set defined in a first-dimensional embedding area from the first data processing model, and an operation of providing the first vector set to a first visualization tool to display a two-dimensional or three-dimensional first data image, the first data image corresponding to the first data set, through the display.
  • GUI Graphic User Interface
  • an operation of obtaining a data set by at least one processor executing at least one instruction stored in a memory, an operation of obtaining a data set, an operation of embedding the data set in a specific dimension to obtain a vector set, the vector set including a plurality of vectors corresponding to a plurality of data included in the data set, an operation of screening the vector set by calculating at least one feature value based on the vectors included in the vector set using a screener having at least one metric set, an operation of identifying at least one vector whose at least one feature value satisfies a predetermined condition, an operation of obtaining a data image including a plurality of data points representing the data set in two dimensions or three dimensions by processing the vector set using at least one visualization tool, an operation of determining a first area including at least one data point corresponding to the at least one vector on the data image and generating tag information associated with the first area, and a first scene including the first area, the first scene being generated by capturing the first area from a specific view
  • the computer program includes an operation of acquiring a data set, an operation of embedding the data set in a specific dimension to acquire a vector set, the vector set including a plurality of vectors corresponding to a plurality of data included in the data set, an operation of screening the vector set by calculating at least one feature value based on the vectors included in the vector set using a screener having at least one metric set, an operation of identifying at least one vector whose at least one feature value satisfies a predetermined condition, an operation of processing the vector set using at least one visualization tool to acquire a data image including a plurality of data points representing the data set in two dimensions or three dimensions, an operation of determining a first region including at least one data point corresponding to the at least one vector on the data image and generating tag information associated with the first region, and a first scene including the first region, the first scene being generated by capturing the first region from
  • a computing device comprising a memory and at least one processor electronically connected to the memory, wherein the at least one processor is configured to execute at least one instruction stored in the memory, and is configured to perform an operation of providing a first data image corresponding to a first data set through a first view port, the first data image including a plurality of data points corresponding to each of data included in the first data set; an operation of generating first snapshot information including a first scene for the first data image being provided through the first view port in response to a user input received through a first GUI and providing the first snapshot information through a second view port; and an operation of generating link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port.
  • a data interaction method including an operation of providing a first data image corresponding to a first data set, the first data image including a plurality of data points corresponding to each of data included in the first data set, through a first view port by at least one processor executing at least one instruction in a memory, an operation of generating first snapshot information including a first scene for the first data image being provided through the first view port in response to a user input received through a first GUI and providing the first snapshot information through a second view port, and an operation of generating link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port.
  • an electronic device comprising a display configured to display at least one GUI (Graphic User Interface), a memory, and at least one processor electronically connected to the memory, wherein at least one processor configured to execute at least one instruction stored in the memory is configured to perform an operation of providing a first data image corresponding to a first data set, the first data image including a plurality of data points corresponding to each of data included in the first data set, through a first view port of the display, an operation of providing first snapshot information including a first scene for the first data image being provided through the first view port through the display in response to a user input received through the first GUI, and an operation of providing link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port.
  • GUI Graphic User Interface
  • a method including an operation of visualizing a data set by at least one processor executing at least one instruction stored in a memory to provide a first data image corresponding to the data set, an operation of receiving a user input for a first area on the first data image, an operation of determining a latent code corresponding to the first area and providing the determined latent code to a generation model to generate synthetic data, an operation of inputting the synthetic data to a data processing model, the data processing model being trained to embed data into a specific dimension, to obtain a synthetic vector corresponding to the synthetic data, and an operation of visualizing the synthetic vector to provide a synthetic point corresponding to the synthetic data on the first data image.
  • a computing device including a display configured to display at least one GUI (Graphic User Interface), a memory, and at least one processor electronically connected to the memory, wherein the at least one processor configured to execute at least one instruction stored in the memory is configured to perform an operation of visualizing a data set and providing a first data image corresponding to the data set through the display, an operation of receiving a user input for a first area on the first data image, an operation of determining a latent code corresponding to the first area and providing the determined latent code to a generation model to generate synthetic data, an operation of inputting the synthetic data into a data processing model, the data processing model being trained to embed data into a specific dimension, to obtain a synthetic vector corresponding to the synthetic data, and an operation of visualizing the synthetic vector and providing a synthetic point corresponding to the synthetic data on the first data image through the display.
  • GUI Graphic User Interface
  • a computing device can automatically diagnose the quality of a large data set and intuitively visualize it. This allows data analysts, AI model developers, and researchers to more easily understand the inherent structure of the data and develop strategies to improve data quality.
  • a user can more efficiently explore and manipulate data during the data visualization process, and utilize it to improve the performance of a machine learning model.
  • FIG. 1 is a diagram illustrating a configuration of a computing device according to various embodiments.
  • FIG. 2 is a diagram illustrating various data processing methods included in a data clinic service provided by a computing device according to various embodiments.
  • FIG. 3 is a diagram illustrating various systems for providing data clinic services according to various embodiments, and artificial intelligence models and algorithms for constructing the systems.
  • FIG. 4 is a diagram illustrating a method for a computing device to provide a data image according to various embodiments.
  • FIG. 5 is a diagram illustrating a method for a computing device to obtain characteristics of a data set according to various embodiments.
  • FIG. 6 is a diagram illustrating a data lens processing system and a data imaging system according to various embodiments.
  • FIG. 7 is a diagram illustrating a system in which a computing device visualizes data and provides user interaction functions, according to various embodiments.
  • FIG. 8 is a diagram illustrating a data visualization and interaction method according to various embodiments.
  • FIG. 9 is a drawing for explaining detailed steps performed in an imaging application step according to various embodiments.
  • FIG. 10 is a diagram illustrating the results of data imaging and diagnosis according to the level of the lens and the visualization tool according to various embodiments.
  • FIG. 11 is a diagram illustrating a lens builder included in a computing device according to various embodiments.
  • FIG. 12 is a diagram illustrating a method for a computing device to build a data processing model for data imaging and diagnosis based on user input, according to various embodiments.
  • FIG. 14 is a diagram illustrating an example of a computing device constructing a first data processing model according to various embodiments.
  • FIG. 17 is a diagram illustrating a method of imaging and visualizing a data set using a first data processing model built on a computing device according to various embodiments.
  • FIG. 20 is a diagram illustrating a method for a computing device to generate synthetic data based on user input according to various embodiments.
  • FIG. 22 is a diagram illustrating a method for a computing device to remove data based on user input, according to various embodiments.
  • FIG. 23 is a diagram illustrating a method for a computing device to perform data improvement in response to a data improvement request and provide visual interaction therefor, according to various embodiments.
  • FIG. 24 is a diagram illustrating an example of a computing device in which a data processing method including a snapshot function is implemented, according to various embodiments.
  • FIG. 25 is a diagram illustrating a method for a computing device to screen a data set to generate snapshot information, according to various embodiments.
  • FIG. 26 is a diagram illustrating a method for a computing device to provide snapshot information based on user input according to various embodiments.
  • Figure 27 is an example of a screen provided by a computing device.
  • FIG. 28 is a diagram illustrating a function of a computing device to reproduce snapshot information according to various embodiments.
  • FIG. 29 is a diagram illustrating a method for a computing device to generate diagnostic reports and snapshot information in conjunction with each other, according to various embodiments.
  • each of the phrases “A or B”, “at least one of A and B”, “at least one of A or B”, “A, B, or C”, “at least one of A, B, and C”, and “at least one of A, B, or C” may include any one of the items listed together in that phrase, or all possible combinations thereof.
  • a device-readable storage medium may be provided in the form of a non-transitory storage medium.
  • non-transitory simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
  • the term 'unit' as used in this disclosure means a software or hardware component such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC).
  • the 'unit' performs specific roles, but is not limited to software or hardware.
  • the 'unit' may be configured to reside on an addressable storage medium and may be configured to play one or more processors. Accordingly, according to some embodiments, the 'unit' includes components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables.
  • the functionality provided within the components and 'units' may be combined into a smaller number of components and 'units' or further separated into additional components and 'units'. Additionally, the components and ' ⁇ parts' may be implemented to activate one or more CPUs within a device or secure multimedia card. Furthermore, according to various embodiments of the present disclosure, the ' ⁇ parts' may include one or more processors.
  • FIG. 1 is a diagram illustrating a configuration of a computing device according to various embodiments.
  • a computing device e.g., an electronic device including a computing means such as a server or client device, hereinafter referred to as a “computing device”
  • a computing device may include a processor (110), a memory (120), a storage device (130), a communication circuit (140), and a bus (not shown).
  • the configuration of the computing device (100) is not limited to the configuration illustrated in FIG. 1 or the configuration described above, and may further include hardware or software configurations included in general computing devices or mobile devices.
  • the processor (110) may include at least one processor, at least some of which are implemented to provide different functions.
  • the processor (110) may execute software (e.g., a program) to control at least one other component (e.g., a hardware or software component) of the computing device (100) connected to the processor (110) and perform various data processing or calculations.
  • the processor (110) may store instructions or data received from other components in the memory (120) (e.g., a volatile memory), process the instructions or data stored in the volatile memory, and store the resulting data in the non-volatile memory.
  • the processor (110) may include a main processor (e.g., a central processing unit or an application processor) or an auxiliary processor (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith.
  • a main processor e.g., a central processing unit or an application processor
  • an auxiliary processor e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor
  • the coprocessor may be configured to use less power than the main processor or to be specialized for a given function.
  • the coprocessor may be implemented separately from the main processor or as a part thereof.
  • the coprocessor may control at least a portion of functions or states associated with at least one component (e.g., the display (240) or a communication circuit) of the computing device (100), for example, on behalf of the main processor while the main processor is in an inactive (e.g., sleep) state, or together with the main processor while the main processor is in an active (e.g., application execution) state.
  • the coprocessor e.g., an image signal processor or a communication processor
  • the coprocessor may be implemented as a part of another functionally related component (e.g., a communication circuit).
  • the coprocessor e.g., a neural network processing device
  • the operation of the computing device (100) described below may be understood as the operation of the processor (110).
  • the memory (120) may include at least one memory, at least some of which are implemented to provide different functions.
  • the memory (120) may store various data used by at least one component (e.g., the processor (110)) of the computing device (100).
  • the data may include, for example, software (e.g., a program) and input data or output data for instructions related thereto.
  • the memory (120) may include volatile memory or non-volatile memory.
  • the memory (120) may be implemented to store an operating system, middleware or applications, and/or the artificial intelligence model described above.
  • the memory (120) may include a plurality of instructions (121) that direct the operations of the processor (110) to implement the functions provided by the service.
  • the processor (110) may execute at least some of the plurality of instructions stored in the memory (120).
  • the computing device (110) may include a software server including the processor (110) that executes the functions provided by the service based on at least some of the plurality of instructions.
  • the storage device (130) can provide a mass storage device to the computing device (100).
  • the storage device (130) can be a computer-readable medium.
  • the storage device (130) can be a floppy disk device, a hard disk device, an optical disk device, a tape device, a flash memory or other similar solid-state memory device, or an array of devices including a storage area network or other configuration device.
  • a computer program product is explicitly embodied in an information medium.
  • the computer program product includes instructions that, when executed, perform one or more methods as described above.
  • the information medium is a computer-readable medium or a machine-readable medium, such as the memory (120), the storage device (130), or the memory of the processor (110).
  • the storage device (130) may include a database (DB).
  • the storage device (130) may include a database having a pre-structured data structure.
  • the computing device (110) may store data sets having interrelated relationships in the database.
  • a computing device can provide services based on various artificial intelligence frameworks performed by at least one processor and a memory electronically connected to at least one processor.
  • the memory (120) or storage device (130) may store at least one artificial intelligence model implementing various types of artificial intelligence (or machine learning) frameworks that can be trained to perform a given task.
  • artificial intelligence or machine learning
  • support vector machines, decision trees, neural networks, etc. are just a few examples of machine learning frameworks used in various applications such as image processing and natural language processing.
  • Some artificial intelligence frameworks, such as neural networks may utilize layers of nodes that perform specific operations.
  • a neural network may include an input layer, an output layer, and one or more intermediate layers. Each node may process its inputs according to a predefined function and provide output to subsequent layers, or in some cases, previous layers. The input to a particular node may be multiplied by a weight value corresponding to the edge between the input and the node. Additionally, each node may have a separate bias value used to generate the output. Various learning procedures can be applied to learn the edge weights and/or bias values (parameters).
  • a neural network architecture may have multiple layers that perform different specific functions. For example, one or more node layers may collectively perform specific operations, such as pooling, encoding, or convolution operations.
  • layer may refer to a group of nodes that share inputs and outputs, such as communicating with external sources or other layers of the network.
  • calculation may refer to a function that can be performed by one or more node layers.
  • model structure may refer to the overall architecture of a layered model, including the number of layers, the connectivity of the layers, and the types of operations performed by individual layers.
  • neural network structure may refer to the model structure of a neural network.
  • trained model and/or “tuned model” may refer to the model structure along with the parameters for the trained or tuned model structure.
  • two trained models may have different values for parameters while sharing the same model structure, such as when they are trained on different training data or when the training process has an underlying probabilistic process.
  • Transfer learning is a broad approach for training models with limited task-specific training data for a specific task.
  • transfer learning a model is first pretrained on another task for which valuable training data is available, and then adapted to a specific task using task-specific training data.
  • pre-training refers to training a model on a pre-training dataset to adjust model parameters in a manner that allows subsequent adjustments of those model parameters to tailor the model to one or more specific tasks.
  • pre-training may involve a self-supervised learning process on unlabeled training data, where the "self-supervised” learning process involves learning from the structure of pre-training examples in the absence of explicit (e.g., manually provided) labels.
  • tuning subsequent modification of the model parameters obtained through pre-training is referred to herein as "tuning.” Tuning may be performed for one or more tasks using supervised learning on explicitly labeled training data, and in some cases, a task different from pre-training may be used for tuning.
  • a communication bus may be a configuration for electronically (or communicatively) connecting multiple components included in a computing device. That is, each component may be interconnected using various buses and mounted on a common motherboard or in another suitable manner.
  • the input/output interface may include an input interface connected to an input device to receive an input signal, or an output interface connected to an output device to output an output signal.
  • the computing device (100) may further include at least one communication circuit (140) for communicating with an external device.
  • the communication circuit (140) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the computing device (100) and an external computing device, and the performance of communication through the established communication channel.
  • the communication circuit may operate independently from the processor (110) (e.g., a program processor) and may include one or more communication processors (e.g., communication chips) that support direct (e.g., wired) communication or wireless communication.
  • the communication circuit (140) may include a wireless communication module (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (e.g., a local area network (LAN) communication module, or a power line communication module).
  • a wireless communication module e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module
  • GNSS global navigation satellite system
  • wired communication module e.g., a local area network (LAN) communication module, or a power line communication module.
  • the corresponding communication module can communicate with an external computing device via a first network (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)).
  • a first network e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)
  • a second network e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)
  • a second network e.g., a long-
  • the wireless communication module can identify or authenticate the computing device (100) within a communication network such as the first network or the second network by using subscriber information stored in the subscriber identification module (e.g., an international mobile subscriber identity (IMSI)).
  • the wireless communication module can support a 5G network subsequent to a 4G network and next-generation communication technologies, such as new radio access technology (NR).
  • NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high-reliability and low-latency communications (URLLC (ultra-reliable and low-latency communications)).
  • the wireless communication module can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate.
  • the wireless communication module can support various technologies to secure performance in the high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna.
  • MIMO massive multiple-input and multiple-output
  • FD-MIMO full dimensional MIMO
  • array antenna analog beam-forming
  • large scale antenna or large scale antenna.
  • the wireless communication module can support various requirements specified in a computing device (100), an endoscope device, or a network system.
  • the wireless communication module can support a peak data rate (e.g., 20 Gbps or more) for eMBB implementation, a loss coverage (e.g., 164 dB or less) for mMTC implementation, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL) each, or 1 ms or less for round trip) for URLLC implementation.
  • a peak data rate e.g., 20 Gbps or more
  • a loss coverage e.g., 164 dB or less
  • U-plane latency e.g., 0.5 ms or less for downlink (DL) and uplink (UL) each, or 1 ms or less for round trip
  • the computing device (100) may be implemented to include at least some of the above-described components (processor, communication circuitry, memory, display).
  • a user device may be implemented to include a processor, communication circuitry, memory, sensors, and a display.
  • a server device may be implemented to include a processor, communication circuitry, and memory.
  • FIG. 2 is a diagram illustrating various data processing methods included in a data clinic service provided by a computing device according to various embodiments.
  • the data clinic service may include various data processing methods. These various data processing methods may be encoded and stored in the memory of the computing device, and at least one processor included in the computing device may be configured to execute at least one encoded instruction. Specifically, at least one processor may process a received input data set based on the various data processing methods and output an output data set.
  • a computing device may perform, but is not limited to, an operating method for data imaging, an operating method for data enhancement, an operating method for data generation, an operating method for data feature extraction, or an operating method for data evaluation.
  • each of the above-described operating methods can be performed based on operating algorithms of at least one processor included in the computing device.
  • a computing device may perform, but is not limited to, a data imaging algorithm, a data enhancement algorithm, a data generation algorithm, a data feature extraction algorithm, or a data evaluation algorithm.
  • each operation method and algorithm are arbitrarily named according to the output results for the convenience of explanation, so each operation method or algorithm is only defined based on the operations performed by the processor, and the name of the operation method or algorithm itself does not limit the invention.
  • a computing device can process an input data set according to a data imaging algorithm to generate an image for the input data set.
  • a computing device can improve data by processing an input data set according to a data improvement algorithm, and can generate a result of the improvement.
  • a computing device can generate synthetic data by processing an input data set according to a data generation algorithm.
  • a computing device can process an input data set according to a data feature extraction algorithm to extract a property of the input data set.
  • a computing device can process an input data set according to a data evaluation algorithm to evaluate the quality of the input data set.
  • the computing device can perform the various operation methods or algorithms described above in parallel, sequentially, or selectively. Specifically, the computing device can use the same input data as input values for different algorithms in parallel, can use the result values output by a specific algorithm sequentially as input values for another algorithm, or can selectively perform some of the algorithms among a plurality of algorithms according to a predetermined method.
  • the various operation methods or algorithms for the data clinic described above can be performed in a deep learning model included in a computing device according to various embodiments of the present disclosure.
  • the computing device according to various embodiments of the present disclosure may include one deep learning model for performing the various operation methods or algorithms described above, but is not limited thereto, and may include multiple deep learning models for performing each of the operation methods or algorithms described above, or may include one or more deep learning models for performing at least some of the various operation methods or algorithms described above.
  • Figure 3 is a diagram illustrating various systems for providing data clinic services according to various embodiments, as well as artificial intelligence models and algorithms for constructing such systems.
  • a system may refer to a system that includes at least one software or hardware configuration to perform a specific function.
  • a computing device may include a data clinic system composed of various artificial intelligence (or neural network, machine learning, etc.) models to provide clinic services.
  • various artificial intelligence or neural network, machine learning, etc.
  • the computing device may include, but is not limited to, a data imaging system, a data diagnostic system, and a data treatment system.
  • the data imaging system may include, but is not limited to, a lens processing model for determining an optimal dimension for representing the characteristics of the data, an imaging model for obtaining a data image reflecting the inherent characteristics of the data, or a visualization model for visually representing the data.
  • the data diagnosis system may include, but is not limited to, a diagnosis model for diagnosing at least one characteristic of data or a quality assessment model for evaluating the quality of data.
  • the data treatment system may include, but is not limited to, a synthetic model (or generation model) for generating targeted virtual data (or synthetic data) as needed, a data diet model for removing at least a portion of the data, or a data correction model for adjusting the characteristics of at least a portion of the data.
  • a synthetic model or generation model
  • a data diet model for removing at least a portion of the data
  • a data correction model for adjusting the characteristics of at least a portion of the data.
  • a module may include multiple hardware components for implementing an artificial intelligence model that performs a specific function.
  • a module may include, but is not limited to, an encoder, a decoder, a generator, a discriminator, an adapter, a natural language processing module, or a large language model (LLM).
  • the computing device can store the plurality of modules described above, and can construct an AI framework based on at least some of the modules to obtain an AI model for a data clinic.
  • a data lens included in a data imaging system can be implemented as an AI model including at least one encoder or at least one adapter, but is not limited thereto.
  • FIG. 4 is a diagram illustrating a method for a computing device to provide a data image according to various embodiments.
  • a computing device may receive a data set and provide an Image of Data (IOD).
  • IOD Image of Data
  • the data set may be data of dimension M (M>0).
  • the data set may be a data set defined on an M-dimensional input space (310).
  • the data set may be a single-modality data set.
  • the data set may be an image data set.
  • the data set may be a text data set.
  • the data set may be a collection of data with different modalities.
  • the data set may be an image data set including annotation information.
  • the data set may be a mixed data set of images and text.
  • a computing device can receive and process data of all modalities that can be used for deep learning, such as time series data sets, sensor data sets, as well as the image data and text data described above, as input data sets.
  • the data image (IOD) provided by the computing device may be an image that processes an input data set and displays it in an imaging space (320).
  • the image does not mean a 2D image, but is a general expression that visually represents data.
  • the imaging space (320) is a concept that includes a 2D space, a 3D space, and an N-dimensional virtual space, and means a space in which a data image provided according to an embodiment appears.
  • the computing device may output an output that displays the data image in a 2D or 3D imaging space, but is not limited thereto.
  • the computing device can provide a data image through the output device.
  • the computing device can provide the data image by outputting the data image through a display.
  • the imaging space (320) may be a screen of a display.
  • the computing device can provide the data image by outputting the data image through a printing device.
  • the imaging space (320) may be paper output by the printing device.
  • the computing device may provide a data image via the external device.
  • the imaging space (320) may be a display screen of the external device.
  • the server device may provide the data image by transmitting the data image to at least one external device that communicates with the server device via a network connected to the server device.
  • a computing device can obtain a data image based on a vector set (or data point set, point data set, etc.) (330) corresponding to an input data set.
  • the computing device can obtain a vector set by mapping the data included in the input data set to an embedding space (or latent space) of a specific dimension.
  • the computing device can obtain a vector set by identifying a manifold formed by the data set in an embedding space of a specific dimension.
  • the manifold may refer to a shape that the input data set represents in an embedding space of a specific dimension.
  • the manifold may refer to an area where a vector set is identified or a shape formed by a vector set when mapping the input data set to a vector set in an embedding space of a specific dimension.
  • An IOD Information Object Descriptor
  • An IOD Information Object Descriptor
  • the shape or color of the visualized point may vary depending on the embodiment, and therefore the term "point" itself is not intended to limit the invention.
  • a point may be expressed using various terms depending on the embodiment. For example, a point may be expressed using terms such as a vector or feature appearing in an embedding space or latent space, but is not limited thereto.
  • the computing device can obtain a vector set by mapping the data set to an N-dimensional embedding space based on a predefined condition defined by a mapping function (e.g., a pre-stored matrix for mapping to an embedding space of a specific dimension). For example, the computing device can obtain a vector set by encoding the data set, but is not limited thereto. For example, the computing device can input the data set to a pre-trained encoder and obtain a vector set through the output layer of the encoder, but is not limited thereto.
  • a mapping function e.g., a pre-stored matrix for mapping to an embedding space of a specific dimension.
  • embedding refers to the process of converting high-dimensional data into low-dimensional vectors, preserving similarities and structural relationships between data points. This embedding is performed in a way that preserves the core information of the data while increasing computational efficiency.
  • each data point is represented as a vector, and these vectors can reflect the distribution and characteristics of the entire data set. This provides a foundation for analyzing the statistical characteristics and inherent patterns of the data set.
  • a computing device may include a data lens (400) for obtaining a data image (IOD) by embedding and visualizing a data set.
  • the data lens (400) may include at least one processing configuration for processing data.
  • the data lens (400) may include at least one neural network model (e.g., an encoder, etc.) for obtaining a vector set based on the data set and at least one visualization model (e.g., PCA, T-SNE, UMAP, etc.) for visualizing the data set based on the vector set to obtain a data image.
  • at least one neural network model e.g., an encoder, etc.
  • at least one visualization model e.g., PCA, T-SNE, UMAP, etc.
  • the data lens (400) may obtain a vector set corresponding to the data set by embedding the data set in an N-dimensional latent space, and may obtain a data image (IOD) corresponding to the data set by representing the vector set in an M-dimensional (e.g., 2-dimensional or 3-dimensional) imaging space (320).
  • IOD data image
  • FIG. 5 is a diagram illustrating a method for a computing device to obtain characteristics of a data set according to various embodiments.
  • the computing device can process the acquired data set to obtain characteristic information corresponding to the data set.
  • the properties of a data set or data may include information related to the distribution (e.g., geometric distribution or statistical distribution) of the data set or data. Specifically, the properties may include the property values of each data included in the data set. For example, a computing device may obtain property information indicating the distribution of the property values of the data included in the data set. Furthermore, the computing device may obtain the property information of the data set based on the statistical distribution, such as the mean, deviation, or variance, of the property values of each data.
  • the properties of a data set or data may include intrinsic characteristics related to the distribution of the data set itself.
  • the properties of a data set or data may include, but are not limited to, the density, homogeneity, bias, or distribution of the data set or data.
  • the properties of a data set or data may include task-dependent properties related to the task for which the data set is utilized (e.g., classification).
  • the properties of a data set or data may include, but are not limited to, the labeling error rate or the proportion of data pairs that are geometrically adjacent (hard-negative) but belong to different classes.
  • the computing device may store computational metrics corresponding to each characteristic of the data set or data in memory. More specifically, the computing device may store, but is not limited to, metrics for computing the density of the data set or data, metrics for computing the homogeneity of the data set or data, metrics for computing the bias of the data set or data, metrics for computing the distribution of the data set or data, etc.
  • the computing device can acquire data set characteristics based on stored operational metrics, using a data feature extraction algorithm built using an artificial neural network.
  • the feature extraction algorithm can be implemented using a feed-forward neural network.
  • the computing device may include, but is not limited to, a separate neural network for computing characteristics of a data set, or may include a neural network including layers for computing characteristics of a data set.
  • a computing device may include an artificial neural network for feature extraction designed to extract features of a data set when inputted with the data set.
  • the artificial neural network for feature extraction may be an artificial neural network that has undergone transfer learning to compute the features of the data.
  • a computing device can acquire characteristics of a data set by constructing an artificial neural network that adds a layer for extracting data characteristics to a neural network model (e.g., a data lens, an encoder, etc.) for providing data images based on the data set.
  • a neural network model e.g., a data lens, an encoder, etc.
  • the computing device can identify a vector set based on the data set and acquire characteristics of the data set or data based on the identified vector set.
  • the computing device can obtain the characteristic value of each data included in the data set by processing each vector included in the vector set with a predetermined algorithm.
  • the computing device can calculate the characteristic value based on the geometric distribution or statistical distribution of each vector included in the vector set, and can assign the calculated characteristic value to the corresponding data.
  • the characteristic value can be calculated based on the distance between vectors.
  • the characteristic value can be obtained based on the number of vectors existing within a predetermined distance from a specific vector (or data point), but is not limited thereto.
  • the characteristic value can be obtained based on the average value of the distances from a specific vector to a predetermined number of nearby vectors, but is not limited thereto.
  • the computing device can calculate the average distance value based on the distance values from a specific vector to K nearby vectors, and can obtain the first characteristic value (e.g., density, etc.) of the specific vector based on the calculated average distance value, but is not limited thereto.
  • the first characteristic value e.g., density, etc.
  • data quality is a concept that includes both quantitative and qualitative quality. Therefore, for successful training of an artificial intelligence model, it is necessary to (i) secure a sufficient amount of training data to train the artificial intelligence model, (ii) secure training data with high-quality inherent characteristics (e.g., unbiased distribution), and (iii) secure training data with characteristics (e.g., task-dependent properties) appropriate for the training purpose (e.g., task of the artificial intelligence model).
  • characteristics e.g., task-dependent properties
  • a computing device can synthesize, modify (or adjust), or remove data in a way that enhances the inherent characteristics and task-dependent characteristics of the data set to obtain high-quality learning data.
  • the computing device can improve the overall quality of a data set by removing at least some data from the data set.
  • a computing device can improve the learning efficiency of an artificial intelligence model that is learned by appropriately removing at least some data from a data set.
  • FIG. 6 is a diagram illustrating a data lens processing system and a data imaging system according to various embodiments.
  • the computing device (3500) can determine a data lens system that processes the data set to preserve inherent characteristics of the data set based on the data set.
  • a computing device can acquire a lens system corresponding to a data set based on a database. Specifically, the computing device can search the database for a lens system corresponding to the input data set based on the characteristics of the input data set.
  • a computing device can obtain a lens system corresponding to a data set based on a lens processing algorithm. Specifically, the computing device can calculate the optimal dimensionality that preserves the inherent characteristics of the input data set.
  • the computing device can obtain a data image representing the intrinsic characteristics of the data set using the imaging system (3510).
  • This disclosure provides a method for precisely identifying the inherent distribution of data by vectorizing it and intuitively visualizing it in two- or three-dimensional space, enabling the effective analysis and visualization of large data sets.
  • interaction and UI/UX implementation technologies are applied to minimize the gap between high-dimensional vector space and visualization space, enabling users to more easily understand and manipulate data quality.
  • FIG. 7 is a diagram illustrating a system in which a computing device visualizes data and provides user interaction functions, according to various embodiments.
  • the computing device (700) can provide a data image (720) through a network environment (e.g., a web or app environment).
  • the user device (701) can check the data image (720) through the network environment and obtain information about the checked data image (720).
  • the user device (701) may be an electronic device such as a laptop computer, a smartphone, a tablet, etc.
  • the user device (701) may be a wearable device such as a smart watch, smart glasses, smart glasses, etc.
  • the user device (701) may be a media device such as a streaming media device, a media player, an automobile entertainment system, etc.
  • the user device (701) may be an XR (mixed reality) device such as a device including a VR device or AR glasses, etc. That is, the user device (701) may be an electronic device that connects to a network to receive a data visualization service according to various embodiments of the present disclosure.
  • XR mixed reality
  • the user device (701) may be an electronic device that connects to a network to receive a data visualization service according to various embodiments of the present disclosure.
  • FIG. 8 is a diagram illustrating a data visualization and interaction method according to various embodiments.
  • the data visualization and interaction method may include a data imaging application step (S810), a data imaging and diagnosis step (S820), a data visualization step (S830), and an interaction step (S840).
  • the electronic device may request imaging, diagnosis, or visualization of a data set from the server. Specifically, the electronic device may transmit the data set and request information about the data set to the server. The server may process the data set based on the received data set and request information.
  • the server can image the data set by vectorizing it. Specifically, the server can identify a vector set corresponding to the data set by embedding the data set in a latent space of a specific dimension.
  • the server can diagnose the intrinsic properties and task-dependent attributes of the data set by analyzing the distance, neighbor relationship, or cluster structure between individual data points included in the data set based on the identified vector set.
  • the server can analyze the distance distribution from a specific data point to the K nearest neighboring vectors, or identify areas overly concentrated in a specific class to determine bias. Furthermore, to assess homogeneity, the server can measure the diversity or dispersion of the region to which each vector belongs.
  • the server can diagnose task-dependent characteristics by identifying how each class is distributed in vector space (e.g., whether there are boundaries or overlaps between classes) for a data set used in a classification task. This allows the server to detect instances of class imbalance or instances where a particular class is excessively close to another class (i.e., a large number of hard negatives).
  • the server can detect outliers or identify factors that degrade data quality (e.g., incorrect labels, duplicate data, noisy samples, etc.) by considering at least one of the Euclidean distance, cosine similarity, or other distances between vectors. Furthermore, the server can pass these analysis results to subsequent steps (e.g., data cleaning or label correction) to improve the quality of the dataset.
  • degrade data quality e.g., incorrect labels, duplicate data, noisy samples, etc.
  • subsequent steps e.g., data cleaning or label correction
  • the server can visualize and express the data set in a two-dimensional or three-dimensional space.
  • the server can reduce the dimensionality of the data set or the vector set corresponding to the data set to obtain a data image (IOD) and provide the data image in a two-dimensional or three-dimensional space.
  • IOD data image
  • the server can apply dimensionality reduction techniques such as PCA (principal component analysis), t-SNE, and UMAP, or use a neural network-based embedding model to map high-dimensional data into a low-dimensional visualization space.
  • PCA principal component analysis
  • t-SNE t-SNE
  • UMAP neural network-based embedding model
  • the server can provide visualized results that reflect the diagnostic results (cluster structure, outliers, bias, class imbalance, etc.) produced in the previous step (e.g., data diagnosis step).
  • the server can display data points identified as outliers with a separate color or special marker (e.g., triangles, stars, etc.), or visually highlight areas with ambiguous class boundaries or high-density areas (e.g., region borders, gradients), thereby enabling users to grasp data distribution and quality issues at a glance.
  • a separate color or special marker e.g., triangles, stars, etc.
  • visually highlight areas with ambiguous class boundaries or high-density areas e.g., region borders, gradients
  • the server can process inputs related to the visualized data image received from an electronic device (e.g., a user device) to provide a dynamic response to the visualized data image.
  • an electronic device e.g., a user device
  • the user can perform inputs such as touch, mouse click, drag, or pinch zoom on any point within the visualized data image (e.g., a specific data point, cluster, or selected area).
  • the computing device (server) of the present disclosure may provide a response by retrieving metadata (e.g., individual feature values, label information, statistical values, etc.) associated with the data point and transmitting the information to the user device, which may be displayed in a pop-up window, tooltip, or separate layer.
  • metadata e.g., individual feature values, label information, statistical values, etc.
  • the server may further analyze the statistical distribution, density, class distribution, etc. of the data points within the designated area and then provide the results (e.g., mean value, deviation, representative image, number of samples, etc.) to the user device.
  • the server reapplies dimensionality reduction techniques (PCA, t-SNE, UMAP, etc.) or color-shape mapping algorithms, or adjusts parameters to generate an updated data image and deliver it to the user's device.
  • the interaction step (S840) of the present disclosure enables users to intuitively and immediately explore visualized data, allowing them to gain a deeper understanding of the data's inherent structure or to detect potential quality issues (e.g., labeling errors, outliers, unclear boundaries between clusters, etc.) early. Furthermore, this interaction feature can be effectively utilized by machine learning model developers during data cleaning or label correction tasks.
  • FIG. 9 is a drawing for explaining detailed steps performed in an imaging application step according to various embodiments.
  • users can input settings for tools to be used for data imaging and diagnosis (e.g., artificial intelligence models for data imaging, such as lenses), data visualization tools, etc.
  • tools to be used for data imaging and diagnosis e.g., artificial intelligence models for data imaging, such as lenses
  • data visualization tools e.g., data visualization tools, etc.
  • the data imaging application step (S810) may include a lens level selection step (S811), a lens property determination step (S813), an other setting step (S815), and a visualization tool selection step (S817).
  • the user can select a level of a lens including at least one artificial intelligence model to be used for data imaging.
  • the level of the lens can be classified and defined according to preset criteria related to the data imaging method.
  • the level of the lens can include, but is not limited to, a first level that diagnoses only quantitative indicators of the data set (basic statistics, missing value ratio, data outlier detection, etc.), a second level that diagnoses by vectorizing the data set using a pre-stored (pre-trained) artificial intelligence model, or a third level that diagnoses by further learning (fine-tuning) or newly learning an artificial intelligence model optimized for the data set and then vectorizing it.
  • the third level can be selected to directly retrain the model to obtain more precise diagnosis results.
  • the user can determine the properties of a lens including at least one artificial intelligence model to be used for data imaging.
  • the properties of the lens may include the structure of the artificial intelligence model, hyperparameters (e.g., number of layers, number of parameters, learning rate, batch size, etc.), or exploration strategies (e.g., type of optimizer, initialization method).
  • hyperparameters e.g., number of layers, number of parameters, learning rate, batch size, etc.
  • exploration strategies e.g., type of optimizer, initialization method.
  • the user can adjust the "model complexity (number of parameters)" or set the "number of learning epochs" to balance analysis and processing time.
  • the user can also select whether to use a pre-trained model suited to a specific domain (e.g., image, text, structured data, etc.) or a generic model (generic AI model).
  • the user device can set whether to perform data diagnosis by selecting whether to skip the diagnosis process and perform only simple visualization or to perform diagnosis as well. Furthermore, the user device can set whether to visualize the diagnosis results by deciding whether to display cluster density or bias indicators on a separate color scale. Furthermore, the user device can set community creation criteria, such as similarity (distance) criteria between data points or the method of creating subgroups or communities based on specific characteristics (e.g., label information). Furthermore, the user device can set whether to create snapshots that separately store and compare or analyze the data distribution (visualization state) at a certain point in time (viewpoint).
  • the user can select at least one of multiple pre-linked visualization tools (e.g., PCT, UMAP, T-SNE, etc.).
  • the user can select the tool based on its characteristics (e.g., dimensionality reduction method, ease of interpretation of results, processing speed, etc.).
  • a preset visualization dimension e.g., 2D, 3D
  • the user can decide whether to view the data in a simple plane (projection) or to interact with it, such as by rotating or zooming in 3D space.
  • PCA is fast and intuitive to interpret
  • t-SNE and UMAP offer the advantage of more precisely reflecting data clustering structures.
  • Users can choose a visualization tool based on a comprehensive consideration of the characteristics of the dataset, analysis objectives, computing resources, and other factors.
  • the data imaging application step (S810), users can directly input various settings and decisions to control the data imaging and diagnostic process so that it is optimized for their analysis purposes and environment.
  • This user-centric configuration step facilitates the vectorization, diagnostics, and visualization processes performed in subsequent steps (S820, S830, etc.), ultimately contributing to improved data quality and enhanced AI model performance.
  • FIG. 10 is a diagram illustrating the results of data imaging and diagnosis according to the level of the lens and the visualization tool according to various embodiments.
  • FIG. 11 is a diagram illustrating a lens builder included in a computing device according to various embodiments.
  • the computing device can build a lens for imaging a data set according to settings input by the user, and can diagnose and visualize the data set using the built lens and visualization tool.
  • a computing device can build a lens for imaging a data set according to settings input by a user, and can diagnose and visualize the data set using the built lens and visualization tool.
  • a computing device e.g., a server
  • the lens builder is a component for building an appropriate lens (e.g., a set of tools for data imaging and diagnosis) by configuring or combining, learning, or tuning tools (e.g., artificial intelligence models, statistical calculators, etc.) for processing the data set, reflecting the user's settings for imaging and diagnosing the data set.
  • the lens builder can be implemented as a separate, independent hardware device or software module, or can be implemented in a form in which at least one processor executes a plurality of instructions stored in memory for building a lens.
  • the lens builder may include, but is not limited to, a model storage unit including a plurality of artificial intelligence models, a model learning unit for learning the artificial intelligence models, an attribute determination unit for determining attributes of the artificial intelligence models, a calculation tool storage unit including a plurality of calculators for calculating statistical characteristics of data, and a lens storage unit for storing constructed lenses.
  • the calculation tool storage unit may store a plurality of calculators (e.g., a first calculator, a second calculator, ⁇ ), and the lens storage unit may store a plurality of lenses (e.g., a first lens, a second lens, a third lens, a fourth lens, ⁇ ), but is not limited thereto.
  • a plurality of calculators e.g., a first calculator, a second calculator, ⁇
  • the lens storage unit may store a plurality of lenses (e.g., a first lens, a second lens, a third lens, a fourth lens, ⁇ ), but is not limited thereto.
  • the computing device can generate a second lens by loading one or more models that meet user requirements or data formats (images, text, structured data, etc.) from among a plurality of pre-learning models stored in a first storage of the model storage (e.g., a CNN model for computer vision, a Transformer model for text embedding, etc.).
  • a CNN model for computer vision
  • a Transformer model for text embedding, etc.
  • the computing device may provide a vector set corresponding to the data set (e.g., a first vector set or a second vector set) to at least one visualization tool, thereby providing a data image (Image of Data, IOD) visualizing the data set to the user.
  • a vector set corresponding to the data set e.g., a first vector set or a second vector set
  • IOD Image of Data
  • the computing device can generate and output at least one of a plurality of different types of data images (IODs) depending on the type of visualization tool or the dimension visualized by the visualization tool.
  • IODs data images
  • the computing device can project the vector set into two dimensions, obtain a two-dimensional first data image (IOD #1), and then provide it.
  • a first visualization tool e.g., PCA, a 2D-based dimensionality reduction algorithm
  • the dimensionality of the data can be reduced and projected into two, three, or other dimensions using a third visualization tool (e.g., t-SNE, UMAP, etc.) to obtain a third data image (IOD #3).
  • a third visualization tool e.g., t-SNE, UMAP, etc.
  • the "Image of Data (IOD)" in the present disclosure refers to the result of visualizing a vector set, but is not limited to a specific standard (2D/3D) or a specific technique (PCA/t-SNE/UMAP/Autoencoder-based visualization, etc.).
  • the computing device may receive user interactions (e.g., zooming in and out, clicking on a specific point, specifying a range, changing a color scale, etc.) for the visualization result in real time, and update the visualized data image or display additional details (e.g., metadata corresponding to each point).
  • These reports can be delivered to users (e.g., data analysts, AI model developers, etc.) in a structured form, and can be immediately utilized in the analysis or model modification stage by being displayed as charts, tables, or heatmaps in a visualization interface (GUI).
  • GUI visualization interface
  • the first data processing model may include at least one artificial intelligence model.
  • the first data processing model may include at least one pre-stored pre-learning model or an artificial intelligence model trained to be optimized for the first data set. For example, based on an input requesting imaging and diagnosis optimized for the first data set, the at least one processor may generate the first data processing model by training the artificial intelligence model optimized for the first data set, but is not limited thereto.
  • the at least one attribute may include an attribute related to an operation for processing data by the first data processing model.
  • the at least one attribute may include, but is not limited to, the number of parameters or hyperparameters of the artificial intelligence model or graphics card information.
  • FIG. 13 is a diagram for explaining information included in dictionary information according to various embodiments.
  • the prior information may include cost information related to the cost required for processing a data set, result preview information that predicts and displays the results of processing the data set in advance, progress information indicating the progress of processing the data set, reference result information indicating the processing results of cases similar to the data set, or attribute information indicating the properties of the processing model used for processing the data. Accordingly, the user can use the service more smoothly by checking the expected processing cost, predicted results, model progress, etc. in advance at the time of requesting data set imaging and diagnosis.
  • cost information may refer to information related to the cost required to process a data set.
  • cost information may include an estimate of the computational cost required to process a data set (e.g., GPU time, CPU core time, memory usage, etc.) or a payment fee (e.g., credits, points, currency, etc.) that a user must pay to use the processing service.
  • a payment fee e.g., credits, points, currency, etc.
  • “result preview information” may refer to information that briefly presents some key results expected after data set processing, allowing the user to estimate the general form or value of the results.
  • the computing device may briefly display the expected accuracy range after AI model training, a preview of cluster distribution, or a sample visualization image.
  • the result preview information may include data images obtained by simply inputting a data set into a visualization tool. While such data images may not accurately reflect the inherent characteristics of the data set, by providing the visualization results of the data set in advance, they can encourage the user to predict the results.
  • the progress information may refer to information about a processing algorithm or processing status of a data set.
  • the progress information may include the learning progress, training time, or remaining time for optimization learning based on a data set.
  • the user can predict the completion time of processing or determine whether additional resources (e.g., computing resources) are allocated.
  • additional resources e.g., computing resources
  • a warning message may be displayed alongside the progress to provide the user with an opportunity to take action in advance.
  • reference result information may refer to information that indirectly estimates expected results or performance indicators by providing examples of processing results previously performed on cases similar to the data set submitted by the user.
  • reference result information may include imaging and diagnostic results data for reference data similar to the input first data set. For example, it may be provided in the form of "When analyzing 100,000 image data of the same category (domain), the average accuracy was 92%, and the processing time was approximately 4 hours.”
  • attribute information may refer to information about attributes associated with a data processing tool.
  • attribute information may include the number of model parameters or graphics card information.
  • the information may be expressed as, "This processing uses a Transformer architecture-based model (approximately 100 million parameters) and an RTX 3090 GPU.” This information allows users to understand model size, resource compatibility (whether internal or cloud GPU), development environment, etc., and predict model utilization and learning performance in advance.
  • At least one processor may be configured to perform an operation (S1250) of constructing a first data processing model based on a user input to the first GUI.
  • the user input to the first GUI may include an approval input for processing the first data set proposed by the computing device.
  • the computing device can process the first data set only if, prior to processing the first data set, it provides the user with prior information about the processing of the first data set, and if an approval input is received from the user who has been provided with the prior information.
  • FIG. 14 is a diagram illustrating an example of a computing device constructing a first data processing model according to various embodiments.
  • a computing device or at least one processor included in the computing device may be configured to perform an operation (S1401) of obtaining a plurality of representative values from a plurality of adapters based on a first data set.
  • the at least one processor may preprocess the first data set to extract a feature value corresponding to the first data set and input the feature value to each of the plurality of adapters.
  • the at least one processor may obtain a plurality of representative values based on data output from each of the plurality of adapters into which the same data has been input.
  • each adapter that receives feature values can output a vector set corresponding to the first data set.
  • at least one processor can obtain a representative value by operating the output vector set in a predetermined manner.
  • the multiple vector sets output from the multiple adapters can be defined based on different dimensions.
  • the multiple adapters can be configured to output vector sets of different dimensions.
  • At least one processor may be configured to perform an operation (S1402) of selecting a first adapter that satisfies a predetermined condition based on a plurality of representative values.
  • the predetermined condition may include a condition set for determining an adapter optimized for the first data set.
  • at least one processor may select the first adapter by identifying at least one adapter corresponding to the lowest (or highest) value among the plurality of representative values.
  • the first adapter may be implemented to embed the first data set in an optimal dimension representing the first data set.
  • the first adapter may be configured to output a first vector set of a first dimension based on the first data set.
  • At least one processor may be configured to perform an operation (S1404) of constructing a first data processing model including a first adapter. Specifically, at least one processor may construct a first data processing model including a foundation model or a pre-trained model and the first adapter by communicatively connecting the first adapter to a pre-stored foundation model or a pre-trained model.
  • Computing devices can build data processing models optimized for a given dataset by further training or fine-tuning only the adapters based on the dataset. Because computing devices only train adapters on pre-stored AI models, they can minimize computational costs.
  • FIG. 15 is a diagram illustrating another example of a computing device constructing a first data processing model according to various embodiments.
  • the computing device or at least one processor included in the computing device may be set to perform an operation (S1501) of inputting a first data set into a plurality of foundation models.
  • At least one processor may be configured to perform an operation (S1502) of obtaining a plurality of representative values based on data output from a plurality of foundation models.
  • each of the plurality of foundation models may output a vector set corresponding to a first data set, and at least one processor may obtain a representative value by operating the output vector set in a predetermined manner.
  • the plurality of vector sets output from the plurality of foundation models may be defined based on different dimensions.
  • the plurality of foundation models may be configured to output vector sets of different dimensions.
  • At least one processor may be configured to perform an operation (S1503) of determining a first foundation model satisfying a predetermined condition based on a plurality of representative values, and an operation (S1504) of constructing a first data processing model including the first foundation model. Since the specific method of determining the model based on the predetermined condition and the plurality of representative values has been described above, a detailed description thereof will be omitted.
  • FIG. 16 is a diagram illustrating another example of a computing device constructing a first data processing model according to various embodiments.
  • a computing device or at least one processor included in the computing device may perform an operation (S1601) of inputting a first data set into a first foundation model.
  • the first foundation model may include a plurality of nodes implemented to receive the same input and output different result values.
  • the first foundation model may include at least one hidden layer including a plurality of nodes into which the same input is input in parallel.
  • At least one processor may perform an operation (S1602) of obtaining a plurality of representative values from a plurality of nodes included in the first foundation model, wherein the plurality of nodes input feature values for the first data set in parallel. Specifically, the plurality of nodes output a plurality of vector sets based on the feature values for the first data set, and at least one processor may obtain a plurality of representative values from the plurality of vector sets based on a predetermined operation.
  • S1602 an operation of obtaining a plurality of representative values from a plurality of nodes included in the first foundation model, wherein the plurality of nodes input feature values for the first data set in parallel. Specifically, the plurality of nodes output a plurality of vector sets based on the feature values for the first data set, and at least one processor may obtain a plurality of representative values from the plurality of vector sets based on a predetermined operation.
  • At least one processor may be configured to perform an operation (S1603) of activating a first node based on a plurality of representative values. Additionally, at least one processor may be configured to perform an operation (S1604) of constructing a first data processing model including a first foundation model in which the first node is activated.
  • At least one processor can determine a first node that satisfies a predetermined condition based on a plurality of representative values, and activate the first node so that input is only input to the first node.
  • at least one processor can select a node optimized for the first data set among the plurality of nodes, and establish a first data processing model by disconnecting communication with nodes other than the selected node.
  • At least one processor may obtain a first vector set defined in a first-dimensional embedding region from a first data processing model (S1270).
  • Each of the plurality of vectors (or data points) included in the first vector set may correspond to each unit data included in the first data set.
  • At least one processor may provide a first data image corresponding to the first data set by providing the first vector set to the first visualization tool (S1290). At this time, at least one processor may also recommend a visualization tool appropriate for visualizing the first data set based on the first vector set or a plurality of feature values.
  • a computing device may include a visualization tool DB or communicate with a DB server that manages information related to visualization tools in advance, and perform an algorithm for searching and determining an optimal visualization tool based on data characteristics.
  • a computing device or at least one processor included in the computing device may perform a data characteristic extraction step (S1810).
  • the computing device may analyze or extract characteristic values such as a domain, data type, data capacity, vector distribution density, number per class, etc. for a data set (or vector set).
  • at least one processor may extract statistics such as a data domain, data size (number of samples) and number of dimensions, average distance between vectors, variance, number of clusters, etc., or model learning specifications (required computing resources, processing time, etc.) based on the data set.
  • the computing device can perform a visualization tool DB query and SQL query generation step (S1820).
  • the computing device (or DB server) can manage metadata for each visualization tool (PCA, T-SNE, UMAP, Autoencoder-based visualization, etc.) in the visualization tool DB (e.g., in table form).
  • the computing device can generate a SQL query based on a matching rule between data characteristics and tool metadata to search for "candidates that satisfy specific conditions (domain, data size, cluster structure importance, operation time constraints, etc.) among available visualization tools.”
  • the computing device can perform the recommendation algorithm execution and result generation step (S1830). Specifically, when the DB search results are returned in multiple visualization tools, the computing device can apply a score calculation or weight-based algorithm. For example, the computing device can be configured to increase the preference for PCA or UMAP when the amount of data is very large, to increase the score of t-SNE or UMAP when cluster accuracy (detailed clustering) is important, or to prioritize PCA with low computational complexity when real-time interaction is required.
  • the final recommendation priority is determined, and a list such as "t-SNE (1st), UMAP (2nd), PCA (3rd)" can be presented to the user.
  • the computing device can perform the final decision step (S1840) on a visualization tool based on a user request. For example, the user may be informed that "t-SNE is the most suitable tool based on data characteristics and priority criteria," and the user can then decide whether to use the recommended tool as is or select a different tool.
  • the pre-recommendation results allow the user to recognize key pros and cons (e.g., visualization quality vs. processing time) in advance.
  • At least one processor can execute the corresponding algorithm on the first vector set to generate and provide a first data image.
  • other candidate tools e.g., the second and third visualization tools
  • a computing device can reflect the type of data set (image, text, structured data, etc.), domain (medical, social networking services, finance, etc.), data volume, vector distribution characteristics (density, variance, number of clusters, etc.), processing time constraints, etc., into a "visualization tool recommendation" algorithm, thereby guiding a user to select a visualization tool more rationally. Furthermore, by querying the visualization tool database using SQL to identify available candidates and calculating priorities among the candidates, the user can obtain visualization results that enable efficient understanding of a high-dimensional embedding space.
  • the computing device provides an environment in which data distribution or cluster structure can be accurately and intuitively understood by appropriately utilizing the strengths and weaknesses of each technique such as PCA, T-SNE, and UMAP through a process of recommending a visualization tool based on data characteristics.
  • FIG. 19 is a diagram illustrating a method for a computing device to generate synthetic data based on data imaging according to various embodiments.
  • a computing device or at least one processor included in the computing device can generate synthetic data based on a vector set corresponding to a data set.
  • At least one processor can obtain a vector set by imaging a dataset using a lens including at least one artificial intelligence model, and can visualize the dataset using at least one visualization tool to display a data image (IOD) corresponding to the dataset in a visualization space.
  • IOD data image
  • FIG. 20 is a diagram illustrating a method for a computing device to generate synthetic data based on user input according to various embodiments.
  • a computing device or at least one processor included in the computing device may visualize a data set and provide a first data image corresponding to the data set (S2010). For example, by utilizing visualization tools such as the aforementioned PCA, t-SNE, and UMAP, each unit data included in the data set may be projected into a two-dimensional or three-dimensional space, and then a data image (IOD) including multiple data points may be displayed on a GUI (Graphical User Interface).
  • IOD Data image
  • At least one processor can receive a user input for a first area on the first data image (S2020).
  • a user's input such as a click, touch, or drag can be recognized for a first area (R) representing a specific point or range on the first data image (IOD).
  • the user input is an input event that occurs on the GUI, and includes various forms such as a left mouse click, a mobile touch, or a pen drawing.
  • At least one processor can determine a latent code corresponding to the first region (S2030).
  • the latent code is utilized as a model internal embedding or noise vector when generating synthetic data.
  • At this time, at least one processor may request feedback regarding the first region specified by user input.
  • the at least one processor may provide the user with a graphical user interface (GUI) (e.g., a pop-up window or highlighting) that visualizes the first region or displays a confirmation message (e.g., "Is this region the target region for generating synthetic data?").
  • GUI graphical user interface
  • At least one processor may define conditions for data generation based on user input. Specifically, at least one processor may define data generation conditions for determining potential code to be input into the data generation model.
  • At least one processor may determine the latent code in various ways based on the properties of the first region. For example, at least one processor may selectively execute an algorithm that determines the latent code based on the presence or absence of data points or the number of data points contained in the first region.
  • At least one processor may determine a latent code based on at least one vector corresponding to at least one point adjacent to the first region.
  • at least one processor may derive a target region including the first region where the user input was made and at least one point adjacent thereto, and then determine a latent code based on at least one vector corresponding to at least one point included in the target region.
  • At least one processor can determine a latent code based on at least two or more vectors. Specifically, at least one processor can identify at least two or more vectors corresponding to at least one point adjacent to the first region, and determine a latent code corresponding to the first region based on the identified at least two or more vectors. For example, by setting a vector interpolated by vectors corresponding to points near the first region as a latent code, synthetic data reflecting the typical (centroid-like) characteristics of the corresponding region can be generated.
  • a single point is included in the first region, it can be defined as a target vector, and a latent code can be generated based on the relationship (distance, density, label, etc.) between adjacent vectors around this target vector.
  • At least one processor may identify at least one area requiring improvement based on diagnostic results for the data set. At least one processor may visually display the identified at least one area via a display of the user device.
  • At least one processor can activate at least one identified region to enable interaction. Specifically, at least one processor can enable user input for at least one region requiring improvement. In this case, a data improvement process for the region requiring improvement can be performed based on user input regarding the at least one region requiring improvement.
  • At least one processor can extract at least one feature of a plurality of unit data included in the data set based on a vector set that embeds the data set in a latent space of a specific dimension. Based on the at least one extracted feature, the at least one processor can identify data in need of improvement and visually display data points corresponding to the data in need of improvement or a specific region containing the data points.
  • At least one processor can generate synthetic data by further reflecting a user's prompt input.
  • at least one processor can receive a user prompt (PROMPT) through an input interface provided on a GUI (e.g., a text field, voice input, a slider, etc.).
  • PROMPT user prompt
  • the user prompt can express the topic, style, properties, constraints, etc. of the synthetic data in natural language or tag-based.
  • the user prompt can include text instructions in the form of "Generate an image of a cat with a blue background” or “Silhouette format with emphasized perspective.” Additionally, if the user prompt specifies a numeric parameter or categorical tag (e.g., "realistic", "cartoon”), the corresponding value can be interpreted as prompt-internal information and passed to the model.
  • a numeric parameter or categorical tag e.g., "realistic", "cartoon”
  • the generative model can determine the basic shape or feature distribution based on the latent code, additionally reflect the subject, style, or detailed properties through prompts (e.g., text, tags, or parameters), and output the final synthetic data. For example, if the latent code is the result of "interpolation between two vectors," it already reflects intermediate properties (e.g., mixing two classes) or spatial characteristics. If the prompt is given as "cat + blue background + cartoon style,” the style or thematic elements can be reflected through the internal conditioning path of the model, resulting in a synthetic image with the corresponding characteristics.
  • prompts e.g., text, tags, or parameters
  • a prompt (PROMPT) entered into an input interface via a GUI along with a determined latent code can be set as a constraint on data generation.
  • At least one processor can input the latent code and the prompt into a generation model, and the generation model can generate synthetic data conditioned by the instructions given by the prompt based on the latent code.
  • At least one processor can convert the user prompt into a vector form based on a text embedding module (such as CLIP, BERT, GPT) or a conditional network (such as a conditional layer or cross-attention), and then cross-attention it to the hidden layer of the generative model or inject it as a conditional input.
  • a text embedding module such as CLIP, BERT, GPT
  • a conditional network such as a conditional layer or cross-attention
  • At least one processor may be implemented to reflect user requests at each step (e.g., diffusion step, GAN upsampling step, etc.) based on logic that combines latent codes and text embeddings (e.g., concat, add, attention).
  • the user can iteratively (refinement synthesis) by modifying the prompt or resetting the latent code (interpolation parameters, selecting additional vectors, etc.).
  • customized synthetic data can be easily obtained by immediately reflecting the conditions desired by the user, compared to when only a latent code obtained by simply interpolating multiple vectors is used.
  • prompt-based generation is advantageous in generating realistic (distribution-preserving) results, since the latent code already reflects the inherent characteristics of the data set.
  • creative uses e.g., content creation, data augmentation, simulation, etc.
  • artificial intelligence models are facilitated by efficiently testing various scenarios (e.g., adversarial examples, artistic expressions, etc.) through prompt changes.
  • At least one processor can automatically extract or generate a prompt from a user-specified first field, thereby utilizing the prompt as a condition for generating synthetic data. This allows for synthetic results that reflect the inherent properties or metadata of the data, without requiring the user to directly input separate text.
  • At least one processor can automatically generate a prompt by analyzing label information, class/category, metadata (e.g., time, location, ID, etc.), visual features (if the data point is an image), etc., for data points belonging to or adjacent to the first region. Specifically, if it is inferred that the region is a "region of cat faces," the at least one processor can generate a simple phrase such as "cat face cluster” or more specific text such as "Multiple cat faces in a close-up shot.”
  • At least one processor can generate prompts using image captioning technology (e.g., vision-language model, optical character recognition (OCR), etc.), and for structured data, it can automatically convert category names or key attribute names into sentence form.
  • image captioning technology e.g., vision-language model, optical character recognition (OCR), etc.
  • At least one processor may use an image captioning model (e.g., a CNN+LSTM architecture, a Transformer-based visual-language model, etc.) to input an image from a first region and automatically output a sentence-based description (caption).
  • an image captioning model e.g., a CNN+LSTM architecture, a Transformer-based visual-language model, etc.
  • the user may be provided with a GUI interface that allows them to "view and edit the extracted phrases," manually modifying the text as desired and then adopting it as the final prompt.
  • At least one processor can input the determined latent code or other vector interpolation results and automatically generated prompts into a generative model (GAN, VAE, Diffusion, etc.).
  • the generative model can output conditional synthetic data that reflects the data distribution inferred from the latent code, as well as context, object information, and properties extracted from the user domain.
  • the computing device generates linguistic constraints along with latent code, thereby enabling "dimensional expansion" (the generation of data with properties not available with existing data).
  • the computing device can construct an N+3-dimensional data lens by additionally learning synthetic data generated based on linguistic constraints for a data lens that outputs vectors on an N-dimensional latent space.
  • a computing device can perform an interaction operation to remove at least some data included in a data set based on a user input.
  • FIG. 22 is a diagram illustrating a method for a computing device to remove data based on user input, according to various embodiments.
  • a computing device or at least one processor included in the computing device may receive a user input regarding a first region including at least one point on a visualized data image (S2210).
  • the first region may be a dense region (a region with very high data density within a cluster) or an region that a user wishes to remove, such as a region containing a large number of outliers.
  • the at least one processor may provide the user with information regarding a region requiring data removal (e.g., a dense region or an outlier region), and the user may provide input regarding the region.
  • the first region may include at least one cluster whose data characteristics satisfy predetermined conditions.
  • At least one processor may remove at least a portion of the data corresponding to a plurality of points included in the first region (S2220).
  • "at least a portion” refers to only data that corresponds to all or specific conditions (e.g., an outlier filter, a specific label, etc.) specified by the user.
  • At least one processor may selectively perform a data removal algorithm based on the properties of the first region. For example, if the at least one processor determines that the data clusters within the first region are excessively dense and cause imbalance, the at least one processor may undersample all data points within the region or remove them at an arbitrary rate. Furthermore, for example, the at least one processor may selectively remove only data that satisfy predefined outlier identification criteria (e.g., distance, density, label mismatch, etc.) among points in the region by determining to remove only outliers within the first region.
  • predefined outlier identification criteria e.g., distance, density, label mismatch, etc.
  • the at least one processor may provide refinement options (e.g., filter conditions) to remove only points with a specific label or outside a specific statistical range within a user-specified region (e.g., "Remove only points that fall within this region and have a label of 0").
  • refinement options e.g., filter conditions
  • At least one processor can provide a data image with at least some data removed (S2230).
  • S2230 By examining the newly visualized distribution, the user can intuitively understand the effects of improved data imbalance or noise reduction. If necessary, deletion history can be maintained or an Undo function can be provided to prevent data loss due to user error.
  • At least one processor can perform reimaging and diagnostics (cluster analysis) on the data set after removal to check again how much the data quality has improved and whether there is an advantage in model learning.
  • Embodiments of the present disclosure enable direct removal through user interaction, thereby reflecting fine-grained domain knowledge that may be missed by existing automated algorithms.
  • immediate feedback can be obtained in a visualized space (2D or 3D), allowing for quick determination of follow-up actions, such as "how much data has been lost” and "how has the distribution changed”.
  • an intuitive and flexible tool for data quality management can be provided by allowing the user to interactively remove data in a specified area (particularly, a dense cluster area, an outlier-rich area, etc.).
  • FIG. 23 is a diagram illustrating a method for a computing device to perform data improvement in response to a data improvement request and provide visual interaction therefor, according to various embodiments.
  • a computing device or at least one processor included in the computing device may receive a data improvement request (S2310).
  • the at least one processor may recognize the data improvement request by receiving user input via at least one GUI that directs data improvement. For example, if a button for improving data in a first manner (e.g., "Resolve data imbalance") or a button for improving data in a second manner (e.g., "Remove noise area”) is selected on the user GUI, the at least one processor may recognize that an improvement request has occurred through the corresponding input. Additionally, the user may recognize that a specific class (or attribute) is lacking and request, for example, "Please create more of that class.”
  • At least one processor can visually represent at least one area requiring data improvement (S2320). Specifically, at least one processor can acquire characteristics of the data set based on a vector set corresponding to the data set, and detect areas requiring data improvement based on the characteristics of the data set.
  • At least one processor can analyze the density of the data set to identify areas where a specific class is underrepresented or where certain regions (clusters) are overly dense. Furthermore, for example, at least one processor can analyze the bias of the data set to automatically identify areas where the class distribution is imbalanced and requires improvement, such as areas where a specific class is underrepresented or overrepresented. Furthermore, for example, at least one processor can detect areas where a large number of outliers exist (areas where noisy data is concentrated) by performing outlier analysis.
  • At least one processor can highlight or outline the detected "areas requiring improvement" on the visualized data image (IOD) to notify the user.
  • at least one processor can provide interactive guidance (e.g., tooltips) on the GUI, explaining the criteria for selecting the areas.
  • At least one processor may perform data improvement by generating or removing data for each of at least one region (S2330). For example, at least one processor may generate synthetic data based on the data generation method according to FIG. 19 for a first region on a data image determined to be lacking in data of a specific class. In addition, for example, at least one processor may remove at least some data based on the data removal method according to FIG. 22 for a second region that is over-dense with data or contains outliers. At this time, the user may select at least one of various detailed options (e.g., execute, cancel, adjust detailed options, etc.) for the automatic suggestion.
  • various detailed options e.g., execute, cancel, adjust detailed options, etc.
  • At least one processor can provide a visualized improved data image corresponding to the results of the data enhancement (S2340). This allows users to visually see how the data distribution has changed, intuitively understand the extent to which imbalances have been resolved, and the extent to which outliers have been reduced.
  • at least one processor can also provide users with a comparison view (before/after visualization) with the previous state or a diagnostic report.
  • users can perform manual corrections by deselecting some of the visualized areas for improvement or by instructing the user to expand the scope of creation or removal.
  • at least one processor can automatically identify areas that require further improvement based on the improvement results, repeatedly performing a loop to improve data quality.
  • FIG. 24 is a diagram illustrating an example of a computing device in which a data processing method including a snapshot function is implemented, according to various embodiments.
  • a computing device (2400) may include multiple components for acquiring snapshot information based on a specific scene on a data image.
  • the multiple components are arbitrarily separated to perform specific operations, and may be physically separate components, or may be separate components based on various operations implemented on a single software program.
  • the computing device (2400) may include a screener for screening a data image or a vector set corresponding to the data image, an event detector for detecting whether an event for capturing a snapshot has occurred, a capture unit for capturing a snapshot, a correction unit for correcting a captured scene, and a generator for generating various information about the snapshot.
  • the screener can screen a vector set or a data image.
  • the screener can be configured with at least one metric for measuring characteristic values (e.g., density, bias, homogeneity, presence of outliers, etc.) of the data set based on the vector set or the data image.
  • the screener can analyze basic characteristics (e.g., density, outlier ratio, etc.) in a 2D/3D data image based on a first metric, or measure high-dimensional distribution characteristics in a high-dimensional vector set based on a second metric.
  • An event detector can detect events occurring during screening. Specifically, the event detector can monitor the measured results from the screener and detect an event occurrence when the monitored results satisfy predetermined conditions. At this time, the event detector can store information about the data in which the event occurred (e.g., data items, vector values corresponding to the data, coordinates of data points corresponding to the data, etc.). For example, if the event detector detects a specific area with a higher data density than a threshold value in the screener results, or a specific area with a bias index exceeding a reference value, the event detector can recognize this as an "event occurrence" and generate a signal.
  • information about the data in which the event occurred e.g., data items, vector values corresponding to the data, coordinates of data points corresponding to the data, etc.
  • the event detector can recognize this as an "event occurrence" and generate a signal.
  • the generator can generate information about events detected by the event detector. Specifically, the generator can generate tag information by synthesizing event information detected by the event detector, the location and metadata of the captured scene, and feature values calculated by the screener. Alternatively, the generator can record additional information related to the snapshot (class label, time, user ID, etc.). That is, the generator can generate tag information based on information about data in which an event occurred, which is pre-saved or generated from the event detector.
  • the tag information generated by the generator can include, for example, identification labels such as "areas with unusual characteristics (concentrated outliers)" or "sections with excessively high density,” the time of snapshot capture, key feature values, etc.
  • the capture unit can capture and save a specific scene on a data image. For example, when a signal notifying the occurrence of an event is transmitted from an event detector or when a user directly commands a snapshot, the capture unit can capture (photograph) a scene containing at least one data point corresponding to an event on the data image at a specific viewpoint.
  • “capture” includes a process of saving the state of an actual GUI screen (2D/3D view) or an internal data structure (vector set + visualization mapping parameters) as an image (or video frame).
  • the correction unit can perform modifications on captured scenes (snapshots). This provides users with highly visible results and, if necessary, can also perform graphical corrections, such as visually highlighting areas or adding labels.
  • a computing device can provide a vector set corresponding to a data set to a screener, and the screener can diagnose the characteristics of the data set based on the vector set.
  • An event detector can identify whether an event has occurred based on the characteristics diagnosed by the screener. For example, if the characteristics diagnosed by the screener satisfy a predetermined condition, the event detector can generate a signal indicating that an event has occurred.
  • a capture unit can capture a snapshot corresponding to an event based on the signal indicating that an event has occurred. Specifically, the capture unit can generate a snapshot by capturing a scene including at least one data point corresponding to at least one vector associated with the event at a specific point in time.
  • a generator can generate tag information based on information about an event that has occurred or information about a location on a data image where an event has occurred.
  • a correction unit can correct the generated snapshot in a predetermined manner (e.g., brightness adjustment, contrast adjustment, etc.).
  • FIG. 25 is a diagram illustrating a method for a computing device to screen a data set to generate snapshot information, according to various embodiments.
  • a computing device or at least one processor included in the computing device may screen a vector set by calculating at least one characteristic value based on at least one vector included in the vector set using a screener having at least one metric set (S2510).
  • at least one processor may diagnose the characteristics of data and measure various indicators such as data density, bias, and homogeneity.
  • screening may proceed along a specific path based on a specific starting point (coordinate location) on a data image (2D/3D) or vector set.
  • at least one processor may sequentially scan and analyze a certain range from a predetermined location or a predetermined reference point, thereby evaluating the density, bias, homogeneity, etc. of the data distribution.
  • At least one processor may perform multiple levels of screening, each level being differentiated based on the screening target. Specifically, at least one processor may perform at least one of a first-level screening for screening a data image or a second-level screening for screening a vector set. The first-level screening performs two-dimensional or three-dimensional data operations based on two-dimensional or three-dimensional data images, while the second-level screening performs high-dimensional operations based on high-dimensional vector sets.
  • At least one processor may apply a second metric to measure the characteristics of the data by performing a calculation (high-dimensional calculation) based on at least one vector included in the vector set. For example, at least one processor may diagnose whether the vectors being screened form a structure (clusters, outliers, etc.) on an actual high-dimensional manifold, how high the bias index is, etc.
  • At least one processor may perform multiple levels of screening, each level being differentiated according to the type of characteristic to be screened and diagnosed. Specifically, at least one processor may perform at least one of a first level of screening in which a metric is set for measuring a first type of characteristic (e.g., missing values, data statistics, etc.), a second level of screening in which a metric is set for measuring a second type of characteristic (e.g., density, etc.), or a third level of screening in which a metric is set for measuring a third type of characteristic (e.g., class distribution, etc.).
  • a first level of screening in which a metric is set for measuring a first type of characteristic (e.g., missing values, data statistics, etc.)
  • a second level of screening in which a metric is set for measuring a second type of characteristic (e.g., density, etc.)
  • a third level of screening in which a metric is set for measuring a third type of characteristic (e.g., class distribution,
  • At least one processor can identify at least one vector whose at least one characteristic value satisfies a predetermined condition (S2520). Through this, at least one processor can identify at least one vector whose first characteristic (e.g., density) satisfies a first condition (e.g., overcrowded area, etc.) or at least one vector whose second characteristic (e.g., bias index) satisfies a second condition (e.g., bias is above a standard).
  • first characteristic e.g., density
  • a first condition e.g., overcrowded area, etc.
  • second characteristic e.g., bias index
  • At least one processor can obtain a data image including a plurality of data points representing the data set in two dimensions or three dimensions by processing the vector set using at least one visualization tool (S2530).
  • At least one processor can identify at least one data point corresponding to at least one vector satisfying a predetermined condition, and determine a target area by determining a predetermined area based on the at least one data point. For example, at least one processor can determine an area having a predetermined radius centered on the at least one data point as the target area.
  • At least one processor may generate tag information associated with the determined target area.
  • the tag information may reflect diagnostic results associated with the target area.
  • the at least one processor may generate tag information based on information about events detected by the event detector and characteristic values diagnosed by the screener, but is not limited thereto.
  • the tag information may include, but is not limited to, the discovery context (e.g., which metric conditions were met), time, analyst ID, and event type (e.g., unusual section, blank section, overcrowded section, etc.).
  • At least one processor may compare multiple candidate scenes (e.g., top view, side view, or 45-degree angle view) and provide them to the user, and determine the scene selected by the user as the first scene.
  • the captured snapshot (first scene) may be corrected (brightness, contrast, highlight, etc.) by the correction unit (MODIFIER) described above, and the final scene corrected in this way and the snapshot information including tag information may be stored in a database or file format.
  • MODIFIER correction unit
  • the computing device (2400) can generate snapshot information based on user input received via a GUI-based input interface. Specifically, the computing device (2400) transmits a signal to the capture unit instructing the generation of a snapshot based on the user input, and the capture unit can capture at least a portion of the data image (IOD).
  • IOD data image
  • FIG. 26 is a diagram illustrating a method for a computing device to provide snapshot information based on user input according to various embodiments.
  • the computing device may provide a first data image corresponding to a first data set through a first view port (2710). At this time, the computing device may also provide preview information for at least some of the plurality of data points included in the first data image.
  • a computing device can provide preview information by outputting the actual data corresponding to a data point.
  • the computing device can configure the preview information so that clicking on a specific data point previews the actual data (e.g., original image, text content, statistical values, etc.) corresponding to that point. This allows the user to immediately confirm the meaning of the point in the data image.
  • At least one processor may generate first snapshot information including a first scene for a first data image being provided through a first view port in response to a user input received through the first GUI and provide the first snapshot information through a second view port (S2620). Specifically, the at least one processor may generate the first snapshot information by capturing a first scene including at least one data point on the first data image.
  • the computing device may receive user input for a first GUI (2720) and provide first snapshot information including a first scene for a first data image through a second view port (2730). Specifically, the computing device may display the first GUI (2720) and prompt the user to specify a desired scene using a button or a specific gesture (such as dragging or selecting a box) that instructs the user to take a snapshot.
  • a button or a specific gesture such as dragging or selecting a box
  • the computing device may transmit a corresponding instruction signal to the capture unit.
  • the capture unit may capture the first scene by synthesizing the current viewpoint, magnification ratio, or range designation of the first viewport (2710) where the first data image is displayed.
  • the computing device may provide the completed first snapshot information to the user through the second viewport (2730).
  • the second viewport (2730) may be used as a "snapshot preview" area.
  • the computing device may display a "tag selection menu” (not shown) in a portion of the second viewport (2730) (e.g., a pop-up, a side panel).
  • At least one processor may analyze the captured scene (e.g., data distribution or point properties within the snapshot) and automatically suggest recommended tags such as "bias,” “dense,” “outlier,” and “label mismatch.”
  • the processor may generate and store final tag information (e.g., "dense,” “bottom-right cluster,” “Class A”>) based on the selected tags.
  • At least one processor may generate link information for connecting the first snapshot information to an external communication network based on user input to the second GUI provided through the second view port (S2630).
  • At least one processor can provide a second GUI (2740) through a second viewport (2720). At least one processor can generate link information based on a user input for the second GUI (2740).
  • the link information can include a link for transmitting the first snapshot information through an external communication network.
  • the link information can be a connection path for transmitting and sharing this snapshot through an external communication network (Internet, company intranet, etc.), and the recipient can check the same scene (snapshot) and tag information or memo information, etc. through the link.
  • the user since the user directly determines and captures an arbitrary point/area, it effectively supports customized scenarios (such as enlarging only a specific cluster, emphasizing only a specific label, etc.) that may be missed in automatic capture.
  • user comments or classification information are organically combined and stored with snapshots through view port switching (first->second) and memo/tag interfaces, thereby providing richer context for future analysis collaboration or document reporting.
  • the link generation function enables real-time sharing and feedback of snapshot information (scenes, tags, notes, etc.) with external team members or other systems, thereby increasing the efficiency of data interpretation in a remote collaboration environment and inducing external exposure to the solution, which may have an economic ripple effect.
  • FIG. 28 is a diagram illustrating a function of a computing device to reproduce snapshot information according to various embodiments.
  • a computing device (or at least one processor) according to an embodiment of the present disclosure can sequentially store a plurality of snapshots captured along a screening path and reproduce these snapshots in the form of a video or slide by sequentially providing these snapshots according to a user reproduction instruction.
  • a computing device or at least one processor included in the computing device may sequentially store a plurality of snapshot information acquired along a screening path (S2810).
  • a screener included in the computing device may search a data image (or a high-dimensional vector set) along a specific path (referred to as a "screening path"), and may capture snapshots at each point when a specific characteristic is detected, an event occurs, or at regular intervals. For example, screening may be performed by "inspecting each block (area) on a 2D view while gradually moving from the upper left to the lower right.” After checking the characteristic value at each inspection point, a snapshot may be taken if it exceeds a threshold value.
  • At least one processor can record the captured snapshots in this manner along with connection information (e.g., "Snap #1 -> #2 -> Snap #3 ") according to the acquisition order or the screening path order.
  • connection information e.g., "Snap #1 -> #2 -> Snap #3 "
  • each snapshot can be stored including metadata such as "view point,” shooting location (coordinates along the screening path or high-dimensional mapping information), event cause (excess density, outlier detection, etc.), shooting time,” etc.
  • At least one processor can play back multiple sequentially stored snapshot information in a predetermined manner according to a playback instruction input by the user (S2820). Specifically, when a user (e.g., an analyst) presses "Play" or inputs a request to sequentially check screening records through a specific interface, at least one processor can retrieve multiple sequentially stored snapshot information.
  • a user e.g., an analyst
  • the computing device may be implemented so that the transmission order of multiple snapshot information during playback differs from the actual snapshot capture order. For example, a user may capture snapshots in a random order at the time of capture, but rearrange them according to specific criteria (e.g., issue priority, reverse chronological order, user-specified order, etc.) during playback.
  • specific criteria e.g., issue priority, reverse chronological order, user-specified order, etc.
  • At least one processor may output snapshots using a predetermined playback method, such as a slideshow method (continuous screen switching at fixed intervals), animation (smooth switching between scenes), or timeline operation (progressing according to the timing of an event in each snapshot).
  • a predetermined playback method such as a slideshow method (continuous screen switching at fixed intervals), animation (smooth switching between scenes), or timeline operation (progressing according to the timing of an event in each snapshot).
  • at least one processor may visually display at least one area where a snapshot was captured on the data image and may also display an animated path connecting the points on the screen.
  • the computing device may provide a dynamic navigation experience to the user by creating a visual effect as if a camera were moving sequentially through the areas where snapshots were captured.
  • the computing device can receive user commands related to playback and control playback based on the received commands.
  • at least one processor can perform actions such as "pause,” “skip,” “reverse playback,” and “add annotations to individual snapshots” based on user interaction actions.
  • At least one processor can encourage the user to utilize the improvement feature by providing improvement information (e.g., the reason for the improvement) when an area requiring improvement is identified during snapshot playback.
  • improvement information e.g., the reason for the improvement
  • a computing device may provide a function for linking snapshot information and a diagnostic report.
  • the computing device or at least one processor included in the computing device may, during a data screening process, generate a diagnostic report describing the characteristics of a data set along with snapshot information for a specific region.
  • FIG. 29 is a diagram illustrating a method for a computing device to generate diagnostic reports and snapshot information in conjunction with each other, according to various embodiments.
  • At least one processor can perform a diagnostic report generation operation (S2920). Specifically, at least one processor can generate a diagnostic report by collating or summarizing various diagnostic results. Since the method for generating a diagnostic report has been described above, a detailed description thereof will be omitted.
  • At least one processor may perform a snapshot information generation operation (S2930). Specifically, if an unusual area (such as an overcrowded area or a cluster of outliers) is discovered during the screening process, at least one processor may capture the scene and generate a snapshot.
  • a snapshot information generation operation S2930. Specifically, if an unusual area (such as an overcrowded area or a cluster of outliers) is discovered during the screening process, at least one processor may capture the scene and generate a snapshot.
  • At least one processor may record snapshot information in the form of an abbreviated or summarized version of the diagnostic results (e.g., "20 noise points found, clusters with a bias index of 0.85 or higher").
  • at least one processor may generate snapshot information based on summary information about a specific diagnostic result and a scene captured from a data image corresponding to the diagnostic result.
  • At least one processor can convert portions of the diagnostic results that meet certain criteria (e.g., above a threshold, a class of interest, etc.) into snapshot information.
  • At least one processor can perform a linking operation (S2940) of a diagnostic report and snapshot information. Specifically, at least one processor can link and store the diagnostic result in the diagnostic report and the snapshot information corresponding to the diagnostic result.
  • a linking operation S2940
  • the two pieces of data can be linked by linking unique IDs (e.g., report ID, snapshot ID) or specifying identical metadata (e.g., time, area coordinates, event type).
  • At least one processor can perform a diagnostic report call operation (S2950) from a snapshot. Specifically, when a user clicks a button instructing to call a diagnostic report in a GUI displaying snapshot information or clicks a label (e.g., class imbalance) displayed on a snapshot, the computing device can quickly call up the diagnostic results corresponding to the snapshot based on the stored information. Thereafter, at least one processor can provide the diagnostic report through a new window (viewport) or pop-up window.
  • a diagnostic report call operation S2950
  • the report can be loaded in a new window (or a secondary viewport) and compared side-by-side with existing snapshots.
  • the computing device stores snapshot information generated during the data screening process and a diagnostic report detailing the characteristics of the data set in conjunction with each other, and then cross-references them based on user input, thereby supporting the improvement and utilization of data quality by flexibly moving between intuitive visual information and precise analysis results.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Software Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Data Mining & Analysis (AREA)
  • Computing Systems (AREA)
  • Evolutionary Computation (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Biomedical Technology (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Biophysics (AREA)
  • Databases & Information Systems (AREA)
  • Human Computer Interaction (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

Selon un mode de réalisation, la présente divulgation concerne un appareil informatique, l'appareil informatique comprenant une mémoire et au moins un processeur connecté électroniquement à la mémoire, l'au moins un processeur mis en œuvre pour exécuter au moins une instruction stockée dans la mémoire étant configuré pour réaliser les étapes suivantes : fournir, par l'intermédiaire d'une première fenêtre d'affichage, une première image de données correspondant à un premier ensemble de données, la première image de données comprenant une pluralité de points de données correspondant respectivement à des éléments de données inclus dans le premier ensemble de données ; générer, en réponse à une entrée d'utilisateur reçue provenant d'une première GUI, des premières informations d'instantané comprenant une première scène de la première image de données fournie par l'intermédiaire de la première fenêtre d'affichage, et fournir celles-ci par l'intermédiaire d'une seconde fenêtre d'affichage ; et générer, sur la base d'une entrée d'utilisateur liée à une seconde GUI fournie par l'intermédiaire de la seconde fenêtre d'affichage, des informations de lien pour connecter les premières informations d'instantané à un réseau de communication externe.
PCT/KR2025/012348 2024-08-16 2025-08-14 Procédé, appareil et système de mise en œuvre d'un outil de visualisation de données Pending WO2026038902A1 (fr)

Applications Claiming Priority (10)

Application Number Priority Date Filing Date Title
KR20240109511 2024-08-16
KR10-2024-0109511 2024-08-16
KR1020250019947A KR102925436B1 (ko) 2024-08-16 2025-02-17 데이터 품질 개선을 위한 사용자 인터랙션의 구현 방법 및 그러한 방법이 구현된 컴퓨팅 장치
KR10-2025-0019945 2025-02-17
KR10-2025-0019947 2025-02-17
KR10-2025-0019944 2025-02-17
KR1020250019945A KR102925421B1 (ko) 2024-08-16 2025-02-17 데이터 진단에 따른 스냅샷 획득 방법 및 그러한 방법이 구현된 컴퓨팅 장치
KR10-2025-0019946 2025-02-17
KR1020250019944A KR20260025723A (ko) 2024-08-16 2025-02-17 데이터 렌즈 수준에 따른 데이터 분석 방법 및 그러한 방법이 구현된 컴퓨팅 장치
KR1020250019946A KR102925425B1 (ko) 2024-08-16 2025-02-17 데이터 스냅샷 획득을 위한 사용자 인터랙션의 구현 방법 및 그러한 방법이 구현되는 컴퓨팅 장치

Publications (1)

Publication Number Publication Date
WO2026038902A1 true WO2026038902A1 (fr) 2026-02-19

Family

ID=98780833

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2025/012348 Pending WO2026038902A1 (fr) 2024-08-16 2025-08-14 Procédé, appareil et système de mise en œuvre d'un outil de visualisation de données

Country Status (1)

Country Link
WO (1) WO2026038902A1 (fr)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20180051367A (ko) * 2016-11-08 2018-05-16 삼성전자주식회사 디바이스가 이미지를 보정하는 방법 및 그 디바이스
KR102029055B1 (ko) * 2013-02-08 2019-10-07 삼성전자주식회사 고차원 데이터의 시각화 방법 및 장치
KR102556766B1 (ko) * 2022-10-12 2023-07-18 주식회사 브이알크루 데이터셋을 생성하기 위한 방법
KR20230121164A (ko) * 2020-05-24 2023-08-17 킥소틱 랩스 인크. 신속한 스크리닝을 위한 도메인-특정 언어 해석기 및대화형 시각적 인터페이스
KR20240002385A (ko) * 2022-06-29 2024-01-05 주식회사 페블러스 데이터 클리닉 방법, 데이터 클리닉 방법이 저장된 컴퓨터 프로그램 및 데이터 클리닉 방법을 수행하는 컴퓨팅 장치

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR102029055B1 (ko) * 2013-02-08 2019-10-07 삼성전자주식회사 고차원 데이터의 시각화 방법 및 장치
KR20180051367A (ko) * 2016-11-08 2018-05-16 삼성전자주식회사 디바이스가 이미지를 보정하는 방법 및 그 디바이스
KR20230121164A (ko) * 2020-05-24 2023-08-17 킥소틱 랩스 인크. 신속한 스크리닝을 위한 도메인-특정 언어 해석기 및대화형 시각적 인터페이스
KR20240002385A (ko) * 2022-06-29 2024-01-05 주식회사 페블러스 데이터 클리닉 방법, 데이터 클리닉 방법이 저장된 컴퓨터 프로그램 및 데이터 클리닉 방법을 수행하는 컴퓨팅 장치
KR102556766B1 (ko) * 2022-10-12 2023-07-18 주식회사 브이알크루 데이터셋을 생성하기 위한 방법

Similar Documents

Publication Publication Date Title
WO2022154471A1 (fr) Procédé de traitement d'image, appareil de traitement d'image, dispositif électronique et support de stockage lisible par ordinateur
WO2020214006A1 (fr) Appareil et procédé de traitement d'informations d'invite
WO2020159232A1 (fr) Procédé, appareil, dispositif électronique et support d'informations lisible par ordinateur permettant de rechercher une image
WO2018093182A1 (fr) Procédé de gestion d'images et appareil associé
WO2020138928A1 (fr) Procédé de traitement d'informations, appareil, dispositif électrique et support d'informations lisible par ordinateur
WO2018117685A1 (fr) Système et procédé de fourniture d'une liste à faire d'un utilisateur
WO2018088794A2 (fr) Procédé de correction d'image au moyen d'un dispositif et dispositif associé
WO2015133699A1 (fr) Appareil de reconnaissance d'objet, et support d'enregistrement sur lequel un procédé un et programme informatique pour celui-ci sont enregistrés
WO2015178716A1 (fr) Procédé et dispositif de recherche
EP1898339A1 (fr) Système et procédé de récuperation
EP3552163A1 (fr) Système et procédé de fourniture d'une liste à faire d'un utilisateur
WO2023080276A1 (fr) Système d'apprentissage profond distribué à liaison de base de données basé sur des interrogations, et procédé associé
WO2024122990A1 (fr) Procédé et dispositif permettant d'effectuer une localisation visuelle
WO2026084477A1 (fr) Procédé, dispositif et programme de classification et de détection de cellules dans une image pathologique sur la base d'ia
WO2021162481A1 (fr) Dispositif électronique et son procédé de commande
WO2026038902A1 (fr) Procédé, appareil et système de mise en œuvre d'un outil de visualisation de données
WO2022080848A1 (fr) Interface utilisateur pour analyse d'image
WO2021230469A1 (fr) Procédé de recommandation d'articles
WO2023080275A1 (fr) Serveur de base de données d'applications de cadre d'apprentissage profond pour classifier le sexe et l'âge, et procédé associé
WO2021246642A1 (fr) Procédé de recommandation de police de caractères et dispositif destiné à le mettre en œuvre
WO2021107360A2 (fr) Dispositif électronique de détermination d'un degré de similarité et son procédé de commande
WO2023080590A1 (fr) Procédé de fourniture de vision par ordinateur
WO2024085552A1 (fr) Dispositif électronique pour fournir un environnement virtuel pour générer des données synthétiques, procédé de fonctionnement d'un dispositif électronique et système comprenant un dispositif électronique
WO2024005464A1 (fr) Procédé de clinique de données, programme informatique dans lequel un procédé de clinique de données est stocké et dispositif informatique qui effectue un procédé de clinique de données
WO2026010411A1 (fr) Dispositif, procédé de fonctionnement de dispositif et support d'enregistrement non transitoire

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25855002

Country of ref document: EP

Kind code of ref document: A1