WO2026038902A1 - 데이터 시각화 도구의 구현 방법, 장치 및 시스템 - Google Patents

데이터 시각화 도구의 구현 방법, 장치 및 시스템

Info

Publication number
WO2026038902A1
WO2026038902A1 PCT/KR2025/012348 KR2025012348W WO2026038902A1 WO 2026038902 A1 WO2026038902 A1 WO 2026038902A1 KR 2025012348 W KR2025012348 W KR 2025012348W WO 2026038902 A1 WO2026038902 A1 WO 2026038902A1
Authority
WO
WIPO (PCT)
Prior art keywords
data
computing device
processor
data set
information
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/KR2025/012348
Other languages
English (en)
French (fr)
Inventor
이주행
이정원
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Pebblous Inc
Original Assignee
Pebblous Inc
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Priority claimed from KR1020250019947A external-priority patent/KR102925436B1/ko
Application filed by Pebblous Inc filed Critical Pebblous Inc
Publication of WO2026038902A1 publication Critical patent/WO2026038902A1/ko
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/904Browsing; Visualisation therefor
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/21Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
    • G06F18/213Feature extraction, e.g. by transforming the feature space; Summarisation; Mappings, e.g. subspace methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/44Arrangements for executing specific programs
    • G06F9/451Execution arrangements for user interfaces
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/042Knowledge-based neural networks; Logical representations of neural networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/10Interfaces, programming languages or software development kits, e.g. for simulating neural networks

Definitions

  • the present disclosure relates to data processing technology for diagnosing and visualizing data. More specifically, it relates to technology for diagnosing data through data imaging and providing interactive functions by visualizing the data.
  • One task of the present disclosure is to diagnose the quality of a large data set.
  • one task of the present disclosure is to visualize a data set so that it can be effectively expressed while maintaining the inherent structure of the data through a vectorization method for precisely identifying the inherent distribution of a large data set.
  • a computing device comprising a memory and at least one processor electronically connected to the memory, wherein the at least one processor is configured to execute at least one instruction stored in the memory, and is configured to perform the following operations: acquiring a first data set; selecting a first level from among a plurality of levels classified according to a data processing method based on a user input; determining at least one property of a first data processing model corresponding to the first level, the first data processing model including at least one artificial intelligence model; providing, through a first GUI (Graphical User Interface), prior information related to a task of processing the first data set using the first data processing model; and constructing the first data processing model based on a user input to the first GUI.
  • GUI Graphic User Interface
  • a data processing method may be provided that is set to perform, by at least one processor executing at least one instruction stored in a memory, an operation of acquiring a first data set, an operation of selecting a first level from among a plurality of levels classified according to a data processing method based on a user input, an operation of determining at least one property of a first data processing model corresponding to the first level, the first data processing model including at least one artificial intelligence model, an operation of providing, through a first GUI (Graphical User Interface), prior information related to a task of processing the first data set using the first data processing model, and an operation of constructing the first data processing model based on a user input to the first GUI.
  • GUI Graphic User Interface
  • an electronic device including a display configured to display at least one GUI (Graphic User Interface), a memory, and at least one processor electronically connected to the memory, wherein the at least one processor configured to execute at least one instruction stored in the memory is configured to perform an operation of receiving a user input for selecting a first level among a plurality of levels classified according to a data processing method, an operation of displaying a first GUI (Graphic User Interface) through the display that indicates prior information related to a task of processing the first data set using a first data processing model corresponding to the first level, the first data processing model including at least one artificial intelligence model, an operation of obtaining a diagnosis result for the first data set based on a first vector set defined in a first-dimensional embedding area from the first data processing model, and an operation of providing the first vector set to a first visualization tool to display a two-dimensional or three-dimensional first data image, the first data image corresponding to the first data set, through the display.
  • GUI Graphic User Interface
  • an operation of obtaining a data set by at least one processor executing at least one instruction stored in a memory, an operation of obtaining a data set, an operation of embedding the data set in a specific dimension to obtain a vector set, the vector set including a plurality of vectors corresponding to a plurality of data included in the data set, an operation of screening the vector set by calculating at least one feature value based on the vectors included in the vector set using a screener having at least one metric set, an operation of identifying at least one vector whose at least one feature value satisfies a predetermined condition, an operation of obtaining a data image including a plurality of data points representing the data set in two dimensions or three dimensions by processing the vector set using at least one visualization tool, an operation of determining a first area including at least one data point corresponding to the at least one vector on the data image and generating tag information associated with the first area, and a first scene including the first area, the first scene being generated by capturing the first area from a specific view
  • the computer program includes an operation of acquiring a data set, an operation of embedding the data set in a specific dimension to acquire a vector set, the vector set including a plurality of vectors corresponding to a plurality of data included in the data set, an operation of screening the vector set by calculating at least one feature value based on the vectors included in the vector set using a screener having at least one metric set, an operation of identifying at least one vector whose at least one feature value satisfies a predetermined condition, an operation of processing the vector set using at least one visualization tool to acquire a data image including a plurality of data points representing the data set in two dimensions or three dimensions, an operation of determining a first region including at least one data point corresponding to the at least one vector on the data image and generating tag information associated with the first region, and a first scene including the first region, the first scene being generated by capturing the first region from
  • a computing device comprising a memory and at least one processor electronically connected to the memory, wherein the at least one processor is configured to execute at least one instruction stored in the memory, and is configured to perform an operation of providing a first data image corresponding to a first data set through a first view port, the first data image including a plurality of data points corresponding to each of data included in the first data set; an operation of generating first snapshot information including a first scene for the first data image being provided through the first view port in response to a user input received through a first GUI and providing the first snapshot information through a second view port; and an operation of generating link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port.
  • a data interaction method including an operation of providing a first data image corresponding to a first data set, the first data image including a plurality of data points corresponding to each of data included in the first data set, through a first view port by at least one processor executing at least one instruction in a memory, an operation of generating first snapshot information including a first scene for the first data image being provided through the first view port in response to a user input received through a first GUI and providing the first snapshot information through a second view port, and an operation of generating link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port.
  • an electronic device comprising a display configured to display at least one GUI (Graphic User Interface), a memory, and at least one processor electronically connected to the memory, wherein at least one processor configured to execute at least one instruction stored in the memory is configured to perform an operation of providing a first data image corresponding to a first data set, the first data image including a plurality of data points corresponding to each of data included in the first data set, through a first view port of the display, an operation of providing first snapshot information including a first scene for the first data image being provided through the first view port through the display in response to a user input received through the first GUI, and an operation of providing link information for connecting the first snapshot information with an external communication network based on a user input for a second GUI provided through the second view port.
  • GUI Graphic User Interface
  • a method including an operation of visualizing a data set by at least one processor executing at least one instruction stored in a memory to provide a first data image corresponding to the data set, an operation of receiving a user input for a first area on the first data image, an operation of determining a latent code corresponding to the first area and providing the determined latent code to a generation model to generate synthetic data, an operation of inputting the synthetic data to a data processing model, the data processing model being trained to embed data into a specific dimension, to obtain a synthetic vector corresponding to the synthetic data, and an operation of visualizing the synthetic vector to provide a synthetic point corresponding to the synthetic data on the first data image.
  • a computing device including a display configured to display at least one GUI (Graphic User Interface), a memory, and at least one processor electronically connected to the memory, wherein the at least one processor configured to execute at least one instruction stored in the memory is configured to perform an operation of visualizing a data set and providing a first data image corresponding to the data set through the display, an operation of receiving a user input for a first area on the first data image, an operation of determining a latent code corresponding to the first area and providing the determined latent code to a generation model to generate synthetic data, an operation of inputting the synthetic data into a data processing model, the data processing model being trained to embed data into a specific dimension, to obtain a synthetic vector corresponding to the synthetic data, and an operation of visualizing the synthetic vector and providing a synthetic point corresponding to the synthetic data on the first data image through the display.
  • GUI Graphic User Interface
  • a computing device can automatically diagnose the quality of a large data set and intuitively visualize it. This allows data analysts, AI model developers, and researchers to more easily understand the inherent structure of the data and develop strategies to improve data quality.
  • a user can more efficiently explore and manipulate data during the data visualization process, and utilize it to improve the performance of a machine learning model.
  • FIG. 1 is a diagram illustrating a configuration of a computing device according to various embodiments.
  • FIG. 2 is a diagram illustrating various data processing methods included in a data clinic service provided by a computing device according to various embodiments.
  • FIG. 3 is a diagram illustrating various systems for providing data clinic services according to various embodiments, and artificial intelligence models and algorithms for constructing the systems.
  • FIG. 4 is a diagram illustrating a method for a computing device to provide a data image according to various embodiments.
  • FIG. 5 is a diagram illustrating a method for a computing device to obtain characteristics of a data set according to various embodiments.
  • FIG. 6 is a diagram illustrating a data lens processing system and a data imaging system according to various embodiments.
  • FIG. 7 is a diagram illustrating a system in which a computing device visualizes data and provides user interaction functions, according to various embodiments.
  • FIG. 8 is a diagram illustrating a data visualization and interaction method according to various embodiments.
  • FIG. 9 is a drawing for explaining detailed steps performed in an imaging application step according to various embodiments.
  • FIG. 10 is a diagram illustrating the results of data imaging and diagnosis according to the level of the lens and the visualization tool according to various embodiments.
  • FIG. 11 is a diagram illustrating a lens builder included in a computing device according to various embodiments.
  • FIG. 12 is a diagram illustrating a method for a computing device to build a data processing model for data imaging and diagnosis based on user input, according to various embodiments.
  • FIG. 14 is a diagram illustrating an example of a computing device constructing a first data processing model according to various embodiments.
  • FIG. 17 is a diagram illustrating a method of imaging and visualizing a data set using a first data processing model built on a computing device according to various embodiments.
  • FIG. 20 is a diagram illustrating a method for a computing device to generate synthetic data based on user input according to various embodiments.
  • FIG. 22 is a diagram illustrating a method for a computing device to remove data based on user input, according to various embodiments.
  • FIG. 23 is a diagram illustrating a method for a computing device to perform data improvement in response to a data improvement request and provide visual interaction therefor, according to various embodiments.
  • FIG. 24 is a diagram illustrating an example of a computing device in which a data processing method including a snapshot function is implemented, according to various embodiments.
  • FIG. 25 is a diagram illustrating a method for a computing device to screen a data set to generate snapshot information, according to various embodiments.
  • FIG. 26 is a diagram illustrating a method for a computing device to provide snapshot information based on user input according to various embodiments.
  • Figure 27 is an example of a screen provided by a computing device.
  • FIG. 28 is a diagram illustrating a function of a computing device to reproduce snapshot information according to various embodiments.
  • FIG. 29 is a diagram illustrating a method for a computing device to generate diagnostic reports and snapshot information in conjunction with each other, according to various embodiments.
  • each of the phrases “A or B”, “at least one of A and B”, “at least one of A or B”, “A, B, or C”, “at least one of A, B, and C”, and “at least one of A, B, or C” may include any one of the items listed together in that phrase, or all possible combinations thereof.
  • a device-readable storage medium may be provided in the form of a non-transitory storage medium.
  • non-transitory simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
  • the term 'unit' as used in this disclosure means a software or hardware component such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC).
  • the 'unit' performs specific roles, but is not limited to software or hardware.
  • the 'unit' may be configured to reside on an addressable storage medium and may be configured to play one or more processors. Accordingly, according to some embodiments, the 'unit' includes components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables.
  • the functionality provided within the components and 'units' may be combined into a smaller number of components and 'units' or further separated into additional components and 'units'. Additionally, the components and ' ⁇ parts' may be implemented to activate one or more CPUs within a device or secure multimedia card. Furthermore, according to various embodiments of the present disclosure, the ' ⁇ parts' may include one or more processors.
  • FIG. 1 is a diagram illustrating a configuration of a computing device according to various embodiments.
  • a computing device e.g., an electronic device including a computing means such as a server or client device, hereinafter referred to as a “computing device”
  • a computing device may include a processor (110), a memory (120), a storage device (130), a communication circuit (140), and a bus (not shown).
  • the configuration of the computing device (100) is not limited to the configuration illustrated in FIG. 1 or the configuration described above, and may further include hardware or software configurations included in general computing devices or mobile devices.
  • the processor (110) may include at least one processor, at least some of which are implemented to provide different functions.
  • the processor (110) may execute software (e.g., a program) to control at least one other component (e.g., a hardware or software component) of the computing device (100) connected to the processor (110) and perform various data processing or calculations.
  • the processor (110) may store instructions or data received from other components in the memory (120) (e.g., a volatile memory), process the instructions or data stored in the volatile memory, and store the resulting data in the non-volatile memory.
  • the processor (110) may include a main processor (e.g., a central processing unit or an application processor) or an auxiliary processor (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith.
  • a main processor e.g., a central processing unit or an application processor
  • an auxiliary processor e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor
  • the coprocessor may be configured to use less power than the main processor or to be specialized for a given function.
  • the coprocessor may be implemented separately from the main processor or as a part thereof.
  • the coprocessor may control at least a portion of functions or states associated with at least one component (e.g., the display (240) or a communication circuit) of the computing device (100), for example, on behalf of the main processor while the main processor is in an inactive (e.g., sleep) state, or together with the main processor while the main processor is in an active (e.g., application execution) state.
  • the coprocessor e.g., an image signal processor or a communication processor
  • the coprocessor may be implemented as a part of another functionally related component (e.g., a communication circuit).
  • the coprocessor e.g., a neural network processing device
  • the operation of the computing device (100) described below may be understood as the operation of the processor (110).
  • the memory (120) may include at least one memory, at least some of which are implemented to provide different functions.
  • the memory (120) may store various data used by at least one component (e.g., the processor (110)) of the computing device (100).
  • the data may include, for example, software (e.g., a program) and input data or output data for instructions related thereto.
  • the memory (120) may include volatile memory or non-volatile memory.
  • the memory (120) may be implemented to store an operating system, middleware or applications, and/or the artificial intelligence model described above.
  • the memory (120) may include a plurality of instructions (121) that direct the operations of the processor (110) to implement the functions provided by the service.
  • the processor (110) may execute at least some of the plurality of instructions stored in the memory (120).
  • the computing device (110) may include a software server including the processor (110) that executes the functions provided by the service based on at least some of the plurality of instructions.
  • the storage device (130) can provide a mass storage device to the computing device (100).
  • the storage device (130) can be a computer-readable medium.
  • the storage device (130) can be a floppy disk device, a hard disk device, an optical disk device, a tape device, a flash memory or other similar solid-state memory device, or an array of devices including a storage area network or other configuration device.
  • a computer program product is explicitly embodied in an information medium.
  • the computer program product includes instructions that, when executed, perform one or more methods as described above.
  • the information medium is a computer-readable medium or a machine-readable medium, such as the memory (120), the storage device (130), or the memory of the processor (110).
  • the storage device (130) may include a database (DB).
  • the storage device (130) may include a database having a pre-structured data structure.
  • the computing device (110) may store data sets having interrelated relationships in the database.
  • a computing device can provide services based on various artificial intelligence frameworks performed by at least one processor and a memory electronically connected to at least one processor.
  • the memory (120) or storage device (130) may store at least one artificial intelligence model implementing various types of artificial intelligence (or machine learning) frameworks that can be trained to perform a given task.
  • artificial intelligence or machine learning
  • support vector machines, decision trees, neural networks, etc. are just a few examples of machine learning frameworks used in various applications such as image processing and natural language processing.
  • Some artificial intelligence frameworks, such as neural networks may utilize layers of nodes that perform specific operations.
  • a neural network may include an input layer, an output layer, and one or more intermediate layers. Each node may process its inputs according to a predefined function and provide output to subsequent layers, or in some cases, previous layers. The input to a particular node may be multiplied by a weight value corresponding to the edge between the input and the node. Additionally, each node may have a separate bias value used to generate the output. Various learning procedures can be applied to learn the edge weights and/or bias values (parameters).
  • a neural network architecture may have multiple layers that perform different specific functions. For example, one or more node layers may collectively perform specific operations, such as pooling, encoding, or convolution operations.
  • layer may refer to a group of nodes that share inputs and outputs, such as communicating with external sources or other layers of the network.
  • calculation may refer to a function that can be performed by one or more node layers.
  • model structure may refer to the overall architecture of a layered model, including the number of layers, the connectivity of the layers, and the types of operations performed by individual layers.
  • neural network structure may refer to the model structure of a neural network.
  • trained model and/or “tuned model” may refer to the model structure along with the parameters for the trained or tuned model structure.
  • two trained models may have different values for parameters while sharing the same model structure, such as when they are trained on different training data or when the training process has an underlying probabilistic process.
  • Transfer learning is a broad approach for training models with limited task-specific training data for a specific task.
  • transfer learning a model is first pretrained on another task for which valuable training data is available, and then adapted to a specific task using task-specific training data.
  • pre-training refers to training a model on a pre-training dataset to adjust model parameters in a manner that allows subsequent adjustments of those model parameters to tailor the model to one or more specific tasks.
  • pre-training may involve a self-supervised learning process on unlabeled training data, where the "self-supervised” learning process involves learning from the structure of pre-training examples in the absence of explicit (e.g., manually provided) labels.
  • tuning subsequent modification of the model parameters obtained through pre-training is referred to herein as "tuning.” Tuning may be performed for one or more tasks using supervised learning on explicitly labeled training data, and in some cases, a task different from pre-training may be used for tuning.
  • a communication bus may be a configuration for electronically (or communicatively) connecting multiple components included in a computing device. That is, each component may be interconnected using various buses and mounted on a common motherboard or in another suitable manner.
  • the input/output interface may include an input interface connected to an input device to receive an input signal, or an output interface connected to an output device to output an output signal.
  • the computing device (100) may further include at least one communication circuit (140) for communicating with an external device.
  • the communication circuit (140) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the computing device (100) and an external computing device, and the performance of communication through the established communication channel.
  • the communication circuit may operate independently from the processor (110) (e.g., a program processor) and may include one or more communication processors (e.g., communication chips) that support direct (e.g., wired) communication or wireless communication.
  • the communication circuit (140) may include a wireless communication module (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (e.g., a local area network (LAN) communication module, or a power line communication module).
  • a wireless communication module e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module
  • GNSS global navigation satellite system
  • wired communication module e.g., a local area network (LAN) communication module, or a power line communication module.
  • the corresponding communication module can communicate with an external computing device via a first network (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)).
  • a first network e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)
  • a second network e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)
  • a second network e.g., a long-
  • the wireless communication module can identify or authenticate the computing device (100) within a communication network such as the first network or the second network by using subscriber information stored in the subscriber identification module (e.g., an international mobile subscriber identity (IMSI)).
  • the wireless communication module can support a 5G network subsequent to a 4G network and next-generation communication technologies, such as new radio access technology (NR).
  • NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high-reliability and low-latency communications (URLLC (ultra-reliable and low-latency communications)).
  • the wireless communication module can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate.
  • the wireless communication module can support various technologies to secure performance in the high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna.
  • MIMO massive multiple-input and multiple-output
  • FD-MIMO full dimensional MIMO
  • array antenna analog beam-forming
  • large scale antenna or large scale antenna.
  • the wireless communication module can support various requirements specified in a computing device (100), an endoscope device, or a network system.
  • the wireless communication module can support a peak data rate (e.g., 20 Gbps or more) for eMBB implementation, a loss coverage (e.g., 164 dB or less) for mMTC implementation, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL) each, or 1 ms or less for round trip) for URLLC implementation.
  • a peak data rate e.g., 20 Gbps or more
  • a loss coverage e.g., 164 dB or less
  • U-plane latency e.g., 0.5 ms or less for downlink (DL) and uplink (UL) each, or 1 ms or less for round trip
  • the computing device (100) may be implemented to include at least some of the above-described components (processor, communication circuitry, memory, display).
  • a user device may be implemented to include a processor, communication circuitry, memory, sensors, and a display.
  • a server device may be implemented to include a processor, communication circuitry, and memory.
  • FIG. 2 is a diagram illustrating various data processing methods included in a data clinic service provided by a computing device according to various embodiments.
  • the data clinic service may include various data processing methods. These various data processing methods may be encoded and stored in the memory of the computing device, and at least one processor included in the computing device may be configured to execute at least one encoded instruction. Specifically, at least one processor may process a received input data set based on the various data processing methods and output an output data set.
  • a computing device may perform, but is not limited to, an operating method for data imaging, an operating method for data enhancement, an operating method for data generation, an operating method for data feature extraction, or an operating method for data evaluation.
  • each of the above-described operating methods can be performed based on operating algorithms of at least one processor included in the computing device.
  • a computing device may perform, but is not limited to, a data imaging algorithm, a data enhancement algorithm, a data generation algorithm, a data feature extraction algorithm, or a data evaluation algorithm.
  • each operation method and algorithm are arbitrarily named according to the output results for the convenience of explanation, so each operation method or algorithm is only defined based on the operations performed by the processor, and the name of the operation method or algorithm itself does not limit the invention.
  • a computing device can process an input data set according to a data imaging algorithm to generate an image for the input data set.
  • a computing device can improve data by processing an input data set according to a data improvement algorithm, and can generate a result of the improvement.
  • a computing device can generate synthetic data by processing an input data set according to a data generation algorithm.
  • a computing device can process an input data set according to a data feature extraction algorithm to extract a property of the input data set.
  • a computing device can process an input data set according to a data evaluation algorithm to evaluate the quality of the input data set.
  • the computing device can perform the various operation methods or algorithms described above in parallel, sequentially, or selectively. Specifically, the computing device can use the same input data as input values for different algorithms in parallel, can use the result values output by a specific algorithm sequentially as input values for another algorithm, or can selectively perform some of the algorithms among a plurality of algorithms according to a predetermined method.
  • the various operation methods or algorithms for the data clinic described above can be performed in a deep learning model included in a computing device according to various embodiments of the present disclosure.
  • the computing device according to various embodiments of the present disclosure may include one deep learning model for performing the various operation methods or algorithms described above, but is not limited thereto, and may include multiple deep learning models for performing each of the operation methods or algorithms described above, or may include one or more deep learning models for performing at least some of the various operation methods or algorithms described above.
  • Figure 3 is a diagram illustrating various systems for providing data clinic services according to various embodiments, as well as artificial intelligence models and algorithms for constructing such systems.
  • a system may refer to a system that includes at least one software or hardware configuration to perform a specific function.
  • a computing device may include a data clinic system composed of various artificial intelligence (or neural network, machine learning, etc.) models to provide clinic services.
  • various artificial intelligence or neural network, machine learning, etc.
  • the computing device may include, but is not limited to, a data imaging system, a data diagnostic system, and a data treatment system.
  • the data imaging system may include, but is not limited to, a lens processing model for determining an optimal dimension for representing the characteristics of the data, an imaging model for obtaining a data image reflecting the inherent characteristics of the data, or a visualization model for visually representing the data.
  • the data diagnosis system may include, but is not limited to, a diagnosis model for diagnosing at least one characteristic of data or a quality assessment model for evaluating the quality of data.
  • the data treatment system may include, but is not limited to, a synthetic model (or generation model) for generating targeted virtual data (or synthetic data) as needed, a data diet model for removing at least a portion of the data, or a data correction model for adjusting the characteristics of at least a portion of the data.
  • a synthetic model or generation model
  • a data diet model for removing at least a portion of the data
  • a data correction model for adjusting the characteristics of at least a portion of the data.
  • a module may include multiple hardware components for implementing an artificial intelligence model that performs a specific function.
  • a module may include, but is not limited to, an encoder, a decoder, a generator, a discriminator, an adapter, a natural language processing module, or a large language model (LLM).
  • the computing device can store the plurality of modules described above, and can construct an AI framework based on at least some of the modules to obtain an AI model for a data clinic.
  • a data lens included in a data imaging system can be implemented as an AI model including at least one encoder or at least one adapter, but is not limited thereto.
  • FIG. 4 is a diagram illustrating a method for a computing device to provide a data image according to various embodiments.
  • a computing device may receive a data set and provide an Image of Data (IOD).
  • IOD Image of Data
  • the data set may be data of dimension M (M>0).
  • the data set may be a data set defined on an M-dimensional input space (310).
  • the data set may be a single-modality data set.
  • the data set may be an image data set.
  • the data set may be a text data set.
  • the data set may be a collection of data with different modalities.
  • the data set may be an image data set including annotation information.
  • the data set may be a mixed data set of images and text.
  • a computing device can receive and process data of all modalities that can be used for deep learning, such as time series data sets, sensor data sets, as well as the image data and text data described above, as input data sets.
  • the data image (IOD) provided by the computing device may be an image that processes an input data set and displays it in an imaging space (320).
  • the image does not mean a 2D image, but is a general expression that visually represents data.
  • the imaging space (320) is a concept that includes a 2D space, a 3D space, and an N-dimensional virtual space, and means a space in which a data image provided according to an embodiment appears.
  • the computing device may output an output that displays the data image in a 2D or 3D imaging space, but is not limited thereto.
  • the computing device can provide a data image through the output device.
  • the computing device can provide the data image by outputting the data image through a display.
  • the imaging space (320) may be a screen of a display.
  • the computing device can provide the data image by outputting the data image through a printing device.
  • the imaging space (320) may be paper output by the printing device.
  • the computing device may provide a data image via the external device.
  • the imaging space (320) may be a display screen of the external device.
  • the server device may provide the data image by transmitting the data image to at least one external device that communicates with the server device via a network connected to the server device.
  • a computing device can obtain a data image based on a vector set (or data point set, point data set, etc.) (330) corresponding to an input data set.
  • the computing device can obtain a vector set by mapping the data included in the input data set to an embedding space (or latent space) of a specific dimension.
  • the computing device can obtain a vector set by identifying a manifold formed by the data set in an embedding space of a specific dimension.
  • the manifold may refer to a shape that the input data set represents in an embedding space of a specific dimension.
  • the manifold may refer to an area where a vector set is identified or a shape formed by a vector set when mapping the input data set to a vector set in an embedding space of a specific dimension.
  • An IOD Information Object Descriptor
  • An IOD Information Object Descriptor
  • the shape or color of the visualized point may vary depending on the embodiment, and therefore the term "point" itself is not intended to limit the invention.
  • a point may be expressed using various terms depending on the embodiment. For example, a point may be expressed using terms such as a vector or feature appearing in an embedding space or latent space, but is not limited thereto.
  • the computing device can obtain a vector set by mapping the data set to an N-dimensional embedding space based on a predefined condition defined by a mapping function (e.g., a pre-stored matrix for mapping to an embedding space of a specific dimension). For example, the computing device can obtain a vector set by encoding the data set, but is not limited thereto. For example, the computing device can input the data set to a pre-trained encoder and obtain a vector set through the output layer of the encoder, but is not limited thereto.
  • a mapping function e.g., a pre-stored matrix for mapping to an embedding space of a specific dimension.
  • embedding refers to the process of converting high-dimensional data into low-dimensional vectors, preserving similarities and structural relationships between data points. This embedding is performed in a way that preserves the core information of the data while increasing computational efficiency.
  • each data point is represented as a vector, and these vectors can reflect the distribution and characteristics of the entire data set. This provides a foundation for analyzing the statistical characteristics and inherent patterns of the data set.
  • a computing device may include a data lens (400) for obtaining a data image (IOD) by embedding and visualizing a data set.
  • the data lens (400) may include at least one processing configuration for processing data.
  • the data lens (400) may include at least one neural network model (e.g., an encoder, etc.) for obtaining a vector set based on the data set and at least one visualization model (e.g., PCA, T-SNE, UMAP, etc.) for visualizing the data set based on the vector set to obtain a data image.
  • at least one neural network model e.g., an encoder, etc.
  • at least one visualization model e.g., PCA, T-SNE, UMAP, etc.
  • the data lens (400) may obtain a vector set corresponding to the data set by embedding the data set in an N-dimensional latent space, and may obtain a data image (IOD) corresponding to the data set by representing the vector set in an M-dimensional (e.g., 2-dimensional or 3-dimensional) imaging space (320).
  • IOD data image
  • FIG. 5 is a diagram illustrating a method for a computing device to obtain characteristics of a data set according to various embodiments.
  • the computing device can process the acquired data set to obtain characteristic information corresponding to the data set.
  • the properties of a data set or data may include information related to the distribution (e.g., geometric distribution or statistical distribution) of the data set or data. Specifically, the properties may include the property values of each data included in the data set. For example, a computing device may obtain property information indicating the distribution of the property values of the data included in the data set. Furthermore, the computing device may obtain the property information of the data set based on the statistical distribution, such as the mean, deviation, or variance, of the property values of each data.
  • the properties of a data set or data may include intrinsic characteristics related to the distribution of the data set itself.
  • the properties of a data set or data may include, but are not limited to, the density, homogeneity, bias, or distribution of the data set or data.
  • the properties of a data set or data may include task-dependent properties related to the task for which the data set is utilized (e.g., classification).
  • the properties of a data set or data may include, but are not limited to, the labeling error rate or the proportion of data pairs that are geometrically adjacent (hard-negative) but belong to different classes.
  • the computing device may store computational metrics corresponding to each characteristic of the data set or data in memory. More specifically, the computing device may store, but is not limited to, metrics for computing the density of the data set or data, metrics for computing the homogeneity of the data set or data, metrics for computing the bias of the data set or data, metrics for computing the distribution of the data set or data, etc.
  • the computing device can acquire data set characteristics based on stored operational metrics, using a data feature extraction algorithm built using an artificial neural network.
  • the feature extraction algorithm can be implemented using a feed-forward neural network.
  • the computing device may include, but is not limited to, a separate neural network for computing characteristics of a data set, or may include a neural network including layers for computing characteristics of a data set.
  • a computing device may include an artificial neural network for feature extraction designed to extract features of a data set when inputted with the data set.
  • the artificial neural network for feature extraction may be an artificial neural network that has undergone transfer learning to compute the features of the data.
  • a computing device can acquire characteristics of a data set by constructing an artificial neural network that adds a layer for extracting data characteristics to a neural network model (e.g., a data lens, an encoder, etc.) for providing data images based on the data set.
  • a neural network model e.g., a data lens, an encoder, etc.
  • the computing device can identify a vector set based on the data set and acquire characteristics of the data set or data based on the identified vector set.
  • the computing device can obtain the characteristic value of each data included in the data set by processing each vector included in the vector set with a predetermined algorithm.
  • the computing device can calculate the characteristic value based on the geometric distribution or statistical distribution of each vector included in the vector set, and can assign the calculated characteristic value to the corresponding data.
  • the characteristic value can be calculated based on the distance between vectors.
  • the characteristic value can be obtained based on the number of vectors existing within a predetermined distance from a specific vector (or data point), but is not limited thereto.
  • the characteristic value can be obtained based on the average value of the distances from a specific vector to a predetermined number of nearby vectors, but is not limited thereto.
  • the computing device can calculate the average distance value based on the distance values from a specific vector to K nearby vectors, and can obtain the first characteristic value (e.g., density, etc.) of the specific vector based on the calculated average distance value, but is not limited thereto.
  • the first characteristic value e.g., density, etc.
  • data quality is a concept that includes both quantitative and qualitative quality. Therefore, for successful training of an artificial intelligence model, it is necessary to (i) secure a sufficient amount of training data to train the artificial intelligence model, (ii) secure training data with high-quality inherent characteristics (e.g., unbiased distribution), and (iii) secure training data with characteristics (e.g., task-dependent properties) appropriate for the training purpose (e.g., task of the artificial intelligence model).
  • characteristics e.g., task-dependent properties
  • a computing device can synthesize, modify (or adjust), or remove data in a way that enhances the inherent characteristics and task-dependent characteristics of the data set to obtain high-quality learning data.
  • the computing device can improve the overall quality of a data set by removing at least some data from the data set.
  • a computing device can improve the learning efficiency of an artificial intelligence model that is learned by appropriately removing at least some data from a data set.
  • FIG. 6 is a diagram illustrating a data lens processing system and a data imaging system according to various embodiments.
  • the computing device (3500) can determine a data lens system that processes the data set to preserve inherent characteristics of the data set based on the data set.
  • a computing device can acquire a lens system corresponding to a data set based on a database. Specifically, the computing device can search the database for a lens system corresponding to the input data set based on the characteristics of the input data set.
  • a computing device can obtain a lens system corresponding to a data set based on a lens processing algorithm. Specifically, the computing device can calculate the optimal dimensionality that preserves the inherent characteristics of the input data set.
  • the computing device can obtain a data image representing the intrinsic characteristics of the data set using the imaging system (3510).
  • This disclosure provides a method for precisely identifying the inherent distribution of data by vectorizing it and intuitively visualizing it in two- or three-dimensional space, enabling the effective analysis and visualization of large data sets.
  • interaction and UI/UX implementation technologies are applied to minimize the gap between high-dimensional vector space and visualization space, enabling users to more easily understand and manipulate data quality.
  • FIG. 7 is a diagram illustrating a system in which a computing device visualizes data and provides user interaction functions, according to various embodiments.
  • the computing device (700) can provide a data image (720) through a network environment (e.g., a web or app environment).
  • the user device (701) can check the data image (720) through the network environment and obtain information about the checked data image (720).
  • the user device (701) may be an electronic device such as a laptop computer, a smartphone, a tablet, etc.
  • the user device (701) may be a wearable device such as a smart watch, smart glasses, smart glasses, etc.
  • the user device (701) may be a media device such as a streaming media device, a media player, an automobile entertainment system, etc.
  • the user device (701) may be an XR (mixed reality) device such as a device including a VR device or AR glasses, etc. That is, the user device (701) may be an electronic device that connects to a network to receive a data visualization service according to various embodiments of the present disclosure.
  • XR mixed reality
  • the user device (701) may be an electronic device that connects to a network to receive a data visualization service according to various embodiments of the present disclosure.
  • FIG. 8 is a diagram illustrating a data visualization and interaction method according to various embodiments.
  • the data visualization and interaction method may include a data imaging application step (S810), a data imaging and diagnosis step (S820), a data visualization step (S830), and an interaction step (S840).
  • the electronic device may request imaging, diagnosis, or visualization of a data set from the server. Specifically, the electronic device may transmit the data set and request information about the data set to the server. The server may process the data set based on the received data set and request information.
  • the server can image the data set by vectorizing it. Specifically, the server can identify a vector set corresponding to the data set by embedding the data set in a latent space of a specific dimension.
  • the server can diagnose the intrinsic properties and task-dependent attributes of the data set by analyzing the distance, neighbor relationship, or cluster structure between individual data points included in the data set based on the identified vector set.
  • the server can analyze the distance distribution from a specific data point to the K nearest neighboring vectors, or identify areas overly concentrated in a specific class to determine bias. Furthermore, to assess homogeneity, the server can measure the diversity or dispersion of the region to which each vector belongs.
  • the server can diagnose task-dependent characteristics by identifying how each class is distributed in vector space (e.g., whether there are boundaries or overlaps between classes) for a data set used in a classification task. This allows the server to detect instances of class imbalance or instances where a particular class is excessively close to another class (i.e., a large number of hard negatives).
  • the server can detect outliers or identify factors that degrade data quality (e.g., incorrect labels, duplicate data, noisy samples, etc.) by considering at least one of the Euclidean distance, cosine similarity, or other distances between vectors. Furthermore, the server can pass these analysis results to subsequent steps (e.g., data cleaning or label correction) to improve the quality of the dataset.
  • degrade data quality e.g., incorrect labels, duplicate data, noisy samples, etc.
  • subsequent steps e.g., data cleaning or label correction
  • the server can visualize and express the data set in a two-dimensional or three-dimensional space.
  • the server can reduce the dimensionality of the data set or the vector set corresponding to the data set to obtain a data image (IOD) and provide the data image in a two-dimensional or three-dimensional space.
  • IOD data image
  • the server can apply dimensionality reduction techniques such as PCA (principal component analysis), t-SNE, and UMAP, or use a neural network-based embedding model to map high-dimensional data into a low-dimensional visualization space.
  • PCA principal component analysis
  • t-SNE t-SNE
  • UMAP neural network-based embedding model
  • the server can provide visualized results that reflect the diagnostic results (cluster structure, outliers, bias, class imbalance, etc.) produced in the previous step (e.g., data diagnosis step).
  • the server can display data points identified as outliers with a separate color or special marker (e.g., triangles, stars, etc.), or visually highlight areas with ambiguous class boundaries or high-density areas (e.g., region borders, gradients), thereby enabling users to grasp data distribution and quality issues at a glance.
  • a separate color or special marker e.g., triangles, stars, etc.
  • visually highlight areas with ambiguous class boundaries or high-density areas e.g., region borders, gradients
  • the server can process inputs related to the visualized data image received from an electronic device (e.g., a user device) to provide a dynamic response to the visualized data image.
  • an electronic device e.g., a user device
  • the user can perform inputs such as touch, mouse click, drag, or pinch zoom on any point within the visualized data image (e.g., a specific data point, cluster, or selected area).
  • the computing device (server) of the present disclosure may provide a response by retrieving metadata (e.g., individual feature values, label information, statistical values, etc.) associated with the data point and transmitting the information to the user device, which may be displayed in a pop-up window, tooltip, or separate layer.
  • metadata e.g., individual feature values, label information, statistical values, etc.
  • the server may further analyze the statistical distribution, density, class distribution, etc. of the data points within the designated area and then provide the results (e.g., mean value, deviation, representative image, number of samples, etc.) to the user device.
  • the server reapplies dimensionality reduction techniques (PCA, t-SNE, UMAP, etc.) or color-shape mapping algorithms, or adjusts parameters to generate an updated data image and deliver it to the user's device.
  • the interaction step (S840) of the present disclosure enables users to intuitively and immediately explore visualized data, allowing them to gain a deeper understanding of the data's inherent structure or to detect potential quality issues (e.g., labeling errors, outliers, unclear boundaries between clusters, etc.) early. Furthermore, this interaction feature can be effectively utilized by machine learning model developers during data cleaning or label correction tasks.
  • FIG. 9 is a drawing for explaining detailed steps performed in an imaging application step according to various embodiments.
  • users can input settings for tools to be used for data imaging and diagnosis (e.g., artificial intelligence models for data imaging, such as lenses), data visualization tools, etc.
  • tools to be used for data imaging and diagnosis e.g., artificial intelligence models for data imaging, such as lenses
  • data visualization tools e.g., data visualization tools, etc.
  • the data imaging application step (S810) may include a lens level selection step (S811), a lens property determination step (S813), an other setting step (S815), and a visualization tool selection step (S817).
  • the user can select a level of a lens including at least one artificial intelligence model to be used for data imaging.
  • the level of the lens can be classified and defined according to preset criteria related to the data imaging method.
  • the level of the lens can include, but is not limited to, a first level that diagnoses only quantitative indicators of the data set (basic statistics, missing value ratio, data outlier detection, etc.), a second level that diagnoses by vectorizing the data set using a pre-stored (pre-trained) artificial intelligence model, or a third level that diagnoses by further learning (fine-tuning) or newly learning an artificial intelligence model optimized for the data set and then vectorizing it.
  • the third level can be selected to directly retrain the model to obtain more precise diagnosis results.
  • the user can determine the properties of a lens including at least one artificial intelligence model to be used for data imaging.
  • the properties of the lens may include the structure of the artificial intelligence model, hyperparameters (e.g., number of layers, number of parameters, learning rate, batch size, etc.), or exploration strategies (e.g., type of optimizer, initialization method).
  • hyperparameters e.g., number of layers, number of parameters, learning rate, batch size, etc.
  • exploration strategies e.g., type of optimizer, initialization method.
  • the user can adjust the "model complexity (number of parameters)" or set the "number of learning epochs" to balance analysis and processing time.
  • the user can also select whether to use a pre-trained model suited to a specific domain (e.g., image, text, structured data, etc.) or a generic model (generic AI model).
  • the user device can set whether to perform data diagnosis by selecting whether to skip the diagnosis process and perform only simple visualization or to perform diagnosis as well. Furthermore, the user device can set whether to visualize the diagnosis results by deciding whether to display cluster density or bias indicators on a separate color scale. Furthermore, the user device can set community creation criteria, such as similarity (distance) criteria between data points or the method of creating subgroups or communities based on specific characteristics (e.g., label information). Furthermore, the user device can set whether to create snapshots that separately store and compare or analyze the data distribution (visualization state) at a certain point in time (viewpoint).
  • the user can select at least one of multiple pre-linked visualization tools (e.g., PCT, UMAP, T-SNE, etc.).
  • the user can select the tool based on its characteristics (e.g., dimensionality reduction method, ease of interpretation of results, processing speed, etc.).
  • a preset visualization dimension e.g., 2D, 3D
  • the user can decide whether to view the data in a simple plane (projection) or to interact with it, such as by rotating or zooming in 3D space.
  • PCA is fast and intuitive to interpret
  • t-SNE and UMAP offer the advantage of more precisely reflecting data clustering structures.
  • Users can choose a visualization tool based on a comprehensive consideration of the characteristics of the dataset, analysis objectives, computing resources, and other factors.
  • the data imaging application step (S810), users can directly input various settings and decisions to control the data imaging and diagnostic process so that it is optimized for their analysis purposes and environment.
  • This user-centric configuration step facilitates the vectorization, diagnostics, and visualization processes performed in subsequent steps (S820, S830, etc.), ultimately contributing to improved data quality and enhanced AI model performance.
  • FIG. 10 is a diagram illustrating the results of data imaging and diagnosis according to the level of the lens and the visualization tool according to various embodiments.
  • FIG. 11 is a diagram illustrating a lens builder included in a computing device according to various embodiments.
  • the computing device can build a lens for imaging a data set according to settings input by the user, and can diagnose and visualize the data set using the built lens and visualization tool.
  • a computing device can build a lens for imaging a data set according to settings input by a user, and can diagnose and visualize the data set using the built lens and visualization tool.
  • a computing device e.g., a server
  • the lens builder is a component for building an appropriate lens (e.g., a set of tools for data imaging and diagnosis) by configuring or combining, learning, or tuning tools (e.g., artificial intelligence models, statistical calculators, etc.) for processing the data set, reflecting the user's settings for imaging and diagnosing the data set.
  • the lens builder can be implemented as a separate, independent hardware device or software module, or can be implemented in a form in which at least one processor executes a plurality of instructions stored in memory for building a lens.
  • the lens builder may include, but is not limited to, a model storage unit including a plurality of artificial intelligence models, a model learning unit for learning the artificial intelligence models, an attribute determination unit for determining attributes of the artificial intelligence models, a calculation tool storage unit including a plurality of calculators for calculating statistical characteristics of data, and a lens storage unit for storing constructed lenses.
  • the calculation tool storage unit may store a plurality of calculators (e.g., a first calculator, a second calculator, ⁇ ), and the lens storage unit may store a plurality of lenses (e.g., a first lens, a second lens, a third lens, a fourth lens, ⁇ ), but is not limited thereto.
  • a plurality of calculators e.g., a first calculator, a second calculator, ⁇
  • the lens storage unit may store a plurality of lenses (e.g., a first lens, a second lens, a third lens, a fourth lens, ⁇ ), but is not limited thereto.
  • the computing device can generate a second lens by loading one or more models that meet user requirements or data formats (images, text, structured data, etc.) from among a plurality of pre-learning models stored in a first storage of the model storage (e.g., a CNN model for computer vision, a Transformer model for text embedding, etc.).
  • a CNN model for computer vision
  • a Transformer model for text embedding, etc.
  • the computing device may provide a vector set corresponding to the data set (e.g., a first vector set or a second vector set) to at least one visualization tool, thereby providing a data image (Image of Data, IOD) visualizing the data set to the user.
  • a vector set corresponding to the data set e.g., a first vector set or a second vector set
  • IOD Image of Data
  • the computing device can generate and output at least one of a plurality of different types of data images (IODs) depending on the type of visualization tool or the dimension visualized by the visualization tool.
  • IODs data images
  • the computing device can project the vector set into two dimensions, obtain a two-dimensional first data image (IOD #1), and then provide it.
  • a first visualization tool e.g., PCA, a 2D-based dimensionality reduction algorithm
  • the dimensionality of the data can be reduced and projected into two, three, or other dimensions using a third visualization tool (e.g., t-SNE, UMAP, etc.) to obtain a third data image (IOD #3).
  • a third visualization tool e.g., t-SNE, UMAP, etc.
  • the "Image of Data (IOD)" in the present disclosure refers to the result of visualizing a vector set, but is not limited to a specific standard (2D/3D) or a specific technique (PCA/t-SNE/UMAP/Autoencoder-based visualization, etc.).
  • the computing device may receive user interactions (e.g., zooming in and out, clicking on a specific point, specifying a range, changing a color scale, etc.) for the visualization result in real time, and update the visualized data image or display additional details (e.g., metadata corresponding to each point).
  • These reports can be delivered to users (e.g., data analysts, AI model developers, etc.) in a structured form, and can be immediately utilized in the analysis or model modification stage by being displayed as charts, tables, or heatmaps in a visualization interface (GUI).
  • GUI visualization interface
  • the first data processing model may include at least one artificial intelligence model.
  • the first data processing model may include at least one pre-stored pre-learning model or an artificial intelligence model trained to be optimized for the first data set. For example, based on an input requesting imaging and diagnosis optimized for the first data set, the at least one processor may generate the first data processing model by training the artificial intelligence model optimized for the first data set, but is not limited thereto.
  • the at least one attribute may include an attribute related to an operation for processing data by the first data processing model.
  • the at least one attribute may include, but is not limited to, the number of parameters or hyperparameters of the artificial intelligence model or graphics card information.
  • FIG. 13 is a diagram for explaining information included in dictionary information according to various embodiments.
  • the prior information may include cost information related to the cost required for processing a data set, result preview information that predicts and displays the results of processing the data set in advance, progress information indicating the progress of processing the data set, reference result information indicating the processing results of cases similar to the data set, or attribute information indicating the properties of the processing model used for processing the data. Accordingly, the user can use the service more smoothly by checking the expected processing cost, predicted results, model progress, etc. in advance at the time of requesting data set imaging and diagnosis.
  • cost information may refer to information related to the cost required to process a data set.
  • cost information may include an estimate of the computational cost required to process a data set (e.g., GPU time, CPU core time, memory usage, etc.) or a payment fee (e.g., credits, points, currency, etc.) that a user must pay to use the processing service.
  • a payment fee e.g., credits, points, currency, etc.
  • “result preview information” may refer to information that briefly presents some key results expected after data set processing, allowing the user to estimate the general form or value of the results.
  • the computing device may briefly display the expected accuracy range after AI model training, a preview of cluster distribution, or a sample visualization image.
  • the result preview information may include data images obtained by simply inputting a data set into a visualization tool. While such data images may not accurately reflect the inherent characteristics of the data set, by providing the visualization results of the data set in advance, they can encourage the user to predict the results.
  • the progress information may refer to information about a processing algorithm or processing status of a data set.
  • the progress information may include the learning progress, training time, or remaining time for optimization learning based on a data set.
  • the user can predict the completion time of processing or determine whether additional resources (e.g., computing resources) are allocated.
  • additional resources e.g., computing resources
  • a warning message may be displayed alongside the progress to provide the user with an opportunity to take action in advance.
  • reference result information may refer to information that indirectly estimates expected results or performance indicators by providing examples of processing results previously performed on cases similar to the data set submitted by the user.
  • reference result information may include imaging and diagnostic results data for reference data similar to the input first data set. For example, it may be provided in the form of "When analyzing 100,000 image data of the same category (domain), the average accuracy was 92%, and the processing time was approximately 4 hours.”
  • attribute information may refer to information about attributes associated with a data processing tool.
  • attribute information may include the number of model parameters or graphics card information.
  • the information may be expressed as, "This processing uses a Transformer architecture-based model (approximately 100 million parameters) and an RTX 3090 GPU.” This information allows users to understand model size, resource compatibility (whether internal or cloud GPU), development environment, etc., and predict model utilization and learning performance in advance.
  • At least one processor may be configured to perform an operation (S1250) of constructing a first data processing model based on a user input to the first GUI.
  • the user input to the first GUI may include an approval input for processing the first data set proposed by the computing device.
  • the computing device can process the first data set only if, prior to processing the first data set, it provides the user with prior information about the processing of the first data set, and if an approval input is received from the user who has been provided with the prior information.
  • FIG. 14 is a diagram illustrating an example of a computing device constructing a first data processing model according to various embodiments.
  • a computing device or at least one processor included in the computing device may be configured to perform an operation (S1401) of obtaining a plurality of representative values from a plurality of adapters based on a first data set.
  • the at least one processor may preprocess the first data set to extract a feature value corresponding to the first data set and input the feature value to each of the plurality of adapters.
  • the at least one processor may obtain a plurality of representative values based on data output from each of the plurality of adapters into which the same data has been input.
  • each adapter that receives feature values can output a vector set corresponding to the first data set.
  • at least one processor can obtain a representative value by operating the output vector set in a predetermined manner.
  • the multiple vector sets output from the multiple adapters can be defined based on different dimensions.
  • the multiple adapters can be configured to output vector sets of different dimensions.
  • At least one processor may be configured to perform an operation (S1402) of selecting a first adapter that satisfies a predetermined condition based on a plurality of representative values.
  • the predetermined condition may include a condition set for determining an adapter optimized for the first data set.
  • at least one processor may select the first adapter by identifying at least one adapter corresponding to the lowest (or highest) value among the plurality of representative values.
  • the first adapter may be implemented to embed the first data set in an optimal dimension representing the first data set.
  • the first adapter may be configured to output a first vector set of a first dimension based on the first data set.
  • At least one processor may be configured to perform an operation (S1404) of constructing a first data processing model including a first adapter. Specifically, at least one processor may construct a first data processing model including a foundation model or a pre-trained model and the first adapter by communicatively connecting the first adapter to a pre-stored foundation model or a pre-trained model.
  • Computing devices can build data processing models optimized for a given dataset by further training or fine-tuning only the adapters based on the dataset. Because computing devices only train adapters on pre-stored AI models, they can minimize computational costs.
  • FIG. 15 is a diagram illustrating another example of a computing device constructing a first data processing model according to various embodiments.
  • the computing device or at least one processor included in the computing device may be set to perform an operation (S1501) of inputting a first data set into a plurality of foundation models.
  • At least one processor may be configured to perform an operation (S1502) of obtaining a plurality of representative values based on data output from a plurality of foundation models.
  • each of the plurality of foundation models may output a vector set corresponding to a first data set, and at least one processor may obtain a representative value by operating the output vector set in a predetermined manner.
  • the plurality of vector sets output from the plurality of foundation models may be defined based on different dimensions.
  • the plurality of foundation models may be configured to output vector sets of different dimensions.
  • At least one processor may be configured to perform an operation (S1503) of determining a first foundation model satisfying a predetermined condition based on a plurality of representative values, and an operation (S1504) of constructing a first data processing model including the first foundation model. Since the specific method of determining the model based on the predetermined condition and the plurality of representative values has been described above, a detailed description thereof will be omitted.
  • FIG. 16 is a diagram illustrating another example of a computing device constructing a first data processing model according to various embodiments.
  • a computing device or at least one processor included in the computing device may perform an operation (S1601) of inputting a first data set into a first foundation model.
  • the first foundation model may include a plurality of nodes implemented to receive the same input and output different result values.
  • the first foundation model may include at least one hidden layer including a plurality of nodes into which the same input is input in parallel.
  • At least one processor may perform an operation (S1602) of obtaining a plurality of representative values from a plurality of nodes included in the first foundation model, wherein the plurality of nodes input feature values for the first data set in parallel. Specifically, the plurality of nodes output a plurality of vector sets based on the feature values for the first data set, and at least one processor may obtain a plurality of representative values from the plurality of vector sets based on a predetermined operation.
  • S1602 an operation of obtaining a plurality of representative values from a plurality of nodes included in the first foundation model, wherein the plurality of nodes input feature values for the first data set in parallel. Specifically, the plurality of nodes output a plurality of vector sets based on the feature values for the first data set, and at least one processor may obtain a plurality of representative values from the plurality of vector sets based on a predetermined operation.
  • At least one processor may be configured to perform an operation (S1603) of activating a first node based on a plurality of representative values. Additionally, at least one processor may be configured to perform an operation (S1604) of constructing a first data processing model including a first foundation model in which the first node is activated.
  • At least one processor can determine a first node that satisfies a predetermined condition based on a plurality of representative values, and activate the first node so that input is only input to the first node.
  • at least one processor can select a node optimized for the first data set among the plurality of nodes, and establish a first data processing model by disconnecting communication with nodes other than the selected node.
  • At least one processor may obtain a first vector set defined in a first-dimensional embedding region from a first data processing model (S1270).
  • Each of the plurality of vectors (or data points) included in the first vector set may correspond to each unit data included in the first data set.
  • At least one processor may provide a first data image corresponding to the first data set by providing the first vector set to the first visualization tool (S1290). At this time, at least one processor may also recommend a visualization tool appropriate for visualizing the first data set based on the first vector set or a plurality of feature values.
  • a computing device may include a visualization tool DB or communicate with a DB server that manages information related to visualization tools in advance, and perform an algorithm for searching and determining an optimal visualization tool based on data characteristics.
  • a computing device or at least one processor included in the computing device may perform a data characteristic extraction step (S1810).
  • the computing device may analyze or extract characteristic values such as a domain, data type, data capacity, vector distribution density, number per class, etc. for a data set (or vector set).
  • at least one processor may extract statistics such as a data domain, data size (number of samples) and number of dimensions, average distance between vectors, variance, number of clusters, etc., or model learning specifications (required computing resources, processing time, etc.) based on the data set.
  • the computing device can perform a visualization tool DB query and SQL query generation step (S1820).
  • the computing device (or DB server) can manage metadata for each visualization tool (PCA, T-SNE, UMAP, Autoencoder-based visualization, etc.) in the visualization tool DB (e.g., in table form).
  • the computing device can generate a SQL query based on a matching rule between data characteristics and tool metadata to search for "candidates that satisfy specific conditions (domain, data size, cluster structure importance, operation time constraints, etc.) among available visualization tools.”
  • the computing device can perform the recommendation algorithm execution and result generation step (S1830). Specifically, when the DB search results are returned in multiple visualization tools, the computing device can apply a score calculation or weight-based algorithm. For example, the computing device can be configured to increase the preference for PCA or UMAP when the amount of data is very large, to increase the score of t-SNE or UMAP when cluster accuracy (detailed clustering) is important, or to prioritize PCA with low computational complexity when real-time interaction is required.
  • the final recommendation priority is determined, and a list such as "t-SNE (1st), UMAP (2nd), PCA (3rd)" can be presented to the user.
  • the computing device can perform the final decision step (S1840) on a visualization tool based on a user request. For example, the user may be informed that "t-SNE is the most suitable tool based on data characteristics and priority criteria," and the user can then decide whether to use the recommended tool as is or select a different tool.
  • the pre-recommendation results allow the user to recognize key pros and cons (e.g., visualization quality vs. processing time) in advance.
  • At least one processor can execute the corresponding algorithm on the first vector set to generate and provide a first data image.
  • other candidate tools e.g., the second and third visualization tools
  • a computing device can reflect the type of data set (image, text, structured data, etc.), domain (medical, social networking services, finance, etc.), data volume, vector distribution characteristics (density, variance, number of clusters, etc.), processing time constraints, etc., into a "visualization tool recommendation" algorithm, thereby guiding a user to select a visualization tool more rationally. Furthermore, by querying the visualization tool database using SQL to identify available candidates and calculating priorities among the candidates, the user can obtain visualization results that enable efficient understanding of a high-dimensional embedding space.
  • the computing device provides an environment in which data distribution or cluster structure can be accurately and intuitively understood by appropriately utilizing the strengths and weaknesses of each technique such as PCA, T-SNE, and UMAP through a process of recommending a visualization tool based on data characteristics.
  • FIG. 19 is a diagram illustrating a method for a computing device to generate synthetic data based on data imaging according to various embodiments.
  • a computing device or at least one processor included in the computing device can generate synthetic data based on a vector set corresponding to a data set.
  • At least one processor can obtain a vector set by imaging a dataset using a lens including at least one artificial intelligence model, and can visualize the dataset using at least one visualization tool to display a data image (IOD) corresponding to the dataset in a visualization space.
  • IOD data image
  • FIG. 20 is a diagram illustrating a method for a computing device to generate synthetic data based on user input according to various embodiments.
  • a computing device or at least one processor included in the computing device may visualize a data set and provide a first data image corresponding to the data set (S2010). For example, by utilizing visualization tools such as the aforementioned PCA, t-SNE, and UMAP, each unit data included in the data set may be projected into a two-dimensional or three-dimensional space, and then a data image (IOD) including multiple data points may be displayed on a GUI (Graphical User Interface).
  • IOD Data image
  • At least one processor can receive a user input for a first area on the first data image (S2020).
  • a user's input such as a click, touch, or drag can be recognized for a first area (R) representing a specific point or range on the first data image (IOD).
  • the user input is an input event that occurs on the GUI, and includes various forms such as a left mouse click, a mobile touch, or a pen drawing.
  • At least one processor can determine a latent code corresponding to the first region (S2030).
  • the latent code is utilized as a model internal embedding or noise vector when generating synthetic data.
  • At this time, at least one processor may request feedback regarding the first region specified by user input.
  • the at least one processor may provide the user with a graphical user interface (GUI) (e.g., a pop-up window or highlighting) that visualizes the first region or displays a confirmation message (e.g., "Is this region the target region for generating synthetic data?").
  • GUI graphical user interface
  • At least one processor may define conditions for data generation based on user input. Specifically, at least one processor may define data generation conditions for determining potential code to be input into the data generation model.
  • At least one processor may determine the latent code in various ways based on the properties of the first region. For example, at least one processor may selectively execute an algorithm that determines the latent code based on the presence or absence of data points or the number of data points contained in the first region.
  • At least one processor may determine a latent code based on at least one vector corresponding to at least one point adjacent to the first region.
  • at least one processor may derive a target region including the first region where the user input was made and at least one point adjacent thereto, and then determine a latent code based on at least one vector corresponding to at least one point included in the target region.
  • At least one processor can determine a latent code based on at least two or more vectors. Specifically, at least one processor can identify at least two or more vectors corresponding to at least one point adjacent to the first region, and determine a latent code corresponding to the first region based on the identified at least two or more vectors. For example, by setting a vector interpolated by vectors corresponding to points near the first region as a latent code, synthetic data reflecting the typical (centroid-like) characteristics of the corresponding region can be generated.
  • a single point is included in the first region, it can be defined as a target vector, and a latent code can be generated based on the relationship (distance, density, label, etc.) between adjacent vectors around this target vector.
  • At least one processor may identify at least one area requiring improvement based on diagnostic results for the data set. At least one processor may visually display the identified at least one area via a display of the user device.
  • At least one processor can activate at least one identified region to enable interaction. Specifically, at least one processor can enable user input for at least one region requiring improvement. In this case, a data improvement process for the region requiring improvement can be performed based on user input regarding the at least one region requiring improvement.
  • At least one processor can extract at least one feature of a plurality of unit data included in the data set based on a vector set that embeds the data set in a latent space of a specific dimension. Based on the at least one extracted feature, the at least one processor can identify data in need of improvement and visually display data points corresponding to the data in need of improvement or a specific region containing the data points.
  • At least one processor can generate synthetic data by further reflecting a user's prompt input.
  • at least one processor can receive a user prompt (PROMPT) through an input interface provided on a GUI (e.g., a text field, voice input, a slider, etc.).
  • PROMPT user prompt
  • the user prompt can express the topic, style, properties, constraints, etc. of the synthetic data in natural language or tag-based.
  • the user prompt can include text instructions in the form of "Generate an image of a cat with a blue background” or “Silhouette format with emphasized perspective.” Additionally, if the user prompt specifies a numeric parameter or categorical tag (e.g., "realistic", "cartoon”), the corresponding value can be interpreted as prompt-internal information and passed to the model.
  • a numeric parameter or categorical tag e.g., "realistic", "cartoon”
  • the generative model can determine the basic shape or feature distribution based on the latent code, additionally reflect the subject, style, or detailed properties through prompts (e.g., text, tags, or parameters), and output the final synthetic data. For example, if the latent code is the result of "interpolation between two vectors," it already reflects intermediate properties (e.g., mixing two classes) or spatial characteristics. If the prompt is given as "cat + blue background + cartoon style,” the style or thematic elements can be reflected through the internal conditioning path of the model, resulting in a synthetic image with the corresponding characteristics.
  • prompts e.g., text, tags, or parameters
  • a prompt (PROMPT) entered into an input interface via a GUI along with a determined latent code can be set as a constraint on data generation.
  • At least one processor can input the latent code and the prompt into a generation model, and the generation model can generate synthetic data conditioned by the instructions given by the prompt based on the latent code.
  • At least one processor can convert the user prompt into a vector form based on a text embedding module (such as CLIP, BERT, GPT) or a conditional network (such as a conditional layer or cross-attention), and then cross-attention it to the hidden layer of the generative model or inject it as a conditional input.
  • a text embedding module such as CLIP, BERT, GPT
  • a conditional network such as a conditional layer or cross-attention
  • At least one processor may be implemented to reflect user requests at each step (e.g., diffusion step, GAN upsampling step, etc.) based on logic that combines latent codes and text embeddings (e.g., concat, add, attention).
  • the user can iteratively (refinement synthesis) by modifying the prompt or resetting the latent code (interpolation parameters, selecting additional vectors, etc.).
  • customized synthetic data can be easily obtained by immediately reflecting the conditions desired by the user, compared to when only a latent code obtained by simply interpolating multiple vectors is used.
  • prompt-based generation is advantageous in generating realistic (distribution-preserving) results, since the latent code already reflects the inherent characteristics of the data set.
  • creative uses e.g., content creation, data augmentation, simulation, etc.
  • artificial intelligence models are facilitated by efficiently testing various scenarios (e.g., adversarial examples, artistic expressions, etc.) through prompt changes.
  • At least one processor can automatically extract or generate a prompt from a user-specified first field, thereby utilizing the prompt as a condition for generating synthetic data. This allows for synthetic results that reflect the inherent properties or metadata of the data, without requiring the user to directly input separate text.
  • At least one processor can automatically generate a prompt by analyzing label information, class/category, metadata (e.g., time, location, ID, etc.), visual features (if the data point is an image), etc., for data points belonging to or adjacent to the first region. Specifically, if it is inferred that the region is a "region of cat faces," the at least one processor can generate a simple phrase such as "cat face cluster” or more specific text such as "Multiple cat faces in a close-up shot.”
  • At least one processor can generate prompts using image captioning technology (e.g., vision-language model, optical character recognition (OCR), etc.), and for structured data, it can automatically convert category names or key attribute names into sentence form.
  • image captioning technology e.g., vision-language model, optical character recognition (OCR), etc.
  • At least one processor may use an image captioning model (e.g., a CNN+LSTM architecture, a Transformer-based visual-language model, etc.) to input an image from a first region and automatically output a sentence-based description (caption).
  • an image captioning model e.g., a CNN+LSTM architecture, a Transformer-based visual-language model, etc.
  • the user may be provided with a GUI interface that allows them to "view and edit the extracted phrases," manually modifying the text as desired and then adopting it as the final prompt.
  • At least one processor can input the determined latent code or other vector interpolation results and automatically generated prompts into a generative model (GAN, VAE, Diffusion, etc.).
  • the generative model can output conditional synthetic data that reflects the data distribution inferred from the latent code, as well as context, object information, and properties extracted from the user domain.
  • the computing device generates linguistic constraints along with latent code, thereby enabling "dimensional expansion" (the generation of data with properties not available with existing data).
  • the computing device can construct an N+3-dimensional data lens by additionally learning synthetic data generated based on linguistic constraints for a data lens that outputs vectors on an N-dimensional latent space.
  • a computing device can perform an interaction operation to remove at least some data included in a data set based on a user input.
  • FIG. 22 is a diagram illustrating a method for a computing device to remove data based on user input, according to various embodiments.
  • a computing device or at least one processor included in the computing device may receive a user input regarding a first region including at least one point on a visualized data image (S2210).
  • the first region may be a dense region (a region with very high data density within a cluster) or an region that a user wishes to remove, such as a region containing a large number of outliers.
  • the at least one processor may provide the user with information regarding a region requiring data removal (e.g., a dense region or an outlier region), and the user may provide input regarding the region.
  • the first region may include at least one cluster whose data characteristics satisfy predetermined conditions.
  • At least one processor may remove at least a portion of the data corresponding to a plurality of points included in the first region (S2220).
  • "at least a portion” refers to only data that corresponds to all or specific conditions (e.g., an outlier filter, a specific label, etc.) specified by the user.
  • At least one processor may selectively perform a data removal algorithm based on the properties of the first region. For example, if the at least one processor determines that the data clusters within the first region are excessively dense and cause imbalance, the at least one processor may undersample all data points within the region or remove them at an arbitrary rate. Furthermore, for example, the at least one processor may selectively remove only data that satisfy predefined outlier identification criteria (e.g., distance, density, label mismatch, etc.) among points in the region by determining to remove only outliers within the first region.
  • predefined outlier identification criteria e.g., distance, density, label mismatch, etc.
  • the at least one processor may provide refinement options (e.g., filter conditions) to remove only points with a specific label or outside a specific statistical range within a user-specified region (e.g., "Remove only points that fall within this region and have a label of 0").
  • refinement options e.g., filter conditions
  • At least one processor can provide a data image with at least some data removed (S2230).
  • S2230 By examining the newly visualized distribution, the user can intuitively understand the effects of improved data imbalance or noise reduction. If necessary, deletion history can be maintained or an Undo function can be provided to prevent data loss due to user error.
  • At least one processor can perform reimaging and diagnostics (cluster analysis) on the data set after removal to check again how much the data quality has improved and whether there is an advantage in model learning.
  • Embodiments of the present disclosure enable direct removal through user interaction, thereby reflecting fine-grained domain knowledge that may be missed by existing automated algorithms.
  • immediate feedback can be obtained in a visualized space (2D or 3D), allowing for quick determination of follow-up actions, such as "how much data has been lost” and "how has the distribution changed”.
  • an intuitive and flexible tool for data quality management can be provided by allowing the user to interactively remove data in a specified area (particularly, a dense cluster area, an outlier-rich area, etc.).
  • FIG. 23 is a diagram illustrating a method for a computing device to perform data improvement in response to a data improvement request and provide visual interaction therefor, according to various embodiments.
  • a computing device or at least one processor included in the computing device may receive a data improvement request (S2310).
  • the at least one processor may recognize the data improvement request by receiving user input via at least one GUI that directs data improvement. For example, if a button for improving data in a first manner (e.g., "Resolve data imbalance") or a button for improving data in a second manner (e.g., "Remove noise area”) is selected on the user GUI, the at least one processor may recognize that an improvement request has occurred through the corresponding input. Additionally, the user may recognize that a specific class (or attribute) is lacking and request, for example, "Please create more of that class.”
  • At least one processor can visually represent at least one area requiring data improvement (S2320). Specifically, at least one processor can acquire characteristics of the data set based on a vector set corresponding to the data set, and detect areas requiring data improvement based on the characteristics of the data set.
  • At least one processor can analyze the density of the data set to identify areas where a specific class is underrepresented or where certain regions (clusters) are overly dense. Furthermore, for example, at least one processor can analyze the bias of the data set to automatically identify areas where the class distribution is imbalanced and requires improvement, such as areas where a specific class is underrepresented or overrepresented. Furthermore, for example, at least one processor can detect areas where a large number of outliers exist (areas where noisy data is concentrated) by performing outlier analysis.
  • At least one processor can highlight or outline the detected "areas requiring improvement" on the visualized data image (IOD) to notify the user.
  • at least one processor can provide interactive guidance (e.g., tooltips) on the GUI, explaining the criteria for selecting the areas.
  • At least one processor may perform data improvement by generating or removing data for each of at least one region (S2330). For example, at least one processor may generate synthetic data based on the data generation method according to FIG. 19 for a first region on a data image determined to be lacking in data of a specific class. In addition, for example, at least one processor may remove at least some data based on the data removal method according to FIG. 22 for a second region that is over-dense with data or contains outliers. At this time, the user may select at least one of various detailed options (e.g., execute, cancel, adjust detailed options, etc.) for the automatic suggestion.
  • various detailed options e.g., execute, cancel, adjust detailed options, etc.
  • At least one processor can provide a visualized improved data image corresponding to the results of the data enhancement (S2340). This allows users to visually see how the data distribution has changed, intuitively understand the extent to which imbalances have been resolved, and the extent to which outliers have been reduced.
  • at least one processor can also provide users with a comparison view (before/after visualization) with the previous state or a diagnostic report.
  • users can perform manual corrections by deselecting some of the visualized areas for improvement or by instructing the user to expand the scope of creation or removal.
  • at least one processor can automatically identify areas that require further improvement based on the improvement results, repeatedly performing a loop to improve data quality.
  • FIG. 24 is a diagram illustrating an example of a computing device in which a data processing method including a snapshot function is implemented, according to various embodiments.
  • a computing device (2400) may include multiple components for acquiring snapshot information based on a specific scene on a data image.
  • the multiple components are arbitrarily separated to perform specific operations, and may be physically separate components, or may be separate components based on various operations implemented on a single software program.
  • the computing device (2400) may include a screener for screening a data image or a vector set corresponding to the data image, an event detector for detecting whether an event for capturing a snapshot has occurred, a capture unit for capturing a snapshot, a correction unit for correcting a captured scene, and a generator for generating various information about the snapshot.
  • the screener can screen a vector set or a data image.
  • the screener can be configured with at least one metric for measuring characteristic values (e.g., density, bias, homogeneity, presence of outliers, etc.) of the data set based on the vector set or the data image.
  • the screener can analyze basic characteristics (e.g., density, outlier ratio, etc.) in a 2D/3D data image based on a first metric, or measure high-dimensional distribution characteristics in a high-dimensional vector set based on a second metric.
  • An event detector can detect events occurring during screening. Specifically, the event detector can monitor the measured results from the screener and detect an event occurrence when the monitored results satisfy predetermined conditions. At this time, the event detector can store information about the data in which the event occurred (e.g., data items, vector values corresponding to the data, coordinates of data points corresponding to the data, etc.). For example, if the event detector detects a specific area with a higher data density than a threshold value in the screener results, or a specific area with a bias index exceeding a reference value, the event detector can recognize this as an "event occurrence" and generate a signal.
  • information about the data in which the event occurred e.g., data items, vector values corresponding to the data, coordinates of data points corresponding to the data, etc.
  • the event detector can recognize this as an "event occurrence" and generate a signal.
  • the generator can generate information about events detected by the event detector. Specifically, the generator can generate tag information by synthesizing event information detected by the event detector, the location and metadata of the captured scene, and feature values calculated by the screener. Alternatively, the generator can record additional information related to the snapshot (class label, time, user ID, etc.). That is, the generator can generate tag information based on information about data in which an event occurred, which is pre-saved or generated from the event detector.
  • the tag information generated by the generator can include, for example, identification labels such as "areas with unusual characteristics (concentrated outliers)" or "sections with excessively high density,” the time of snapshot capture, key feature values, etc.
  • the capture unit can capture and save a specific scene on a data image. For example, when a signal notifying the occurrence of an event is transmitted from an event detector or when a user directly commands a snapshot, the capture unit can capture (photograph) a scene containing at least one data point corresponding to an event on the data image at a specific viewpoint.
  • “capture” includes a process of saving the state of an actual GUI screen (2D/3D view) or an internal data structure (vector set + visualization mapping parameters) as an image (or video frame).
  • the correction unit can perform modifications on captured scenes (snapshots). This provides users with highly visible results and, if necessary, can also perform graphical corrections, such as visually highlighting areas or adding labels.
  • a computing device can provide a vector set corresponding to a data set to a screener, and the screener can diagnose the characteristics of the data set based on the vector set.
  • An event detector can identify whether an event has occurred based on the characteristics diagnosed by the screener. For example, if the characteristics diagnosed by the screener satisfy a predetermined condition, the event detector can generate a signal indicating that an event has occurred.
  • a capture unit can capture a snapshot corresponding to an event based on the signal indicating that an event has occurred. Specifically, the capture unit can generate a snapshot by capturing a scene including at least one data point corresponding to at least one vector associated with the event at a specific point in time.
  • a generator can generate tag information based on information about an event that has occurred or information about a location on a data image where an event has occurred.
  • a correction unit can correct the generated snapshot in a predetermined manner (e.g., brightness adjustment, contrast adjustment, etc.).
  • FIG. 25 is a diagram illustrating a method for a computing device to screen a data set to generate snapshot information, according to various embodiments.
  • a computing device or at least one processor included in the computing device may screen a vector set by calculating at least one characteristic value based on at least one vector included in the vector set using a screener having at least one metric set (S2510).
  • at least one processor may diagnose the characteristics of data and measure various indicators such as data density, bias, and homogeneity.
  • screening may proceed along a specific path based on a specific starting point (coordinate location) on a data image (2D/3D) or vector set.
  • at least one processor may sequentially scan and analyze a certain range from a predetermined location or a predetermined reference point, thereby evaluating the density, bias, homogeneity, etc. of the data distribution.
  • At least one processor may perform multiple levels of screening, each level being differentiated based on the screening target. Specifically, at least one processor may perform at least one of a first-level screening for screening a data image or a second-level screening for screening a vector set. The first-level screening performs two-dimensional or three-dimensional data operations based on two-dimensional or three-dimensional data images, while the second-level screening performs high-dimensional operations based on high-dimensional vector sets.
  • At least one processor may apply a second metric to measure the characteristics of the data by performing a calculation (high-dimensional calculation) based on at least one vector included in the vector set. For example, at least one processor may diagnose whether the vectors being screened form a structure (clusters, outliers, etc.) on an actual high-dimensional manifold, how high the bias index is, etc.
  • At least one processor may perform multiple levels of screening, each level being differentiated according to the type of characteristic to be screened and diagnosed. Specifically, at least one processor may perform at least one of a first level of screening in which a metric is set for measuring a first type of characteristic (e.g., missing values, data statistics, etc.), a second level of screening in which a metric is set for measuring a second type of characteristic (e.g., density, etc.), or a third level of screening in which a metric is set for measuring a third type of characteristic (e.g., class distribution, etc.).
  • a first level of screening in which a metric is set for measuring a first type of characteristic (e.g., missing values, data statistics, etc.)
  • a second level of screening in which a metric is set for measuring a second type of characteristic (e.g., density, etc.)
  • a third level of screening in which a metric is set for measuring a third type of characteristic (e.g., class distribution,
  • At least one processor can identify at least one vector whose at least one characteristic value satisfies a predetermined condition (S2520). Through this, at least one processor can identify at least one vector whose first characteristic (e.g., density) satisfies a first condition (e.g., overcrowded area, etc.) or at least one vector whose second characteristic (e.g., bias index) satisfies a second condition (e.g., bias is above a standard).
  • first characteristic e.g., density
  • a first condition e.g., overcrowded area, etc.
  • second characteristic e.g., bias index
  • At least one processor can obtain a data image including a plurality of data points representing the data set in two dimensions or three dimensions by processing the vector set using at least one visualization tool (S2530).
  • At least one processor can identify at least one data point corresponding to at least one vector satisfying a predetermined condition, and determine a target area by determining a predetermined area based on the at least one data point. For example, at least one processor can determine an area having a predetermined radius centered on the at least one data point as the target area.
  • At least one processor may generate tag information associated with the determined target area.
  • the tag information may reflect diagnostic results associated with the target area.
  • the at least one processor may generate tag information based on information about events detected by the event detector and characteristic values diagnosed by the screener, but is not limited thereto.
  • the tag information may include, but is not limited to, the discovery context (e.g., which metric conditions were met), time, analyst ID, and event type (e.g., unusual section, blank section, overcrowded section, etc.).
  • At least one processor may compare multiple candidate scenes (e.g., top view, side view, or 45-degree angle view) and provide them to the user, and determine the scene selected by the user as the first scene.
  • the captured snapshot (first scene) may be corrected (brightness, contrast, highlight, etc.) by the correction unit (MODIFIER) described above, and the final scene corrected in this way and the snapshot information including tag information may be stored in a database or file format.
  • MODIFIER correction unit
  • the computing device (2400) can generate snapshot information based on user input received via a GUI-based input interface. Specifically, the computing device (2400) transmits a signal to the capture unit instructing the generation of a snapshot based on the user input, and the capture unit can capture at least a portion of the data image (IOD).
  • IOD data image
  • FIG. 26 is a diagram illustrating a method for a computing device to provide snapshot information based on user input according to various embodiments.
  • the computing device may provide a first data image corresponding to a first data set through a first view port (2710). At this time, the computing device may also provide preview information for at least some of the plurality of data points included in the first data image.
  • a computing device can provide preview information by outputting the actual data corresponding to a data point.
  • the computing device can configure the preview information so that clicking on a specific data point previews the actual data (e.g., original image, text content, statistical values, etc.) corresponding to that point. This allows the user to immediately confirm the meaning of the point in the data image.
  • At least one processor may generate first snapshot information including a first scene for a first data image being provided through a first view port in response to a user input received through the first GUI and provide the first snapshot information through a second view port (S2620). Specifically, the at least one processor may generate the first snapshot information by capturing a first scene including at least one data point on the first data image.
  • the computing device may receive user input for a first GUI (2720) and provide first snapshot information including a first scene for a first data image through a second view port (2730). Specifically, the computing device may display the first GUI (2720) and prompt the user to specify a desired scene using a button or a specific gesture (such as dragging or selecting a box) that instructs the user to take a snapshot.
  • a button or a specific gesture such as dragging or selecting a box
  • the computing device may transmit a corresponding instruction signal to the capture unit.
  • the capture unit may capture the first scene by synthesizing the current viewpoint, magnification ratio, or range designation of the first viewport (2710) where the first data image is displayed.
  • the computing device may provide the completed first snapshot information to the user through the second viewport (2730).
  • the second viewport (2730) may be used as a "snapshot preview" area.
  • the computing device may display a "tag selection menu” (not shown) in a portion of the second viewport (2730) (e.g., a pop-up, a side panel).
  • At least one processor may analyze the captured scene (e.g., data distribution or point properties within the snapshot) and automatically suggest recommended tags such as "bias,” “dense,” “outlier,” and “label mismatch.”
  • the processor may generate and store final tag information (e.g., "dense,” “bottom-right cluster,” “Class A”>) based on the selected tags.
  • At least one processor may generate link information for connecting the first snapshot information to an external communication network based on user input to the second GUI provided through the second view port (S2630).
  • At least one processor can provide a second GUI (2740) through a second viewport (2720). At least one processor can generate link information based on a user input for the second GUI (2740).
  • the link information can include a link for transmitting the first snapshot information through an external communication network.
  • the link information can be a connection path for transmitting and sharing this snapshot through an external communication network (Internet, company intranet, etc.), and the recipient can check the same scene (snapshot) and tag information or memo information, etc. through the link.
  • the user since the user directly determines and captures an arbitrary point/area, it effectively supports customized scenarios (such as enlarging only a specific cluster, emphasizing only a specific label, etc.) that may be missed in automatic capture.
  • user comments or classification information are organically combined and stored with snapshots through view port switching (first->second) and memo/tag interfaces, thereby providing richer context for future analysis collaboration or document reporting.
  • the link generation function enables real-time sharing and feedback of snapshot information (scenes, tags, notes, etc.) with external team members or other systems, thereby increasing the efficiency of data interpretation in a remote collaboration environment and inducing external exposure to the solution, which may have an economic ripple effect.
  • FIG. 28 is a diagram illustrating a function of a computing device to reproduce snapshot information according to various embodiments.
  • a computing device (or at least one processor) according to an embodiment of the present disclosure can sequentially store a plurality of snapshots captured along a screening path and reproduce these snapshots in the form of a video or slide by sequentially providing these snapshots according to a user reproduction instruction.
  • a computing device or at least one processor included in the computing device may sequentially store a plurality of snapshot information acquired along a screening path (S2810).
  • a screener included in the computing device may search a data image (or a high-dimensional vector set) along a specific path (referred to as a "screening path"), and may capture snapshots at each point when a specific characteristic is detected, an event occurs, or at regular intervals. For example, screening may be performed by "inspecting each block (area) on a 2D view while gradually moving from the upper left to the lower right.” After checking the characteristic value at each inspection point, a snapshot may be taken if it exceeds a threshold value.
  • At least one processor can record the captured snapshots in this manner along with connection information (e.g., "Snap #1 -> #2 -> Snap #3 ") according to the acquisition order or the screening path order.
  • connection information e.g., "Snap #1 -> #2 -> Snap #3 "
  • each snapshot can be stored including metadata such as "view point,” shooting location (coordinates along the screening path or high-dimensional mapping information), event cause (excess density, outlier detection, etc.), shooting time,” etc.
  • At least one processor can play back multiple sequentially stored snapshot information in a predetermined manner according to a playback instruction input by the user (S2820). Specifically, when a user (e.g., an analyst) presses "Play" or inputs a request to sequentially check screening records through a specific interface, at least one processor can retrieve multiple sequentially stored snapshot information.
  • a user e.g., an analyst
  • the computing device may be implemented so that the transmission order of multiple snapshot information during playback differs from the actual snapshot capture order. For example, a user may capture snapshots in a random order at the time of capture, but rearrange them according to specific criteria (e.g., issue priority, reverse chronological order, user-specified order, etc.) during playback.
  • specific criteria e.g., issue priority, reverse chronological order, user-specified order, etc.
  • At least one processor may output snapshots using a predetermined playback method, such as a slideshow method (continuous screen switching at fixed intervals), animation (smooth switching between scenes), or timeline operation (progressing according to the timing of an event in each snapshot).
  • a predetermined playback method such as a slideshow method (continuous screen switching at fixed intervals), animation (smooth switching between scenes), or timeline operation (progressing according to the timing of an event in each snapshot).
  • at least one processor may visually display at least one area where a snapshot was captured on the data image and may also display an animated path connecting the points on the screen.
  • the computing device may provide a dynamic navigation experience to the user by creating a visual effect as if a camera were moving sequentially through the areas where snapshots were captured.
  • the computing device can receive user commands related to playback and control playback based on the received commands.
  • at least one processor can perform actions such as "pause,” “skip,” “reverse playback,” and “add annotations to individual snapshots” based on user interaction actions.
  • At least one processor can encourage the user to utilize the improvement feature by providing improvement information (e.g., the reason for the improvement) when an area requiring improvement is identified during snapshot playback.
  • improvement information e.g., the reason for the improvement
  • a computing device may provide a function for linking snapshot information and a diagnostic report.
  • the computing device or at least one processor included in the computing device may, during a data screening process, generate a diagnostic report describing the characteristics of a data set along with snapshot information for a specific region.
  • FIG. 29 is a diagram illustrating a method for a computing device to generate diagnostic reports and snapshot information in conjunction with each other, according to various embodiments.
  • At least one processor can perform a diagnostic report generation operation (S2920). Specifically, at least one processor can generate a diagnostic report by collating or summarizing various diagnostic results. Since the method for generating a diagnostic report has been described above, a detailed description thereof will be omitted.
  • At least one processor may perform a snapshot information generation operation (S2930). Specifically, if an unusual area (such as an overcrowded area or a cluster of outliers) is discovered during the screening process, at least one processor may capture the scene and generate a snapshot.
  • a snapshot information generation operation S2930. Specifically, if an unusual area (such as an overcrowded area or a cluster of outliers) is discovered during the screening process, at least one processor may capture the scene and generate a snapshot.
  • At least one processor may record snapshot information in the form of an abbreviated or summarized version of the diagnostic results (e.g., "20 noise points found, clusters with a bias index of 0.85 or higher").
  • at least one processor may generate snapshot information based on summary information about a specific diagnostic result and a scene captured from a data image corresponding to the diagnostic result.
  • At least one processor can convert portions of the diagnostic results that meet certain criteria (e.g., above a threshold, a class of interest, etc.) into snapshot information.
  • At least one processor can perform a linking operation (S2940) of a diagnostic report and snapshot information. Specifically, at least one processor can link and store the diagnostic result in the diagnostic report and the snapshot information corresponding to the diagnostic result.
  • a linking operation S2940
  • the two pieces of data can be linked by linking unique IDs (e.g., report ID, snapshot ID) or specifying identical metadata (e.g., time, area coordinates, event type).
  • At least one processor can perform a diagnostic report call operation (S2950) from a snapshot. Specifically, when a user clicks a button instructing to call a diagnostic report in a GUI displaying snapshot information or clicks a label (e.g., class imbalance) displayed on a snapshot, the computing device can quickly call up the diagnostic results corresponding to the snapshot based on the stored information. Thereafter, at least one processor can provide the diagnostic report through a new window (viewport) or pop-up window.
  • a diagnostic report call operation S2950
  • the report can be loaded in a new window (or a secondary viewport) and compared side-by-side with existing snapshots.
  • the computing device stores snapshot information generated during the data screening process and a diagnostic report detailing the characteristics of the data set in conjunction with each other, and then cross-references them based on user input, thereby supporting the improvement and utilization of data quality by flexibly moving between intuitive visual information and precise analysis results.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Software Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Data Mining & Analysis (AREA)
  • Computing Systems (AREA)
  • Evolutionary Computation (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Computational Linguistics (AREA)
  • Mathematical Physics (AREA)
  • Biomedical Technology (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Biophysics (AREA)
  • Databases & Information Systems (AREA)
  • Human Computer Interaction (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Evolutionary Biology (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

본 개시의 일 실시예에 따르면, 메모리 및 상기 메모리에 전자적으로 연결되는 적어도 하나의 프로세서를 포함하고, 상기 메모리에 저장된 적어도 하나의 인스트럭션을 실행하도록 구현되는 적어도 하나의 프로세서는, 제1 데이터 셋에 대응되는 제1 데이터 이미지 - 상기 제1 데이터 이미지는 상기 제1 데이터 셋에 포함되는 데이터 각각에 대응되는 복수의 데이터 포인트들을 포함함 - 를 제1 뷰 포트를 통해 제공하는 동작, 제1 GUI를 통해 수신된 사용자 입력에 대응하여, 상기 제1 뷰 포트를 통해 제공 중인 상기 제1 데이터 이미지에 대한 제1 장면을 포함하는 제1 스냅 샷 정보를 생성하여 제2 뷰 포트를 통해 제공하는 동작 및 상기 제2 뷰 포트를 통해 제공되는 제2 GUI에 대한 사용자 입력을 기초로, 상기 제1 스냅 샷 정보를 외부 통신 네트워크와 연결하기 위한 링크 정보를 생성하는 동작을 수행하도록 설정되는 컴퓨팅 장치가 제공될 수 있다.

Description

데이터 시각화 도구의 구현 방법, 장치 및 시스템
본 개시는 데이터를 진단하고 시각화하기 위한 데이터 처리 기술에 관한 것이다. 보다 구체적으로, 데이터의 이미징을 통해 데이터를 진단하고, 데이터를 시각화함으로써 인터랙션 기능을 제공하는 기술에 관한 것이다.
최근 인공지능(AI) 및 머신러닝(ML) 기술이 다양한 산업 분야에서 널리 활용됨에 따라, 대규모 데이터 셋을 효과적으로 처리하고 분석하는 기술의 중요성이 증대되고 있다. 데이터 품질은 인공지능 모델(또는 신경망)의 성능에 직접적인 영향을 미치는 중요한 요소이며, 이를 진단하고 분석하는 과정에서 벡터화(vectorization) 및 시각화 기술이 필수적으로 요구된다. 특히, 딥러닝 기반 모델의 학습을 위해 대량의 이미지, 텍스트 및 구조화된 데이터가 사용됨에 따라, 데이터의 내재적 특성을 정확하게 이해하는 것이 중요하다.
기존의 데이터 품질 진단 방법은 정형 데이터(structured data)에 대한 무결성 검증 및 기초적인 통계 분석에 주로 의존하고 있으며, 비정형 데이터(unstructured data)에 대한 정밀한 분석이 어렵다. 또한, 대규모 데이터 셋의 내재적 분포를 파악하는 기존 기술들은 고차원 공간에서의 데이터 구조를 효과적으로 반영하지 못하며, 데이터 품질을 평가하는 데 있어 시각적인 이해를 제공하는 기능이 미흡하다.
기존의 데이터 시각화 기술 역시 한계가 존재한다. 고차원 데이터의 경우 차원 축소 기법을 활용하여 2차원 또는 3차원 공간에서 표현하더라도 원래 데이터의 특성을 유지하는 것이 어렵다. 또한, 사용자가 데이터 셋의 품질을 분석하고 조작하는 과정에서 직관적인 인터랙션이 부족하여, 데이터 탐색 및 품질 개선 과정이 비효율적으로 이루어지고 있다.
본 개시의 일 과제는 대량의 데이터 셋의 품질을 진단하는 것이다.
또한, 본 개시의 일 과제는 대량의 데이터 셋의 내재적 분포를 정밀하게 식별하기 위한 벡터화(vectorization) 방법을 통해, 데이터의 내재적 구조를 유지하면서도 효과적으로 표현할 수 있도록 데이터 셋을 시각화하는 것이다.
또한, 본 개시의 일 과제는 고차원 벡터 공간과 시각화 공간 간의 간극을 최소화하여 직관적인 사용자 인터랙션을 달성하는 것이다.
본 개시의 일 실시예에 따르면, 메모리 및 상기 메모리에 전자적으로 연결되는 적어도 하나의 프로세서를 포함하고, 상기 메모리에 저장된 적어도 하나의 인스트럭션을 실행하도록 구현되는 적어도 하나의 프로세서는, 제1 데이터 셋을 획득하는 동작, 사용자 입력을 기초로, 데이터 처리 방식에 따라 분류된 복수의 수준들 중 제1 수준을 선택하는 동작, 상기 제1 수준에 대응되는 제1 데이터 처리 모델 -상기 제1 데이터 처리 모델은 적어도 하나의 인공지능 모델을 포함함-에 대한 적어도 하나의 속성을 결정하는 동작, 상기 제1 데이터 처리 모델을 이용하여 상기 제1 데이터 셋을 처리하는 작업에 연관되는 사전 정보를 제1 GUI(Graphic User Interface)를 통해 제공하는 동작 및 상기 제1 GUI에 대한 사용자 입력을 기초로, 상기 제1 데이터 처리 모델을 구축하는 동작을 수행하도록 설정되는 컴퓨팅 장치가 제공될 수 있다.
또한, 본 개시의 일 실시예에 따르면, 메모리에 저장된 적어도 하나의 인스트럭션을 실행하는 적어도 하나의 프로세서에 의해, 제1 데이터 셋을 획득하는 동작, 사용자 입력을 기초로, 데이터 처리 방식에 따라 분류된 복수의 수준들 중 제1 수준을 선택하는 동작, 상기 제1 수준에 대응되는 제1 데이터 처리 모델 -상기 제1 데이터 처리 모델은 적어도 하나의 인공지능 모델을 포함함- 에 대한 적어도 하나의 속성을 결정하는 동작, 상기 제1 데이터 처리 모델을 이용하여 상기 제1 데이터 셋을 처리하는 작업에 연관되는 사전 정보를 제1 GUI(Graphic User Interface)를 통해 제공하는 동작 및 상기 제1 GUI에 대한 사용자 입력을 기초로, 상기 제1 데이터 처리 모델을 구축하는 동작을 수행하도록 설정되는 데이터 처리 방법이 제공될 수 있다.
또한, 본 개시의 일 실시예에 따르면, 적어도 하나의 GUI(Graphic User Interface)를 표시하도록 구현되는 디스플레이, 메모리 및 상기 메모리에 전자적으로 연결되는 적어도 하나의 프로세서를 포함하고, 상기 메모리에 저장된 적어도 하나의 인스트럭션을 실행하도록 구현되는 적어도 하나의 프로세서는, 데이터 처리 방식에 따라 분류된 복수의 수준들 중 제1 수준을 선택하는 사용자 입력을 수신하는 동작, 상기 제1 수준에 대응되는 제1 데이터 처리 모델 -상기 제1 데이터 처리 모델은 적어도 하나의 인공지능 모델을 포함함-을 이용하여 상기 제1 데이터 셋을 처리하는 작업에 연관되는 사전 정보를 나타내는 제1 GUI(Graphic User Interface)를 상기 디스플레이를 통해 표시하는 동작, 상기 제1 데이터 처리 모델로부터 제1 차원의 임베딩 영역에서 정의되는 제1 벡터 셋을 기초로 상기 제1 데이터 셋에 대한 진단 결과를 획득하는 동작 및 상기 제1 벡터 셋을 제1 시각화 도구에 제공하여 2차원 또는 3차원의 제1 데이터 이미지 - 상기 제1 데이터 이미지는 상기 제1 데이터 셋에 대응됨- 를 상기 디스플레이를 통해 표시하는 동작을 수행하도록 설정되는 전자 장치가 제공될 수 있다.
또한, 본 개시의 일 실시예에 따르면, 메모리에 저장된 적어도 하나의 인스트럭션을 실행하는 적어도 하나의 프로세서에 의하여, 데이터 셋을 획득하는 동작, 상기 데이터 셋을 특정 차원으로 임베딩하여 벡터 셋 - 상기 벡터 셋은 상기 데이터 셋에 포함되는 복수의 데이터에 대응되는 복수의 벡터들을 포함함 -을 획득하는 동작, 적어도 하나의 메트릭이 설정된 스크리너를 이용하여, 상기 벡터 셋에 포함되는 벡터를 기초로 적어도 하나의 특성 값을 산출함으로써 상기 벡터 셋을 스크리닝하는 동작, 상기 적어도 하나의 특성 값이 미리 정해진 조건을 만족하는 적어도 하나의 벡터를 식별하는 동작, 적어도 하나의 시각화 도구를 이용하여, 상기 벡터 셋을 처리함으로써 상기 데이터 셋을 2차원 또는 3차원으로 나타내는 복수의 데이터 포인트들을 포함하는 데이터 이미지를 획득하는 동작, 상기 데이터 이미지 상에서 상기 적어도 하나의 벡터에 대응되는 적어도 하나의 데이터 포인트를 포함하는 제1 영역을 결정하여, 상기 제1 영역에 연관되는 태그 정보를 생성하는 동작 및 상기 제1 영역을 포함하는 제1 장면 - 상기 제1 장면은 특정 시점(view-point)로부터 상기 제1 영역을 캡쳐함으로써 생성됨- 및 상기 태그 정보를 포함하는 스냅 샷 정보를 저장하는 동작을 포함하는 방법이 제공될 수 있다.
또한, 본 개시의 일 실시예에 따르면, 상기 컴퓨터 프로그램에 저장된 적어도 하나의 인스트럭션을 실행하는 적어도 하나의 프로세서에 의해, 데이터 셋을 획득하는 동작, 상기 데이터 셋을 특정 차원으로 임베딩하여 벡터 셋 - 상기 벡터 셋은 상기 데이터 셋에 포함되는 복수의 데이터에 대응되는 복수의 벡터들을 포함함 -을 획득하는 동작, 적어도 하나의 메트릭이 설정된 스크리너를 이용하여, 상기 벡터 셋에 포함되는 벡터를 기초로 적어도 하나의 특성 값을 산출함으로써 상기 벡터 셋을 스크리닝하는 동작, 상기 적어도 하나의 특성 값이 미리 정해진 조건을 만족하는 적어도 하나의 벡터를 식별하는 동작, 적어도 하나의 시각화 도구를 이용하여, 상기 벡터 셋을 처리함으로써 상기 데이터 셋을 2차원 또는 3차원으로 나타내는 복수의 데이터 포인트들을 포함하는 데이터 이미지를 획득하는 동작, 상기 데이터 이미지 상에서 상기 적어도 하나의 벡터에 대응되는 적어도 하나의 데이터 포인트를 포함하는 제1 영역을 결정하여, 상기 제1 영역에 연관되는 태그 정보를 생성하는 동작 및 상기 제1 영역을 포함하는 제1 장면 - 상기 제1 장면은 특정 시점(view-point)로부터 상기 제1 영역을 캡쳐함으로써 생성됨- 및 상기 태그 정보를 포함하는 스냅 샷 정보를 저장하는 동작을 수행하도록 설정되는 컴퓨터 프로그램이 제공될 수 있다.
또한, 본 개시의 일 실시예에 따르면, 메모리 및 상기 메모리에 전자적으로 연결되는 적어도 하나의 프로세서를 포함하고, 상기 메모리에 저장된 적어도 하나의 인스트럭션을 실행하도록 구현되는 적어도 하나의 프로세서는, 제1 데이터 셋에 대응되는 제1 데이터 이미지 - 상기 제1 데이터 이미지는 상기 제1 데이터 셋에 포함되는 데이터 각각에 대응되는 복수의 데이터 포인트들을 포함함 - 를 제1 뷰 포트를 통해 제공하는 동작, 제1 GUI를 통해 수신된 사용자 입력에 대응하여, 상기 제1 뷰 포트를 통해 제공 중인 상기 제1 데이터 이미지에 대한 제1 장면을 포함하는 제1 스냅 샷 정보를 생성하여 제2 뷰 포트를 통해 제공하는 동작 및 상기 제2 뷰 포트를 통해 제공되는 제2 GUI에 대한 사용자 입력을 기초로, 상기 제1 스냅 샷 정보를 외부 통신 네트워크와 연결하기 위한 링크 정보를 생성하는 동작을 수행하도록 설정되는 컴퓨팅 장치가 제공될 수 있다.
또한, 본 개시의 일 실시예에 따르면, 메모리에 적어도 하나의 인스트럭션을 실행하는 적어도 하나의 프로세서에 의해, 제1 데이터 셋에 대응되는 제1 데이터 이미지 - 상기 제1 데이터 이미지는 상기 제1 데이터 셋에 포함되는 데이터 각각에 대응되는 복수의 데이터 포인트들을 포함함 - 를 제1 뷰 포트를 통해 제공하는 동작, 제1 GUI를 통해 수신된 사용자 입력에 대응하여, 상기 제1 뷰 포트를 통해 제공 중인 상기 제1 데이터 이미지에 대한 제1 장면을 포함하는 제1 스냅 샷 정보를 생성하여 제2 뷰 포트를 통해 제공하는 동작 및 상기 제2 뷰 포트를 통해 제공되는 제2 GUI에 대한 사용자 입력을 기초로, 상기 제1 스냅 샷 정보를 외부 통신 네트워크와 연결하기 위한 링크 정보를 생성하는 동작을 포함하는 데이터 인터랙션 방법이 제공될 수 있다.
또한, 본 개시의 일 실시예에 따르면, 전자 장치에 있어서, 적어도 하나의 GUI(Graphic User Interface)를 표시하도록 구현되는 디스플레이, 메모리 및 상기 메모리에 전자적으로 연결되는 적어도 하나의 프로세서를 포함하고, 상기 메모리에 저장된 적어도 하나의 인스트럭션을 실행하도록 구현되는 적어도 하나의 프로세서는, 제1 데이터 셋에 대응되는 제1 데이터 이미지 - 상기 제1 데이터 이미지는 상기 제1 데이터 셋에 포함되는 데이터 각각에 대응되는 복수의 데이터 포인트들을 포함함 - 를 제1 뷰 포트를 상기 디스플레이를 통해 제공하는 동작, 제1 GUI를 통해 수신된 사용자 입력에 대응하여, 상기 제1 뷰 포트를 통해 제공 중인 상기 제1 데이터 이미지에 대한 제1 장면을 포함하는 제1 스냅 샷 정보를 제2 뷰 포트를 상기 디스플레이를 통해 제공하는 동작 및 상기 제2 뷰 포트를 통해 제공되는 제2 GUI에 대한 사용자 입력을 기초로, 상기 제1 스냅 샷 정보를 외부 통신 네트워크와 연결하기 위한 링크 정보를 제공하는 동작;을 수행하도록 설정되는 전자 장치가 제공될 수 있다.
또한, 본 개시의 일 실시예에 따르면, 메모리에 저장된 적어도 하나의 인스트럭션을 실행하는 적어도 하나의 프로세서에 의해, 데이터 셋을 시각화하여 상기 데이터 셋에 대응되는 제1 데이터 이미지를 제공하는 동작, 상기 제1 데이터 이미지 상에서의 제1 영역에 대한 사용자 입력을 수신하는 동작, 상기 제1 영역에 대응되는 잠재 코드를 결정하고, 결정된 상기 잠재 코드를 생성 모델에 제공하여, 합성 데이터를 생성하는 동작, 상기 합성 데이터를 데이터 처리 모델 -상기 데이터 처리 모델은 데이터를 특정 차원으로 임베딩하도록 학습됨-에 입력하여, 상기 합성 데이터에 대응되는 합성 벡터를 획득하는 동작 및 상기 합성 벡터를 시각화하여 상기 합성 데이터에 대응되는 합성 포인트를 상기 제1 데이터 이미지 상에 제공하는 동작을 포함하는 방법이 제공될 수 있다.
또한, 본 개시의 일 실시예에 따르면, 적어도 하나의 GUI(Graphic User Interface)를 표시하도록 구현되는 디스플레이, 메모리 및 상기 메모리에 전자적으로 연결되는 적어도 하나의 프로세서를 포함하고, 상기 메모리에 저장된 적어도 하나의 인스트럭션을 실행하도록 구현되는 적어도 하나의 프로세서는, 데이터 셋을 시각화하여 상기 데이터 셋에 대응되는 제1 데이터 이미지를 상기 디스플레이를 통해 제공하는 동작, 상기 제1 데이터 이미지 상의 제1 영역에 대한 사용자 입력을 수신하는 동작, 상기 제1 영역에 대응되는 잠재 코드를 결정하고, 결정된 상기 잠재 코드를 생성 모델에 제공하여, 합성 데이터를 생성하는 동작, 상기 합성 데이터를 데이터 처리 모델 -상기 데이터 처리 모델은 데이터를 특정 차원으로 임베딩하도록 학습됨-에 입력하여, 상기 합성 데이터에 대응되는 합성 벡터를 획득하는 동작 및 상기 합성 벡터를 시각화하여 상기 합성 데이터에 대응되는 합성 포인트를 상기 디스플레이를 통해 상기 제1 데이터 이미지 상에 제공하는 동작을 수행하도록 설정되는 컴퓨팅 장치가 제공될 수 있다.
본 발명의 과제의 해결 수단이 상술한 해결 수단들로 제한되는 것은 아니며, 언급하지 아니한 해결 수단들은 본 명세서 및 첨부된 도면으로부터 본 발명이 속하는 기술 분야에서 통상의 지식을 가진 자에게 명확하게 이해될 수 있을 것이다.
본 개시의 일 실시예에 따르면, 컴퓨팅 장치는 대량의 데이터 셋의 품질을 자동으로 진단하고, 이를 직관적으로 시각화할 수 있다. 이를 통해 데이터 분석가, AI 모델 개발자 또는 연구자 등은 데이터의 내재적 구조를 보다 쉽게 이해하고, 데이터 품질을 향상시키기 위한 전략을 수립할 수 있다.
또한, 본 개시의 일 실시예에 따르면, 사용자는 데이터 시각화 과정에서 보다 효율적으로 데이터를 탐색하고 조작할 수 있으며, 머신러닝 모델의 성능을 개선하는 데 활용할 수 있다.
본 개시의 효과들이 상술한 효과들로 제한되는 것은 아니며, 언급되지 아니한 효과들은 본 명세성 및 첨부된 도면으로부터 본 발명이 속하는 기술 분야에서 통상의 지식을 가진 자에게 명확하게 이해될 수 있을 것이다.
도 1은, 다양한 실시예들에 따른, 컴퓨팅 장치의 구성을 도시한 도면이다.
도 2는, 다양한 실시예들에 따른, 컴퓨팅 장치에 의해 제공되는 데이터 클리닉 서비스에 포함되는 다양한 데이터 처리 방법들을 설명하기 위한 도면이다.
도 3은, 다양한 실시예들에 따른, 데이터 클리닉 서비스를 제공하기 위한 다양한 시스템 및 시스템을 구축하기 위한 인공지능 모델 및 알고리즘들을 도시한 도면이다.
도 4는, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터 이미지를 제공하는 방법을 설명하기 위한 도면이다.
도 5는, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터 셋의 특성을 획득하는 방법을 설명하기 위한 도면이다.
도 6은, 다양한 실시예들에 따른, 다양한 실시예들에 따른, 데이터 렌즈 가공 시스템(Lens Processing System) 및 데이터 이미징 시스템(Imaging System)을 도시한 도면이다.
도 7은, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터를 가시화하고, 사용자 인터랙션 기능을 제공하는 시스템을 도시한 도면이다.
도 8은, 다양한 실시예들에 따른, 데이터 시각화 및 인터랙션 방법을 설명하기 위한 도면이다.
도 9는, 다양한 실시예들에 따른, 이미징 신청 단계에서 수행되는 세부적인 단계들을 설명하기 위한 도면이다.
도 10은, 다양한 실시예들에 따른, 렌즈의 수준 및 시각화 도구에 따른 데이터 이미징 및 진단의 결과를 설명하기 위한 도면이다.
도 11은, 다양한 실시예들에 따른, 컴퓨팅 장치에 포함되는 렌즈 빌더를 설명하기 위한 도면이다.
도 12는, 다양한 실시예들에 따른, 컴퓨팅 장치가 사용자 입력에 따라 데이터 이미징 및 진단을 위한 데이터 처리 모델을 구축하는 방법을 도시한 도면이다.
도 13은, 다양한 실시예들에 따른, 사전 정보에 포함되는 정보를 설명하기 위한 도면이다.
도 14는, 다양한 실시예들에 따른, 컴퓨팅 장치가 제1 데이터 처리 모델을 구축하는 일 예시를 도시한 도면이다.
도 15는, 다양한 실시예들에 따른, 컴퓨팅 장치가 제1 데이터 처리 모델을 구축하는 다른 예시를 도시한 도면이다.
도 16은, 다양한 실시예들에 따른, 컴퓨팅 장치가 제1 데이터 처리 모델을 구축하는 또 다른 예시를 도시한 도면이다.
도 17은, 다양한 실시예들에 따른, 컴퓨팅 장치가 구축된 제1 데이터 처리 모델을 이용하여 데이터 셋을 이미징 및 가시화하는 방법을 설명하기 위한 도면이다.
도 18은, 다양한 실시예들에 따른, 컴퓨팅 장치가 시각화 도구를 추천하는 방법을 설명하기 위한 도면이다.
도 19는, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터 이미징을 기반으로 합성 데이터를 생성하는 방법을 설명하기 위한 도면이다.
도 20은, 다양한 실시예들에 따른, 컴퓨팅 장치가 사용자 입력을 기초로 합성 데이터를 생성하는 방법을 설명하기 위한 도면이다.
도 21은, 다양한 실시예들에 따른, 컴퓨팅 장치가 사용자 입력을 기초로 합성 데이터를 생성하는 방법의 일 예시를 도시한 도면이다.
도 22는, 다양한 실시예들에 따른, 컴퓨팅 장치가 사용자 입력을 기초로 데이터를 제거하는 방법을 설명하기 위한 도면이다.
도 23은, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터 개선 요청에 대응하여 데이터 개선을 수행하고, 이에 대한 시각적인 인터랙션을 제공하는 방법을 설명하기 위한 도면이다.
도 24는, 다양한 실시예들에 따른, 컴퓨팅 장치가 스냅 샷(snapshot) 기능을 포함하는 데이터 처리 방법이 구현된 컴퓨팅 장치의 일 예시를 도시한 도면이다.
도 25는, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터 셋을 스크리닝하여 스냅 샷 정보를 생성하는 방법을 설명하기 위한 도면이다.
도 26은, 다양한 실시예들에 따른, 컴퓨팅 장치가 사용자 입력을 기반으로 스냅샷 정보를 제공하는 방법을 설명하기 위한 도면이다.
도 27은 컴퓨팅 장치에 의해 제공되는 화면의 일 예시이다.
도 28은, 다양한 실시예들에 따른, 컴퓨팅 장치가 스냅 샷 정보를 재생하는 기능을 설명하기 위한 도면이다.
도 29는, 다양한 실시예들에 따른, 컴퓨팅 장치가 진단 레포트와 스냅 샷 정보를 연동하여 생성하는 방법을 도시한 도면이다.
이하, 본 개시의 실시예를 첨부의 도면을 참조하여 상세하게 설명한다. 실시예를 설명함에 있어서 본 개시가 속하는 기술 분야에 익히 알려져 있고 본 개시와 직접적으로 관련이 없는 기술 내용에 대해서는 설명을 생략한다. 이는 불필요한 설명을 생략함으로써 본 개시의 요지를 흐리지 않고 더욱 명확히 전달하기 위함이다.
본 명세서에 기재된 실시예는 본 발명이 속하는 기술 분야에서 통상의 지식을 가진 자에게 본 발명의 사상을 명확히 설명하기 위한 것이므로, 본 발명이 본 명세서에 기재된 실시예에 한정되는 것은 아니며, 본 발명의 범위는 본 발명의 사상을 벗어나지 아니하는 수정예 또는 변형예를 포함하는 것으로 해석되어야 한다.
본 명세서에서 사용되는 용어는 본 발명에서의 기능을 고려하여 가능한 현재 널리 사용되고 있는 일반적인 용어를 선택하였으나 이는 본 발명이 속하는 기술 분야에서 통상의 지식을 가진 자의 의도, 판례 또는 새로운 기술의 출현 등에 따라 달라질 수 있다. 다만, 이와 달리 특정한 용어를 임의의 의미로 정의하여 사용하는 경우에는 그 용어의 의미에 관하여 별도로 기재할 것이다. 따라서 본 명세서에서 사용되는 용어는 단순한 용어의 명칭이 아닌 그 용어가 가진 실질적인 의미와 본 명세서의 전반에 걸친 내용을 토대로 해석되어야 한다.
본 명세서에 첨부된 도면은 본 발명을 용이하게 설명하기 위한 것으로 도면에 도시된 형상은 본 발명의 이해를 돕기 위하여 필요에 따라 과장되어 표시된 것일 수 있으므로 본 발명이 도면에 의해 한정되는 것은 아니다.
본 명세서에서, "A 또는 B", "A 및 B 중 적어도 하나", "A 또는 B 중 적어도 하나", "A, B 또는 C", "A, B 및 C 중 적어도 하나", 및 "A, B, 또는 C 중 적어도 하나"와 같은 문구들 각각은 그 문구들 중 해당하는 문구에 함께 나열된 항목들 중 어느 하나, 또는 그들의 모든 가능한 조합을 포함할 수 있다.
본 명세서에서 본 발명에 관련된 공지의 구성 또는 기능에 대한 구체적인 설명이 본 발명의 요지를 흐릴 수 있다고 판단되는 경우에 이에 관한 자세한 설명은 필요에 따라 생략하기로 한다. 또한, 본 명세서의 설명 과정에서 이용되는 숫자(예를 들어, 제1, 제2 등)는 하나의 구성요소를 다른 구성요소와 구분하기 위한 식별기호에 불과하다.
또한, 이하의 설명에서 사용되는 구성요소에 대한 접미사 "부분" 및 "부"는 명세서 작성의 용이함만이 고려되어 부여되거나 혼용되는 것으로서, 그 자체로 서로 구별되는 의미 또는 역할을 갖는 것은 아니다.
즉, 본 개시의 실시예들은 본 개시가 완전하도록 하고, 본 개시가 속하는 기술분야에서 통상의 지식을 가진 자에게 본 개시의 범주를 알려주기 위해 제공되는 것이며, 본 개시의 발명은 청구항의 범주에 의해 정의될 뿐이다. 명세서 전체에 걸쳐 동일 참조 부호는 동일 구성 요소를 지칭한다.
“제1" 및/또는 "제2" 등의 용어는 다양한 구성 요소들을 설명하는데 사용될 수 있지만, 상기 구성 요소들은 상기 용어들에 의해 한정되어서는 안 된다. 상기 용어들은 하나의 구성 요소를 다른 구성 요소로부터 구별하는 목적으로만, 예컨대 본 개시의 개념에 따른 권리 범위로부터 이탈되지 않은 채, 제1 구성요소는 제2 구성요소로 명명될 수 있고, 유사하게 제2 구성요소는 제1 구성요소로도 명명될 수 있다.
어떤 구성요소가 다른 구성요소에 "연결되어" 있다거나 "접속되어" 있다고 언급된 때에는, 그 다른 구성요소에 직접적으로 연결되어 있거나 또는 접속되어 있을 수도 있지만, 중간에 다른 구성요소가 존재할 수도 있다고 이해되어야 할 것이다. 반면에, 어떤 구성요소가 다른 구성요소에 "직접 연결되어" 있다거나 "직접 접속되어" 있다고 언급된 때에는, 중간에 다른 구성요소가 존재하지 않는 것으로 이해되어야 할 것이다. 구성요소들 간의 관계를 설명하는 다른 표현들, 즉 "~사이에"와 "바로 ~사이에" 또는 "~에 이웃하는"과 "~에 직접 이웃하는" 등도 마찬가지로 해석되어야 한다.
도면에서 처리 흐름도 도면들의 각 블록과 흐름도 도면들의 조합들은 컴퓨터 프로그램 인스트럭션들에 의해 수행될 수 있다. 이들 컴퓨터 프로그램 인스트럭션들은 범용 컴퓨터, 특수용 컴퓨터 또는 기타 프로그램 가능한 데이터 프로세싱 장비의 프로세서에 탑재될 수 있으므로, 컴퓨터 또는 기타 프로그램 가능한 데이터 프로세싱 장비의 프로세서를 통해 수행되는 그 인스트럭션들이 흐름도 블록(들)에서 설명된 기능들을 수행하는 수단을 생성하게 된다. 이들 컴퓨터 프로그램 인스트럭션들은 특정 방식으로 기능을 구현하기 위해 컴퓨터 또는 기타 프로그램 가능한 데이터 프로세싱 장비를 지향할 수 있는 컴퓨터 이용 가능 또는 컴퓨터 판독 가능 메모리에 저장되는 것도 가능하므로, 그 컴퓨터 이용가능 또는 컴퓨터 판독 가능 메모리에 저장된 인스트럭션들은 흐름도 블록(들)에서 설명된 기능을 수행하는 인스트럭션 수단을 내포하는 제조 품목을 생산하는 것도 가능할 수 있다. 컴퓨터 프로그램 인스트럭션들은 컴퓨터 또는 기타 프로그램 가능한 데이터 프로세싱 장비 상에 탑재되는 것도 가능하므로, 컴퓨터 또는 기타 프로그램 가능한 데이터 프로세싱 장비 상에서 일련의 동작 단계들이 수행되어 컴퓨터로 실행되는 프로세스를 생성해서 컴퓨터 또는 기타 프로그램 가능한 데이터 프로세싱 장비를 수행하는 인스트럭션들은 흐름도 블록(들)에서 설명된 기능들을 실행하기 위한 단계들을 제공하는 것도 가능할 수 있다.
또한, 기기로 읽을 수 있는 저장 매체는, 비일시적(non-transitory) 저장 매체의 형태로 제공될 수 있다. 여기서, '비일시적'은 저장 매체가 실재(tangible)하는 장치이고, 신호(signal)(예: 전자기파)를 포함하지 않는다는 것을 의미할 뿐이며, 이 용어는 데이터가 저장 매체에 반영구적으로 저장되는 경우와 임시적으로 저장되는 경우를 구분하지 않는다.
또한, 각 블록은 특정된 논리적 기능(들)을 실행하기 위한 하나 이상의 실행 가능한 인스트럭션들을 포함하는 모듈, 세그먼트 또는 코드의 일부를 나타낼 수 있다. 또, 몇 가지 대체 실행 예들에서는 블록들에서 언급된 기능들이 순서를 벗어나서 발생하는 것도 가능함을 주목해야 한다. 예컨대, 잇달아 도시되어 있는 두 개의 블록들은 사실 실질적으로 동시에 수행되는 것도 가능하고 또는 그 블록들이 때때로 해당하는 기능에 따라 역순으로 수행되는 것도 가능하다. 예를 들어, 모듈, 프로그램 또는 다른 구성요소에 의해 수행되는 동작들은 순차적으로, 병렬적으로, 반복적으로, 또는 휴리스틱하게 실행되거나, 상기 동작들 중 하나 이상이 다른 순서로 실행되거나, 생략되거나, 또는 하나 이상의 다른 동작들이 추가될 수 있다.
본 개시에서 사용되는 '~부(unit)'라는 용어는 소프트웨어 또는 FPGA(Field Programmable Gate Array) 또는 ASIC(Application Specific Integrated Circuit)과 같은 하드웨어 구성요소를 의미한다. '~부'는 특정한 역할들을 수행하지만 소프트웨어 또는 하드웨어에 한정되는 의미는 아니다. '~부'는 어드레싱할 수 있는 저장 매체에 있도록 구성될 수도 있고 하나 또는 그 이상의 프로세서들을 재생시키도록 구성될 수도 있다. 따라서, 일부 실시예에 따르면 '~부'는 소프트웨어 구성요소들, 객체지향 소프트웨어 구성요소들, 클래스 구성요소들 및 태스크 구성요소들과 같은 구성요소들과, 프로세스들, 함수들, 속성들, 프로시저들, 서브루틴들, 프로그램 코드의 세그먼트들, 드라이버들, 펌웨어, 마이크로코드, 회로, 데이터, 데이터베이스, 데이터 구조들, 테이블들, 어레이들, 및 변수들을 포함한다. 구성요소들과 '~부'들 안에서 제공되는 기능은 더 작은 수의 구성요소들 및 '~부'들로 결합되거나 추가적인 구성요소들과 '~부'들로 더 분리될 수 있다. 뿐만 아니라, 구성요소들 및 '~부'들은 디바이스 또는 보안 멀티미디어카드 내의 하나 또는 그 이상의 CPU들을 재생시키도록 구현될 수도 있다. 또한 본 개시의 다양한 실시예에 따르면, '~부'는 하나 이상의 프로세서를 포함할 수 있다.
이하 첨부된 도면을 참조하여 본 개시의 동작 원리를 상세히 설명한다. 하기에서 본 개시를 설명함에 있어 관련된 공지 기능 또는 구성에 대한 구체적인 설명이 본 개시의 요지를 불필요하게 흐릴 수 있다고 판단되는 경우에는 그 상세한 설명을 생략할 것이다. 그리고 후술되는 용어들은 본 개시에서의 기능을 고려하여 정의된 용어들로서 이는 사용자, 운용자의 의도 또는 관례 등에 따라 달라질 수 있다. 그러므로 그 정의는 본 명세서 전반에 걸친 내용을 토대로 내려져야 할 것이다.
도 1은, 다양한 실시예들에 따른, 컴퓨팅 장치의 구성을 도시한 도면이다.
도 1을 참조하면, 일 실시예에 따른 컴퓨팅 장치(예: 서버 또는 클라이언트 디바이스 등의 컴퓨팅 수단을 포함하는 전자 장치, 이하 "컴퓨팅 장치"라 함)(100)는 프로세서(110), 메모리(120), 저장 장치(130), 통신 회로(140) 및 버스(미도시)를 포함할 수 있다. 컴퓨팅 장치(100)의 구성이 도 1에 도시된 구성이나 상술한 구성에 한정되는 것은 아니고, 일반적인 컴퓨팅 장치 또는 모바일 디바이스에 포함되는 하드웨어 또는 소프트웨어 구성을 더 포함할 수 있음은 물론이다.
프로세서(110)는 적어도 일부가 서로 다른 기능을 제공하도록 구현되는 적어도 하나의 프로세서를 포함할 수 있다. 예를 들면, 소프트웨어(예: 프로그램)를 실행하여 프로세서(110)에 연결된 컴퓨팅 장치(100)의 적어도 하나의 다른 구성요소(예: 하드웨어 또는 소프트웨어 구성요소)를 제어할 수 있고, 다양한 데이터 처리 또는 연산을 수행할 수 있다. 일 실시예에 따르면, 데이터 처리 또는 연산의 적어도 일부로서, 프로세서(110)는 다른 구성요소로부터 수신된 명령 또는 데이터를 메모리(120)(예: 휘발성 메모리)에 저장하고, 휘발성 메모리에 저장된 명령 또는 데이터를 처리하고, 결과 데이터를 비휘발성 메모리에 저장할 수 있다. 일 실시예에 따르면, 프로세서(110)는 메인 프로세서(예: 중앙 처리 장치 또는 어플리케이션 프로세서) 또는 이와는 독립적으로 또는 함께 운영 가능한 보조 프로세서(예: 그래픽 처리 장치, 신경망 처리 장치(NPU: neural processing unit), 이미지 시그널 프로세서, 센서 허브 프로세서, 또는 커뮤니케이션 프로세서)를 포함할 수 있다. 예를 들어, 컴퓨팅 장치(100)가 메인 프로세서 및 보조 프로세서를 포함하는 경우, 보조 프로세서는 메인 프로세서보다 저전력을 사용하거나, 지정된 기능에 특화되도록 설정될 수 있다. 보조 프로세서는 메인 프로세서와 별개로, 또는 그 일부로서 구현될 수 있다. 보조 프로세서는, 예를 들면, 메인 프로세서가 인액티브(예: 슬립) 상태에 있는 동안 메인 프로세서를 대신하여, 또는 메인 프로세서가 액티브(예: 어플리케이션 실행) 상태에 있는 동안 메인 프로세서와 함께, 컴퓨팅 장치(100)의 구성요소들 중 적어도 하나의 구성요소(예: 디스플레이(240) 또는 통신 회로)와 관련된 기능 또는 상태들의 적어도 일부를 제어할 수 있다. 일 실시예에 따르면, 보조 프로세서(예: 이미지 시그널 프로세서 또는 커뮤니케이션 프로세서)는 기능적으로 관련 있는 다른 구성요소(예: 통신 회로)의 일부로서 구현될 수 있다. 일 실시예에 따르면, 보조 프로세서(예: 신경망 처리 장치)는 인공지능 모델의 처리에 특화된 하드웨어 구조를 포함할 수 있다. 한편, 이하에서 기술되는 컴퓨팅 장치(100)의 동작은, 프로세서(110)의 동작으로 이해될 수 있다.
다양한 실시예들에 따르면, 메모리(120)는 적어도 일부가 서로 다른 기능을 제공하도록 구현되는 적어도 하나의 메모리를 포함할 수 있다. 메모리(120)는 컴퓨팅 장치(100)의 적어도 하나의 구성요소(예: 프로세서(110))에 의해 사용되는 다양한 데이터를 저장할 수 있다. 데이터는, 예를 들어, 소프트웨어(예: 프로그램) 및, 이와 관련된 명령에 대한 입력 데이터 또는 출력 데이터를 포함할 수 있다. 메모리(120)는, 휘발성 메모리 또는 비휘발성 메모리를 포함할 수 있다. 메모리(120)는 운영 체제, 미들웨어 또는 어플리케이션, 및/또는 전술한 인공지능 모델을 저장하도록 구현될 수 있다.
또한, 메모리(120)는 서비스에 의해 제공되는 기능들을 구현하기 위한 프로세서(110)의 동작들을 지시하는 복수의 지시 사항들(Instructions, 121)을 포함할 수 있다. 이때, 프로세서(110)는 메모리(120)에 저장된 복수의 지시 사항들 중 적어도 일부를 실행할 수 있다. 컴퓨팅 장치(110)는 복수의 지시 사항들 중 적어도 일부를 기초로 서비스에 의해 제공되는 기능들을 실행하는 프로세서(110)를 포함하는 소프트웨어 서버를 포함할 수 있다.
저장 장치(130)는 컴퓨팅 디바이스(100)에 대용량 저장 디바이스를 제공할 수 있다. 저장 장치(130)는 컴퓨터 판독 가능 매체일 수 있다. 예를 들어, 저장 장치(130)는 플로피 디스크 디바이스, 하드 디스크 디바이스, 광학 디스크 디바이스, 테이프 디바이스, 플래시 메모리 또는 기타 유사 솔리드 스테이트 메모리 디바이스, 또는 저장 영역 네트워크나 기타 구성의 디바이스를 포함한 디바이스 어레이일 수 있다. 또한, 컴퓨터 프로그램 제품은 정보 매체에 명백하게 구현된다. 컴퓨터 프로그램 제품에는 실행 시 위에 설명된 것과 같은 하나 이상의 방법을 수행하는 명령들이 포함되어 있다. 정보 매체는 메모리(120), 저장 장치(130), 또는 프로세서(110)의 메모리와 같은 컴퓨터 판독 가능 매체 또는 기계 판독 가능 매체이다.
또한, 저장 장치(130)는 데이터베이스(DB)를 포함할 수 있다. 저장 장치(130)는 미리 구조화된 데이터 구조를 가지는 데이터베이스를 포함할 수 있다. 컴퓨팅 장치(110)는 상호 간에 연관 관계를 가지는 데이터 셋을 데이터베이스에 저장할 수 있다.
본 개시에 따른 컴퓨팅 장치는 적어도 하나의 프로세서 및 적어도 하나의 프로세서에 전자적으로 연결된 메모리에 의해 수행되는 다양한 인공지능 프레임워크들을 기반으로 서비스를 제공할 수 있다.
이와 관련하여, 메모리(120) 또는 저장 장치(130)는 주어진 작업(task)을 수행하도록 학습시킬 수 있는 다양한 유형의 인공지능(또는 머신러닝) 프레임워크가 구현된 적어도 하나의 인공지능 모델을 저장할 수 있다. 예를 들어, 서포트 벡터 머신, 의사 결정 트리, 신경망 등은 이미지 처리 및 자연어 처리와 같은 다양한 애플리케이션에서 사용되는 머신 러닝 프레임워크의 몇몇 예시에 불과하다. 신경망과 같은 일부 인공지능 프레임워크는 특정 연산을 수행하는 노드들의 계층들을 이용할 수 있다.
신경망에서 노드는 하나 이상의 에지(edge)를 통해 서로 연결된다. 신경망은 입력 계층, 출력 계층 및 하나 이상의 중간 계층들을 포함할 수 있다. 개별 노드는 미리 정의된 함수에 따라 각각의 입력을 처리하고 후속 계층 또는 경우에 따라 이전 계층에 출력을 제공할 수 있다. 특정 노드에 대한 입력에는 입력과 노드 사이의 에지에 해당하는 가중치 값을 곱할 수 있다. 또한, 노드는 출력을 생성하는 데 사용되는 개별 바이어스 값을 가질 수 있다. 에지 가중치 및/또는 바이어스 값(파라미터)을 학습하기 위해 다양한 학습 절차를 적용할 수 있다.
신경망 구조는 서로 다른 특정 기능을 수행하는 여러 계층들을 가질 수 있다. 예를 들어, 하나 이상의 노드 레이어는 풀링, 인코딩 또는 컨볼루션 연산과 같은 특정 연산을 집합적으로 수행할 수 있다. 본 개시에서 "계층(layer)"이라는 용어는 외부 소스 또는 네트워크의 다른 레이어와 주고받는 등 입력과 출력을 공유하는 노드 그룹을 의미할 수 있다. "연산(calculation)"이라는 용어는 하나 이상의 노드 레이어에서 수행할 수 있는 기능을 의미할 수 있다. "모델 구조(model structure)"라는 용어는 레이어 수, 레이어의 연결성 및 개별 레이어가 수행하는 작업 유형을 포함하여 계층화된 모델의 전반적인 아키텍처를 의미할 수 있다. "신경망 구조(neural network structure)"라는 용어는 신경망의 모델 구조를 의미할 수 있다. "학습된 모델" 및/또는 "튜닝된 모델(tuned model)"이라는 용어는 학습 또는 튜닝된 모델 구조에 대한 매개변수와 함께 모델 구조를 의미할 수 있다. 예를 들어, 두 모델이 서로 다른 훈련 데이터에 대해 훈련되거나 훈련 프로세스에 기본 확률론적 프로세스가 있는 경우와 같이, 훈련된 두 모델은 동일한 모델 구조를 공유하면서도 매개변수에 대해 서로 다른 값을 가질 수 있다.
"전이 학습"은 특정 작업에 대한 제한된 작업 별 훈련 데이터로 모델을 훈련하는 한 가지 광범위한 접근 방식이다. 전이 학습에서, 모델은 먼저 중요한 훈련 데이터를 사용할 수 있는 다른 작업에 대해 사전 훈련된 다음, 작업별 훈련 데이터를 사용하여 특정 작업에 맞게 조정할 수 있다.
본 개시에서 사용되는 "사전 훈련"이라는 용어는, 하나 이상의 특정 작업에 대해 모델을 조정하기 위해 해당 모델 파라미터의 후속 조정을 허용하는 방식으로 모델 파라미터를 조정하기 위한 사전 훈련 데이터 세트에 대한 모델 훈련을 지칭한다. 경우에 따라 사전 학습에는 레이블이 지정되지 않은 학습 데이터에 대한 자기 지도 학습 프로세스가 포함될 수 있으며, 여기서 '자기 지도' 학습 프로세스는 명시적인(예: 수동으로 제공된) 레이블이 없는 경우 사전 학습 예제의 구조에서 학습하는 것을 포함한다. 사전 학습을 통해 얻은 모델 파라미터의 후속 수정을 여기서는 "튜닝"이라고 한다. 튜닝은 명시적으로 레이블이 지정된 학습 데이터에서 지도 학습을 사용하여 하나 이상의 작업에 대해 수행할 수 있으며, 경우에 따라 사전 학습과 다른 작업을 튜닝에 사용할 수도 있다.
통신 버스(미도시)는 컴퓨팅 장치에 포함되는 복수의 구성들 사이를 전자적으로(또는 통신) 연결하기 위한 구성일 수 있다. 즉, 각각의 구성 요소는 다양한 버스를 사용하여 상호 연결되고, 공통 마더보드에 장착되거나 적절한 다른 방식으로 장착될 수 있다.
입출력 인터페이스(미도시)는 입력 장치에 연결되어 Input 신호를 수신하는 입력 인터페이스 또는 출력 장치에 연결되어 output 신호를 출력하는 출력 인터페이스 등을 포함할 수 있다.
또한, 컴퓨팅 장치(100)는 외부 장치와 통신하기 위한 적어도 하나의 통신 회로(140)를 더 포함할 수 있다.
통신 회로(140)는 컴퓨팅 장치(100)와 외부 컴퓨팅 장치 간의 직접(예: 유선) 통신 채널 또는 무선 통신 채널의 수립, 및 수립된 통신 채널을 통한 통신 수행을 지원할 수 있다. 통신 회로는 프로세서(110)(예: 프로그램 프로세서)와 독립적으로 운영되고, 직접(예: 유선) 통신 또는 무선 통신을 지원하는 하나 이상의 커뮤니케이션 프로세서(예: 통신 칩)를 포함할 수 있다. 일 실시예에 따르면, 통신 회로(140)는 무선 통신 모듈(예: 셀룰러 통신 모듈, 근거리 무선 통신 모듈, 또는 GNSS(global navigation satellite system) 통신 모듈) 또는 유선 통신 모듈(예: LAN(local area network) 통신 모듈, 또는 전력선 통신 모듈)을 포함할 수 있다. 이들 통신 모듈 중 해당하는 통신 모듈은 제1 네트워크(예: 블루투스, WiFi(wireless fidelity) direct 또는 IrDA(infrared data association)와 같은 근거리 통신 네트워크) 또는 제2 네트워크(예: 레거시 셀룰러 네트워크, 5G 네트워크, 차세대 통신 네트워크, 인터넷, 또는 컴퓨터 네트워크(예: LAN 또는 WAN)와 같은 원거리 통신 네트워크)를 통하여 외부의 컴퓨팅 장치와 통신할 수 있다. 이런 여러 종류의 통신 모듈들은 하나의 구성요소(예: 단일 칩)로 통합되거나, 또는 서로 별도의 복수의 구성요소들(예: 복수 칩들)로 구현될 수 있다. 무선 통신 모듈은 가입자 식별 모듈에 저장된 가입자 정보(예: 국제 모바일 가입자 식별자(IMSI))를 이용하여 제1 네트워크 또는 제2 네트워크와 같은 통신 네트워크 내에서 컴퓨팅 장치(100)를 확인 또는 인증할 수 있다. 무선 통신 모듈은 4G 네트워크 이후의 5G 네트워크 및 차세대 통신 기술, 예를 들어, NR 접속 기술(new radio access technology)을 지원할 수 있다. NR 접속 기술은 고용량 데이터의 고속 전송(eMBB(enhanced mobile broadband)), 단말 전력 최소화와 다수 단말의 접속(mMTC(massive machine type communications)), 또는 고신뢰도와 저지연(URLLC(ultra-reliable and low-latency communications))을 지원할 수 있다. 무선 통신 모듈은, 예를 들어, 높은 데이터 전송률 달성을 위해, 고주파 대역(예: mmWave 대역)을 지원할 수 있다. 무선 통신 모듈은 고주파 대역에서의 성능 확보를 위한 다양한 기술들, 예를 들어, 빔포밍(beamforming), 거대 배열 다중 입출력(massive MIMO(multiple-input and multiple-output)), 전차원 다중입출력(FD-MIMO: full dimensional MIMO), 어레이 안테나(array antenna), 아날로그 빔형성(analog beam-forming), 또는 대규모 안테나(large scale antenna)와 같은 기술들을 지원할 수 있다. 무선 통신 모듈은 컴퓨팅 장치(100), 내시경 장치 또는 네트워크 시스템에 규정되는 다양한 요구사항을 지원할 수 있다. 일 실시예에 따르면, 무선 통신 모듈은 eMBB 실현을 위한 Peak data rate(예: 20Gbps 이상), mMTC 실현을 위한 손실 Coverage(예: 164dB 이하), 또는 URLLC 실현을 위한 U-plane latency(예: 다운링크(DL) 및 업링크(UL) 각각 0.5ms 이하, 또는 라운드 트립 1ms 이하)를 지원할 수 있다.
컴퓨팅 장치(100)는 상술한 구성들(프로세서, 통신 회로, 메모리, 디스플레이) 중 적어도 일부만 포함하도록 구현될 수 있다. 예를 들어, 사용자 디바이스는 프로세서, 통신 회로, 메모리, 센서 및 디스플레이를 포함하도록 구현될 수 있다. 또한, 예를 들어, 서버 장치는 프로세서, 통신 회로 및 메모리를 포함하도록 구현될 수 있다.
도 2는, 다양한 실시예들에 따른, 컴퓨팅 장치에 의해 제공되는 데이터 클리닉 서비스에 포함되는 다양한 데이터 처리 방법들을 설명하기 위한 도면이다.
도 2룰 참조하면, 데이터 클리닉 서비스는 다양한 데이터 처리 방법들을 포함할 수 있다. 이때, 상기 다양한 데이터 처리 방법들은 코드화되어 상기 컴퓨팅 장치의 메모리에 저장되어 있을 수 있고, 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 코드화된 적어도 하나의 인스트럭션을 실행하도록 설정될 수 있다. 구체적으로, 적어도 하나의 프로세서는 입력받은 인풋 데이터 셋을 상기 다양한 데이터 처리 방법들을 기초로 처리하여 아웃풋 데이터 셋을 출력할 수 있다.
예를 들어, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치는 데이터 이미징을 위한 동작 방법, 데이터 개선을 위한 동작 방법, 데이터 생성을 위한 동작 방법, 데이터 특성 추출을 위한 동작 방법, 또는 데이터 평가를 위한 동작 방법을 수행할 수 있으나, 이에 한정되지 않는다.
또한, 상술한 각각의 동작 방법들은 상기 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서의 동작 알고리즘들을 기초로 수행될 수 있다.
예를 들어, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치는 데이터 이미징 알고리즘, 데이터 개선 알고리즘, 데이터 생성 알고리즘, 데이터 특성 추출 알고리즘, 또는 데이터 평가 알고리즘 등을 수행할 수 있으나, 이에 한정되지 않는다.
이때, 각각의 동작 방법 및 알고리즘의 명칭은 설명의 편의를 위해 출력되는 결과에 따라 임의로 명명한 것이므로, 각각의 동작 방법 또는 알고리즘은 프로세서에 의해 수행되는 동작들을 기초로 정의될 뿐 동작 방법 또는 알고리즘의 명칭 자체로서 발명을 한정하는 것은 아니다.
보다 구체적으로, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치는, 데이터 이미징 알고리즘에 따라, 인풋 데이터 셋을 처리하여 상기 인풋 데이터 셋에 대한 이미지를 생성할 수 있다.
또한, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치는, 데이터 개선 알고리즘에 따라 인풋 데이터 셋을 처리하여 데이터를 개선할 수 있고, 상기 개선의 결과를 생성할 수 있다.
또한, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치는, 데이터 생성 알고리즘에 따라, 인풋 데이터 셋을 처리하여 가상데이터(synthetic data)를 생성할 수 있다.
또한, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치는, 데이터 특성 추출 알고리즘에 따라, 인풋 데이터 셋을 처리하여 상기 인풋 데이터 셋의 특성(property)을 추출할 수 있다.
또한, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치는, 데이터 평가 알고리즘에 따라, 인풋 데이터 셋을 처리하여 상기 인풋 데이터 셋의 품질을 평가할 수 있다.
상술한 각각의 알고리즘에 대한 상세한 내용은 아래에서 설명한다.
또한, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치는 상술한 다양한 동작 방법들 또는 알고리즘들을 병렬적으로, 연속적으로, 또는 선택적으로 수행할 수 있다. 구체적으로, 상기 컴퓨팅 장치는 동일한 인풋 데이터를 서로 상이한 알고리즘의 입력 값으로 병렬적으로 이용할 수도 있고, 특정 알고리즘에 따라 출력된 결과 값을 다른 알고리즘의 입력 값으로 연속적으로 이용할 수도 있고, 미리 정해진 방식에 따라 복수의 알고리즘들 중 일부의 알고리즘을 선택적으로 수행할 수도 있다.
또한, 상술한 데이터 클리닉을 위한 다양한 동작 방법 또는 알고리즘은 본 개시의 다양한 실시예에 따른 컴퓨팅 장치에 포함되는 딥러닝 모델에서 수행될 수 있다. 구체적으로, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치는 상술한 다양한 동작 방법 또는 알고리즘을 수행하기 위한 하나의 딥러닝 모델을 포함할 수 있으나, 이에 한정되지 않고, 상술한 각각의 동작 방법 또는 알고리즘을 수행하기 위한 다수의 딥러닝 모델들을 포함할 수도 있고, 상술한 다양한 동작 방법 또는 알고리즘들 중 적어도 일부를 수행하기 위한 하나 이상의 딥러닝 모델들을 포함할 수도 있다.
도 3은, 다양한 실시예들에 따른, 데이터 클리닉 서비스를 제공하기 위한 다양한 시스템 및 시스템을 구축하기 위한 인공지능 모델 및 알고리즘들을 도시한 도면이다. 여기서, 시스템은 특정 기능을 수행하기 위해 적어도 하나의 소프트웨어적 구성 또는 하드웨어적 구성을 포함하는 시스템을 의미할 수 있다.
도 3을 참조하면, 본 개시에 따른 컴퓨팅 장치는 클리닉 서비스를 제공하기 위해, 다양한 인공지능(또는 신경망, 머신러닝 등) 모델들로 구성된 데이터 클리닉 시스템을 포함할 수 있다.
예를 들어, 컴퓨팅 장치는 데이터 이미징 시스템, 데이터 진단 시스템 및 데이터 치료 시스템 등을 포함할 수 있으나, 이에 한정되지 않는다.
여기서, 데이터 이미징 시스템은 데이터의 특성을 나타내기 위한 최적의 차원을 결정하기 위한 렌즈 처리 모델, 데이터의 내재적 특성을 반영하는 데이터 이미지를 획득하기 위한 이미징 모델 또는 데이터를 시각적으로 나타내기 위한 시각화 모델 등을 포함할 수 있으나, 이에 한정되지 않는다.
또한, 데이터 진단 시스템은 데이터의 적어도 하나의 특성을 진단하기 위한 진단 모델 또는 데이터의 품질을 평가하기 위한 품질 평가 모델 등을 포함할 수 있으나, 이에 한정되지 않는다.
또한, 데이터 치료 시스템은 필요에 따라 타겟팅된 가상 데이터(또는 합성 데이터)를 생성하기 위한 합성 모델(또는 생성 모델), 데이터의 적어도 일부를 제거하기 위한 데이터 다이어트 모델, 또는 데이터의 적어도 일부의 특성을 조정하기 위한 데이터 보정 모델 등을 포함할 수 있으나, 이에 한정되지 않는다.
컴퓨팅 장치가 포함하는 다양한 머신 러닝 모델들은 메모리에 저장된 복수의 모듈들로 구성될 수 있다. 본 개시에서 모듈(Module)은 특정 기능을 수행하는 인공지능 모델을 구현하기 위한 복수의 하드웨어적 구성들을 포함할 수 있다. 예를 들어, 모듈은 인코더, 디코더, 생성기, 구별기(Discriminator), 어댑터, 자연어 처리 모듈, 또는 거대 언어 모델(LLM) 등을 포함할 수 있으나, 이에 한정되지 않는다.
컴퓨팅 장치는 상술한 복수의 모듈들을 저장할 수 있고, 복수의 모듈들 중 적어도 일부를 기초로 인공지능 프레임워크를 구성하여 데이터 클리닉을 위한 인공지능 모델을 획득할 수 있다. 예를 들어, 데이터 이미징 시스템에 포함되는 데이터 렌즈(Data lens)는 적어도 하나의 인코더 또는 적어도 하나의 어댑터를 포함하는 인공지능 모델로 구현될 수 있으나, 이에 한정되지 않는다.
도 4는, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터 이미지를 제공하는 방법을 설명하기 위한 도면이다.
도 4를 참조하면, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치는 데이터 셋을 입력 받아 데이터 이미지(Image of Data, IOD)를 제공할 수 있다.
이때, 상기 데이터 셋은 M(M>0)차원의 데이터일 수 있다. 다시 말해, 상기 데이터 셋은 M차원의 인풋 스페이스(310)상에서 정의되는 데이터 셋일 수 있다.
또한, 상기 데이터 셋은 단일 모달리티(modality)의 데이터 셋일 수 있다. 예를 들어, 상기 데이터 셋은 이미지 데이터 셋일 수 있다. 또한, 상기 데이터 셋은 텍스트 데이터 셋일 수 있다. 또한, 이에 한정되지 않고, 상기 데이터 셋은 모달리티가 서로 상이한 데이터의 집합일 수 있다. 예를 들어, 상기 데이터 셋은 주석(annotation) 정보를 포함하는 이미지 데이터 셋일 수 있다. 또한, 상기 데이터 셋은 이미지 및 텍스트의 혼합 데이터 셋일 수 있다.
본 개시의 다양한 실시예에 따른 컴퓨팅 장치는 상술한 이미지 데이터 및 텍스트 데이터뿐만 아니라 시계열 데이터 셋, 센서 데이터 셋 등 딥러닝 학습에 이용될 수 있는 모든 모달리티의 데이터를 인풋 데이터 셋으로 입력 받아 처리할 수 있다.
본 개시의 다양한 실시예에 따른 컴퓨팅 장치가 제공하는 데이터 이미지(IOD)는 입력받은 데이터 셋을 처리하여 이미징 공간(320)에 나타낸 것일 수 있다. 여기서 이미지는 2D 이미지를 의미하는 것이 아닌, 데이터를 시각적으로 표현한 것을 통칭하는 표현이다. 구체적으로, 상기 이미징 공간(320)은 2D 공간, 3D 공간, N차원의 가상의 공간을 모두 포함하는 개념으로, 실시예에 따라 제공된 데이터 이미지가 나타나는 공간을 의미한다. 예를 들어, 컴퓨팅 장치가 입력 받은 데이터 셋을 처리하여 상기 데이터 이미지를 PDF 형태로 출력하는 경우, 2D 또는 3D 이미징 공간에 상기 데이터 이미지를 나타낸 아웃풋을 출력할 수 있으나, 이에 한정되지 않는다.
본 개시의 다양한 실시예에 따른 컴퓨팅 장치가 출력 장치(미도시)를 포함하는 경우, 상기 컴퓨팅 장치는 데이터 이미지를 상기 출력 장치를 통해 제공할 수 있다. 예를 들어, 컴퓨팅 장치는 디스플레이를 통해 데이터 이미지를 출력함으로써 데이터 이미지를 제공할 수 있다. 이 경우, 이미징 공간(320)은 스플레이의 화면일 수 있다. 또한, 예를 들어, 컴퓨팅 장치는 프린팅 장치를 통해 데이터 이미지를 출력함으로써 데이터 이미지를 제공할 수 있다. 이 경우, 이미징 공간(320)은 프린팅 장치에 의해 출력되는 용지(paper)일 수 있다.
또한, 본 개시의 다양한 실시예에 따른 컴퓨팅 장치가 통신부를 통해 외부 디바이스와 통신하는 경우, 컴퓨팅 장치는 상기 외부 디바이스를 통해 데이터 이미지를 제공할 수 있다. 이 경우, 이미징 공간(320)은 외부 디바이스의 디스플레이 화면일 수 있다. 예를 들어, 컴퓨팅 장치가 서버 장치인 경우, 상기 서버 장치는 서버 장치에 연결된 네트워크를 통해 서버 장치와 통신하는 적어도 하나의 외부 디바이스로 데이터 이미지를 송신함으로써 데이터 이미지를 제공할 수 있다.
본 개시의 다양한 실시예에 따른 컴퓨팅 장치는 입력된 데이터 셋에 대응되는 벡터 셋(또는 데이터 포인트 셋, 포인트 데이터 셋 등)(330)을 기초로 데이터 이미지를 획득할 수 있다.
이때, 컴퓨팅 장치는 입력받은 데이터 셋에 포함되는 데이터를 특정 차원의 임베딩(embedding) 공간(또는 잠재(latent) 공간)으로 매핑함으로써 벡터 셋을 획득할 수 있다. 구체적으로, 컴퓨팅 장치는 데이터 셋이 특정 차원의 임베딩(embedding) 공간에서 형성하는 매니폴드(manifold)를 확인함으로써 벡터 셋을 획득할 수 있다. 여기서, 매니폴드(manifold)는 입력받은 인풋 데이터 셋이 특정 차원의 임베딩 공간에서 나타내는 형상을 의미할 수 있다. 다시 말해, 매니폴드는 입력받은 인풋 데이터 셋을 특정 차원의 임베딩 공간상의 벡터 셋으로 매핑하는 경우, 벡터 셋이 확인되는 영역 또는 벡터 셋이 형성하는 형상을 의미할 수 있다.
데이터 이미지(IOD)는 데이터 셋에 포함되는 각각의 데이터를 포인트로 시각화한 데이터 셋일 수 있다. 이 경우, 시각화되는 포인트의 형상 또는 색상 등은 실시예에 따라 다양한게 선택될 수 있으므로 '포인트'라는 용어 자체로 발명을 한정하려는 것은 아니다. 또한, 포인트는 실시예에 따라 다양한 용어로 표현될 수 있다. 예를 들어, 포인트는 임베딩 공간 또는 잠재 공간에 나타나는 벡터(vector) 또는 특징(feature) 등의 용어로 표현될 수 있으나, 이에 한정되지 않는다.
본 개시의 다양한 실시예에 따른 컴퓨팅 장치가 데이터 이미지를 제공하기 위해서는, 상술한 바와 같이, 입력받은 데이터 셋에 대응되는 벡터 셋을 식별할 필요가 있다.
컴퓨팅 장치는 매핑 함수로 정의된 미리 정해진 조건(예를 들어, 특정 차원의 임베딩 공간에 매핑하기 위해 미리 저장된 행렬(matrix))을 기초로 데이터 셋을 N차원의 임베딩 공간에 매핑함으로써 벡터 셋을 획득할 수 있다. 예를 들어, 컴퓨팅 장치는 데이터 셋을 인코딩함으로써 벡터 셋을 획득할 수 있으나, 이에 한정되지 않는다. 예를 들어, 컴퓨팅 장치는 데이터 셋을 미리 훈련된 인코더에 입력하고, 인코더의 출력 레이어를 통해, 벡터 셋을 획득할 수 있으나, 이에 한정되지 않는다.
여기서 임베딩이란, 고차원의 데이터를 저차원의 벡터로 변환하여 데이터 포인트 간의 유사성과 구조적 관계를 유지하는 과정을 의미한다. 이러한 임베딩은 데이터의 핵심 정보를 보존하면서도 연산 효율성을 높일 수 있는 방식으로 수행된다. 임베딩 과정에서 각 데이터 포인트는 벡터로 표현되며, 이 벡터들은 데이터 셋 전체의 분포와 특성을 반영할 수 있다. 이를 통해 데이터 셋의 통계적 특성 및 내재된 패턴을 분석할 수 있는 기초를 마련할 수 있는 것이다.
컴퓨팅 장치는 데이터 셋을 임베딩하고 시각화함으로써 데이터 이미지(IOD)를 획득하기 위한 데이터 렌즈(400)를 포함할 수 있다. 이때, 데이터 렌즈(400)는 데이터를 처리하기 위한 적어도 하나의 처리 구성을 포함할 수 있다. 구체적으로, 데이터 렌즈(400)는 데이터 셋을 기초로 벡터 셋을 획득하기 위한 적어도 하나의 신경망 모델(예: 인코더 등) 및 벡터 셋을 기초로 데이터 셋을 시각화하여 데이터 이미지를 획득하기 위한 적어도 하나의 시각화 모델(예: PCA, T-SNE, 또는 UMAP 등)을 포함할 수 있다. 구체적으로, 데이터 렌즈(400)는 데이터 셋을 N차원의 잠재 공간에 임베딩함으로써 데이터 셋에 대응되는 벡터 셋을 획득할 수 있고, 벡터 셋을 M차원(예: 2차원 또는 3차원)의 이미징 공간(320)에 나타냄으로써 데이터 셋에 대응되는 데이터 이미지(IOD)를 획득할 수 있다.
도 5는, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터 셋의 특성을 획득하는 방법을 설명하기 위한 도면이다.
도 5를 참조하면, 컴퓨팅 장치는 획득한 데이터 셋을 처리하여 데이터 셋에 대응되는 특성 정보를 획득할 수 있다.
데이터 셋 또는 데이터의 특성(property)은 데이터 셋 또는 데이터의 분포(예: 기하학적 분포 또는 통계적 분포 등)와 연관되는 정보를 포함할 수 있다. 구체적으로, 특성 정보는 데이터 셋에 포함되는 데이터 각각의 특성 값을 포함할 수 있다. 예를 들어, 컴퓨팅 장치는 데이터 셋에 포함되는 데이터의 특성 값들의 분포를 나타내는 특성 정보를 획득할 수 있다. 또한, 컴퓨팅 장치는 데이터들 각각의 특성 값들의 평균, 편차, 또는 분산 등의 통계적 분포를 기초로 데이터 셋의 특성 정보를 획득할 수 있다.
일 예로, 데이터 셋 또는 데이터의 특성(property)은 데이터 셋 자체의 분포에 연관되는 내재적(intrinsic) 특성을 포함할 수 있다. 예를 들어, 데이터 셋 또는 데이터의 특성은 데이터 셋 또는 데이터의 밀도, 균질도, 편향성 또는 분포 등을 포함할 수 있으나, 이에 한정되지 않는다.
다른 예로, 데이터 셋 또는 데이터의 특성(property)은 데이터 셋이 활용되는 태스크(예를 들어, Classification)에 관련되는 태스크에 의존된 특성(task-dependent property)을 포함할 수 있다. 예를 들어, 데이터 셋 또는 데이터의 특성은 라벨링 오류(labeling error) 비율 또는 클래스가 상이하면서 기하학적으로 인접한 데이터 쌍(하드-네거티브(Hard-negative))의 비율 등을 포함할 수 있으나, 이에 한정되지 않는다.
또한, 컴퓨팅 장치는 데이터 셋 또는 데이터의 특성들 각각에 대응되는 연산 메트릭들을 메모리에 저장하고 있을 수 있다. 보다 구체적으로, 컴퓨팅 장치는 데이터 셋 또는 데이터의 밀도를 연산하기 위한 메트릭, 데이터 셋 또는 데이터의 균질도를 연산하기 위한 메트릭, 데이터 셋 또는 데이터의 편향을 연산하기 위한 메트릭 또는 데이터 셋 또는 데이터의 분포를 연산하기 위한 메트릭 등을 저장하고 있을 수 있으나, 이에 한정되지 않는다.
또한, 컴퓨팅 장치는 인공 신경망으로 구축된 데이터 특성 추출 알고리즘에 따라 저장된 연산 메트릭을 기초로 데이터 셋의 특성을 획득할 수 있다. 구체적으로, 특성 추출 알고리즘은 피드 포워드(feed-forward) 신경망으로 구현될 수 있다.
예를 들어, 컴퓨팅 장치는 데이터 셋의 특성을 연산하기 위한 별도의 인경 신경망을 포함하거나, 데이터 셋의 특성을 연산하기 위한 레이어(layer)를 포함하는 인경 신경망을 포함할 수 있으나, 이에 한정되지 않는다.
일 예로, 컴퓨팅 장치는 데이터 셋을 입력받으면 데이터 셋의 특성을 추출하도록 설계된 특성 추출을 위한 인공 신경망을 포함할 수 있다. 이때, 특성 추출을 위한 인공 신경망은 데이터의 특성을 연산하도록 전이 학습된 인공신경망일 수 있다.
다른 예로, 컴퓨팅 장치는 데이터 셋을 기초로 데이터 이미지를 제공하기 위한 신경망 모델(예: 데이터 렌즈, 인코더 등)에 데이터 특성 추출을 위한 레이어(layer)를 추가한 인공 신경망을 구축함으로써 데이터 셋의 특성을 획득할 수 있다. 구체적으로, 컴퓨팅 장치는 데이터 셋을 기초로 벡터 셋을 식별하고, 식별한 벡터 셋을 기초로 데이터 셋 또는 데이터의 특성을 획득할 수 있다.
이때, 컴퓨팅 장치는 벡터 셋에 포함되는 각각의 벡터들을 미리 정해진 알고리즘으로 처리함으로써 데이터 셋에 포함되는 각각의 데이터의 특성 값을 획득할 수 있다. 이 경우, 컴퓨팅 장치는 벡터 셋에 포함되는 각각의 벡터들의 기하학적 분포 또는 통계적 분포를 기초로 특성 값(value)을 산출할 수 있고, 산출된 특성 값을 대응되는 데이터에 할당할 수 있다. 이때, 특성 값은 벡터들 사이의 거리를 기초로 계산될 수 있다. 예를 들어, 특성 값은 특정 벡터(또는 데이터 포인트)를 중심으로 미리 정해진 거리 이내에 존재하는 벡터의 수를 기초로 획득될 수 있나, 이에 한정되지 않는다. 또한, 예를 들어, 특성 값은 특정 벡터로부터 가까운 미리 정해진 개수의 벡터들까지의 거리의 평균 값을 기초로 획득될 수 있으나, 이에 한정되지 않는다. 예를 들어, 컴퓨팅 장치는 특정 벡터로부터 근접한 K개의 벡터들까지의 거리 값들을 기초로 평균 거리 값을 계산할 수 있고, 계산된 평균 거리 값을 기초로 특정 벡터의 제1 특성 값(예: 밀도 등)을 획득할 수 있으나, 이에 한정되지 않는다.
인공 지능 모델의 프레임워크를 최적화하고, 높은 정확도의 결과 값을 도출해내기 위해서, 모델의 학습에 활용되는 데이터의 품질(quality)이 매우 중요하다.
데이터의 품질은 전술한 바와 같이, 양적 품질 및 질적 품질을 모두 포함하는 개념이므로, 인공지능 모델의 성공적인 학습을 위해 (i) 인공 지능 모델을 학습시키기에 충분한 양의 학습 데이터의 확보 및 (ii) 고품질의 내재적 특성(예: 편향 없는 분포)을 가지는 학습 데이터의 확보 및 (iii) 학습 목적(예: 인공지능 모델의 태스크)에 적합한 특성(예: task-dependent property)을 가지는 학습 데이터의 확보가 필요하다.
본 개시의 일 실시예에 따른 컴퓨팅 장치는, 고품질의 학습 데이터를 획득하기 위해, 데이터 셋의 내재적 특성 및 태스크에 의존된 특성을 향상시키는 방향으로 데이터를 합성(synthesis)하거나 수정(또는 조정)하거나 제거할 수 있다.
이와 더불어, 본 개시에 따른 컴퓨팅 장치는, 데이터 셋에서 적어도 일부의 데이터를 제거함으로써 데이터 셋의 전체적인 품질을 향상시킬 수 있다.
일반적으로, 데이터의 다운 샘플링(down-sampling) 또는 언더 샘플링(under-sampling) 기법은 데이터의 불균형 문제를 해결하기 위해 이용된다. 다만, 기존 언더 샘플링 방법에 따르면, 머신 러닝 데이터 셋의 특성을 고려하지 않고 데이터를 제거함에 따라, 머신 러닝 모델의 학습에 악영향을 끼치는 문제가 있다.
본 개시에 따른 컴퓨팅 장치는, 데이터 셋에서 적어도 일부 데이터를 적절히 제거함으로써 학습되는 인공 지능 모델의 학습 효율을 향상시킬 수 있다.
도 6은, 다양한 실시예들에 따른, 데이터 렌즈 가공 시스템(Lens Processing System) 및 데이터 이미징 시스템(Imaging System)을 도시한 도면이다.
도 6의 (a)를 참조하면, 컴퓨팅 장치는 데이터 셋을 기초로 데이터 렌즈 시스템(System of Lens)을 획득할 수 있다. 이때, 데이터 렌즈 시스템은 데이터 셋을 특정 임베딩 공간에 매핑하기 위한 적어도 하나의 구성을 의미하는 용어일 수 있다. 예를 들어, 데이터 렌즈 시스템은 데이터 셋을 특정 차원의 임베딩 공간(또는 잠재 공간(latent space))에 매핑하기 위한 적어도 하나의 인코더(encoder) 및/또는 파라미터를 조정하기 위한 적어도 하나의 어댑터(adapter)를 포함할 수 있다. 또한, 예를 들어, 데이터 렌즈 시스템은 데이터 셋에 대응되는 잠재 변수(latent variable 또는 잠재 특징 벡터)를 식별하기 위한 적어도 하나의 노드로 구성된 신경망 레이어를 포함할 수 있다.
컴퓨팅 장치(3500)는 데이터 셋을 기초로 데이터 셋의 내재적 특성을 보존하도록 데이터 셋을 처리하는 데이터 렌즈 시스템을 결정할 수 있다.
일 예로, 컴퓨팅 장치는 데이터베이스를 기초로 데이터 셋에 대응되는 렌즈 시스템을 획득할 수 있다. 구체적으로, 컴퓨팅 장치는 입력된 데이터 셋의 특성을 기초로 데이터베이스에서 상기 데이터 셋에 대응되는 렌즈 시스템을 검색할 수 있다.
다른 예로, 컴퓨팅 장치는 렌즈 가공 알고리즘을 기초로 데이터 셋에 대응되는 렌즈 시스템을 획득할 수 있다. 구체적으로, 컴퓨팅 장치는 입력된 데이터 셋의 내재적 특성을 보존하는 최적의 차원(dimensionality)을 연산할 수 있다.
도 6의 (b)를 참조하면, 컴퓨팅 장치는 결정된 데이터 렌즈 시스템을 포함하는 데이터 이미징 시스템(3510)을 기초로 데이터 셋을 처리하여 데이터 이미지(Image of Data)를 획득할 수 있다. 이때, 이미징 시스템(3510)은 적어도 하나의 모듈(예: 인코더, 어댑터 등)로 구성된 렌즈 시스템을 포함할 수 있다.
이 경우, 컴퓨팅 장치는 이미징 시스템(3510)을 이용하여 데이터 셋의 내재적 특성을 나타내는 데이터 이미지를 획득할 수 있다.
본 개시는 대량의 데이터 셋을 효과적으로 분석하고 시각적으로 표현하기 위해, 데이터의 내재적 분포를 정밀하게 식별하도록 벡터화(vectorization)하고, 이를 2차원 또는 3차원 공간에서 직관적으로 시각화할 수 있는 방법을 제공한다. 특히, 고차원 벡터 공간과 시각화 공간 간의 간극을 최소화할 수 있도록 인터랙션 및 UI/UX 구현 기술을 적용하여, 사용자가 데이터 품질을 보다 쉽게 이해하고 조작할 수 있도록 한다.
또한, 본 개시의 방법은 대량의 데이터 셋을 입력으로 받아, 해당 데이터 셋의 내재적 특징(예: 기하학적 분포 또는 데이터 사이의 연관 관계)을 유지하면서도 보다 명확한 시각적 표현을 제공할 수 있다. 이를 통해, 사용자는 데이터의 분포를 직관적으로 파악할 수 있으며, 인공지능의 학습 또는 빅데이터 분석 과정에서 데이터 품질을 개선하는 데 효과적으로 활용할 수 있다.
도 7은, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터를 가시화하고, 사용자 인터랙션 기능을 제공하는 시스템을 도시한 도면이다.
도 7을 참조하면, 컴퓨팅 장치(700)는 데이터 셋(710)을 이미징하여, 데이터 셋(710)에 대응되는 벡터 셋을 획득할 수 있다. 이때, 컴퓨팅 장치(700)는 벡터 셋을 기초로 데이터 이미지(720)를 획득할 수 있다. 데이터 이미지(720)는 데이터 셋(710)을 2차원 또는 3차원으로 가시화한 데이터일 수 있다. 컴퓨팅 장치(100)는 데이터 셋(710)을 기초로 N(N>3)차원의 벡터 셋을 식별하고, N차원의 벡터 셋을 2차원 또는 3차원으로 차원 축소함으로써 데이터 이미지(720)를 획득할 수 있다. 즉, 컴퓨팅 장치(700)는 데이터 셋(700)에 대응되는 데이터 이미지(720)를 2차원 이미징 공간 또는 3차원 이미징 공간에 표현할 수 있다.
컴퓨팅 장치(700)는 네트워크 환경(예: 웹(WEB) 또는 앱(APP) 환경 등)을 통해 데이터 이미지(720)를 제공할 수 있다. 사용자 장치(701)는 네트워크 환경을 통해 데이터 이미지(720)를 확인할 수 있고, 확인된 데이터 이미지(720)에 대한 정보를 획득할 수 있다. 이때, 사용자 장치(701)는 노트북 컴퓨터, 스마트폰, 태블릿 등과 같은 전자 장치일 수 있다. 또한, 예를 들어, 사용자 장치(701)는 스마트 워치, 스마트 글래스, 스마트 안경 등과 같은 웨어러블 디바이스일 수 있다. 또한, 예를 들어, 사용자 장치(701)는 스트리밍 미디어 디바이스, 미디어 플레이어, 자동차 엔터테인먼트 시스템 등과 같은 미디어 디바이스일 수도 있다. 또한, 예를 들어, 사용자 장치(701)는 VR 기기 또는 AR 글라스를 포함하는 기기 등과 같은 XR(혼합 현실) 디바이스일 수도 있다. 즉, 사용자 장치(701)는 본 개시의 다양한 실시예들에 따른 데이터 시각화 서비스를 제공받기 위해 네트워크에 접속하는 전자적 장치일 수 있다.
도 8은, 다양한 실시예들에 따른, 데이터 시각화 및 인터랙션 방법을 설명하기 위한 도면이다.
도 8을 참조하면, 데이터 시각화 및 인터랙션 방법은 데이터 이미징 신청 단계(S810), 데이터 이미징 및 진단 단계(S820), 데이터 시각화 단계(S830) 및 인터랙션 단계(S840)를 포함할 수 있다.
데이터 이미징 신청 단계(S810)에서, 전자 장치는 데이터 셋에 대한 이미징, 진단 또는 시각화를 서버에 요청할 수 있다. 구체적으로, 전자 장치는 데이터 셋 및 데이터 셋에 대한 요청 정보를 서버에 전송할 수 있다. 서버는 수신한 데이터 셋 및 요청 정보를 기초로 데이터 셋을 처리할 수 있다.
데이터 이미징 및 진단 단계(S820)에서, 서버는 데이터 셋을 벡터화(vectorization)함으로써 데이터 셋을 이미징할 수 있다. 구체적으로, 서버는 데이터 셋을 특정 차원의 잠재 공간에 임베딩함으로써 데이터 셋에 대응되는 벡터 셋을 식별할 수 있다.
서버는 식별된 벡터 셋을 기초로 데이터 셋에 포함된 개별 데이터 포인트 간의 거리(distance), 인접도(neighbor relationship), 또는 군집(cluster) 구조 등을 분석함으로써, 데이터 셋의 내재적 특성(intrinsic property) 및 태스크 의존 특성(task-dependent attribute)을 진단할 수 있다.
예를 들어, 내재적 특성으로서 데이터의 밀도(density)를 측정하기 위해, 서버는 특정 데이터 포인트를 중심으로 근접한 K개의 이웃 벡터들까지의 거리 분포를 분석하거나, 편향성(bias)을 판별하기 위해 특정 클래스로 과도하게 밀집된 부분을 식별할 수 있다. 또한, 균질도(homogeneity)를 평가하기 위해 각 벡터가 속한 지역의 다양성(diversity) 또는 분산 정도를 측정할 수도 있다.
또한, 예를 들어, 서버는 분류(Classification) 태스크에 사용되는 데이터 셋에 대하여, 각 클래스가 벡터 공간상에서 어떻게 분포되는지(클래스 간 경계나 중첩 여부 등)를 식별함으로써 태스크 의존 특성을 진단할 수 있다. 이를 통해, 서버는 클래스 불균형(imbalanced classes)이 존재하거나, 특정 클래스가 다른 클래스와 과도하게 가까운 지점(하드 네거티브; hard negative)이 많은 경우 등을 찾아낼 수 있다.
이 과정에서 서버는 벡터들 간의 유클리디안 거리(Euclidean distance), 코사인 유사도(cosine similarity), 또는 기타 거리 중 적어도 하나를 고려하여, 이상 값(outlier)을 탐지하거나, 데이터 품질 저하 요인(예: 잘못된 라벨, 중복 데이터, 노이즈가 심한 샘플 등)을 식별할 수 있다. 또한, 서버는 이러한 분석 결과를 후속 단계(예: 데이터 정제(data cleaning)나 라벨 정정(label correction) 등)에 전달하여 데이터 셋의 품질 개선에 활용할 수 있다.
데이터 시각화 단계(S830)에서, 서버는 데이터 셋을 2차원 또는 3차원 공간에 가시화하여 표현할 수 있다. 서버는 데이터 셋 또는 데이터 셋에 대응되는 벡터 셋을 차원 축소하여 데이터 이미지(IOD)를 획득하고, 데이터 이미지를 2차원 또는 3차원 공간에 제공할 수 있다. 예를 들어, 서버는 PCA(주성분 분석), t-SNE, UMAP 등의 차원 축소 기법을 적용하거나, 신경망 기반 임베딩 모델을 사용하여 고차원 데이터를 낮은 차원의 시각화 공간으로 매핑할 수 있다.
여기서, 서버는 전 단계(예: 데이터 진단 단계)에서 산출된 진단 결과(군집 구조, 이상치(outlier), 편향성, 불균형 클래스 등)를 반영하여 시각화된 결과를 제공할 수 있다. 구체적으로, 서버는 이상치로 판별된 데이터 포인트를 별도의 색상이나 특수 마커(예: 삼각형, 별 모양 등)로 표시하거나, 클래스 간 경계가 모호한 영역이나 밀도가 높은 영역을 시각적으로 강조(예: 영역 테두리 표시, 그라데이션)함으로써, 사용자가 데이터의 분포 및 품질 문제를 한눈에 파악할 수 있도록 유도할 수 있다.
인터랙션 단계(S840)에서, 서버는 전자 장치(예: 사용자 장치)로부터 수신된 시각화된 데이터 이미지에 연관된 입력을 처리하여, 시각화된 데이터 이미지에 대한 동적 반응(dynamic response)을 제공할 수 있다. 구체적으로, 사용자는 시각화된 데이터 이미지 내의 임의 지점(예: 특정 데이터 포인트, 클러스터, 또는 선택된 영역)에 대하여 터치, 마우스 클릭, 드래그, 핀치 줌(pinch zoom) 등의 입력을 수행할 수 있다.
예를 들어, 사용자가 데이터 포인트 하나를 클릭하는 경우, 본 개시의 컴퓨팅 장치(서버)는 해당 데이터 포인트에 연관된 메타데이터(예: 개별 특성값, 라벨 정보, 통계값 등)를 조회하여 팝업창, 툴팁, 또는 별도의 레이어 상에 표시할 수 있는 정보를 사용자 장치로 전송함으로써 응답을 제공할 수 있다. 또 다른 예로, 사용자가 시각화된 공간에서 특정 구간을 드래그하여 영역을 지정하면, 서버는 지정된 영역 내 데이터 포인트들의 통계적 분포나 밀도, 클래스 분포 등을 추가 분석한 후 그 결과(예: 평균값, 편차, 대표 이미지, 샘플 수 등)를 사용자 장치에 제공할 수 있다.
또한, 사용자는 시각화된 데이터 이미지의 축소, 확대, 회전(3차원 시각화의 경우), 혹은 색상 스케일(color scale) 변경과 같은 인터랙션을 수행함으로써, 다양한 방식으로 데이터를 관찰할 수 있다. 서버는 이러한 변경 요청에 대응하여, 차원 축소 기법(PCA, t-SNE, UMAP 등)이나 색상·형상 매핑 알고리즘 등을 다시 적용하거나 파라미터를 조정하여, 갱신된 데이터 이미지를 생성하고 이를 사용자 장치에 제공한다.
이처럼, 본 개시의 인터랙션 단계(S840)를 통해 사용자는 시각화된 데이터에 대한 직관적이고 즉각적인 탐색이 가능해지며, 데이터의 내재적 구조를 보다 깊이 이해하거나 잠재적 품질 문제(예: 라벨 오류, 이상치, 클러스터 간 경계 불명확성 등)를 조기에 발견할 수 있다. 나아가, 이러한 인터랙션 기능은 머신러닝 모델 개발자가 데이터 클리닝(cleaning) 또는 레이블 정정(label correction) 작업을 수행하는 과정에서도 효과적으로 활용될 수 있다.
본 단계에서 설명된 각종 인터랙션 기능은 실시예를 보다 명확히 하기 위한 예시에 불과하며, 상술되지 않은 다른 형태의 사용자 입력(예: 제스처 입력, 음성·텍스트 명령 등)에 대한 응답 방식도 본 발명이 속하는 기술 분야에서 통상의 지식을 가진 자에게 자명하게 변형·응용 가능함은 물론이다. 따라서 본 명세서에서 사용된 특정 예시들에 의해 본 발명의 권리범위가 제한되지 않는다는 점이 이해되어야 한다.
도 9는, 다양한 실시예들에 따른, 이미징 신청 단계에서 수행되는 세부적인 단계들을 설명하기 위한 도면이다.
사용자는 데이터의 이미징 및 진단을 신청하는 과정에서, 데이터 이미징 및 진단에 사용될 도구(예: 렌즈 등의 데이터 이미징을 위한 인공지능 모델 등) 데이터 시각화 도구 등에 대한 설정을 입력할 수 있다.
도 9를 참조하면, 데이터 이미징 신청 단계(S810)는 렌즈 수준 선택 단계(S811), 렌즈 속성 결정 단계(S813), 기타 설정 단계(S815) 및 시각화 도구 선택 단계(S817)를 포함할 수 있다.
렌즈 수준 선택 단계(S811)에서, 사용자는 데이터 이미징에 사용될 적어도 하나의 인공지능 모델을 포함하는 렌즈의 수준(level)을 선택할 수 있다. 본 개시에서 렌즈의 수준(level)은, 데이터 이미징 방식과 관련된 미리 설정된 기준에 따라 분류되어 정의될 수 있다. 예를 들어, 렌즈의 수준(level)은 데이터 셋의 정량적 지표만을 진단 (기초 통계, 결측치 비율, 데이터 이상치 검출 등)하는 제1 수준, 데이터 셋을 미리 저장된(사전 훈련된) 인공지능 모델을 이용하여 벡터화하여 진단하는 제2 수준 또는 데이터 셋에 최적화된 인공지능 모델을 추가 학습(fine-tuning)하거나 새롭게 학습한 후 벡터화하여 진단하는 제3 수준 등을 포함할 수 있으나, 이에 한정되지 않는다. 예컨대 데이터 셋 규모가 매우 크거나 높은 정밀도가 요구되는 경우에는 제3 수준을 선택하여 모델을 직접 재학습함으로써 보다 정교한 진단 결과를 얻을 수 있다.
렌즈 속성 결정 단계(S813)에서, 사용자는 데이터 이미징에 사용될 적어도 하나의 인공지능 모델을 포함하는 렌즈의 속성을 결정할 수 있다. 본 개시에서 렌즈의 속성이란, 인공지능 모델의 구조, 하이퍼파라미터(예: 레이어 수, 파라미터 수, 학습률, 배치 크기 등), 또는 탐색 전략(예: 옵티마이저 종류, 초기화 방식) 등을 포함할 수 있다. 예를 들어, 사용자는 "모델 복잡도(파라미터 수)"를 조절하거나, "학습 에포크 수"를 설정하여 분석과 처리 시간 간의 균형을 맞출 수 있다. 또한, 특정 도메인(이미지, 텍스트, 구조화된 데이터 등)에 맞는 사전 학습 모델을 사용할지, 혹은 범용 모델(Generic AI Model)을 사용할지도 선택 가능하다.
기타 설정 단계(S815)에서, 사용자 요구사항을 반영하기 위한 세부 옵션들이 제공될 수 있다. 예를 들어, 사용자 장치는 진단 과정을 스킵하고 단순 시각화만 할지, 아니면 진단까지 수행할지 여부 선택함으로써 데이터 진단 여부를 설정할 수 있다. 또한, 예를 들어, 사용자 장치는 군집 밀도나 편향성 지표를 별도의 색상 스케일로 표시할지 결정함으로써 진단 결과 시각화 여부를 설정할 수 있다. 또한, 예를 들어, 사용자 장치는 데이터 포인트 간의 유사도(거리) 기준이나, 특정 특성(라벨 정보 등)에 따라 서브 그룹 또는 커뮤니티의 생성 방식에 연관된 커뮤니티(community) 생성 기준을 설정할 수 있다. 또한, 예를 들어, 사용자 장치는 일정 시점(view-point)에서의 데이터 분포(시각화 상태)를 별도로 저장하여 비교 또는 분석하는 스냅샷(snapshot)을 생성할지 여부를 설정할 수 있다.
시각화 도구 선택 단계(S817)에서, 사용자는 미리 연동된 복수의 시각화 도구들(예: PCT, UMAP, T-SNE 등) 중 적어도 하나를 선택할 수 있다. 이 경우, 사용자는 각 도구의 특징(차원 축소 방식, 결과 해석 용이성, 처리 속도 등)을 고려하여 선택할 수 있다. 또한, 미리 설정된 가시화 차원(예: 2차원, 3차원)을 선택함으로써, 단순 평면(projection)에서 확인할지, 3차원 공간에서 회전 또는 확대 등의 상호 작용을 할지 결정할 수 있다.
예를 들어, PCA 방식은 속도가 빠르고 해석이 직관적인 반면, t-SNE나 UMAP은 데이터 군집 구조를 더 정교하게 반영하는 장점이 있다. 사용자는 데이터 셋의 특성, 분석 목표, 컴퓨팅 리소스 등을 종합적으로 고려하여 시각화 도구를 선택할 수 있다.
이상 설명한 바와 같이, 데이터 이미징 신청 단계(S810)에서 사용자는 다양한 설정 및 결정 사항을 직접 입력하여, 데이터 이미징 및 진단 프로세스가 자신의 분석 목적과 환경에 최적화되도록 제어할 수 있다. 이러한 사용자 중심의 설정 단계는 후속 단계(S820, S830 등)에서 수행될 벡터화, 진단, 시각화 과정을 보다 효과적으로 진행하도록 지원하며, 궁극적으로 데이터 품질 개선 및 인공지능 모델 성능 향상에 기여한다.
도 10은, 다양한 실시예들에 따른, 렌즈의 수준 및 시각화 도구에 따른 데이터 이미징 및 진단의 결과를 설명하기 위한 도면이다.
도 11은, 다양한 실시예들에 따른, 컴퓨팅 장치에 포함되는 렌즈 빌더를 설명하기 위한 도면이다.
컴퓨팅 장치는 사용자에 의해 입력된 설정에 따라 데이터 셋을 이미징하기 위한 렌즈를 구축하고, 구축된 렌즈 및 시각화 도구를 이용하여 데이터 셋을 진단 및 시각화할 수 있다.
컴퓨팅 장치는 사용자에 의해 입력된 설정에 따라 데이터 셋을 이미징하기 위한 렌즈를 구축하고, 구축된 렌즈 및 시각화 도구를 이용하여 데이터 셋을 진단 및 시각화할 수 있다. 도 10을 참조하면, 컴퓨팅 장치(예: 서버)는 데이터 셋(dataset)을 렌즈 빌더(LENS BUILDER)에 제공할 수 있다. 여기서 렌즈 빌더는 데이터 셋의 이미징 및 진단에 대한 사용자 설정을 반영하여, 데이터 셋을 처리하기 위한 도구(예: 인공지능 모델, 통계 연산기 등)를 구성 또는 조합하거나 학습 또는 튜닝함으로써 적절한 렌즈(예: 데이터 이미징, 진단용 도구 세트)를 구축하기 위한 구성 요소이다. 렌즈 빌더는 별도의 독립적인 하드웨어 장치나 소프트웨어 모듈로 구현될 수도 있고, 적어도 하나의 프로세서가 렌즈 구축을 위해 메모리에 저장된 복수의 인스트럭션을 실행하는 형태로 구현될 수도 있다.
예를 들어, 도 11을 참조하면, 렌즈 빌더는 복수의 인공지능 모델들을 포함하는 모델 저장부, 인공지능 모델을 학습하기 위한 모델 학습부, 인공지능 모델의 속성을 결정하기 위한 속성 결정부, 데이터의 통계적 특성을 연산하기 위한 복수의 연산기들을 포함하는 연산 도구 저장부 및 구축된 렌즈들을 저장하기 위한 렌즈 저장부를 포함할 수 있으나, 이에 한정되지 않는다.
구체적으로, 모델 저장부는 복수의 사전학습 모델들(예: 제1 사전학습 모델, 제2 사전학습 모델, 쪋)을 저장하는 제1 저장부, 복수의 파운데이션 모델들(예: 제1 파운데이션 모델, 제2 파운데이션 모델, 쪋)을 저장하는 제2 저장부 및 복수의 어댑터들(예: 제1 어댑터, 제2 어댑터, 쪋)을 저장하는 제3 저장부 등을 포함할 수 있으나, 이에 한정되지 않는다. 여기서, 어댑터는 기존 인공지능 모델을 변환하기 위하여 인공지능 모델에 전자적으로 연결되는 구성으로, 데이터 변환기, 인공지능 변환기, 저랭크 어댑터 또는 변환 모듈 등으로 지칭될 수 있다. 또한, 연산 도구 저장부는 복수의 연산기들(예: 제1 연산기, 제2 연산기, 쪋)을 저장할 수 있고, 렌즈 저장부는 복수의 렌즈들(예: 제1 렌즈, 제2 렌즈, 제3 렌즈, 제4 렌즈, 쪋)을 저장할 수 있으나, 이에 한정되지 않는다.
다시 도 10을 참조하면, 컴퓨팅 장치는 렌즈 빌더를 이용하여, 데이터 셋을 이미징하거나 진단하기 위한 적어도 하나의 렌즈를 구축할 수 있다. 이때, 적어도 하나의 렌즈는 데이터 셋을 처리하여 데이터 셋에 연관된 특성을 산출하기 위한 적어도 하나의 인스트럭션이 저장된 연산 도구를 포함할 수 있다.
일 예로, 컴퓨팅 장치는 데이터 셋을 통계적으로 분석함으로써 데이터 셋을 진단할 수 있다.
구체적으로, 컴퓨팅 장치는 제1 수준의 렌즈에 대한 요청에 기초하여, 제1 렌즈(LENS #1)를 생성할 수 있다. 이때, 제1 렌즈는 데이터 셋의 통계적 특성을 진단하기 위한 적어도 하나의 연산 도구를 포함할 수 있다.
예를 들어, 도 10 및 도 11을 참조하면, 컴퓨팅 장치는 렌즈 빌더를 이용하여, 연산도구 저장부에 저장된 복수의 연산기들 중 적어도 하나를 포함하는 제1 렌즈를 생성할 수 있다.
예를 들어, 컴퓨팅 장치는 연산 도구 저장부에 저장된 복수의 연산기(예: 결측치 검출기, 이상치 감지기, 분산·편차 계산기, 상관관계 분석기 등) 중 적어도 하나를 선택 또는 조합하여 제1 렌즈를 구성할 수 있다. 예컨대, 사용자가 "기초 통계 수준 진단"을 요청하면, 컴퓨팅 장치는 평균값, 최솟값, 최댓값 또는 표준편차 등을 산출하는 연산기를 활성화하고, "결측치 비율 진단"을 추가로 선택한 경우, 결측치 검출기를 포함하도록 제1 렌즈를 구축할 수 있다.
또한, 컴퓨팅 장치는 데이터 셋을 제1 렌즈에 입력하고, 제1 렌즈를 통해 데이터 셋에 대한 제1 진단 결과(1st DIAGNOSIS RESULT)를 획득할 수 있다. 제1 진단 결과에는, 예를 들어, 전체 데이터 수, 결측치 비율, 이상치 비율 등 데이터 품질 상태를 나타내는 기초 지표, 특정 속성 별 통계치(평균, 분산, 왜도(skewness), 첨도(kurtosis) 등), 상호 연관관계(피어슨 상관계수, 스피어만 상관계수 등) 분석 결과, 또는 클래스별 데이터 건수, 편향성 지표(bias measure) 등이 포함될 수 있다.
이를 통해 사용자는 해당 데이터 셋의 전반적인 통계적적 특성을 빠르게 파악할 수 있으며, 결측치 처리나 이상치 제거와 같은 후속 작업을 결정할 수 있다.
다른 예로, 컴퓨팅 장치는 데이터 셋을 벡터화하여 벡터화된 데이터를 분석함으로써 보다 심층적인 특성 진단을 수행할 수 있다.
구체적으로, 컴퓨팅 장치는 제2 수준의 렌즈에 대한 요청에 기초하여, 적절한 사전학습 인공지능 모델(pre-trained AI model)을 사용하는 제2 렌즈(LENS #2)를 생성할 수 있다.
예를 들어, 도 10 및 도 11을 참조하면, 컴퓨팅 장치는 모델 저장부의 제1 저장부에 저장된 복수의 사전학습 모델들(예: 컴퓨터 비전용 CNN 모델, 텍스트 임베딩용 Transformer 모델 등) 중 사용자 요구 사항이나 데이터 형태(이미지, 텍스트, 구조화 데이터 등)에 맞는 모델 하나 이상을 불러와 제2 렌즈를 생성할 수 있다.
또한, 컴퓨팅 장치는 데이터 셋을 적어도 하나의 사전학습 모델에 입력하여, 인공지능 모델의 적어도 하나의 레이어로부터 데이터 셋에 대응되는 제1 벡터 셋(1st VECTOR SET)을 획득할 수 있다. 컴퓨팅 장치는 제1 벡터 셋을 기초로 데이터 셋에 포함되는 데이터의 적어도 하나의 특성 값을 연산함으로써 데이터 셋에 대한 제2 진단 결과(2nd DIAGNOSIS RESULT)를 획득할 수 있다. 즉, 컴퓨팅 장치는 제1 벡터 셋을 분석(예: 유클리디안 거리, 군집화, 밀도 계산 등)하여 데이터 셋의 내재적 특성이나 클래스 분포, 편향성(bias) 등 더 높은 수준의 통찰을 제공하는 제2 진단 결과(2nd DIAGNOSIS RESULT)를 산출할 수 있다.
제2 진단 결과에는, 예를 들어, 임베딩 공간에서의 데이터 간 군집(cluster) 구조(클러스터 개수, 대표 중심점), 벡터 사이의 밀도/분산(특정 구역에 데이터가 과도하게 밀집되어 있는지), 클래스별 경계 모호성(하드 네거티브(hard negative) 비율, 클래스 간 거리 분포 등), 또는 특정 속성값을 기준으로 본 부분군(subgroup) 간 차이등을 포함할 수 있으나, 이에 한정되지 않는다.
이를 통해, 단순 통계적 특성과 달리, 데이터의 내재적 구조나 잠재적 편향 문제를 포착할 수 있으며, 모델 학습 시 발생할 수 있는 오류나 성능 저하의 원인을 파악하는 데 용이하다.
또한, 컴퓨팅 장치는 제3 수준의 렌즈에 대한 요청에 기초하여, 데이터 셋에 최적화된 인공지능 모델을 학습한 뒤 이를 통해 벡터화를 수행하기 위한 제3 렌즈(LENS #3)를 생성할 수 있다.
예를 들어, 도 10 및 11을 참조하면, 컴퓨팅 장치는 모델 저장부의 제2 저장부에 저장된 복수의 파운데이션 모델들 중 적어도 하나 또는 제3 저장부에 저장된 복수의 어댑터들 중 적어도 하나를 포함하는 제3 렌즈를 생성할 수 있다. 구체적으로, 컴퓨팅 장치는 모델 학습부를 이용하여, 복수의 파운데이션 모델들 중 적어도 하나 또는 복수의 어댑터들 중 적어도 하나를 데이터 셋에 최적화되도록 학습함으로써 제3 렌즈를 생성할 수 있다.
구체적인 예로, 컴퓨팅 장치는 데이터 셋을 기초로 적어도 하나의 파운데이션 모델을 미세 조정(fine-tuning)함으로써 제3 렌즈를 생성할 수 있다. 또는 컴퓨팅 장치는 데이터 셋을 이용하여 적어도 하나의 어댑터를 학습하고, 학습된 어댑터 및 적어도 하나의 파운데이션 모델을 포함하는 제3 렌즈를 생성할 수 있다.
예를 들어, 컴퓨팅 장치는 데이터 셋의 내재적 특성을 산출하는 데에 최적화된 차원(dimension)을 결정하고, 결정된 차원의 잠재 공간에 벡터화하도록 적어도 하나의 인공지능 모델을 학습할 수 있다. 구체적인 예로, 컴퓨팅 장치는 데이터 분포를 추정하기 위한 다양한 지표(예: 재구성 오류(Reconstruction Error), 군집 품질 지표(Silhouette score, Davies-Bouldin Index 등), 분류 또는 회귀 정확도 등)를 계산하여, 임베딩 차원이 변함에 따라 이들 지표가 어떻게 달라지는지 분석할 수 있다. 예컨대, 오토인코더 방식으로 잠재 공간의 차원을 16, 32, 64 등으로 바꿔가며 재학습하여, 재구성 오류가 최소화되는 지점이나, 특정 태스크(분류·군집 등)의 성능이 가장 높은 차원을 탐색함으로써, 최적 차원을 결정할 수 있으나, 이에 한정되지 않는다.
또한, 예를 들어, 컴퓨팅 장치는 대규모 사전학습된 파운데이션 모델(예: Vision Transformer, GPT 계열 모델 등) 중 특정 도메인(예: 의료영상, SNS 텍스트 등)에 적합한 후보 모델을 골라, 사용자 데이터 셋으로 미세 조정(fine-tuning)을 수행할 수 있다. 또는, 파운데이션 모델은 고정해두고, 어댑터(예: LoRA, Prompt Tuning 등)를 학습시켜 모델 파라미터 일부만 업데이트함으로써, 계산량을 줄이면서도 사용자의 데이터 분포에 더 정밀하게 맞출 수 있다.
이렇게 학습 또는 튜닝된 모델을 포함하는 제3 렌즈는 해당 데이터 셋에 특화된 임베딩을 생성하므로, 더욱 높은 진단 정밀도와 통찰을 제공한다.
컴퓨팅 장치는 데이터 셋을 제3 렌즈에 입력하여, 데이터 셋을 최적 차원의 잠재 공간에 임베딩함으로써 제2 벡터 셋(2nd VECTOR SET)을 획득할 수 있다. 또한, 컴퓨팅 장치는 제2 벡터 셋을 기초로 데이터 셋에 포함되는 데이터의 적어도 하나의 특성 값을 연산함으로써 데이터 셋에 대한 제3 진단 결과(3rd DIAGNOSIS RESULT)를 획득할 수 있다. 즉, 컴퓨팅 장치는 제2 벡터 셋을 분석(예: 유클리디안 거리, 군집화, 밀도 계산 등)하여 데이터 셋의 내재적 특성이나 클래스 분포, 편향성(bias) 등 더 높은 수준의 통찰을 제공하는 제3 진단 결과를 산출할 수 있다.
제2 렌즈(사전학습된 인공지능 모델 포함)에 의해 산출되는 제1 벡터 셋은, 사전학습 모델의 구조 및 파라미터에 따라 미리 정해진 차원(이하, "제1 차원")을 기반으로 정의될 수 있다. 반면, 제3 렌즈(데이터 셋에 최적화된 인공지능 모델 포함)에 의해 산출되는 제2 벡터 셋은, 데이터 셋의 내재적 특성(intrinsic property)을 분석하여 적응적으로 결정된 차원(이하, "제2 차원")을 기반으로 정의될 수 있다.
이에 따라, 제1 벡터 셋이 정의되는 제1 차원은 제2 벡터 셋이 정의되는 제2 차원과 상이할 수 있다. 제1 벡터 셋이 정의되는 제1 차원은 인공지능 모델에 의해 결정되고, 제2 벡터 셋이 정의되는 제2 차원은 데이터 셋에 따라 적응적으로 결정될 수 있다. 예컨대, 제2 렌즈에 따른 제1 차원은 특정 사전학습 모델(예: 768차원, 1024차원 등)의 고정 구조를 반영하지만, 제3 렌즈에 따른 제2 차원은 데이터 셋의 매니폴드 구조, 재구성 오류(reconstruction error), 혹은 군집 품질 지표 등을 고려하여 최적화될 수 있다.
나아가, 본 개시의 다른 실시예들에 따르면, 제1 차원과 제2 차원의 결정 방식이 상이하기 때문에, 제1 벡터 셋과 제2 벡터 셋이 나타내는 데이터 분포 및 특성도 서로 다를 수 있다. 즉, 제2 수준의 렌즈를 이용한 사전학습 모델 기반 임베딩은 일반적이고 보편화된 특징을 반영하는 반면, 제3 수준의 렌즈를 이용한 데이터 셋에 맞춤형으로 학습된 임베딩은 대상 데이터 셋의 특수성(편향 구조, 클래스 경계, 노이즈 양상 등)을 더 정교하게 반영할 수 있다.
이에 의해, 데이터 진단 혹은 시각화 절차에서, 제1 벡터 셋을 이용한 결과와 제2 벡터 셋을 이용한 결과 간에 분포 형태, 클러스터 구조, 편향 지표 등이 달라질 수 있으며, 본 개시에 따른 컴퓨팅 장치는 두 벡터 셋을 상호 비교·분석함으로써 데이터 셋의 고유 특성과 보편적 특징을 균형 있게 파악할 수 있도록 지원할 수도 있다.
또한, 컴퓨팅 장치는 데이터 셋에 대응되는 벡터 셋(예: 제1 벡터 셋 또는 제2 벡터 셋)을 적어도 하나의 시각화 도구(VISUALIZATION TOOL)에 제공하여, 데이터 셋을 가시화한 데이터 이미지(Image of Data, IOD)를 사용자에게 제공할 수 있다.
이때, 컴퓨팅 장치는 시각화 도구의 종류 또는 시각화 도구에 의해 가시화되는 차원에 따라 복수의 서로 상이한 형태의 데이터 이미지(IOD) 중 적어도 하나를 생성하여 출력할 수 있다.
예를 들어, 제1 시각화 도구(예: PCA, 2D 기반 차원 축소 알고리즘)를 이용하는 경우, 컴퓨팅 장치는 벡터 셋을 2차원으로 투영(projection)하여, 2차원의 제1 데이터 이미지(IOD #1)를 획득한 뒤 이를 제공할 수 있다.
또한, 예를 들어, 3차원 시각화가 가능한 제2 시각화 도구(예: 특정 3D 차원 축소 알고리즘, WebGL 기반 3D 뷰어 등)를 사용하는 경우, 컴퓨팅 장치는 3차원의 제2 데이터 이미지(IOD #2)를 획득해 사용자 디바이스(예: PC, 스마트폰, XR 기기 등)에 송신함으로써, 입체적 관점에서 데이터 분포를 확인할 수 있게 유도할 수 있다.
또한, 예를 들어, 제3 시각화 도구(예: t-SNE, UMAP 등)를 통해서도 마찬가지로 2차원, 3차원 또는 그 외 차원으로 데이터의 차원을 축소 및 투영하여, 제3 데이터 이미지(IOD #3)를 획득할 수 있다.
이와 같이, 본 발명에 따른 컴퓨팅 장치는 벡터 셋을 다양한 시각화 도구에 제공해 복수 형태의 데이터 이미지를 생성함으로써, 사용자가 데이터 분포의 특성, 클러스터 구조, 이상치(outlier)의 위치, 클래스 간 경계 등을 직관적으로 확인할 수 있도록 지원한다. 특히, 사용자는 원하는 시각화 도구와 차원을 자유롭게 선택 또는 변환함으로써, 동일 데이터 셋에 대해서도 상이한 시각적 표현 결과(IOD #1, IOD #2, IOD #3 등)를 비교 및 검토할 수 있다.
예를 들어, 사용자가 초기에는 2차원 시각화를 통해 전체적인 분포를 관찰한 뒤, 3차원 시각화로 전환하여 특정 군집(Cluster) 사이의 공간적 인접 관계를 더욱 명확하게 파악할 수 있다. 또는, t-SNE와 UMAP을 번갈아 적용하며 지역적 구조(local structure)와 전역적 구조(global structure) 간의 차이를 비교하여, 데이터 품질상 문제가 되는 이상치나 편향성을 조기에 식별할 수 있다.
이와 같이, "데이터 이미지(Image of Data, IOD)"는 본 개시에서 벡터 셋을 시각화한 결과물을 지칭하되, 특정 표준(2D·3D), 특정 기법(PCA·t-SNE·UMAP·Autoencoder 기반 시각화 등)에 국한되지 않는다. 나아가, 컴퓨팅 장치는 시각화 결과에 대한 사용자 인터랙션(예: 확대 및 축소, 특정 지점 클릭, 범위 지정, 색상 스케일 변경 등)을 실시간으로 수신하여, 가시화된 데이터 이미지를 갱신하거나, 세부 정보(예: 각각의 포인트에 대응되는 메타데이터) 등을 추가로 표시해줄 수도 있다.
또한, 컴퓨팅 장치는 제1 진단 결과, 제2 진단 결과 또는 제3 진단 결과를 기초로 진단 레포트(DIAGNOSIS REPORT)를 생성하여 제공할 수 있다. 예를 들어, 컴퓨팅 장치는 제1 진단 결과(통계 지표, 결측치 또는 이상치 비율 등), 제2 진단 결과(사전학습 모델 기반 임베딩 분석 결과, 군집 분포도, 클래스 간 거리 등), 또는 제3 진단 결과(맞춤형 모델 또는 어댑터를 통해 얻은 고정밀 임베딩 및 특성 분석 결과)를 각각 데이터베이스에 저장하거나, 별도의 JSON 또는 XML 등 구조화 포맷으로 관리한 뒤, 이를 대시보드 형태나 문서 형태로 결합 또는 배치하여 종합 리포트를 생성할 수 있다.
진단 레포트는, 예를 들어, 데이터 품질 등급에 대한 정보, 데이터의 품질을 저하시키는 요인에 대한 요약 정보(예: 주요 이상치 또는 편향 원인, 잠재적 라벨 불일치 가능성 지표), 데이터 개선에 대한 추천 정보(예: 데이터 다이어트 또는 벌크업 추천 정보), 군집 별 시각화 정보(예: 임베딩 공간에서 주요 클러스터를 시각적으로 표시, 밀집도 또는 거리 측정 결과 등), 또는 학습 모델 성능 예측 정보(예: 특정 모델로 가정 시 예측 정확도 또는 F1 Score 추정치 등)를 포함할 수 있으나, 이에 한정되지 않는다.
이러한 레포트는 사용자(예: 데이터 분석가, AI 모델 개발자 등)에게 정형화된 형태로 전달될 수 있으며, 곧바로 시각화 인터페이스(GUI)에서 차트, 테이블 또는 히트맵 등으로 표현되어 분석 단계나 모델 수정 단계에 즉각 활용될 수도 있다.
도 12는, 다양한 실시예들에 따른, 컴퓨팅 장치가 사용자 입력에 따라 데이터 이미징 및 진단을 위한 데이터 처리 모델을 구축하는 방법을 도시한 도면이다.
도 12를 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 제1 데이터 셋을 획득하는 동작(S1210)을 수행하도록 설정될 수 있다. 이때, 컴퓨팅 장치는 제1 데이터 셋을 사용자 디바이스로부터 수신하거나, 저장된 제1 데이터 셋을 불러옴으로써 제1 데이터 셋을 획득할 수 있다.
또한, 적어도 하나의 프로세서는 사용자 입력을 기초로, 데이터 처리 방식에 따라 분류된 복수의 수준들 중 제1 수준을 선택하는 동작(S1220)을 수행하도록 설정될 수 있다. 구체적으로, 적어도 하나의 프로세서는 복수의 수준들 중 제1 수준을 선택하는 사용자 입력에 따라 제1 수준을 선택할 수 있다. 이 경우, 적어도 하나의 프로세서는 제1 수준에 대응되는 제1 데이터 처리 방식을 결정할 수 있다. 적어도 하나의 프로세서는 제1 데이터 셋을 처리하기 위한 처리 알고리즘이 저장된 도구들(예: 연산기 또는 인공지능 모델 등) 중 적어도 하나를 결정함으로써 제1 데이터 처리 방식을 결정할 수 있다.
또한, 적어도 하나의 프로세서는 제1 수준에 대응되는 제1 데이터 처리 모델에 대한 적어도 하나의 속성을 결정하는 동작(S1230)을 수행하도록 설정될 수 있다. 이때, 제1 데이터 처리 모델은 적어도 하나의 인공지능 모델을 포함할 수 있다. 구체적으로, 제1 데이터 처리 모델은 미리 저장된 적어도 하나의 사전학습 모델 또는 제1 데이터 셋에 최적화되어 학습된 인공지능 모델을 포함할 수 있다. 예를 들어, 제1 데이터 셋에 최적화된 이미징 및 진단을 요청하는 입력에 기초하여, 적어도 하나의 프로세서는 제1 데이터 셋에 최적화된 인공지능 모델을 학습함으로써 제1 데이터 처리 모델을 생성할 수 있으나, 이에 한정되지 않는다. 또한, 적어도 하나의 속성은 제1 데이터 처리 모델이 데이터를 처리하는 연산에 관련된 속성을 포함할 수 있다. 예를 들어, 적어도 하나의 속성은 인공지능 모델의 파라미터 또는 하이퍼 파라미터 개수 또는 그래픽 카드 정보 등을 포함할 수 있으나, 이에 한정되지 않는다.
또한, 적어도 하나의 프로세서는 제1 데이터 처리 모델을 이용하여, 제1 데이터 셋을 처리하는 작업에 연관되는 사전 정보를 제1 GUI(Graphic User Interface)를 통해 제공하는 동작(S1240)을 수행하도록 설정될 수 있다. 이때, 사전 정보는 컴퓨팅 장치가 데이터 셋을 처리하여 데이터 셋의 이미징 및 진단에 대한 결과를 산출하는 것과 관련하여, 사용자에게 알리도록 설정된 정보일 수 있다.
도 13은, 다양한 실시예들에 따른, 사전 정보에 포함되는 정보를 설명하기 위한 도면이다.
도 13을 참조하면, 사전 정보는 데이터 셋의 처리에 소요되는 비용과 연관되는 비용 정보, 데이터 셋 처리에 대한 결과를 예측하여 미리 나타낸 결과 프리뷰 정보, 데이터 셋의 처리의 진행 상황을 나타내는 진행 정보, 데이터 셋과 유사한 케이스의 처리 결과를 나타내는 레퍼런스 결과 정보, 또는 데이터 처리에 이용되는 처리 모델의 속성을 나타내는 속성 정보 등을 포함할 수 있다. 이에 따라 사용자는 데이터 셋 이미징 및 진단 신청 시점에서 처리 예상 비용, 예측 결과, 모델 진행도 등을 미리 확인함으로써, 보다 원활하게 서비스를 이용할 수 있다
본 개시에서 "비용 정보"는 데이터 셋의 처리에 소요되는 비용과 연관된 정보를 의미할 수 있다. 예를 들어, 비용 정보는 데이터 셋의 처리를 위한 연산 소요 비용(예: GPU 시간, CPU 코어 시간, 메모리 사용량 등)에 대한 추정값을 포함하거나, 사용자가 해당 처리 서비스를 이용하기 위해 지불해야 하는 결제 비용(예: 크레딧, 포인트, 화폐 단위 등)등을 포함할 수 있다. 구체적으로, "예상 연산 시간 2시간, GPU 1대 사용 시 비용 10달러"와 같은 형태로 사용자에게 제시됨으로써, 사용자는 데이터 셋의 크기, 모델 복잡도 등에 따라 어느 정도의 자원 및 비용이 소모되는지 미리 파악할 수 있다.
본 개시에서 "결과 프리뷰 정보"는 데이터 셋 처리 후 예상되는 일부 핵심 결과를 간략하게 미리 제시함으로써, 사용자가 결과의 대략적 형태나 가치를 가늠하게 해주는 정보를 의미할 수 있다. 이때, 컴퓨팅 장치는 인공지능 모델 학습 후 예상되는 정확도 범위, 클러스터 분포 미리보기 또는 샘플 시각화 이미지 등을 간략히 표시해줄 수 있다. 또한, 예를 들어, 결과 프리뷰 정보는 데이터 셋을 시각화 도구에 단순 입력하여 획득된 데이터 이미지를 포함할 수 있다. 이러한 데이터 이미지는 데이터 셋의 내재적 특성을 정확히 반영하지는 못하지만, 데이터 셋의 시각화 결과를 사전에 제공함으로써 사용자가 결과를 예측할 수 있도록 유도할 수 있다.
이를 통해 사용자는 최종 결과물을 얻기 전에도, 본 개시의 데이터 이미징 및 진단 프로세스가 자신의 목적에 부합하는지 판단할 수 있으며, 경우에 따라 분석 방향 변경이나 옵션 재설정을 신속히 결정할 수 있다.
본 개시에서 "진행 정보"는 데이터 셋의 처리 알고리즘 또는 처리 상황 등에 대한 정보를 의미할 수 있다. 예를 들어, 진행 정보는 데이터 셋 기반의 최적화 학습 진행에 대한 학습 진행도, 학습 소요 시간 또는 남은 시간 등을 포함할 수 있다. 사용자는 현재 학습 단계(초기화, 피처 추출, 파인튜닝, 검증 등)와 남은 예상 시간 등을 실시간으로 확인함으로써, 처리 완료 시점을 예측하거나, 추가로 필요한 자원(예: 컴퓨팅 리소스) 할당 여부를 판단할 수 있다. 또한, 처리 도중 문제(예: 결측치 과다, 이상치 폭증 등)가 감지된 경우, 진행도와 함께 경고 메시지를 병행 표시함으로써 사용자에게 사전에 조치 기회를 제공할 수도 있다.
본 개시에서 "레퍼런스 결과 정보"는 사용자가 제출한 데이터 셋과 유사한 케이스에 대해 과거에 수행되었던 처리 결과를 예시로 제시함으로써, 예상 결과나 성능 지표 등을 간접적으로 가늠하게 하는 정보를 의미할 수 있다. 예를 들어, 레퍼런스 결과 정보는 입력된 제1 데이터 셋과 유사한 레퍼런스 데이터를 이미징 및 진단한 결과 데이터를 포함할 수 있다. 예컨대, "동일한 카테고리(도메인) 이미지 데이터 10만 건을 분석했을 때, 평균 정확도는 92% 수준이었고, 처리 시간은 약 4시간 소요"와 같은 형태로 제공될 수 있다.
이를 통해 사용자는 자신의 데이터 셋 분석 결과를 예측하거나, 타사/타인의 유사 데이터 처리 사례에서 얻어진 인사이트를 참고해 처리 전략을 세울 수 있다.
본 개시에서 "속성 정보"는 데이터 처리 도구에 연관되는 속성에 대한 정보를 의미할 수 있다. 예를 들어, 속성 정보는 모델 파라미터 개수 또는 그래픽 카드 정보 등을 포함할 수 있다. 예를 들어, "이 처리에는 트랜스포머 아키텍처 기반 모델(파라미터 약 1억 개)과 RTX 3090 GPU가 사용된다"라는 식으로 표시될 수 있다. 사용자는 이러한 정보를 통해 모델 규모, 자원 호환성(내부 GPU 또는 클라우드 GPU 여부), 개발 환경 등을 이해하여, 모델 활용도 및 학습 성능을 사전에 예측할 수 있다.
다시 도 12를 참조하면, 적어도 하나의 프로세서는 제1 GUI에 대한 사용자 입력을 기초로, 제1 데이터 처리 모델을 구축하는 동작(S1250)을 수행하도록 설정될 수 있다. 이때, 제1 GUI에 대한 사용자 입력은 컴퓨팅 장치에 의해 제안된 제1 데이터 셋의 처리에 대한 승인(approvement) 입력을 포함할 수 있다.
즉, 컴퓨팅 장치는 제1 데이터 셋에 대한 처리를 하기 이전에, 사용자에게 제1 데이터 셋의 처리에 대한 사전 정보를 제공하고, 사전 정보를 제공받은 사용자에 의해 승인 입력이 수신된 경우에 한하여, 제1 데이터 셋을 처리할 수 있다.
도 14는, 다양한 실시예들에 따른, 컴퓨팅 장치가 제1 데이터 처리 모델을 구축하는 일 예시를 도시한 도면이다.
도 14를 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 제1 데이터 셋을 기초로 복수의 어댑터들로부터 복수의 대표 값들을 획득하는 동작(S1401)을 수행하도록 설정될 수 있다. 이때, 적어도 하나의 프로세서는 제1 데이터 셋을 전처리하여 제1 데이터 셋에 대응되는 특징 값을 추출하고, 특징 값을 복수의 어댑터들 각각에 입력할 수 있다. 또한, 적어도 하나의 프로세서는 동일한 데이터가 입력된 복수의 어댑터들 각각으로부터 출력된 데이터를 기초로 복수의 대표 값들을 획득할 수 있다.
또한, 특징 값을 입력받은 각각의 어댑터는 제1 데이터 셋에 대응되는 벡터 셋을 출력할 수 있다. 이 경우, 적어도 하나의 프로세서는 출력된 벡터 셋을 미리 정해진 방식으로 연산하여 대표 값을 획득할 수 있다. 복수의 어댑터들로부터 출력되는 복수의 벡터 셋들은 서로 상이한 차원을 기초로 정의될 수 있다. 복수의 어댑터들은 서로 상이한 차원의 벡터 셋을 출력하도록 설정될 수 있다.
또한, 적어도 하나의 프로세서는 복수의 대표 값들을 기초로 미리 정해진 조건을 만족하는 제1 어댑터를 선택하는 동작(S1402)을 수행하도록 설정될 수 있다. 이때, 미리 정해진 조건은 제1 데이터 셋에 최적화된 어댑터를 결정하기 위해 설정된 조건을 포함할 수 있다. 구체적으로, 적어도 하나의 프로세서는 복수의 대표 값들 중 가장 낮은(또는 가장 높은) 값에 대응되는 적어도 하나의 어댑터를 식별함으로써 제1 어댑터를 선택할 수 있다. 이때, 제1 어댑터는 제1 데이터 셋을 나타내는 최적의 차원으로 제1 데이터 셋을 임베딩하도록 구현될 수 있다. 예를 들어, 제1 어댑터는 제1 데이터 셋을 기초로 제1 차원의 제1 벡터 셋을 출력하도록 설정될 수 있다.
또한, 적어도 하나의 프로세서는 제1 어댑터를 포함하는 제1 데이터 처리 모델을 구축하는 동작(S1404)을 수행하도록 설정될 수 있다. 구체적으로, 적어도 하나의 프로세서는 미리 저장된 파운데이션 모델 또는 사전학습 모델에 제1 어댑터를 통신적으로 연결함으로써 파운데이션 모델 또는 사전학습 모델 및 제1 어댑터를 포함하는 제1 데이터 처리 모델을 구축할 수 있다.
컴퓨팅 장치는 데이터 셋을 기초로 어댑터만 추가 학습 또는 미세 조정(fine-tuning)함으로써 데이터 셋에 최적화된 데이터 처리 모델을 구축할 수 있다. 컴퓨팅 장치는 미리 저장된 인공지능 모델에 어댑터만 추가 학습하기 때문에, 최소한의 연산 비용을 소모할 수 있다.
도 15는, 다양한 실시예들에 따른, 컴퓨팅 장치가 제1 데이터 처리 모델을 구축하는 다른 예시를 도시한 도면이다.
도 15를 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 제1 데이터 셋을 복수의 파운데이션 모델들에 입력하는 동작(S1501)을 수행하도록 설정될 수 있다.
또한, 적어도 하나의 프로세서는 복수의 파운데이션 모델들로부터 출력된 데이터를 기초로 복수의 대표 값들을 획득하는 동작(S1502)을 수행하도록 설정될 수 있다. 구체적으로, 복수의 파운데이션 모델들 각각은 제1 데이터 셋에 대응되는 벡터 셋을 출력할 수 있고, 적어도 하나의 프로세서는 출력된 벡터 셋을 미리 정해진 방식으로 연산하여 대표 값을 획득할 수 있다. 복수의 파운데이션 모델들로부터 출력되는 복수의 벡터 셋들은 서로 상이한 차원을 기초로 정의될 수 있다. 복수의 파운데이션 모델들은 서로 상이한 차원의 벡터 셋을 출력하도록 설정될 수 있다.
또한, 적어도 하나의 프로세서는 복수의 대표 값들을 기초로 미리 정해진 조건을 만족하는 제1 파운데이션 모델을 결정하는 동작(S1503) 및 제1 파운데이션 모델을 포함하는 제1 데이터 처리 모델을 구축하는 동작(S1504)을 수행하도록 설정될 수 있다. 미리 정해진 조건 및 복수의 대표 값들을 기초로 모델을 결정하는 구체적인 방법은 상술하였으므로, 생략하기로 한다.
도 16은, 다양한 실시예들에 따른, 컴퓨팅 장치가 제1 데이터 처리 모델을 구축하는 또 다른 예시를 도시한 도면이다.
도 16을 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 제1 데이터 셋을 제1 파운데이션 모델에 입력하는 동작(S1601)을 수행할 수 있다. 이때, 제1 파운데이션 모델은 동일한 인풋을 입력받아서 서로 상이한 결과 값을 출력하도록 구현된 복수의 노드들을 포함할 수 있다. 예를 들어, 제1 파운데이션 모델은 동일한 인풋이 병렬적으로 입력되는 복수의 노드들을 포함하는 적어도 하나의 히든 레이어를 포함할 수 있다.
적어도 하나의 프로세서는 제1 파운데이션 모델에 포함되는 복수의 노드들 - 복수의 노드들은 제1 데이터 셋에 대한 특징 값이 병렬적으로 입력됨 -로부터 복수의 대표 값들을 획득하는 동작(S1602)을 수행할 수 있다. 구체적으로, 복수의 노드들은 제1 데이터 셋에 대한 특징 값을 기초로 복수의 벡터 셋들을 출력하고, 적어도 하나의 프로세서는 복수의 벡터 셋을 미리 정해진 연산을 기초로 복수의 대표 값들을 획득할 수 있다.
또한, 적어도 하나의 프로세서는 복수의 대표 값들을 기초로 제1 노드를 활성화하는 동작(S1603)을 수행하도록 설정될 수 있다. 또한, 적어도 하나의 프로세서는 제1 노드가 활성화된 제1 파운데이션 모델을 포함하는 제1 데이터 처리 모델을 구축하는 동작(S1604)을 수행하도록 설정될 수 있다.
구체적으로, 적어도 하나의 프로세서는 복수의 대표 값들을 기초로 미리 정해진 조건을 만족하는 제1 노드를 결정할 수 있고, 제1 노드로만 인풋이 입력되도록 제1 노드를 활성화할 수 있다. 즉, 적어도 하나의 프로세서는 복수의 노드들 중 제1 데이터 셋에 최적화된 노드를 선정할 수 있고, 선정된 노드를 제외한 노드에는 통신 연결을 해제함으로써 제1 데이터 처림 모델을 구축할 수 있다.
도 17은, 다양한 실시예들에 따른, 컴퓨팅 장치가 구축된 제1 데이터 처리 모델을 이용하여 데이터 셋을 이미징 및 가시화하는 방법을 설명하기 위한 도면이다.
도 17을 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 제1 데이터 처리 모델의 구축 이후에, 제1 데이터 셋을 제1 데이터 처리 모델에 입력할 수 있다(S1260). 구체적으로, 컴퓨팅 장치는 제1 데이터 셋에 최적화된 벡터화를 위한 인공지능 모델의 학습이 완료된 경우, 제1 데이터 처리 모델의 구축이 완료된 것으로 판단하여 트리거 신호를 발생시킬 수 있다. 적어도 하나의 프로세서는 생성된 트리거 신호를 기초로 제1 데이터 셋을 제1 데이터 처리 모델에 입력할 수 있다.
또한, 적어도 하나의 프로세서는 제1 데이터 처리 모델로부터 제1 차원의 임베딩 영역에서 정의되는 제1 벡터 셋을 획득할 수 있다(S1270). 제1 벡터 셋에 포함되는 복수의 벡터(또는 데이터 포인트)들 각각은 제1 데이터 셋에 포함되는 각각의 단위 데이터에 대응될 수 있다.
또한, 적어도 하나의 프로세서는 제1 벡터 셋을 기초로 제1 데이터 셋에 포함되는 복수의 데이터 각각에 대응되는 복수의 특성 값들을 획득할 수 있다(S1280). 구체적으로, 적어도 하나의 프로세서는 제1 벡터 셋에 포함되는 적어도 둘 이상의 벡터들 사이의 거리 값을 기초로 복수의 데이터 각각에 대응되는 복수의 특성 값들을 획득할 수 있다.
또한, 적어도 하나의 프로세서는 제1 벡터 셋을 제1 시각화 도구에 제공하여 제1 데이터 셋에 대응되는 제1 데이터 이미지를 제공할 수 있다(S1290). 이때, 적어도 하나의 프로세서는 제1 벡터 셋 또는 복수의 특성 값들을 기초로, 제1 데이터 셋을 시각화하는 데 적절한 시각화 도구를 추천할 수도 있다.
예를 들어, 본 개시의 일 실시예에 따르면, 컴퓨팅 장치는 시각화 도구 DB를 포함하거나, 시각화 도구와 관련된 정보를 미리 관리하는 DB 서버와 통신하여, 데이터 특성에 기반한 최적의 시각화 도구를 검색·결정하는 알고리즘을 수행할 수 있다.
도 18은, 다양한 실시예들에 따른, 컴퓨팅 장치가 시각화 도구를 추천하는 방법을 설명하기 위한 도면이다.
도 18을 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 데이터 특성 추출 단계(S1810)를 수행할 수 있다. 구체적으로, 컴퓨팅 장치는 데이터 셋(또는 벡터 셋)에 대해 도메인(domain), 데이터 종류, 데이터 용량, 벡터 분포 밀도, 클래스별 개수 등과 같은 특성 값들을 분석 또는 추출할 수 있다. 예를 들어, 적어도 하나의 프로세서는 데이터 셋을 기초로 데이터 도메인, 데이터 크기(샘플 수) 및 차원 수, 벡터 간 평균 거리, 분산, 클러스터 수 등의 통계량 또는 모델 학습 사양(필요 컴퓨팅 자원, 처리 시간 등) 등을 추출할 수 있다.
이러한 특성 값들은 내부적으로 SQL 쿼리 생성 시 인자로 사용되거나, 알고리즘 로직을 수행하기 위한 매개변수가 될 수 있다.
또한, 컴퓨팅 장치는 시각화 도구 DB 조회 및 SQL 쿼리 생성 단계(S1820)를 수행할 수 있다. 컴퓨팅 장치(또는 DB 서버)는 시각화 도구 DB(예: 테이블 형태)에서, 각 시각화 도구(PCA, T-SNE, UMAP, Autoencoder-based 시각화 등)에 대한 메타정보를 관리할 수 있다. 이때, 컴퓨팅 장치는 데이터 특성과 도구 메타정보 간의 매칭 규칙을 바탕으로 SQL 쿼리를 생성하여, "사용 가능한 시각화 도구 중 특정 조건(도메인, 데이터 크기, 군집 구조 중요도, 연산 시간 제약 등)을 만족하는 후보"를 검색할 수 있다.
또한, 컴퓨팅 장치는 추천 알고리즘 실행 및 결과 산출 단계(S1830)를 수행할 수 있다. 구체적으로, DB 검색 결과가 복수의 시각화 도구로 반환되는 경우, 컴퓨팅 장치는 점수(score) 계산 또는 가중치(weight) 기반의 알고리즘을 적용할 수 있다. 예컨대, 컴퓨팅 장치는 데이터 양이 매우 크면 PCA 또는 UMAP 선호도 증가하도록 설정되거나, 클러스터 정확도(세밀한 군집화)가 중요하면 t-SNE 또는 UMAP의 점수 증가하도록 설정되거나, 실시간 상호작용 필요 시, 연산 복잡도가 낮은 PCA 우선 고려하도록 설정될 수 있다.
이를 종합하여 최종 추천 우선순위가 결정되며, 예컨대 "t-SNE(1위), UMAP(2위), PCA(3위)"와 같은 목록이 사용자에게 제시될 수 있다.
또한, 컴퓨팅 장치는 사용자 요청에 따른 시각화 도구 최종 결정 단계(S1840)를 수행할 수 있다. 예컨대, 사용자에게는 "데이터 특성 및 우선순위 기준에 따라 t-SNE가 가장 적합하다"는 안내가 제공될 수 있고, 사용자는 추천 도구를 그대로 사용할지, 다른 도구를 선택할지를 결정할 수 있다.
본 개시의 일 실시예에 따르면, 사용자가 특정 도구를 수동으로 지정하더라도, 사전 추천 결과를 통해 주요 장단점(예: 시각화 품질 vs. 처리 시간)을 사전에 인지할 수 있다.
또한, 최종적으로 결정된 시각화 도구(예: 제1 시각화 도구)를 적용하면, 적어도 하나의 프로세서는 제1 벡터 셋에 해당 알고리즘을 실행하여 제1 데이터 이미지를 생성 및 제공할 수 있다. 필요에 따라 다른 후보 도구(제2, 제3 시각화 도구 등)를 순차적으로 적용해 복수의 데이터 이미지를 비교 또는 분석하는 것도 가능하다.
본 개시의 일 실시예에서, 컴퓨팅 장치는 데이터 셋의 종류(이미지, 텍스트, 구조화 데이터 등), 도메인(의료, SNS, 금융 등), 데이터 용량, 벡터 간 분포 특성(밀도, 분산, 군집 수 등), 처리 시간 제약 등이 "시각화 도구 추천" 알고리즘에 반영되어, 사용자가 보다 합리적으로 시각화 도구를 선택하게 유도할 수 있다. 나아가, 시각화 도구 DB를 SQL로 질의하여 사용 가능한 후보를 식별하고, 해당 후보들 간 우선순위를 산출함으로써, 사용자는 고차원 임베딩 공간을 효율적으로 이해할 수 있는 시각화 결과를 얻을 수 있다.
이와 같이, 본 개시의 실시예에 따른 컴퓨팅 장치는 데이터 특성에 기반한 시각화 도구 추천 과정을 통해, PCA, T-SNE, UMAP 등의 기법들 각각이 갖는 장단점을 적절히 활용할 수 있게 함으로써, 데이터 분포나 군집 구조를 정확하고 직관적으로 파악할 수 있는 환경을 제공한다.
도 19는, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터 이미징을 기반으로 합성 데이터를 생성하는 방법을 설명하기 위한 도면이다.
도 19를 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 데이터 셋(dataset)에 대응되는 벡터 셋(VECOTOR SET)을 기초로 합성 데이터를 생성할 수 있다.
구체적으로, 적어도 하나의 프로세서는 적어도 하나의 인공지능 모델을 포함하는 렌즈(LENS)를 이용하여 데이터 셋(dataset)을 이미징함으로써 벡터 셋(VECTOR SET)을 획득하고, 적어도 하나의 시각화 도구(VISUALIZATION TOOL)를 이용하여 데이터 셋을 가시화함으로써 데이터 셋에 대응되는 데이터 이미지(IOD)를 가시화 공간에 나타낼 수 있다.
이때, 적어도 하나의 프로세서는 벡터 셋을 기초로 합성 데이터(Synthetic data)를 생성할 수 있다. 구체적으로, 적어도 하나의 프로세서는 미리 정해진 연산 조건이 설정된 연산기(CACULATOR)에 벡터 셋에 포함되는 적어도 하나의 벡터를 입력할 수 있다. 이 경우, 연산기는 적어도 하나의 벡터를 기초로 잠재 코드(LATENT CODE)를 출력할 수 있다. 또한, 적어도 하나의 프로세서는 출력된 잠재 코드(LATENT CODE)를 생성 모델(GENERATOR)에 입력할 수 있다. 이때, 잠재 코드는 잠재 공간(또는 임베딩 공간)에서 정의되는 임의의 벡터일 수 있다. 생성 모델은 데이터 생성을 위한 적어도 하나의 인공지능 모델(예: Diffusion 모델, GAN, VAE 등)을 포함할 수 있다. 생성 모델은 입력된 잠재 코드를 기초로 합성 데이터를 생성하여 출력할 수 있다.
적어도 하나의 프로세서는 생성된 합성 데이터를 시각화함으로써 합성 데이터와 기존 데이터 셋 사이의 연관관계를 시각적으로 표현할 수 있다. 구체적으로, 적어도 하나의 프로세서는 합성 데이터를 시각화 도구에 입력하여, 합성 데이터에 대응되는 합성 포인트(SP)를 데이터 이미지 상에 나타낼 수 있다. 또는, 적어도 하나의 프로세서는 합성 데이터를 렌즈에 입력하여 합성 데이터에 대응되는 합성 벡터(SYNTHETIC VECTOR, 미도시)를 획득하고, 합성 벡터를 시각화 도구에 입력하여 합성 데이터에 대응되는 합성 포인트를 데이터 이미지 상에 나타낼 수 있다.
기존 데이터 셋의 데이터 이미지에서 생성된 합성 데이터의 위치를 시각적으로 표현함으로써, 사용자가 생성된 데이터와 기존 데이터의 연관성을 직관적으로 인지하도록 유도할 수 있다.
도 20은, 다양한 실시예들에 따른, 컴퓨팅 장치가 사용자 입력을 기초로 합성 데이터를 생성하는 방법을 설명하기 위한 도면이다.
도 21은, 다양한 실시예들에 따른, 컴퓨팅 장치가 사용자 입력을 기초로 합성 데이터를 생성하는 방법의 일 예시를 도시한 도면이다.
도 20을 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 데이터 셋을 시각화하여 데이터 셋에 대응되는 제1 데이터 이미지를 제공할 수 있다(S2010). 예를 들어, 전술한 PCA, t-SNE, UMAP 등의 시각화 도구를 활용하여, 데이터 셋에 포함된 각 단위 데이터를 2차원 또는 3차원 공간으로 투영한 뒤, GUI(Graphic User Interface) 상에 복수의 데이터 포인트들을 포함하는 데이터 이미지(IOD)를 표시할 수 있다
또한, 적어도 하나의 프로세서는 제1 데이터 이미지 상에서의 제1 영역에 대한 사용자 입력을 수신할 수 있다(S2020). 를 들어, 도 21을 참조하면, 제1 데이터 이미지(IOD) 상의 특정 지점, 혹은 범위를 나타내는 제1 영역(R)에 대해 사용자의 클릭, 터치, 드래그 등의 입력을 인식할 수 있다. 여기서, 사용자 입력은 GUI 상에서 발생한 입력 이벤트로서, 마우스 좌클릭, 모바일 터치, 펜 드로잉 등 다양한 형태를 포괄한다.
다시 도 20을 참조하면, 적어도 하나의 프로세서는 제1 영역에 대응되는 잠재 코드를 결정할 수 있다(S2030). 잠재 코드는 합성 데이터 생성 시 모델 내부 임베딩이나 노이즈 벡터로 활용된다.
이때, 적어도 하나의 프로세서는 사용자 입력에 의해 지정된 제1 영역에 대해 피드백을 요청할 수 있다. 구체적으로, 적어도 하나의 프로세서는 사용자에게 제1 영역을 시각화하거나, 확인 메시지(예: "이 영역이 합성 데이터를 생성하는 대상 영역이 맞습니까?" 등)를 나타내는 GUI(예: 팝업창 또는 하이라이트 표시 등)를 제공할 수 있다. 이를 통해 사용자 의도에 맞지 않는 지점을 잘못 클릭했을 경우, 취소 또는 재설정 절차를 밟아 오류 방지가 가능하다.
이때, 적어도 하나의 프로세서는 사용자 입력을 기초로 데이터 생성의 조건을 정의할 수 있다. 구체적으로, 적어도 하나의 프로세서는 데이터 생성 모델에 입력될 잠재 코드를 결정하기 위한 데이터 생성 조건을 정의할 수 있다.
또한, 적어도 하나의 프로세서는 제1 영역의 속성에 따라 다양한 방식으로 잠재 코드를 결정할 수 있다. 예를 들어, 적어도 하나의 프로세서는 제1 영역에 포함되는 데이터 포인트의 유무 또는 데이터 포인트의 개수 등을 기초로 잠재 코드를 결정하는 알고리즘을 선택적으로 수행할 수 있다.
일 예로, 제1 영역에 데이터 포인트가 포함되지 않는 경우, 적어도 하나의 프로세서는 제1 영역에 인접한 적어도 하나의 포인트에 대응되는 적어도 하나의 벡터를 기초로 잠재 코드를 결정할 수 있다. 또는, 적어도 하나의 프로세서는 사용자 입력이 이루어진 제1 영역 및 이와 인접한 적어도 하나의 포인트를 포함하는 대상 영역을 도출해낸 뒤, 대상 영역에 포함되는 적어도 하나의 포인트에 대응되는 적어도 하나의 벡터를 기초로 잠재 코드를 결정할 수 있다.
다른 예로, 제1 영역에 적어도 하나의 데이터 포인트가 포함되는 경우, 적어도 하나의 프로세서는 제1 영역에 포함되는 적어도 하나의 데이터 포인트에 대응되는 적어도 하나의 벡터를 기초로 잠재 코드를 결정할 수 있다.
이때, 적어도 하나의 프로세서는 적어도 둘 이상의 벡터들을 기초로 잠재 코드를 결정할 수 있다. 구체적으로, 적어도 하나의 프로세서는 제1 영역에 인접한 적어도 하나의 포인트에 대응되는 적어도 둘 이상의 벡터를 식별하고, 식별된 적어도 둘 이상의 벡터들을 기초로 제1 영역에 대응되는 잠재 코드를 결정할 수 있다. 예컨대, 제1 영역 근방 포인트들에 대응되는 벡터들에 의해 보간된 벡터를 잠재 코드로 설정하면, 해당 영역의 전형적(centroid-like) 특성을 반영하는 합성 데이터가 생성될 수 있다.
또한, 제1 영역에 단일 포인트가 포함된 경우, 이를 타겟 벡터로 정의할 수 있으며, 이 타겟 벡터 주변의 인접 벡터들 간 관계(거리, 밀도, 라벨 등)를 기초로 잠재 코드를 생성할 수 있다.
또 다른 예로, 제1 영역에 복수의 데이터 포인트들의 클러스터(cluster)가 존재하는 경우, 적어도 하나의 프로세서는 클러스터의 속성을 기초로 잠재 코드를 결정할 수 있다. 구체적으로, 도 21을 참조하면, 사용자가 IOD 상의 특정 밀집 군집(cluster) 일부를 드래그-선택했다면, 그 군집에 속한 데이터 포인트들의 중심점(centroid) 또는 경계(boundary) 정보를 기초로 잠재 코드를 결정할 수 있다.
다시 도 20을 참조하면, 적어도 하나의 프로세서는 결정된 잠재 코드를 생성 모델에 제공하여, 합성 데이터를 생성할 수 있다(S2040). 또한, 적어도 하나의 프로세서는 생성된 합성 데이터를 데이터 처리 모델에 입력하여, 합성 데이터에 대응되는 합성 벡터를 획득할 수 있다(S2050). 또한, 적어도 하나의 프로세서는 합성 벡터를 시각화하여 합성 데이터에 대응되는 합성 포인트를 제1 데이터 이미지 상에 제공할 수 있다(S2060).
적어도 하나의 프로세서는 데이터 셋에 대한 진단 결과를 기초로 개선이 필요한 적어도 하나의 영역을 식별할 수 있다. 적어도 하나의 프로세서는 식별된 적어도 하나의 영역을 사용자 장치의 디스플레이를 통해 시각적으로 나타낼 수 있다.
적어도 하나의 프로세서는 식별된 적어도 하나의 영역을 인터랙션 가능한 상태로 활성화할 수 있다. 구체적으로, 적어도 하나의 프로세서는 개선이 필요한 적어도 하나의 영역에 대하여, 사용자에 의한 입력이 가능한 상태로 설정할 수 있다. 이 경우, 적어도 하나의 개선 필요 영역에 대한 사용자 입력을 기초로, 개선 필요 영역에 대한 데이터 개선 프로세스를 수행할 수 있다.
적어도 하나의 프로세서는 데이터 셋을 특정 차원의 잠재 공간에 임베딩한 벡터 셋을 기초로 데이터 셋에 포함되는 복수의 단위 데이터의 적어도 하나의 특성을 추출할 수 있다. 적어도 하나의 프로세서는 추출된 적어도 하나의 특성을 기초로 개선이 필요한 데이터를 식별할 수 있고, 개선이 필요한 데이터에 대응되는 데이터 포인트 또는 데이터 포인트를 포함하는 특정 영역을 시각적으로 나타낼 수 있다.
도 21을 참조하면, 적어도 하나의 프로세서는 사용자의 프롬프트 입력을 더 반영하여 합성 데이터를 생성할 수 있다. 구체적으로, 적어도 하나의 프로세서는 GUI 상에 마련된 입력 인터페이스(예: 텍스트 필드, 음성 입력, 슬라이더 등)를 통해, 사용자 프롬프트(PROMPT)를 수신할 수 있다. 사용자 프롬프트는 합성 데이터에 대한 주제, 스타일, 속성, 제약 조건 등을 자연어 또는 태그 기반으로 표현할 수 있다. 예를 들어, 사용자 프롬프트는 “파란색 배경의 고양이 이미지를 생성해줘" 또는 “강조된 원근감을 가진 실루엣 형식"같은 형태의 텍스트 지시 사항이 포함될 수 있다. 추가로, 사용자 프롬프트가 숫자형 파라미터나 범주형 태그(예: "리얼리스틱(realistic)", "카툰(cartoon)") 등을 지정하면, 해당 값을 프롬프트 내부 정보로 해석하여 모델에 전달할 수도 있다.
적어도 하나의 프로세서는 결정된 잠재 코드와, 사용자로부터 입력된 프롬프트를 조합하여, 생성 모델(예: GAN, VAE, Diffusion Model 등)에 대한 조건(condition)으로 설정할 수 있다. 예컨대, 잠재 코드는 이미지 생성을 위한 기본 latent vector로 사용되며, 프롬프트는 텍스트 임베딩(예: CLIP, BERT 기반 인코더) 또는 조건 레이어(conditional layer)로 변환되어 모델에 입력될 수 있다. 이러한 구조는, 예컨대 "text-to-image" 형태의 디퓨전 모델(Stable Diffusion 등)에서, 텍스트 임베딩을 통해 합성 과정 전반을 조건화하는 것과 유사하게 구현될 수 있다.
이 경우, 생성 모델(GENERATOR)은 잠재 코드에 기초하여 기본 형상이나 특성 분포를 결정하고, 프롬프트(텍스트, 태그 또는 파라미터 등)에 의해 주제, 스타일 또는 세부 속성을 추가적으로 반영하며, 최종적인 합성 데이터를 출력할 수 있다. 예를 들어, 잠재 코드가 "두 벡터 사이의 보간" 결과라면, 이미 중간적 속성(예: 두 클래스 혼합)이나 공간적 특성을 반영하고, 프롬프트가 "고양이 + 파란색 배경 + 카툰 스타일" 등으로 주어졌다면, 그 스타일 또는 주제 요소가 모델 내부의 컨디셔닝(Conditioning) 경로를 통해 반영되어, 해당 특성을 가진 합성 이미지가 출력될 수 있다.
결정된 잠재 코드와 함께 GUI를 통해 입력 인터페이스에 입력된 프롬프트(PROMPT)를 데이터 생성의 제약으로 설정할 수 있다. 적어도 하나의 프로세서는 잠재 코드 및 프롬프트를 생성 모델에 입력할 수 있고, 생성 모델은 잠재 코드에 기초하여, 프롬프트에 의해 지시된 사항에 의해 조건화된 합성 데이터를 생성할 수 있다.
또한, 적어도 하나의 프로세서는 텍스트 임베딩 모듈(CLIP, BERT, GPT 등) 또는 조건화 네트워크(조건 레이어, cross-attention 등)를 기초로, 사용자 프롬프트를 벡터 형태로 변환한 뒤, 생성 모델의 히든 레이어(hidden layer) 와 교차 연결(cross-attention)하거나 조건 입력(conditional input)으로 주입할 수 있다.
또한, 적어도 하나의 프로세서는 잠재 코드와 텍스트 임베딩을 결합(예: concat, add, attention)하는 로직을 기초로 매 단계(예: diffusion step, GAN upsampling step 등)에서 사용자 요청 사항이 반영되도록 구현될 수 있다.
도 21에서 도시된 일 실시예에 따르면, 사용자는 합성 결과물을 확인한 뒤, 프롬프트를 수정하거나, 잠재 코드를 재설정(보간 파라미터, 추가 벡터 선택 등)함으로써 반복적(refinement) 합성이 가능하다.
본 개시의 실시예에 의하면, 단순히 복수 벡터의 보간으로 얻어진 잠재 코드만을 사용했을 때보다, 사용자가 원하는 조건을 즉각 반영함으로써, 맞춤형(customized) 합성 데이터를 쉽게 얻을 수 있다.
또한, 본 개시의 실시예에 의하면, 잠재 코드가 이미 데이터 셋의 내재적 특성을 반영하므로, 프롬프트 기반 생성이 현실적인(distribution-preserving) 결과물을 생성하는 데 유리하다.
또한, 본 개시의 실시예에 의하면, 프롬프트 변경을 통해 여러 시나리오(적대 예시, 예술적 표현 등)를 효율적으로 시험함으로써, 인공지능 모델을 이용한 창의적 활용(예: 콘텐츠 제작, 데이터 보강, 시뮬레이션 등)이 용이해진다.
또한, 적어도 하나의 프로세서는 사용자가 지정한 제1 영역으로부터 프롬프트(Prompt)를 자동으로 추출(Extraction)하거나 생성(Generation)함으로써, 해당 프롬프트를 합성 데이터 생성 시의 조건으로 활용할 수 있다. 이를 통해, 사용자가 별도의 텍스트를 직접 입력하지 않아도, 데이터의 내재적 속성이나 메타정보를 반영한 합성 결과물을 얻을 수 있다.
예를 들어, 적어도 하나의 프로세서는 제1 영역에 속하거나 인접한 데이터 포인트들에 대해, 라벨 정보, 클래스/카테고리, 메타데이터(예: 시간, 위치, ID 등), 시각적 특징(이미지인 경우) 등을 분석하여, 자동으로 프롬프트를 생성할 수 있다. 구체적으로, "고양이 얼굴이 모여 있는 영역"이란 점이 추론되면, 적어도 하나의 프로세서는"cat face cluster"와 같은 간단한 문구 또는 "Multiple cat faces in a close-up shot"과 같은 좀 더 구체적인 텍스트를 생성할 수 있다.
적어도 하나의 프로세서는 이미지 캡셔닝(captioning) 기술(예: 비전-언어 모델, OCR(문자 인식) 등)을 이용하여 프롬프트를 생성할 수 있으며, 구조화된 데이터에 대해서는 카테고리명 또는 주요 속성명 등을 자동 변환해 문장 형태로 만들 수도 있다.
구체적인 예로, 적어도 하나의 프로세서는 이미지 캡셔닝 모델(예: CNN+LSTM 구조, Transformer 기반 시각-언어 모델 등)을 사용하여, 제1 영역의 이미지를 입력으로 넣고 문장 형태의 설명(캡션)을 자동으로 산출할 수 있다. 사용자에게는 "추출된 문구를 확인 및 편집"할 수 있는 GUI 인터페이스가 제공될 수 있으며, 사용자는 원하는 대로 텍스트를 수동 수정한 뒤, 최종 프롬프트로 채택할 수 있다.
적어도 하나의 프로세서는, 결정된 잠재 코드 또는 기타 벡터 보간 결과와, 자동 생성된 프롬프트를 결합하여 생성 모델(GAN, VAE, Diffusion 등)에 입력할 수 있다.
이 경우, 생성 모델은 잠재 코드로부터 추론된 데이터 분포를 반영하면서도, 사용자 영역에서 추출된 콘텍스트나 사물(객체) 정보, 속성 등을 반영해 조건화(conditional)된 합성 데이터를 출력할 수 있다.
이처럼 본 개시의 실시예에 따른 컴퓨팅 장치는 잠재 코드와 더불어 언어적 제약을 생성하므로 "차원의 확장"(기존 데이터로는 획득 불가능한 속성의 데이터가 생성)이 가능하다. 예를 들어, 컴퓨팅 장치는 N차원의 잠재 공간 상에 벡터를 출력하는 데이터 렌즈에 대하여, 언어적 제약을 기반으로 생성된 합성 데이터를 추가 학습함으로써 N+3 차원의 데이터 렌즈를 구축할 수 있다.
본 개시의 실시예에 의하면, 자동 프롬프트 생성 기능을 통해, 사용자 입력이 간소화되고 직관적인 합성 데이터 생성 흐름이 보장된다.
또한, 본 개시의 실시예에 의하면, 사용자 지정 영역의 메타정보, 시각 정보, 라벨 등을 "텍스트 조건"으로 전환함으로써, 데이터의 내재적 의미가 자연스럽게 반영된다.
또한, 본 개시의 실시예에 의하면, 캡셔닝이나 OCR 등 추가 모델을 연계하여, 고도화된 객체 인식이나 상황 설명을 수행함으로써, 다양한 도메인(예: 의료 영상, 위성 영상, 문서 처리 등)에 유연하게 적용 가능하다.
따라서, 본 개시의 실시예에 의하면, 사용자 지정 영역으로부터 프롬프트를 자동으로 추출하거나 생성하여, 기존에 결정된 잠재 코드와 함께 생성 모델에 입력함으로써, 해당 영역의 속성이 반영된 조건화(conditional) 합성 데이터를 얻을 수 있도록 한다.
본 개시의 일 실시예에 따른 컴퓨팅 장치는 사용자 입력을 기초로 데이터 셋에 포함되는 적어도 일부의 데이터를 제거하는 인터랙션 동작을 수행할 수 있다.
도 22는, 다양한 실시예들에 따른, 컴퓨팅 장치가 사용자 입력을 기초로 데이터를 제거하는 방법을 설명하기 위한 도면이다.
도 22를 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 시각화된 데이터 이미지 상에서의 적어도 하나의 포인트를 포함하는 제1 영역에 대한 사용자 입력을 수신할 수 있다(S2210). 예컨대, 제1 영역은 밀집된 영역(클러스터 내 데이터 밀도가 매우 높은 부분)일 수도 있고, 이상치가 다수 포함된 영역 등 사용자가 제거를 희망하는 영역일 수도 있다. 적어도 하나의 프로세서는 데이터의 제거가 필요한 영역(예: 밀집 영역 또는 이상치 영역)에 대한 정보를 사용자에게 제공할 수 있고, 사용자는 해당 영영에 대해 입력을 제공할 수 있다. 제1 영역은 데이터의 특성이 미리 정해진 조건을 만족하는 적어도 하나의 클러스터를 포함할 수 있다.
또한, 적어도 하나의 프로세서는 제1 영역에 포함되는 복수의 포인트들에 대응되는 데이터 중 적어도 일부를 제거할 수 있다(S2220). 여기서 "적어도 일부"란, 사용자가 지정한 전체 또는 특정 조건(예: 이상치 필터, 특정 라벨 등)에 해당하는 데이터만을 말한다.
적어도 하나의 프로세서는 제1 영역의 속성을 기초로 데이터 제거 알고리즘을 선택적으로 수행할 수 있다. 예를 들어, 적어도 하나의 프로세서는 제1 영역에 포함된 데이터 클러스터가 너무 과도하게 밀집되어 불균형을 일으킨다고 판단하는 경우, 해당 영역 내 모든 데이터 포인트를 언더샘플링(undersampling)하거나, 임의 비율로 제거할 수 있다. 또한, 예를 들어, 적어도 하나의 프로세서는 제1 영역에 있는 이상치들만 제거하도록 결정함으로써, 해당 영역의 포인트 중 사전에 정의된 이상치 판별 기준(거리, 밀도, 라벨 불일치 등)을 만족하는 데이터만 선택적으로 제거할 수도 있다. 이 경우, 적어도 하나의 프로세서는 사용자 지정 영역 내에서도 특정 라벨, 특정 통계 범위를 벗어난 포인트만 제거하도록 세분화 옵션(필터 조건 등)을 제공할 수 있다. (예: "이 영역에 속하면서, 라벨이 0인 포인트만 제거 등)
또한, 적어도 하나의 프로세서는 적어도 일부의 데이터가 제거된 데이터 이미지를 제공할 수 있다(S2230). 사용자는 새롭게 시각화된 분포를 확인하면서, 데이터 불균형 개선 또는 잡음(noise) 감소 효과를 직관적으로 파악할 수 있다. 필요 시, 삭제 이력을 남기거나, 복원 기능(Undo) 등을 제공하여, 사용자 실수에 따른 데이터 손실을 방지할 수도 있다.
또한, 적어도 하나의 프로세서는 제거가 완료된 데이터 셋을 대상으로 리이미징 및 진단(클러스터 분석) 등을 수행하여, 데이터 품질이 얼마나 향상됐는지, 모델 학습에 이점이 있는지 등을 다시 확인할 수 있다.
본 개시의 실시예에 의하면, 사용자 인터랙션을 통해 직접적인 제거가 가능해지므로, 기존 자동화된 알고리즘이 놓칠 수 있는 세밀한 도메인 지식을 반영할 수 있다.
또한, 본 개시의 실시예에 의하면, 밀집 영역의 언더샘플링 또는 이상치 제거 등을 통해, 데이터 불균형 해소나 학습 시 과적합 방지 등에 실질적인 이점을 제공한다.
또한, 본 개시의 실시예에 의하면, 시각화(2D 또는 3D)된 공간에서 즉각적인 피드백을 얻을 수 있어, "어느 정도 데이터가 사라졌는지", "어떻게 분포가 변화했는지" 등 후속 조치를 빠르게 결정할 수 있다.
따라서, 본 개시의 실시예에 의하면, 사용자가 지정한 영역의 데이터(특히 클러스터 밀집 구역, 이상치 다발 구역 등)를 상호작용(interactive) 방식으로 제거할 수 있도록 함으로써, 데이터 품질 관리를 위한 직관적이고 유연한 도구를 제공할 수 있는 것이다.
도 23은, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터 개선 요청에 대응하여 데이터 개선을 수행하고, 이에 대한 시각적인 인터랙션을 제공하는 방법을 설명하기 위한 도면이다.
도 23을 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 데이터 개선 요청을 수신할 수 있다(S2310). 구체적으로, 적어도 하나의 프로세서는 데이터 개선을 지시하는 적어도 하나의 GUI를 통해 사용자 입력을 수신함으로써 데이터 개선 요청을 인지할 수 있다. 예컨대, 사용자 GUI에서 데이터를 제1 방식으로 개선하기 위한 버튼(예: "데이터 불균형 해소")이 선택되거나, 제2 방식으로 개선하기 위한 버튼(예: "노이즈 영역 제거")이 선택되는 경우, 적어도 하나의 프로세서는 해당 입력을 통해 개선 요청이 발생했음을 알 수 있다. 또한, 특정 클래스(혹은 속성)가 부족하다는 것을 사용자가 인지하여 "해당 클래스를 더 많이 생성해달라"는 식으로 요청할 수도 있다.
또한, 적어도 하나의 프로세서는 데이터 개선이 필요한 적어도 하나의 영역을 시각적으로 표현할 수 있다(S2320). 구체적으로, 적어도 하나의 프로세서는 데이터 셋에 대응되는 벡터 셋을 기초로 데이터 셋에 대한 특성을 획득하고, 데이터 셋에 대한 특성을 기초로 데이터 개선이 필요한 영역을 탐지할 수 있다.
예를 들어, 적어도 하나의 프로세서는 데이터 셋의 밀도를 분석함으로써, 특정 클래스가 지나치게 적거나 혹은 일부 구역(클러스터)이 과도하게 밀집된 영역을 식별할 수 있다. 또한, 예를 들어, 적어도 하나의 프로세서는 데이터 셋의 편향성을 분석함으로써, 클래스 분포가 불균형해서 개선이 필요한 경우, 특정 클래스의 분포가 부족한 지점 또는 특정 클래스가 과밀한 지점 등을 자동으로 식별할 수 있다. 또한, 예를 들어, 적어도 하나의 프로세서는 이상치 분석을 함으로써, 이상치가 다량 존재하는 부분(노이즈 데이터가 집중된 구역)을 감지할 수 있다.
적어도 하나의 프로세서는 이렇게 탐지된 "개선이 필요한 영역"을 시각화된 데이터 이미지(IOD) 상에서 하이라이트하거나 테두리 표시 등으로 강조해 사용자에게 알릴 수 있다. 이 경우, 적어도 하나의 프로세서는 GUI 상에서 어떤 기준으로 영역이 선정되었는지 등에 대한 인터랙티브 가이드(예: 툴팁 등)을 제공할 수 있다. 사용자는 이 표시를 확인해, 데이터 개선이 필요한 근거(예: "어떤 부분에서 데이터가 부족한지, 과잉한지 또는 노이즈가 발생했는지" 등)를 직관적으로 파악할 수 있다.
또한, 적어도 하나의 프로세서는 적어도 하나의 영역 각각에 대하여, 데이터를 생성하거나 제거함으로써 데이터 개선을 수행할 수 있다(S2330). 예를 들어, 적어도 하나의 프로세서는 특정 클래스의 데이터가 부족하다고 판단되는 데이터 이미지 상의 제1 영역에 대하여, 도 19에 따른 데이터 생성 방법을 기초로 합성 데이터를 생성할 수 있다. 또한, 예를 들어, 적어도 하나의 프로세서는 데이터가 과밀하거나, 이상치가 포함되어있는 제2 영역에 대하여, 도 22에 따른 데이터 제거 방법을 기초로 적어도 일부의 데이터를 제거할 수 있다. 이때, 사용자는 자동 제안에 대해 다양한 세부 옵션들(예: 실행, 취소, 세부 옵션 조정 등) 중 적어도 하나를 선택할 수 있다.
또한, 적어도 하나의 프로세서는 데이터 개선이 수행된 결과에 대응되는 개선된 데이터 이미지를 시각화하여 제공할 수 있다(S2340). 이를 통해 사용자는 데이터 분포가 어떻게 변했는지를 한눈에 확인하고, 불균형이 어느 정도 해소되었는지, 이상치가 얼마나 줄었는지 등을 직관적으로 파악할 수 있다. 선택적으로, 적어도 하나의 프로세서는 사용자에게 바로 이전 상태와의 비교 뷰(before/after 시각화) 또는 진단 리포트를 함께 제공할 수도 있다.
또한, 사용자는 시각화된 개선 영역 중 일부를 선택 해제하거나, 생성 또는 제거의 범위를 더 확장하도록 지시함으로써 수동 보정을 수행할 수 있다. 또한, 적어도 하나의 프로세서는 개선 결과를 기초로 다시 필요한 부분을 자동 식별하여, 여러 차례 반복적으로 데이터 품질을 높이는 루프를 수행할 수 있다.
도 24는, 다양한 실시예들에 따른, 컴퓨팅 장치가 스냅 샷(snapshot) 기능을 포함하는 데이터 처리 방법이 구현된 컴퓨팅 장치의 일 예시를 도시한 도면이다.
도 24를 참조하면, 컴퓨팅 장치(2400)는 데이터 이미지 상의 특정 장면을 기초로 스냅 샷(snapshot) 정보를 획득하기 위한 복수의 구성들을 포함할 수 있다. 여기서, 복수의 구성들은 특정 동작을 수행하기 위하여 임의적으로 구분한 것으로, 실제로 물리적으로 분리된 구성일 수도 있으나, 하나의 소프트웨어 프로그램 상에서 구현되는 다양한 동작들을 기반으로 분리된 구성일 수도 있다.
구체적으로, 컴퓨팅 장치(2400)는 데이터 이미지 또는 데이터 이미지에 대응되는 벡터 셋을 스크리닝하기 위한 스크리너(SCREENER), 스냅샷을 촬상할 이벤트의 발생 여부를 감지하기 위한 이벤트 디텍터(EVENT DETECTOR), 스냅샷을 촬영하기 위한 캡쳐부(CAPTURE), 촬상된 장면을 보정하기 위한 보정부(MODIFIER) 및 스냅샷에 대한 다양한 정보를 생성하기 위한 생성기(GENERATOR)를 포함할 수 있다.
스크리너는 벡터 셋(VECTOR SET) 또는 데이터 이미지를 스크리닝할 수 있다. 구체적으로, 스크리너는 벡터 셋(VECTOR SET) 또는 데이터 이미지를 기반으로 데이터 셋의 특성 값(예: 밀도, 편향, 균질도, 이상치 여부 등)을 측정하기 위한 적어도 하나의 메트릭이 설정될 수 있다. 예를 들어, 스크리너는 제1 메트릭을 기반으로 2D/3D 데이터 이미지에서 기초적 특성(밀집도, 이상치 비율 등)을 분석하거나, 제2 메트릭을 기반으로 고차원 벡터 셋에서 고차원 분포 특성을 측정할 수 있다.
이벤트 디텍터는 스크리닝 도중에 발생하는 이벤트를 감지할 수 있다. 구체적으로, 이벤트 디텍터는 스크리너에 의한 측정 결과 값을 모니터링할 수 있고, 모니터링된 결과 값이 미리 정해진 조건을 만족하면 이벤트가 발생한 것으로 감지할 수 있다. 이때, 이벤트 디텍터는 이벤트가 발생한 데이터에 대한 정보(예: 데이터 항목, 해당 데이터에 대응되는 벡터 값, 해당 데이터에 대응되는 데이터 포인트의 좌표 등)를 저장할 수 있다. 예를 들어, 이벤트 디텍터는 스크리너 결과, 데이터의 밀집도가 임계값보다 큰 특정 영역 또는 편향 지표가 기준치 이상인 특정 영역 등이 감지되면, 이벤트 디텍터는 이를 "이벤트 발생"으로 인식하고, 신호를 생성할 수 있다.
생성기는 이벤트 감지부에 의해 감지된 이벤트에 대한 정보를 생성할 수 있다. 구체적으로, 생성기는 이벤트 디텍터가 감지한 이벤트 정보, 캡쳐된 장면의 위치 및 메타정보, 스크리너가 산출한 특성 값 등을 종합하여 태그 정보를 생성할 수 있다. 또는, 생성기는 스냅샷 관련 부가 정보(클래스 라벨, 시간, 사용자 ID 등)를 기록할 수 있다. 즉, 생성기는 미리 저장되거나, 이벤트 디텍터로부터 생성된 이벤트가 발생한 데이터에 대한 정보를 기초로 태그 정보를 생성할 수 있다. 생성기에 의해 생성된 태그 정보는, 예를 들어, "특성이 일반적이지 않은 영역(이상치 밀집)" 또는 "밀도가 과도하게 높아진 구간"과 같은 식별 라벨, 스냅샷 촬영 시점, 주요 특성 값 등을 포함할 수 있다.
이로써, 컴퓨팅 장치는 추후 사용자가 스냅샷을 불러올 때, 맥락(context)을 즉시 파악할 수 있게 유도할 수 있다.
캡쳐부는 데이터 이미지 상에서의 특정 장면을 촬상하여 저장할 수 있다. 예를 들어, 캡쳐부는 이벤트 디텍터로부터 이벤트 발생을 알리는 신호가 전송되거나, 혹은 사용자가 스냅샷을 직접 촬영하도록 명령하면, 데이터 이미지 상에서 이벤트에 대응되는 적어도 하나의 데이터 포인트를 포함하는 장면(scene)을 특정 시점(view point)에서 캡쳐(촬영)할 수 있다. 이때, "캡쳐"는 실제 GUI 화면(2D/3D 뷰) 혹은 내부 데이터 구조(벡터 셋 + 시각화 매핑 파라미터)의 상태를 이미지(또는 영상 프레임)로 저장하는 과정을 포함한다.
보정부는 촬상된 장면(스냅샷)에 대해 보정(modification)을 수행할 수 있다. 이를 통해 사용자에게 가시성이 높은 결과물을 제공할 수 있으며, 필요 시 영역을 시각적으로 강조하거나, 레이블을 추가하는 등의 그래픽 보정을 수행할 수도 있다.
예를 들어, 컴퓨팅 장치는 데이터 셋에 대응되는 벡터 셋(VECTOR SET)을 스크리너에 제공할 수 있고, 스크리너는 벡터 셋을 기초로 데이터 셋의 특성을 진단할 수 있다. 이벤트 디텍터는 스크리너에 의해 진단된 특성을 기초로 이벤트 발생 여부를 식별할 수 있다. 예를 들어, 이벤트 디텍터는 스크리너에 의해 진단된 특성이 미리 정해진 조건을 만족하는 경우, 이벤트가 발생을 알리는 신호를 생성할 수 있다. 이 경우, 캡쳐부는 이벤트가 발생한 것을 알리는 신호를 기반으로 이벤트에 대응되는 스냅샷을 촬상할 수 있다. 구체적으로, 캡쳐부는 이벤트에 연관되는 적어도 하나의 벡터에 대응되는 적어도 하나의 데이터 포인트를 포함하는 장면을 특정 시점에서 촬상함으로써 스냅샷을 생성할 수 있다. 또한, 이와 병렬적으로 생성기는 발생한 이벤트에 대한 정보 또는 이벤트가 발생한 데이터 이미지 상의 위치에 대한 정보 등을 기초로 태그 정보를 생성할 수 있다. 또한, 보정부는 생성된 스냅샷을 미리 정해진 방식으로 보정(예: 밝기 조정, 명도 조절 등)할 수 있다.
도 25는, 다양한 실시예들에 따른, 컴퓨팅 장치가 데이터 셋을 스크리닝하여 스냅 샷 정보를 생성하는 방법을 설명하기 위한 도면이다.
도 25를 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 적어도 하나의 메트릭이 설정된 스크리너를 이용하여, 벡터 셋에 포함되는 적어도 하나의 벡터를 기초로 적어도 하나의 특성 값을 산출함으로써 벡터 셋을 스크리닝할 수 있다(S2510). 구체적으로, 적어도 하나의 프로세서는 데이터의 특성을 진단할 수 있고, 예컨대 데이터의 밀도, 편향, 균질도 등의 다양한 지표를 측정할 수 있다.
본 개시에서 스크리닝은 데이터 이미지(2D/3D) 또는 벡터 셋 상의 특정 출발점(좌표 위치)을 기준으로 특정 경로를 따라 진행될 수 있다. 예컨대, 적어도 하나의 프로세서는 미리 결정된 위치 또는 결정된 참조 지점(reference point)으로부터 일정 범위를 순차적으로 스캔(scan)하여 분석함으로써, 데이터 분포의 밀도, 편향, 균질도 등을 평가할 수 있다.
또한, 적어도 하나의 프로세서는 스크리닝 대상에 따라 구분된 복수의 수준들의 스크리닝을 수행할 수 있다. 구체적으로, 적어도 하나의 프로세서는 데이터 이미지를 스크리닝하기 위한 제1 수준의 스크리닝 또는 벡터 셋을 스크리닝하기 위한 제2 수준의 스크리닝 중 적어도 하나를 수행할 수 있다. 제1 수준의 스크리닝은 2차원 또는 3차원의 데이터 이미지를 기반으로 2차원 또는 3차원 데이터의 연산을 수행하는 반면, 제2 수준의 스크리능은 고차원의 벡터 셋을 기반으로 고차원 연산을 수행한다는 차이가 있다.
예를 들어, 제1 수준의 스크리닝이 지시되는 경우, 적어도 하나의 프로세서는 데이터 이미지에 포함되는 적어도 하나의 데이터 포인트를 기초로 연산(2차원 또는 3차원 데이터 연산)하여 데이터의 특성을 측정하기 위한 제1 메트릭을 적용할 수 있다.
또한, 예를 들어, 제2 수준의 스크리닝 지시되는 경우, 적어도 하나의 프로세서는 벡터 셋에 포함되는 적어도 하나의 벡터를 기초로 연산 (고차원 연산)하여 데이터의 특성을 측정하기 위한 제2 메트릭을 적용할 수 있다. 예를 들어, 적어도 하나의 프로세서는 스크리닝 대상 벡터들이 실제 고차원 매니폴드상에서 어떤 구조(클러스터, 이상치 등)를 형성하는지, 편향 지표가 얼마나 높은지 등을 진단할 수 있다.
또한, 이에 한정되지 않고, 적어도 하나의 프로세서는 스크리닝하여 진단하려는 특성의 종류에 따라 구분된 복수의 수준들의 스크리닝을 수행할 수 있다. 구체적으로, 적어도 하나의 프로세서는 제1 종류의 특성(예: 결측치, 데이터 통계 등)을 측정하기 위한 메트릭이 설정된 제1 수준의 스크리닝 또는 제2 종류의 특성(예: 밀도 등)을 측정하기 위한 메트릭이 설정된 제2 수준의 스크리닝 또는 제3 종류의 특성(예: 클래스 분포 등)을 측정하기 위한 메트릭이 설정된 제3수준의 스크리닝 중 적어도 하나를 수행할 수 있다.
또한, 적어도 하나의 프로세서는 적어도 하나의 특성 값이 미리 정해진 조건을 만족하는 적어도 하나의 벡터를 식별할 수 있다(S2520). 이를 통해, 적어도 하나의 프로세서는 제1 특성(예: 밀도)이 제1 조건(예: 과밀 지역 등)을 만족하는 적어도 하나의 벡터 또는 제2 특성(예: 편향 지표)이 제2 조건(예: 편향도가 기준 이상)을 만족하는 적어도 하나의 벡터를 식별할 수 있다.
또한, 적어도 하나의 프로세서는 적어도 하나의 시각화 도구를 이용하여, 상기 벡터 셋을 처리함으로써 상기 데이터 셋을 2차원 또는 3차원으로 나타내는 복수의 데이터 포인트들을 포함하는 데이터 이미지를 획득할 수 있다(S2530).
또한, 적어도 하나의 프로세서는 데이터 이미지 상에서 적어도 하나의 벡터에 대응되는 적어도 하나의 데이터 포인트를 포함하는 타겟 영역을 결정하여, 타겟 영역에 연관되는 태그 정보를 생성할 수 있다(S2540).
구체적으로, 적어도 하나의 프로세서는 미리 정해진 조건을 만족하는 적어도 하나의 벡터에 대응되는 적어도 하나의 데이터 포인트를 식별할 수 있고, 적어도 하나의 데이터 포인트를 기준점으로 하는 소정의 영역을 결정함으로써 타겟 영역을 결정할 수 있다. 예를 들어, 적어도 하나의 프로세서는 적어도 하나의 데이터 포인트를 중심으로 소정의 반경을 가지는 영역을 타겟 영역으로 결정할 수 있다.
또한, 적어도 하나의 프로세서는 결정된 타겟 영역에 연관되는 태그 정보를 생성할 수 있다. 이때, 태그 정보는 타겟 영역에 연관되는 진단 결과를 반영할 수 있다. 예를 들어, 적어도 하나의 프로세서는 이벤트 디텍터에 의해 감지된 이벤트에 대한 정보 및 스크리너에 의해 진단된 특성 값을 기초로 태그 정보를 생성할 수 있으나, 이에 한정되지 않는다. 예를 들어, 태그 정보는 발견 배경(예: 어떤 메트릭 조건을 달성했는지), 시간, 분석자 ID, 이벤트 종류(예: 특이 구간, 공백 구간, 과밀 구간 등) 등을 포함할 수 있으나, 이에 한정되지 않는다.
또한, 적어도 하나의 프로세서는 타겟 영역을 포함하는 제1 장면 - 제1 장면은 특정 시점(view-point)로부터 타겟 영역을 캡쳐함으로써 생성됨- 및 태그 정보를 포함하는 스냅 샷 정보를 저장할 수 있다(S2550). 구체적으로, 적어도 하나의 프로세서는 타겟 영역이 시각적으로 가장 명확하게 보이는 뷰를 캡쳐함으로써 제1 장면을 획득할 수 있다. 적어도 하나의 프로세서는 타겟 영역을 캡쳐하기 위한 가상의 카메라의 시점을 조정해가며 촬상된 복수의 장면들 중 최적의 장면을 선정함으로써 제1 장면을 획득할 수 있다. 또는, 컴퓨팅 장치가 자동으로 타겟 영역이 가장 잘 드러나는 카메라 각도/배율을 계산하여, 시점(view point)를 조정한 뒤, 스냅샷을 촬영할 수도 있다(예: 3D 점 구름 회전).
경우에 따라, 적어도 하나의 프로세서는 다수의 후보 장면들(ex: 상단뷰사이드뷰 또는 45도 각도 뷰)을 비교하여 사용자에게 제공할 수 있고, 사용자에 의해 선택된 장면을 제1 장면으로 결정할 수 있다. 또한, 촬영된 스냅샷(제1 장면)은 앞서 설명한 보정부(MODIFIER)에 의해 보정(밝기, 명도, 강조 표시 등)될 수 있으며, 이렇게 보정된 최종 장면과 태그 정보를 포함하는 스냅샷 정보가 DB나 파일 형태로 저장될 수 있다.
다시 도 24를 참조하면, 컴퓨팅 장치(2400)는 GUI에 기반한 입력 인터페이스를 통해 수신된 사용자 입력을 기반으로 스냅샷 정보를 생성할 수 있다. 구체적으로, 컴퓨팅 장치(2400)는 사용자 입력을 기초로 스냅샷의 생성을 지시하는 신호를 캡쳐부에 전달하고, 캡쳐부는 데이터 이미지(IOD)의 적어도 일부를 촬상할 수 있다.
도 26은, 다양한 실시예들에 따른, 컴퓨팅 장치가 사용자 입력을 기반으로 스냅샷 정보를 제공하는 방법을 설명하기 위한 도면이다.
도 27은 컴퓨팅 장치에 의해 제공되는 화면의 일 예시이다.
도 26을 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 제1 데이터 셋에 대응되는 제1 데이터 이미지를 제1 뷰 포트를 통해 제공할 수 있다(S2610).
예를 들어, 도 27을 참조하면, 컴퓨팅 장치는 제1 뷰 포트(2710)를 통해 제1 데이터 셋에 대응되는 제1 데이터 이미지를 제공할 수 있다. 이때, 컴퓨팅 장치는 제1 데이터 이미지에 포함되는 복수의 데이터 포인트들 중 적어도 일부에 대한 프리뷰 정보를 함께 제공할 수 있다.
예를 들어, 컴퓨팅 장치는 데이터 포인트에 대응되는 실제 데이터를 함께 출력함로써 프리뷰 정보를 제공할 수 있다. 구체적으로, 컴퓨팅 장치는 특정 데이터 포인트를 클릭하면 해당 포인트에 대응되는 실제 데이터(예: 원본 이미지, 텍스트 내용, 통계값 등)를 미리 볼 수 있도록 프리뷰 정보를 구성할 수 있다. 이를 통해, 사용자는 데이터 이미지 상 포인트가 나타내는 의미를 즉시 확인할 수 있다.
다시 도 26을 참조하면, 적어도 하나의 프로세서는 제1 GUI를 통해 수신된 사용자 입력에 대응하여, 제1 뷰 포트를 통해 제공 중인 제1 데이터 이미지에 대한 제1 장면을 포함하는 제1 스냅 샷 정보를 생성하여 제2 뷰 포트를 통해 제공할 수 있다(S2620). 구체적으로, 적어도 하나의 프로세서는 제1 데이터 이미지 상에서 적어도 하나의 데이터 포인트를 포함하는 제1 장면을 캡쳐함으로써 제1 스냅 샷 정보를 생성할 수 있다.
예를 들어, 도 27을 참조하면, 컴퓨팅 장치는 제1 GUI(2720)에 대한 사용자 입력을 수신하여, 제1 데이터 이미지에 대한 제1 장면을 포함하는 제1 스냅 샷 정보를 제2 뷰 포트(2730)를 통해 제공할 수 있다. 구체적으로, 컴퓨팅 장치는 제1 GUI(2720)를 표시하여, 스냅 샷 촬영을 지시하는 버튼이나 특정 제스처(드래그 또는 박스 선택 등)를 이용해 사용자가 원하는 장면을 지정하도록 유도할 수 있다.
또한, 사용자가 "스냅 샷"을 요청하면, 컴퓨팅 장치는 캡쳐부에 이에 대한 지시 신호를 전달할 수 있다. 이 경우, 캡쳐부는 제1 데이터 이미지가 표시된 제1 뷰 포트(2710)의 현재 시점(view point), 확대 비율, 또는 사용자의 범위 지정 등을 종합하여 제1 장면을 촬상할 수 있다. 이후, 컴퓨팅 장치는 완성된 제1 스냅 샷 정보를 제2 뷰 포트(2730)를 통해 사용자에게 제공할 수 있다. 도 27의 예시에서, 제2 뷰 포트(2730)는 "스냅 샷 미리보기" 영역으로 사용될 수 있다.
또한, 적어도 하나의 프로세서는 사용자로부터 메모 입력을 수신하기 위한 인터페이스를 제2 뷰 포트를 통해 더 제공할 수 있다. 예를 들어, 컴퓨팅 장치는 제2 뷰 포트(2730) 내에 "메모 입력 창"(미도시)을 활성화할 수 있다. 사용자는 스냅 샷을 확인하면서, 문자열 코멘트(메모) (예: "여기서는 Class A가 과도하게 밀집", "이상치 의심" 등)를 직접 작성할 수 있다. 또한, 적어도 하나의 프로세서는 사용자 메모를 스냅 샷 정보와 매핑하여 저장함으로써, 추후 해당 스냅 샷을 재확인할 때 사용자에 의해 작성된 메모 정보를 함께 불러올 수 있다.
또한, 적어도 하나의 프로세서는 제1 스냅 샷 정보에 대한 태그 정보를 생성하기 위한 인터페이스를 제2 뷰 포트를 통해 더 제공할 수 있다. 구체적으로, 적어도 하나의 프로세서는 사용자에 의해 제1 스냅 샷 정보가 생성된 경우, 제2 뷰 포트의 적어도 일부를 통해 복수의 태그들 중 적어도 하나를 선택할 것을 요청하는 인터페이스를 활성화할 수 있다. 이 경우, 적어도 하나의 프로세서는 촬상된 장면을 기초로 추천 태그를 결정할 수 있고, 추천 태그에 대한 정보를 사용자에게 제공할 수 있다. 또한, 적어도 하나의 프로세서는 선택된 태그를 기초로 태그 정보를 생성하여 저장할 수 있다. 적어도 하나의 프로세서는 태그 종류에 따라 스냅 샷 정보를 구조화하여 저장할 수 있다.
예를 들어, 컴퓨팅 장치는 제2 뷰 포트(2730) 일부 영역(예: 팝업, 사이드 패널)에 "태그 선택 메뉴"(미도시)를 표시할 수 있다. 적어도 하나의 프로세서는 촬영된 장면(해당 스냅 샷 내 데이터 분포나 포인트 속성 등)을 분석해, "편향", "밀집", "이상치", "라벨 불일치" 등 추천 태그를 자동 제시할 수도 있다. 사용자가 하나 이상의 태그를 선택하면, 프로세서는 이를 기반으로 최종 태그 정보(예: <"밀집", "우하단 클러스터", "Class A">)를 생성하여 저장할 수 있다.
다시 도 26을 참조하면, 적어도 하나의 프로세서는 제2 뷰 포트를 통해 제공되는 제2 GUI에 대한 사용자 입력을 기초로, 제1 스냅 샷 정보를 외부 통신 네트워크와 연결하기 위한 링크 정보를 생성할 수 있다(S2630).
예를 들어, 도 27을 참조하면, 적어도 하나의 프로세서는 제2 GUI(2740)를 제2 뷰포트(2720)를 통해 제공할 수 있다. 적어도 하나의 프로세서는 제2 GUI(2740)에 대한 사용자 입력을 기초로 링크 정보를 생성할 수 있다. 링크 정보는 제1 스냅 샷 정보를 외부 통신 네트워크를 통해 전달하기 위한 링크를 포함할 수 있다. 링크 정보는 외부 통신 네트워크(인터넷, 회사 내부망 등)를 통해 이 스냅 샷을 전달 및 공유하기 위한 접속 경로가 될 수 있으며, 수신자는 해당 링크를 통해 동일 장면(스냅 샷) 및 태그 정보 또는 메모 정보 등을 확인할 수 있다.
본 개시의 실시예에 의해, 사용자가 임의 시점/영역을 직접 결정하여 캡쳐하므로, 자동 캡쳐에서 놓칠 수 있는 맞춤형 시나리오(특정 군집만 크게 확대, 특정 라벨만 강조 등)를 효과적으로 지원한다.
또한, 본 개시의 실시예에 의해, 뷰 포트 전환(제1->제2)과 메모/태그 인터페이스를 통해, 사용자 코멘트나 분류 정보를 스냅 샷과 유기적으로 결합하여 저장함으로써, 향후 분석 협업이나 문서 보고 시 맥락을 더욱 풍부하게 제공한다.
또한, 본 개시의 실시예에 의해, 링크 생성 기능은, 스냅 샷 정보(장면, 태그, 메모 등)를 외부 팀원 또는 다른 시스템과 실시간으로 공유 및 피드백받을 수 있게 하며, 원격 협업 환경에서의 데이터 해석 효율을 높이고, 해당 솔루션에 대한 외부 노출을 유도하여 경제적 파급 효과를 가질 수 있다.
도 28은, 다양한 실시예들에 따른, 컴퓨팅 장치가 스냅 샷 정보를 재생하는 기능을 설명하기 위한 도면이다.
본 개시의 실시예에 따른 컴퓨팅 장치(또는 적어도 하나의 프로세서)는 스크리닝 경로를 따라 촬상된 복수의 스냅 샷을 순차적으로 저장하고, 사용자 재생 지시에 따라 이러한 스냅 샷들을 연속적으로 제공함으로써 영상 또는 슬라이드 형태로 재생할 수 있다.
도 28을 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 스크리닝 경로에 따라 획득된 복수의 스냅 샷 정보를 순차적으로 저장할 수 있다(S2810). 구체적으로, 컴퓨팅 장치에 포함된 스크리너(SCREENER)는 데이터 이미지(또는 고차원 벡터 셋)를 특정 경로로 탐색("스크리닝 경로"라 칭함)하면서, 각 지점에서 특정 특성이 감지되거나, 이벤트가 발생하거나, 혹은 일정 간격마다 스냅 샷을 촬상할 수 있다. 예컨대, "좌측 상단에서 우하단으로 점진적으로 이동하면서, 2D 뷰 상의 각 블록(영역)을 검사"하는 방식으로 스크리닝이 진행될 수 있다. 각 검사 지점에서 특성 값을 확인한 뒤, 임계값을 초과하면 스냅 샷을 촬영할 수 있다.
또한, 적어도 하나의 프로세서는 이렇게 촬상된 스냅 샷들을 획득 순서 또는 스크리닝 경로 상 순서에 따라 연결 정보(예: "Snap #1 -> #2 -> Snap #3 ...)와 함께 기록할 수 있다. 이때, 각각의 스냅 샷은 "시점(view point), 촬상 위치(스크리닝 경로 상 좌표 또는 고차원 매핑 정보), 이벤트 원인(밀도 초과, 이상치 감지 등), 촬영 시각" 등의 메타데이터를 포함하여 저장될 수 있다.
이로써, 스캐닝이 진행된 흐름대로 스냅 샷들이 나열되어 저장될 수 있으며, 나중에 사용자가 해당 시점들의 시퀀스(sequence)를 따라 재생함으로써 연속된 탐색 과정을 시각적으로 복기할 수 있다.
또한, 적어도 하나의 프로세서는 사용자에 의해 입력된 재생 지시에 따라, 순차적으로 저장된 복수의 스냅 샷 정보를 미리 정해진 방식으로 재생할 수 있다(S2820). 구체적으로, 사용자(분석가 등)가 "재생(Play)"을 누르거나, 특정 인터페이스를 통해 스크리닝 기록을 순차적으로 확인하는 것을 지시하는 요청을 입력하는 경우, 적어도 하나의 프로세서는 순차적으로 저장된 복수의 스냅 샷 정보를 불러올 수 있다.
이때, 컴퓨팅 장치는 복수의 스냅 샷 정보의 재생 시의 송출 순서와 실제 스냅 샷 촬상 순서가 상이하도록 구현될 수 있다. 예컨대, 사용자는 촬영 당시에는 임의 순서로 스냅 샷을 촬영했지만, 재생 시에는 특정 기준(예: 문제 우선도, 시간 역순, 사용자 지정 순서 등)에 따라 다시 정렬할 수 있다.
또한, 적어도 하나의 프로세서는 슬라이드쇼 방식(고정 간격으로 연속 화면 전환), 애니메이션(장면 사이를 부드럽게 전환), 타임라인 동작(각 스냅 샷의 이벤트 발생 시점에 맞춰 진행) 등 미리 정해진 재생 방식을 적용해 스냅 샷들을 출력할 수 있다. 이때, 적어도 하나의 프로세서는 스냅 샷이 촬상된 적어도 하나의 영역을 데이터 이미지 상에서 시각적으로 표시하고, 해당 지점들을 잇는 이동 경로를 화면 상에 애니메이션 형태로 보여줄 수도 있다. 예를 들어, 컴퓨팅 장치는 스냅 샷이 촬영된 영역들을 순서대로 카메라가 이동하는 듯한 시각효과를 연출하여, 사용자에게 동적 탐색 경험을 제공할 수 있다.
이때, 컴퓨팅 장치는 재생에 연관된 사용자 명령을 수신할 수 있고, 수신된 명령에 따라 재생을 제어할 수 있다. 예를 들어, 적어도 하나의 프로세서는 사용자에 의한 인터랙션 동작을 기초로 "일시 정지(pause)", "건너뛰기(skip)", "역순 재생(reverse)", "개별 스냅 샷에 주석 추가" 등의 동작들을 수행할 수 있다.
이를 통해, 재생 중 특정 스냅 샷을 더 자세히 살펴보고, 필요한 경우 해당 시점에서 데이터 제거하거나 합성하는 등의 데이터 개선 동작을 결정할 수도 있다. 구체적으로, 적어도 하나의 프로세서는 스냅 샷의 재생 도중, 개선이 필요한 영역이 제공되는 경우, 개선 정보(예: 개선이 필요한 이유 등)를 제공함으로써 사용자로 하여금 개선 기능의 사용을 유도할 수 있다.
본 개시의 실시예에 따른 컴퓨팅 장치는 스냅 샷 정보와 진단 레포트의 연동 기능을 제공할 수 있다. 구체적으로, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 데이터 스크리닝 과정에서, 특정 영역에 대한 스냅 샷 정보와 함께 데이터 셋의 특성을 설명하는 진단 레포트를 생성할 수 있다.
도 29는, 다양한 실시예들에 따른, 컴퓨팅 장치가 진단 레포트와 스냅 샷 정보를 연동하여 생성하는 방법을 도시한 도면이다.
도 29를 참조하면, 컴퓨팅 장치 또는 컴퓨팅 장치에 포함되는 적어도 하나의 프로세서는 데이터 스크리닝 동작(S2910)을 수행할 수 있다. 구체적으로, 적어도 하나의 프로세서는 스크리너(SCREENER)를 이용하여, 데이터 셋(또는 벡터 셋)에 대해 특정 메트릭(밀도, 편향, 이상치 비율 등)을 측정하여, 진단 결과(예: 특정 영역의 과밀, 클래스 불균형)를 산출할 수 있다.
또한, 적어도 하나의 프로세서는 진단 레포트 생성 동작(S2920)을 수행할 수 있다. 구체적으로, 적어도 하나의 프로세서는 다양한 진단 결과를 취합하거나 요약함으로써 진단 레포트를 생성할 수 있다. 진단 레포트를 생성하는 방법에 대해서는 상술하였으므로 자세한 설명은 생략하기로 한다.
또한, 적어도 하나의 프로세서는 스냅 샷 정보 생성 동작(S2930)을 수행할 수 있다. 구체적으로, 스크리닝 과정에서 특이한 구역(과밀 영역, 이상치 밀집 등)이 발견되면, 적어도 하나의 프로세서는 해당 장면을 캡쳐하여 스냅 샷을 생성할 수 있다.
이때, 자세한 진단 결과가 이미 레포트 형태로 준비되어 있다면, 적어도 하나의 프로세서는 진단 결과를 축약하거나 요약한 형태로 스냅 샷 정보를 기록할 수도 있다(예: "노이즈 포인트 20개 발견, 편향 지표 0.85 이상인 클러스터" 등). 구체적인 예로, 적어도 하나의 프로세서는 특정 진단 결과에 대한 요약 정보 및 해당 진단 결과에 대응되는 데이터 이미지 상의 영역을 촬상한 장면을 기초로 스냅 샷 정보를 생성할 수 있다.
또는, 스크리닝 과정에서 진단 레포트가 생성될 때, 적어도 하나의 프로세서는 진단 결과들 중 특정 기준을 만족하는 부분(예: 임계치 이상, 관심 클래스 등)을 스냅 샷 정보로 전환할 수 있다.
적어도 하나의 프로세서는 진단 레포트와 스냅샷 정보의 연동 동작(S2940)을 수행할 수 있다. 구체적으로, 적어도 하나의 프로세서는 진단 레포트 상의 진단 결과와 해당 진단 결과에 대응되는 스냅 샷 정보를 연동하여 저장할 수 있다. 구체적인 예로, 적어도 하나의 프로세서는 두 자료(스냅 샷 정보, 진단 레포트)를 동시에 데이터베이스나 파일 시스템에 저장할 때, 고유 ID(예: 레포트 ID, 스냅 샷 ID)를 상호 링크하거나, 메타데이터(예: 시간, 영역 좌표, 이벤트 종류)를 동일하게 지정함으로써 연동할 수 있다.
적어도 하나의 프로세서는 스냅 샷에서 진단 레포트 호출 동작(S2950)을 수행할 수 있다. 구체적으로, 스냅 샷 정보가 표시되는 GUI에서, 사용자가 진단 레포트를 불러오는 것을 지시하는 버튼을 클릭하거나, 스냅 샷 위에 표시된 라벨(클래스 불균형 등)을 클릭하면, 컴퓨팅 장치는 연동 저장된 정보를 바탕으로 해당 스냅 샷에 대응하는 진단 결과를 신속하게 불러올 수 있다. 이후, 적어도 하나의 프로세서는 새 창(뷰포트) 또는 팝업 등을 통해 해당 진단 레포트를 제공할 수 있다.
이처럼 사용자는 스냅 샷 이미지를 확인하다가, 더 자세한 분석 결과가 궁금한 경우, 연동된 진단 레포트를 불러올 수 있다. 이때, 새로운 창(또는 제2 뷰포트)을 띄워서 레포트를 로드할 수 있으며, 기존 스냅 샷과 나란히 비교 표시도 가능하다.
결국, 본 개시의 실시예에 따른 컴퓨팅 장치는 데이터 스크리닝 과정에서 생성된 스냅 샷 정보와, 데이터 셋의 특성을 상세히 기술한 진단 레포트를 연동하여 저장하고, 이후 사용자 입력에 따라 상호 참조함으로써, 직관적인 시각 정보와 정밀 분석 결과 사이를 유연하게 오가며 데이터 품질을 개선 및 활용할 수 있도록 지원한다.
이상과 같이 실시예들이 비록 한정된 실시예와 도면에 의해 설명되었으나, 해당 기술분야에서 통상의 지식을 가진 자라면 상기의 기재로부터 다양한 수정 및 변형이 가능하다. 예를 들어, 설명된 기술들이 설명된 방법과 다른 순서로 수행되거나, 및/또는 설명된 시스템, 구조, 장치, 회로 등의 구성요소들이 설명된 방법과 다른 형태로 결합 또는 조합되거나, 다른 구성요소 또는 균등물에 의하여 대치되거나 치환되더라도 적절한 결과가 달성될 수 있다.
그러므로, 다른 구현들, 다른 실시예들 및 특허청구범위와 균등한 것들도 후술하는 특허청구범위의 범위에 속한다.

Claims (15)

  1. 컴퓨팅 장치에 있어서,
    메모리; 및
    상기 메모리에 전자적으로 연결되는 적어도 하나의 프로세서;를 포함하고,
    상기 메모리에 저장된 적어도 하나의 인스트럭션을 실행하도록 구현되는 적어도 하나의 프로세서는,
    제1 데이터 셋에 대응되는 제1 데이터 이미지 - 상기 제1 데이터 이미지는 상기 제1 데이터 셋에 포함되는 데이터 각각에 대응되는 복수의 데이터 포인트들을 포함함 - 를 제1 뷰 포트를 통해 제공하는 동작;
    제1 GUI를 통해 수신된 사용자 입력에 대응하여, 상기 제1 뷰 포트를 통해 제공 중인 상기 제1 데이터 이미지에 대한 제1 장면을 포함하는 제1 스냅 샷 정보를 생성하여 제2 뷰 포트를 통해 제공하는 동작; 및
    상기 제2 뷰 포트를 통해 제공되는 제2 GUI에 대한 사용자 입력을 기초로, 상기 제1 스냅 샷 정보를 외부 통신 네트워크와 연결하기 위한 링크 정보를 생성하는 동작;을 수행하도록 설정되는 컴퓨팅 장치.
  2. 제1항에 있어서,
    상기 제1 뷰 포트를 통해 제공하는 동작은,
    상기 제1 데이터 이미지에 포함되는 복수의 데이터 포인트들 중 적어도 일부에 대한 프리뷰 정보 - 상기 프리뷰 정보는 데이터 포인트에 대응되는 실제 데이터를 출력함으로써 획득됨 - 를 제공하는 동작;을 포함하는 컴퓨팅 장치.
  3. 제1항에 있어서,
    상기 제1 스냅샷 정보는, 제1 데이터 이미지가 표시된 상기 제1 뷰 포트의 시점(view point), 확대 비율, 또는 사용자에 의해 지정된 범위 중 적어도 하나를 고려하여 제1 장면을 촬상함으로써 생성되는 것을 특징으로 하는 컴퓨팅 장치.
  4. 제1항에 있어서,
    상기 제2 뷰 포트를 통해 제공하는 동작은,
    메모 입력을 수신하기 위한 인터페이스를 제공하는 동작;을 포함하는 컴퓨팅 장치.
  5. 제1항에 있어서,
    상기 적어도 하나의 프로세서는,
    상기 제1 스냅 샷 정보가 생성된 경우, 상기 제2 뷰 포트의 적어도 일부를 통해 복수의 태그들 중 적어도 하나를 선택할 것을 요청하는 인터페이스를 활성화하는 동작;을 수행하도록 더 설정되는 컴퓨팅 장치.
  6. 제1항에 있어서,
    상기 적어도 하나의 프로세서는,
    상기 제1 장면에 대해 보정을 수행함으로써 보정된 장면을 획득하는 동작;을 수행하도록 더 설정되는 컴퓨팅 장치.
  7. 제1항에 있어서,
    상기 스냅 샷 정보는, 복수의 시점(view-point)들에 대응되는 복수의 후보 장면들 중 사용자에 의해 선택된 장면을 제1 장면을 기초로 생성되는 것을 특징으로 하는 컴퓨팅 장치.
  8. 제1항에 있어서,
    상기 적어도 하나의 프로세서는,
    상기 제1 데이터 셋에 대한 진단 결과를 획득함으로써 진단 레포트를 생성하는 동작; 및
    상기 진단 레포트 상의 적어도 하나의 진단 결과와 상기 적어도 하나의 진단 결과에 대응되는 스냅 샷 정보를 연동하여 저장하는 동작;을 수행하도록 더 설정되는 컴퓨팅 장치.
  9. 제8항에 있어서,
    상기 적어도 하나의 프로세서는,
    상기 제2 뷰 포트를 통해 제공되는 제3 GUI에 대한 사용자 입력을 기초로, 상기 제1 스냅 샷 정보에 대응되는 상기 진단 레포트에 포함되는 진단 결과를 불러오는 동작;을 수행하도록 더 설정되는 컴퓨팅 장치.
  10. 제9항에 있어서,
    상기 적어도 하나의 프로세서는,
    제3 뷰 포틀르 통해, 상기 진단 레포트에서 상기 제1 스냅샷 정보에 대응되는 진단 결과를 제공하는 동작;을 수행하도록 더 설정되는 컴퓨팅 장치.
  11. 데이터 인터랙션 방법에 있어서,
    메모리에 적어도 하나의 인스트럭션을 실행하는 적어도 하나의 프로세서에 의해,
    제1 데이터 셋에 대응되는 제1 데이터 이미지 - 상기 제1 데이터 이미지는 상기 제1 데이터 셋에 포함되는 데이터 각각에 대응되는 복수의 데이터 포인트들을 포함함 - 를 제1 뷰 포트를 통해 제공하는 동작;
    제1 GUI를 통해 수신된 사용자 입력에 대응하여, 상기 제1 뷰 포트를 통해 제공 중인 상기 제1 데이터 이미지에 대한 제1 장면을 포함하는 제1 스냅 샷 정보를 생성하여 제2 뷰 포트를 통해 제공하는 동작; 및
    상기 제2 뷰 포트를 통해 제공되는 제2 GUI에 대한 사용자 입력을 기초로, 상기 제1 스냅 샷 정보를 외부 통신 네트워크와 연결하기 위한 링크 정보를 생성하는 동작;을 포함하는 데이터 인터랙션 방법.
  12. 제11항에 있어서,
    상기 제1 뷰 포트를 통해 제공하는 동작은,
    상기 제1 데이터 이미지에 포함되는 복수의 데이터 포인트들 중 적어도 일부에 대한 프리뷰 정보 - 상기 프리뷰 정보는 데이터 포인트에 대응되는 실제 데이터를 출력함으로써 획득됨 - 를 제공하는 동작;을 포함하는 데이터 인터랙션 방법.
  13. 제11항에 있어서,
    상기 스냅 샷 정보는, 복수의 시점(view-point)들에 대응되는 복수의 후보 장면들 중 사용자에 의해 선택된 장면을 제1 장면을 기초로 생성되는 것을 특징으로 하는 데이터 인터랙션 방법.
  14. 제11항에 있어서,
    상기 제1 스냅 샷 정보가 생성된 경우, 상기 제2 뷰 포트의 적어도 일부를 통해 복수의 태그들 중 적어도 하나를 선택할 것을 요청하는 인터페이스를 활성화하는 동작;을 더 포함하는 데이터 인터랙션 방법.
  15. 전자 장치에 있어서,
    적어도 하나의 GUI(Graphic User Interface)를 표시하도록 구현되는 디스플레이;
    메모리; 및
    상기 메모리에 전자적으로 연결되는 적어도 하나의 프로세서;를 포함하고,
    상기 메모리에 저장된 적어도 하나의 인스트럭션을 실행하도록 구현되는 적어도 하나의 프로세서는,
    제1 데이터 셋에 대응되는 제1 데이터 이미지 - 상기 제1 데이터 이미지는 상기 제1 데이터 셋에 포함되는 데이터 각각에 대응되는 복수의 데이터 포인트들을 포함함 - 를 제1 뷰 포트를 상기 디스플레이를 통해 제공하는 동작;
    제1 GUI를 통해 수신된 사용자 입력에 대응하여, 상기 제1 뷰 포트를 통해 제공 중인 상기 제1 데이터 이미지에 대한 제1 장면을 포함하는 제1 스냅 샷 정보를 제2 뷰 포트를 상기 디스플레이를 통해 제공하는 동작; 및
    상기 제2 뷰 포트를 통해 제공되는 제2 GUI에 대한 사용자 입력을 기초로, 상기 제1 스냅 샷 정보를 외부 통신 네트워크와 연결하기 위한 링크 정보를 제공하는 동작;을 수행하도록 설정되는 전자 장치.
PCT/KR2025/012348 2024-08-16 2025-08-14 데이터 시각화 도구의 구현 방법, 장치 및 시스템 Pending WO2026038902A1 (ko)

Applications Claiming Priority (10)

Application Number Priority Date Filing Date Title
KR20240109511 2024-08-16
KR10-2024-0109511 2024-08-16
KR1020250019947A KR102925436B1 (ko) 2024-08-16 2025-02-17 데이터 품질 개선을 위한 사용자 인터랙션의 구현 방법 및 그러한 방법이 구현된 컴퓨팅 장치
KR10-2025-0019945 2025-02-17
KR10-2025-0019947 2025-02-17
KR10-2025-0019944 2025-02-17
KR1020250019945A KR102925421B1 (ko) 2024-08-16 2025-02-17 데이터 진단에 따른 스냅샷 획득 방법 및 그러한 방법이 구현된 컴퓨팅 장치
KR10-2025-0019946 2025-02-17
KR1020250019944A KR20260025723A (ko) 2024-08-16 2025-02-17 데이터 렌즈 수준에 따른 데이터 분석 방법 및 그러한 방법이 구현된 컴퓨팅 장치
KR1020250019946A KR102925425B1 (ko) 2024-08-16 2025-02-17 데이터 스냅샷 획득을 위한 사용자 인터랙션의 구현 방법 및 그러한 방법이 구현되는 컴퓨팅 장치

Publications (1)

Publication Number Publication Date
WO2026038902A1 true WO2026038902A1 (ko) 2026-02-19

Family

ID=98780833

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2025/012348 Pending WO2026038902A1 (ko) 2024-08-16 2025-08-14 데이터 시각화 도구의 구현 방법, 장치 및 시스템

Country Status (1)

Country Link
WO (1) WO2026038902A1 (ko)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR20180051367A (ko) * 2016-11-08 2018-05-16 삼성전자주식회사 디바이스가 이미지를 보정하는 방법 및 그 디바이스
KR102029055B1 (ko) * 2013-02-08 2019-10-07 삼성전자주식회사 고차원 데이터의 시각화 방법 및 장치
KR102556766B1 (ko) * 2022-10-12 2023-07-18 주식회사 브이알크루 데이터셋을 생성하기 위한 방법
KR20230121164A (ko) * 2020-05-24 2023-08-17 킥소틱 랩스 인크. 신속한 스크리닝을 위한 도메인-특정 언어 해석기 및대화형 시각적 인터페이스
KR20240002385A (ko) * 2022-06-29 2024-01-05 주식회사 페블러스 데이터 클리닉 방법, 데이터 클리닉 방법이 저장된 컴퓨터 프로그램 및 데이터 클리닉 방법을 수행하는 컴퓨팅 장치

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
KR102029055B1 (ko) * 2013-02-08 2019-10-07 삼성전자주식회사 고차원 데이터의 시각화 방법 및 장치
KR20180051367A (ko) * 2016-11-08 2018-05-16 삼성전자주식회사 디바이스가 이미지를 보정하는 방법 및 그 디바이스
KR20230121164A (ko) * 2020-05-24 2023-08-17 킥소틱 랩스 인크. 신속한 스크리닝을 위한 도메인-특정 언어 해석기 및대화형 시각적 인터페이스
KR20240002385A (ko) * 2022-06-29 2024-01-05 주식회사 페블러스 데이터 클리닉 방법, 데이터 클리닉 방법이 저장된 컴퓨터 프로그램 및 데이터 클리닉 방법을 수행하는 컴퓨팅 장치
KR102556766B1 (ko) * 2022-10-12 2023-07-18 주식회사 브이알크루 데이터셋을 생성하기 위한 방법

Similar Documents

Publication Publication Date Title
WO2022154471A1 (en) Image processing method, image processing apparatus, electronic device and computer-readable storage medium
WO2020214006A1 (en) Apparatus and method for processing prompt information
WO2020159232A1 (en) Method, apparatus, electronic device and computer readable storage medium for image searching
WO2018093182A1 (en) Image management method and apparatus thereof
WO2020138928A1 (en) Information processing method, apparatus, electrical device and readable storage medium
WO2018117685A1 (en) System and method of providing to-do list of user
WO2018088794A2 (ko) 디바이스가 이미지를 보정하는 방법 및 그 디바이스
WO2015133699A1 (ko) 객체 식별 장치, 그 방법 및 컴퓨터 프로그램이 기록된 기록매체
WO2015178716A1 (en) Search method and device
EP1898339A1 (en) Retrieval system and retrieval method
EP3552163A1 (en) System and method of providing to-do list of user
WO2023080276A1 (ko) 쿼리 기반 데이터베이스 연동 딥러닝 분산 시스템 및 그 방법
WO2024122990A1 (ko) 비주얼 로컬라이제이션을 수행하기 위한 방법 및 장치
WO2026084477A1 (ko) Ai 기반 병리 이미지의 세포 분류 및 검출 방법, 장치 및 프로그램
WO2021162481A1 (en) Electronic device and control method thereof
WO2026038902A1 (ko) 데이터 시각화 도구의 구현 방법, 장치 및 시스템
WO2022080848A1 (ko) 영상 분석을 위한 사용자 인터페이스
WO2021230469A1 (ko) 아이템 추천 방법
WO2023080275A1 (ko) 성별 및 나이를 분류하는 딥러닝 프레임워크 응용 데이터베이스 서버 및 그 방법
WO2021246642A1 (ko) 폰트를 추천하는 방법 및 이를 구현하는 장치
WO2021107360A2 (ko) 유사도를 판단하는 전자 장치 및 그 제어 방법
WO2023080590A1 (en) Method for providing computer vision
WO2024085552A1 (ko) 합성 데이터 생성을 위한 가상 환경을 제공하는 전자 장치, 전자 장치의 동작 방법 및 전자 장치를 포함하는 시스템
WO2024005464A1 (ko) 데이터 클리닉 방법, 데이터 클리닉 방법이 저장된 컴퓨터 프로그램 및 데이터 클리닉 방법을 수행하는 컴퓨팅 장치
WO2026010411A1 (ko) 장치, 장치의 동작 방법 및 비일시적 기록매체

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25855002

Country of ref document: EP

Kind code of ref document: A1