AI training data, with expert review

Data Intelligence × Human Experts

Prepare video, audio, images, and documents for AI training, with experts checking the data quality.

19 datasets AI Hub projects 10 as lead, 9 as participant
6 years Consecutive NIA data projects 2020–2025
3 stages Quality review Design, processing, and final dataset

How it helps your team

For companies and research institutions that need data for their own AI development.

Design for your training goal

Define data types, labeling rules, and output formats around what your AI needs to learn.

Prepare source materials

Add transcripts, translations, image descriptions, or labels, and bring files into a consistent format.

Check quality at each stage

Experts review the design, processing, and final dataset for errors and omissions.

Delivered data projects

Across NIA AI training-data projects in 2020–2025, we delivered 19 datasets: 10 as lead organization and 9 as participant. These selected projects show our roles and outputs.

Lead organization 2025 · image · multimodal

K-Stock content data

Approx. 52K images · approx. 1.15M labeled sentences

Lead organization 2024 · video · image

Korean-context video understanding data

41K training images · 410K descriptions · approx. 7.8K reference videos

Lead organization 2024 · education · video

Question generation data from educational video material

3.9K videos and 3.9K scripts · approx. 12K question–answer pairs

Lead organization 2023 · speech · interpretation

Generative-AI interpretation & translation data for international academic conferences

Approx. 2K hours of Korean/English speech · approx. 900K transcribed/translated sentences

Participant 2023 · finance · multilingual

Generative-AI multilingual parallel corpus for the financial domain

Approx. 2.54M parallel sentences across 5 language pairs

Participant 2022 · patents · classification

Patent data mapped to the national science and technology standard classification

Approx. 300K labeled patents · 17 main categories · 188 subcategories

Volumes describe the datasets as reported by AI Hub. Joint projects identify our participation role; dataset totals are not solely TWIGFARM deliveries. Source, processed, and evaluation volumes should not be added together.

See all 19 datasets, roles, and annual volumes →

Representative build volumes by domain

DomainRepresentative volume
Translation & language corporaThree datasets of 1.5M Korean–English sentence pairs each (2020–2021), 2.54M sentences across five financial language pairs, 2.5M Korean–Chinese / Korean–Japanese patent and engineering parallel items
Speech & interpretation2,017 hours of specialist Korean–English conference interpretation speech with 900K transcribed and translated sentences, plus 118 hours of defence-domain speech
Video & image multimodal41,000 images with 410K descriptions for Korean-context video understanding, 52,463 K-Stock images with 1.15M captions, 3,900 educational videos with 12,078 question–answer pairs
Specialist domains200,830 LLM training items for civil law and IP law, 300,240 labeled patents mapped to the national S&T classification, 400,452 multilingual tourism menu translations

The full list, with the year, role, and build volume of each dataset, is on the R&D achievements page. Volumes follow the detail published on AI Hub; because source, labeled, and evaluation data are counted separately for some datasets, summing every figure would double-count.

Dataset design and processing for your training goal

Multimodal dataset construction

We design and process advanced multimodal training data that combines video, audio, images, and text. The critical part is securing cross-modal alignment quality that single-modality datasets cannot provide.

Media-specific datasets

Broadcast subtitles, webtoon dialogue, image descriptions, multilingual translation, and custom metadata — datasets built on media intelligence that reflect the data shapes actually used in media operations.

Enterprise-grade refinement

Source data from specialized domains such as patents, law, environment (ESG), and engineering is refined to agreed training and delivery requirements, including noise identification and removal, schema normalization, and label taxonomy design.

Three-stage quality management

Training data requires more than collecting source material. We manage source quality, label definitions, consistency between workers, and delivery specifications. AI supports repetitive processing, with experts reviewing the results.

StageWhat is checkedBasis for the next stage
Design reviewTraining objective, data types, labels, and sample compositionAgreed instructions and sample results
Processing reviewErrors and omissions in transcripts, translations, descriptions, and labels; source alignmentCorrected, reviewed data in consistent formats
Final dataset reviewAgreed volume, format, and quality criteriaFinal dataset incorporating review results

Dataset quality is checked against agreed criteria. The performance of a model trained on the data is evaluated separately. Error tolerances, review methods, and sampling scope are defined at the start.

Experience operating large data projects

Six consecutive years of NIA AI training-data work, from 2020 to 2025, covered media, language, education, law, finance, defence, and mobility. Consortium collaboration, workforce management, large-volume data handling, and security and infrastructure operations inform delivery.

Clients can follow design, processing, and review as one connected process. Specialist work includes agreeing on the expertise needed to review terminology and context, while keeping source materials, outputs, and corrections distinct.

Delivery criteria agreed before work starts

  • Goal and scope: intended model and task, data types, languages, domains, and required volume.
  • Samples and instructions: label definitions, exception handling, and formats checked on a small batch.
  • Quality and review: inspection items, sample scope, error decisions, and rework criteria.
  • Schedule and data handling: milestones, interim reviews, security requirements, and transfer methods.

See the complete R&D project register for dataset names, years, roles, and public build volumes.

How it works

  1. 01 Define the goal and sources

    Agree on the training goal, source materials, required volume, and quality criteria.

  2. 02 Design and process a sample

    Check labels and formats on a small batch before applying them to the full dataset.

  3. 03 Review and deliver

    Check results against agreed criteria and deliver the training dataset.

See how it fits your team

Tell us about your work and materials. We will help you define the right scope.