R&D Achievements

Government-funded R&D projects and outcomes

Technical capabilities developed through national R&D and AI training datasets

Key metric

Selected as an operator of the national AI training-data program (AI Hub) six years running, 2020–2025

19 AI Hub training datasets 10 as lead · 9 as participant (2020–2025)
5 National R&D projects Lead or participant in MSIT and MSS programs (2017–2024)
Approx. KRW 3.7B Government funding for participating projects 2017–2024 grant agreements · includes all consortium members

Main R&D areas

We apply multimodal content understanding, media asset management, and data refinement to industry needs.

Multimodal content understanding and language AI

We research AI and machine translation that understand the visual and cultural context of webtoons and video.

  • Spatially aware translation of text in webtoon images
  • LLMs that account for narrative and cultural context
  • Multimodal translation using background, tone, and facial expression

Related project: MSS Startup Growth Technology Development program (strategic track)

Media asset management and semantic search

We extract people, dialogue, and mood from video and audio, then structure the information for natural-language search.

  • Automatic metadata generation and semantic search
  • Subtitle production and editing pipelines for broadcast and OTT content

Related project: MSIT broadcasting and communications technology R&D with industry partners

AI data refinement and industry AX

We refine unstructured data and design training datasets for specialized domains including law, finance, and patents.

  • Automated noise detection and data refinement
  • Dataset modeling for LLM training and evaluation
  • Legal LLM data, financial parallel corpora, and patent classification data

Related work: AI Hub datasets for specialized domains

Government R&D projects & national AI dataset record

National R&D since 2017 has built our language and media AI capabilities. Through AI Hub, we have developed 10 datasets as lead organization and participated in 9 more.

Major national R&D projects (2017–2024)

Technology development of subtitle production and editing for colloquial broadcasting contents MSIT · IITP Apr 2022 – Dec 2024
R&D of a multimodal machine translator for K-content localization and communication in the metaverse MSS · TIPA Apr 2022 – Jun 2024
Source-text correction system for improving machine translation performance MSS · TIPA Jul 2020 – Jul 2021
Personalized Korean–English neural machine translation system MSS · TIPA Jun 2018 – Jun 2019
GCon Studio — a cloud-based translation document management tool for business MSS · KISED Jul 2017 – Dec 2017

AI training-data program — 10 datasets as lead organization (2020–2025)

K-Stock content data 2025 · image · multimodal Approx. 52K images · approx. 1.15M labeled sentences
Korean-context video understanding data 2024 · video · image 41K training images · 410K descriptions · approx. 7.8K reference videos
Question generation data from educational video material 2024 · education · video 3.9K videos and 3.9K scripts · approx. 12K question–answer pairs
Generative-AI interpretation & translation data for international academic conferences 2023 · speech · interpretation Approx. 2K hours of Korean/English speech · approx. 900K transcribed/translated sentences
Generative-AI multilingual translation quality evaluation data 2023 · translation quality Approx. 307K source/human-translated sentences · approx. 123K post-edited/evaluated sentences
Machine translation app, translator evaluation and new corpus construction for AI Hub data 2023 · corpus · MT evaluation 1.4M sentences · approx. 2.86M deliverables
Machine translation app, translator evaluation and new corpus construction for AI Hub data 2022 · corpus · MT evaluation 1.09M sentences · approx. 2.95M deliverables
Korean–English parallel corpus for engineering and science 2021 · parallel corpus 1.5M Korean–English parallel sentences
Korean–English translation corpus (social sciences) 2020 · parallel corpus 1.5M Korean–English parallel sentences
Korean–English translation corpus (engineering and science) 2020 · parallel corpus 1.5M Korean–English parallel sentences

AI training-data program — 9 datasets as participant (2022–2024)

Intellectual property law LLM pre-training and instruction tuning data 2024 · law · LLM Approx. 95K source/Q&A items · 5.5K legal summaries
Civil law LLM pre-training and instruction tuning data 2024 · law · LLM Approx. 95K source/Q&A items · 5.6K legal summaries
AI instructor data 2024 · defence · speech Approx. 54K military training materials · approx. 12K Q&A items · 118 hours of navy speech
Comprehensive trip-chain data for public transport users 2024 · mobility · speech 5K trip chains · approx. 31K speech segments · 5K transcriptions
Individual trip-chain data for passenger car users 2024 · mobility · speech 5K trip chains · 5K GPS records · approx. 31K speech segments
Generative-AI multilingual parallel corpus for the financial domain 2023 · finance · multilingual Approx. 2.54M parallel sentences across 5 language pairs
Generative-AI Korean–Chinese / Korean–Japanese patent and engineering parallel corpus 2023 · patents · engineering 2.5M Korean–Chinese/Korean–Japanese parallel corpus items
Patent data mapped to the national science and technology standard classification 2022 · patents · classification Approx. 300K labeled patents · 17 main categories · 188 subcategories
Tourism food menu data 2022 · tourism · multilingual OCR Approx. 100K menu images · approx. 400K multilingual translations

Volumes follow public AI Hub information. Rounded figures are marked “approx.” Source, annotation, and evaluation items are counted separately and should not be summed across categories.

Applying research in industry

We apply research and validation results in LETR WORKS and use field feedback to inform further research.

Core research

Advancing the latest AI models and multimodal architecture algorithms.

Validation on national projects

Verifying data and technical stability through large government programs and industry–academia collaboration.

Deployment to the commercial platform

Commercialized as the core engine of the LETR WORKS media intelligence platform.

Open research partnerships

We work with research institutions, universities, and local governments on content AI and validation for industry-specific AI transformation.