Content search and management

LETR WORKS

Find what you need in videos, images, documents, and audio, then translate, dub, and reuse it.

95.8% Korean document retrieval recall MIRACL-ko · 300K documents · R@10
16 Translation languages Availability depends on content and feature
4 types Content supported Video, images, documents, and audio

How it helps your team

For teams that regularly reuse a content library and deliver it in multiple languages.

Describe the scene you need

Search for a scene such as “the logo appears while someone explains the price” and review the relevant segments.

Put existing content to work

Continue from the same source into subtitling, translation, and dubbing. Keep originals and outputs together.

Share with buyers

Give selected buyers access to an online screening room and see which content they watch.

Search evaluation results and conditions

Korean image and document retrieval were evaluated on public datasets in May 2026. Video results come from a separate product approximation evaluation.

ModalityDatasetResult
ImageXM3600-ko (3,600 images)R@1 0.582 · R@5 0.813 · R@10 0.879
DocumentMIRACL-ko (300K documents)With reranking, R@10 0.944 → 0.958
Video (product-like)120 clips · multi-frame max-poolR@1 0.625 · R@10 0.917

R@K is a recall measure for the top K retrieved results. These figures describe specific datasets and evaluation conditions, not overall service accuracy or performance on every customer library. Adoption includes evaluating retrieval on representative customer assets.

Four modalities, one asset model

Using a separate tool per modality fragments permissions, history, and search. LETR WORKS keeps versions, folders, access rights, and index state together in a single asset model.

  • Video — transcoding and previews, STT (Korean, English, Japanese, Chinese and more), speaker diarization, subtitle translation and dubbing, scene-level search
  • Image & webtoon — general/document/webtoon-specific OCR, PSD layer parsing that extracts text without OCR, layer re-typesetting with no background regeneration, speech-bubble shape recognition, font mapping, vertical writing
  • Document — PDF, DOCX, PPTX, XLSX and HWPX parsing, TM/TB-based translation, DOCX and HWPX round-trip export
  • Audio — per-speaker tracks, workspace-level voice profiles, speaker re-identification across videos

Connecting people, terminology, and scenes

When auto-tagging ends as a list of strings, the same person scatters under a different name in every asset and terminology gets decided again with each translation. LETR WORKS puts extracted elements onto a workspace-level ontology so they hold relationships to each other.

EntityHow it connects
People · speakersVoice embeddings plus face clusters merge the same person across videos
Terminology · proper nounsLinked to translation memory and glossaries (TM/TB) so translation, subtitles, and dubbing keep one spelling
Scenes · objectsCaptions, OCR, and detections are bound to a scene — the smallest unit for search and re-editing
Assets · derivativesOriginals and their translated, dubbed, and re-typeset versions stay in one lineage

Because of that structure, a search returns connected meaning rather than a handful of files, and the generation stage reuses the same definitions.

One scene, several clues, searched at once

Video has two different axes: what was said and what is shown. LETR WORKS makes the scene the unit of indexing, builds a different kind of chunk for each aspect of that scene, and embeds each with the model best suited to it. Results come back as meaningful segments rather than a bare timecode.

ChunkContent
visualKeyframe imagery
transcriptSpoken subtitles (with speaker labels)
ocrText burned into the frame
captionAI-generated scene description
objects · facesDetected objects / face clusters

Text and visual indexes are queried in parallel, merged, re-ranked, and grouped back into scenes. A router estimates whether a query is visual or textual in intent and weights the merge accordingly.

Outputs from the content you find

Search results become inputs for subtitles, dubbed audio, document outputs, and metadata, using the same original assets and terminology.

KindWhat gets made
GenerationScene captions, summaries, tags, retrieval metadata, sales-kit curation
Regeneration · languageSubtitle translation into 16 languages, TM/TB-consistent terminology, SDH subtitles
Regeneration · voiceAI voices that carry emotion, plus clone-voice dubbing based on the original performer
Regeneration · layoutImage and webtoon layer re-typesetting (no background regeneration), HWPX and DOCX document round-trips

Structured assets can be reused for additional languages and formats. Cost and turnaround are assessed against the scope of each job. Short-form generation and a RAG chatbot are roadmap items being built on the same asset.

What you get today

We separate what is ready to deploy now from what is rolling out in stages.

AreaStatusDetail
Video translationAvailableSTT → translation, waveform in parallel → project creation
Image & webtoon translationAvailableOCR → translation → reconstruction, including the PSD layer path
Document translationAvailableParsing → translation, with HWPX round-trip
DubbingAvailableTTS from the translation plus alignment data
DRM sales kitAvailableKMS/HLS encryption + allowlist + viewer analytics
Multimodal RAG indexing & searchAvailableIn-house GPU inference, benchmarked
AI Signal (discoverability diagnostics)AvailableEight brand, content, and discoverability categories with multi-LLM probes
Broadcast QC · burn-in render · video compositingPartially availableAvailable in some product lines
Short-form generation · RAG chatbot · audio separation · external connectorsRoadmapDesigned, being implemented in stages

Cost stays under your control Video indexing is expensive to run, so it is opt-in per workspace. Turn it on only for the assets you need to search; un-indexed video is still fully usable inside a project. Costs are reviewed against the selected assets and processing volume.

Sharing and operating your content library

Original, translated, and dubbed assets stay connected. Online screening rooms use email allowlists and KMS/HLS encryption; per-viewer analytics support buyer reviews and follow-up conversations.

Embedding, reranking, and OCR use in-house GPU inference. Feature scope, external processing, access rights, and indexing targets are agreed during adoption. If you also need translators and reviewers, see our professional translation and dubbing service.

How it works

  1. 01 Add your content

    Upload videos, images, documents, or audio and choose which assets to analyze and index.

  2. 02 Find and work

    Locate scenes and information, then create subtitles, translations, or dubbing in the languages you need.

  3. 03 Review and share

    Review outputs, download them, or share them with defined access permissions.

See how it fits your team

Tell us about your work and materials. We will help you define the right scope.