Describe the scene you need
Search for a scene such as “the logo appears while someone explains the price” and review the relevant segments.
Content search and management
Find what you need in videos, images, documents, and audio, then translate, dub, and reuse it.
For teams that regularly reuse a content library and deliver it in multiple languages.
Search for a scene such as “the logo appears while someone explains the price” and review the relevant segments.
Continue from the same source into subtitling, translation, and dubbing. Keep originals and outputs together.
Give selected buyers access to an online screening room and see which content they watch.
Korean image and document retrieval were evaluated on public datasets in May 2026. Video results come from a separate product approximation evaluation.
| Modality | Dataset | Result |
|---|---|---|
| Image | XM3600-ko (3,600 images) | R@1 0.582 · R@5 0.813 · R@10 0.879 |
| Document | MIRACL-ko (300K documents) | With reranking, R@10 0.944 → 0.958 |
| Video (product-like) | 120 clips · multi-frame max-pool | R@1 0.625 · R@10 0.917 |
R@K is a recall measure for the top K retrieved results. These figures describe specific datasets and evaluation conditions, not overall service accuracy or performance on every customer library. Adoption includes evaluating retrieval on representative customer assets.
Using a separate tool per modality fragments permissions, history, and search. LETR WORKS keeps versions, folders, access rights, and index state together in a single asset model.
When auto-tagging ends as a list of strings, the same person scatters under a different name in every asset and terminology gets decided again with each translation. LETR WORKS puts extracted elements onto a workspace-level ontology so they hold relationships to each other.
| Entity | How it connects |
|---|---|
| People · speakers | Voice embeddings plus face clusters merge the same person across videos |
| Terminology · proper nouns | Linked to translation memory and glossaries (TM/TB) so translation, subtitles, and dubbing keep one spelling |
| Scenes · objects | Captions, OCR, and detections are bound to a scene — the smallest unit for search and re-editing |
| Assets · derivatives | Originals and their translated, dubbed, and re-typeset versions stay in one lineage |
Because of that structure, a search returns connected meaning rather than a handful of files, and the generation stage reuses the same definitions.
Video has two different axes: what was said and what is shown. LETR WORKS makes the scene the unit of indexing, builds a different kind of chunk for each aspect of that scene, and embeds each with the model best suited to it. Results come back as meaningful segments rather than a bare timecode.
| Chunk | Content |
|---|---|
visual | Keyframe imagery |
transcript | Spoken subtitles (with speaker labels) |
ocr | Text burned into the frame |
caption | AI-generated scene description |
objects · faces | Detected objects / face clusters |
Text and visual indexes are queried in parallel, merged, re-ranked, and grouped back into scenes. A router estimates whether a query is visual or textual in intent and weights the merge accordingly.
Search results become inputs for subtitles, dubbed audio, document outputs, and metadata, using the same original assets and terminology.
| Kind | What gets made |
|---|---|
| Generation | Scene captions, summaries, tags, retrieval metadata, sales-kit curation |
| Regeneration · language | Subtitle translation into 16 languages, TM/TB-consistent terminology, SDH subtitles |
| Regeneration · voice | AI voices that carry emotion, plus clone-voice dubbing based on the original performer |
| Regeneration · layout | Image and webtoon layer re-typesetting (no background regeneration), HWPX and DOCX document round-trips |
Structured assets can be reused for additional languages and formats. Cost and turnaround are assessed against the scope of each job. Short-form generation and a RAG chatbot are roadmap items being built on the same asset.
We separate what is ready to deploy now from what is rolling out in stages.
| Area | Status | Detail |
|---|---|---|
| Video translation | Available | STT → translation, waveform in parallel → project creation |
| Image & webtoon translation | Available | OCR → translation → reconstruction, including the PSD layer path |
| Document translation | Available | Parsing → translation, with HWPX round-trip |
| Dubbing | Available | TTS from the translation plus alignment data |
| DRM sales kit | Available | KMS/HLS encryption + allowlist + viewer analytics |
| Multimodal RAG indexing & search | Available | In-house GPU inference, benchmarked |
| AI Signal (discoverability diagnostics) | Available | Eight brand, content, and discoverability categories with multi-LLM probes |
| Broadcast QC · burn-in render · video compositing | Partially available | Available in some product lines |
| Short-form generation · RAG chatbot · audio separation · external connectors | Roadmap | Designed, being implemented in stages |
Cost stays under your control Video indexing is expensive to run, so it is opt-in per workspace. Turn it on only for the assets you need to search; un-indexed video is still fully usable inside a project. Costs are reviewed against the selected assets and processing volume.
Original, translated, and dubbed assets stay connected. Online screening rooms use email allowlists and KMS/HLS encryption; per-viewer analytics support buyer reviews and follow-up conversations.
Embedding, reranking, and OCR use in-house GPU inference. Feature scope, external processing, access rights, and indexing targets are agreed during adoption. If you also need translators and reviewers, see our professional translation and dubbing service.
Upload videos, images, documents, or audio and choose which assets to analyze and index.
Locate scenes and information, then create subtitles, translations, or dubbing in the languages you need.
Review outputs, download them, or share them with defined access permissions.
Tell us about your work and materials. We will help you define the right scope.