At some point almost everyone runs into a PDF they can't fully read: an English research paper, a contract from an overseas partner, a product manual, a visa or job application form. The usual workaround is clunky - open the PDF, select and copy text in chunks, paste it into a translator, then copy the translated text back into a document and try to recreate the layout by hand. It is slow, error-prone, and especially painful for longer documents with many short paragraphs and headings. This guide walks through a simpler approach: upload the PDF as-is, let the tool detect paragraphs and font sizes automatically, translate them one by one using an open-source AI model that runs locally in your browser, and then export either plain text or a new PDF that tries to keep each paragraph close to its original position.
1. First, check whether your PDF actually contains text
Before you try to translate anything, it helps to understand what kind of PDF you're dealing with. Most PDFs created from Word, Google Docs, Hangul (HWP), or a web page retain real text data internally - the letters, their positions, and their font sizes are all stored as structured data, even though visually it just looks like a printed page. These are called "text PDFs," and they can be processed directly: a computer program can read every character, its coordinates, and its size without needing to guess anything.
On the other hand, a PDF made by scanning a paper document with a scanner or photographing it with a phone camera is fundamentally different. Even though it looks exactly like a normal page of text, what's actually stored inside the file is a picture - a grid of pixels - with no character information attached at all. A computer has no way to know that a particular cluster of pixels represents the word "Agreement" unless it runs optical character recognition (OCR) first. The quickest way to tell the difference yourself: open the PDF in any viewer and try to click-and-drag to select some of the text. If a highlighted selection box appears around the words, it's a text PDF. If nothing gets selected no matter where you drag, it's a scanned image PDF.
2. If it's a text PDF: use Dagochim's PDF Translator directly
If your document is a text PDF, you can go straight to Dagochim's PDF Translator and drop the file in, or click to select it. Once it's loaded, the tool samples a few paragraphs and automatically guesses whether the document is mostly English or mostly Korean, suggesting a translation direction accordingly; you can also override this and pick "English → Korean" or "Korean → English" manually if the auto-detection gets it wrong (which can happen with documents that mix both languages heavily, like a bilingual contract). For longer documents, you can choose a specific page range to translate instead of the whole file - this defaults to the first five pages, which is usually enough to preview the quality before committing to a full run. If you do need the entire document translated, you can widen the range, but keep in mind that translation time scales with the number of paragraphs, so a 40-page report will take noticeably longer than a two-page memo.
3. If it's a scanned PDF: run OCR first
If you try to load a scanned PDF into the translator, it will detect that there's essentially no extractable text and show a message pointing you toward OCR instead of silently producing an empty or broken result. In that case, the recommended workflow is to first open Dagochim's Image to Text (OCR) tool, which reads the printed characters out of each page image and turns them into actual selectable, copyable text. OCR does a solid job with clearly printed text at reasonable resolution, but it can struggle with handwriting, low-quality or skewed scans, unusual fonts, or very small print - so it's worth skimming the OCR output for obvious errors (a misread character here and there, merged words, or garbled lines) before copying it over for translation. Once you have clean extracted text, you can paste it into the translator's text-based workflow or treat it as your working document going forward.
4. Know what to expect: a one-time ~630MB download
It's worth understanding how this translator actually works under the hood, because it affects what to expect the first time you use it. Rather than sending your document's text to a remote server to be translated and sending the result back (which is how most "free online translator" websites operate), this tool downloads an open-source multilingual translation model called M2M100, released by Meta AI, and runs it entirely inside your own browser using WebAssembly. That means the actual translation computation happens on your device, not on Dagochim's servers or anyone else's.
The tradeoff is that model files are not small - expect roughly 630MB to download the first time you click "Start translating." This only has to happen once per browser; after the initial download, the model is cached locally (using the browser's own storage) so that every subsequent translation, even days or weeks later, starts up almost instantly without re-downloading anything. If you're on a mobile data plan or a slow connection, it's worth running that first translation while connected to Wi-Fi so the download doesn't eat into your data allowance or take an unexpectedly long time.
5. Watch the paragraph-by-paragraph progress, and use pause/cancel if needed
Once the model is ready, translation proceeds one paragraph at a time, with the interface showing how many paragraphs are done out of the total, plus a rough estimated time remaining based on how long recent paragraphs took (typically somewhere around two to four seconds per paragraph, depending on your device's processing power and the paragraph's length). For a document with a handful of paragraphs this finishes in well under a minute; for a long document with dozens of pages, it can take several minutes. If something comes up partway through, you can pause the process and resume later without losing progress, or cancel entirely if you realize you selected the wrong file or page range.
6. Review the translation and edit it directly
When translation finishes, you'll see the original text and the translated text side by side, paragraph by paragraph, matching how they appeared in the source PDF. Machine translation is remarkably useful but it isn't infallible - sentences involving specialized terminology, proper nouns, idioms, or numbers embedded in complex phrases are the most likely spots to need a second look. The translated text boxes are directly editable, so if you spot something awkward or incorrect, you can fix it right there in the browser, and whatever you type becomes the final version used in every export format below. This review step is especially important if the document will be shared with other people or used for anything beyond your own quick reference.
7. Choose the right export format for your purpose
Dagochim's PDF Translator offers three different output formats, and which one makes sense depends on what you're going to do with the translation next. If you just need the translated words themselves - to paste into an email, a chat message, or another document you're already writing - the plain text (.txt) download is the simplest and lightest option. If you want to preserve basic structure like paragraph breaks and headings so you can drop the content into a word processor or CMS with minimal reformatting, the HTML download works well, since headings are wrapped in heading tags and paragraphs in paragraph tags.
If, however, you need something that visually resembles the original PDF - for example, to send back a bilingual version of a form, flyer, or short report while keeping it recognizable - the "format-preserving PDF" option is the one to use. Under the hood, this works by covering each detected paragraph's original bounding box with a white rectangle and then drawing the translated text back into that same area, shrinking the font size automatically if the translation is noticeably longer than the source text (which is common, since Korean and English text often have different lengths for an equivalent sentence). This approach works best on relatively simple, single-column documents. For PDFs with tables, multiple text columns, or text layered directly on top of images or diagrams, the result can look visibly different from the original - paragraph boxes might not line up perfectly, or shrunk text might look cramped - so it's worth opening the exported PDF and checking a few pages before relying on it for anything formal.
8. Can you use the translated PDF for official purposes?
The short answer is: treat it as a high-quality first draft, not a final, certified translation, especially for anything with legal, medical, or official weight. M2M100 is a genuinely capable open-source translation model and does a good job with everyday language, general correspondence, articles, and straightforward instructions. However, like any machine translation system, it can misinterpret domain-specific legal terminology, subtle contractual language, medical jargon, or phrases whose correct meaning depends heavily on surrounding context that the model doesn't fully capture. For casual use - understanding the gist of an email, skimming a foreign-language article, drafting a rough reply - the output is typically good enough to use directly. For anything that will be signed, submitted to a government agency, used in a medical context, or otherwise carries real consequences if mistranslated, it's strongly recommended to have a qualified human translator or a fluent native speaker review the final text before it's used.
Quick step-by-step summary
- Open the PDF and try selecting text to check if it's a text PDF or a scanned image.
- If it's scanned, run it through Image to Text (OCR) first.
- If it's a text PDF, upload it directly to the PDF Translator.
- Confirm the translation direction and page range, then start translating (expect a one-time ~630MB model download).
- Read through the translated paragraphs and fix anything that looks off.
- Download as plain text, HTML, or a format-preserving PDF, depending on what you need next.
Frequently asked questions
Can it translate scanned PDFs (images)?
This tool only translates text-based PDFs directly. For a PDF made from scanned pages or photos, first extract the text with Dagochim's Image-to-Text (OCR) tool, then copy that text and translate it.
How good is the translation quality? Is it as accurate as a professional translator?
This produces a machine-translation draft using the open-source M2M100 model. It handles everyday sentences reasonably well, but technical terminology, complex sentences, and formal documents like contracts can come out wrong or awkward. Always have a human review important documents after translation.
Is my file sent to a server?
No. Reading the PDF, extracting text, AI translation, and generating the new PDF all happen locally in your browser.