OCR PDF: Make a Scanned PDF Searchable
PDFs made with a scanner or a phone store text as images, so you can't search or copy it. Upload one here and every page is run through text recognition (OCR), giving you a PDF with an invisible text layer on top of the unchanged original pages. Find words with Ctrl+F and drag to copy text in the result. Files are never sent to a server.
Loading the tool...
How to use
- Drop or choose a scanned or photo PDF whose text can't be selected. If the PDF is password-protected, unlock it first with Unlock PDF.
- Choose the document language (Korean + English, Korean, or English) and the recognition quality (Normal 200 dpi or Accurate 300 dpi), and set a page range for long documents. For English-only documents, pick English. Pages that already have text are skipped by default.
- Click 'Start OCR'. Each page is read in turn and progress is shown. You can cancel at any time.
- When it's done, the searchable, copyable PDF is ready to download. You can also download just the recognized text (.txt), or pass the result straight to a next tool such as Compress PDF.
How search and copy work
Each page is rendered as an image and run through OCR, which returns every word with its text and position (bounding box). Dagochim leaves the original page as it is and writes each word in the same position as invisible text (the PDF's transparent text layer, text rendering mode 3). You see the original scan on screen, but when you search or drag to select, the hidden text underneath is picked up. Because the font size is set to match each word's width, the selection lines up closely with the original text.
Good to know: limits
- Misread words won't show up in search. Blurry scans, skewed pages, very small print, handwriting, text inside tables and vertical text are recognized less accurately.
- To support both Korean and English characters, the full Nanum Gothic font is embedded, which makes the file larger (up to about 2 MB). If needed, shrink the result again with Compress PDF (light).
- Each page takes a few seconds to a few tens of seconds. The first time, OCR language data is downloaded once (about 6 MB for Korean + English).
- If you want to turn phone photos straight into a PDF, Scan to PDF is easier (corner correction and black-and-white document mode).
Open-source credits
- tesseract.js (Apache License 2.0) and tessdata_fast Korean and English data (Apache License 2.0) — recognize the text in the page images.
- PDF.js (Mozilla, Apache License 2.0) — renders PDF pages as images.
- pdf-lib (MIT License) and fontkit (MIT License) — write the invisible text layer into the PDF.
- Nanum Gothic (NAVER, SIL Open Font License 1.1) — the font used for the hidden Korean and Latin text.
FAQ
Why can't I search a scanned PDF?
A PDF made with a scanner or phone camera is like one photo per page. People see text, but to a computer it's just an image. OCR (optical character recognition) reads the text in the image and adds it to the PDF as real text, so you can search and copy it.
Does it change how my PDF looks?
No. The original page images and content are left untouched; the recognized words are only added as invisible text in the same positions. The PDF looks and prints exactly like the original. The embedded font (Nanum Gothic) makes the file larger, by up to about 2 MB.
How accurate is the OCR?
Clean scans of printed documents are mostly read correctly, but blurry or skewed scans, small print, handwriting, text inside tables and vertical text can be misread or missed. Misread words won't show up in search, so check the result for important documents. If there's a lot of small text, raise the quality to 300 dpi.
How long does it take?
It depends on your computer and the amount of text, but usually a few seconds to a few tens of seconds per page. The first time, OCR language data is downloaded once (about 6 MB for Korean + English). For long documents, it's easier to process the pages in several ranges.
Is my file sent to a server?
No. Reading the PDF, recognizing text and creating the new PDF all happen inside this browser. Only the OCR language data and font files are downloaded; your PDF and the recognized text are never sent anywhere.
Related tools
Image to Text (OCR)
Extract text from screenshots and photos.
PDF & documentsPDF to HWP (HWPX)
Convert PDF text and tables into an editable Hangul document.
PDF & documentsPDF to Excel
Extract tables from a PDF or photo into an XLSX file.
PDF & documentsTranslate PDF
Translate English PDFs to Korean, or Korean PDFs to English.
PDF & documentsScan to PDF
Scan documents or book pages, straighten corners, add OCR.
PDF & documents