A lot of the tables we need to work with never started life as a spreadsheet. A supplier emails an invoice as a PDF. A school sends a transcript as a scanned image. A shop's price list only exists printed on paper, so you take a photo of it. Retyping every row and column by hand into Excel is slow, and it is easy to mistype a number, especially in a long table. This guide walks through turning a table inside a PDF or a photo into a real, usable Excel (.xlsx) file, without uploading the document anywhere.
When you need to convert a PDF table to Excel
- You received an invoice or a statement as a PDF and need to re-enter the line items into bookkeeping software.
- You photographed a printed price list or inventory sheet and want it as a spreadsheet you can sort, filter, and total.
- A government agency or a research paper published statistics only as a PDF table, and you need the numbers for analysis.
- You have an old scanned document with a table and want to preserve it digitally as a spreadsheet instead of a flat image.
- A client sent a schedule or a comparison table as a photo in a chat app, and you need it in a spreadsheet to share with a team.
Step by step with Dagochim
- Open the PDF/Photo Table to Excel tool.
- Drop in your PDF file, or a photo that contains a table.
- If the PDF has selectable text, its exact position is read directly. If it is a scan or a photo, on-device text recognition (OCR) reads it first.
- Check the table that appears in the editable preview. Click any cell to retype it, and use the row/column delete, merge, and split buttons to fix anything that looks off.
- Download the result as a real .xlsx workbook, download just the current table as .csv, or copy the table and paste it straight into a spreadsheet you already have open.
Text PDFs vs. scanned PDFs: why the results differ
A PDF that still has selectable text - for example, one exported directly from Excel, Word, or an accounting system - already stores the exact position of every letter and number on the page. That makes it possible to read the position of each piece of text, line text up into rows by vertical position, and line it up into columns by horizontal position, with high accuracy. If the original document also drew visible ruling lines for the table (the thin lines separating cells), those lines can be used to make the column boundaries even more precise.
A scanned PDF or a photo has no text information at all - it is just a picture of a document. To extract anything from it, the page (or the photo) first goes through on-device optical character recognition (OCR), which tries to recognize each word and remembers roughly where on the page it sits. The recognized words are then grouped into rows and columns the same way as with a text PDF. Because OCR is pattern recognition rather than an exact readout, it is not 100% accurate - a blurry photo or an unusual font can cause it to confuse similar-looking characters, such as the digit 0 with the letter O, or the digit 1 with a lowercase l. Always re-check numbers and names after converting a scan.
Fixing a table that comes out misaligned
Complex tables - especially ones where a single cell spans multiple rows or columns (a "merged cell") - will not always convert perfectly on the first try, because a PDF or an image has no real concept of "this is one combined cell." The editable preview in Dagochim's tool gives you a few simple ways to fix this without starting over:
- Merge cells: pick two cells that were split by mistake and combine their text into one.
- Split a cell: if several values got joined into one cell (for example separated by a comma or a tab), split them out into the next cells in that row.
- Delete a row or column: remove an empty or unwanted row or column with one click.
- Toggle header row: decide whether the first row should be treated as a header when it is exported.
If a PDF has several pages with tables, each page's table appears as its own tab in the preview, so you can clean each one up separately before exporting everything together.
How the Excel file is actually built
This tool writes the .xlsx file itself - it does not rely on a third-party spreadsheet library. Internally, an .xlsx file is a zip archive containing a set of XML files that describe the workbook, the sheets, and the styles; this tool assembles that XML directly and packages it into a zip using a small open-source compression library. Each detected table becomes its own worksheet inside a single workbook. Values that look numeric - such as 1,234, -12.5, 30%, or an amount written with a currency symbol like ₩12,000 - are automatically detected and stored as real numbers (with the thousands separator, percent sign, or currency symbol stripped), so you can sum or average a column in Excel immediately without having to reformat it first. Column widths are set automatically based on how much text ends up in each column.
If you only need one of the tables, "Download as .csv" saves that single table as a CSV file encoded in UTF-8 with the byte-order mark that Excel expects, so text in Korean or any other non-Latin script displays correctly when it is opened. "Copy this table" copies the table as tab-separated text to your clipboard, which is the fastest option when you already have a spreadsheet open and just want to paste a table into it.
Alternatives: other ways to do this
Recent versions of Excel can sometimes open a PDF directly through Data > Get Data > From File > From PDF, and Hancom Office has a similar "import" feature for some file types. These built-in importers can work well for very simple, clean tables, but they tend to struggle in the same way with complex layouts, merged cells, or scanned documents, because the underlying problem - a PDF has no real table structure - is the same regardless of which tool reads it. There are also many upload-based online converters, but for documents that contain financial figures, personal information, or business data (like an invoice or a transcript), it is safer to use a method that processes the file locally instead of sending it to someone else's server.
Common problems and how to fix them
- Columns are misaligned. A PDF does not actually contain the concept of a "table" - only the position of each piece of text. Tables with irregular spacing between columns can be split incorrectly. Use the column delete or cell merge buttons in the preview to fix this quickly.
- Numbers come out as text. If a unit or currency word is directly attached to a number without a recognized symbol (for example "1234won" instead of "1,234" or "₩1,234"), it will not be recognized as numeric and will stay as text. Retype the value without the attached unit if you need to calculate with it.
- OCR gets some characters wrong. Low scan quality, a tilted photo, or an unusual font all reduce OCR accuracy. Use a sharp, flat, well-lit photo whenever possible, and always double-check figures that matter, like totals and dates.
- A large PDF takes a while. Pages with a lot of text, or scanned pages at high resolution, take longer to process because OCR has to analyze every pixel. Converting a smaller page range first can help you confirm the result looks right before processing the whole document.
Is my data safe?
Yes. Dagochim's PDF/photo-to-Excel tool never uploads your file to a server - everything, including finding the table and building the .xlsx file, happens inside this browser tab. The only thing fetched from outside is the language data that text recognition (OCR) needs the first time you use it; the contents of your document are never sent anywhere.
Tips for better accuracy
- Use the highest-resolution scan or photo you can. A blurry or low-resolution image reduces OCR accuracy noticeably.
- Keep the document as flat and straight as possible when photographing it. A tilted or curved page makes it much harder to correctly separate rows and columns.
- Handwritten entries inside a table are recognized less reliably than printed text; it is usually faster and safer to retype those cells by hand after conversion.
- For tables with amounts or dates that matter, compare the exported spreadsheet side-by-side with the original PDF or photo once before relying on it.
Frequently asked questions
Does it work on PDFs without a visible table?
If the text is laid out in rows and columns it will be detected as a table. Plain paragraph text falls back to one line per spreadsheet row.
Can it read scanned PDFs or photos?
Yes, those are read with on-device text recognition (OCR). Low-quality scans can produce mistakes, so always check the result.
My numbers show up as text in Excel
Most number-looking cells are stored as real numbers automatically. If a unit is glued to the number (like 1234won), it may stay as text; retype that cell or reformat it as a number in Excel.