Drag & Drop a PDF file here
or click to choose a PDF with tables or structured text from your device
✅ .PDF · Selectable-Text DocumentsExtract table-like data from a text-based PDF, review detected rows and columns, and download a clean CSV file directly in your browser. Ideal for statements, reports, price lists, research tables and other PDFs that contain selectable text.
Upload a PDF containing selectable text. The converter reads text positions from each chosen page, groups nearby text into rows and cells, previews the detected table data and prepares a CSV download.
Drag & Drop a PDF file here
or click to choose a PDF with tables or structured text from your device
✅ .PDF · Selectable-Text DocumentsA practical browser-based workflow for turning text-based PDF tables into reusable spreadsheet data without manually retyping every row.
Use the default settings first, then fine-tune row alignment or column spacing only when the preview needs improvement.
A PDF usually stores visual text at page coordinates, while CSV stores a logical sequence of rows and fields. Conversion must infer structure that may not exist explicitly in the PDF.
A PDF to CSV converter has to bridge two very different document models. PDF is primarily a fixed-layout page description format: words, numbers and drawing commands are placed at exact positions. CSV is a plain-text data format where each record is a row and a delimiter separates fields. A PDF may look like a perfect table to a human even though the underlying file contains nothing that literally says “this is column 3.”
This page uses Mozilla PDF.js to read selectable text and its positions. It then groups text items that sit at nearly the same vertical coordinate into rows and uses larger horizontal spaces as probable cell boundaries. This coordinate-based approach works well for many statements, reports and simple tables, but it is still heuristic extraction rather than a guarantee that every visual grid will become a perfect spreadsheet.
CSV files commonly use commas between fields and quotation marks around values that themselves contain commas, quotes or line breaks. The widely referenced RFC 4180 documents a conventional CSV format. This converter follows standard double-quote escaping so values such as Smith, Jones & Co. remain one field rather than being split into two columns.
Some PDFs are generated by spreadsheet or reporting software and preserve text in predictable horizontal bands. Others position each word independently, draw table lines as vector graphics, or convert every page to an image. Two documents can look identical on screen but have completely different internal structures. That is why extraction quality depends more on the PDF's underlying text layout than on how clean the table looks visually.
Use this guide to predict whether direct text extraction should work or whether OCR/manual cleanup is more appropriate.
| PDF Type | Direct CSV Extraction | Expected Quality | Best Next Step |
|---|---|---|---|
| Digitally generated table with selectable text | Yes | Usually good with aligned rows/columns | Use this converter and review preview |
| Bank/financial statement with regular columns | Often | Good to moderate | Adjust column gap if descriptions merge |
| Multi-column report with paragraphs and tables | Possible | Mixed | Extract only table pages and clean CSV |
| Scanned statement or photographed table | No text layer | No direct table text | Run OCR first, then extract |
| Table with merged cells / nested headers | Partial | Manual cleanup likely | Review headers and realign columns |
| Password-protected or restricted PDF | Depends | May fail to open/extract | Use an authorized unlocked copy |
The two most useful controls are row tolerance and minimum column gap, because PDF text is positioned visually rather than stored as spreadsheet cells.
If a report contains a cover page, narrative sections and appendices, extracting everything at once can mix unrelated text into the CSV. Enter a page range such as 4-7 or 2,5,8-10 so the parser focuses on pages that actually contain the table.
PDF text items on the same visual row may not share exactly the same Y coordinate, especially when fonts have different sizes or baselines. The row tolerance allows small differences. If two separate lines collapse into one CSV row, reduce the tolerance. If one visual row is being split into two, increase it slightly.
The column-gap setting determines how much horizontal white space must appear between adjacent text items before a new cell begins. If a company name like “North Valley Supplies” becomes three CSV fields, increase the gap. If two separate numeric columns are being merged, lower the gap.
Transaction descriptions, product names and notes often wrap across multiple visual lines. A human recognizes the continuation, but the PDF may store it as a new row. For these files, extraction can still save substantial time, but you may need to merge continuation rows in your spreadsheet after download.
Always compare a sample of CSV rows against the original PDF, especially before using the data for accounting, research, compliance or financial analysis. Parentheses, minus signs, decimal separators and thousands separators can affect downstream calculations even when the visual table appears correct.
CSV is useful when fixed document tables need to become sortable, filterable, searchable or machine-readable data.
Most problems come from how the PDF encodes text, not from CSV itself.
If you cannot select individual words in a normal PDF viewer, the page may be an image. OCR is required before this text-based converter can detect rows.
Some generators store each word or character separately. Increase the column gap so ordinary spaces do not become extra CSV columns.
Reduce row alignment tolerance. Smaller values require text items to be closer vertically before they are grouped together.
Increase row tolerance slightly. Different font sizes, superscripts or baseline offsets can make one visual row appear at multiple Y positions.
Merged or nested headers do not map naturally to flat CSV fields. Extract the table, then normalize header names manually in a spreadsheet.
Visually positioned text can have an internal order different from what your eyes see. Coordinate sorting helps, but very complex designs may need a dedicated table-extraction tool.
Statements, invoices and reports may contain sensitive business or personal information, so local processing can be valuable.
The converter loads the selected PDF into browser memory and uses PDF.js to read its text layer and coordinates. The generated CSV is assembled locally and downloaded as a browser Blob.
Understand page ranges, delimiters, quoting, numeric formats, scanned PDFs, spreadsheet imports and the limits of automated table recognition.
CSV is intentionally simple: it stores rows and text fields without formulas, formatting, merged cells, colors, multiple worksheets or embedded charts. That makes CSV ideal for importing data into many systems, but it also means a visually rich PDF table cannot retain its original styling. If you need spreadsheet formatting or multiple sheets, an XLSX workflow may be more appropriate after the data has been extracted.
Comma-separated values are common in English-language workflows, while semicolons are sometimes easier in locales where commas are used as decimal separators. Tab-delimited output can be convenient for pasting into spreadsheet software, and pipe delimiters can help when the data contains many commas and tabs. Choose the delimiter expected by your target application.
Quoting prevents delimiters inside data from being interpreted as column separators. In minimal mode, the converter adds quotes only when a field contains the selected delimiter, a quote mark or a line break. In “quote every cell” mode, every value is enclosed in double quotes, which can be useful for predictable downstream parsing.
A value such as 1,234.56 can represent one number in one locale, while 1.234,56 can represent the same quantity elsewhere. This converter preserves extracted text rather than guessing numeric meaning. That avoids silently changing values, but your spreadsheet import settings must use the correct locale and delimiter.
Dates such as 03/04/2026 are ambiguous across regions. CSV contains no built-in date type, so spreadsheet software may interpret a text date automatically. When accuracy matters, import columns as text first and explicitly convert them using the intended date format.
Long PDF tables often repeat the same column headings on every page. This browser converter preserves what it detects, so repeated headers can appear as repeated CSV rows. That is safer than automatically deleting content that only looks duplicated. You can remove repeated header rows after reviewing the preview or downloaded file.
A scanner normally creates page images. Without an OCR text layer, there are no words and coordinates for PDF.js to extract. OCR software can recognize characters and add text, after which table reconstruction becomes possible. OCR quality depends on scan resolution, skew, font clarity, language, handwriting and table borders, so scanned data should be checked carefully.
After extracting table data, you may also need to split source pages, combine reports, compress documents or convert other file types.
Answers to common questions about PDF table extraction, CSV formatting, scanned documents, privacy and data accuracy.
Upload a text-based PDF above, preview the detected rows and columns, adjust extraction settings if needed, and download clean CSV data.
📊 Open PDF to CSV Converter ↑