United Kingdom. Extraction routes, and the data protection question nobody asks first
Convert PDF to CSV: which route works for the file you have?
There is no such thing as a table in a PDF. The format describes marks on a page, not rows and columns, so every conversion is a reconstruction and how well it works depends entirely on what the file has underneath it. Three files that look identical on screen need three different methods, and one of them needs the data typed. Two questions settle which you are holding, and a third settles a question people usually skip: whether uploading this particular file to a website is a decision somebody should have made deliberately.
Question 1
Open the PDF and try to select a line of text. What happens?
This one test separates the three fundamentally different cases, and it takes five seconds. Everything downstream depends on the answer.
Email me this result
We will send you these figures and a link back to them. One email, nothing else unless you tick the box.
PDF to CSV: which route works for which kind of PDF, 2026
Last updated
PDF has no concept of a row or a column. Every conversion reconstructs a table that was never stored as one, and how well that goes is decided by how the file was made rather than by which tool you use.
The categories in this table are the structural cases a PDF can be in, distinguished by a test anybody can run in five seconds: whether text selects. The accessibility statements are quoted from gov.uk's guidance on publishing accessible documents, read on 15 August 2026, including its warning that a scanned document converted to PDF cannot be searched and will not be read by a screen reader, which is the same property that stops a converter reading it. The data protection row is taken from the text of UK GDPR Article 28 on legislation.gov.uk, which requires processing by a processor to be governed by a contract setting out the subject matter, duration, nature and purpose of the processing, the type of personal data and the categories of data subjects. No product is named and no accuracy rate is published anywhere on this page: extraction accuracy depends on the individual file, no authority measures it, and a percentage here would be a guess given the appearance of a benchmark. What the table gives instead is the specific defect each route produces, because that is what you check the output for.
| What you have | How to tell | What works | The defect to look for |
|---|---|---|---|
| Text layer, ruled table | Text selects cleanly and the cells have lines around them | Any table extraction tool. The rules are objects in the file, so the tool has real evidence about the column boundaries | Merged cells and headers spanning two columns, which usually arrive shifted by one field |
| Text layer, whitespace-aligned table | Text selects cleanly and the columns are held apart by spacing alone | The same tools, with column positions checked by hand before the full run | A wide value closing the gap between two columns, merging them for that row only. Look for rows with fewer fields than the rest |
| Text layer, awkward | Text selects but grabs whole blocks or jumps around the page | Extraction with the reading order corrected, or a rule set written for the layout | Reading order. The output is complete and in the wrong sequence, which is easy to miss and hard to detect automatically |
| Scanned image | Nothing selects. It behaves like a photograph | Optical character recognition first, then extraction, then verification by eye | Character confusions in numeric columns: 0 and O, 1 and l, 5 and S, 8 and B, and decimal points that move or disappear |
| Fillable form | There are boxes you can click into and type in | Export the field data. It is already structured as name and value pairs, so nothing has to be inferred from position | Checkbox and radio values, whose stored value is frequently not what the form displays |
| Statement or invoice layout | It is not a grid: records wrap, totals sit between rows, amounts change column by sign | A rule set written for that specific layout, or getting the data from its source system instead | Plausible but unevenly wrong output. Subtotals arriving as transactions is the classic case |
| Any of the above, containing personal data | Names, addresses, account numbers, anything about an identifiable person | The same method, performed on a machine you control, or by a service your organisation has a processing contract with | Uploading it to a converter found in a search result. That service becomes a processor, and Article 28 requires a contract before the processing starts |
- PDF does not store tables: it stores marks on a page, so every PDF to CSV conversion reconstructs a structure that was never recorded.
- One test separates the three fundamentally different cases: if text selects, the characters are in the file; if nothing selects, the page is an image.
- gov.uk warns that a scanned document converted to PDF cannot be searched by users and will not be read by a screen reader, which is the same property that stops a converter reading it.
- Whitespace-aligned tables fail row by row rather than wholesale: a value wide enough to close the gap merges two columns for that row only.
- Optical character recognition confuses a predictable set of characters, and 0 with O, 1 with l, 5 with S and 8 with B are all invisible errors inside a numeric column.
- A fillable PDF already stores structured data as named field and value pairs, so exporting it is easier and more reliable than extracting text from any document.
- Uploading a file containing personal data to an online converter makes that service a processor, and UK GDPR Article 28 requires the processing to be governed by a contract.
Cite this page
“PDF to CSV: which route works for which kind of PDF, 2026”, CSV Convert, https://csvconvert.uk/ (updated 2026-08-15). The categories in this table are the structural cases a PDF can be in, distinguished by a test anybody can run in five seconds: whether text selects. The accessibility statements are quoted from gov.uk's guidance on publishing accessible documents, read on 15 August 2026, including its warning that a scanned document converted to PDF cannot be searched and will not be read by a screen reader, which is the same property that stops a converter reading it. The data protection row is taken from the text of UK GDPR Article 28 on legislation.gov.uk, which requires processing by a processor to be governed by a contract setting out the subject matter, duration, nature and purpose of the processing, the type of personal data and the categories of data subjects. No product is named and no accuracy rate is published anywhere on this page: extraction accuracy depends on the individual file, no authority measures it, and a percentage here would be a guess given the appearance of a benchmark. What the table gives instead is the specific defect each route produces, because that is what you check the output for.
Scope of this checker
- Text-layer PDFs, scanned image PDFs and the awkward middle case
- Ruled tables against whitespace-aligned columns, which need different tools
- PDFs that are really forms, where the data can be exported directly
- Statement and invoice layouts, which are not tables at all
- The data protection question: what uploading a file to a web converter actually is
- Route selection only. Nothing is uploaded to this site and nothing is converted here
CSV Convert is an independent site operated by Ellul Solutions Ltd. It is not affiliated with, endorsed by or connected to any software vendor, the Information Commissioner's Office or any government body, and it is not a law firm or a data protection adviser. This site does not convert, process or store your documents and never asks you to upload one: it tells you which route your file needs so you can choose a tool deliberately. Nothing here is legal advice on a data protection question. We publish no extraction accuracy figure anywhere, because accuracy depends on the individual document and no authority measures it. Every requirement quoted is taken from the source cited beside it and read on the date shown at the top of this page.
Have a lot of these, or one that will not behave?
Tell us what the documents are and how many. Data extraction suppliers who handle that kind of file will contact you directly.
Frequently asked
How do I convert a PDF to CSV?
It depends on how the PDF was made, and one test tells you which case you are in. Open it and try to select a line of text. If it selects cleanly the characters are in the file and a table extraction tool can reconstruct the columns. If nothing selects, the page is an image and optical character recognition has to invent the characters first, with verification required afterwards. If there are boxes you can type into, it is a form and the field data can be exported directly, which is the easiest case of all.
Why do PDF tables convert badly?
Because PDF does not store tables. It records where characters and lines go on a page, not that a group of characters is a cell or that a set of cells is a row. Every conversion infers a grid that was never recorded. Ruled tables convert better because the lines are objects in the file and give the tool real evidence about column boundaries. Whitespace-aligned tables are inferred from horizontal positions, which works until a value is wide enough to close the gap between two columns.
Is it safe to upload a bank statement to an online PDF converter?
It is a decision worth making deliberately rather than by clicking. A statement contains personal and financial data, and uploading it to a service to convert makes that service a processor acting on your behalf. UK GDPR Article 28 requires that processing to be governed by a contract setting out its subject matter, duration, nature and purpose, the type of personal data and the categories of data subjects. A free converter found in a search result has none of that. Converting locally, on a machine you control, removes the question.
How accurate is OCR on a scanned table?
We publish no accuracy figure, because it depends entirely on the individual document and no authority measures it, so a percentage here would be a guess given the appearance of a benchmark. What is predictable is which characters go wrong: 0 and O, 1 and l, 5 and S, 8 and B, and decimal points that move or vanish. In a numeric column those errors are invisible. Treat recognition output as a draft requiring verification, check any totals, and validate fields with known formats such as dates and account numbers.
My PDF is a bank statement, not a table. What now?
Statement and invoice layouts defeat general purpose converters for a structural reason: there is no consistent grid. Records wrap onto second lines, amounts change column by sign, and running totals sit between records and look like records. A converter produces plausible output that is unevenly wrong, which is worse than an obvious failure. What works is a rule set written for that specific layout, or getting the data from its source. For UK bank data in particular, asking whether a machine-readable feed exists is worth doing before writing a parser.
How do I check whether the conversion worked?
Six checks, in about five minutes. Count the rows against the source, including any that wrapped. Look for rows with fewer fields than the rest, which is the signature of a merged column. Confirm the numeric columns are numbers rather than text by sorting one. Sum them against any total in the source. Check the first and last rows, where headers arriving as data and off-by-one range errors both live. And spot check five records by eye from the middle of the document, which is the only check that catches reading-order errors.
Is there a better option than converting at all?
Frequently, and it is the first thing to try. A PDF was generated by something, and that something held the data in a database, a spreadsheet or a report. Asking the supplier, the bank or the public body that produced it for a CSV or a data feed takes one email and succeeds more often than people expect. Every hour spent on extraction is an hour spent reconstructing something that already exists somewhere in a better form, and the reconstruction is never as good as the original.
The detail
Sourced, dated, kept current.
- Why converting a PDF table to CSV is harder than it looks
PDF stores marks on a page, not rows and columns. What that means for extraction, and why two identical-looking files behave completely differently.
- Scanned PDFs: what OCR gets wrong, and how to catch it
Recognition errors in numeric columns do not look like errors. The predictable confusions, the checks that catch them, and why input quality dominates.
- Before you upload a PDF to an online converter
A free converter processing your file is a processor under UK GDPR, and Article 28 requires a contract. What that means in practice, and the simpler alternative.
- Checking a PDF to CSV conversion before you trust it
Every route has a signature defect. Six checks that take five minutes and catch nearly all of them, in the order worth doing them.
Sources
- gov.uk, publishing accessible documents
- UK GDPR Article 28, processor (legislation.gov.uk)
- UK GDPR Article 5, principles relating to processing (legislation.gov.uk)
- UK GDPR Article 32, security of processing (legislation.gov.uk)
- ICO, UK GDPR guidance and resources
- gov.uk, accessibility requirements for public sector websites and apps
One selection test, and you know the route
What happens when you select the text, what the data looks like, and whether it is personal. Three answers.
Put this on your site
<script src="https://csvconvert.uk/embed.js" async></script>