United Kingdom. Extraction routes, and the data protection question nobody asks first

Convert PDF to CSV: which route works for the file you have?

There is no such thing as a table in a PDF. The format describes marks on a page, not rows and columns, so every conversion is a reconstruction and how well it works depends entirely on what the file has underneath it. Three files that look identical on screen need three different methods, and one of them needs the data typed. Two questions settle which you are holding, and a third settles a question people usually skip: whether uploading this particular file to a website is a decision somebody should have made deliberately.

Question 1

Open the PDF and try to select a line of text. What happens?

This one test separates the three fundamentally different cases, and it takes five seconds. Everything downstream depends on the answer.

PDF to CSV: which route works for which kind of PDF, 2026

Last updated

PDF has no concept of a row or a column. Every conversion reconstructs a table that was never stored as one, and how well that goes is decided by how the file was made rather than by which tool you use.

The categories in this table are the structural cases a PDF can be in, distinguished by a test anybody can run in five seconds: whether text selects. The accessibility statements are quoted from gov.uk's guidance on publishing accessible documents, read on 15 August 2026, including its warning that a scanned document converted to PDF cannot be searched and will not be read by a screen reader, which is the same property that stops a converter reading it. The data protection row is taken from the text of UK GDPR Article 28 on legislation.gov.uk, which requires processing by a processor to be governed by a contract setting out the subject matter, duration, nature and purpose of the processing, the type of personal data and the categories of data subjects. No product is named and no accuracy rate is published anywhere on this page: extraction accuracy depends on the individual file, no authority measures it, and a percentage here would be a guess given the appearance of a benchmark. What the table gives instead is the specific defect each route produces, because that is what you check the output for.

PDF to CSV: which route works for which kind of PDF, 2026
What you haveHow to tellWhat worksThe defect to look for
Text layer, ruled tableText selects cleanly and the cells have lines around themAny table extraction tool. The rules are objects in the file, so the tool has real evidence about the column boundariesMerged cells and headers spanning two columns, which usually arrive shifted by one field
Text layer, whitespace-aligned tableText selects cleanly and the columns are held apart by spacing aloneThe same tools, with column positions checked by hand before the full runA wide value closing the gap between two columns, merging them for that row only. Look for rows with fewer fields than the rest
Text layer, awkwardText selects but grabs whole blocks or jumps around the pageExtraction with the reading order corrected, or a rule set written for the layoutReading order. The output is complete and in the wrong sequence, which is easy to miss and hard to detect automatically
Scanned imageNothing selects. It behaves like a photographOptical character recognition first, then extraction, then verification by eyeCharacter confusions in numeric columns: 0 and O, 1 and l, 5 and S, 8 and B, and decimal points that move or disappear
Fillable formThere are boxes you can click into and type inExport the field data. It is already structured as name and value pairs, so nothing has to be inferred from positionCheckbox and radio values, whose stored value is frequently not what the form displays
Statement or invoice layoutIt is not a grid: records wrap, totals sit between rows, amounts change column by signA rule set written for that specific layout, or getting the data from its source system insteadPlausible but unevenly wrong output. Subtotals arriving as transactions is the classic case
Any of the above, containing personal dataNames, addresses, account numbers, anything about an identifiable personThe same method, performed on a machine you control, or by a service your organisation has a processing contract withUploading it to a converter found in a search result. That service becomes a processor, and Article 28 requires a contract before the processing starts
  • PDF does not store tables: it stores marks on a page, so every PDF to CSV conversion reconstructs a structure that was never recorded.
  • One test separates the three fundamentally different cases: if text selects, the characters are in the file; if nothing selects, the page is an image.
  • gov.uk warns that a scanned document converted to PDF cannot be searched by users and will not be read by a screen reader, which is the same property that stops a converter reading it.
  • Whitespace-aligned tables fail row by row rather than wholesale: a value wide enough to close the gap merges two columns for that row only.
  • Optical character recognition confuses a predictable set of characters, and 0 with O, 1 with l, 5 with S and 8 with B are all invisible errors inside a numeric column.
  • A fillable PDF already stores structured data as named field and value pairs, so exporting it is easier and more reliable than extracting text from any document.
  • Uploading a file containing personal data to an online converter makes that service a processor, and UK GDPR Article 28 requires the processing to be governed by a contract.

Cite this page

“PDF to CSV: which route works for which kind of PDF, 2026”, CSV Convert, https://csvconvert.uk/ (updated 2026-08-15). The categories in this table are the structural cases a PDF can be in, distinguished by a test anybody can run in five seconds: whether text selects. The accessibility statements are quoted from gov.uk's guidance on publishing accessible documents, read on 15 August 2026, including its warning that a scanned document converted to PDF cannot be searched and will not be read by a screen reader, which is the same property that stops a converter reading it. The data protection row is taken from the text of UK GDPR Article 28 on legislation.gov.uk, which requires processing by a processor to be governed by a contract setting out the subject matter, duration, nature and purpose of the processing, the type of personal data and the categories of data subjects. No product is named and no accuracy rate is published anywhere on this page: extraction accuracy depends on the individual file, no authority measures it, and a percentage here would be a guess given the appearance of a benchmark. What the table gives instead is the specific defect each route produces, because that is what you check the output for.

Scope of this checker

  • Text-layer PDFs, scanned image PDFs and the awkward middle case
  • Ruled tables against whitespace-aligned columns, which need different tools
  • PDFs that are really forms, where the data can be exported directly
  • Statement and invoice layouts, which are not tables at all
  • The data protection question: what uploading a file to a web converter actually is
  • Route selection only. Nothing is uploaded to this site and nothing is converted here

CSV Convert is an independent site operated by Ellul Solutions Ltd. It is not affiliated with, endorsed by or connected to any software vendor, the Information Commissioner's Office or any government body, and it is not a law firm or a data protection adviser. This site does not convert, process or store your documents and never asks you to upload one: it tells you which route your file needs so you can choose a tool deliberately. Nothing here is legal advice on a data protection question. We publish no extraction accuracy figure anywhere, because accuracy depends on the individual document and no authority measures it. Every requirement quoted is taken from the source cited beside it and read on the date shown at the top of this page.

Have a lot of these, or one that will not behave?

Tell us what the documents are and how many. Data extraction suppliers who handle that kind of file will contact you directly.

Tick the box above to send your enquiry.

  • Free, no obligation
  • We never ask you to upload a document here
  • No marketing lists, ever

Frequently asked

How do I convert a PDF to CSV?

It depends on how the PDF was made, and one test tells you which case you are in. Open it and try to select a line of text. If it selects cleanly the characters are in the file and a table extraction tool can reconstruct the columns. If nothing selects, the page is an image and optical character recognition has to invent the characters first, with verification required afterwards. If there are boxes you can type into, it is a form and the field data can be exported directly, which is the easiest case of all.

Why do PDF tables convert badly?

Because PDF does not store tables. It records where characters and lines go on a page, not that a group of characters is a cell or that a set of cells is a row. Every conversion infers a grid that was never recorded. Ruled tables convert better because the lines are objects in the file and give the tool real evidence about column boundaries. Whitespace-aligned tables are inferred from horizontal positions, which works until a value is wide enough to close the gap between two columns.

Is it safe to upload a bank statement to an online PDF converter?

It is a decision worth making deliberately rather than by clicking. A statement contains personal and financial data, and uploading it to a service to convert makes that service a processor acting on your behalf. UK GDPR Article 28 requires that processing to be governed by a contract setting out its subject matter, duration, nature and purpose, the type of personal data and the categories of data subjects. A free converter found in a search result has none of that. Converting locally, on a machine you control, removes the question.

How accurate is OCR on a scanned table?

We publish no accuracy figure, because it depends entirely on the individual document and no authority measures it, so a percentage here would be a guess given the appearance of a benchmark. What is predictable is which characters go wrong: 0 and O, 1 and l, 5 and S, 8 and B, and decimal points that move or vanish. In a numeric column those errors are invisible. Treat recognition output as a draft requiring verification, check any totals, and validate fields with known formats such as dates and account numbers.

My PDF is a bank statement, not a table. What now?

Statement and invoice layouts defeat general purpose converters for a structural reason: there is no consistent grid. Records wrap onto second lines, amounts change column by sign, and running totals sit between records and look like records. A converter produces plausible output that is unevenly wrong, which is worse than an obvious failure. What works is a rule set written for that specific layout, or getting the data from its source. For UK bank data in particular, asking whether a machine-readable feed exists is worth doing before writing a parser.

How do I check whether the conversion worked?

Six checks, in about five minutes. Count the rows against the source, including any that wrapped. Look for rows with fewer fields than the rest, which is the signature of a merged column. Confirm the numeric columns are numbers rather than text by sorting one. Sum them against any total in the source. Check the first and last rows, where headers arriving as data and off-by-one range errors both live. And spot check five records by eye from the middle of the document, which is the only check that catches reading-order errors.

Is there a better option than converting at all?

Frequently, and it is the first thing to try. A PDF was generated by something, and that something held the data in a database, a spreadsheet or a report. Asking the supplier, the bank or the public body that produced it for a CSV or a data feed takes one email and succeeds more often than people expect. Every hour spent on extraction is an hour spent reconstructing something that already exists somewhere in a better form, and the reconstruction is never as good as the original.

The detail

Sourced, dated, kept current.

Sources

  1. gov.uk, publishing accessible documents
  2. UK GDPR Article 28, processor (legislation.gov.uk)
  3. UK GDPR Article 5, principles relating to processing (legislation.gov.uk)
  4. UK GDPR Article 32, security of processing (legislation.gov.uk)
  5. ICO, UK GDPR guidance and resources
  6. gov.uk, accessibility requirements for public sector websites and apps

One selection test, and you know the route

What happens when you select the text, what the data looks like, and whether it is personal. Three answers.

Put this on your site

Embed it anywhere; attribution stays on. Copy this:

<script src="https://csvconvert.uk/embed.js" async></script>
Talk to a specialist