ImgWell.com

Why Can't I Select or Copy Text From This PDF?

ImgWell's batch PDF-to-text tool showing a results list with one successfully extracted PDF (2 pages, ~340 characters, with a download button) and another flagged with a warning: no extractable text found, possibly scanned or saved with outlined text

You open a PDF, the text is perfectly legible — a contract clause, a form field, a paragraph of body copy — and then you try to select it, copy it, or run it through a text-extraction tool, and nothing happens. Either the selection tool just draws an empty box, or the extracted output comes back blank.

The instinct is to assume the tool is broken. Most of the time, it isn’t — the PDF simply doesn’t contain any actual text data, and no tool, however good, can extract characters that were never stored as characters in the first place.

The 30-second check

Open the PDF in any reader — your browser’s built-in viewer, Adobe Acrobat, whatever — and try dragging your mouse across the text like you’re selecting it.

If it’s the second case, one of two things is going on.

Cause 1: it’s a scanned page

The most common cause. The “page” is actually an image — a photo of a printed document, a fax, an old paper scan — wrapped in a PDF container.

The characters you’re reading are just pixels arranged to look like letters; the PDF file format has no idea what any of it actually says. Getting real text out of a scanned page requires OCR (optical character recognition) — genuine recognition, not extraction, the same technology used to read text out of a photo.

Cause 2: the text was converted to outlines

Less common than a scan, but far more confusing when it happens — because this kind of PDF isn’t scanned and doesn’t contain any images at all, and the text still won’t extract.

Some certificates, official forms, and posters get their text “converted to outlines” (sometimes called “creating outlines” or flattening) when they’re exported — every character gets turned into a filled vector shape instead of being drawn with an actual font. This is usually done so the document looks identical everywhere, regardless of whether the viewer’s computer has the right font installed.

The tradeoff: once that happens, the PDF no longer stores any information about which character each shape represents — just lines and fills. This isn’t a compatibility quirk in any particular tool. Open the same file in Adobe Acrobat and try to select that text with the mouse — it won’t select there either. The character data is genuinely gone from the file; no tool can read what was never saved, including AI-based tools (short of running OCR on it, which treats it exactly like a scanned image).

What you can actually do about each one

Checking a batch of PDFs at once

Rather than manually drag-selecting through every file, you can run a whole batch through at once and get a clear answer for each one — including which files have no text layer at all, instead of a blank result you have to guess about:

  1. Open batch PDF to text and drop in every PDF you want to check at once.
  2. Click “Start extraction.” Everything runs locally in your browser — nothing is uploaded to a server.
  3. Files with a real text layer show their page and character counts and download as .txt. Files with no extractable text are flagged directly in the results — “no extractable text found, this may be scanned or have outlined text” — instead of a silent, unexplained empty file.