Can I extract text from a scanned PDF?
Yes. When a PDF has no text layer (common for scans), this tool automatically switches to OCR — you don't need to figure that out yourself or go find a different tool. That step is slower than plain extraction (it recognizes page by page) and less accurate than direct extraction, so proofread anything important.
Why can't I extract text that's clearly visible on the page?
Besides scanned PDFs, there's another common cause: the PDF was saved with its text converted to curves/vector shapes (common in PDFs exported from design software — certificates, posters, forms). This is usually done so the document looks the same even if the reader's computer doesn't have the right font installed, but the trade-off is that the PDF no longer contains any information about which character each shape represents — just lines and fills. In this case, not just this tool but any tool, including Adobe Acrobat, will fail to select that text with the mouse — it's not a software limitation, the PDF itself simply no longer contains the text data. The good news is this tool's OCR fallback kicks in either way — OCR reads the page visually, so it doesn't matter whether the cause was a scan or converted-to-curves text.
Does the extracted text keep the original formatting?
Only basic line breaks are kept — font, size, color, and table structure aren't preserved. This tool is built for pulling out the text content, not turning a PDF into an editable Word document. If you need full formatting preserved for editing, use a dedicated PDF-to-Word tool instead.
Is my PDF uploaded to a server?
No. Text extraction runs entirely in your browser — the PDF's content is never sent to any server, and nothing is kept once you close the page.