CarePDF Logo
CarePDF Logo

OCR PDF

从 PDF 中提取文本图层,以便您可以搜索、复制或重复使用内容。

Select PDF files

or drag and drop files here

Got a password-protected PDF? Unlock it here first

How to OCR PDF – Extract Text

1

Upload your scanned document

Select a PDF that consists of scanned pages or images of text.

2

Select output format

Choose whether you want a fully searchable PDF or a plain, editable text file.

3

Extract text

The Optical Character Recognition (OCR) engine reads the image and converts it back into machine-readable text.

Frequently Asked Questions

Does this work on scanned documents that are just photos of text?
该工具使用 pdf.js 引擎读取已嵌入 PDF 中的文本图层,因此它在已包含真实文本的 PDF 上效果最佳。没有文本层的纯图像扫描需要完整的光学字符识别,这需要服务器端处理(如 Tesseract OCR),而不是该工具执行的浏览器内提取。
How do I know if my PDF already has a text layer?
尝试在普通 PDF 查看器中选择或突出显示页面上的文本 - 如果您可以选择单个单词,则存在文本层,并且该工具将干净地提取它。
What output formats can I choose?
Pick a searchable PDF if you want to keep the original look while making the content selectable and searchable, or a plain text file if you just need the raw words to paste elsewhere.
这可以处理手写笔记吗?
No, this tool extracts existing digital text layers rather than performing image-based character recognition, so handwriting isn't something it can read.
Is my document processed on your servers?
No, text extraction happens entirely inside your browser using pdf.js, so your document's content is never uploaded anywhere.

Explore More PDF Tools