# PDF OCR — Extract Text from Scanned PDFs

> Pull text out of a scanned PDF in your browser. Each page is rendered at high resolution and recognized locally, with per-page text and a combined export. No upload, no signup.

URL: https://uttir.com/pdf-ocr
Categories: pdf-tools, image-tools, text-tools, developer-tools
Privacy: The PDF is rendered and recognized in your browser. Nothing is uploaded to a server.

## About

A scanned PDF is a PDF whose pages are images, not real text — open it in a viewer and you cannot copy the text because there is none, only pixels. This tool renders each page, recognizes the text in your browser, and returns the combined result with a clear per-page boundary.

Because the recognition runs locally, the PDF never leaves your device. Useful for old books, scanned contracts, government forms, and any other PDF where you need the text without the formatting.

## How to use

1. **Choose a PDF** — Up to 50 MB. The first run on a language downloads the model (~10 MB), cached after.
2. **Pick the language** — Match the language in the document. English is the default.
3. **Extract text** — Each page is rendered at 300 DPI equivalent, then recognized. Large PDFs may take a few minutes.
4. **Copy or download** — Per-page text in a tab, combined text in a single block, or download as .txt.

## Examples

### A scanned book chapter

Take a PDF of a scanned book chapter, run the tool, and get clean text you can paste into a note-taking app or a search box. The combined export gives you one block; the per-page view lets you copy a single page.

Input:

```
A 12-page scanned PDF of a chapter in English.
```

Output:

```
12 page results, each with the recognized text. Combined: ~15,000 characters with 90%+ average confidence on printed text.
```

### A scanned receipt or invoice

Snap or scan a paper receipt, save as PDF, drop it in. The tool returns the text with a per-page boundary, ready to paste into a spreadsheet or accounting tool.

Input:

```
A 1-page scanned receipt.
```

Output:

```
1 page result with the merchant, line items, total, and date. Confidence is usually 90%+ on printed receipts.
```

### A government form in another language

Match the language dropdown to the form. Switching languages downloads a new model on first use.

Input:

```
A 3-page French form.
```

Output:

```
3 page results, recognized in French. Pick French from the dropdown before extracting.
```

## FAQ

### Is my PDF uploaded?

No. The PDF is rendered and recognized in your browser. Nothing is sent to a server, and the file is gone when you close the tab.

### What size PDF can I process?

Up to 50 MB. A typical scanned page at 300 DPI is ~1 MB, so 50 MB is roughly 50 pages of dense scans. For larger documents, split the PDF first using the PDF Split tool.

### How long does it take?

A typical page takes 1-3 seconds after the language model is loaded. A 20-page PDF runs in 1-2 minutes. The first page on a new language is slower (10-30s) because the model downloads.

### The PDF already has text. Should I use this?

No — if you can already copy text from the PDF, the recognition is unnecessary and may even degrade the result. Use a regular PDF text extractor or just open the file in a viewer and copy. This tool is for PDFs whose pages are images (scans, faxes, image-only exports).

### What languages are supported?

12 common languages: English, Spanish, French, German, Italian, Portuguese, Russian, Japanese, Chinese (Simplified and Traditional), Korean, and Arabic. Match the dropdown to the document language for the best accuracy.

### Does it work on handwriting?

It is trained on printed text. Neat handwriting may work; messy handwriting usually will not. For serious handwriting recognition, a dedicated model trained on handwriting works much better, but no reliable in-browser option exists today.

### What about complex layouts?

Multi-column pages, tables, and mixed text-and-figure layouts are recognized for content but the structure may not be preserved — column 2 may end up above column 1, captions may mix with body text. For structured output, post-process the result.

### Why is the confidence low on some pages?

Low confidence means the recognition is unsure. Common causes: low-resolution scan, faded ink, stamps or watermarks over the text, or a font the model has not seen often. Try rescanning at higher resolution or pre-processing the page (increase contrast, convert to grayscale) before extraction.

## Related tools

- [Image to Text (OCR)](https://uttir.com/ocr) — Extract text from any image — screenshots, scanned documents, photos of signs. Choose a language, see the recognized text with confidence, and copy or download it. Free, in browser, no signup.
- [PDF to JPG](https://uttir.com/pdf-to-jpg) — Render every page of a PDF as a high-resolution JPG image.
- [PDF Merge](https://uttir.com/pdf-merge) — Combine multiple PDFs into one, in the order you choose.
- [PDF Split](https://uttir.com/pdf-split) — Extract a page range from a PDF into a new, smaller document.
- [PDF Compressor](https://uttir.com/pdf-compressor) — Repack a PDF with compressed object streams to reduce its file size.
- [PDF to Markdown](https://uttir.com/pdf-to-markdown) — Convert PDF text into clean Markdown with headings and lists.
- [Word Counter](https://uttir.com/word-counter) — Count words, characters, sentences, and paragraphs in your text instantly.
- [Character Counter](https://uttir.com/character-counter) — Count characters, words, lines, and UTF-8 bytes as you type.

---

For the full HTML page with the live tool, visit https://uttir.com/pdf-ocr.
This file is the markdown rendering at https://uttir.com/pdf-ocr.md. See https://uttir.com/llms.txt for a site-wide summary.
