# Scanned PDF OCR to Word

Rasterize scanned PDF pages, recognize text with Tesseract, and export an editable DOCX.

> Canonical page: https://elysiatools.com/en/tools/scanned-pdf-ocr-to-word

- **Category:** PDF Tools

- **Keywords:** PDF, OCR, Tesseract, Word, DOCX

## Overview

Render selected pages with Poppler or Ghostscript, run Tesseract locally, and place recognized text into editable Word paragraphs. Tables and handwriting may need manual cleanup.

## Inputs

- **Source PDF File** (file): Upload a scanned PDF
- **OCR Languages** (select)
- **Page Selection** (text): 1,3-5 (empty = all pages)
- **DPI** (number)
- **Page Segmentation Mode** (number)
- **OCR Engine Mode** (number)
- **Include Page Separators** (checkbox)

## When to use

- When a scanned PDF contains text that cannot be selected or edited.
- When you need to extract text from selected pages into an editable Word document.
- When you want to convert scanned documents in supported languages into Word paragraphs.

## How it works

- Upload a PDF file containing scanned pages.
- Choose the OCR language and optionally enter a page selection such as 1,3-5.
- Set the DPI, page segmentation mode, OCR engine mode, and page separator option if needed.
- The tool rasterizes the selected pages, runs Tesseract OCR locally, and exports the recognized text as a DOCX file.

## Use cases

- Turn scanned reports, letters, and archived documents into editable Word text.
- Extract text from selected pages of a multipage scanned PDF.
- Prepare OCR text for proofreading, editing, and reuse in a DOCX document.

## Frequently asked questions

### What file can I upload?

Upload one PDF file containing scanned pages.

### What does the tool create?

It creates an editable DOCX containing recognized text in Word paragraphs.

### Can I convert only certain pages?

Yes. Enter page numbers and ranges, such as 1,3-5. Leave the field empty to process all pages.

### Which OCR languages are supported?

Supported choices include English, Simplified Chinese, Traditional Chinese, Japanese, Korean, Arabic, Hindi, Russian, French, German, Spanish, Portuguese, Italian, Vietnamese, and Thai.

### Will tables and handwriting be preserved accurately?

Tables and handwriting may require manual cleanup after OCR.

## Related tools

- [PDF Add Signature Online](https://elysiatools.com/en/tools/pdf-add-signature-online): Add a visible typed or image signature block to selected PDF pages and download the signed copy
- [PDF Batch Watermark](https://elysiatools.com/en/tools/pdf-batch-watermark): Apply text watermarks to multiple PDF files with rotation, opacity, and tiling options
- [PDF Denoise](https://elysiatools.com/en/tools/pdf-denoise): Remove visual noise from scanned PDF pages — salt-and-pepper speckle, random grain, and faint background haze — using real image-processing algorithms. Text pages are preserved as searchable vector content.
- [PDF OCR Text Layer](https://elysiatools.com/en/tools/pdf-ocr-text-layer): Add searchable/copyable OCR text layer to scanned PDF using Tesseract
- [Encrypted PDF Converter](https://elysiatools.com/en/tools/encrypted-pdf-converter): Open password-protected PDFs with OpenDataLoader and export them as Markdown, JSON, or text once the correct password is provided
- [Markdown to PDF Converter - Markdown转PDF转换器](https://elysiatools.com/en/tools/markdown-to-pdf-converter): Convert Markdown files to PDF documents with proper formatting, syntax highlighting, and styling
- [PDF Header/Footer Snippets](https://elysiatools.com/en/tools/pdf-header-footer-snippets): Convert HTML to PDF with reusable logo, title, and date snippets
- [PDF to Clean Text for LLM](https://elysiatools.com/en/tools/pdf-to-clean-text-for-llm): Extract clean text from PDFs with OpenDataLoader for summarization, translation, embedding, and other LLM workflows

## Samples

- [PDF Samples](https://elysiatools.com/en/samples/pdf-samples): Generated PDF samples from tools dated 2026-02-01 to 2026-02-10
- [Markdown Slide Deck Samples](https://elysiatools.com/en/samples/md-slide-deck-to-pdf): Remark/Marp style Markdown slide decks for testing PDF export layouts
- [DOCX Samples](https://elysiatools.com/en/samples/docx-samples): Sample DOCX documents for word-processing testing and parser validation
- [Time Zone Workflow Scheduler ICS Samples](https://elysiatools.com/en/samples/time-zone-workflow-scheduler-ics-samples): ICS files generated in the same structure returned by the Time Zone Workflow Scheduler, with multiple VEVENT meeting candidates exported from overlap windows

## Related content

- [Text Case, Encoding, and Normalization Conversion Tools](https://elysiatools.com/en/hubs/text-convert): Compare text case conversion, character-width conversion, encoding conversion, quoted-printable handling, and inline text normalization tools in one hub.
- [PDF Conversion and Document Export Tools](https://elysiatools.com/en/hubs/pdf-convert): Compare tools that convert documents, images, and structured extractions into or out of PDF in one hub for publishing, sharing, and downstream processing.
- [Text Tools](https://elysiatools.com/en/hubs/text-utility): Explore 33 text tools for utility workflows and compare closely related utilities quickly.
- [PDF Assembly, Layout, and Protection Tools](https://elysiatools.com/en/hubs/pdf-utility): Compare PDF page assembly, layout control, watermarking, stationery overlays, anonymization, password protection, and redaction helper tools in one hub.
