# PDF to Text Advanced

Advanced PDF to text converter with page selection, formatting options, and metadata extraction

> Canonical page: https://elysiatools.com/en/tools/pdf-to-text-advanced

- **Category:** Document Tools

- **Keywords:** pdf, text, extract, convert, pages, metadata, formatting, advanced

## Overview

Advanced PDF to text conversion with extensive customization options.

**Features:**
- Extract text from PDF with high fidelity
- Select specific pages or page ranges
- Include PDF metadata (title, author, creation date)
- Add page headers and/or line numbers
- Preserve paragraph structure
- Multiple output formats (plain text, structured, JSON)
- Aggressive or gentle text cleaning
- Character encoding handling
- Batch processing support

## Inputs

- **PDF File** (file): Upload a PDF file
- **Page Range** (text): e.g., '1-5,7,10-12' or 'all'
- **Output Format** (select)
- **Text Cleaning** (select)
- **Include PDF Metadata** (checkbox)
- **Add Page Headers** (checkbox)
- **Add Line Numbers** (checkbox)
- **Preserve Paragraph Structure** (checkbox)

## When to use

- When you need to extract text from specific page ranges of a PDF rather than the entire document.
- When you want to convert PDF content into structured JSON format for data analysis or programmatic ingestion.
- When you need to clean up extracted text by removing unwanted formatting or preserving paragraph structures.

## How it works

- Upload your PDF document using the file input field.
- Configure extraction settings such as page range, output format (plain, structured, or JSON), and text cleaning level.
- Toggle options to include metadata, page headers, line numbers, or preserve paragraph structure.
- Click convert to process the file and download the extracted text output.

## Use cases

- Converting academic papers or reports into clean plain text for research and analysis.
- Parsing PDF invoices or books into structured JSON format for automated database entry.
- Extracting specific chapters or sections from large PDF manuals using custom page ranges.

## Frequently asked questions

### Can I extract text from specific pages only?

Yes, you can specify individual pages or ranges, such as '1-5,7,10-12', in the Page Range field.

### What output formats are supported?

The tool supports Plain Text, Structured text with separators, and JSON formats.

### Can I extract PDF metadata like author and title?

Yes, checking the 'Include PDF Metadata' option will append the document's metadata to the output.

### What does the text cleaning option do?

It offers gentle or aggressive cleaning to remove unwanted artifacts, or 'none' to keep the raw extracted text.

### Does the tool preserve paragraph layouts?

Yes, enabling the 'Preserve Paragraph Structure' option helps maintain the original paragraph formatting.

## Related tools

- [PDF Text Extractor](https://elysiatools.com/en/tools/pdf-text-extractor): Extract text content from PDF documents with support for page selection, formatting options, and multi-language processing
- [PDF to Clean Text for LLM](https://elysiatools.com/en/tools/pdf-to-clean-text-for-llm): Extract clean text from PDFs with OpenDataLoader for summarization, translation, embedding, and other LLM workflows
- [PDF Header/Footer Noise Remover](https://elysiatools.com/en/tools/pdf-header-footer-noise-remover): Compare extraction with and without repeated page furniture to spot header/footer noise before using PDF text in RAG, summarization, or editing workflows
- [PDF to Word](https://elysiatools.com/en/tools/pdf-to-word): Convert PDF files to editable Word documents (.docx) by extracting text and preserving basic formatting
- [PDF Denoise](https://elysiatools.com/en/tools/pdf-denoise): Remove visual noise from scanned PDF pages — salt-and-pepper speckle, random grain, and faint background haze — using real image-processing algorithms. Text pages are preserved as searchable vector content.
- [PDF to Markdown Converter](https://elysiatools.com/en/tools/pdf-to-markdown): Convert PDF documents to Markdown format with text extraction and formatting preservation
- [HTML Mail to PDF](https://elysiatools.com/en/tools/libreoffice-html-mail-to-pdf): Render email HTML content to PDF with preserved basic formatting using Puppeteer
- [Markdown to PDF Converter - Markdown转PDF转换器](https://elysiatools.com/en/tools/markdown-to-pdf-converter): Convert Markdown files to PDF documents with proper formatting, syntax highlighting, and styling

## Samples

- [PDF Samples](https://elysiatools.com/en/samples/pdf-samples): Generated PDF samples from tools dated 2026-02-01 to 2026-02-10
- [Markdown Slide Deck Samples](https://elysiatools.com/en/samples/md-slide-deck-to-pdf): Remark/Marp style Markdown slide decks for testing PDF export layouts
- [Text with Date Samples](https://elysiatools.com/en/samples/text-with-date-samples): Text containing various date formats for testing date extraction and parsing
- [Text Case Format Samples](https://elysiatools.com/en/samples/text-case-formats): Examples of different text case formats: camelCase, snake_case, kebab-case, PascalCase, and more

## Related content

- [PDF to Office, Web, Ebook, and Structured Conversion](https://elysiatools.com/en/hubs/pdf-to-office-and-ebook-converters): Turn PDFs into Word, Excel, PowerPoint, HTML, EPUB, image, text, and XML outputs with a clear format-selection workflow.
