Document Tools
Extract text content from PDF documents with support for page selection, formatting options, and multi-language processing
Call this tool from your code in three languages.
# 1) Request a presigned URL → returns { uploadUrl, storageKey }
curl -X POST 'https://api.elysiatools.com/api/upload/presign/pdf-text-extractor' \
-H 'Content-Type: application/json' \
-d '{"filename":"pdfFile.ext","contentType":"application/octet-stream","size":12345}'
# 2) PUT the file bytes directly to the presigned uploadUrl
curl -X PUT '<presigned uploadUrl>' \
--data-binary @/path/to/file.ext
# 3) Call the tool, passing the returned storageKey for each file field
curl -X POST 'https://api.elysiatools.com/en/api/tools/pdf-text-extractor' \
-F 'pdfFile=uploads/2026/01/01/your-tool-1700000000000-abc123.ext' \
-F 'pageRange=e.g., 1-5 or 3 or 1,3,5' \
-F 'outputFormat=plain' \
-F 'preserveFormatting=true' \
-F 'removeExtraWhitespace=false' \
-F 'includeLineNumbers=false' \
-F 'encoding=utf-8'Send a POST request with your inputs as JSON. File parameters require a separate upload first.
POST https://api.elysiatools.com/en/api/tools/pdf-text-extractor| Name | Type | Required | Description |
|---|---|---|---|
| pdfFile | fileupload required | Yes | Supports PDF files up to 100MB |
| pageRange | text | No | Specify pages to extract (1-5 for range, 3 for single page, 1,3,5 for multiple). Leave empty for all pages. |
| outputFormat | select | No | — |
| preserveFormatting | checkbox | No | Keep original layout, spacing, and formatting as much as possible |
| removeExtraWhitespace | checkbox | No | Clean up excessive spaces and line breaks |
| includeLineNumbers | checkbox | No | Add line numbers to the extracted text |
| encoding | select | No | — |
Text result
{
"result": "Processed text content",
"error": "Error message (optional)",
"message": "Notification message (optional)",
"metadata": {
"key": "value"
}
}Add this tool to your Model Context Protocol server so AI agents can list and call it.
Add this block to your MCP client configuration:
{
"mcpServers": {
"elysiatools-pdf-text-extractor": {
"name": "pdf-text-extractor",
"description": "Extract text content from PDF documents with support for page selection, formatting options, and multi-language processing",
"baseUrl": "https://api.elysiatools.com/mcp/sse?toolId=pdf-text-extractor",
"command": "",
"args": [],
"env": {},
"isActive": true,
"type": "sse"
}
}
}After connecting to the SSE endpoint, list the exposed tools:
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list"
}Invoke the tool by its id, passing arguments built from its parameters:
{
"jsonrpc": "2.0",
"id": 2,
"method": "tools/call",
"params": {
"name": "pdf-text-extractor",
"arguments": {
"pdfFile": "https://example.com/file.ext",
"pageRange": "e.g., 1-5 or 3 or 1,3,5",
"outputFormat": "plain",
"preserveFormatting": true,
"removeExtraWhitespace": false,
"includeLineNumbers": false,
"encoding": "utf-8"
}
}
}Questions or issues? Contact [email protected]