Write Python. Put PDF processing to work.
pip install pdfrest gives you a typed Python client for the pdfRest API. Call Python methods for conversion, OCR, compression, redaction, and more, with PDF processing handled by your selected pdfRest deployment.
Your first call
Start with a small PDF that contains selectable text and your pdfRest Cloud API key. The SDK reads the key from PDFREST_API_KEY, keeping it out of your Python source.
1. Install the SDK and set your key
In your Python environment, run these commands in macOS or Linux:
pip install pdfrest
export PDFREST_API_KEY="your-api-key-here"
Using Windows PowerShell?
Install the same package, then set the environment variable in the terminal where you will run the script:
pip install pdfrest
$env:PDFREST_API_KEY="your-api-key-here"Need a key? Open API Lab and sign in or create your pdfRest account. For a virtual environment, uv, or Poetry, follow the getting started guide.
2. Upload a PDF and extract its text
Save the following as quickstart.py. Put your test PDF in the same working directory and name it input.pdf.
from pathlib import Path
from pdfrest import PdfRestClient
input_pdf = Path("input.pdf")
if not input_pdf.exists():
raise FileNotFoundError(
"Place a test PDF at ./input.pdf before running this script."
)
with PdfRestClient() as client:
uploaded = client.files.create_from_paths([input_pdf])[0]
document = client.extract_pdf_text(uploaded, full_text="document")
full_text = ""
if document.full_text is not None and document.full_text.document_text is not None:
full_text = document.full_text.document_text
print(f"Input file id: {uploaded.id}")
print("Extracted text preview:")
print(full_text[:500] if full_text else "(no text returned)")
3. Run the script
python quickstart.py
The script prints the uploaded file’s resource ID and up to 500 characters of extracted text. Image-only scans need OCR before text extraction; this example does not run OCR automatically.
create_from_paths uploads the file and returns a resource with an id. extract_pdf_text processes that resource and returns a typed response you can inspect in your editor.This example follows the official SDK quickstart. Free-plan restrictions can affect returned content; use the response and SDK warnings when checking results.
When a Python library is enough, and when it isn’t
If you already use pypdf, pdfplumber, or ReportLab, keep them for the jobs they handle well. Depending on the library, that may include reading metadata, rearranging pages, extracting text from digital PDFs, or creating and stamping documents. Local libraries also suit tasks that must run without a network service.
Reach for pdfRest when you want to integrate PDF processing without maintaining the processing service yourself, or deploy that service in your own infrastructure. Common needs include:
- Converting Office documents, images, and HTML to PDF while preserving layout
- Producing PDF/A for archival or PDF/X for print workflows
- Recognizing text in scanned documents with optical character recognition (OCR)
- Redacting content instead of simply covering it
- Creating and validating ZUGFeRD and Factur-X e-invoices through the API
- Compressing and linearizing PDFs for storage and delivery
These are pdfRest API capabilities. Available operations depend on your deployment and plan, and not every endpoint has an SDK helper. The coverage section below explains how to handle that distinction.
Chain operations with resource IDs
When an operation produces a file, reuse the output resource in the next operation instead of downloading and uploading it again. In the SDK, a single-file result exposes output_file; results with multiple files expose output_files.
Keep intermediate files in the same pdfRest deployment and pass the resulting file resource into the next supported SDK method.
Text extraction in the first example returns text in a response model. It does not need to create an output PDF, so an output file ID is not universal.
See the SDK file and chaining guide for resource handling. The Python API samples include a Complex Flow Examples folder with chained workflows. Those API samples are separate from the SDK examples.
One Python interface. Three deployment options.
The client defaults to pdfRest Cloud at https://api.pdfrest.com. Set base_url when creating the client to target the EU region, your pdfRest Container, or your AWS deployment.
- Cloud: use the US or EU endpoint and your Cloud API key.
- Container: direct requests to the pdfRest server you operate, with licensing and access configured for that deployment.
- AWS: use the URL and configuration for your deployed pdfRest service.
The SDK keeps the same client interface across deployments. Before moving a workflow, confirm that the target deployment supports its endpoints and options, and configure its URL, access, and licensing. Existing resource IDs stay with the deployment that created them.
Running Python locally does not make Cloud processing local. For on-premises or isolated environments, the processing service must also be deployed within your infrastructure.
The client configuration guide covers base_url, timeouts, retries, and asynchronous calls with AsyncPdfRestClient.
What the SDK covers
The SDK provides typed helpers for a broad set of processing and file-management operations. Use the API guide to find a method by task, then check its parameters and return model in the API reference.
For an endpoint without a helper, build the request in API Lab and copy the generated Python sample. This is a direct API request, not an SDK method. Configure it for the same endpoint and authentication, and reuse resource IDs where the operation supports them.
The current SDK guide does not list dedicated ZUGFeRD or Factur-X helpers. Use the API directly for those operations and check the guide as coverage expands.
Beyond Python
If part of your pipeline lives in self-hosted n8n, the pdfRest community node brings supported PDF operations into visual workflows. Install @pdfrest/n8n-nodes-pdfrest from Settings → Community Nodes, then connect to pdfRest Cloud or your self-hosted deployment.
Follow the n8n setup guide for installation, connection settings, and workflow templates. Node coverage can differ from SDK and API coverage.
Next steps
- Python SDK documentation: setup, configuration, methods, and models.
- SDK source and examples: explore the package and example code.
- API Lab: test requests and generate starter code.
- Python API samples: browse the Complex Flow Examples folder for connected operations.
