
One-stop PDF processing with high-precision text and structured table extraction, page merging and splitting, and easy conversion between PDF and Excel/Word formats.
PDF Processor is a comprehensive, high-precision document operations and data extraction toolkit on EasyClaw. Designed specifically for full-stack developers, financial analysts, and administrators, it executes advanced headless PDF tasks locally — including high-speed document merging, page splitting, table-to-Excel extraction, dynamic page-numbering injection, and automatic table of contents (TOC) compiling in a single background operation.
The skill is built for analysts extracting data tables from multi-page financial reports, coordinators compiling joint-agreement PDFs, and developers building automated print-to-file pipelines locally.
The expected outcome is a successfully compiled PDF or spreadsheet dataset: written directly to your workspace imports, paired with an instant, local download link, complete with professional structural formatting.
1. Analyze input documents and action. Provide the file paths of the PDFs you want to process, and specify your operational task (e.g., "extract the table on page 3 of this financial report" or "merge these two PDFs").
2. Local PDF engine execution. The skill launches secure, local document processing libraries (such as Python `pypdf`, `pdfplumber`, or localized C++ PDF pools) headlessly in the background.
3. Structured table extraction. For table queries, the engine parses page geometries, isolates cellular borders and text runs, and reconstructs the data grid cleanly into a local CSV or Excel spreadsheet.
4. Dynamic merging and page-numbering. For merges, the skill concatenates the source files sequentially, registers standard running headers, generates page numbers, and compiles an automated TOC.
5. Output delivery. Converted spreadsheets, merged PDFs, or split documents are written directly to your local workspace exports folder (`public/data/pdf_exports/`) and returned with a direct local link.
- All-in-one PDF organizer: Merge, split, rotate, and compress PDF documents headlessly.
- Structured table extractor: Extracts page tables with high precision and exports them to Excel.
- Dynamic page-numbering: Injects running headers and automated "Page X of Y" footers.
- Table of contents (TOC) builder: Compiles an automated, clickable table of contents for merged PDFs.
- Frictionless local read/write: Read and write documents directly across any local filesystem folder.
- Clean Markdown reports: Outputs extraction summaries and merge logs with structured data grids.
1. Extracting a financial table from a PDF to Excel
An investment analyst is reviewing a 100-page corporate financial report and needs to run calculations on a dense table displayed on page 3. Copy-pasting the values manually scrambles the cells. The skill reads the PDF, isolates the page-3 table grid precisely, and exports the dataset as a clean, ready-to-use Excel (`.xlsx`) sheet in seconds.
2. Merging two contracts and adding a table of contents
An administrator needs to combine a 10-page main service agreement with a 5-page SOW attachment into a single PDF document. They ask the skill to merge. The tool concatenates the files, injects a unified page-numbering scheme, compiles an automated table of contents page at the front, and delivers the finalized PDF.
3. Splitting a bulk receipt PDF into separate files
An accounting assistant receives a single, compiled 50-page PDF containing all company receipts from last month. They want to split it. The skill parses the file, extracts each page as a separate, individually named receipt PDF, and writes them into an organized local directory on their desktop.
4. Compressing a heavy PDF report for email delivery
An executive has written an image-heavy corporate presentation PDF that is 25MB (too large for email limits). They ask the skill to compress the file. The tool runs a non-destructive image-sampling compression routine headlessly, reducing the file size to 3MB while keeping text and diagrams crystal clear.
5. Automatically prompting for missing target metrics
A user types a vague instruction: *"Merge these PDFs"* without providing files or formats. The skill's context logic detects the missing variables and prompts politely: *"To merge your documents, could you please share the absolute file paths of the PDFs? And would you like us to inject page numbers or generate a table of contents?"* preventing failed runs.
An analyst needs to merge two contracts and inject page numbering.
1. They open EasyClaw and activate PDF Processor.
2. They run: *"Merge 'D:\\project\\contracts\\MSA.pdf' and 'D:\\project\\contracts\\SOW.pdf'. Add page numbers and a TOC."*
3. The skill validates the local file paths, launches the local PDF engine, and executes the merge.
4. It compiles the pages, injects the running page-numbering footers, and builds the TOC.
5. It writes the merged document to: `public/data/pdf_exports/Merged_Agreement.pdf` and returns the download link.
Multi-page PDFs merged, split, and extracted headlessly in under 30 seconds.
Add this skill to your EasyClaw workspace
Describe your task in a chat message
Review the output and iterate if needed
Export or share the results directly from EasyClaw
Combine with other skills to build automated workflows
The skill natively supports: Merging multiple PDFs, splitting pages into separate documents, rotating pages, compressing file sizes, extracting structured tables to Excel/CSV, and injecting custom running headers, footers, page numbers, and tables of contents.
The table extractor is highly optimized to parse vector-based, selectable PDFs. For locked scanned PDFs or image screenshots, run the document through PDF OCR to Word Converter first to extract the text structure cleanly.
No. The core merging, splitting, table parsing, and PDF compiling pipelines are executed entirely locally in your workspace container, requiring no external paid API configurations or cloud subscriptions.
Yes. In compliance with strict corporate security and data privacy standards, all document reading, table parsing, and file writes are executed entirely locally on your machine, ensuring your data is never leaked.
The engine parses the PDF's vector coordinates to identify border lines and text run positions. It reconstructs the cells, rows, and columns of the table, writing the data directly to a clean Excel `.xlsx` or CSV file.
Yes. The PDF engine can automatically write custom running page numbers (such as "Page X of Y") in the footer or header of any merged or processed document, adjusting the alignment cleanly.
Yes. You can specify precise page ranges, such as "split page 1 to 5 as a new PDF" or "extract pages 12, 15, and 18 from this report."
All merged PDFs, split documents, and extracted spreadsheets are written directly to `public/data/pdf_exports/` inside your local workspace exports directory.
If you request a PDF operation without providing the specific file path or target page ranges, the skill's context logic will politely prompt you for the required parameters to prevent failed compiles.
No. All merging, splitting, page-numbering, and table extracting are executed entirely locally on your machine. You can process your PDF files completely offline.
Browse more in General Tools or all skills.
Get EasyClaw, add this skill, and start building AI agent workflows in minutes.
Get EasyClaw Free →