
Converts scanned PDFs, screenshots, and other image-based documents into editable PDF, Word, or TXT files with one click
OCR Document Converter (OCR文档转换器) is an automated, local document extraction and file compilation skill on EasyClaw. Designed specifically for lawyers, researchers, and database administrators, it converts scanned PDFs, screenshot images, and un-selectable documents into fully editable Word (`.docx`), PDF, or text (`.txt`) files with a single conversational command, while maintaining the original visual layouts, column grids, and paragraph margins with high fidelity.
The skill is built for legal paralegals digitizing signed paper agreements, business researchers compiling data tables from textbook scans, and admin offices converting bulk image batches locally.
The expected outcome is a compiled, editable corporate document: written directly to your workspace exports directory, paired with an instant local download link, completed entirely offline.
1. Analyze input format and paths. Provide local absolute paths to your scanned PDFs, screenshots, or folders of images, and specify your target output format (Word, editable PDF, or TXT).
2. Local OCR engine initialization. The skill launches a secure, local optical character recognition (OCR) engine (such as Tesseract OCR or localized layout extraction libraries) in the background.
3. Multi-element layout parsing. Bypasses basic line-by-line reading. The engine analyzes the document structure, identifying and separating paragraphs, tables, header grids, and image blocks.
4. Text extraction and conversion. The parsed text is extracted, corrected for common OCR typos, and compiled into the target file format, recreating tables and margins with high fidelity.
5. Output delivery. The completed, fully editable document is written directly to `public/data/ocr_exports/` in your local workspace and returned with a direct download link.
- Scanned PDF to Word converter: Converts un-selectable scanned PDFs into editable `.docx` files.
- Bulk image OCR queue: Processes folders of screenshots and scans, extracting text into structured `.txt` logs.
- Fidelity layout preservation: Recreates tables, headers, and column margins in the final Word files.
- 100% local processing: Converts and extracts text offline without external cloud API key requirements.
- Image-PDF to editable PDF: Converts locked PDFs into searchable, selectable PDF documents.
- Multi-language OCR support: High-accuracy extraction for Chinese (Simplified/Traditional) and English text.
1. Converting a scanned paper agreement into an editable Word document
A legal assistant is handed a printed, signed 10-page contract that needs several revisions. Instead of re-typing the entire agreement manually, they point the skill to the scanned PDF file. The OCR engine reads the pages, parses the paragraphs, tables, and signature lines, and converts the document into a clean, editable Word (`.docx`) file with high layout fidelity in seconds.
2. Extracting text from a bulk folder of screenshots
A researcher has captured 50 screenshot images of research tables and textbook pages. They ask the skill to process the batch. The converter runs the OCR queue headlessly over the folder, extracting the text from each image, and compiles them into structured, searchable local text files grouped by category.
3. Converting a locked image PDF into a searchable PDF
A student downloads an academic ebook that is stored entirely as image scans, preventing them from highlighting or searching text. The skill processes the book: running full-text OCR, injecting a searchable text layer underneath the original images, and generating a selectable, searchable PDF while keeping the original design 100% intact.
4. Digitizing structured billing tables from receipts
An accountant has a scanned invoice with complex billing tables. Standard OCR tools scramble the column data into messy text blocks. The skill’s table-parsing engine detects the column borders, maps the cells, and reconstructs the database grid cleanly into an editable Word table or Excel CSV ledger.
5. Automatically prompting for missing target formats
A user types a vague instruction: *"Convert this file"* without providing details. The skill's context logic detects the missing variables and prompts politely: *"To run the OCR conversion, could you please specify: Are we converting a scanned PDF or an image? And what target format — Word, editable PDF, or TXT — should we export?"* preventing failed compiles.
An administrator needs to convert a scanned PDF contract to an editable Word document.
1. They open EasyClaw and activate OCR Document Converter.
2. They run: *"Convert 'D:\\project\\scans\\Signed_Agreement.pdf' into an editable Word document."*
3. The skill validates the file path, launches the local OCR engine, and parses the document layouts.
4. It extracts the Chinese and English text, structures the tables, and writes the `.docx` file.
5. It saves the completed document to: `public/data/ocr_exports/Signed_Agreement_Editable.docx` and returns the download link.
Scanned documents converted to editable formats headlessly in under 45 seconds.
Add this skill to your EasyClaw workspace
Describe your task in a chat message
Review the output and iterate if needed
Export or share the results directly from EasyClaw
Combine with other skills to build automated workflows
No. The core image parsing, text extraction, and document recompilings are executed entirely locally in your workspace container, requiring no external paid API configurations or cloud subscriptions.
The skill natively supports scanned PDFs, screenshot images (`.jpg`, `.jpeg`, `.png`, `.webp`), and image-based PDFs, and can export them as editable Word (`.docx`), searchable PDF, or plain text (`.txt`) files.
Yes. The OCR engine is highly optimized to read and extract both Simplified and Traditional Chinese characters, as well as English text, with high character-recognition accuracy.
Yes. In compliance with strict corporate security and data privacy standards, all document reading, OCR parsing, and file writes are executed entirely locally on your machine, ensuring your data is never leaked.
If you request an OCR conversion without specifying the target file or format, the skill's context logic will politely prompt you for the specific file path and target output format to prevent failed compiles.
The engine parses the raw image layout before extracting text, identifying page zones (headers, paragraphs, tables). When compiling the final Word file, it places the extracted text inside matching CSS-like style containers to recreate the original design.
Yes. All OCR parsing, text extraction, and file writes are executed entirely locally on your machine, allowing you to run all document tasks completely offline.
A searchable PDF is a document where the scanned page images are preserved visually, but a hidden, selectable text layer is injected underneath each page, allowing users to search, highlight, and copy text.
Yes. You can instruct the skill: "convert all images in D:\\project\\scans\\ to TXT files." The background queue will parse and extract text from each file sequentially, saving them in your workspace.
Yes. Every successful document conversion, image OCR, and file export is logged locally under `public/data/logs/` in your workspace, creating a clear local history of your operations.
Browse more in General Tools or all skills.
Get EasyClaw, add this skill, and start building AI agent workflows in minutes.
Get EasyClaw Free →