GiliSoft Formathor

Get reusable text from a PDF

Convert Office, PDF, image, and scanned files locally, then prepare output for delivery.

GiliSoft Formathor document workflow illustration

Convert PDF to Text for Editing on Windows

By GiliSoft • Windows document workflow

Plain text is useful for rewriting or indexing content, but it will not preserve columns, fonts, graphics, or page design. First determine whether the PDF already contains selectable text.

Quick answer

Try selecting a sentence. For a text-based PDF, use the PDF-to-text output in GiliSoft Formathor. For image-only pages, use an OCR-oriented task first and proofread the result before editing.

Before You Start

  • Confirm you have permission to reuse the document text.
  • Keep an original PDF copy and identify headings, tables, and reading order.
  • Test a short passage for selectable text; copied gibberish can indicate an encoding or OCR issue.

Choose the Right Route

SituationApproachWhat to verify
Selectable text PDFExport to plain textParagraph and reading order make sense
Image-only PDFRun OCR before text exportRecognized words match the scan
Complex tables or headingsConsider Word instead of TXTStructure survives well enough to edit

Convert PDF to Text for Editing on Windows: Step by Step

  1. Choose the source type

    A normal text PDF can export directly; a scan needs recognition before meaningful text output.

  2. Select the conversion task

    In Formathor, choose PDF to File or the relevant text/OCR route and write into a separate folder.

  3. Open the text output

    Inspect paragraph order, columns, bullets, symbols, accented characters, and table data.

  4. Edit in the right format

    Use a text editor for plain prose, or choose Word output instead when rich formatting and tables matter.

GiliSoft Formathor document task workspace
Choose the relevant task in Formathor and inspect its saved output.

Check the Result Before Sharing

  • Do names, dates, amounts, and punctuation match the source?
  • Did a two-column page get read in the correct sequence?
  • Were table cells or footnotes lost when formatting was removed?

Limits and Common Mistakes

Exporting text does not edit the original PDF. OCR can misread low-resolution scans and is never a substitute for proofreading critical material.

Frequently Asked Questions

Will plain text keep the PDF layout?

No. It keeps characters and some line structure but not the original page design.

Why is the exported file empty?

The source may be an image-only scan; use OCR and inspect the recognition result.

Should I convert to Word instead?

Choose Word when you need headings, tables, and a more editable document structure.

Continue with GiliSoft Formathor

Get reusable text from a PDF. Keep the source, check the output, and choose the delivery format that fits the task.