FIXLGSTOOLBOX
PDF Tools

035 · PDF

PDF Text & Image Extractor

Extract the PDF text layer and actual embedded images by page, then save them as TXT, individual images, or ZIP files.

LOCALYour PDF and extracted results are processed in the current browser and are not uploaded or stored on a server.
LOCALThe PDF, extracted text, and images stay in the current browser and are not uploaded or stored on a server.
Drop or choose a PDFOne PDF · extract the text layer and actual embedded raster images.Service limits: 1 PDF · 50MB · 200 pages · warning after 500 images · safety stop at 1,000

NEXT WORK

Next work

RELATED TOOLS

Related tools

HOW TO USE

How to use

  1. 01

    Choose a PDF file.

  2. 02

    Select Text, Images, or Text + Images.

  3. 03

    Choose all pages, checked pages, or a custom page range.

  4. 04

    Run extraction and review page-level text and embedded-image results.

  5. 05

    Download text as TXT, images individually or as ZIP, or save both under text/images folders in one ZIP.

USE CASES

Use cases

01

Organize report text by page

Save a report as TXT with explicit page markers when normal copy/paste mixes line or column order.

02

Collect photos embedded in a PDF

Find raster image objects used inside the PDF instead of turning the full visible page into a screenshot.

03

Deliver all extracted material in one ZIP

Bundle text and images under separate folders while keeping deterministic page-based filenames for downstream work.

EXPERT POST

Practical standards for extracting PDF text and embedded images

Text layers, reading order, PDF image objects, PNG fallback, duplicate resources, scanned PDFs, and OCR boundaries all affect how extraction results should be interpreted.

Text extraction is not OCR

A text PDF stores character information as document objects that PDF.js can read directly. A scan may contain only page images and no text layer. Tool 035 reports that state rather than silently running OCR.

Visual reading order can differ from internal order

Columns, tables, floating text boxes, and unusual font encodings can make PDF text-item order differ from the way a person reads the page. The tool therefore reconstructs spacing conservatively and does not promise perfect layout recovery.

Image extraction is different from full-page rendering

Tool 027 rasterizes the complete visible PDF page. Tool 035 instead follows PDF operators and image objects to find embedded raster content.

The original encoded stream is not always safely recoverable

DCTDecode JPEG resources are exported from their original compressed stream as JPG without re-encoding. Other raster images that cannot be preserved safely because of filters, masks, or color spaces use a decoded PNG fallback.

Major images and all images serve different needs

PDFs can contain photos alongside tiny icons, masks, decorative fragments, and repeated logos. The default Major images view applies conservative size, mask, and duplicate rules; All images exposes as many successfully decoded objects as possible.

Repeated XObjects are deduplicated automatically

Identical bitmaps are removed in both Major images and All images to reduce result count and ZIP size. Images with different dimensions or actual pixel content remain separate results.

Sequential page processing is safer for large PDFs

Keeping many text structures, bitmaps, blobs, and previews alive at once can cause sharp memory growth. Processing one page at a time and releasing page objects is better suited to local browser tools.

ZIP paths and deterministic filenames matter

Names such as page-001-image-001.png preserve page order in file explorers. Separating text and images also makes the bundle easier to reuse in scripts and document workflows.

IMPORTANT NOTES

Important notes

Your PDF and extracted results are processed in the current browser and are not uploaded or stored on a server.

FAQ

FAQ

01Can it extract text from scanned PDFs?+

Only when the PDF already contains a text layer. Image-only scans require OCR, which Tool 035 does not run automatically.

02Does it extract original photos from the PDF?+

It extracts embedded raster image objects rather than full-page screenshots. Depending on the PDF structure, an image may be delivered as a decoded PNG fallback instead of the original encoded stream.

03Can it turn every PDF page into JPG?+

No. Full-page conversion belongs to Tool 027 PDF to Image Converter.

04Can I extract only certain pages?+

Yes. Use all pages, checked pages, or a custom range such as 1-3,5,8.

05Are tiny icons included?+

The default Major images view hides very small decorative objects and masks. Switch to All images to inspect small successfully decoded resources.

Your PDF and extracted results are processed in the current browser and are not uploaded or stored on a server.