Why Is My PDF So Large, and What Will Actually Shrink It?
Short answer
There are five causes, and the fix only works if it matches the cause. Scans need downsampling, print-resolution images need a re-export from the source, embedded fonts need subsetting, form and layer data needs flattening, and leftover edit history needs a lossless rebuild. Identify which one you have first — compressing the wrong thing achieves nothing.
On this page
A 2MB PDF and a 60MB PDF can look identical on screen. The difference is in what the file is actually storing, and there are only five realistic answers. Each one responds to a completely different fix, and applying the wrong fix either does nothing measurable or quietly damages the document.
So before you compress anything, spend a minute working out which case you have. The companion guide why is my PDF file so large covers the per-page arithmetic that gets you started. This guide is the decision table: cause on the left, the fix that actually helps on the right.
The diagnostic table
| Cause | How to recognise it | What actually fixes it | Typical saving |
|---|---|---|---|
| Scanned pages | Text will not select; over 1MB per page | Downsample to 150 DPI, greyscale | 85–97% |
| Print-resolution images | Text selects; photos look sharp when zoomed to 400% | Re-export from source at screen quality | 50–80% |
| Embedded fonts | Small file, many fonts, few images | Subset fonts on re-export | 10–40% |
| Form or layer data | Fillable fields, or layers from a design tool | Flatten the document | 10–30% |
| Edit history | File grew after editing; size does not match content | Lossless rebuild or “save as” | 20–60% |
The rest of this guide works through each row.
Cause 1: scanned pages
This is the single most common reason a PDF is enormous, and it is also the one with the most dramatic fix available.
A scanner does not produce text. It produces a photograph of a page, and every pixel of that photograph is stored. An A4 page scanned in colour at 600 DPI is roughly 4960 x 7016 pixels — about 35 million pixels. Even after the scanner compresses it, you are looking at several megabytes per page.
How to recognise it: try to select a word. If nothing highlights, or the whole page highlights as one block, it is a scan. The per-page arithmetic confirms it — anything over 1MB per page is almost certainly scanned.
What fixes it: downsampling the page images, and converting colour scans of black-and-white documents to greyscale. Resolution savings scale with the square of the DPI, so halving the resolution removes three quarters of the pixels. Going from 600 DPI colour to 150 DPI greyscale is typically a 90–97% reduction.
What does not fix it: ordinary lossless compression. Lossless compression works by removing redundancy — duplicate fonts, unused objects, an inefficient cross-reference table. A photograph has none of that. You can run a scan through a lossless optimiser all day and save almost nothing.
Be honest about the tradeoff: downsampling is lossy and irreversible. Keep the original until you have checked the result at 100% zoom on the page with the smallest text. The details are in how to compress a scanned PDF.
Cause 2: images at print resolution
The second most common case, and the one that catches people exporting from Word, PowerPoint or InDesign.
A PDF built for a printing press keeps images at 300 DPI or higher. Displayed on a screen at 96–150 DPI, everything above that is data nobody can see. The document looks exactly the same and weighs four times as much.
How to recognise it: text selects normally, so it is not a scan, but the file is still 300KB–1MB per page. Zoom a photo to 400% — if it is still sharp, you are carrying print resolution you do not need.
What fixes it: re-exporting from the source document with a screen-oriented preset. In Word this is Save As → PDF → Options → Minimum size. In InDesign it is the Smallest File Size export preset. Re-exporting beats compressing the PDF afterwards, because you avoid creating the excess data rather than discarding it after the fact.
One specific trap. Cropping an image in Word or PowerPoint hides the cropped area but keeps storing it. A photo cropped to a tenth of its frame still carries all of the original pixels into the PDF. In Word, select the image, go to Picture Format → Compress Pictures, and tick Delete cropped areas of pictures. That one checkbox can halve a document.
If you no longer have the source file, downsample the PDF instead — our PDF compressor does this in the browser. For individual images you are placing into a document, the image compressor handles them before they ever reach the PDF.
Cause 3: embedded fonts
Every font family embedded in full adds roughly 100–500KB. On a normal document that is a rounding error you can ignore.
It stops being a rounding error on documents assembled from many sources — a report combining sections from different authors, each with their own template. A dozen embedded font families can add several megabytes to a file with almost no images in it.
How to recognise it: the file is mostly text, has few or no images, and is still 200KB+ per page. In Acrobat Reader, File → Properties → Fonts lists everything embedded.
What fixes it: font subsetting, which stores only the characters actually used rather than the complete typeface. Most modern exporters subset by default. Older exporters, some LaTeX configurations, and anything that embeds a full CJK font do not — and a full CJK font is measured in megabytes, not kilobytes.
The honest limit here: subsetting is worth 10–40% on a font-heavy document and nothing at all on anything else. Do not reach for it unless the fonts list is genuinely long.
Cause 4: forms, layers and attachments
Less common, but it explains the files that should be small and inexplicably are not.
- Interactive form fields carry field definitions, appearance streams for each state, and sometimes JavaScript.
- Layered content from Illustrator or InDesign keeps hidden layers stored in full. You cannot see them; the file still carries them.
- Embedded file attachments. A PDF can carry arbitrary files inside it — spreadsheets, other PDFs, images.
- Page thumbnails cached by older versions of Acrobat.
What fixes it: flattening. This merges form fields and layers into fixed page content, dropping the interactive machinery.
Warn yourself before you do this. Flattening is irreversible. The values in a filled form are preserved visually but can no longer be edited, extracted or validated. If anyone downstream needs to process the form data programmatically, flattening breaks that permanently. Keep an unflattened master copy.
Cause 5: leftover edit history
The invisible one, and the cause people find most surprising.
Most PDF editors implement saving as an incremental update: new content is appended to the end of the file and the old content is marked superseded but left in place. The file grows with every save. Ten rounds of edits can mean ten near-complete copies of the document inside one PDF.
How to recognise it: the file grew noticeably after editing, and its size bears no relationship to what you can see on the pages.
What fixes it: a full rewrite rather than an incremental save. In most editors, Save As to a new filename does this where plain Save does not. A lossless rebuild through a compressor achieves the same thing.
This is also a privacy matter, not just a size one. Text you deleted may still be recoverable from the superseded data in the file. Redacting by drawing a black box over text is the classic version of this mistake — the text is still there, underneath. If a document has been through several editing rounds and is going to someone outside your organisation, rebuild it.
Myths worth discarding
- “Zip it.” PDFs are already internally compressed. Zipping saves a few percent and makes the file harder for the recipient to open.
- “Compress it twice for extra savings.” Each pass re-encodes already-degraded images. Two rounds look considerably worse than one for almost no additional saving. Go back to the original and compress once, correctly.
- “Delete some pages.” If images are the problem, removing pages mutilates the document for a fraction of the saving that fixing the images would give.
- “Print to PDF will shrink it.” Sometimes, by accident. It also rasterises text, destroys any text layer, and breaks links and bookmarks. It is not a compression tool.
- “Any online compressor will do.” Contracts, payslips, medical letters and anything carrying a signature should not be uploaded to an unknown server. Our compressor runs entirely in your browser for exactly this reason.
Where to go next
Once you know the cause, the matching guide has the detail:
- Scans — how to compress a scanned PDF
- General size reduction — how to reduce PDF size
- Keeping quality intact — compress PDF without losing quality
The one thing that consistently wastes people’s time is compressing before diagnosing. Thirty seconds on the table above saves an hour of trying tools that were never going to help.