Most tools that claim to "remove duplicate pages" from a PDF only catch the easy case: two pages that are byte-for-byte identical. That's trivial. The harder, more realistic case is two pages that are duplicates in every way that matters to a human, same document, scanned twice, exported from a different program, recompressed, or rotated, but aren't identical at the pixel or byte level.
We built a 19-page test file to measure the difference, ran our own tool against it, and are publishing the file, the answer key, and the results, including what we don't catch yet.
Download the challenge set (PDF) — 19 pages, 6 categories. Download the answer key — what should and shouldn't be flagged, and why. Try it yourself in PDFHaul →
What's in the test file
Category | Tests | True matches | Deliberate non-matches |
|---|---|---|---|
A -- Cross-generator equivalence | Same wording rendered by three different PDF pipelines (reportlab, pdfLaTeX, LibreOffice) | Pages 1-3 | Page 4 (different wording, same pipeline) |
B -- Recompressed image | Same source image saved at two JPEG quality levels | Pages 5-6 | Page 7 (different source image) |
C -- Shift/rotation | Same text shifted, and the same text with a real 90-degree page rotation applied | Pages 8-10 | Page 11 (different content, same rotation) |
D -- Blank pages | Three genuinely empty pages | Pages 12-14 | -- |
E -- Visually similar, substantively different | Two near-identical invoice templates differing only in invoice number and a fifty-cent amount | -- | Pages 15-16 (both are non-matches) |
F -- Simulated rescan | Same page content run through two independent noise/tilt/blur simulations | Pages 17-18 | Page 19 (different source content, same noise treatment) |
Every category was built with genuinely different source material, not simulated differences layered onto identical files. The full methodology for each section is below.
What we found when we ran it through PDFHaul
Category | Result |
|---|---|
A -- Cross-generator | Caught. Flagged as "likely duplicate," held for review. |
B -- Recompressed image | Caught. Flagged as "likely duplicate," held for review. |
C -- Shift/rotation | Caught. Flagged as "likely duplicate," held for review. |
D -- Blank pages | Caught. Flagged as an exact duplicate. |
E -- Near-identical invoices | Correctly not flagged. No false positive. |
F -- Simulated rescan | Not caught. Known limitation, not an oversight, see below. |
This isn't a marketing number. It's what our own detector actually does when pointed at our own test file, today.
Why Section F doesn't pass yet
Scanned-document duplicates are the hardest case, and we're not going to claim we've solved it when we haven't. We tried two different approaches, comparing pixel-level ink patterns with a tilt and noise tolerance, and a sub-pixel patch-alignment method, and tested both against true rescans and deliberately hard near-duplicates: the same page with a single digit changed, the kind of difference that matters most on a real scanned invoice.
Both methods failed to reliably tell a true rescan apart from a one-digit near-duplicate. The tolerance needed to absorb normal scanner noise turns out to be roughly the same size as the gap between strokes in a small digit, so a changed digit can hide inside what looks like ordinary scan noise.
Rather than ship a method we measured and couldn't trust, we're leaving scanned-page matching exact-match-only for now. We're exploring a different approach, reading the page content directly rather than comparing pixel geometry, as a possible fix, and we'll publish results if and when it actually passes this same test.
Methodology
Section A was built by rendering identical wording through reportlab in Python, pdfLaTeX, and LibreOffice, three genuinely different generation pipelines, verified byte-different but text-identical.
Section B used one procedurally generated source image, not a flat color, so compression artifacts are real, saved at JPEG quality 95 and quality 30, genuine lossy recompression rather than simulated.
Section C combined a base text page, the same text redrawn at different margins, and the same base page with a real PDF Rotate 90 attribute applied.
Section D used three pages with no content stream at all, genuinely rather than simulated empty.
Section E used two invoice-template pages with an identical layout, differing only in invoice number and a fifty-cent total, the kind of near-duplicate a detector should not merge.
Section F rasterized a base page to an image, then ran it through two independent noise simulations with different random tilt, brightness, blur, and per-pixel noise.
A note on what this is for
We're publishing this because we think duplicate-page detection as an industry mostly isn't measured at all. Tools claim to do it, and nobody checks how well. We'd rather show our actual numbers, including the gap, than make a claim we can't back up. If you run this file through another tool and get a different result, we'd genuinely like to know. Tell us at [email protected].
Peter
Founder of PDFHaul and Bultech
Building tools that make working with documents faster and simpler.

