Optical character recognition gets talked about as if it is one uniform technology that either works or does not. In practice it is closer to a spectrum. Some content types come back essentially perfect, and others come back unreliable no matter which tool you use, including PDFHaul's. Knowing which is which before you rely on OCR output for something important, like a contract or a research citation, saves you from finding out the hard way.
What OCR Reads Well
Clean, printed text in a standard font is the easy case, and it is also the most common case. Typed reports, books, invoices, contracts printed from a word processor, and most everyday business documents fall into this category. If the scan is reasonably high resolution, the text is dark against a light background, and the font is a common one like Times New Roman, Arial, or Calibri, OCR engines recognize the characters at a very high accuracy rate. This is true across most modern OCR tools, not just PDFHaul's, because printed text in standard fonts is the scenario that OCR models have been trained on most heavily for decades.
Multiple languages are also handled well by modern OCR, provided the language is correctly identified. PDFHaul's OCR runs automatic language detection across 15 supported languages rather than requiring you to select one manually, which matters because guessing wrong on the language setting is one of the more common causes of bad OCR output on tools that require a manual selection.
Numbers, Symbols, and Special Characters
Digits and common symbols are generally read accurately, but a few specific confusions come up repeatedly across every OCR engine. The letter O and the number 0 look nearly identical in many fonts, as do the lowercase letter l, the uppercase letter I, and the number 1. Currency symbols and less common punctuation are read correctly most of the time but are worth a second glance in financial documents where a misread symbol changes the meaning of a number. Mathematical notation is a genuine weak point. Formulas, superscripts, subscripts, and specialized mathematical symbols are frequently misread or dropped entirely, because they do not follow the left-to-right character sequence that OCR is fundamentally built around.
Handwriting: The Hard Limit
This is the one category where it is worth being completely direct rather than optimistic. Printed block handwriting, the kind used on forms where each letter is written separately, is sometimes readable, but accuracy varies enormously depending on how neat the individual writer is. Cursive handwriting is, at the time of writing, still an unsolved problem for OCR broadly. This is not a limitation specific to PDFHaul or to free tools. Even the most advanced commercial and research OCR systems struggle with cursive handwriting, because unlike printed characters, cursive strokes connect and vary continuously from writer to writer with no fixed shape to pattern-match against.
If you have a scanned document that is mostly or entirely handwritten in cursive, such as an old letter or handwritten notes, do not expect OCR to produce a usable transcription. It may pull a few isolated words correctly, but treat any handwritten OCR output as a rough guess that needs a full manual read-through, not a starting point you can lightly edit.
Stamps, Seals, and Overlapping Marks
Official stamps, embossed seals, signature overlays, and watermarks placed directly on top of text create a specific problem: the OCR engine has to distinguish the actual character shapes from the visual noise sitting on top of them. Where a stamp or seal only touches the edge of a line of text, the rest of that line usually still reads correctly. Where it sits directly over the middle of a word, that word is often partially or completely lost. This comes up often in notarized documents, government forms, and older business records where a rubber stamp was applied after the document was typed.
Low-Quality and Skewed Scans
Resolution matters more than most people expect. A scan saved at a low resolution, a photo taken from an angle rather than straight on, uneven lighting that casts a shadow across part of the page, or a page that was scanned slightly crooked all reduce accuracy, sometimes significantly. Skew correction and basic image cleanup happen automatically as part of most modern OCR pipelines, including PDFHaul's, but there is a limit to how much a badly degraded source image can be recovered. If you have a choice in how a document gets scanned, a flatbed scan at a reasonable resolution will consistently outperform a phone photo taken at an angle under mixed lighting.
Tables and Structured Layouts
This deserves separating from plain paragraph text, because it fails differently. OCR is generally fine at recognizing the individual characters inside a scanned table. What it does not do on its own is understand that those characters belong to a grid, which row and column each value sits in, and which header a number relates to. Basic OCR run on a table often returns the correct words and numbers but as an unstructured wall of text with the spatial relationships lost. This is why PDFHaul's Extract Tables tool layers table-structure detection on top of OCR rather than relying on OCR text output alone, since rebuilding the grid is a genuinely separate step from reading the characters.
A Real Example
Picture a scanned notarized affidavit. The body of the document is typed in a standard font, there is a notary stamp partially overlapping the last paragraph, and a handwritten signature underneath. Run this through OCR and the typed body text comes back essentially perfect, since it is exactly the clean, printed, standard-font scenario OCR handles best. The paragraph the stamp touches is a different story: any words directly under the stamp's ink are likely to come back garbled or missing entirely, while the rest of that same paragraph outside the stamp's edges reads correctly. The handwritten signature will not be transcribed at all, because a signature is essentially cursive handwriting, and that is the one category OCR cannot reliably read regardless of which tool processes it.
The practical result is a document that is now searchable for everything in the typed body, which is usually the part people actually need to search, while the stamp area and signature still require a human to read directly from the original scan if their exact content matters.
Quick Reference
Content Type | Typical Result | Worth a Manual Check? |
Clean printed text, standard font | Near-perfect | No |
Block (printed) handwriting | Inconsistent, varies by writer | Yes |
Cursive handwriting | Unreliable across all engines | Always |
Low-resolution or blurry scan | Degraded, more errors | Yes |
Text over a stamp, seal, or watermark | Partial or missed characters | Yes |
Tables and grids | Characters read fine, structure often lost without table detection | Yes, for structure |
Math notation and formulas | Poor, frequently misread | Always |
What This Means for Your Workflow
The practical rule of thumb: the closer your source document is to clean, typed, single-language text on a plain background, the more you can trust OCR output as-is. The further it drifts from that, handwriting, heavy stamps, math notation, badly degraded scans, the more you should treat the output as a first draft that needs a human check before you rely on it for anything with real consequences, like a legal filing, a financial record, or an academic citation. OCR is a genuine time-saver for the majority of everyday scanned documents. It is not a substitute for reading the document yourself when accuracy actually matters.
Frequently Asked Questions
Does PDFHaul's OCR handle handwriting at all?
It can sometimes pick up isolated words from clear, printed block handwriting, but cursive handwriting is unreliable across every OCR engine currently available, not just PDFHaul's. Treat handwritten OCR output as a rough draft, not a finished transcription.
Will OCR misread numbers in a way that changes their meaning?
It can, particularly with characters that look alike in certain fonts, like 0 and O, or 1 and l. For financial or legal documents where an exact number matters, spot-check the OCR output against the original rather than trusting it blindly.
Why does OCR sometimes miss text near a stamp or signature?
The stamp or signature visually overlaps the character shapes underneath it, which makes it harder for the engine to separate the actual letters from the mark sitting on top of them. Text at the edge of a stamp usually survives; text directly underneath it often does not.
Does image quality actually matter that much?
Yes. A flatbed scan at a reasonable resolution will consistently outperform a phone photo taken at an angle or under uneven lighting, even though both formats are technically supported.
Why does a scanned table come back as jumbled text instead of a table?
Plain OCR reads the characters correctly but does not understand the grid structure on its own. PDFHaul's Extract Tables tool adds a separate table-structure detection step on top of OCR specifically to solve this, rather than relying on OCR text alone.
Peter
Founder of PDFHaul and Bultech
Building tools that make working with documents faster and simpler.