We built a 19-page test file to measure how well PDF tools actually detect duplicate pages, then published our results, including what we don't catch yet.
MCP’s critics have a point: a bad server can make an AI workflow harder than a good CLI. But Azure, Notion, Figma, GitHub, PDFHaul, and others show why the protocol still matters. The question is whether an assistant can reach the right tools and help finish the work.
How we measured our own table extraction accuracy against a real, diverse document corpus, including the parts that didn't go well. Full methodology, real code, and the files that still aren't perfect.
PDFHaul's OCR runs on Tesseract, the open source engine behind free and unlimited scanning, without the per page fees that limit most free competitors.