Datalab Launches OmniExtractBench to Fix Benchmark Bias
Datalab introduces OmniExtractBench, a shared evaluation yardstick designed to fix bias and opacity in document extraction benchmarks.

Stock photo for illustration only, not from the actual event
- Datalab launches OmniExtractBench as a shared document extraction benchmark
- Addresses bias and opacity issues found in vendor-published leaderboards
- Utilizes content-based pairing via the Hungarian algorithm instead of position
- Available for installation from PyPI (v0.1.7) with open-source code on GitHub
As extraction vendors continue to publish their own individual leaderboards, Datalab points out that these metrics remain difficult to compare or audit objectively. To address this issue, the company developed OmniExtractBench as a shared yardstick designed to establish a transparent and standardized evaluation framework across the industry.
The benchmark operates by providing a system with a PDF document and a target JSON schema. The system then returns a JSON output, which is scored value-by-value against a gold standard file. The underlying source code is hosted on GitHub, while the dataset is publicly available on Hugging Face under a CC BY 4.0 license. Developers can easily install the scorer from PyPI as omni-extract-bench version 0.1.7, which requires Python 3.11 or higher and relies exclusively on SciPy under the Apache 2.0 license.

Photo by Jaime Lopes / Unsplash
Datalab's launch announcement highlights four recurring issues currently affecting existing extraction benchmarks. Regulatory filing forms form the largest category with 88 documents, alongside 128 single-page documents. On the other end of the spectrum, 33 documents exceeding 100 pages account for a full 40% of all pages combined, while Datalab's own synthetic suite represents the second largest share.
Tables present the most significant challenge in data extraction. When evaluated by position, a single missing row shifts all subsequent rows, resulting in a positional comparison score of 0%. In contrast, OmniExtractBench achieved a 99% score on the exact same test case. The solution relies on content-based pairing using the Hungarian algorithm, with OmniExtractBench introducing an additional verdict layer on top of existing approaches like ExtractBench and LongArray-Extract.
The introduction of a standardized benchmark like OmniExtractBench represents a crucial milestone for automated document processing. Vendor-specific leaderboards often lack reproducibility and transparency. By adopting a content-matching methodology via the Hungarian algorithm, the framework mitigates table-shift errors and provides developers with an auditable yardstick to measure true extraction performance.
In evaluations across the full corpus covering 10 system configurations, Datalab's accurate mode led the rankings with an accuracy of 93.85%. Datalab balanced at 93.48% and Reducto deep_extract v2 at 93.47% finished in a close tie, while precision and recall metrics help illustrate the specific failure modes of each evaluated system.
Source: MarkTechPost
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment