Close Menu
MyAppsPlus

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Apple Announces ‘Full Disk Access’ Changes on macOS Due to AI Agents

    October 2, 2026

    What’s the difference between artificial intelligence & superintelligence

    October 2, 2026

    I’m watching a horror movie every day in October 2026 — and today’s pick is Mike Flanagan’s mind-bending nightmare (Oct. 2)

    October 2, 2026
    Facebook X (Twitter) Instagram
    Facebook X (Twitter) Instagram
    MyAppsPlusMyAppsPlus
    Friday, October 2
    • Home
    • Breaking Tech
    • Apps & Software
    • AI & Automation
    • Android
    • iPhone & iOS
    • More
      • Reviews
      • How-To Guides
      • Deals & Discounts
      • Shop
    MyAppsPlus
    Home»AI & Automation»Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks
    AI & Automation

    Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks

    myappsplusBy myappsplusOctober 2, 2026005 Mins Read
    Share Facebook Twitter Pinterest Copy Link LinkedIn Tumblr Email Telegram WhatsApp
    Follow Us
    Google News Flipboard
    Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks
    Share
    Facebook Twitter LinkedIn Pinterest Email Copy Link

    Datalabhas released OmniExtractBench, an open benchmark for structured document extraction. It tests how accurately a system fills a JSON schema from a PDF. The benchmark pools 620 documents from 4 existing benchmarks. One deterministic scorer grades all of them and explains each decision.

    The release lands while extraction vendors publish their own leaderboards. Datalab argues those leaderboards are hard to compare or audit. OmniExtractBench is its attempt at a shared yardstick.

    Is it deployable?Yes, the scorer installs from PyPI as omni-extract-bench (v0.1.7, Python 3.11+, SciPy only) under Apache 2.0. Rerunning vendors requires your own API keys and paid credits.

    What is OmniExtractBench?

    OmniExtractBench is a structured extraction benchmark built by Datalab. Each task gives a system a PDF and a JSON schema. The system returns JSON, which is scored value by value against a gold file. The code is on GitHub, and the data is on Hugging Face under CC BY 4.0.

    The 4 flaws it targets

    Datalab’s launch post names 4 recurring problems with existing extraction benchmarks:

    • Bias: documents and scoring can favor the vendor that built the benchmark.
    • Opaque harnesses: a low score may reflect a broken harness, not a weak model.
    • Unclear scoring: readers cannot tell why a given document scored low.
    • Narrow document variety: some suites hold only dense tables, others only clean, unscanned files.

    Where the 620 documents come from

    Regulatory filing forms are the largest category, at 88 documents. 128 documents are a single page. At the other end, 33 documents over 100 pages hold 40% of all pages. Datalab’s own synthetic suite is the second largest share.

    How the scorer works

    The scorer flattens prediction and gold JSON into addresses, which are paths to single values. It normalizes each value first, so “03/31/2024” matches “2024-03-31”.

    Tables are the hard part. Compared by position, one missed row shifts every row after it. We reran the scorer on a 100-row table missing its first row. Positional comparison scored 0%, while OmniExtractBench scored 99%.

    The fix is content-based pairing with the Hungarian algorithm. ExtractBench and LongArray-Extract already align rows this way. OmniExtractBench adds a verdict layer on top.

    6 verdicts per value

    • matched: paired, and the values agree.
    • misread: paired, but the values differ.
    • unfound: gold has a value, the prediction does not.
    • fabricated: the schema allows it, gold is silent, the prediction fills it.
    • invented_item: part of a predicted row that pairs with nothing.
    • invented_field: an address the schema never declared.

    Accuracy is matched values over all verdicts. Precision divides matched values by predicted values. Recall divides them by gold values. The full rules are in the metric spec.

    The null rule

    Empty strings, None and whitespace count as omissions, so those addresses are dropped. Strings like “NA” or “-” remain real answers. This blocks a quiet exploit: padding a schema with empty optional fields to earn free matches. In our test, padded null fields added 0 verdicts.

    How it compares with other extraction benchmarks

    Sources linked on each benchmark name. Checked September 27, <a href="https://myappsplus.com/use-of-ai-automation-and-predictive-analytics-2026/” title=”Use of AI, Automation and Predictive Analytics 2026″>2026.

    Who misses fields and who invents values

    Datalab scored 10 system configurations on the full corpus. Its accurate mode led at 93.85 accuracy. Datalab balanced (93.48) and Reducto deep_extract v2 (93.47) are effectively tied. Precision and recall then show how each system fails.

    • Balanced: Datalab (both modes) and Reducto keep precision and recall within 0.6 points.
    • Leans to misses: GPT 5.6-sol posts 95.11 precision but 84.99 recall. It loses 11.88% to unfound values. Gemini and Claude show the same pattern, less sharply.
    • Leans to invented values: LlamaExtract has 93.13 recall but 86.57 precision, losing 9.03% to fabricated values. Extend loses 4.01% to invented items.
    • Low on both: Mistral OCR 4.1 and Azure Content Understanding trail on both metrics, with recall lower still.

    How each system falls short of 100%, by verdict type

    Run it yourself

    uv pip install omni-extract-bench
    oeb score --pred pred.json --gt gold.json --schema schema.json --verdicts

    Install the [benchmark]extra and run oeb benchmark to rerun vendors. Runs are resumable, and each provider needs its own credentials. Datalab also suggests testing its playground on your own documents.

    Key Takeaways

    • OmniExtractBench pools 620 documents from LlamaIndex, micro1, Extend and Datalab suites.
    • A deterministic scorer gives every value 1 of 6 auditable verdicts.
    • Content-based row pairing stops one missed row from zeroing a table.
    • Dropping null addresses stops schema padding from inflating scores.
    • Datalab accurate leads at 93.85; Datalab balanced and Reducto tie near 93.5.
    • What does OmniExtractBench measure? It measures how accurately a system extracts values from a PDF into a JSON schema, scored per value.
    • Is OmniExtractBench open source? Yes. The scorer is Apache 2.0 on GitHub and PyPI. The dataset is CC BY 4.0 on Hugging Face.
    • Who built OmniExtractBench? Datalab built it. Datalab also maintains the open-source Marker and Surya document tools.

    Thanks to the Datalab team for the resources behind this article. Datalab supported and sponsored this content.

    The post Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks appeared first on MarkTechPost.

    agentic ai AI Shorts applications Artificial intelligence Computer Vision Editors Pick Embedding Model For Devs Generative AI Machine Learning New Releases OCR open source Promote Python Software Engineering Sponsored staff Tech News technology Vision Language Model
    Follow on Google News Follow on Flipboard
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email Copy Link
    myappsplus
    • Website

    Related Posts

    What’s the difference between artificial intelligence & superintelligence

    October 2, 2026

    This new ChatGPT scam tricks you into installing malware – how to spot the trap

    October 2, 2026

    NVIDIA Announces DGX Spark 64GB: A 1-PetaFLOP Grace Blackwell Desktop for Local AI Agents, Fine-Tuning, and Inference

    October 2, 2026
    Add A Comment
    Leave A Reply Cancel Reply

    Top Posts

    Can reviews settle disputes that marked first two seasons?

    September 21, 20268 Views

    Experts call for leveraging AI, breaking key tech bottlenecks to propel advanced manufacturing

    September 19, 20265 Views

    Here’s what happens to your old computer parts you donate to Goodwill

    September 28, 20264 Views
    Latest Reviews

    Apple’s latest Mac Mini runs on a new M6 chip, and starts at $899

    myappsplusAugust 25, 2026

    New M6 and M5 Pro Mac mini pre-orders now live at Amazon

    myappsplusAugust 25, 2026

    How to watch LASK vs Celtic: Free Streams & TV Channels for Champions League 2026/27 Play-Offs

    myappsplusAugust 25, 2026
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Most Popular

    Apple’s latest Mac Mini runs on a new M6 chip, and starts at $899

    August 25, 20260 Views

    New M6 and M5 Pro Mac mini pre-orders now live at Amazon

    August 25, 20260 Views

    How to watch LASK vs Celtic: Free Streams & TV Channels for Champions League 2026/27 Play-Offs

    August 25, 20260 Views
    Our Picks

    Apple Announces ‘Full Disk Access’ Changes on macOS Due to AI Agents

    October 2, 2026

    What’s the difference between artificial intelligence & superintelligence

    October 2, 2026

    I’m watching a horror movie every day in October 2026 — and today’s pick is Mike Flanagan’s mind-bending nightmare (Oct. 2)

    October 2, 2026

    Subscribe to Updates

    Subscribe to our newsletter and get the latest tech news, app updates, AI trends, smartphone reviews, and exclusive deals delivered straight to your inbox.

    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Get In Touch
    • Disclaimer
    • Privacy Policy
    • Terms & Conditions
    © 2026 MyAppsPlus. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.