ExtractBench goes live on Kaggle

LlamaIndex@llama_indexSep 2, 2026

ExtractBench is now live on @Kaggle. It tests schema-guided document extraction on the documents most likely to break downstream agents and workflows, including long record lists, noisy scans, handwriting, and complex tables. The benchmark covers 370 enterprise documents across https://t.co/rTDYMUABll

Views10.0k
Comments5
Reposts6
Likes32
Launched Sep 2, 2026View post

More from LlamaIndex

1

LiteParse v2.1 is here, and its bringing the fastest markdown output possible. In this release, we are fulfilling our top request: markdown output. But in the spirit of "lite"-ness, we are doing this completely LLM-free and fast. Not only is it fast, it also beats all other https://t.co/bdnSdNsMhA

Jun 18, 2026
Views308.1k
Comments14
Reposts29
View post
2

We just built a Private Equity Assistant with LlamaAgents and the newly released LlamaCloud SDK. It can: 📊 Turn portfolio spreadsheets into structured, LLM-ready data with LlamaSheets 📂 Classify investor decks and extract key details with LlamaClassify and LlamaExtract 🤖 https://t.co/udycH9Fnng

Jan 30, 2026
Views27.6k
Comments1
Reposts8
View post
3

Let's talk parsing tables. Two days ago we launched ParseBench,the first document OCR benchmark built for AI agents. This deep dive breaks down TableRecordMatch (GTRM), our metric for evaluating complex tables the way your pipeline actually consumes them: as records keyed by https://t.co/7ZQOUqo3hb

Apr 15, 2026
Views26.5k
Comments1
Reposts11
View post
4

🚀 The team at @Google just released the Agents API, a service for building and running custom agents inside a sandboxed Linux environment, and we built a template that gives these agents access to LlamaParse / LiteParse, enabling them to process unstructured documents https://t.co/cS6Ydyt9Kt

May 19, 2026
Views22.5k
Comments9
Reposts12
View post

See all LlamaIndex launches →