
If you're building pipelines or datasets, this shows how to use Docling to preserve structure, tables, and layout when converting messy PDFs and scanned documents into clean JSON/Markdown — the difference between a working retrieval system and garbage-in-garbage-out.
“I think we can all agree that context is the most important aspect to building an AI application or an agent, right?”
Cedric Clyburn
“20 scientific papers now feature a new nonsensical term that doesn't exist because AI misinterpreted a very old article that was scanned and taken to a PDF”
Cedric Clyburn
“that data and the way you process it is the key determining factor in whether your answer is going to be correct or incorrect for the user or customer at the end of the day”
Cedric Clyburn
“we're doing RAG but without having to use a chunker or embedding model or vector database, etc., etc. So the index ends up being the markdown outline of the document.”
Cedric Clyburn
“It's fast, it's cheap, and most importantly, it's open-source.”
Cedric Clyburn
videoWhy AI Agents Need Million-Token Context — Thomas Wolf & Olive Song, MiniMax
videoYour company brain will leak secrets: how we stopped it for big banks — Tanmai Gopal, PromptQL
videoTethered: Our Agents Are Us — Shu Fang, Two Sigma
videoAgents' next frontier: agent-to-agent and network effects — Jean-Denis Greze, TownChecking sign-in…
Loading comments…