Session

Documents Are Harder Than Models

The model is the easy part. The hard part arrives as a 340 page scanned PDF, rotated ninety degrees, in two languages, containing a table that spans eleven pages and a signature sitting directly on top of the number that matters.

This is a deep dive into the unglamorous work behind every document-answering agent: getting real enterprise files into a state where retrieval and reasoning can do anything useful at all.

We will cover ingestion of formats nobody chose, and OCR quality as the upstream determinant of everything downstream. Why layout is semantics: a number in a table cell means something different from the same number in a paragraph, and flattening a document to plain text throws that distinction away for good. Chunking strategies that survive real documents rather than blog-post documents. Tables, forms and multilingual corpora. And what to do about the small percentage of files that will never parse cleanly, because there will always be some.

On the retrieval side: why naive vector search underperforms on structured business documents, where hybrid approaches earn their complexity, metadata as a first-class retrieval signal, and citation back to a page and a region so that a human can verify the answer. For most enterprise use cases that verifiability is not a nice extra, it is the actual requirement.

Built and operated in production, on documents I did not get to choose. Live demo, with genuinely awful inputs.

Takeaways

- An ingestion pipeline for messy real-world enterprise documents
- Why OCR and layout extraction quality dominates final answer quality
- Chunking and table handling strategies that hold up on real files
- Retrieval tuning for structured business documents, beyond naive vector search
- Verifiable citation: getting an answer back to a page, a region and a source


Preferred duration: 45 minutes including Q&A. Can be delivered in 30 or 60 minutes on request.

Target audience: engineers and architects building retrieval or document processing systems. No machine learning background required.

Level: intermediate.

Language: the patterns are language-agnostic; code examples are in Go.

Technical requirements: my own laptop (USB-C / HDMI) and internet access for the live demo. A recorded fallback is always available.

First public delivery: not yet delivered.

Marc Arndt

VP Engineering and Architecture at Evana AG

Heidelberg, Germany

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top