AI Deploy Network
AI Evaluation FrameworkCross-Industry / Horizontal 9/28/2026

Hybrid retrieval over structure-aware PDF chunks for section-level question answering

Give it a PDF and a natural-language query and it returns the pages and section headings that answer it. EmbeddingGemma vectors fused with BM25, evaluated on 45 graded queries with 12 held out: nDCG@5 0.645 to 0.728, hit@3 82% to 93%.

Hours Automated
0
Cost Savings
Currency not specified
Revenue Impact
Currency not specified

Business Challenge

Search inside long technical PDFs returns documents, not answers. Users need the specific pages and section headings that address their question, and a naive chunk-and-embed pipeline fails on exactly the content that matters: tables get OCR'd from a rendered image and lose an entire symbol column, and a running page header makes all 40 pages match every query equally well. The failure presents as a ranking problem, so teams reach for reranking, which does not fix it.

Solution Delivered

Built a structure-aware pipeline. Layout-aware partitioning types every block as title, prose, table or image, then three chunking paths that share no code so structure survives: prose grouped by section title, tables rendered to markdown from the embedded text layer via pdfplumber rather than OCR, and figures as caption plus OCR. Generic boilerplate detection drops lines appearing on 60% or more of pages. Retrieval fuses dense cosine over EmbeddingGemma-300m in ChromaDB with BM25 at alpha 0.6, chosen from a sweep showing a plateau across 0.5 to 0.8. Chunk hits collapse to pages within 55% of the top score. Deployed on Cloud Run with a PDF-viewer UI. HyDE, cross-encoder reranking and metadata scoring were each built and measured, and each lost ground, so each was rejected and reported rather than buried.

Outcomes Achieved

- nDCG@5 improved from 0.645 to 0.728 and hit@3 from 82% to 93%, measured on a 12-query held-out split never inspected during tuning. - Reading tables from the embedded text layer instead of OCR recovered a symbol column the index was missing entirely. - Boilerplate detection cleared 226 of 852 elements and 57 false section titles. - Three candidate optimisations measured and rejected on evidence: HyDE (0.710 to 0.684), cross-encoder reranking (0.728 to 0.727), metadata scoring (0.710 to 0.704). - Queries tagged by failure class (word, symbol, table, figure, multi), because those classes fail for different reasons and an average hides it. - Live on Cloud Run.

Measurable Business Outcome

Retrieval quality improved from nDCG@5 0.645 to 0.728 and hit@3 from 82% to 93% on a held-out query set, with every gain attributable to a specific extraction fix rather than an unmeasured ranking change.

Business Outcome Categories

AI Performance ImprovementProductivity Improvement