The Business Challenge
The business situation the organisation faced before the project began.
Search inside long technical PDFs returns documents, not answers. Users need the specific pages and section headings that address their question, and a naive chunk-and-embed pipeline fails on exactly the content that matters: tables get OCR'd from a rendered image and lose an entire symbol column, and a running page header makes all 40 pages match every query equally well. The failure presents as a ranking problem, so teams reach for reranking, which does not fix it.
Why AI Was the Right Solution
Why AI was the right approach — and what alternatives were considered.
The query is natural language and the target is a semantic span, so lexical search alone cannot close the gap: a user asking about a concept rarely uses the document's exact wording.
Alternatives considered. Keyword and BM25 search alone was rejected as a complete solution because it misses paraphrase entirely, but retained as half of a hybrid, since it is precise on symbols and identifiers where embeddings are weak. Dense embeddings alone were rejected for the mirror-image reason: strong on paraphrase, unreliable on exact symbols and codes. Having a general LLM read the whole document per query was rejected on cost and latency at 40-plus pages, and because it removes the traceability that makes an answer checkable.
The chosen approach was hybrid retrieval, dense vectors fused with BM25, over chunks that preserve document structure, with the LLM kept out of the ranking path entirely so results stay deterministic and inspectable.
How the Solution Was Delivered
Discovery, design, development, testing and rollout — the journey, not the tooling.
- 1Built the graded evaluation set first: 45 queries, 12 held out and never inspected during tuning
- 2Partitioned documents layout-aware, typing every block as title, prose, table or image
- 3Implemented three separate chunking paths that share no code, so structure survives
- 4Switched table extraction to the embedded text layer via pdfplumber instead of OCR
- 5Added generic boilerplate detection to strip repeated page furniture
- 6Fused dense cosine over EmbeddingGemma-300m in ChromaDB with BM25, alpha chosen by sweep
- 7Added section expansion and page aggregation so chunk hits resolve to pages
- 8Built, measured and rejected three ranking upgrades against the held-out split
- 9Deployed on Cloud Run behind a PDF-viewer UI
Key Technical & Architecture Decisions
Architecture, model selection, workflow and trade-offs.
The decision that mattered most was building the graded evaluation set before touching retrieval, with 12 of 45 queries held out. Without it, every later change would have been argued on intuition, and three of them would have been adopted wrongly.
Three chunking paths that deliberately share no code: prose grouped by section title, tables rendered to markdown from the embedded text layer, figures as caption plus OCR. Sharing code across these is what collapses a table into unusable prose.
Hybrid fusion at alpha 0.6, picked from a sweep that showed a plateau across 0.5 to 0.8, so the exact value is not load-bearing, which is itself worth knowing.
Page aggregation collapses chunk hits to pages within 55 percent of the top score, capped at eight. A fixed top-k is wrong in both directions: too many pages for a precise query, too few for a broad one. Section expansion lets a chunk inherit a decayed score from the best chunk under its heading, capped at four, so one weak hit cannot drag in an entire section.
Challenges & How They Were Solved
Obstacles hit along the way and how they were overcome.
The central challenge was misdiagnosis. Poor results looked like a ranking problem, and the industry-standard response is reranking. A cross-encoder reranker was built and measured: 0.728 to 0.727, no gain. HyDE: 0.710 to 0.684, an active loss. Metadata as a separate scoring channel: 0.710 to 0.704, also a loss. The failure was never in the ranker.
Meanwhile, reading tables from the embedded text layer rather than OCRing a picture of them recovered a symbol column the index had been missing entirely, and boilerplate detection, dropping lines that appear on 60 percent or more of pages, cleared 226 of 852 elements and 57 false section titles. The wins were all in extraction.
A secondary challenge was hardware. A 4 GB VRAM ceiling ruled out llava-1.5-7b for figure captioning, so llava-interleave-qwen-0.5b was used instead, a constraint that shaped the architecture rather than being worked around.
User Adoption & Change Management
How users responded — training, change management and feedback loops.
This was built and deployed as an independent technical project with a public live demo on Cloud Run, not as an organisational rollout. Adoption was designed around inspectability: the PDF-viewer UI shows the retrieved pages in context, so a reader can verify the answer against the source rather than trusting a ranking.
No adoption-volume or user-productivity claim is made. There is no usage instrumentation behind this record.
Business Outcomes & Impact
The measurable outcomes from the underlying AI Deployment — with the story behind them.
Retrieval quality improved from nDCG@5 0.645 to 0.728 and hit@3 from 82% to 93% on a held-out query set, with every gain attributable to a specific extraction fix rather than an unmeasured ranking change.
Business Outcome Categories
The story behind the numbers
Retrieval quality improved from nDCG@5 0.645 to 0.728 and hit@3 from 82 percent to 93 percent, measured on a 12-query held-out split never inspected during tuning. Every gain traces to a specific extraction fix rather than an unmeasured ranking change.
A 13 percent lift in nDCG@5 is the headline, but the more useful number is three: the count of reasonable-sounding improvements that made things worse. HyDE with Qwen2.5-1.5B-Instruct cost 0.026 nDCG, taking 0.710 down to 0.684. Cross-encoder reranking with BAAI/bge-reranker-base moved 0.728 to 0.727, nothing, while adding a model to the serving path. Metadata as a separate scoring channel cost 0.006. Each is a standard recommendation, each was implemented properly, and each was dropped because the held-out split said so. Had the evaluation set been built after the optimisations rather than before, at least one would have shipped on the strength of the idea alone.
The gains came instead from unglamorous extraction work. Reading tables from the embedded text layer recovered a symbol column the index was missing entirely, and boilerplate detection cleared 226 of 852 elements and 57 false section titles.
This is a builder-reported technical outcome. No client-confirmed hours saved, cost saving or revenue impact is claimed.
Lessons Learned
What surprised the team, what worked well, and what would be done differently.
- A ranking problem is often an extraction problem in disguise; inspect what is in the index before tuning how it is scored
- Build the graded evaluation set before the optimisations, and hold part of it back
- Report the changes that lost ground; three rejected optimisations say more about a system than one adopted one
- Tables must be read from the embedded text layer, never OCRed from a rendered image
- Repeated page furniture is not harmless noise; it flattens the entire relevance distribution
- Chunking paths for prose, tables and figures should share no code
- A parameter plateau is a finding: alpha anywhere from 0.5 to 0.8 performed the same
- Hardware constraints belong in the architecture, not in a list of regrets
Future Opportunities
Where this solution could go next.
The clearest next step is broadening the evaluation corpus. 45 graded queries over one document class is enough to reject an optimisation but not enough to generalise the alpha choice or the 55 percent page-aggregation threshold.
Beyond that, per-class failure reporting in production, since queries already carry tags (word, symbol, table, figure, multi) and those classes fail for different reasons that an average hides. Then a confidence signal on the returned pages, so a low-confidence result can say this document may not contain the answer rather than returning its best eight pages regardless.
Professional Reflections
Guidance for another builder tackling a similar problem.
Build the evaluation harness first. It is the least interesting part of a retrieval project and the only part that tells you whether anything else worked.
When results are poor, read your index before you tune your ranker. Open the actual chunks. A table rendered as garbled prose, or a header repeated on every page, is visible in ten minutes and invisible in any metric dashboard.
And publish the negative results. Three measured rejections are more credible evidence of rigour than a single reported improvement, and they save the next person the same three weeks.