Multi-tenant Enterprise RAG assistant with role-based access control: employees ask questions in natural language and get cited answers only from documents their role and tenant are authorized to see.
Enterprise knowledge was fragmented across PDFs, policies, and SOPs. Employees spent 30–60 minutes daily searching documents, often finding outdated or irrelevant versions. Worse, a naive search/AI tool would leak restricted content — management and confidential documents were visible to everyone. The business needed fast, accurate answers with strict enforcement of who can see what, across tenants, roles, and departments, without sending sensitive data to external AI vendors (local models required).
Built a production-style RAG platform: FastAPI backend, PostgreSQL (metadata + full-text search), Qdrant (vector search over 2,000+ chunks), Celery/Redis background ingestion, and local LLM answering. Retrieval combines dense vector search + Postgres keyword search fused by Reciprocal Rank Fusion, refined by a cross-encoder reranker, with per-tenant Redis caching. Every query is filtered by tenant, role visibility (general/management/confidential), and department — unauthorized content is excluded before generation, not after. A Streamlit dashboard demonstrates login, tenant-scoped Q&A with citations, latency breakdown, and answer evaluation. Measured performance: retrieval ~1s (server-side), answers with full citations and audit-ready sourcing.
Knowledge workers get cited answers in seconds instead of manually searching documents. Access-control violations by the AI are structurally impossible since filtering happens at retrieval time. IT retains full data sovereignty with local embeddings and models — no sensitive documents leave the infrastructure. The evaluation module lets teams grade answer quality against ground truth before rollout, and per-stage latency metrics make performance bottlenecks visible and provable.
Reduced document-search time from ~30–45 minutes to under 1 minute per query (retrieval measured at ~460ms embed+hybrid, 500ms rerank, server-side). Zero unauthorized disclosures in role-based testing: same question returns full answer for admin/manager and a permission denial for employee/outsider roles. Answer cache serves repeat queries in 11ms. (Pilot-scale estimates — replace with production measurements after rollout.)