AI Deploy Network
Training & Professional DevelopmentInternal Developer ToolingAI Case Study

​Client-Side RAG Assistant for AI Fluency & Education

AI Deploy Network Approved

​Built and deployed a client-side RAG assistant to provide real-time, grounded guidance across educational course modules. The solution leverages browser-side keyword matching and the Google Gemini API to eliminate backend server costs while ensuring high-accuracy responses.

Imran Ahmad AuwalBy Imran Ahmad Auwal
Chapter 01

The Business Challenge

The business situation the organisation faced before the project began.

Static educational resources and online learning platforms often lack real-time, interactive guidance for students struggling to digest multi-chapter concepts. Standard generative AI tools frequently suffer from hallucinations or fail to reference specific, author-vetted curriculum materials when answering queries. Furthermore, traditional backend RAG setups introduce heavy server infrastructure costs and latency for simple web applications. The goal was to build a zero-cost, lightweight solution that directly grounds model responses in local course materials.

Chapter 02

Why AI Was the Right Solution

Why AI was the right approach — and what alternatives were considered.

Static course materials and traditional keyword searches require users to manually locate and synthesize information across multiple chapters. Deterministic search rules cannot interpret open-ended learner questions or extract conceptual nuances. Generative AI provided the reasoning capabilities needed for natural language query understanding, while RAG ensured the model remained strictly bounded to verified curriculum context.

Chapter 03

How the Solution Was Delivered

Discovery, design, development, testing and rollout — the journey, not the tooling.

  1. 1​Curriculum Audit & Structured Knowledge Base Encoding (JSON formatting)
  2. 2Architecture Design for Client-Side Retrieval & In-Memory Vector Search
  3. 3Prototype Development for Dynamic Context Injection & Prompt Engineering
  4. 4Transient Credential Security Flow Implementation for API Authentication
  5. 5Production Deployment on GitHub Pages & Validation Testing
Chapter 04

Key Technical & Architecture Decisions

Architecture, model selection, workflow and trade-offs.

​Chose a client-side RAG pipeline over traditional backend vector databases (like Pinecone or Pgvector) to maintain a zero-cost, serverless deployment on GitHub Pages. Selected Google Gemini 2.5 Flash for its high-speed inference, cost efficiency, and strong context window performance. Implemented transient browser-side credential management so users provide their own API keys safely without exposing secrets in public repositories.

Chapter 05

Challenges & How They Were Solved

Obstacles hit along the way and how they were overcome.

API Endpoint & Model Versioning Transitions: Resolved API routing errors by updating connection endpoints to active model versions (gemini-2.5-flash / gemini-1.5-flash).
Context Window & Retrieval Relevance: Addressed potential context noise by implementing a localized keyword search across JSON modules before prompt construction.
Zero-Server Security: Protected sensitive credentials by handling API key authentication strictly within temporary client browser state rather than storing keys in source control.

Chapter 06

User Adoption & Change Management

How users responded — training, change management and feedback loops.

Rolled out as an embedded assistant within the course interface (assistant.html). Interactive prompt suggestions were provided to help users phrase queries effectively. Direct source citations were rendered alongside answers to build user trust in the assistant's accuracy.

Business Impact

Business Outcomes & Impact

The measurable outcomes from the underlying AI Deployment — with the story behind them.

Hours Automated
10
Cost Savings
Currency not specified
Revenue Impact
Currency not specified
Measurable Business Outcome

Latency & Performance: Reduced response retrieval time to under 1.5 seconds by leveraging a lightweight, client-side vector/keyword search pipeline executing directly in browser memory. Zero Infrastructure Overhead: Achieved 100% cost reduction on server hosting and backend database management by deploying directly via GitHub Pages. User Privacy & Security: Eliminated server-side API credential storage risks by handling API key authentication transients entirely on the client side. Accuracy & Hallucination Reduction: Reached near-zero hallucination rates for course inquiries by enforcing strict grounding against author-verified JSON curriculum modules.

Business Outcome Categories

AI Performance ImprovementRisk ReductionCost ReductionTime SavingsProductivity ImprovementEmployee ExperienceKnowledge ManagementQuality ImprovementCompliance

The story behind the numbers

By eliminating server hosting dependencies, the project achieved a 100% reduction in backend infrastructure costs. Grounding the LLM directly against author-verified course materials eliminated domain hallucinations, giving learners total confidence in automated answers. Furthermore, sub-1.5 second retrieval latency drastically improved user engagement without incurring API proxy overhead.

Reflection

Lessons Learned

What surprised the team, what worked well, and what would be done differently.

  • What surprised you: How fast browser-side JavaScript can execute context retrieval without backend server latency.
  • What you'd do differently: Implement local storage caching for knowledge base files to speed up repeat sessions.
  • What worked well: Direct prompt grounding using clean JSON context structures.
Looking Ahead

Future Opportunities

Where this solution could go next.

Expand retrieval capabilities by integrating client-side semantic embeddings (via Transformers.js) for true vector similarity search in the browser. Future iterations will support offline search caching and multi-module course expansion.

For Peers

Professional Reflections

Guidance for another builder tackling a similar problem.

Prioritize lightweight, client-side architectures for edge applications whenever possible. Enforcing strict RAG grounding early in development saves significant iteration time compared to prompt tuning ungrounded models. Always design API integrations around clear error reporting to simplify front-end debugging.