The Business Challenge
The business situation the organisation faced before the project began.
Static educational resources and online learning platforms often lack real-time, interactive guidance for students struggling to digest multi-chapter concepts. Standard generative AI tools frequently suffer from hallucinations or fail to reference specific, author-vetted curriculum materials when answering queries. Furthermore, traditional backend RAG setups introduce heavy server infrastructure costs and latency for simple web applications. The goal was to build a zero-cost, lightweight solution that directly grounds model responses in local course materials.
Why AI Was the Right Solution
Why AI was the right approach — and what alternatives were considered.
Static course materials and traditional keyword searches require users to manually locate and synthesize information across multiple chapters. Deterministic search rules cannot interpret open-ended learner questions or extract conceptual nuances. Generative AI provided the reasoning capabilities needed for natural language query understanding, while RAG ensured the model remained strictly bounded to verified curriculum context.
How the Solution Was Delivered
Discovery, design, development, testing and rollout — the journey, not the tooling.
- 1Curriculum Audit & Structured Knowledge Base Encoding (JSON formatting)
- 2Architecture Design for Client-Side Retrieval & In-Memory Vector Search
- 3Prototype Development for Dynamic Context Injection & Prompt Engineering
- 4Transient Credential Security Flow Implementation for API Authentication
- 5Production Deployment on GitHub Pages & Validation Testing
Key Technical & Architecture Decisions
Architecture, model selection, workflow and trade-offs.
Chose a client-side RAG pipeline over traditional backend vector databases (like Pinecone or Pgvector) to maintain a zero-cost, serverless deployment on GitHub Pages. Selected Google Gemini 2.5 Flash for its high-speed inference, cost efficiency, and strong context window performance. Implemented transient browser-side credential management so users provide their own API keys safely without exposing secrets in public repositories.
Challenges & How They Were Solved
Obstacles hit along the way and how they were overcome.
API Endpoint & Model Versioning Transitions: Resolved API routing errors by updating connection endpoints to active model versions (gemini-2.5-flash / gemini-1.5-flash).
Context Window & Retrieval Relevance: Addressed potential context noise by implementing a localized keyword search across JSON modules before prompt construction.
Zero-Server Security: Protected sensitive credentials by handling API key authentication strictly within temporary client browser state rather than storing keys in source control.
User Adoption & Change Management
How users responded — training, change management and feedback loops.
Rolled out as an embedded assistant within the course interface (assistant.html). Interactive prompt suggestions were provided to help users phrase queries effectively. Direct source citations were rendered alongside answers to build user trust in the assistant's accuracy.
Business Outcomes & Impact
The measurable outcomes from the underlying AI Deployment — with the story behind them.
Latency & Performance: Reduced response retrieval time to under 1.5 seconds by leveraging a lightweight, client-side vector/keyword search pipeline executing directly in browser memory. Zero Infrastructure Overhead: Achieved 100% cost reduction on server hosting and backend database management by deploying directly via GitHub Pages. User Privacy & Security: Eliminated server-side API credential storage risks by handling API key authentication transients entirely on the client side. Accuracy & Hallucination Reduction: Reached near-zero hallucination rates for course inquiries by enforcing strict grounding against author-verified JSON curriculum modules.
Business Outcome Categories
The story behind the numbers
By eliminating server hosting dependencies, the project achieved a 100% reduction in backend infrastructure costs. Grounding the LLM directly against author-verified course materials eliminated domain hallucinations, giving learners total confidence in automated answers. Furthermore, sub-1.5 second retrieval latency drastically improved user engagement without incurring API proxy overhead.
Lessons Learned
What surprised the team, what worked well, and what would be done differently.
- What surprised you: How fast browser-side JavaScript can execute context retrieval without backend server latency.
- What you'd do differently: Implement local storage caching for knowledge base files to speed up repeat sessions.
- What worked well: Direct prompt grounding using clean JSON context structures.
Future Opportunities
Where this solution could go next.
Expand retrieval capabilities by integrating client-side semantic embeddings (via Transformers.js) for true vector similarity search in the browser. Future iterations will support offline search caching and multi-module course expansion.
Professional Reflections
Guidance for another builder tackling a similar problem.
Prioritize lightweight, client-side architectures for edge applications whenever possible. Enforcing strict RAG grounding early in development saves significant iteration time compared to prompt tuning ungrounded models. Always design API integrations around clear error reporting to simplify front-end debugging.