Production AI agent inside EmailZap, an AI-native Gmail app. Zap answers questions about a user's inbox and automates Gmail and Google Calendar tasks, backed by a memory layer, PII pseudonymisation and a multi-model eval harness.
EmailZap needed an assistant that could answer questions about a user's email and act on it (drafting replies, organising mail, scheduling) across Gmail and Google Calendar, at consumer scale and consumer margins. The constraints were real: 150K+ email updates a day, sensitive personal data that could not be sent to third-party LLMs as-is, provider models that change behaviour without notice, and an LLM bill that had to stay well under a dollar per active user per month.
Designed Zap as a LangGraph multi-agent system using ReAct-style reasoning, tool calling and persistent state across turns. Agents call Gmail and Google Calendar tools, and the system exposes MCP access through an architecture that extends to custom CRMs and other services. A personalised memory and context layer pulls facts from email history to tailor automated drafts. PII/SPII is pseudonymised before any LLM call, and the app passed CASA Tier 2. Supporting infrastructure: GCP Pub/Sub for Gmail notifications and split storage that caches email bodies in S3. For quality, I built a golden eval dataset and benchmarked 15 models on classification behaviour, response validity, latency and cost, added three-run majority voting to stabilise routing, and set up automated weekly drift detection. Inference cost came down through provider changes, prompt caching, batching and Rspamd-based pre-filtering with custom rules that keeps about 70% of emails away from the LLM.
- Average weekly unique app users rose about 3.7x across four calendar months. - Weekly sends of memory-enabled drafts rose about 4.4x across six-week comparison periods. - LLM cost cut by 86% with classification accuracy maintained, reaching $0.56 average ($0.38 median) monthly cost per active user. - Three-run majority voting resolved 92.5% of routing flips in evaluation. - Drift detection caught two unannounced provider model changes within six days, including a 40% drop in response length and 28% higher latency. - CASA Tier 2 certification achieved.
86% reduction in LLM inference cost at maintained classification accuracy ($0.56 average monthly cost per active user), about $6,900 a month (about $82,800 a year) at 2,000+ monthly active users. The AI-native engineering workflow also let EmailZap reduce the team from six to three engineers while keeping the same delivery cadence, saving about $120K a year. About 3.7x growth in average weekly unique users and 4.4x growth in weekly memory-enabled draft sends.