The Business Challenge
The business situation the organisation faced before the project began.
EmailZap needed an assistant that could answer questions about a user's email and act on it (drafting replies, organising mail, scheduling) across Gmail and Google Calendar, at consumer scale and consumer margins. The constraints were real: 150K+ email updates a day, sensitive personal data that could not be sent to third-party LLMs as-is, provider models that change behaviour without notice, and an LLM bill that had to stay well under a dollar per active user per month.
Why AI Was the Right Solution
Why AI was the right approach — and what alternatives were considered.
Email is unstructured, personal and endlessly varied. The questions users ask ("find my boarding pass", "what did Acme say about the invoice", "set up the call this email asks for") need language understanding, retrieval across threads, and judgement about what counts as needing attention. A rules engine breaks on the long tail within days, and off-the-shelf filters sort mail but cannot answer questions or draft a reply in the user's voice.
Headcount is not an option for a consumer product. Nobody pays for a human to read their inbox inside a consumer subscription.
We were deliberate about where AI does not belong. About 70% of mail is filtered out with deterministic rules before any model sees it, because a newsletter does not need an LLM to be recognised as a newsletter. In the rebuild the split is explicit: the model proposes, the harness decides. Anything the system can know for itself (IDs, whether the user is a recipient, whether a later reply exists, whether a quote really appears in the source) is computed in code, and the model is used only for judgement: what matters, what to say, and which tool to call next. Safety rules like "never send without an explicit instruction" live in code, not in prompts.
How the Solution Was Delivered
Discovery, design, development, testing and rollout — the journey, not the tooling.
- 1Shipped Zap v1: a LangGraph multi-agent system with ReAct-style tool calling over Gmail and Calendar, persistent state across turns, and MCP access
- 2Added a memory and context layer that pulls facts from email history into drafts, with PII/SPII pseudonymised before every LLM call; passed CASA Tier 2
- 3Built a golden eval set, benchmarked 15 models, added three-run majority voting for unstable routing, and set up weekly drift detection
- 4Brought cost down with provider changes, prompt caching, batching and a rules-based pre-filter that keeps about 70% of mail away from the LLM
- 5Stepped back for v2 and redefined the agent: one-line goal, two entry points (proactive and on-demand), three first jobs, and read / prepare / commit authority tiers
- 6Pulled 30 real production emails from Langfuse traces and annotated them before designing anything; five design rules came straight from that data
- 7Fixed the design order from hardest to change to easiest: output and evidence contract, tool contracts, context, harness loop, and framework last
- 8Wrote 14 contracts (core plus Gmail and Calendar as plugins) and reviewed the spec like code; review found 9 real defects on paper
- 9Turned every open question into a numbered, owned decision with an example and a recommendation (13 decided)
- 10Phase A: implemented the contracts as code with no model, using fake providers and hand-scripted golden runs in the model's slot
- 11Ran Phase A through four approval gates and 11 vertical slices, test-first with mutation testing, using parallel coding agents on disjoint slices
- 12Result: 70 findings folded back into the spec, 591 tests, 3 end-to-end golden runs, 8 misbehaviour cases refused
- 13Next: human-reviewed eval set, then framework and model bake-offs scored against the same contracts, then one real slice end to end
Key Technical & Architecture Decisions
Architecture, model selection, workflow and trade-offs.
Single agent before multi-agent. v1 used a LangGraph multi-agent design. For v2 the lesson was to start with one agent and good tools, and add agents only when measured need demands it, since every extra agent adds cost and new failure modes.
Design from hardest to change to easiest. Output and evidence contract first (UI, evals and feedback all depend on it), then tool contracts, context layer, harness loop, and the framework last. A framework is just an engine for the loop; you cannot compare frameworks fairly until you have tools to plug in and tasks to grade them on.
The model proposes, the harness decides. The model is treated like an untrusted browser client. It works with short handles issued by the harness instead of real IDs, so it cannot invent or cross-reference records. It quotes, and the harness locates the quote in the source and records verified evidence. Answer sentences carry claim markers tied to verified claims. Result fields are split into model-authored proposals and harness-authored facts, so the model cannot say "sent" or "I checked your replies" when that did not happen.
No write tools for the model. Send, archive and calendar changes only come from a validated result shown on a confirm card, and the harness executes after the user taps. An injected email cannot talk its way into a send.
Core plus plugins. Gmail and Calendar are integrations that bring their own tools under one contract, so adding a CRM is a plugin, not a core change.
No model calls inside tools. Search returns provider metadata, not generated summaries, which keeps latency, cost and tracing honest.
Ledger vs trace. Full run traces go to Langfuse; a thin database ledger keeps what must never be lost: approvals, what the card showed, send receipts, memory changes.
Greenfield architecture, reused infrastructure. Existing drafts, search and send are reused behind adapters, after checking in code what they actually do.
Build vs buy on file safety: a managed malware scanner, and never a public scanning service that shares uploaded files.
Challenges & How They Were Solved
Obstacles hit along the way and how they were overcome.
Provider models change without notice. We saw a 40% drop in response length and 28% higher latency from unannounced updates. Fix: a golden eval set and automated weekly drift detection, so changes are caught in days with evidence.
The same model disagreed with itself on routing. Three-run majority voting resolved 92.5% of routing flips in evaluation.
Inference cost scaled with inbox size. Fixed with provider changes, prompt caching, batching, and a rules-based pre-filter that keeps about 70% of mail away from the LLM, for an 86% cost cut at the same accuracy.
Sensitive data. Personal email cannot go to third-party models as-is. PII/SPII is pseudonymised before every call, which also carried us through CASA Tier 2.
Prompt injection and untrusted content. Emails can contain instructions aimed at the agent, links, and attachments of any type. Rules now live in code: never fetch a URL from an email, scan and parse files only in a sandboxed pipeline, mark sender-controlled text as untrusted, and give the model no write tools.
Real mail is messier than invented examples. Requests addressed to someone other than the inbox owner, conditional actions ("if you didn't grant this access..."), deadlines that have passed, and a matching phrase from last year's flight. Designing from 30 real production emails turned each into a rule and a test fixture.
A spec that looks finished is not. After two review rounds, making the contracts executable still found 70 gaps. One let an approval card appear before the result was validated, so a draft with an invented address could reach the user. Another, found only by mutation testing, would have misreported a sent email on a double tap. All were fixed in code and in the spec before any model was involved.
Reuse claims were wrong in the details. "Reuse search" hid an AI query-builder call; "reuse drafts" came with constraints (one live draft per thread, replies only). Checking the code before recording the decision kept these out of production.
User Adoption & Change Management
How users responded — training, change management and feedback loops.
For an agent that touches someone's inbox, adoption is trust. v1 earned it through usage: weekly unique users rose about 3.7x, and weekly sends of memory-enabled drafts rose about 4.4x. People send drafts they trust, and drafts grounded in their own history got sent.
v1 also showed where trust is fragile: an agent that sounds confident but cannot show its evidence, or that might act on the wrong email. The rebuild designs trust in from the start:
Prepare by default, act only when asked. Reading proceeds, drafts and actions are shown, and anything that commits (send, archive, calendar change) needs a tap on a card bound to exactly what was shown. No auto-send; the hard target is zero sends without an explicit instruction.
Answers carry evidence. Claims link to verified quotes from the user's own mail, so users can check an answer instead of taking it on faith.
Memory is visible. Low-risk preferences save automatically but are shown; facts about people, sensitive topics, or anything that changes Zap's behaviour need confirmation. Old facts expire, so "my flight is tomorrow" never becomes a permanent belief.
Ask, don't guess. Mail that looks meant for someone else surfaces as "Is this yours?", and the answer shapes how similar mail is treated.
Feedback is several separate signals rather than one "Zap score": explicit feedback, opens, draft edits and sends, provider receipts, and human-reviewed eval cases, each with a known limitation.
Inside the team, the change was in how we build. Every open question became a numbered decision with an example and an owner, and the rebuild runs on written briefs and approval gates, so engineers and coding agents work from the same record instead of chat history.
Business Outcomes & Impact
The measurable outcomes from the underlying AI Deployment — with the story behind them.
86% reduction in LLM inference cost at maintained classification accuracy ($0.56 average monthly cost per active user), about $6,900 a month (about $82,800 a year) at 2,000+ monthly active users. The AI-native engineering workflow also let EmailZap reduce the team from six to three engineers while keeping the same delivery cadence, saving about $120K a year. About 3.7x growth in average weekly unique users and 4.4x growth in weekly memory-enabled draft sends.
Business Outcome Categories
The story behind the numbers
The numbers on the deployment are the visible part. What changed underneath matters as much.
Zap turned EmailZap from an inbox with AI features into an assistant people come back to. Average weekly unique users rose about 3.7x across four calendar months after launch, and weekly sends of memory-enabled drafts rose about 4.4x across six-week comparison periods. Drafts that use facts from a user's own history get sent, not just read, which was the signal we cared about most.
Cost was the other half. A consumer subscription app cannot carry an LLM bill that scales with inbox size. Cutting inference cost 86% at the same classification accuracy, to about $0.56 per active user per month ($0.38 median), changed the pricing conversation from "can we afford an agent" to "which tier gets which features". Pre-filtering about 70% of mail out of LLM processing did more for margin than any model swap.
Trust and compliance opened doors. Pseudonymising PII/SPII before any LLM call and passing CASA Tier 2 was a hard requirement for restricted Gmail scopes. Without it the product does not ship.
The rebuild changed how the team works. Weekly drift detection caught two unannounced provider model changes within six days, including a 40% drop in response length and 28% higher latency. That moved the team from "the model seems off today" to measured, dated evidence. The contract-first rebuild did the same for design: 70 gaps found before any model was involved, each fixed in minutes, each one that would otherwise have shown up in production as "the model is flaky". It also gave the company a repeatable method for the next agent, not just a better Zap.
Lessons Learned
What surprised the team, what worked well, and what would be done differently.
- Define goal, authority, correctness and speed before choosing any tech
- Design from real production data, and treat production labels as proposals, not ground truth
- Design from hardest to change to easiest: output, tools, context, harness, framework
- The model proposes, the harness decides: if the system can know it, the model does not get to say it
- Safety rules live in code; prompts are for judgement; evals measure both
- Deterministic pre-filtering often saves more cost than switching models
- Watch for provider model drift weekly; it happens without notice
- Review specs like code, then make them executable; code found 70 gaps two reviews missed
- Make decisions explicit, numbered and owned, with a concrete example and a recommendation
- Reuse infrastructure, not architecture, and check what the reused code actually does
- Separate what must never be lost (ledger) from what is nice to have (trace)
- Keep credentials in the environment, never in chat or prompts
- One owner per shared file when several agents or people work in parallel
Future Opportunities
Where this solution could go next.
The rebuild is mid-flight. Phase A (executable contracts, no model) is done. Next:
Phase B: a human-reviewed eval set covering search, answers, drafts and calendar cases, loaded into Langfuse datasets with score configs.
Phase C: bake-offs scored against the same contracts and evals. Runtime first (own loop vs Pydantic AI vs LangGraph vs an agent SDK), then model, then the method for judging importance. The contracts make this a fair comparison instead of a vibe check.
Phase D: one real slice end to end, with real Gmail and Calendar adapters, the chosen runtime, prompts tuned against evals, and a single surface.
After that: a CRM integration as the first plugin beyond Google; typed or voice approval (never the model interpreting words as approval); more proactive jobs from the same core; and applying the same contract-first method to the next agent the company builds.
Professional Reflections
Guidance for another builder tackling a similar problem.
Define the agent before you build it. Write down the goal, what it may read, prepare and change, how you will know it worked, and how fast it must be. Most agent bugs are unclear authority, not bad prompts.
Look at real data first. Production traces are the fastest source of hard, realistic examples. But production labels are not ground truth: our labels agreed with the old classifier on 26 of 27 cases, and that is agreement, not accuracy.
Pick the framework last. Output contracts, tools and context are expensive to change. The framework is cheap to swap if the layers under it are clean. Teams that start with the framework end up designing around it.
Put safety in code. A prompt that says "never auto-send" can be talked around by one malicious email. The harness decides; the model proposes.
Make the spec executable before making it smart. Scripted golden runs in the model's slot, fake providers with hard-case fixtures, and mutation testing found far more than two rounds of human review. Every gap found this way costs minutes. Found after a model is plugged in, it looks like a flaky model and costs days.
Treat confusion as a design signal. When a clear explanation was hard to give, the concept was usually wrong: a vague name, two ideas in one field, a missing boundary.
Measure cost and drift from day one. Consumer AI economics are unforgiving, and providers change models under you.
Parallel coding agents work well once one lead has built the core and the slices touch disjoint files. Give each shared file one owner and keep decisions in one written table.