Deployed support copilot for an e-commerce and subscription demo: Claude tool-calling, policy RAG, reliability scoring, human review, and resilient background refund workflows with simulated commerce integrations.
Support workflows combine order lookups, policy questions, refunds and subscription changes. An LLM must ground responses in tools and policy evidence, handle unreliable APIs, and defer risky or uncertain actions to a person. This independent project models those workflows with seeded data and simulated Shopify/Stripe integrations; it is not a claimed client engagement.
Built a TypeScript monorepo with a NestJS API and Next.js dashboard. Claude calls order, refund-eligibility and knowledge-base tools; PostgreSQL/pgvector supports policy retrieval. Zod validates structured final responses. A reliability scorer combines model confidence with tool success, while high-risk requests route to human review. BullMQ/Redis handles idempotent, resumable refund jobs and inbound webhooks are deduplicated. Added retries, usage/cost tracking, tests and CI with real database/queue services. Deployed the containerised API to Google Cloud Run and the UI to Vercel, backed by Supabase and Upstash. Commerce/payment providers are simulated.
Delivered a working support application with chat, human escalation review and reliability/cost analytics. The documented deployment exposed and resolved background-worker CPU starvation and Redis IPv6 connection hangs on Cloud Run. The project demonstrates grounded support workflows and explicit human control over risky requests. No customer adoption, support-time savings, revenue impact or production commerce transaction volume is claimed.
Demonstrated three operational surfaces: chat testing, escalation review, and cost/tool-reliability analytics. The documented scorer blends model confidence and tool success 60/40; risky subscription actions escalate regardless of confidence. Deployment evidence records fixes for stalled refund jobs and hanging Redis connections. These are functional and engineering outcomes from an independent demo, not measured customer-business KPIs. Evidence: https://github.com/ns-0437/reformly-support-copilot and https://github.com/ns-0437/reformly-support-copilot/blob/master/docs/CASE-STUDY.md . Live UI documented in the repository: https://web-mu-kohl-77.vercel.app .