I build LLM systems and the machinery that keeps them honest: agentic orchestration that runs unattended, retrieval that actually grounds an answer, and the evaluation harnesses that tell you when a model has quietly gotten worse.
Most of my work starts from the same observation. ML systems rarely fail with a crash, they fail quietly. A recommendation drifts across sessions, a detector's per-class recall collapses after a retrain, an agent reasons confidently from a false premise. So I spend as much time on benchmarks, failure analysis and supervision as I do on the models themselves.
Read more
In production I built an agentic stock intelligence platform on GPT-4o with a custom MCP orchestration layer, shipped as a containerised five-service system on GCP, producing investment signals across 24 equity buckets daily with no human in the loop. Alongside it I fine-tuned an RT-DETR detector for real-time drone perception to over 90% mAP at under 33 ms per frame, and built the benchmark suite that regressed model quality version over version.
Independently I build agent reliability and retrieval tooling: AgentFuse, an open-source circuit breaker that detects tool loops, goal drift, logic traps and runaway spend in long-running agents, and a hybrid retrieval system that lifted nDCG@5 from 0.645 to 0.728 on a held-out split by repairing extraction rather than reranking.
B.Tech in Computer Science and Engineering from IIIT Vadodara. Currently an AI Engineer at BBI Inc. I work on agent orchestration, retrieval and model evaluation, and I am open to implementation work in those areas.