The Business Challenge
The business situation the organisation faced before the project began.
Thistle, a $100M+ revenue food-tech business, needed more engineering throughput to keep pace with growth, without adding to its engineering headcount. The answer was a new engineering pod in India, run as a fractional engagement through Studio Management. That brought its own risks: a new, small team had to ramp up on an existing production codebase, work across time zones from the core team, and ship at the pace the business needed while keeping review quality high. With only two engineers, adding people or review rounds was not an option, so the process itself had to get faster.
Why AI Was the Right Solution
Why AI was the right approach — and what alternatives were considered.
The constraint was fixed headcount. Hiring more engineers would have added cost, ramp-up time and cross-time-zone coordination, and adding review rounds to a two-person team would have made it slower, not safer.
The work suited AI. A large share of each ticket was implementation inside an existing codebase with established patterns, which coding agents handle well once the intent is pinned down. The scarce resource was senior judgement, so the goal was to spend human attention only where it changed outcomes (product intent, architecture and business-critical logic) and let agents do the rest.
AI also made the process cheap to change. When we learned something, we updated a skill and every future agent session picked it up, which is much faster than retraining people on a new checklist.
How the Solution Was Delivered
Discovery, design, development, testing and rollout — the journey, not the tooling.
- 1Discovery: mapped the existing codebase, release process and where the three days of ticket-to-PR time actually went
- 2Encoded the SDLC as skills: product spec, architecture, design and implementation slices, each producing an artifact a human could approve quickly
- 3Set up Conductor so each engineer ran several Claude Code and Codex sessions in parallel, one isolated workspace per slice
- 4Routed models through OpenRouter so a different model reviewed what another had written
- 5Piloted a lights-out factory: agents specced, built, reviewed and merged with no human reading the code
- 6Rolled lights-out back after it let through changes that passed their tests but were wrong for the business
- 7Introduced smart review: every PR carries proofs and a core/leaf split, and humans review the core
- 8Made it the pod's default workflow and tracked ticket-to-PR lead time as the headline metric
Key Technical & Architecture Decisions
Architecture, model selection, workflow and trade-offs.
Spec first, code last. We built a chain of skills that mirrors how a senior team works: a product spec skill to pin down behaviour and acceptance criteria, an architecture skill for data model, contracts and boundaries, a design skill to map Figma screens to components and states, and an implementation skill that breaks the work into small, independently shippable slices. Each stage produced an artifact a human could approve in minutes before code was written.
Parallel, isolated agents. Conductor let each engineer run several Claude Code and Codex sessions at once, each in its own git worktree, so slices were built in parallel without stepping on each other.
Multi-model through OpenRouter. We did not bet on one model. OpenRouter gave us one interface to route work across models, and the model that reviewed a change was never the one that wrote it. Cross-model review caught more than self-review.
Proofs over trust. Instead of asking a reviewer to believe a PR, the agent had to show it: tests mapped to each acceptance criterion, output from actually running the change, screenshots for UI work, and a short argument for why the edge cases were covered. A PR without proof went back to the agent, not to a human.
Core versus leaf. The agent classified every change. Core was code where a mistake is expensive or hard to spot: business logic, data models and migrations, order and pricing flows, auth and external contracts. Leaf was code that is easy to verify and cheap to fix: UI components, glue code, adapters, tests and copy. Humans reviewed core line by line; leaf was accepted on the strength of its proofs and the automated review.
The trade-off: smart review is slower than lights-out on paper, but it kept the team's understanding of the system intact, and that is what let two engineers debug production quickly.
Challenges & How They Were Solved
Obstacles hit along the way and how they were overcome.
Lights-out did not work. Our first design had agents take a ticket through spec, build, AI review and merge without a human reading the code. It was fast, and it failed in ways that are easy to miss. Agents wrote tests that confirmed their own misunderstanding, so a change could be green in CI and still wrong for the business. Ambiguous tickets were resolved by the agent guessing instead of asking. Reviewer agents were good at style and obvious bugs but weak at catching a subtly wrong rule in an order or pricing flow. And because nobody had read the code, the engineers gradually lost their mental model of the system, so fixes took longer when something broke. We rolled it back.
The answer was not a return to full manual review, which would have given up the speed. It was targeted review. The spec skill forced ambiguity to surface before code was written. Proofs made agents demonstrate behaviour against acceptance criteria instead of against their own tests. The core/leaf split meant humans still read every line that mattered, while most of the diff by volume was verified by evidence.
Getting the classification right took iteration. Early on, agents labelled too much as leaf. We tightened the rules in the skill with concrete examples from the codebase, and any uncertain case defaulted to core.
Context was the other recurring problem. Agents did not know the codebase's conventions out of the box, so we encoded them in skills and repository instructions, and updated them whenever review caught the same mistake twice.
User Adoption & Change Management
How users responded — training, change management and feedback loops.
The pod was new, so there was no old process to unlearn, but trust had to be built in both directions. Engineers worried that reviewing only core code meant owning bugs they had not read, and the core team wanted to know that faster did not mean riskier.
We introduced the workflow in stages. Skills and parallel agents came first, with full human review, so the team could see what agents produced before anything was relaxed. Lights-out was run openly as an experiment, and when it failed we said so and explained what we were changing. Smart review then rolled out with one simple rule: when in doubt, it is core.
Proofs did most of the trust-building. A PR that lists the acceptance criteria, the tests that cover them and the result of running it is far easier to approve with confidence than a large diff. Recurring review findings were fed back into the skills, so the engineers could see the system improving. It became the pod's default way of working rather than a tool people had to remember to use.
Business Outcomes & Impact
The measurable outcomes from the underlying AI Deployment — with the story behind them.
Ticket-to-PR lead time reduced from about 3 days to about 1 day (about 67%) after introducing AI-assisted development and automated code review, saving about 160 engineering hours a month (about 40 tickets a month at about 4 hours saved each). Engineering supported 30% revenue growth with no additional engineering hires, avoiding two SWE hires and saving about $250K per year, and contributed about $1M in revenue impact.
Business Outcome Categories
The story behind the numbers
Ticket-to-PR lead time fell from about three days to about one. For a two-engineer pod working across time zones from the core team, that changed how much could ship in a sprint. Most of the saving did not come from agents typing faster. It came from removing waiting and rework: specs and architecture were agreed before any code existed, so fewer PRs came back, and reviewers spent their time on the small part of each change that carried real risk instead of reading every line.
The pod contributed to 30% revenue growth at Thistle without additional engineering hires. The business got more delivery capacity while keeping the cost and coordination overhead of a two-person team.
Smart review also gave the business a quality story it could trust. Every change arrived with evidence that it worked, and every line of business-critical code still had a human who had read and owned it. That mattered to stakeholders as much as the speed-up, because it meant shipping faster did not mean shipping blind.
Lessons Learned
What surprised the team, what worked well, and what would be done differently.
- Lights-out code factories optimise for speed and quietly erode quality and the team's understanding of the system
- Tests written by the agent that wrote the code are not independent evidence
- Most human review time was going to code that did not need it; the core/leaf split fixed that
- When classification is uncertain, default to core
- Spec, architecture and design skills saved more time than faster code generation did
- A different model reviewing the code caught more than self-review
Future Opportunities
Where this solution could go next.
Make proofs machine-checked instead of agent-reported, so CI confirms every acceptance criterion has a passing test before a human is asked to look.
Turn core/leaf classification into a CI step with a risk score based on code ownership, change history and past incidents, instead of relying only on the agent's judgement.
Add role-based agents with different competencies, such as a staff-level reviewer for core changes and an SRE agent for deploy and rollback checks.
Bring voice into the loop, so grooming and review conversations feed straight into specs instead of being retyped.
Track stability alongside speed, with change failure rate and time to restore next to lead time, to show the speed-up is not being paid for in incidents.
Professional Reflections
Guidance for another builder tackling a similar problem.
Do not remove humans from review. Move them. Lights-out is tempting because agents can write, test and review their own code, but an agent grading its own work is not evidence. The leverage comes from deciding where human judgement matters and protecting it there.
Put the effort into the steps before code. Most bad agent output traced back to an unclear spec or a missing architectural decision, not a weak model. A clear spec and a small slice beat a stronger model with a vague ticket.
Ask for proof, not confidence. Make agents show that a change works against acceptance criteria, and send anything without proof back to the agent.
Encode your process as skills and keep editing them. Every repeated review comment is a missing line in a skill.
Use more than one model, and never let the same model write and review the same change.