The Business Challenge
The business situation the organisation faced before the project began.
The offer (fractional engineering leadership plus a managed senior team in India: I take over infrastructure, delivery and hiring, then hand the team over with its own lead) sells to people who do not search for it. Most do not know "fractional CTO" is a category. What they do have is a public footprint: a seed round on a funding site, a founding engineer post on Work at a Startup, a CTO departure, an acquisition by a search fund, or a CEO-written post in "Ask HN: Who is hiring?".
No single database holds these signals. A person could read all of it, but not every week and not consistently.
Constraints I set:
- No cold email leaves without a human reading the row and approving it.
- No invented facts in any email. The first line must state something checkable.
- No guessed email addresses passed off as found. Guesses only if verified, and labelled otherwise.
- Every limit adjustable without opening n8n. The whole system operable from a spreadsheet.
- Cheap: free tiers where they exist, model spend proportional to leads, not to searches.
Why AI Was the Right Solution
Why AI was the right approach — and what alternatives were considered.
The inputs are unstructured and scattered: funding announcements, news about executive departures, job posts and free-text Hacker News comments. Turning a news snippet into "this company, this domain, this segment, this dated fact" is reading comprehension, which a rules engine cannot do reliably and a person cannot do at weekly volume.
Qualification is a judgement call across several weak signals: is there a CTO, how recent and relevant is the signal, is there budget, does the stack fit. A language model with a strict rubric and the search results in context does this well enough to be worth reviewing, and much faster than manual research.
At the same time, AI was deliberately kept out of every step that can be done deterministically. Scoring totals, email construction and verification, de-duplication, status transitions and sending are all code. The model reads, extracts, judges and drafts; code enforces the rules; a person decides. Hiring an SDR or buying a lead database was the alternative, but no database covers these signals and neither option fits a one-person budget.
How the Solution Was Delivered
Discovery, design, development, testing and rollout — the journey, not the tooling.
- 1Defined the four buyer segments, exclusions and non-negotiable constraints (human approval, no invented facts, no unverified emails)
- 2Designed the Google Sheet as database and control plane: Leads, Config and Runs tabs, with a status state machine
- 3Built lead generation: targeted Tavily searches, snippet cleaning and batching, Haiku extraction, de-duplication by domain and name
- 4Added the Hacker News "Who is hiring?" source with regex pre-filtering before model calls
- 5Built qualification: two searches per lead, Sonnet scoring against a rubric, sourced contacts, first-line drafting, code-side score recomputation
- 6Added email discovery: candidates generated in code, verified with Reoon, labelled by confidence, within a per-run budget
- 7Built outreach: deterministic push of approved rows to Instantly with per-segment campaigns and explicit failure states
- 8Hardened every workflow after early runs: retries, continue-on-error, run logging with notes naming what needs attention
- 9Set up sending infrastructure: dedicated domain inbox, SPF, DKIM and DMARC, warmup, conservative daily caps, tracking off
Key Technical & Architecture Decisions
Architecture, model selection, workflow and trade-offs.
Split models by job. Claude Haiku handles high-volume extraction from snippets and HN posts (about one call per twenty snippets). Claude Sonnet handles qualification, one call per lead with about 4 to 6K input tokens. This kept the careful prompt careful and the volume step cheap.
Google Sheet as database and control plane. Every cap, model name and campaign ID is a Config row read at the start of each run; every run writes a Runs row. The operator already has it open, and it works from a phone.
Human approval as the design, not a compromise. The approve column is the code review. Outreach only pushes rows that are approved and complete.
No model in the sending path. Outreach is deterministic, with explicit In sequence or Push failed states and the reason recorded.
Email addresses are never produced by the model. The model may only report an address that literally appears in a source. Otherwise code builds candidates (first@, first.last@, flast@), verifies them with Reoon, stops at the first safe result and labels catch-all or unknown results as guesses.
Rubric scored in parts, totalled in code. no_cto 0 to 3, signal 0 to 3, budget 0 to 2, stack 0 to 1, warm 0 to 1 (set only by the human). Hard disqualifiers override the score.
Blank beats generic. If there is no checkable fact from the last twelve months, the first line stays empty and the row cannot be sent.
Deliverability as part of the system: dedicated inbox on the site's domain, SPF, DKIM (2048-bit) and DMARC at p=quarantine, gradual warmup, a 17/day campaign cap, open and link tracking off, and a booking link on a short redirect so it can change without editing sequences.
Accepted duplication. The parse node exists twice, identically, with a comment saying so. A shared sub-workflow would have cost more than it saved at this scale.
Challenges & How They Were Solved
Obstacles hit along the way and how they were overcome.
Each guardrail was added after a specific failure on an early run.
- Score totals that did not match the breakdown. The code now sums the parts and ignores the model's reported total.
- Empty replies. Extended-thinking responses put a thinking block first, so reading the first content block returned nothing and every lead went to Review. The parser now joins only text blocks.
- Invalid request bodies. A lone emoji fragment in a search snippet produced invalid JSON and a 400 from the API. A helper now strips lone UTF-16 surrogates from every string that enters a prompt.
- Overwriting human work. A later run could replace a contact a reviewer had typed in. The parser now only fills empty contact fields.
- Silent failures. Every HTTP node has three retries with a 5 s wait and continue-on-error, and every run logs a row whose notes name the lead IDs that need attention. Nothing fails silently and nothing retries forever.
- Token waste on irrelevant posts. Regex keep/skip filters cut the HN thread from 255 posts to 38 before any model call.
- Extraction accuracy. It improved noticeably once each snippet carried a hint about which search produced it.
- Rows skipped without explanation. Rows missing a campaign or first line are now named in the run notes instead of being dropped quietly.
User Adoption & Change Management
How users responded — training, change management and feedback loops.
This is a one-person system, so adoption meant designing for a single operator's morning routine. The review takes about ten minutes: read the Qualified rows, check the evidence, tick approve or leave it.
Trust comes from legibility. The score breakdown, CTO evidence, email confidence and drafted first line sit side by side in the row, so a decision can be made without opening any other tool.
Control stays with the operator. Every cap lives in the sheet, so a run can be throttled or stopped from a phone. Fields a human types are never overwritten by a later run, and warm-lead scoring is set only by the human.
The review loop also improves the system: reading rows every morning is how the qualification prompt and the extraction rules got better.
Business Outcomes & Impact
The measurable outcomes recorded on this AI Case Study — with the story behind them.
Live since 28 September 2026. Automated sequences start mid October once inbox warmup completes; results will be added as replies arrive. Measured so far: - Build time: two days (28 and 29 September 2026), including false starts. - Operating time: about ten minutes a morning to review Qualified rows and tick approve, replacing manual research that was not feasible weekly. - Model cost: cents per qualified lead. Extraction is roughly 7 to 10 Haiku calls per lead-generation run; qualification is one Sonnet call per lead, capped at 15 a day. - Token spend controlled by pre-filtering: the September 2026 Hacker News thread went from 255 posts to 38 before a single model call. - Seven emails sent by hand to top-scored leads in the first two days, with follow-ups scheduled. Metrics to be reported: leads per run by source, qualification rate, share of Qualified rows approved by the human (a direct measure of model precision), verified-email rate, reply rate and booked calls per hundred sends, by segment.
Business Outcome Categories
The story behind the numbers
The practical outcome is that a prospecting job I could not do consistently by hand now runs on a schedule, and my part of it is a ten-minute review each morning.
Every lead arrives with its evidence next to it: the signal and source URL, a score breakdown, whether a CTO is present, how the email was obtained (found, verified, catch-all guess or unconfirmed) and a drafted first line. That makes the review fast, and it makes errors legible. When a row is wrong, I can see which input was wrong and fix the prompt or the rule rather than guess.
Because the send step is deterministic and gated on a human approval column, the system can run unattended without risk of an unreviewed or fabricated email going out. A bad batch can be throttled to zero from a phone before the next run by editing a cell.
Cost stays proportional to value: cheap models and regex filters handle volume, the expensive model only sees leads that survived extraction, and no model runs in the sending path at all.
Pipeline performance numbers (approval rate, verified-email rate, replies and booked calls by segment) will be added once sequences have been running for a few weeks.
Lessons Learned
What surprised the team, what worked well, and what would be done differently.
- Extraction and judgement are different jobs; splitting them across a fast and a smart model kept costs low and quality high
- The approval gate is the design, not a compromise
- Make the model's work legible in the row so the reviewer can see which input was wrong
- A spreadsheet is a fine control plane for a one-person system
- Blank beats generic: an empty first line should block sending
- Copy-paste is acceptable duplication at small scale
Future Opportunities
Where this solution could go next.
- Add a reply-classification workflow (Instantly webhook to model to sheet) so replies land in the same row as the send.
- Add a LinkedIn connections export as a warm-lead source through the same qualification path, with warm set by the human.
- Move the shared parse logic into an n8n sub-workflow once a third source appears.
- Report results by segment: leads per run by source, qualification rate, human approval rate, verified-email rate, reply rate and booked calls per hundred sends.
Professional Reflections
Guidance for another builder tackling a similar problem.
Use the same principle you would use for engineering with agents: agents draft, code checks, a person decides. Put a human gate at the one point where a mistake leaves the building, and make everything upstream legible enough that the gate takes minutes, not hours.
Keep the model away from anything code can enforce. Recompute totals, build and verify identifiers in code, and let the model do only the reading and judging.
Pre-filter with cheap rules before paying for tokens, and log every run, including partial failures. A spreadsheet is a perfectly good control plane for a one-person system.