Local engineering copilot that turns tickets into reviewable implementation briefs, linking existing behavior to code and proposed changes to requirements and tests. Includes human clarification, deterministic verification and retrieval evaluation.
Engineering tickets often omit current behavior, affected components, product decisions and acceptance tests. Engineers must rediscover that context before implementation. ReadySpec is an independent local tool that investigates a repository, exposes unresolved decisions and produces a traceable brief for human review. Demonstration and benchmark repositories are fictional; no client adoption or production business impact is claimed.
Built a TypeScript/Next.js workflow with bounded repository snapshots, lexical retrieval, SQLite persistence and Zod-validated stages. The workflow investigates code, obtains consent for selected excerpts, asks clarification questions, generates a brief, verifies citations and requirement-to-test traceability, then supports editing, named approval and export. Observed behavior, proposals, assumptions and unresolved decisions stay distinct. Provider adapters support Gemini, Anthropic and Groq; a labelled deterministic fixture mode works without an API key. Repository reads exclude secrets and do not execute or modify repository code. Call/token budgets, cancellation and resumable failures bound model usage.
Delivered the full local investigate-to-export workflow with a deterministic verifier and reproducible evaluation harness. The repository reports 200 tests plus lint, typecheck, build and CI. Gemini completed the staged workflow in a live smoke run and through the browser UI. Anthropic was tested only against a mock server; Groq integration was live-tested but did not complete a benchmark. No model-quality benchmark or measured engineer time savings exists yet.
Repository-reported deterministic evaluation covers 30 hand-authored tickets over four fictional repositories, with 13 held out in two cohorts. On the six-ticket clean held-out v2 cohort, required-file recall was 100% and retrieval precision 48%; development recall was 97% with 52% precision. These measure evidence retrieval, not generated-brief quality or business productivity. Model-dependent comparisons against a single prompt remain unmeasured. Evidence: https://github.com/ns-0437/readyspec and https://github.com/ns-0437/readyspec/blob/main/evals/REPORT.md . Live validation and limitations are documented in the repository README and docs/decisions.md.