AI Deploy Network
AI AgentSoftware & SaaS 9/27/2026

Polish: self-hosted UI review for code, copy, and rendered screens

Polish is a self-hosted command-line tool that reviews UI code and rendered screens against a 136-rule rubric for usability, design craft, and accessibility. It runs on your own API key with no quotas.

Hours Automated
0
Cost Savings
Currency not specified
Revenue Impact
Currency not specified

Business Challenge

AI-generated UI code ships with problems a senior designer catches on sight but automated linters miss: inline hex colors, clickable divs instead of buttons, delete actions with no confirmation, images with no alt text. Reviewing this by hand is slow and inconsistent, and it does not scale across a team or an agent workflow. The existing options are backend-focused linters that ignore design craft, or hosted services that need bot access to your repository and cap how many reviews you can run. I wanted a way to check UI quality that stays local, uses my own model key, and returns the same verdict whether a person or a coding agent runs it.

Solution Delivered

Polish is a Node script you clone and run locally. It reads UI source files and optional screenshots, then sends them to a language model with a rubric encoded as plain data. The rubric has three layers: ten usability heuristics, eight design craft categories, and one accessibility pass, covering 136 checkable rules. Because the rubric is data, a team can swap in its own design standards. Each run returns a score out of 100, plus findings tagged by severity and category, with a file and line reference and a suggested fix. Scoring is deterministic, so the number is easy to reason about. A verify command re-checks only prior findings against updated code at a fraction of the cost of a full review, which closes the fix loop. It works with Groq, OpenAI, Anthropic, Gemini, OpenRouter, or any OpenAI-compatible endpoint, using one provider per run or a fallback chain. It runs as a CLI you wire into pre-commit or CI, exiting non-zero when critical findings exist, and as an MCP server over stdio so coding agents call the same engine. An init-agent command writes an AGENTS.md file that teaches agents when to review and how to verify, and every run emits a machine-readable receipt for CI and agent loops.

Outcomes Achieved

Polish turns subjective UI review into a repeatable, scriptable step. The same rubric produces the same verdict whether a person or an agent runs it, so review quality no longer depends on who is looking. Running it stays local and cheap: a dry run of a full HTML page estimates about 29,000 tokens, under a cent on most models, with no monthly cap. The verify command re-checks only prior findings, so confirming a fix costs far less than a fresh review. The project is open source under MIT and doubles as a working reference for encoding design judgment as data a model can check.

Measurable Business Outcome

Reduced UI review to a single command that returns a score, severity-ranked findings, and fixes, replacing an inconsistent manual pass. A full-page dry run estimates about 29,000 tokens, under one cent per review on most models, with no monthly quota. The verify command re-checks only previous findings, cutting the cost of confirming a fix to a fraction of a full review. Codified 136 checkable rules across usability, design craft, and accessibility into swappable data.

Business Outcome Categories

Quality ImprovementProductivity ImprovementAI Performance ImprovementTime Savings