{"_id":"@akins20/ui-critic","_rev":"4-d12d50ee332d1debc8c9933e5ab5fdcb","name":"@akins20/ui-critic","dist-tags":{"latest":"0.3.2"},"versions":{"0.2.0":{"name":"@akins20/ui-critic","version":"0.2.0","keywords":["gemini","claude","claude-code","ui","design-review","screenshots","playwright","critique","agent"],"author":{"name":"Otesc","email":"otesc.co@gmail.com"},"license":"MIT","_id":"@akins20/ui-critic@0.2.0","maintainers":[{"name":"akins20","email":"ogunbiye@gmail.com"}],"homepage":"https://github.com/Akins20/ui-critic#readme","bugs":{"url":"https://github.com/Akins20/ui-critic/issues"},"bin":{"ui-critic":"bin/ui-critic.mjs"},"dist":{"shasum":"64c9d3187c20b9e5e85282007e102e02ce1edbab","tarball":"https://registry.npmjs.org/@akins20/ui-critic/-/ui-critic-0.2.0.tgz","fileCount":18,"integrity":"sha512-eG1JMGHaOcrH8wj68YW0rD6WEpUYG/AtG+EHmvzcQV3AF8fW4R5+s+gVcqyr8AsCehd3UiMvUYqCadYOLmiT1A==","signatures":[{"sig":"MEUCIQCQ1H6VqxM9gUYAIbXMxAEPO8MfBIamFHV79g+xacfK1wIgUiOUOZqYlHQrDdeRncLk0XkpgypUqEznc9yiaMZS3nU=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":130324},"type":"module","engines":{"node":">=20"},"gitHead":"adcda3af9069afd4695bc0aab376a65a338524b5","scripts":{"test":"node --test","prepublishOnly":"npm test"},"_npmUser":{"name":"akins20","email":"ogunbiye@gmail.com"},"repository":{"url":"git+https://github.com/Akins20/ui-critic.git","type":"git"},"_npmVersion":"10.9.2","description":"A second pair of eyes for a UI: capture screenshots, get a ranked visual critique from Gemini, and compare before and after. Built so one coding agent (Claude Code) can ask another model for design review.","directories":{},"_nodeVersion":"22.17.0","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"peerDependencies":{"playwright":">=1.40"},"peerDependenciesMeta":{"playwright":{"optional":true}},"_npmOperationalInternal":{"tmp":"tmp/ui-critic_0.2.0_1789132642009_0.9933235799467486","host":"s3://npm-registry-packages-npm-production"}},"0.3.0":{"name":"@akins20/ui-critic","version":"0.3.0","keywords":["gemini","claude","claude-code","ui","design-review","screenshots","playwright","critique","agent","openai","chatgpt","gpt","accessibility","visual-regression"],"author":{"name":"Otesc","email":"otesc.co@gmail.com"},"license":"MIT","_id":"@akins20/ui-critic@0.3.0","maintainers":[{"name":"akins20","email":"ogunbiye@gmail.com"}],"homepage":"https://github.com/Akins20/ui-critic#readme","bugs":{"url":"https://github.com/Akins20/ui-critic/issues"},"bin":{"ui-critic":"bin/ui-critic.mjs"},"dist":{"shasum":"3e97987a667db4ce7a3e447f25bb597c94387eeb","tarball":"https://registry.npmjs.org/@akins20/ui-critic/-/ui-critic-0.3.0.tgz","fileCount":24,"integrity":"sha512-JGrNUQI5zcudfGWcEha3Wt2o6vXUVTZcSgKU7j3aM9Y9I9SbPlDiHU5paLh4mqVskoGOB1dmT3gs+/W3RKVlSw==","signatures":[{"sig":"MEUCIQDoKg8Y5mLM4Izy9kXYFvkXxC9Qz6utBwplEny2+RU6GAIgHlSbGWFEhD1uvH7yrfn9/ibRBsLqs7Bt1RMUlse8gik=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"},{"sig":"MEUCICi2EOjFHGSCFvLy4kPoEEb4sSdYjWVF8YZCVcyu1rH9AiEA/bz852dLkPPh/bVWCsU++YaZnmNBAjaUZWelIZvcKLM=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":168055},"type":"module","engines":{"node":">=20"},"gitHead":"42ab8ffce6a641124ef60f4024c3ed05b25967ab","scripts":{"test":"node --test","prepublishOnly":"npm test"},"_npmUser":{"name":"akins20","email":"ogunbiye@gmail.com"},"repository":{"url":"git+https://github.com/Akins20/ui-critic.git","type":"git"},"_npmVersion":"10.9.2","description":"A second pair of eyes for a UI: capture screenshots and measured facts, get a ranked visual critique from Gemini or OpenAI, and verify before and after. Built so one coding agent (Claude Code) can ask another model for design review.","directories":{},"_nodeVersion":"22.17.0","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"peerDependencies":{"playwright":">=1.40"},"peerDependenciesMeta":{"playwright":{"optional":true}},"_npmOperationalInternal":{"tmp":"tmp/ui-critic_0.3.0_1789137112370_0.8747600491320042","host":"s3://npm-registry-packages-npm-production"}},"0.3.1":{"name":"@akins20/ui-critic","version":"0.3.1","keywords":["gemini","claude","claude-code","ui","design-review","screenshots","playwright","critique","agent","openai","chatgpt","gpt","accessibility","visual-regression"],"author":{"name":"Otesc","email":"otesc.co@gmail.com"},"license":"MIT","_id":"@akins20/ui-critic@0.3.1","maintainers":[{"name":"akins20","email":"ogunbiye@gmail.com"}],"homepage":"https://github.com/Akins20/ui-critic#readme","bugs":{"url":"https://github.com/Akins20/ui-critic/issues"},"bin":{"ui-critic":"bin/ui-critic.mjs"},"dist":{"shasum":"dd44ff8e03da5d561306c39bc1b8456d1a8167eb","tarball":"https://registry.npmjs.org/@akins20/ui-critic/-/ui-critic-0.3.1.tgz","fileCount":24,"integrity":"sha512-uzk77Y4edM1Q5Czf0V2HxAO0FX9UFYOLZ/Vy52EDXmCt2jzR6TUXICa9a0GQI/BKjPK9rXtHT08d5uBRXjbrmw==","signatures":[{"sig":"MEUCIHeTFrajGAjSghzMKuXY8fyeWDHimbC1MqrQur05Ikj/AiEAk7kx/Dualj2TYqaZRwOekuneETWc4yk6P6LDf2uMQKk=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"},{"sig":"MEUCIFAV2/+Q/FbdoU6/L1onhpAif7XeE+1g/RjnjD1XmcjYAiEAvb+FDLAusSPgeDhFumfoGFp9AROZf+VD5kWdEZ2hO9M=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"}],"unpackedSize":168524},"type":"module","engines":{"node":">=20"},"gitHead":"eafc02b77c00a68a5809735438157b09d04aac76","scripts":{"test":"node --test","prepublishOnly":"npm test"},"_npmUser":{"name":"akins20","email":"ogunbiye@gmail.com"},"repository":{"url":"git+https://github.com/Akins20/ui-critic.git","type":"git"},"_npmVersion":"10.9.2","description":"A second pair of eyes for a UI: capture screenshots and measured facts, get a ranked visual critique from Gemini or OpenAI, and verify before and after. Built so one coding agent (Claude Code) can ask another model for design review.","directories":{},"_nodeVersion":"22.17.0","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"peerDependencies":{"playwright":">=1.40"},"peerDependenciesMeta":{"playwright":{"optional":true}},"_npmOperationalInternal":{"tmp":"tmp/ui-critic_0.3.1_1789139151644_0.6854328474614297","host":"s3://npm-registry-packages-npm-production"}},"0.3.2":{"_id":"@akins20/ui-critic@0.3.2","bin":{"ui-critic":"bin/ui-critic.mjs"},"bugs":{"url":"https://github.com/Akins20/ui-critic/issues"},"dist":{"shasum":"8ae10d8de7df66b1acd403ae1a22d8b2faa3fc70","tarball":"https://registry.npmjs.org/@akins20/ui-critic/-/ui-critic-0.3.2.tgz","fileCount":24,"integrity":"sha512-4SGEerpVNAqfOPewcpkCsRkhTv4EJC1yexDEvLw2wmPKRVkATg1/T4ChW6WYeKL7PJwG/4EwRW2d/X/OH/gvdA==","signatures":[{"sig":"MEUCIQDcEeUYfksJQt5mEj6rhpp4RerpFuEY5Xv08LPEEAiJXgIgQKFYyvccyKnYbXf2Yw6iJYhAIKqzlKSkB33wXLBrm84=","keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U"},{"keyid":"SHA256:DhQ8wR5APBvFHLF/+Tc+AYvPOdTpcIDqOhxsBHRwC7U","sig":"MEUCIBn8dbxLk9A3FcaH4ZeeQwsOsSrJNv8cCk/bWFI8m2tGAiEAuLiv0qkJPQJ/EQZIXHGOu1urFpi6rHnkLa/uXwcv7v0="}],"attestations":{"url":"https://registry.npmjs.org/-/npm/v1/attestations/@akins20%2fui-critic@0.3.2","provenance":{"predicateType":"https://slsa.dev/provenance/v1"}},"unpackedSize":165784},"name":"@akins20/ui-critic","type":"module","author":{"name":"Otesc","email":"otesc.co@gmail.com"},"engines":{"node":">=20"},"gitHead":"5caa79ffe3187269fff1c83e1795ec3d644e4f3e","license":"MIT","scripts":{"test":"node --test","prepublishOnly":"npm test"},"version":"0.3.2","_npmUser":{"name":"GitHub Actions","email":"npm-oidc-no-reply@github.com","trustedPublisher":{"id":"github","oidcConfigId":"oidc:2ab06ca1-c41f-464b-b338-cd480c5451d0"}},"homepage":"https://github.com/Akins20/ui-critic#readme","keywords":["gemini","claude","claude-code","ui","design-review","screenshots","playwright","critique","agent","openai","chatgpt","gpt","accessibility","visual-regression"],"repository":{"url":"git+https://github.com/Akins20/ui-critic.git","type":"git"},"_npmVersion":"12.0.2","description":"A second pair of eyes for a UI: capture screenshots and measured facts, get a ranked visual critique from Gemini or OpenAI, and verify before and after. Built so one coding agent (Claude Code) can ask another model for design review.","directories":{},"maintainers":[{"name":"akins20","email":"ogunbiye@gmail.com"}],"_nodeVersion":"22.23.2","publishConfig":{"access":"public"},"_hasShrinkwrap":false,"peerDependencies":{"playwright":">=1.40"},"peerDependenciesMeta":{"playwright":{"optional":true}},"_npmOperationalInternal":{"host":"s3://npm-registry-packages-npm-production","tmp":"tmp/ui-critic_0.3.2_1789142470084_0.87633308174882"}}},"time":{"created":"2026-09-11T13:17:21.403Z","modified":"2026-09-11T16:01:10.575Z","0.2.0":"2026-09-11T13:17:22.162Z","0.3.0":"2026-09-11T14:31:52.452Z","0.3.1":"2026-09-11T15:05:51.724Z","0.3.2":"2026-09-11T16:01:10.188Z"},"bugs":{"url":"https://github.com/Akins20/ui-critic/issues"},"author":{"name":"Otesc","email":"otesc.co@gmail.com"},"license":"MIT","homepage":"https://github.com/Akins20/ui-critic#readme","keywords":["gemini","claude","claude-code","ui","design-review","screenshots","playwright","critique","agent","openai","chatgpt","gpt","accessibility","visual-regression"],"repository":{"url":"git+https://github.com/Akins20/ui-critic.git","type":"git"},"description":"A second pair of eyes for a UI: capture screenshots and measured facts, get a ranked visual critique from Gemini or OpenAI, and verify before and after. Built so one coding agent (Claude Code) can ask another model for design review.","maintainers":[{"name":"akins20","email":"ogunbiye@gmail.com"}],"readme":"# ui-critic\n\nA second pair of eyes for a UI. Capture screenshots and measured facts from a running\nsite, get a ranked visual critique from Gemini or OpenAI that judges against your product's\npurpose and audience across every design discipline, triage it, ship the fixes, and verify\nwith a before/after comparison.\n\nIt exists so one coding agent can ask another model for design review: Claude Code does the\nbuilding and the judgement, Gemini does the looking. It works just as well for a human at a\nterminal. Zero required dependencies beyond Node 20; Playwright is optional, for capture.\n\n## Quick start\n\n```bash\nnpm i -g @akins20/ui-critic                   # or run it with npx @akins20/ui-critic ...\nexport GEMINI_API_KEY=...                     # or OPENAI_API_KEY; read from the environment only\nui-critic init --base https://your.site       # writes ui-critic.config.json, ui-critic/brief.md, ui-critic/decisions.md\n# fill in the brief: what the product is for and who it is for are required\nui-critic run --label before --follow-requests\n# make changes, then either deploy or run locally\nui-critic verify --before ui-critic-out/before --base http://localhost:3000\n```\n\nThe package is `@akins20/ui-critic` on npm; the command it installs is `ui-critic`.\n\n## Which model looks\n\nGemini is the default critic (`gemini-3.8-flash`). Any OpenAI model with vision works too:\npass `--model gpt-5.4-mini` (the provider is inferred from the id) or set `\"provider\":\n\"openai\"` in the config, and export `OPENAI_API_KEY`. Reasoning models take the thinking\nlevel as their reasoning effort; `includeThoughts` keeps the reasoning summary. OpenAI has\nno explicit context cache, so the tool relies on its automatic prefix caching; the prompt\nis already ordered stable-prefix first for that. `ui-critic models` lists the provider's\nvision models with their prices, and every report and the cost ledger name the provider.\n\n`capture`, `run` and `verify` need Playwright with Chromium in the project\n(`npm i -D playwright && npx playwright install chromium`; `@playwright/test` and\n`playwright-core` are accepted too). `critique` and `compare` need only Node and the key.\n\n## What the critic gets\n\nThe critic never judges a generic store. Every request carries, in this order:\n\n1. **The review rules** and the **design disciplines** it must sweep on every page: layout\n   and grid, spacing and rhythm, typography, colour and contrast, surfaces and dividers,\n   imagery and icons, component states, navigation, microcopy, conversion, trust, motion,\n   accessibility, responsiveness, and consistency across pages. Each page review accounts\n   for every discipline (fine, issue, or not applicable), so nothing is skipped because a\n   louder problem caught its eye. The list is configurable (`disciplines`). Alongside them,\n   a short list of **interaction principles** every screen must satisfy (`principles`,\n   also configurable), starting with the one screenshots miss most: every action a user\n   takes gets immediate, visible feedback (pressed, loading, success, error). A violation\n   is a finding; where a screenshot cannot show a state, the critic must ask rather than\n   assume.\n2. **The brief** (required): what the product is for, who it is for, the brand system,\n   what to ignore, what matters most, benchmarks. `init` writes the template; a critique\n   refuses to run while the brief is empty, still the template, or missing Product or\n   Audience.\n3. **Extra context** you choose: `context.files` (design tokens, copy decks, policies) and\n   the answers file, where your team answers what the critic asked for last time.\n   **Settled decisions** (`ui-critic/decisions.md`, one bullet each with the reason) are\n   closed: the critic is told not to reopen them, every finding names the decision it would\n   reopen or none, and those findings are withheld from the report but counted, so the\n   filter stays auditable.\n4. **Every page's first screen** at every viewport, then per page the **full-page\n   capture** and the **measured facts** the capture gathered in the browser: fonts and the\n   base size, the text size histogram, the heading outline, landmarks, image alt coverage,\n   interactive targets under 24px, the lowest-contrast visible text with its WCAG AA\n   result, and the **runtime facts** a screenshot cannot show: console errors, uncaught\n   exceptions, failed requests (not the ones the browser cancelled itself, such as\n   abandoned prefetches), HTTP errors and cumulative layout shift. Measured facts are\n   ground truth for the critic, so it does not guess a contrast ratio or a font size.\n\n## States, mini-features and signed-in pages\n\nScreenshots of resting pages miss what happens when someone acts. **Scenarios** open a\nroute, run a few steps and shoot the result, which is reviewed as its own page named\n`route [scenario]`: a filter sheet opened, a plan length pressed, a wishlist toggled, a\ncard hovered, a link focused from the keyboard, an invalid form submitted, an empty search.\nThe step vocabulary is `goto`, `click`, `hover`, `focus`, `fill`, `press`, `wait`,\n`waitFor`, `waitForURL` and `scroll`, with Playwright selectors. A scenario that only makes\nsense at some sizes (a mobile filter sheet, a desktop hover) lists them in `viewports`.\nEvery scenario runs in a fresh browser context, so a saved wishlist or a switched theme\nnever leaks into the next capture; signed-in scenarios reuse one login session.\n\n```json\n{\n  \"routes\": [\"/\", \"/shop\", \"/shop?q=wig\", { \"path\": \"/account/plans\", \"auth\": true }],\n  \"scenarios\": [\n    { \"name\": \"filters-open\", \"route\": \"/shop\", \"steps\": [{ \"click\": \"text=Filters\" }, { \"wait\": 400 }] },\n    { \"name\": \"card-focus\", \"route\": \"/\", \"steps\": [{ \"focus\": \".pcard .title\" }] },\n    { \"name\": \"invalid-submit\", \"route\": \"/login\", \"steps\": [{ \"fill\": { \"selector\": \"input[name=email]\", \"value\": \"not-an-email\" } }, { \"click\": \"button[type=submit]\" }, { \"wait\": 500 }] }\n  ],\n  \"auth\": {\n    \"mode\": \"form\",\n    \"login\": \"/login\",\n    \"envFile\": \"ui-critic/auth.env\",\n    \"steps\": [\n      { \"fill\": { \"selector\": \"input[name=email]\", \"envVar\": \"UI_CRITIC_AUTH_USER\" } },\n      { \"fill\": { \"selector\": \"input[name=password]\", \"envVar\": \"UI_CRITIC_AUTH_PASS\" } },\n      { \"click\": \"button[type=submit]\" }\n    ],\n    \"success\": \"**/account**\"\n  }\n}\n```\n\nRoutes and scenarios marked `auth: true` are captured in a signed-in context. The tool\nsigns in with credentials it reads **by variable name** from the environment or from\n`auth.envFile` (keep that file out of git); a literal password in the config is rejected.\n`storageState` mode loads a Playwright storage state you exported after signing in\nyourself. When the credentials are missing the signed-in pages are skipped and the report\nsays so under \"Not captured\", instead of quietly reviewing a login redirect.\n\n## What the critic can ask for\n\nEach page review and the site review end with `requests`: pages, files, answers or\nmeasurements the critic needs to judge better, each with the reason. Page requests for the\nsame origin are fulfilled automatically with `--follow-requests` (or\n`followRequests.enabled` in the config): the tool captures the page, reviews it and adds it\nto the report, up to `maxPages`; a requested page that redirects a visitor to sign in is\ncaptured again in the signed-in context when `auth` is configured. Everything else is\nlisted under \"Critic's requests\" in the report; answer it in `ui-critic/answers.md` (or add the page to `routes`) and rerun, and the\nanswers travel with the next critique.\n\n## What you get\n\n- `critique.md` / `critique.json`: a score per page and for the site, strengths, ranked\n  findings with evidence and a specific recommendation each, the measured facts and\n  discipline coverage per page, a revamp-or-polish verdict, the five highest-leverage\n  changes, and the critic's requests.\n- `compare.md` / `compare.json`: per page and viewport, what improved, what regressed, what\n  is still open, with measured facts from both sides. Every reported regression gets a\n  second, stricter look at the same captures; only confirmed ones count, each tagged\n  `measured` (a fact proves it) or `judged` (visual judgement), and the unconfirmed ones are\n  listed with the reason. A verdict that rested only on unconfirmed regressions is revised.\n\nFindings are typed (`hierarchy`, `typography`, `conversion`, `accessibility`, ...) and each\nis marked `ui`, `placeholder-content` or `needs-engineering-judgement`, so the reader can\ntriage instead of obeying.\n\n## Commands\n\n| command | does |\n| --- | --- |\n| `init [--base url]` | write a starter config with every default spelled out, and a brief from the template |\n| `models [--filter flash]` | list vision-capable models for the key |\n| `capture --base url --label name` | above-the-fold PNG, full-page JPEG and measured audit per route, scenario and viewport (signed in where configured), plus `manifest.json` |\n| `critique --in dir` | per-page scores, strengths, ranked findings, coverage, requests; a site-level verdict and top five priorities |\n| `compare --before dir --after dir` | per page and viewport: improved, regressed, still open |\n| `run --base url --label name` | capture then critique |\n| `verify --before dir --base url` | capture \"after\" then compare, in one step |\n| `cost [--out dir]` | total the usage ledger per run at today's prices |\n\n`--json` prints a machine-readable summary to stdout (for an agent to parse); the full\nreports are always written next to the screenshots. `--fail-on measured` makes `compare`\nand `verify` exit with code 2 when a confirmed regression is backed by a measured fact\n(contrast, target size, landmarks, headings, layout shift, errors), which is the choice\nfor a pull-request gate because it cannot flake on taste; `--fail-on regressed` trips on\nany confirmed regression and `--fail-on worse` on a worse verdict. `--no-confirm` skips\nthe second look.\n\nCalls run a few at a time (`concurrency`, default 3, `--concurrency N`,\n`UI_CRITIC_CONCURRENCY`); each finished page or pair is checkpointed as it lands and the\nreport keeps the capture order.\n\n## Configuration\n\nEvery knob has a default. Resolution order, lowest to highest: built-in defaults,\n`ui-critic.config.json` (working directory or `--config`), environment, flags.\n\n```json\n{\n  \"base\": \"https://your.site\",\n  \"routes\": [\"/\", \"/shop\", \"/products/example\", \"/how-it-works\", \"/login\"],\n  \"viewports\": {\n    \"desktop\": { \"width\": 1366, \"height\": 900, \"deviceScaleFactor\": 1 },\n    \"mobile\": { \"width\": 390, \"height\": 844, \"deviceScaleFactor\": 2, \"isMobile\": true }\n  },\n  \"brief\": \"ui-critic/brief.md\",\n  \"context\": { \"files\": [\"app/globals.css\"], \"answers\": \"ui-critic/answers.md\", \"decisions\": \"ui-critic/decisions.md\" },\n  \"followRequests\": { \"enabled\": true, \"maxPages\": 3 },\n  \"concurrency\": 3,\n  \"compare\": { \"confirmRegressions\": true },\n  \"provider\": \"gemini\",\n  \"disciplines\": [\"layout and grid: ...\", \"typography: ...\"],\n  \"principles\": [\"Every action a user takes gets immediate, visible feedback: ...\"],\n  \"out\": \"ui-critic-out\",\n  \"model\": \"gemini-3.8-flash\",\n  \"hideSelectors\": [\"nextjs-portal\"],\n  \"thinking\": { \"level\": \"high\", \"includeThoughts\": false },\n  \"cache\": { \"enabled\": true, \"ttlSeconds\": 3600, \"minTokens\": 2048, \"keep\": false },\n  \"generation\": { \"temperature\": 0.3, \"maxOutputTokens\": 32768 },\n  \"pricing\": {},\n  \"ledger\": \"usage.jsonl\"\n}\n```\n\nEnvironment: `GEMINI_API_KEY` or `OPENAI_API_KEY` (one is required), `GEMINI_MODEL`,\n`UI_CRITIC_PROVIDER`, `UI_CRITIC_THINKING`, `UI_CRITIC_CACHE=0`, `UI_CRITIC_OUT`,\n`UI_CRITIC_BRIEF`, `UI_CRITIC_CONCURRENCY`. Flags: `--provider`, `--model`,\n`--thinking-level`, `--include-thoughts`, `--no-cache`, `--ttl`, `--temperature`,\n`--routes`, `--out`, `--brief`, `--context`, `--answers`, `--decisions`,\n`--follow-requests`, `--max-pages`, `--concurrency`, `--no-confirm`, `--config`.\n\n### Thinking\n\n`thinking.level` is `off`, `low`, `medium` or `high` (Gemini 3.x `thinkingLevel`, or the\nreasoning effort of an OpenAI reasoning model; the default is `high`, since a critique is\njudgement work). `thinking.budget` sets a token\nbudget for models that use `thinkingBudget` instead. `includeThoughts: true` keeps the\nmodel's reasoning in `thoughts.md` beside the report, so a reviewer can see why a finding\nwas made. An agent can set all of these per run with flags or env without touching the file.\n\n### Caching\n\nThe prompt is built stable-prefix-first: the review rules, the brief, the extra context and\nevery page's above-the-fold capture come first and are byte-identical across calls, so the\nAPI's implicit prefix caching applies on its own. With `cache.enabled` (the default) that\nprefix is also stored as an explicit context cache for the run (created when it is at least\n`minTokens`, reused by every page and site call, deleted at the end unless `keep`), so the\nshared screenshots and brief are paid for once at the cached rate instead of once per\ncall. Any cache failure falls back to inline, with the reason recorded in the report.\n\n### Resilience\n\nThinking tokens count against the output budget on Gemini 3.x, so the default\n`generation.maxOutputTokens` is 32768 and a response cut off at the budget is retried\nwith double the budget up to 65536. Every finished page is checkpointed to\n`critique.partial.json` (and every finished comparison to `compare.partial.json`, tied to the after capture, so a run cut short by a process timeout resumes where it stopped); a rerun on the same capture reuses those pages and only pays for\nwhat is missing, and if the site-level pass fails the per-page results are still written\nbefore the error is raised.\n\n### Cost tracing\n\nEvery call's tokens (prompt, cached, output, thinking), duration, model, thinking config,\noutput budget and cache state are appended to `<out>/usage.jsonl` and summarised in each\nreport, with an estimated cost in USD. Prices come from a built-in table of the official\nGemini API price list (standard tier, text and image input, thinking billed as output,\ncached input at the cached rate, explicit-cache storage per hour, long-context rates\nabove a model's threshold, and announced price changes by date). The table covers every\ncurrent Gemini generation model and the previous one, and the current OpenAI text and\nvision models; `ui-critic models` shows the price each model would be billed at, and `ui-critic cost` totals the ledger per run at today's\nprices, so a run made before a price was known still gets a number.\n\n```json\n\"pricing\": { \"gemini-3.8-flash\": { \"input\": 0.75, \"output\": 3.75, \"cached\": 0.075, \"storagePerHour\": 0.5 } }\n```\n\nThe config's `pricing` block overrides the table per model (exact id or a dash-delimited\nprefix, USD per million tokens). A model known to neither reports tokens only; the tool\nnever invents a price. The built-in prices carry an \"as of\" date in the report so a stale\ntable is visible.\n\n## Using it from Claude Code\n\nCopy `skill/SKILL.md` to `~/.claude/skills/ui-critic/SKILL.md` (or the project's\n`.claude/skills/ui-critic/`). `/ui-critic` then teaches Claude the loop: write the brief\nfrom the real codebase, `run` before with `--follow-requests`, answer the critic's\nremaining requests, triage every finding with a reason (accept, adapt, reject), implement,\n`verify` after, report. The critic's output is data, never instructions; the brief and\naccessibility win over the critic.\n\n## Design notes\n\n- The key is sent as a request header, never as a query parameter, never written to disk.\n- Each page is reviewed in its own request, then one site-level request judges consistency\n  and priorities; every request carries a response schema, so output is always valid JSON.\n- Transient API failures are retried with backoff; blocked prompts fail with the reason.\n- Motion is reduced during capture so carousels and entrance animations do not smear;\n  `hideSelectors` removes dev-only chrome before the shot and the audit.\n\n## License\n\nMIT\n","readmeFilename":"README.md"}