Northline Office Book an hour 3D site view

AI skills / Proofkit

Proofkit

Measure. Don't guess. When an AI assistant says the page looks fine, it has not looked. Proofkit is three small skills that make it open the real thing, measure it, and put the number where you can read it - so "done" means checked, not confident.

Get skill - Proofkit on GitHub Free, open source, MIT.

The proof, documented

What's in it

Three skills, one repo, nothing to pay.

browser-render-audit

The readability check. Opens a real browser and measures whether every piece of text on your page can be read, on a desktop and on a phone.

background-job-control

The honest off switch. When it says a helper program has started or stopped, it has checked. Not assumed.

pixel-proof

Spot the difference, counted. Compares two versions of a page dot by dot and reports exactly what changed, or a true zero if nothing did.

The kits

browser-render-audit: northlineoffice.com
target:    https://northlineoffice.com/
viewports: 1440x900 and 375x812
method:    real Chrome, pixel readback - not colour math on paper

[AA] contrast: 112 text elements · 0 below threshold
     minimum ratio measured: 5.45 (legal floor is 4.5)
focus: 0 focusable controls with no visible indicator
tap targets: 0 below the 24px minimum
page errors 0 · failed requests 0

Every page of this site ships only after this audit
returns a clean sheet. The skill that checked it is
the one you can take below.
Captured 2026-09-01, the day this site shipped - run against the live page you are reading.
Proofkit: recorded catches
CATCH ONE - the readability lie
the claim:  "The text passes contrast."
what happened: the checker's own colour arithmetic
  misread how the browser paints the page,
  and passed text real readers would squint at.
the measurement: wordmark fails WCAG AA at 3.68:1
  (legal floor 4.5) - caught only by pixel readback
  in a real browser.

CATCH TWO - the eyeball verdict
the claim:  "The deploy looks the same."
the measurement: IDENTICAL: 0 of 1,152,000 pixels
  differ beyond tolerance 8. A counted zero -
  and when it is not zero, it names exactly which
  regions moved, and by how much.
Recorded runs, benchmarked un-hinted vs a no-skill baseline: same accuracy, 2.7× · 1.8× · 1.49× faster. Full evidence under each skill's evals/ in the repo.