Northline Office Book an hour 3D site view

AI skills / The receipts

The receipts

Three claims, three receipts. Each story below follows the same shape: what an AI assistant used to say, what it prints now, and the recorded number. Nothing here is a mock-up. Every figure came from a real run, and the last section on the page says exactly where each one lives.

-91.91%the bill on one job, same answer 3.68contrast caught below the 4.5 minimum 0 of 1,152,000pixels differ: a counted zero 0programs stopped without your OK

You pay for AI by the word. Left to itself, by default, it wastes them.

End of the month. The work was good. The bill is triple what you guessed, and nobody can say why.

An assistant will happily read an entire document to answer one small question, and redo yesterday’s work from scratch because nobody reminded it what it already knew. Then it calls itself efficient, with no number to back that up.

Why?

Because the bill compounds. An AI does not remember a conversation the way you do: every time you send a message, it re-reads everything that came before it. Every question, every answer, every document you ever pasted in. And you are billed for all of it, again, on every turn.

So the longer a conversation runs, the more each reply costs. Message fifty can cost many times what message one did, even if message fifty is a single sentence.

Keep feeding it whole files and the bill does not just add up. It multiplies. That is why a discipline that trims what the AI reads pays for itself on day one.

ThriftKit marblenyc/thriftkit

It used to sayIt now prints
“Let me read the log.”To answer one question, it read the whole file: tens of thousands of words. The answer was three lines long. You paid for every word. slice hands the assistant only the lines that matter. Same question, same answer, a bill 91.91% smaller on that exact job.1
“This workflow is cheaper now.”A savings claim with nothing added up. No before, no after, no receipt. meter adds up the whole job, every step and every retry, and stamps it NOT A SAVING if the total went up. Nothing counts as cheaper until the meter says so.

Receipt: -91.91% on record. Anyone can re-run the test that produced it. GitHub

“Looks fine” is a verdict nobody measured.

Friday, ship day. “Everything looks fine,” says the assistant. Nobody measured anything.

The mistakes that hurt are the ones a quick look can’t catch: text that is genuinely too faint to read, a “stopped” program still running in the background, a page that quietly shifted overnight.

Why?

Because when an assistant says “looks fine,” it has not looked. It cannot see the page the way your visitor does; it is predicting that you will accept the answer.

The only honest check is the one a careful human would do: open the real thing, measure it, and compare the number against the standard. So that is what this skill makes it do.

ProofKit marblenyc/proofkit

It used to sayIt now prints
“The text passes contrast.”The assistant checked the colours with its own arithmetic. Its arithmetic misreads how browsers paint the screen, so it passed text that real readers would squint at. audit opens a real browser and measures the page it drew: this text fails the readability standard - 3.68, minimum 4.5. The assistant’s own check had passed it.2
“The update didn’t change anything.”Two screenshots and a shrug. Nobody counted anything. diff counts every pixel: IDENTICAL: 0 of 1,152,000 pixels differ. A counted zero. And when it is not zero, it shows exactly where, and what moved.

Receipt: tested head-to-head against an assistant working without them: same accuracy, 2.7× · 1.8× · 1.49× faster.3 GitHub

The work ends. The programs keep running.

Monday your laptop was quick. By Friday it crawls, and nothing on screen explains it.

A long AI work session starts dozens of little helper programs along the way. When the work is done, half of them are still going, quietly eating your computer's memory until everything feels slow.

Why?

Because computers do not clean up after an AI session. Every helper program it starts stays running until something deliberately stops it.

On an ordinary machine that is the difference between fast on Monday and crawling by Friday, and most people never find out where the memory went.

The fix has to ask first: a tool that kills programs on its own is scarier than the problem. This one lists, asks, and only then acts.

Wind Down marblenyc/wind-down

It used to sayIt now does
“I stopped the server.”It shut the front door and left the engine running. The part that was using your memory never stopped. wind-down lists everything still running, in plain sight, and stops only what you approve, one item at a time. Your work, your history and your place in it all stay open.4

Receipt: looks before it touches. Never stops anything without your OK. Your browser and editor are off-limits. GitHub

Where every number on this page comes from

  1. The -91.91% is the recorded receipt for slicing one 85 KB log, in token-thrift/docs/receipts.md; the bench that produced it ships in the repo. Nothing counts as a saving until job_meter.py compare says the whole-job total fell - moving work to a cheaper model included.
  2. 3.68:1 is a real failure caught on a real wordmark. Chrome serves color-mix() backgrounds as color(srgb r g b); an audit that only parses rgb() resolves them against the page ground and reports the text as passing at 4.5:1.
  3. Each instrument was run un-hinted against a top-tier no-skill baseline in a context-clean harness - isolated config, neutral paths, external grading. Correctness tied in all three cases; the multiples are the measured wall-clock deltas. Full runs live under each tree's evals/. On a strong model these buy amortization and pinned determinism; on a weaker one they buy correctness outright.
  4. Wind Down stops operating-system processes, which can lose unsaved work in the process it stops. Nothing is stopped without explicit per-item approval, and each process is re-verified immediately before it is stopped. Safeguards, not guarantees - provided as-is; see the repository's disclaimer.

Convinced, or at least curious?