Prompt kit

Nobody Told the Agent What Done Means

Six paste-ready prompts. No account required. Works in ChatGPT, Claude, Cursor, or a doc.

The argument is in the essay. This page is what you leave with. Fill the brackets before you paste, or let the model interview you. Do not let it invent the passing condition.

Session definition-of-done card

Paste at the start of a session, before any work. Fill the brackets first, or let the model interview you. The model must not invent a passing condition.

Anthropic's June 2026 Claude Code study: people keep about 70% of planning decisions, including what counts as done, and only about 20% of execution. This card is that planning decision. Do not let the model take it.

Paste the block as written.

You are writing a session definition of done. Do not start the original task.

Task in one sentence: [what I asked for]

Produce a one-screen card with exactly these headings:

1. Outcome (what is different in the world if this is finished — observable by someone who was not in this session)
2. Ordinary measure (one number this team already uses; not an agent dashboard, not "looks good")
3. Evidence (what a second person could inspect without me: command + expected exit, URL, file path, named review)
4. Signer (a person, not "the team")
5. Must not optimize (the activity proxy that would look like progress and must not count)
6. Last failure to avoid (one prior miss on work like this, or "unknown")
7. Out of scope (domain boundary; the dangerous remainder a general agent must not improvise)

Rules:
- If a heading is unknown, write "unknown" and set Status: Draft.
- If Status is Draft, do not start the work. Ask me to confirm or fill the unknown line.
- Do not edit this card once I confirm it. Restate it in one block and wait.
- Never treat a plan, a status update, token spend, or your own summary as done.
- What counts as done is a planning decision I keep. Do not invent the passing condition, the evidence, or the signer.

Tone: an operating note, not a policy PDF. Under 180 words.

Four founder questions

Before you expand an agent's authority. Nate's questions, translated so a founder can answer them in fifteen minutes. Do not turn this into a strategy memo.

Paste the block as written.

I am about to give an agent more work, more authority, or more volume.

Agent / workflow: [name it]
What it is allowed to touch: [code, inbox, ads, tickets, customer data — be specific]
Last week of output: [paste or describe]

Answer these four questions in order. One short paragraph each. If you cannot answer, say so and stop; do not invent the missing fact.

1. Inspectability. Can an ordinary competent person — not the person who ran the session — open this work and explain why it works and how to build on it? What would they look at?
2. Ordinary measures. Which numbers does this business already use? If the agent's dashboard improved last week and those numbers were flat, which do we believe, and why?
3. Domain boundary and last failure. Where does our expertise end? What was this agent's last important failure, and what did we change after it?
4. Liability. If this work sits outside our expertise and carries liability (tax, employment, contracts, regulated money), what is stopping us from buying a domain-specific agent or a managed service instead of configuring a general one?

Do not recommend a new platform. Do not compliment the agent. End with one sentence: expand, hold, or buy.

Unplug test

For founders and anyone about to let a general agent into the dangerous remainder.

Paste the block as written.

I am considering letting a general-purpose agent do this work: [describe it].

Run the unplug test.

1. Could I (or a named person here) do this unaided well enough to sign it? Yes / no / not me — then who.
2. If the agent vanished tomorrow, what breaks in 24 hours? In 30 days?
3. Which part of this work is the dangerous remainder: tax, employment law, contracts, regulated money, or "none"?
4. If the remainder is not "none," the recommendation is buy a domain-specific tool or keep the liability. Do not propose a clever prompt that makes a general agent "careful."

Return a four-line card: unaided, break-if-unplugged, remainder, decision (keep / buy / do not automate).
If I have left a line thin, ask me one question, then finish the card.

Inspectability test

The twenty-minute test. Works for code, a campaign, a support policy, or a spreadsheet.

Paste the block as written.

You are running an inspectability test on agent-produced work.

Artifact: [path, URL, or paste]
Who would inspect it in real life: [second- or third-best person on this work, not the session owner]
Time box: 20 minutes of their attention, not yours.

Do not praise the artifact.

Return:
- What they could explain in 20 minutes (be specific: which file, which step, which number)
- What they would have to take on trust
- Whether a new person could build on this without a meeting
- Pass / fail on inspectability, and the smallest change that would make a fail into a pass (a comment, a test, a named owner — not a rewrite)

If you cannot see the artifact, say what evidence is missing and stop.

Revenue measure, not activity measure

For anything that touches leads, inbox, ads, or support. Paste before the agent starts.

Paste the block as written.

You are grading work that is allowed to touch the cash register.

Workflow: [speed-to-lead / booked meetings / conversion / CAC / revenue / other]
Activity people will be tempted to count: [emails sent, leads scraped, tickets closed, ads prepared]
Ordinary measure we already use: [name the number and where it lives]

Write a one-screen grading card:

- Done means (the ordinary measure moved, or a named person accepted a blocked external step as theirs)
- Not done means (activity completed, artifacts exist, dashboard is green, ordinary measure is flat)
- Evidence the check ran (query, report, screenshot, CRM field — not the agent's count)
- Forbidden optimizations (the activity metric, and one milder cheat: weakening the definition of a lead, a close, or a conversion)

If the ordinary measure is unnamed, Status: Draft. Do not start outreach, sending, or closing.

Grader-hack diagnostic

One team, thirty minutes. The evaluation incident at business volume: the agent found a score it could move.

Paste the block as written.

You are diagnosing whether we are grading activity and calling it work.

I will paste: the agent's stated objective, what the dashboard celebrates, and 5 recent outputs (or I will describe them).

For each output, score 0–2:
- 0 = only process exists (plan, report, status, request for approval)
- 1 = an artifact exists, but the ordinary business measure did not move
- 2 = an ordinary measure moved, or a named person signed a blocked step

Then tell me:
- the two outputs most likely to be a grader hack (emails sent, tickets vanished, tests weakened, demo that stops at an unconnected account)
- the ordinary measure we should have used
- the smallest definition-of-done card that would have caught the worst one

Do not recommend a new platform. Do not invent a metric we do not already collect.