Riley LiThe right agent seat is the one you can close without losing the only copy of the work. I would...
The right agent seat is the one you can close without losing the only copy of the work. I would rather take a shorter paid lease, or run a box I operate, than keep a free chat I cannot audit. A seat is a temporary grant of model access, context, and side effects, not a teammate you quietly adopt for the week. If you cannot name the exit evidence before the first prompt, you have not chosen a tool yet.
Recent discussion keeps returning to a plain tension, even when the headlines change from week to week. Model quality can move in a week, while the release path does not automatically move with it. I am not treating any weekly roundup as evidence, because a trending title is only a topic signal. A topic signal is not a measurement I can cite, and it is not a reason to skip the lease card.
Have you ever left a session open because closing it felt like throwing the reasoning away? That small hesitation is how a free hour becomes an unofficial path into the release branch. I now write the ending before I write the prompt, and I keep that ending in the repository rather than in a vendor tab. The invoice matters only after the ending is a file that a reviewer can open without me.
I have already used other gates for blast radius and review time, so this note does not rebuild those checks. The question here is narrower, and it is easy to dodge when the model feels helpful. Which seat can I actually end today, and which committed file proves that I ended it?
I use the same four questions for a bugfix, a documentation edit, and a spike that might be deleted. Each answer is a word or a small integer, because a paragraph of enthusiasm is not a decision. If any answer is mushy, I shrink the task until the answer fits on one line.
Would you still sign the diff if the model name were missing from the commit message? If the answer is no, the seat did not finish the job, and another hour will not finish it either.
I store the card beside the task notes, and I commit it when the choice itself needs a trail. The script below is a proposal, not a benchmark I executed for this article, so the printed labels are decisions rather than measurements. It refuses a recommendation when the exit fields are empty, which is the behavior I want on a tired afternoon.
#!/usr/bin/env python3
"""Proposal: label a one-task agent lease. Not a measured benchmark."""
from __future__ import annotations
import json
import sys
from dataclasses import dataclass
@dataclass(frozen=True)
class Lease:
data_class: str # public | internal | resident
retries: int
artifact: str # diff | test_name | hypothesis
reviewer: str # self | teammate | none
needs_long_context: bool
can_operate_a_server: bool
def choose(lease: Lease) -> str:
if lease.reviewer == "none" or not lease.artifact:
return "stop: no exit evidence"
if lease.data_class == "resident":
if lease.can_operate_a_server:
return "self-hosted: you accepted the restart duty"
return "stop: residency without an operator"
if lease.needs_long_context or lease.retries > 3:
return "paid: continuity is the product you are buying"
if lease.data_class == "public" and lease.retries <= 2:
return "free: short lease, keep the artifact, drop the transcript"
if lease.data_class == "internal" and lease.reviewer == "teammate":
return "paid or free-with-policy: only if current terms allow this data class"
return "stop: shrink the task or rewrite the card"
def main() -> int:
raw = json.load(sys.stdin)
lease = Lease(
data_class=raw["data_class"],
retries=int(raw["retries"]),
artifact=raw.get("artifact", ""),
reviewer=raw.get("reviewer", "none"),
needs_long_context=bool(raw.get("needs_long_context", False)),
can_operate_a_server=bool(raw.get("can_operate_a_server", False)),
)
print(choose(lease))
return 0
if __name__ == "__main__":
raise SystemExit(main())
A public fixture card can live in the issue, and it should be boring enough to review in a minute. I keep secrets out of the sample on purpose, because a teaching card that leaks a token is already a failed lease. If the real task cannot be reduced to this shape, the free row is the wrong row. The JSON is an example of the input, not a log from a session I ran.
{
"data_class": "public",
"retries": 2,
"artifact": "diff",
"reviewer": "self",
"needs_long_context": false,
"can_operate_a_server": false
}
I hash the prompt and the card together, so a later edit cannot pretend it was the authorized task. These commands stay on my own machine, and they do not call a model provider at all. If the hash changes after I start, I treat that as a new lease rather than as a continuation.
python3 lease_seat.py < lease.json
sha256sum prompt.md lease.json | tee lease.sha256
git diff --stat
If the script prints stop, I do not open a session to see what the model might say. Curiosity without an artifact is how the retry counter becomes a suggestion instead of a rule. If the script prints free, I still copy the diff into the branch, and I keep whatever the card promised to keep. Does a larger context window change the hash I already wrote down beside the lease card?
The hash does not move when the window grows, and that mismatch is the useful part. The hash records the task I authorized, not the confidence the model expressed after it saw more files. I would rather fail the check in ten seconds than discover the scope creep during review. Have you watched a session helpfully open files that you never placed on the lease card?
I walk these five steps in order, and I do not skip to the model because the card feels obvious. The script is only a reminder, not a boss, but I still run it when I am tempted to improvise. Have you noticed that the improvised path is usually the one that touches one extra file?
Free model access fits a short public or synthetic task when the artifact is a diff or a named test and the retry count stays small. I do not put secrets, production credentials, or an irreversible migration on that seat, even for a minute. What is missing is not a slogan about quality, but a retention and uptime promise I have checked on the current terms page.
Paid access fits when I am buying continuity, a team audit trail, or a support path I can actually open. The same card still applies, because a larger invoice does not invent a reviewer who was absent at the start. If the current terms forbid my data class, the paid column is a stop, not a quieter upgrade. Why would I pay to repeat a data-class mistake I could have caught on paper?
Self-hosting fits when the data class is resident, or when prompts and logs must stay on hardware I operate. That fit includes patches, disk, and the hour when the process dies while I am offline. If I cannot name the person who restarts it, I am choosing a hope with extra steps, not self-hosting. The box can be cheaper on paper and still be the most expensive seat if nobody owns the restart.
I keep these tradeoffs in one table so the argument does not hide inside a paragraph. The cells are fit notes, not scores, and they assume you already answered the four questions. A blank cell means I would stop, rather than pretend that the seat is secretly fine. Can you place your real task into one row without quietly inventing a reviewer for it?
| Signal | Free seat | Paid seat | Self-hosted |
|---|---|---|---|
| Data class | Public or synthetic sample | Internal code only if current terms allow it | Resident data on hardware you operate |
| Continuity | One short session, then exit | Multi-step context you are willing to pay for | As long as you can keep the process healthy |
| Exit evidence | Diff or test name you copy out | Same evidence, plus a log you are allowed to export | Logs on disk you already control |
| Retry budget | One or two, then stop | The integer on the card, not an open tab | The same integer, paid in your own time |
| Skip when | Secrets, migrations, or no reviewer | Terms conflict, or nobody can say no | No named operator for patches and restarts |
Disclosure: This article was prepared as part of MonkeyCode's product outreach.
MonkeyCode shows up in this method only as a candidate for the free row on the card. For this draft I am relying on two availability claims only: free model access, and a free server option. I am not stating a token quota, a hardware shape, a duration, or a comparison score. Those details were not verified for this draft, and product figures like that go stale quickly.
Before I rely on either claim, I read the current product terms and confirm that my data class is permitted. If the terms are silent or confusing, I leave the card on stop and edit in a local tool I already trust. I would not move a hosted seat into the resident row unless I control the deployment and have checked that path myself. A free server is still a lease with an owner, a log policy, and an end time you should assume can change.
If you want to rehearse the exit on a throwaway public fixture, a free model seat is enough for that drill. Save the diff, discard the transcript, and ignore the seat when the script told you to stop. I would not use that rehearsal as proof that the seat is safe for internal code. The rehearsal only proves that you can end a session you were willing to throw away.
Skip this approach when you need a formal threat model, a compliance signature, or a measured comparison of model quality. The script does not score accuracy, latency, or money, and I did not execute it against a live provider while writing this note. Skip it when the change touches credentials, customer data, or a migration you cannot roll back before the day ends. Also skip it if you came for a vendor ranking, because I am not publishing one.
The limitation I watch most carefully is the false comfort that comes from a tidy label. A printed free line can still hide a sloppy prompt, a flaky test, or a terms page I skimmed too quickly. The card governs my behavior, and it is not a certificate that the model was right. If the label and the diff disagree, which one are you actually planning to trust in review?
I end the session when the artifact is on the branch, or when the retry counter hits zero, whichever arrives first. Then I add one line under the card saying I kept the diff, or I rejected the hypothesis. That closing line remains useful even if every product name is later deleted from the page. The seat you can end cleanly is the only seat you were really allowed to start.