Your Agent Rereads Every Tool Result. Build a Tiny Context Compactor in TypeScript.

Your Agent Rereads Every Tool Result. Build a Tiny Context Compactor in TypeScript.

# ai# typescript# agents# llm
Your Agent Rereads Every Tool Result. Build a Tiny Context Compactor in TypeScript.Bobby Hall Jr

Strands, DeepSeek and two new papers point at the same lever: context management. Build a tiny compactor that caps tool output, keeps failure lines, and elides old results. No API key.

An agent loop has a quiet habit.

Every time it calls the model, it sends the whole conversation again.

The system prompt. The task. Every tool call. Every tool result.

So a 20,000-token test log is not paid for once.

It is paid for on every turn after it lands.

That is the part I keep thinking about, because the last three weeks of agent news kept pointing at it.

  • On September 17, 2026, researchers published An Empirical Study of Harness Design for Coding Agents. Across 176 matched settings, they found context management matters more as the window gets tighter, and most of its benefit comes from preventing context-overflow failures. Rule-based elision before LLM summarization gave the best overall efficiency.
  • On September 21, 2026, the Strands Agents team released Strands harness and reported 28% lower token cost across six benchmarks on the same Claude or GPT models. In their words, "Our default context management largely drove the token-efficiency and accuracy": tool results over about 1,500 tokens get truncated, and compaction triggers above 85% of the window. That is their benchmark, not an independent one.
  • On September 22, 2026, CliffCompaction reported up to 50% lower cost under a bounded context. Its rule is strict: only truncate or drop content, never rephrase it, and never compact a compaction.
  • On October 3, 2026, DeepSeek published Harness v0.2.1-alpha.1, the latest build of its open-source, everything-is-a-plugin harness. In the Strands benchmark, DeepSeek Harness was the most token-efficient overall, and it also typically had the lowest accuracy.

Different teams. Same lever.

That last point matters too.

Spending fewer tokens is easy. Spending fewer tokens without losing the one line that mattered is the actual job.

So let's build a tiny version.

By the end, you'll run one command:

npx tsx compact.ts
Enter fullscreen mode Exit fullscreen mode

And watch one scripted agent run replayed under four context policies, with the tokens each one actually sends to the model.

No API key.

No model.

Just TypeScript.

One honesty note: this is my small model of the idea, not the Strands or DeepSeek implementation. The tools, the outputs and the token counter are mocked, and every number in the output comes from example inputs.

Table of Contents

  1. What We Are Building
  2. Project Setup
  3. Step 1: A Scripted Agent Run
  4. Step 2: Cap Big Tool Results
  5. Step 3: Elide Old Results
  6. Step 4: Build the View
  7. Step 5: Replay and Count
  8. Where It Breaks Down
  9. The Bigger Idea

What We Are Building

Cap, keep failures, elide, drop at 85%

The agent keeps a full log. That never changes.

What changes is the view: the slice of that log the model reads on each call.

Full log (never edited)
  ↓
Cap: big tool results become a preview + a ref
  ↓
Keep: failure lines always survive the cut
  ↓
Elide: old tool results become one-line stubs
  ↓
Drop: over 85% of the window, drop the oldest whole messages
  ↓
View (what the model reads this turn)
Enter fullscreen mode Exit fullscreen mode

Four rules. No summarizer.

Project Setup

mkdir tiny-context-compactor && cd tiny-context-compactor
npm init -y
npm install -D tsx typescript @types/node
Enter fullscreen mode Exit fullscreen mode

Save the following TypeScript blocks in order as compact.ts.

Step 1: A Scripted Agent Run

// compact.ts: a tiny context compactor for an agent loop.
// Everything is mocked: the tools, their output, and the token counter (about 4 characters per token).
// The run, the 32,000-token window and the thresholds are example inputs. No model, no API key.

// Step 1: a scripted agent run and a rough token counter
type Msg = { turn: number; role: "system" | "user" | "assistant" | "tool"; tool?: string; text: string };

const tokens = (s: string) => Math.ceil(s.length / 4); // a heuristic, not a real tokenizer
const lines = (n: number, f: (i: number) => string) => Array.from({ length: n }, (_, i) => f(i)).join("\n");
const FAIL = "FAIL src/checkout.test.ts > applies 10% coupon: expected 90, received 100";

const testLog = (failAt: number | null) =>
  lines(1800, (i) => (i === failAt ? FAIL : `PASS src/suite-${i % 97}.test.ts > case ${i} (${(i * 37) % 90 + 3} ms)`));
const source = (name: string, n: number) =>
  lines(n, (i) => `  const ${name}${i} = applyRule(cart.items[${i % 12}], rules.${name}); // line ${i + 1}`);

const RUN: [tool: string, call: string, output: string][] = [
  ["list_files", "list_files src/", lines(40, (i) => `src/module-${i}.ts`)],
  ["run_tests", "run_tests", testLog(1137)],
  ["read_file", "read_file src/checkout.ts", source("checkout", 220)],
  ["grep", "grep -n coupon src/", lines(30, (i) => `src/module-${i}.ts:${i * 7 + 3}: coupon`)],
  ["read_file", "read_file src/coupon.ts", source("coupon", 160)],
  ["read_file", "read_file CHANGELOG.md", lines(900, (i) => `- v2.${900 - i}.0: internal release notes, item ${i}`)],
  ["edit_file", "edit_file src/coupon.ts", "ok: 1 line changed"],
  ["run_tests", "run_tests", testLog(null)],
  ["git_diff", "git diff", "- return price;\n+ return price * (1 - coupon.percent / 100);"],
];

const LOG: Msg[] = [
  { turn: 0, role: "system", text: "You are a coding agent. Use tools. Verify before you finish." },
  { turn: 0, role: "user", text: "The checkout test is failing. Find the bug and fix it." },
  ...RUN.flatMap(([tool, call, output], i): Msg[] => [
    { turn: i + 1, role: "assistant", text: `call ${call}` },
    { turn: i + 1, role: "tool", tool, text: output },
  ]),
];
const CALLS = RUN.length + 1; // the model is called once per turn, plus once to write the answer
Enter fullscreen mode Exit fullscreen mode

The run is a coding agent fixing a failing checkout test. Nine tool calls.

Three of them are big: two run_tests logs of 1,800 lines each and a 900-line CHANGELOG. Only one line in the first test log matters.

The token counter is a rough heuristic, about four characters per token. Real tokenizers differ, so treat every count here as relative.

Step 2: Cap Big Tool Results

// Step 2: cap big tool results and keep the full text out of the context
const store = new Map<string, string>();
const CAP = 1500; // offload a tool result above this many tokens
const PREVIEW = 750; // keep this many tokens of it in context

function headTail(text: string, keep: string[] = []): string {
  const all = text.split("\n");
  const pick = (from: string[]) => {
    const out: string[] = [];
    for (const l of from) {
      if (tokens(out.join("\n")) + tokens(l) > PREVIEW / 2) break;
      out.push(l);
    }
    return out;
  };
  const head = pick(all);
  const tail = pick([...all].reverse()).reverse();
  const hidden = all.length - head.length - tail.length;
  return [...head, `... ${hidden} lines offloaded ...`, ...keep, ...tail].join("\n");
}
// Same budget, but lines that look like failures always survive the cut.
const isFailure = (line: string) => /FAIL|ERROR/.test(line);
const errorsFirst = (text: string) =>
  headTail(text, text.split("\n").filter(isFailure).slice(0, 5));

function capped(m: Msg, preview: (t: string) => string): string {
  if (m.role !== "tool" || tokens(m.text) <= CAP) return m.text;
  const ref = `offload://turn-${m.turn}`;
  store.set(ref, m.text);
  return `${preview(m.text)}\n[full result: ${tokens(m.text)} tokens at ${ref}]`;
}

// A tool the model can call to get back what was cut.
const retrieve = (ref: string, pattern: RegExp) =>
  (store.get(ref) ?? "").split("\n").filter((l) => pattern.test(l)).slice(0, 5);
Enter fullscreen mode Exit fullscreen mode

The most important function is errorsFirst().

headTail() keeps the start and the end of a big result, which is what most truncation does.

But a failing test does not politely sit at the top or the bottom of the log.

errorsFirst() spends the same budget and pins anything that looks like FAIL or ERROR into the preview.

The full text goes into a store, with a ref the model can pass to retrieve().

Step 3: Elide Old Results

// Step 3: elide old tool results with a rule, before any summarizing
function elided(m: Msg, now: number, keepRecent: number): string | null {
  if (m.role !== "tool" || now - m.turn <= keepRecent) return null;
  store.set(`offload://turn-${m.turn}`, m.text);
  return `[elided: ${m.tool} from turn ${m.turn}, ${tokens(m.text)} tokens, ref offload://turn-${m.turn}]`;
}
Enter fullscreen mode Exit fullscreen mode

Once a tool result is more than two turns old, it becomes one line: which tool, which turn, how big, and where it lives.

The model still knows the result existed.

It just stops paying for it.

Step 4: Build the View

// Step 4: build the view for each model call. Truncate or drop. Never rewrite.
type Policy = { name: string; preview?: (t: string) => string; keepRecent?: number; window?: number };
const WINDOW = 32_000;

function view(call: number, p: Policy): Msg[] {
  // Always rebuilt from the original log,
  // so a compaction never compacts a compaction.
  let v = LOG.filter((m) => m.turn < call).map((m) => {
    const stub = p.keepRecent === undefined ? null : elided(m, call, p.keepRecent);
    return { ...m, text: stub ?? (p.preview ? capped(m, p.preview) : m.text) };
  });
  const size = () => v.reduce((n, m) => n + tokens(m.text), 0);
  const pinned = (m: Msg) => m.turn === 0 || call - m.turn <= (p.keepRecent ?? 0);
  const over = () => !!p.window && size() > p.window * 0.85;
  while (over() && v.some((m) => !pinned(m))) {
    v.splice(v.findIndex((m) => !pinned(m)), 1); // drop oldest, whole
  }
  return v;
}
Enter fullscreen mode Exit fullscreen mode

The important part is the first line of view().

The view is rebuilt from the original log on every call. Nothing is edited in place, so a stub is never made from another stub.

Then, if the view is still above 85% of the window, it drops the oldest unpinned message, whole. The system prompt, the task and the last two turns are pinned.

Truncate or drop. Never rewrite.

Step 5: Replay and Count

// Step 5: replay the same run under each policy and count what the model actually reads
const POLICIES: Policy[] = [
  { name: "naive" },
  { name: "cap", preview: headTail },
  { name: "cap+errors", preview: errorsFirst },
  { name: "full", preview: errorsFirst, keepRecent: 2, window: WINDOW },
];

console.log(`${CALLS} model calls, ${RUN.length} tool results, window ${WINDOW.toLocaleString("en-US")} tokens (example run)\n`);
console.log(`${"policy".padEnd(12)} ${"billed".padStart(9)}   ${"peak".padStart(9)}   ${"fits".padEnd(14)} FAIL seen at call 3`);
for (const p of POLICIES) {
  let billed = 0, peak = 0, overflow = 0;
  let sawFail = false;
  for (let call = 1; call <= CALLS; call++) {
    const v = view(call, p);
    const size = v.reduce((n, m) => n + tokens(m.text), 0);
    billed += size;
    peak = Math.max(peak, size);
    if (size > WINDOW && !overflow) overflow = call;
    if (call === 3) sawFail = v.some((m) => m.text.includes(FAIL));
  }
  const fits = overflow ? `no (call ${overflow})` : "yes";
  const n = (x: number) => x.toLocaleString("en-US").padStart(9);
  console.log(`${p.name.padEnd(12)} ${n(billed)}   ${n(peak)}   ${fits.padEnd(14)} ${sawFail ? "yes" : "no"}`);
}

console.log(`\nretrieve("offload://turn-2", /FAIL/):`);
for (const l of retrieve("offload://turn-2", /FAIL/)) console.log(`  ${l}`);
Enter fullscreen mode Exit fullscreen mode

Run it:

npx tsx compact.ts
Enter fullscreen mode Exit fullscreen mode

Real output from my run:

10 model calls, 9 tool results, window 32,000 tokens (example run)

policy          billed        peak   fits           FAIL seen at call 3
naive          290,208      58,205   no (call 7)    yes
cap             22,905       4,242   yes            no
cap+errors      23,057       4,261   yes            yes
full             9,367       1,637   yes            yes

retrieve("offload://turn-2", /FAIL/):
  FAIL src/checkout.test.ts > applies 10% coupon: expected 90, received 100
Enter fullscreen mode Exit fullscreen mode

Here is how I read that table.

Billed input tokens by policy

naive would read 290,208 input tokens across ten calls, and it never gets the chance. It passes the 32,000-token window on call 7, right after the CHANGELOG read lands on top of the first test log.

cap fits easily. But the head-and-tail preview cut the one FAIL line out of the first test log. On the very next call, the model cannot see why the test failed.

cap+errors costs 152 more tokens than cap across the run, and the failure line is back.

full adds elision and the 85% drop rule. It reads 9,367 tokens across the run, with a peak of 1,637.

And here is a detail I like. In this run, the 85% rule never fires. Capping and elision already kept every view small.

That lines up with the harness study. Cheap rules first. The expensive machinery is a safety net.

Where It Breaks Down

This demo is small on purpose. Here is what sits right outside it.

  1. Recoverable is not the same as recovered. My retrieve() tool can get the FAIL line back. The harness study found that making elided content recoverable "adds machinery that models rarely use and yields no accuracy gain." A ref helps only if the model thinks to follow it. Put the signal in the preview.

  2. Regex is a guess about what matters. FAIL|ERROR works for this test runner. A stack trace, a warning that turns into an outage, or a JSON field named status will not match. Each tool deserves its own preview rule.

  3. Compaction can break your prompt cache. Providers cache a request's reused prefix. When a result turns into a stub, everything after it changes, and the cache has to warm up again. Strands ships prompt caching and context management as defaults side by side, and the two pull against each other. Elide at stable boundaries, not on every turn.

  4. Dropping whole messages has rules. Real chat APIs expect each tool result to follow its tool call. Drop one without the other and the request can fail. My demo drops plain text.

  5. The counter is fake. Four characters per token is a heuristic. Use your provider's token counting before you trust a threshold.

  6. Summaries drift. I left LLM summarization out on purpose. CliffCompaction's rule of never rephrasing exists because a summary of a summary slowly stops being true.

The Bigger Idea

Tool result
  ↓
Full log (the truth, never edited)
  ↓
Policy (cap, keep, elide, drop)
  ↓
View (what the model reads)
  ↓
Model call
Enter fullscreen mode Exit fullscreen mode

The model provides reasoning.

The tools provide evidence.

The log provides the truth.

The view decides what the model is allowed to pay attention to.

That is why the harness is becoming the product. Two agents on the same model can read very different conversations.

A cheap agent that forgot the failing line is not cheap. It is wrong at a discount.

Your context window is a budget. Spend it on evidence.


Try Roster

I'm building Roster around this idea: AI employees with real responsibilities, tools, memory and schedules, and a harness that decides what they need to see on every step.

If the same follow-ups, handoffs, and waiting loops keep eating your week, give them to an AI employee.

Try Roster →