Elatify
AI Architecture

Leaked Claude Code Deep Dive: Context Management for Production AI Agents

Long-running agents don’t fail because the model is “weak.” They fail because the system has no guardrails for tokens. This post distills layered context defenses into patterns you can use in enterprise agent platforms.

9 min read
April 2, 2026
Budget tool outputs (before you summarize anything)
Tool results are the fastest way to blow your context window. Start with caps, previews, and persisted storage.
  • Per-turn aggregate budget (not per-tool only)
  • Per-tool maximum size
  • Preview in-context + full result stored for on-demand reads
  • Treat cache stability as a first-class constraint
Prune history cheaply
If you can remove stale turns from the API payload without harming current work, do it. Keep full UI history, project a smaller API view.
  • Non-destructive UI history for rollback
  • API view omits pruned messages
  • Track “tokens freed” so downstream logic sees reality
Micro-compaction: surgical cleanup
A set of small heuristics that clean old tool results when it’s “cheap” to do so (e.g., after cache TTL windows).
  • Clear older tool results when caches are already cold
  • Prefer server-side cache edits (keep local replay deterministic)
  • Never split tool-use / tool-result pairs
Projection summaries (collapse) instead of destructive rewrites
Maintain an append-only “collapse log” and generate a compact projection each turn. UI retains truth; API gets the view.
  • Commit log of collapses
  • Per-turn projection (read model) for API calls
  • Gating to avoid conflicts with expensive full summarizers
Last resort: LLM summarization + circuit breakers
When you must summarize, use a structured summary and guard it with retries, fallbacks, and a circuit breaker to prevent infinite failure loops.
  • Structured summary template (intent, files, errors, pending tasks)
  • Fit-to-window retry strategy (drop oldest groups until it fits)
  • Circuit breaker after consecutive failures
Implementation snippet: budget tool results
A simple pattern: cap tool output, keep a preview, persist the full result for later retrieval.
type ToolResult = { tool: string; content: string; persistedPath?: string };

function enforceToolBudget(
  results: ToolResult[],
  { perToolMaxChars, perTurnMaxChars }: { perToolMaxChars: number; perTurnMaxChars: number }
) {
  let remaining = perTurnMaxChars;

  return results.map((r) => {
    const cap = Math.min(perToolMaxChars, remaining);
    if (r.content.length <= cap) {
      remaining -= r.content.length;
      return r;
    }

    const preview = r.content.slice(0, Math.min(2000, cap));
    remaining -= preview.length;
    return {
      ...r,
      content: preview + "\n\n[full tool result persisted; read on demand]",
      persistedPath: persistToDisk(r),
    };
  });
}

Building enterprise agent systems?

Elatify helps teams design reliable, governable AI architectures—beyond prompt tweaks.