Headroom
Open-source context compression layer for AI coding agents and LLM workflows, cutting repetitive tool outputs and JSON tokens with reversible retrieval.

Dhanji Bhagat
Founder, Emiote
Fully hosted platform. Automated backups and SLA.
Enterprise agent platforms from $40-$500/user/mo; raw API token bills
Private compute. Zero seat taxes; team runs ops.
$0/mo local machine / $5-$15/mo VPS (+ raw LLM API tokens)
Headroom is an open-source, local-first context compression layer for AI agents and LLM applications, created by Headroom Labs. Built with Rust and Python, it sits between coding assistants and model providers to compress repetitive tool outputs, logs, and JSON data. Using reversible retrieval (CCR), it shrinks token payloads while allowing models to fetch original content on demand.
1. Why Headroom Matters: The Context Bottleneck
Modern coding agents run iterative feedback loops. Claude Code, Codex, Cursor, Aider, and OpenClaw inspect files, run tests, query language servers, and execute shell commands. Each action appends raw output into the session history. Within five turns, a debugging session accumulates tens of thousands of tokens.
Most of this data is repetitive ceremony:
- Search results return hundreds of file paths where the model only references three.
- Test runners print thousands of passing test lines surrounding one assertion failure.
- Database queries and API calls return wide JSON arrays containing dozens of unused schema keys.
- Diff outputs repeat entire unmodified file contexts.
When raw tool payloads fill the context window, two problems occur. First, inference bills scale linearly with input tokens. On frontier models such as Claude 3.7 Sonnet or Claude Opus, input tokens cost between $3.00 and $15.00 per million tokens. Long-running agent sessions routinely spend dollars per task on unread boilerplate. Second, attention dilution degrades reasoning quality. Models fail to locate critical signals when buried under hundreds of lines of passing test logs.
Standard Agent Flow (Uncompressed Token Waste):
[Agent Action] -> [Raw Tool Output: 50,000 Tokens] -> [Full Prompt Sent to LLM]
(Bloated JSON, logs, paths) (Expensive, slow, noisy)
Headroom Architecture (Reversible Local Compression):
[Agent Action] -> [Raw Tool Output] -> [Headroom ContentRouter] -> [Compressed: 15,000 Tokens]
| |
+-> [Local SQLite CCR Cache] v
| [LLM Receives Focus]
| |
+<-- headroom_retrieve -+ (if needed)
Earlier tools attempted to solve this with lossy token pruning or arbitrary rolling windows. Microsoft’s LLMLingua applied small language models to drop low-perplexity words, but unpredictable token removal often corrupted structured JSON and code syntax. Simple rolling-window context managers discarded older conversation turns, causing agents to forget initial requirements.
Headroom approaches context optimization through structured, reversible transformation. Instead of guessing token importance with an external model, it identifies the data type and applies deterministic algorithms: statistical filtering on JSON arrays, deduplication on log lines, AST parsing on code, and local caching of original text.
2. Architecture & Compression Pipeline
Headroom is distributed as a Python package (headroom-ai), a native Rust core (crates/headroom-core exposed via PyO3 as headroom._core), an OpenAI and Anthropic compatible HTTP proxy (headroom proxy), and a TypeScript client SDK (sdk/typescript).
+---------------------------------------------------------------+
| YOUR APPLICATION |
| (Claude Code · Codex · Cursor · OpenClaw · SDK) |
+---------------------------------------------------------------+
|
v
+---------------------------------------------------------------+
| HEADROOM |
| FastAPI Proxy (:8787) · Inline SDK · MCP Server |
| | |
| v |
| Transform Pipeline |
| 1. CacheAligner (prefix stability monitoring) |
| 2. ContentRouter (Magika classifier + deterministic sniff) |
| | |
| +--> SmartCrusher (JSON: Kneedle, SimHash, zlib) |
| +--> LogCompressor (Logs: error preservation) |
| +--> SearchCompressor (Paths: structure dedup) |
| +--> CodeCompressor (Tree-sitter AST outlines) |
| +--> Kompress (ModernBERT ONNX fallback) |
| | |
| v |
| 3. CCR Store (Local SQLite: hash key + original content) |
| 4. Output Shaper (Verbosity steering + effort routing) |
+---------------------------------------------------------------+
|
v
+---------------------------------------------------------------+
| MODEL PROVIDERS (Anthropic, OpenAI, Bedrock) |
+---------------------------------------------------------------+
Core Subsystems
The pipeline evaluates each message block independently and fails open. If any compression step raises an exception, the payload passes through unchanged.
| Component | Implementation | Function |
|---|---|---|
| Proxy Control Plane | Python (FastAPI / Uvicorn) | Listens on port 8787, intercepts /v1/messages and /v1/chat/completions, and handles provider routing. |
| Rust Acceleration Core | crates/headroom-core | High-throughput parsing, Blake3 hashing, SimHash fingerprinting, Aho-Corasick matching, and DashMap storage. |
| ContentRouter | Hybrid Rust / Python | Detects payload formats using Google Magika ONNX classifier and regex heuristics, routing each block to one compressor. |
| SmartCrusher | Rust / Python | Compresses JSON arrays of dicts, strings, and numbers using statistical variance and Kneedle elbow detection. |
| Log & Search Compressors | Rust native | Strips repetitive file prefixes and passing logs while protecting anomaly lines matching errors or exceptions. |
| CodeCompressor | Tree-sitter (11 languages) | AST-aware code compression. Gated behind strict safety rules and disabled during active code editing. |
| CCR Store | SQLite / DashMap | Caches full uncompressed content locally, keyed by 24-character Blake3 hashes for on-demand retrieval. |
| Output Shaper | Proxy middleware | Injects tail system prompts for verbosity control and routes reasoning effort on thinking models. |
SmartCrusher: Statistical JSON Reduction
Tool outputs from database queries, Kubernetes APIs, and REST endpoints consist primarily of JSON arrays. SmartCrusher optimizes these payloads through a multi-stage statistical pipeline:
- Schema and Array Detection: Validates JSON structure. Arrays containing fewer than five items or under 200 tokens pass through untouched.
- Kneedle Elbow Sizing: Calculates bigram coverage curves to determine the exact retention threshold where additional items provide diminishing semantic return.
- SimHash Deduplication: Groups near-duplicate dictionary entries and retains representative prototypes.
- Zlib Diversity Validation: Compresses candidate subsets with DEFLATE to verify that semantic entropy matches the full dataset.
- Mandatory Anomaly Gates: Regardless of the target compression budget, SmartCrusher never drops entries containing error keywords (“error”, “exception”, “failed”, “critical”), numeric anomalies exceeding two standard deviations from the mean, or length outliers.
Live-Zone-Only Context Management
Earlier releases of Headroom included experimental rolling-window context managers that dropped older turns to fit token limits. That architecture was removed. Headroom version 0.37.0 enforces live-zone-only processing:
- It never deletes, reorders, or summarizes historical conversation turns.
- It compresses tool results and file reads in place within the newest turn.
- System prompts and user instructions pass through untouched, maintaining full compatibility with provider KV prefix caches.
3. Reversible Compression (CCR)
The central architectural compromise of traditional prompt compression is lossiness. If an optimizer drops a file path or log line that the model later requires, the agent fails.
Headroom addresses this with Compress-Cache-Retrieve (CCR):
Step 1: Compression
SmartCrusher compresses 1,000 JSON items down to 20 items.
Step 2: Caching
The original 1,000 items are written to a local SQLite database (ccr_store.db)
keyed by Blake3 hash 'a7f9c2'.
Step 3: Marker Injection
Headroom appends a marker to the compressed tool output:
[1000 items compressed to 20. Retrieve more: hash=a7f9c2]
Step 4: Tool Injection
Headroom injects the 'headroom_retrieve' function into the request payload:
{
"name": "headroom_retrieve",
"description": "Retrieve original uncompressed data from Headroom cache",
"parameters": { "hash": "The hash key from the compression marker" }
}
Step 5: Resolution
- Scenario A: The model answers using the 20 representative items. 90% token savings realized.
- Scenario B: The model needs the full dataset and calls headroom_retrieve('a7f9c2').
The Headroom proxy intercepts the tool call, fetches the original content from SQLite in ~1ms,
and supplies the full data back to the model without human intervention.
On Anthropic and OpenAI proxy paths, Headroom handles CCR resolution transparently. The client agent never sees intermediate retrieval round trips.
4. Visual Tour & Interface Workflows
Headroom visualizes token reduction through local metrics, terminal telemetry, and an integrated real-time dashboard.
Context Compression Pipeline
The diagram above documents a real SRE incident payload. A 55,957-token prompt containing thousands of container log lines is compressed to 24,340 tokens, representing a 57% input reduction. The critical FATAL error log at line 67 is detected by the anomaly gate and preserved byte for byte.
Live Cache and Savings Dashboard
Running headroom dashboard opens the local browser UI served directly from the proxy process:

The dashboard surfaces three critical operational metrics:
- Total Input Tokens Saved: Aggregate volume and dollar savings calculated against model list prices.
- Cache Hit Rates: Ratio of turns that successfully matched provider prefix caches.
- CCR Store Utilization: Active keys, storage footprint, and cache eviction cycles in the local SQLite database.
5. Total Cost of Ownership (TCO) Comparison
Headroom is free, open-source software under the Apache 2.0 license. Operating costs are limited to local machine compute or low-cost server infrastructure.
Workload Modeling: 10-Engineer Agentic Team
Consider a software engineering team of 10 developers using Claude Code or Cursor. Each developer runs 30 agent turns per day, averaging 60,000 input tokens and 1,500 output tokens per turn on Claude 3.7 Sonnet ($3.00/M input, $15.00/M output).
- Baseline Daily Input Volume: 10 devs x 30 turns x 60,000 tokens = 18,000,000 tokens/day ($54.00/day).
- Baseline Monthly Input Cost: ~22 working days = $1,188.00/month.
- Headroom Compression Savings: Measured 35% average reduction on coding agent workflows (JSON, search, logs).
- Net Input Tokens Saved: 6,300,000 tokens/day (~$415.80/month saved).
- Annual Token Cost Avoidance: ~$4,989.60/year.
| Dimension | Managed Enterprise Agent Platforms | Headroom (Self-Hosted OSS) |
|---|---|---|
| Software License | $40 to $500 / user / month (Devin, etc.) | $0 / month (Apache 2.0) |
| Proxy & Storage Infrastructure | Bundled in vendor markup | $0 (Local laptop) or $10/mo (Shared team VPS) |
| Model Token Pricing | 15% to 50% vendor platform markup | Raw Provider Rates (Direct API keys) |
| Context Retention Policy | Proprietary cloud storage | Local SQLite (~/.headroom/ccr_store.db) |
| Data Privacy Boundaries | Transcripts stored on third-party cloud | 100% On-Device (Payloads never leave machine) |
| Annual Software Cost (10 Devs) | $4,800 to $60,000+ / year | $0 to $120 / year (Compute only) |
| Net Annual Token Savings | $0 (Vendor captures margin) | ~$4,900+ saved in direct API expenses |
6. The Bad: What to Know Before Adopting
Operating Headroom in real development environments reveals distinct constraints:
1. HEADROOM_BEACON is Enabled by Default
Headroom contains two separate telemetry systems. While HEADROOM_TELEMETRY is off by default, HEADROOM_BEACON is enabled by default (opt-out). On each request, it transmits an anonymous summary payload to remote servers. This payload includes token counts, compression ratios, model identifiers, skip reasons, OS, and architecture.
Although the collector allowlists counters and strips message contents, teams handling regulated data or operating in air-gapped environments must explicitly disable this flag at startup:
export HEADROOM_BEACON=off
# or set the cross-tool convention:
export DO_NOT_TRACK=1
2. Code Compression is Bypassed in Most Coding Sessions
Headroom includes an AST-based CodeCompressor using tree-sitter grammars. In practice, code compression rarely fires during interactive development.
Two protective safety gates prevent it:
protect_recent_code=4: Source code blocks in the four most recent messages are never compressed.protect_analysis_context=True: If the user message contains words like “analyze”, “review”, “explain”, “fix”, or “debug”, all code compression is disabled across the entire conversation.
Because almost every developer prompt contains these terms, code passes through uncompressed. Headroom savings on coding agents come almost entirely from tool outputs, directory listings, grep results, and test logs, not from compressing source code.
3. Small Payloads Produce Negative ROI
Requests with tool outputs below 50 tokens are bypassed by default. For small prompts, parsing JSON, computing SimHash fingerprints, and executing the FastAPI proxy middleware introduces 1ms to 3ms of overhead without saving meaningful tokens. Headroom is optimized for heavy multi-turn agent workflows, not single-turn chatbots.
4. CCR Roundtrip Latency on Model Misses
When SmartCrusher compresses an array and the LLM determines that it cannot answer without the missing items, the model issues a headroom_retrieve tool call. While local SQLite retrieval takes less than 2 milliseconds, the resulting network round trip back to Anthropic or OpenAI costs between 800ms and 2,500ms. If a workload frequently triggers CCR retrieval, the added inference latency can outweigh the token savings.
5. Native Extension Build Requirements
Headroom relies on a compiled Rust extension (headroom._core) and ONNX Runtime. Prebuilt wheels are published on PyPI for standard architectures, but custom container builds (Alpine Linux musl or older Python versions) require a full Rust toolchain, maturin, and C compilers to build tree-sitter bindings.
7. Quickstart & Deployment
Headroom can be deployed as an agent wrapper, a standalone proxy, or an inline SDK.
Option A: Wrap an Existing Agent (Zero Config)
The headroom wrap command starts the local proxy in the background, configures environment variables, and launches the coding agent:
# 1. Install CLI and optional extras
pip install "headroom-ai[all]"
# 2. Wrap Claude Code
headroom wrap claude
# 3. Wrap Codex or Cursor
headroom wrap codex
headroom wrap cursor
# 4. Remove wrappers when finished
headroom unwrap claude
Option B: Standalone HTTP Proxy
Run the proxy as a local service and point your API clients to port 8787:
# Start proxy with default settings
export HEADROOM_BEACON=off
headroom proxy --port 8787
# Configure tools to route through proxy
export ANTHROPIC_BASE_URL="http://127.0.0.1:8787"
export OPENAI_BASE_URL="http://127.0.0.1:8787/v1"
Option C: Inline Python Compression
To integrate context optimization directly into custom applications:
from headroom import compress
from openai import OpenAI
messages = [
{"role": "system", "content": "You are an automated code reviewer."},
{"role": "user", "content": "Inspect this test log output:"},
{"role": "user", "content": open("heavy_test_run.log").read()}
]
# Compress payload before inference
result = compress(messages, model="gpt-4o")
client = OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=result.messages
)
print(f"Original tokens: {result.original_tokens}")
print(f"Tokens sent: {result.compressed_tokens}")
print(f"Saved: {result.tokens_saved} ({result.compression_ratio:.1%})")
Option D: TypeScript SDK
import { compress } from "headroom-ai";
const messages = [
{ role: "system", content: "You are an automated triage assistant." },
{ role: "user", content: JSON.stringify(largeApiPayload) }
];
const result = await compress(messages, { model: "claude-3-7-sonnet" });
console.log(`Saved ${result.tokens_saved} tokens`);
8. ReframeHub Architectural Insight
The primary engineering insight in Headroom is that reversible compression decouples token budgets from reasoning loss.
Traditional prompt compression treated token reduction as an information-loss problem. Optimizers attempted to predict which words an LLM might need, leading to conservative compression ratios (10% to 20%) or catastrophic hallucinations when essential tokens were dropped.
By introducing a local, low-latency key-value store (CCR) and injecting a standardized retrieval tool, Headroom transforms prompt optimization into a hierarchical cache problem:
- The prompt sent to the LLM functions as an index or abstract.
- The local SQLite database functions as primary storage.
- The model itself acts as the cache-invalidation agent, retrieving full payloads only when the index proves insufficient.
This architecture proves that token optimization does not require larger context windows or smaller models. It requires treating the LLM context window as a transient CPU cache rather than a permanent database.
9. Who Should Use This?
Good Fit
- Teams running high-volume coding agents: Engineering teams spending thousands of dollars monthly on Claude Code, Cursor, Codex, or OpenClaw sessions.
- Data-heavy agent pipelines: Applications that regularly pass database records, JSON dumps, API responses, or CI logs into agent context.
- Privacy-sensitive deployments: Organizations that need token reduction without routing prompts through third-party cloud optimization proxies.
Bad Fit
- Short, single-turn conversational chatbots: Customer-facing chat widgets where prompts average under 200 tokens.
- Code-only editing without tool execution: Workflows that only read and write small source files without invoking tests, build scripts, or terminal commands.
- Zero-latency real-time voice agents: Applications where a 2-millisecond proxy evaluation or potential CCR retrieval round trip violates strict sub-second response budgets.
Evaluating agent context compression or proxy architectures?
Reframe ($199) audits your coding agent infrastructure, token unit economics, CCR retrieval boundaries, and prompt cache hit rates. Diagnosis only.
Fixed $199 fee · 100% vendor-neutral review · 3-day delivery guarantee
