Edgee logo

Edgee

Freemium
Visit

Edgee is an agent gateway for coding teams that reduces token usage, routes requests across models, and provides usage visibility. Its Compression V2 works as a drop-in CLI layer for Claude Code, Codex, OpenCode, Cursor, and other supported agents.

What is Edgee?

Edgee is an agent gateway for coding agents and engineering teams. It sits between supported coding tools and language-model providers to provide token compression, model routing, usage tracking, and controls for credentials and budgets. Compression V2 focuses specifically on reducing the tokens sent to and generated by coding-agent workflows without requiring application code changes.

Compression V2 uses two layers and three independently configurable techniques. On the input side, it trims tool results and reduces the tool surface shown to the model by selecting task-relevant MCP servers, skills, and tool definitions. On the output side, it can make model responses more concise with selectable light, medium, or hard settings. Edgee describes the compression as semantically lossless for code-oriented tasks, while its published measurements distinguish between production results and controlled benchmarks.

The published production aggregate reports 15–20% lower token bills from compression alone across active customers over a rolling 30-day period. A separate SWE-bench Lite evaluation measured a 50% reduction with all three techniques enabled; this is a controlled benchmark rather than a forecast for every workload. Edgee also reports less than 12 ms P50 gateway overhead.

What can Edgee do?

Tool result trimming

Filters tool and CLI results before they reach the model, removing boilerplate, pagination markers, ANSI escape sequences, repeated headers, and verbose framing while preserving code-task context.

Task-aware tool surface reduction

Scores MCP servers, skills, and tool definitions against the classified task, then strips or down-scopes irrelevant tools for the model request. The full tool set remains available in the IDE.

Configurable output brevity

Reduces response verbosity without intentionally removing technical content. Light, medium, and hard settings let teams choose the desired level of concision.

Drop-in agent gateway

Runs through the Edgee CLI with supported coding agents and existing provider subscriptions or API keys, avoiding changes to agent application code.

Routing and fallback controls

Routes requests across configured models under budget rules, retries eligible transient provider errors, and can use fallback routes when credentials, models, and spending controls permit.

Usage observability

Provides dashboards, logs, usage and cost metrics, session views, debug logs, activity exports, and team-level attribution features such as repository and pull-request tracking.

Use Cases

“Reduce coding-agent token bills”

Teams can enable compression independently of routing to reduce input and output token consumption across live coding-agent sessions. Edgee's published production aggregate reports 15–20% lower token bills from compression alone, while actual results depend on workload and enabled techniques.

“Run MCP-heavy development workflows”

Projects that expose many MCP servers, skills, or tools can use task-aware tool surface reduction so the model receives a smaller, task-relevant tool definition set instead of the full surface on every request.

“Keep agents operating across providers”

Engineering teams can place supported agents behind one gateway, apply budget-driven routing, and configure retries or fallback routes for eligible transient provider failures.

“Standardize team usage and attribution”

Organizations can manage developer and squad usage, model access, plugins, spending controls, and usage exports, with GitHub integration for per-repository and per-pull-request attribution.

“Connect application code to a gateway API”

Developers building their own agent or application can use Edgee's Gateway API and compatible client interfaces, then verify requests and inspect traffic in the gateway logs.

Frequently Asked Questions

Do I need to change my coding agent or application code?

Supported coding agents can be launched through the Edgee CLI, for example with an `edgee launch` command, without changing the agent code. Existing API clients can use Edgee's Gateway base URL and key according to the relevant integration guide.

Can I keep my existing provider subscription or use my own API keys?

Yes. Edgee states that it works with existing Claude, Codex, and Cursor subscriptions, and it supports bring-your-own-provider credentials. Fallback routes and strategy destinations still need eligible credentials when BYOK-only access is used.

How much does Compression V2 save?

Edgee reports a 15–20% reduction in token bills from compression alone across active customers in a rolling 30-day production aggregate. Its separate SWE-bench Lite benchmark measured a 50% reduction with all three techniques enabled. The benchmark is not presented as a forecast for every fleet, and routing savings should be evaluated separately.

Does compression remove information needed for coding tasks?

Edgee describes Compression V2 as semantically lossless for code-oriented tasks. Its techniques target tool-result formatting, irrelevant tool definitions, and response verbosity; users can toggle the techniques independently and choose an output-brevity level.

Can Edgee run on our own infrastructure?

The documentation says Edgee can be deployed on an organization's own infrastructure through a licensed deployment of its proprietary gateway. Connected and headless modes have different configuration-synchronization and usage-enforcement behavior.

Quick Facts

Category
Developer Tool; AI Infrastructure
Primary users
Software developers, engineering teams, and coding-agent operators
Supported workflow
CLI-based coding-agent gateway and compatible application API
Named agent integrations
Claude Code, Codex, OpenCode, Cursor, and Copilot
Compression model
Two layers: input and output; three independently configurable techniques
Published production result
15–20% lower token bills from compression alone in Edgee's rolling 30-day customer aggregate

Edgee Alternatives

deadeye logo

deadeye

deepaksinghcs14.github.io

deadeye is a plugin for coding agents that selects a suitable model and effort level for each task, while trimming verbose command output before it enters context. It is built for Claude Code and also supports Codex CLI, Gemini CLI, Cursor, and Windsurf on an experimental basis.

HAR HQ logo

HAR HQ

harproject.dev

HAR HQ is a team governance and observability layer for organizations running coding agents with HAR. It standardizes verification, tracks AI-related work and spend, and helps engineering teams review and improve agent workflows.

GMI Cloud logo

GMI Cloud

gmicloud.ai

AI infrastructure for production inference, training, and fine-tuning on NVIDIA GPUs.

Semantic Kernel logo

Semantic Kernel

learn.microsoft.com

Semantic Kernel is Microsoft’s public GitHub software project for integrating large language model technology into applications. The repository includes implementation and documentation directories for .NET, Python, and Java development.

OneCLI logo

OneCLI

onecli.sh

OneCLI is an open-source platform for giving employees personal AI agents that can work across company tools. It provides isolated sandboxes, gateway-enforced policies, scoped credentials, approvals, and audit logs for team use.

Fireworks AI logo

Fireworks AI

fireworks.ai

Generative AI platform for serving, training, fine-tuning, and deploying open-source models