Per-step model routing
Predicts which model to use for each prompt or step so teams can balance quality and cost across agent workflows.
Not Diamond is an intelligent model router for coding agents that improves quality and lowers cost with secure, privacy-preserving enterprise deployment options.
Not Diamond is an intelligent model router for coding agents. It predicts which AI model to use for each prompt or step, with the stated goal of improving quality while lowering cost for engineering teams working on long-horizon agent workloads.
The product is positioned as a secure integration layer that works with existing harnesses and gateway infrastructure rather than acting as a gateway itself. The site says it is designed for production-grade workflows, supports continuous learning from usage signals, and offers privacy-preserving routing options for teams that need to keep payload data local.
Predicts which model to use for each prompt or step so teams can balance quality and cost across agent workflows.
Uses signals such as payload semantics, KV cache state, implicit feedback, session outcomes, sub-agent architecture, compaction events, and prior turns to make recommendations.
Integrates with existing gateway infrastructure rather than replacing it, and is described as stack agnostic through a secure API.
Supports continuous learning from harness usage and session outcomes, with reinforcement learning used to improve recommendations over time.
Offers privacy-preserving routing that relies on derived request metadata and locally computed features, without seeing payloads, inputs, or outputs.
Provides organization-wide analytics for enterprise customers, along with advanced security and admin controls, MDM deployments, SSO with SAML, SLAs, and 24/7 support.
Use Not Diamond when a team wants to lower inference spend across coding-agent workflows without manually choosing a model for every step.
Use it when agent quality matters and routing decisions need to account for prompt context, cache state, prior turns, and session outcomes.
Use it when you already have a gateway and harness in place and want routing to fit into that stack rather than replace it.
Use it when teams need privacy-preserving routing for sensitive workloads and want decisioning based on derived metadata instead of raw payloads.
Use it when an enterprise team needs centralized analytics, admin controls, SSO, MDM deployment support, and SLAs around routing infrastructure.
Intelligent model routing predicts which AI model to use for each prompt so a team can improve accuracy while reducing cost. Not Diamond applies that approach to coding-agent workflows.
No. The pricing page says Not Diamond is not a gateway; it integrates with your existing gateway infrastructure and works with your model gateway and harness of choice.
The product is described as purpose-built for long-horizon coding agent workloads and works with a variety of coding agent harnesses. The pricing page also notes that custom harness integration needs can be discussed with the team.
Yes. The pricing FAQ says Not Diamond works with Claude Code and is consistent with Anthropic's commercial terms of service when used solely with Anthropic models.
The pricing page says teams using Not Diamond achieve at least 20-40% cost savings, and often more, without degradation in quality relative to the frontier. It also says the router adds about 100-200ms per step on average.