Multi-turn conversational coherence
Trinity is positioned for multi-turn interactions that keep goals and constraints intact across long sessions, reducing the need to restate context.
Arcee AI open-weight foundation models for multi-turn conversation, tool use, and structured outputs. Deploy locally, on-prem, or in the cloud.
Arcee AI is a U.S.-based open-intelligence lab focused on open-weight foundation models. Its Trinity family is the main product line shown on the site, with models designed to run locally, on-prem, or in cloud environments while keeping the same general skill profile across sizes.
The product pages position Trinity around practical production use: multi-turn conversation, tool use, structured outputs, and long-context work. Arcee also emphasizes open-weight licensing, portable weights, and the ability to move workloads between deployment footprints without rebuilding prompts or workflows.
Trinity is positioned for multi-turn interactions that keep goals and constraints intact across long sessions, reducing the need to restate context.
The model family is trained for accurate function selection, valid parameters, and schema-compliant JSON, which makes it suitable for agentic tool workflows.
Arcee says the same skill profile is shared across sizes, so teams can test on smaller models and move to larger deployments without reworking prompts or playbooks.
The line includes open-weight releases and a managed API, giving teams a choice between local control and hosted deployment.
The Trinity page highlights long context support and efficient attention, aiming to lower the cost of extended-context workloads.
The model family is available in variants aimed at edge, cloud, and reasoning-oriented deployments, including Nano, Mini, and Large Thinking.
Use Nano when you need a model that can run locally on consumer GPUs, edge servers, mobile-class devices, or other latency-sensitive environments.
Use Mini for customer-facing applications, agent backends, and high-throughput services that need cloud or VPC deployment with open weights.
Use Large Preview or Large Thinking when the workload depends on complex toolchains, longer reasoning chains, or more demanding agent behavior.
Use the family for applications that need schema-compliant JSON, function calling, or other structured outputs that can be consumed by downstream systems.
Use the hosted API when you want managed access, or download the weights when you need to run the model with local control and infrastructure ownership.
Trinity is a family of open-weight language models for multi-turn conversations, tool use, and structured outputs. The source describes it as usable either through a hosted API or by downloading weights to run locally.
The Trinity page lists Nano for edge and privacy-sensitive deployments, Mini for cloud and on-prem production, and Large variants for cloud deployment. The family keeps the same general skill profile across sizes so you can move between footprints without changing prompts.
The source says Trinity supports accurate function selection, valid parameters, schema-true JSON, graceful recovery when tools fail, and coherent long-session conversation. It is designed for agent workflows that depend on structured outputs and tool orchestration.
Arcee describes Trinity Nano and Mini as having a 128K context window, while Trinity Large variants are listed with a 512K context window on the product page. The broader site also highlights long-context usage and sustained multi-turn coherence.
The source does not provide a public pricing page. It does mention a managed hosted API, open weights, and an enterprise support/contact-sales path on the Trinity page.