NovaSynth logo

NovaSynth

Freemium
訪問

NovaSynth is a pre-production testing platform for AI agents that uses synthetic users across real SIP phone calls, LiveKit audio, and chat. It helps teams rehearse multi-turn interactions, surface failures, and validate fixes before releasing agents to customers.

NovaSynthとは?

NovaSynth is a pre-production testing platform for AI agents. It creates synthetic users with defined goals, moods, accents, languages, interruptions, and adversarial behaviors, then sends them through voice or chat interactions with an agent. The purpose is to rehearse complete conversations, identify failures before release, and validate fixes without calling a live production system.

The platform supports real SIP telephony, LiveKit audio, and chat testing. Runs can be organized across persona and scenario combinations, captured as traces, and scored to help teams investigate problems in multi-turn agent behavior.

NovaSynthでできること

Synthetic personas and scenarios

Define callers by personality, background, language, behavior, goals, moods, accents, and interruptions. NovaSynth can also suggest edge cases from agent context such as prompts, product requirements, workflows, and documentation.

Real voice and chat interactions

Test agents through SIP calls using Twilio, Plivo, or Vobiz, LiveKit real-time audio, and chat over HTTP. Voice tests use real telephony, audio, and latency rather than only mocked transcripts.

Batch and concurrency testing

Run persona-by-scenario matrices under a shared batch and compare results together. Concurrency can be increased for load testing, with the site documenting support for up to 2,000 concurrent lines when provisioned for that use.

Sandboxed tool behavior

Reads can be replayed and writes can be sandboxed, allowing teams to exercise agent decisions end to end without sending write operations to production systems.

Trace capture and scoring

Synthetic calls are captured as full traces, including audio, and scored with more than 30 voice AI metrics. Noveum's wider evaluation platform lists 112 calibrated metrics across areas such as task progress, safety, hallucination, and conversational quality.

Closed-loop failure investigation

The workflow connects persona and scenario setup with scorers, root-cause investigation, recommendations, and a later validation run so teams can check whether a change resolved the observed failure.

利用シーン

“Pre-release voice-agent rehearsal”

A team preparing a phone-based agent can call a test number or staging SIP endpoint with synthetic customers before launch, exposing failures in turn-taking, de-escalation, interruptions, silence handling, and recovery.

“Adversarial safety testing”

Teams can run prompt-injection probes and other deliberately difficult conversations to check whether an agent refuses unsafe requests and preserves its intended boundaries.

“Regression testing after agent changes”

After changing a prompt, model, workflow, or tool, teams can rerun comparable persona and scenario batches to identify behavioral regressions across complete multi-turn journeys.

“Load and capacity rehearsal”

Engineering teams can raise concurrency to examine how an agent behaves under many simultaneous calls. The source notes that large load runs should be provisioned in advance.

“Chat-agent quality checks”

Teams testing HTTP chat agents can use synthetic conversations and credit-metered text runs to evaluate scenarios that do not require telephony, including topic switches and adversarial intent.

よくある質問

Does NovaSynth call a live production system?

No. NovaSynth is described as a pre-production product. It calls the agent endpoint supplied by the team, such as a test number, staging SIP endpoint, or unreleased agent. Live production evaluation is described as a separate post-production module.

What agent context is needed to get started?

For a single-agent architecture, the site recommends providing the system prompt along with a PRD, BRD, workflow documentation, or other material describing the agent's intended behavior. If the prompt cannot be exposed, teams can provide a description of the agent instead, though recommendations may be less precise. For multi-agent systems, the site recommends SDK integration so context can come from traces.

Which connection types are supported?

The site lists SIP through Twilio, Plivo, and Vobiz; LiveKit; Pipecat; LangChain and LangGraph; OpenAI and Anthropic SDKs; chat over HTTP; and custom or ground-up agents connected through an SDK or lightweight hooks around LLM calls.

Can NovaSynth test tools without changing production data?

Yes. The product describes replaying reads and sandboxing writes so a changed decision can be exercised end to end while write operations do not reach production systems.

How is NovaSynth usage metered?

NovaSynth is available on all listed Noveum plans but is credit-gated. The pricing page states that voice testing uses 25 credits per call minute and text testing uses 10 credits per conversation; monthly credits reset each billing cycle, while purchased add-on credits do not expire.

クイック情報

Product type
Pre-production AI agent testing platform
Testing modes
SIP voice, LiveKit audio, and HTTP chat
Supported connections
Twilio, Plivo, Vobiz, Pipecat, LangChain, LangGraph, OpenAI SDK, Anthropic SDK, and custom agents
Primary workflow
Persona and scenario setup → run → scoring → root-cause analysis → fix validation
Voice capacity
Up to 2,000 concurrent lines for load testing, with provisioning required for large runs
Usage model
Credit-gated; 25 credits per voice call minute and 10 credits per text conversation

NovaSynthの代替品

DeepEval logo

DeepEval

deepeval.com

DeepEval is an open-source LLM evaluation framework for testing and benchmarking AI applications. It helps developers run pytest-native evaluations, score outputs and agent traces, and iterate on systems across text, image, audio, and voice workflows.

Galileo logo

Galileo

www.galileo.ai

Galileo is an AI observability and evaluation platform for testing, debugging, and governing LLM and agent systems across development and production. It helps teams turn evaluation results into production guardrails and monitor AI behavior at scale.

Jev State logo

Jev State

jev-state.vercel.app

Jev State is a workspace for defining, testing, and regression-checking conversational decisions powered by Jev. It helps teams inspect why an agent takes a step and export runnable TypeScript and tests for an application.

Giskard logo

Giskard

www.giskard.ai

Giskard is an AI security and evaluation platform for testing conversational LLM agents before and after deployment. It combines automated red teaming, quality evaluation, runtime guardrails, and remediation workflows for teams responsible for reliable AI systems.

Promptfoo logo

Promptfoo

www.promptfoo.dev

Promptfoo is an AI security and testing platform for evaluating LLM applications, agents, models, and workflows. It helps developers and security teams find vulnerabilities, validate guardrails, map findings to security frameworks, and track remediation through development and deployment.

MCPJam logo

MCPJam

www.mcpjam.com

MCPJam is a testing and evaluation platform for MCP servers. It helps developers inspect servers locally, run user and model-based tests, and add behavior checks to CI/CD workflows.