In Closed Beta

Test and improve your voice agent prompts.

Track prompt versions, test conversations with simulated voice, text, or human testers, and score each run against your own criteria. See how your prompt changes affect the results.

ExampleRefund request · Prompt v5 · Custom scorecard00:50 / 01:10
640 ms responseBarge-inPolicy mismatch1.9 s pause
Human testerMaya R., frustrated
Your agentOpenAI Realtime · Ava
Your agent

Yes. Refunds post within 30 days.

Follows instructionsPass
Asks a clarifying questionPass
Refund policy accurateFail
Completes the taskNeeds review
Core benefits

One workspace for every prompt test

01

Track prompt versions

Keep a history of your prompt versions and test how each one behaves in conversation. Compare results to see where a change helps and where it introduces new problems.

02

Score runs against your criteria

Set evaluation criteria for your use case, from following instructions to completing a task or answering accurately. Score each conversation against those criteria to see where your prompt needs work.

03

Test with text, voice, and people

Run quick text checks, simulate spoken conversations, and test with people. Use each mode to investigate different issues and refine your prompt before putting it in front of users.

Bring prompt versions, test conversations, and custom scores together so your team can find issues and decide what to improve.

How it works

From prompt to feedback in four steps

01

Choose a model and add your prompt

Select a supported voice model, add your system prompt, and adjust the voice settings for the behavior you want to test.

02

Define a scenario and success criteria

Describe the conversation you want to test and set criteria for a successful run, such as instruction following, task completion, or answer accuracy.

03

Choose how to test

Run a text conversation, simulate a voice session, or test with a human participant. Choose the mode that helps you investigate the issue.

04

Review, revise, and compare

Review scores and flagged issues, save a new prompt version, and test again. Compare results across versions using the same scenarios and criteria.

Features

Tools for testing and improving your voice agent prompts

Human testing and feedback

See how people across accents, ages, and languages interact with your agent. Collect ratings and feedback on what felt clear, confusing, or frustrating.

Text and simulated voice tests

Run quick text checks or test spoken conversations with simulated callers. Create personas with their own voice, pace, background noise, and temperament, then run them in parallel.

Prompt version history

Save each prompt and voice configuration with a note. Review earlier versions, compare their scores, and restore a previous configuration.

Custom scoring

Define criteria for task completion, instruction following, and answer accuracy. Add policies or reference documents when checking factual responses.

Recordings and transcripts

Dual-channel audio with synced transcripts. Share a clip of the exact failing turn.

Latency and interruptions

Turn-by-turn response time, barge-in recovery and dead air, measured on the audio itself.

Regression tests in CI

Run a suite on every prompt or model change. Block the merge when pass rate drops.

Example results

See how your prompt versions perform

Review run scores, inspect flagged issues, and decide what to test next. Example data shown for illustration.

Pass rate92.4%+3.1 pts vs previous prompt
p50 response latency610 ms−140 ms
Barge-in recovery97%+5 pts
Hallucination rate1.8%−0.9 pts
Pass rate by runLast 14 runs
Run 1271Run 1284
ScenarioPass ratep95 latencyStatus
Refund request, duplicate charge88%1.24 sFailing
Change delivery address98%0.71 sPassing
Cancel subscription, retention offer94%0.88 sPassing
Caller speaks Spanish mid-call81%1.02 sFlaky
Angry caller asks for a human96%0.64 sPassing
Voice platforms

Supported models and integrations

Choose from supported voice models and speech providers, or connect an existing agent. Compare how your prompt behaves across different model and voice configurations.

Speech-to-speech
OpenAI Realtime
Gemini Live
GPT Live
Text-to-speech
ElevenLabs
Cartesia
OpenAI TTS
Agent platforms
Vapi
Retell
LiveKit
Pipecat
Telephony and CI
Twilio
SIP
WebRTC
GitHub Actions
Pricing

Plans for your testing workflow

Prices shown billed annually. Monthly billing is 20% more.

Starter

For one agent getting ready to launch

$199/ month
Request a demo
1 agent1,000 simulated calls25 human-tester callsRecordings and transcripts

Team

Most teams

For teams shipping agent changes weekly

$799/ month
Request a demo
Up to 5 agents10,000 simulated calls200 human-tester callsCI regression suitesAccuracy scoring against your docs

Enterprise

For regulated and high-volume contact centers

Custom
Talk to sales
Unlimited agentsDedicated tester pool and languagesSSO, audit log, data residencyVPC deployment

Questions

Who are the human testers?

A vetted, paid pool of testers screened for clear audio and attention to detail. You choose languages, accents and demographics per run, and every tester follows your scenario brief.

Can I test an existing voice agent?

Yes. If your agent answers a phone number, SIP trunk or WebRTC room, Cullman can call it. Platform integrations add richer logs but are optional.

How are test runs scored?

Each run is evaluated against the criteria you define, such as instruction following, task completion, and answer accuracy. For accuracy checks, add policies, FAQs, or a knowledge base to review responses against your source material.

How long does a human-tested run take?

Most runs of up to 50 calls finish within a few hours during business hours. Simulated runs finish in minutes.

Can I run tests from CI?

Yes. Trigger a suite from GitHub Actions, GitLab or any CI with one API call, and fail the build on a pass-rate threshold.

What happens to call recordings?

Recordings are encrypted at rest and kept for 90 days by default. Enterprise plans can set retention and data region.

See how your voice agent prompt performs

Bring a prompt to a live demo. We'll walk through a test conversation, score it against your criteria, and show you how to compare results across prompt versions.

Request a demo