LatentEval

For builders

Hamming vs Cekura: voice agent eval platforms compared

A source-checked Hamming vs Cekura comparison: published pricing against sales-call pricing, a self-refereed 90% win rate, judge agreement quoted without a base rate, and how to run your own bake-off.

For builders

In brief

5 POINTS
  • In the search we ran on 2026-08-02, the top result for this query was published by Coval, a third voice-agent eval vendor selling against both companies being compared.
  • Cekura publishes self-serve pricing at $0.25 per voice testing minute and $500 per month for its Startup plan, checked at source 2026-08-02 and again 2026-08-06.
  • Hamming lists Startup, Agency, and Enterprise tiers with no figures; every price requires booking a sales call with the CEO.
  • Hamming's homepage claims a 90% win rate in head-to-head bake-offs, a fraction only Hamming is positioned to assemble.
  • A fair bake-off fixes the scenario set and sample size first, then reports pass rates per intent with confidence intervals.

The top result for this comparison, in the search we ran on 2026-08-02, was a page published by Coval, a third voice-agent eval platform that sells against both companies in its own title.1 Neither Hamming’s site nor Cekura’s showed up with an answer of its own.

Coval published that page in March 2026, it closes by inviting the reader to consider Coval, and one of its facts had drifted by the time we read it. It describes a $30 per month Developer plan for Cekura. Cekura’s pricing page carries no plan under that name: the three plans listed there are pay-as-you-go, Startup, and Enterprise. The only $30 per month on that page buys an additional seat on top of the first free one.2 So on the day we looked, the best-ranked answer to this question was a competitor’s comparison aging in place.

Pick Cekura if you want to start today under published prices: $0.25 per voice testing minute, $0.05 per monitored call, a $500 per month Startup plan, and 300 free credits to start with no card required.2 Pick Hamming if you are buying voice-specialist depth for an enterprise or agency deployment and you accept that every price requires a sales call.3 Treat both vendors’ headline quality numbers, Hamming’s “90% win rate” above all, as the vendors’ own counts until your own pre-registered bake-off replaces them. Skip that last step and your vendor decision rests on a number produced by the vendor it favors.

Both platforms sell the same loop

Strip the adjectives from both homepages and the two products describe the same loop: generate test scenarios from the agent’s own prompt, place simulated phone calls against it in parallel, score the resulting conversations with LLM judges, then keep scoring live traffic once the agent ships.

Where they part is emphasis. Hamming’s page adds the depth you would expect from a voice specialist: DTMF and IVR emulation, speech-level analysis past the transcript, red-teaming for jailbreaks and prompt injection, claimed support for 65+ languages and regional accents, and chat agents covered alongside voice.4 Cekura’s page leans operational: observability with interruption tracking, latency, sentiment, and gibberish detection, alerting into Slack and webhooks, judge-prompt tuning against your own recordings, and named integrations with Vapi, Retell, Pipecat, LiveKit, ElevenLabs, Synthflow, Cisco, and Five9.5

Both claim SOC 2. Hamming states Type II, certified December 2025, with a HIPAA business associate agreement available; Cekura lists SOC 2, HIPAA, and GDPR.45

We have not run either platform. This is a documents comparison: every cell below comes from a page the vendor published, read on 2026-08-02, and the last column records whether you could check the claim without talking to anyone. Every price and headline figure in the table was read again on 2026-08-06, and none of them had changed.

TABLEShow full table (5 rows)Showing full table (5 rows)
What a buyer asksHammingCekuraCheckable before a sales call?
What does it cost?Startup, Agency, and Enterprise tiers, all “Contact us”3$0.25/min testing, $0.05/monitored call, $500/mo Startup plan2Cekura yes; Hamming no
Can I try it today?The pricing page books a call with the CEO3300 free credits, no card required2Cekura yes; Hamming no
How many calls at once?“50K+ concurrent test calls”, a homepage ceiling4Plan caps: 10 concurrent self-serve, 50 on Startup2Cekura’s caps are plan terms; Hamming’s figure is homepage copy until a contract states it
Is the judge any good?“95-96% agreement with human evaluators”4Judge tuning offered, no headline accuracy figure5Neither publishes a methodology you could rerun
Will I win with it?“90% win rate in head-to-head bake-offs"4"Trusted by 70+ Conversational AI companies”5Both are the vendor counting for itself

Read that last column on its own and the five rows fall into two groups. On price, trial, and concurrency the two vendors sit on opposite sides of the line, because one publishes and the other does not. On judge quality and win rate the line stops sorting them: both claims land on the far side together, and the question a buyer can still answer from the outside shrinks to which sentences can be checked at all.

Five buyer questions split by whether the answer is published: Cekura's price, free credits, and concurrency caps sit on the checkable side and Hamming's three tiers, sales call, and 50K+ homepage figure on the not-checkable side, while on judge quality and win rate both vendors fall on the not-checkable side together.
Five buyer questions sorted by one thing: whether the answer is published. Cekura's prices, free credits, and plan caps can be read on its pricing page; Hamming's sit behind a sales call. On the last two questions the split disappears: neither vendor publishes a method you could rerun, and each is counting for itself. The page recommends either platform depending on what you are buying. Structural diagram. Every cell is a vendor's own published page, read 2026-08-02. No benchmark data.

The first two rows deserve their own section, because they decide who can even shop here without a meeting.

Cekura’s prices are on the page. Hamming’s are behind a sales call.

Cekura’s pricing page reads like a utility bill. Pay-as-you-go bills at $0.25 per voice testing minute, $0.05 per monitored call, and $0.025 per chat reply. One seat is free and each additional one costs $30 per month. Self-serve logs are kept 30 days, and the 300 free credits you start with come to about 60 minutes of testing by the page’s own reckoning. The $500 per month Startup plan bundles roughly 2,000 testing minutes and 10,000 monitored calls, with 50 concurrent calls, ten seats, five projects, and 90-day retention. At the pay-as-you-go rates printed on the same page, either allowance on its own would bill at $500: 2,000 minutes at $0.25, or 10,000 monitored calls at $0.05. Enterprise moves to annual contracts and adds VPC or on-prem hosting, SSO, SCIM, and audit logs.2

Hamming’s pricing page lists Startup, Agency, and Enterprise tiers and no figures at all. Each tier’s button books a call with the CEO, and the Startup tier promises “startup-friendly pricing from a YC S24 company” without saying what that pricing is.3 What can be said about Hamming’s cost is that it sits behind a sales call. That is a legitimate way to sell enterprise software. It also means one side of this comparison can be priced from your desk while the other cannot, and it means any third-party page quoting a Hamming price is quoting something Hamming’s own site does not state.

Prices in this market rot fast. Coval’s March comparison already names a Cekura plan that is not on Cekura’s page, this page will drift the same way in time, and that is why every figure above carries the date we read it.

The pricing asymmetry is visible from your desk. The quality claims run the other way: only one party can see the data underneath them.

A win rate only the winner can count

Hamming’s homepage puts a number on its own competitive record: “With a 90% win rate in head-to-head bake-offs, Hamming is the proven choice.”4 The sentence carries no denominator: no count of bake-offs, no time window, no list of who was on the other side, and no referee outside Hamming. A win rate is a fraction, and the only entity that could assemble this one is Hamming itself, because the bake-offs a vendor knows about are the ones that ran through its own funnel. A team that evaluated two rivals and never contacted Hamming at all is a bake-off no such counter can see.

The figure may be true. As published, it is a conversion statistic about the vendor’s own sales pipeline, scored by the party it flatters, which is the same structural problem as a benchmark graded by the team it advertises.

The same discount applies across the aisle, at lower stakes: Cekura’s “Trusted by 70+ Conversational AI companies” is also the vendor counting for itself, though a customer count claims far less than a win rate over named rivals.5

Judge agreement, quoted without a base rate

Hamming also publishes an evaluator quality figure: “95-96% agreement with human evaluators”, attributed to higher-quality models and a two-step evaluation pipeline.4 Raw percent agreement is the weakest form an evaluator claim can take, for a reason that is arithmetic rather than suspicion. If 9 in 10 test calls pass, a judge that answers pass every single time agrees with humans 90% of the time while measuring nothing. Push the pass rate to 95 in 100 and the same do-nothing judge scores 95%, which is the bottom of the range Hamming publishes.

Cohen’s kappa exists to strip that chance agreement out, which is what our inter-rater reliability calculator computes from a labeled slice of calls, and the recurring failure patterns of LLM judges have names and specific tests of their own. Neither vendor publishes the pass base rate, the sample size, or a kappa, so the claim cannot be told apart from a judge coasting on an easy class balance.

Three aligned 100 unit strips showing 90 pass and 10 fail, a judge answering pass on all 100, and 90 units where the two overlap, which scores 90 percent raw agreement while catching none of the 10 failures.
Raw percent agreement is a form a judge can satisfy without reading anything. The strips are the hypothetical stated above, not Hamming's data: if 9 in 10 test calls pass, a judge that answers pass every time agrees with the humans on 90 of 100 calls and catches none of the 10 failures. A real evaluator could score the same, which is why the published form cannot tell the two apart without the base rate, the sample size, and a kappa. Structural diagram of a worked hypothetical stated in the body. The quoted vendor figure is Hamming homepage copy, read 2026-08-02, and is not plotted as data.

Cekura publishes no headline judge accuracy at all, and instead sells tooling for tuning the judge against your own recordings.5 On the merits that is the more defensible position of the two: a judge has to be calibrated per agent, on your own traffic, and a single global accuracy figure would not survive the move from one call mix to another.

Zoom out from the two vendors and the same question lands on the comparison you are reading: who wrote it, and what do they sell.

The best-ranked comparison came from a rival, and this page has a stake too

Coval’s comparison is detailed, current-looking, and ends where a competitor’s comparison tends to end: “If you are evaluating both and need those capabilities, consider also looking at Coval.”1 It was published on March 20, 2026 under the byline of Coval’s founding growth engineer, and neither Hamming nor Cekura had a page answering it in the results we pulled. Search results differ by user, by location and by date, so take that as one dated observation rather than a measurement of the market.

Our stake belongs under the same light. LatentEval is building in agent evaluation, the market both of these vendors sell into, though not in voice-agent QA specifically. The instrument this site is building toward is a fault-injection profiler, designed to seed a known fault into an agent run and report the contained share with a confidence interval, which is the shape this page has been asking both vendors for: a fraction with a denominator and an interval around it. It is an unshipped instrument, so this page claims no containment number of its own. The standard we have been holding Hamming and Cekura to is one we have already had to meet in public: our three-way benchmark of Claude deployments prints the sample every cell stands on, publishes the alternative constructions of its own headline index, and keeps a ledger of the corrections it has had to issue. Every figure above belongs to a page Hamming or Cekura published, and every footnote carries the URL and the date we read it.

That leaves you one reliable move: run the comparison yourself, on rules you fix before the first call.

Run the bake-off yourself, and fix the rules before the first call

A fair bake-off is mostly tedium, which is why vendors’ own numbers keep standing in for it. Four rules cover most of it:

  1. Write the scenario set from your own production intents before you contact either vendor, and freeze it. Score the path as well as the endpoint: a call that reaches the right booking through three misheard turns differs from a clean one, which is the argument for AI agent evaluation metrics that score the path alongside the outcome.
  2. Fix the sample size first. How many test calls you need to detect a given gap between two platforms is a power question, and our sample size and power calculator answers it before you spend a minute of testing credit. Step 1 runs one frozen scenario set against both platforms, which is a paired design. Use the calculator’s McNemar paired mode and give it a discordance rate; the independent-arms default will oversize the run.
  3. Report every pass rate per intent segment, with an interval. A blended score across password resets and billing disputes hides whichever segment is doing the failing, the same trap that sinks blended chatbot containment benchmarks, and the pass-rate confidence interval calculator turns each k-of-n into a Wilson interval.
  4. Count what the agent declined alongside what it failed. A voice agent that dodges the hard intents scores clean on the easy ones, and that dodged share is its refusal rate.

Either platform will happily run this experiment; simulated calls at scale are the actual product, and a pre-registered run of your own produces the denominator neither homepage gives you. If the agent you are evaluating types rather than talks, you are shopping in a different aisle: Datadog LLM Observability vs Langfuse covers the enterprise observability split, and Langfuse vs LangSmith the open-source three-way.

Whichever aisle you end up in, buy on a number you produced yourself.

Footnotes

  1. Coval, “Hamming vs Cekura: Voice AI QA Compared (2026)”, published March 20, 2026, bylined by Coval’s founding growth engineer. It closes with “If you are evaluating both and need those capabilities, consider also looking at Coval, which evaluates agents built on any platform”, and states that “Cekura publishes a $30/month Developer plan with 750 credits, one seat, 10 concurrent calls, one project, and email support. A 7-day trial with 300 credits is available with no credit card required.” Cekura’s own pricing page lists no Developer plan and no $30/month plan on either check; its $30/month figure is the price of an additional seat, and its 300 free starting credits carry no stated 7-day limit. https://www.coval.ai/blog/hamming-vs-cekura (checked 2026-08-02, re-read 2026-08-06 with the passage unchanged) 2

  2. Cekura pricing page: three plans are listed, “Pay as you go”, “Startup Plan”, and “Enterprise”, and none is named Developer. Pay-as-you-go is $0.25 per voice testing minute, $0.05 per monitored call, $0.025 per reply for chat, 10 concurrent calls, “1 seat free · $30/mo per additional seat”, 30-day log retention, and “300 free credits to start ≈ 60 min” with no card required. Startup is $500/month with about 2,000 minutes of voice testing, about 10,000 monitored calls, 50 concurrent calls, 10 seats, 5 projects, and 90-day retention. Enterprise is custom, on annual contracts, with VPC/on-prem hosting, SSO, SCIM, and audit logs. https://www.cekura.ai/pricing (checked 2026-08-02, re-read 2026-08-06 with every figure unchanged) 2 3 4 5 6

  3. Hamming AI pricing page: Startup, Agency, and Enterprise tiers, all listed as “Contact us” with no dollar figures or metering axes; each tier’s call-to-action books a sales call with the CEO; the Startup tier is described as “startup-friendly pricing from a YC S24 company”. https://hamming.ai/pricing (checked 2026-08-02, re-read 2026-08-06 with no figures added) 2 3 4

  4. Hamming AI homepage: “With a 90% win rate in head-to-head bake-offs, Hamming is the proven choice” (under the FAQ heading “How is Hamming different from other voice agent QA tools?”); “95-96% agreement with human evaluators” via “higher-quality models and a two-step evaluation pipeline”; simulations said to “predict reality with 95%+ accuracy”; “50K+ concurrent test calls”; 65+ languages and regional accents; DTMF/IVR emulation; red-teaming for jailbreak, prompt injection, and PII; production call replay; SOC 2 Type II certified December 2025 with a HIPAA BAA available. https://hamming.ai (checked 2026-08-02, re-read 2026-08-06 with every quoted figure unchanged) 2 3 4 5 6 7

  5. Cekura homepage: scenario testing, parallel calling, voice evals scoring empathy, responsiveness, and hallucinations, observability with gibberish detection, interruption tracking, latency, sentiment, and pitch, judge-prompt tuning against recordings, alerting to Slack, email, and webhooks; integrations named include Vapi, Retell, Pipecat, LiveKit, ElevenLabs, Synthflow, Cisco, and Five9; “Trusted by 70+ Conversational AI companies”; SOC 2, HIPAA, and GDPR listed. https://www.cekura.ai (checked 2026-08-02, re-read 2026-08-06 with the customer count unchanged). Cekura launched under the name Vocera, per Y Combinator’s launch page for “Cekura (formerly Vocera)”. https://www.ycombinator.com/launches/M57-cekura-formerly-vocera-testing-monitoring-for-ai-voice-agents (checked 2026-08-02) 2 3 4 5 6