The Voice Agent Maturity Curve
Most voice-AI comparisons are apples-to-orangutans because the category is not a category — it is a curve. Five stages from IVR replacement to agentic operator, every named vendor placed, and the bet I would write today: stage-three capability is mispriced and the EU Council's March 13 position just bought operators an extra year of arbitrage.
It is the last Friday of March, the rain on the back-office window has been steady since lunch, and I have spent the better part of the week reading voice-AI pitch decks. Twelve of them, give or take. Two from founders I have known since the pre-LLM IVR days, three from companies I had not heard of on Monday morning, the rest from the cluster of well-funded entrants that have been working the QSR and reservations corners since the autumn. The decks have one thing in common, and it is the thing that prompted this essay. Every one of them describes their product as voice AI for restaurants. Every one of them lands on the same operator pain points — missed calls, drive-thru throughput, after-hours coverage, multilingual guests. And every one of them, when you push past the cover slide, is solving a fundamentally different problem at a fundamentally different layer of capability. Calling all twelve voice AI for restaurants is like calling a moped, a delivery van, a long-haul truck, and a tank vehicles. The word is technically correct. It is also analytically useless.
The contrarian read, and the one I have been wanting to write for a month, is this: the category bifurcation matters more than the funding rounds. PolyAI’s $86 million Series D in December, Slang’s $36 million Series B last month, Presto’s $10 million in January, Vox AI’s $8.7 million seed last August — those are real cheques and they are signaling real demand. But the cheques do not tell you what an operator is actually buying. The cheques sort by who raised, not what stage of capability the product reflects. And the operator buying decision in 2026 is increasingly going to be governed by a stage question, not a vendor question. That is the argument.
A note on framing before the frame. This piece is the first draft of a maturity curve I expect to refine over the next several weeks. The May refinement, forthcoming in this column, will tighten the boundaries between stages and add evidence from the spring earnings cycle. The same way the January Four Margins frame needed a February rewrite once the print cycle landed, this March frame is the working version, not the museum piece. Read it as a tool to put against your own desk, not as a finished taxonomy. If the boundaries shift in May, the shift will be visible in this column.
Why a curve, not a category
The reason the voice-AI conversation has been confused for two years is that most operator-facing comparisons treat the category as a bag of vendors, sorted by funding, logo prominence, or whichever brand happens to be in the most recent press cycle. That is a useful sort for a marketing audit. It is the wrong sort for a capability audit. If you ask which voice-AI product should I deploy, the answer depends almost entirely on which stage of the capability curve you are buying — and most of the operators I have talked to this winter have been buying a different stage than the one their procurement deck implies.
The five-stage curve below is my working frame. It is not exhaustive, the boundaries between stages are porous, and a single vendor can plausibly sit at two adjacent stages depending on which feature you are looking at. But the shape is real. The capability gap between stage two and stage three is wider than the gap between stage one and stage two, and the capability gap between stage four and stage five is wider still. Operators procuring at stage two and expecting stage three are why so many voice AI deployments have under-performed against the deck. The stage mismatch is the failure mode. Naming the stages is the cure.
Five stages. Stage one: IVR replacement. Stage two: menu memorization. Stage three: recommendation engine. Stage four: loyalty-aware concierge. Stage five: agentic operator. The rest of this essay walks each stage as an H2 — definition, capability test, vendors who live there, operator deployments that prove it, and what the transition out of the stage actually looks like in the wild.
Stage one: IVR replacement
Stage one is the floor of the category and the place every voice-AI deployment starts. The product replaces an interactive voice response menu — press one for reservations, press two for hours, press three to speak to a manager — with a synthetic voice that can hold a brief, scripted, intent-classified conversation. The capability test is single-intent recognition with a finite slot fill. Can the system tell that the caller wants to make a reservation, capture the party size and time, and confirm or escalate to a human? If the answer is yes, you are at stage one. If the answer is no, you are not yet at stage one, and you should not be in market.
The interpretation that matters is this: stage one is not interesting on its own merits, but it is the necessary floor. Every voice agent in production today passes the stage-one capability test or it would not have made it past the first week of a pilot. Stage one is table stakes, in the literal sense — you cannot host the conversation at any higher stage without passing it. The vendors who are only at stage one in 2026 are mostly the regional and white-label IVR shops who repackaged a hosted speech-to-text and intent-classification stack with the word AI added to the brochure. They are not the venture-funded names. They are the line items on the operator’s existing telephony invoice with the suffix .ai appended.
The operator proof points at stage one are everywhere and they are not the story. Every QSR with a missed-call deflection workflow, every full-service reservation line that routes to a synthetic host after the third ring, every drive-thru that funnels overflow to a hold queue with a brief intent capture — those are stage-one deployments, and they are functionally invisible to the operator’s P&L. They reduce a small percentage of call-handling cost. They do not move the labor or the throughput line meaningfully. The transition out of stage one happens when the operator realises that answering the call is not the same problem as running the call. Once that realisation lands, the operator is shopping at stage two.
Stage two: menu memorization
Stage two is where most of the venture-funded voice-AI category currently lives, and it is the stage that operators most consistently mistake for something further up the curve. The capability test is full menu and policy comprehension. Can the system answer arbitrary questions about every item on the menu, handle modifier combinations, enforce upsell rules, route allergen questions correctly, and complete a multi-item order without dropping the cart? Stage two is the can-actually-take-the-order stage. Stage one was can-route-the-call.
The vendor concentration at stage two is dense. PolyAI, SoundHound, Vox AI, Slang, Presto, ConverseNow, the various infrastructure layers (Retell, Vapi, Rime), and the newer entrants — Newo, Hostie, Loman — are all, at the substance level, competing on the quality of their stage-two capability. The marketing language differs. The category language differs. The pricing pages diverge. But the actual deliverable, the thing the operator is paying for, is the ability to take the order accurately, in the operator’s voice, on the operator’s menu, at the operator’s volume. That is stage two.
The operator deployments that prove stage two are the ones the popular press has been treating as the headline story for two years. Wendy’s FreshAI is the canonical case. The publicly disclosed accuracy numbers — 86% of orders completed without employee intervention, 99% order accuracy on the orders the agent completed — are stage-two performance numbers. They are the cleanest public reporting any operator has done on a voice-AI deployment, and the operator wisdom inside the disclosure is the recognition that order accuracy is the deliverable. Domino’s voice ordering, the long-running Wingstop AI pilots, the McDonald’s drive-thru program before its 2024 pause and after its 2025 restart with the new infrastructure layer — all stage-two. Each of those deployments is measured, correctly, on order completion and order accuracy at peak. That is the right measurement for stage two. It is also, and this is the key, the ceiling of stage two’s value contribution.
The transition out of stage two is the transition that almost no vendor has actually made and that the entire category is mispricing. It is the move from the system knows the menu to the system knows the guest, the moment, and the upside. That move is stage three.
Stage three: recommendation engine
Stage three is where the curve starts to bend, and it is the stage I think is most mispriced today. The capability test is contextual recommendation. Given the caller’s history, the time of day, the weather, the inventory state, the LTO calendar, and the conversation so far, can the system make a recommendation that increases the check size or improves the guest experience in a way the menu-only stage-two agent cannot? Stage three is where the voice agent stops being a faster IVR and starts being a host with judgement.
The vendors who can credibly claim stage three are a much smaller list than the stage-two cohort, and even among that smaller list the claim is mostly aspirational rather than deployed. PolyAI’s recent product disclosures imply stage-three capability is on the roadmap. SoundHound’s deeper QSR integrations — the ones that read POS state in real time and reason about availability and upsell — are stage-three in some configurations. Vox AI’s autonomous voice AI language is explicitly stage-three positioning — the 90-plus-languages line is the headline, but the substance claim is the system reasoning about the order in real context. Slang’s Series B materials gesture at stage-three guest intent, though the production deployments I have seen are mostly stage-two with stage-three on the disclosure roadmap. The honest reading: stage three exists as a deployed capability in narrow configurations at three or four vendors, and as a marketing claim at most of the rest.
The operator deployments that prove stage three are harder to point at because the operators who have actually achieved it are also the ones least likely to describe it as a voice AI deployment. They describe it as our menu engineering tool, or our personalized loyalty surface, or our drive-thru optimizer. The voice agent is one channel of the underlying capability. The capability itself is the recommendation engine that reads the guest, the moment, and the operator state and produces a better next-best-action than the human on the headset would have produced under the same conditions. That is a different product. It is, more honestly, a different company — the substance work is the recommendation model and the data pipelines that feed it, not the speech surface.
The transition out of stage three is the transition into loyalty awareness. The recommendation engine that reasons about the menu and the moment becomes the concierge that reasons about the guest across visits. That is stage four.
Stage four: loyalty-aware concierge
Stage four is where the voice agent stops being a transaction layer and becomes a relationship layer. The capability test is multi-visit memory and preference modelling. Does the agent know that this caller usually orders the spicy chicken sandwich without mayo, has a standing Friday-night reservation, mentioned a peanut allergy on the August call, and joined the loyalty program in November? Can the agent reason about all four facts in the same conversation, in a way the human host has neither the time nor the memory to do? Stage four is the stage at which the voice agent starts to do work a good host with a notebook used to do — and at scale.
The vendor concentration at stage four is currently near-zero in deployed production. The roadmap claims are everywhere; the audited deployments are few. The two surfaces I would point at as proof of the direction — not proof of the destination — are the chain-level concierges the major hotel brands have shipped this year. Hilton’s AI Planner, launched on the Hilton Honors app in late January and examined in this column in February, is the cleanest proof point because Hilton has been explicit about the underlying loyalty integration. The planner reads Honors data, prior stays, and stated preferences, and renders a recommendation that the brand believes a Diamond-tier guest will recognize as preference-aware. That is stage-four positioning, on a chat surface rather than a voice surface, at a hotel brand rather than a restaurant. Marriott’s RENAI is the equivalent on the Marriott side, with less public disclosure. Neither is a voice product. Both are stage-four capability deployed in the channel where stage-four capability happens to be most procurement-ready inside the brand today.
The interpretation is this: stage four is happening in hospitality, but it is happening in the chat and app surfaces first, not in voice. The voice version of stage four — the concierge that picks up the phone, recognises the caller, remembers them across visits, and reasons about preferences in real time — is the product no vendor has shipped in restaurant volume at the time of this writing. The vendor who ships it first will hold pricing power for as long as the rest of the cohort takes to catch up. That window is the bet I will come back to in the closing section.
A clarification on one named deployment that has been mis-described in adjacent press, because the conflation is doing damage to the category’s vocabulary. Chipotle’s Ava Cado is not a customer voice agent. Ava Cado is the company’s HR-side recruiting and onboarding assistant — internal hiring workflow tooling, not guest-facing ordering. Operators who have seen the Chipotle name in voice-AI category roundups should treat the inclusion as a category error. The substance product is HR automation, the surface is internal-facing, and the relevant comparison set is workforce-tech, not voice. I flag it here because I have heard the conflation three times this week and the maturity curve should not import the confusion.
The transition out of stage four is the transition into action. The agent that knows the guest becomes the agent that acts on the operator’s behalf. That is stage five.
Stage five: agentic operator
Stage five is the destination of the curve and the stage at which the language voice AI becomes meaningfully insufficient. The capability test is autonomous operational decision-making. Can the agent decide that the kitchen is over capacity and stop accepting new orders for the next eight minutes, on its own authority? Can the agent recognize a fraud pattern in the call queue and route the next three calls to a human supervisor? Can the agent re-price an under-performing LTO at the moment of the call, within the bounds the operator has pre-authorized? Stage five is the agent that decides, not the agent that answers.
There is, at the time of this writing, no shipped stage-five voice deployment in restaurant volume that I would defend in public. The closest published gestures are research-level disclosures from the infrastructure cohort — Retell, Vapi, Rime — about agent autonomy primitives that could compose into stage-five capability if a thick application layer were built on top. None of those infrastructure disclosures is a shipped operator deployment. They are toolkits, not products. The application layer that would turn them into stage-five voice agents is the gap. The operators with the most mature data infrastructure — the ones who could, on paper, deploy stage-five capability today — are mostly choosing not to, because the regulatory environment for autonomous AI decisioning is still consolidating and the brand risk of an agentic mis-decision is high relative to the operating-margin gain.
The interpretation is this: stage five is the destination, the timeline to broad deployment is longer than the venture cycle is currently pricing, and the operators who get there first will get there by building the capability in-house rather than buying it from a vendor. The vendor business at stage five looks more like enterprise AI platform than voice AI for restaurants. The category language will have to change to accommodate it. By the May refinement of this frame I expect to have at least one named operator-side stage-five deployment to point at; today I do not.
The pricing dynamic: stage three is mispriced
The argument that follows from the curve is a pricing argument, and it is the part of this essay that I expect to age the most. Here it is, plainly: stage-three capability is currently mispriced low, and the 24-month window before pricing power evaporates is the most actionable arbitrage operators have available in the voice-AI category today.
Why stage three, specifically. Stage one is commoditised — every IVR-replacement product is approximately interchangeable, the pricing has converged on per-call-minute economics, and the vendor pricing power is near zero. Stage two is the densest competitive field — the dozen named vendors are competing on incremental quality at roughly the same price point, the operator can play vendors against each other on RFP, and the vendor pricing power is bounded by the next quote. Stage four and stage five are not yet shipped in restaurant volume, so the pricing has not yet been tested in the market. Stage three is the stage where the capability is meaningful enough to move the operator P&L, the vendor count is small enough to allow pricing power, and the market understanding is still primitive enough that operators are paying stage-two prices for stage-three capability when they can find it. That gap is the arbitrage.
The operator move, if I were running a restaurant group of more than fifty units today, would be the following. I would identify the two or three vendors in market who can credibly demonstrate stage-three capability against my menu, my POS state, and my LTO calendar — a demonstration in production, not in a sandbox. I would negotiate a multi-year commercial deal at current pricing, with a price protection clause that holds the vendor to stage-two-adjacent economics for the term of the deal. I would expect, by the second year of the deal, that the same vendor’s market pricing for the same capability has tripled, because by then the operator cohort will have figured out what stage three is worth. That is the arbitrage. The window closes when the category vocabulary catches up to the capability stack. My read is that it closes in 24 months.
The companion observation is that the guest-data margin lever is most pronounced at stage three and above. A stage-two agent collects transaction data. A stage-three agent collects intent data, preference data, and contextual decision data — the same dataset class the platform acquirers (AmEx-Resy-Tock, DoorDash-SevenRooms) are paying multi-billion-dollar valuations to assemble. The operator who deploys stage-three voice without negotiating the data-ownership terms is contributing rows to the vendor’s asset for free. The contract should specify, in writing, who owns the structured transcripts, the preference signals, and the recommendation telemetry. Vendors who refuse the clause are signaling which margin they are actually building toward. Operators who accept the refusal are paying twice.
The regulatory tailwind: the Council position bought operators a year
The other reason to write this essay this week, rather than three weeks ago or three weeks forward, is that the regulatory clock for voice-AI deployment in EU markets just moved — and the move was net favourable for operators in a way the press cycle has not fully metabolised.
On March 13, the European Council adopted a position on the EU AI Act implementation timeline that effectively extends compliance deadlines for the obligations most relevant to voice agents — Article 50 transparency, the synthetic-content marking rules, the high-risk-system audit requirements adjacent to them. The full mechanics are worked out in my colleague Juliet’s coverage from earlier this month and in the column’s running EU AI Act explainer. The substance, for the maturity-curve argument, is this: the compliance cost line that EU operators were budgeting against an August 2 hard date has now been deferred, and the deferral is meaningful enough to change the timing of stage-three voice deployments in EU markets.
The interpretation is that the deferral is the most underrated tailwind of the year so far for voice-AI procurement. Operators who were waiting for compliance clarity before deploying stage-three capability in EU markets can now move forward with a longer runway. Vendors who were budgeting their EU go-to-market against an August deadline now have additional quarters to invest in capability rather than compliance. The cost shift is asymmetric — it falls on vendors more than on operators, because the operators were going to ask the vendors to bear the compliance burden anyway — and the asymmetry is what makes the deferral a real procurement input rather than a press-cycle artefact.
The risk on the other side is that operators interpret the deferral as a delay, not a deferral, and slow their stage-three procurement on the assumption that the pricing window will also slide. That is the wrong read. The pricing window is governed by the vendor cohort’s ability to defend stage-three capability against competitive entry, and the EU compliance timeline is not the binding constraint on that defense. The binding constraint is the vendor count at stage three, which I expect to grow substantially in 2026 regardless of what Brussels does. The operator who waits is the operator who pays the higher price in 2027. The Council position bought time on compliance, not time on procurement.
The bet: a 24-month window
The bet, in one paragraph. Most named vendors in the voice-AI-for-hospitality category are stuck at stage two. The capability gap to stage three is wider than the marketing language admits and narrower than the vendor cohort thinks. The operator who buys stage-three capability today, on a multi-year commercial structure with explicit guest-data terms, will hold a meaningful margin advantage over the operator who waits — and the size of the advantage is bounded above by how quickly the category catches up to the vocabulary. My read is 24 months. By the May refinement of this frame I will revise the number if the spring earnings cycle gives me reason to. By the autumn I expect to be writing a different essay entirely — one about which operators successfully captured the stage-three position and which ones spent 2026 paying for stage-two capability dressed up in stage-three pricing decks.
The category vocabulary, finally. Voice AI for restaurants is too broad to be useful and conversational AI is too vague to be actionable. Operators evaluating vendors in 2026 should ask the stage question first. Which stage is the vendor’s product actually at, in production, against my menu and my data? Which stage is the next-quarter roadmap promising? Which stage is the procurement contract pricing? If the three answers do not agree, the deal is mispriced — sometimes in the operator’s favor, frequently in the vendor’s. Stage discipline is the cure.
This is the working frame for the spring. The May refinement will tighten the boundaries and add evidence from the Q1 earnings cycle. The framework holds. The labels may shift. Mise is a working draft, not a finished document, and the draft is what is on the desk this week.
— Eitan is editor-in-chief of TableTransfers. Tips: [email protected].
The Voice Agent Maturity Curve
mise
·12 min read
The Four Margins of a Restaurant
mise
·14 min read
The AI Premium in Hospitality M&A: Broker Story or Real Number?
the bottom line
·9 min read
What the DoorDash/SevenRooms Deal Actually Buys
the bottom line
·11 min read
Related posts
mise
·20 min read
The Voice Agent Maturity Curve
mise
·26 min read
The Four Tables: Why the reservation system is the most contested square foot in hospitality
mise
·22 min read