The Voice-Agent Maturity Curve: A Framework for Operators
A four-stage framework (Pilot → Drive-Thru → Phone-Channel → CRM) for evaluating voice AI deployments, tied to the January 2025 fundraise wave and Wendy's deployment data. The Mise framework essay for the month.
It is the last Friday of January, and I am sitting at my desk with three browser tabs open, a printed deck from a PR agency, and a cup of coffee that has gone cold twice. The first tab is the Bland announcement from Wednesday — forty million dollars, voice agents for enterprise, the press release written in the kind of sentence-fragments that try to sound like a manifesto. The second tab is the ElevenLabs Series C, posted yesterday — one hundred and eighty million dollars at a 3.3 billion dollar valuation. TechCrunch did the math out loud in the headline body: “37 times ARR.” The third tab is a Wendy’s corporate blog post about something they call FreshAI, with a single number that has been quoted at me by four different vendors this month: 86 percent pilot accuracy.
I have been an editor long enough to know what this combination of tabs means. It means the market has decided voice AI for restaurants is a category, and it has decided this without agreeing on what voice AI for restaurants actually is. The Bland deck talks about phone agents. The ElevenLabs round talks about a voice model platform that powers, among other things, phone agents. Wendy’s talks about a drive-thru system. PolyAI, in a press release from last fall I have pinned to my noticeboard, talks about reservations and missed calls — “between 30 and 60 percent of phone calls,” they wrote, citing their own customer data. Each of these is being called “voice AI.” Each of these is being priced into the same wave of capital. And each of these is, in operator terms, a completely different product, sold to a completely different buyer, solving a completely different problem.
This is the first Mise essay of 2025, and I want to use it to set down a framework that I expect to refer back to all year. The thesis is simple. Voice AI for restaurants is not one product on one curve. It is four products on one curve, and where any given vendor sits on that curve tells you almost everything about whether their valuation makes sense, whether their pilot data should be believed, and whether the operator across the table from you is buying the thing they think they are buying.
I am going to call this the Voice-Agent Maturity Curve. The stages, in order, are Pilot, Drive-Thru, Phone-Channel, and CRM. The rest of this essay is about what each stage means, what the January fundraise wave priced into them, and how an operator should use the curve when a vendor walks in the door.
The framework
The four stages are not technology generations. They are deployment depths. A vendor can be technically capable of stage four and still only have stage one deployments in the wild. An operator can sign a stage four contract and still be running a stage one pilot eighteen months later. The point of the curve is to separate what is being sold from what is being run, and to make the distance between those two things visible.
Here are the four stages, stated plainly.
Stage 1: Pilot. A single location, or a small number of locations, running a voice agent in a controlled environment. Success is measured in accuracy on a known menu, hours of uptime, and the absence of catastrophic failures. The agent is not yet load-bearing. A human can take over at any time, and frequently does. The operator is buying a measurement, not a service.
Stage 2: Drive-Thru. The voice agent is the primary order-taker at a real drive-thru lane, during real business hours, on a real menu, with real customers who did not consent to being part of a pilot. Success is measured in order accuracy at scale, average handle time, upsell capture, and the percentage of orders that escalate to a human. The agent is load-bearing for the lane. If it falls over, the lane falls over.
Stage 3: Phone-Channel. The voice agent owns the inbound phone line for an entire restaurant or chain — reservations, takeout, modifications, basic FAQs, callbacks. Success is measured in the percentage of calls that get answered (the baseline being the 30 to 60 percent that PolyAI’s customer data says go unanswered today), the percentage that complete without a human, and the conversion rate on the calls that matter (reservations booked, orders placed). The agent is load-bearing for the channel. The phone is the agent’s domain, and a human is the exception.
Stage 4: CRM. The voice agent is no longer a channel. It is a layer. It knows the guest from prior visits, prior calls, prior reservations. It can recognize a returning caller, pull their preferences, route them to the right action, and write the outcome of the conversation back to the customer record. The phone call becomes a structured event in the guest’s history, not a transient interaction. Success is measured in guest-level outcomes — repeat visit lift, lifetime value, recovery on complaint calls — not in call-level metrics. This stage is, as of the last Friday of January 2025, vendor whitespace. Nobody is shipping it at scale. Several are pricing as if they will.
That is the framework. Pilot, Drive-Thru, Phone-Channel, CRM. The rest of this essay walks each stage with an example, then turns to the question of what the January fundraise wave priced in, and what an operator should do about it.
Stage 1: Pilot
The pilot stage is where almost every voice AI deployment in the restaurant industry currently lives. This is not a criticism. It is a description. Pilots are how operators learn whether a vendor’s claims survive contact with their menu, their staff, their network, and their guests. A pilot that runs for six months and produces a clean accuracy number is a successful pilot. A pilot that runs for six weeks and gets quietly shelved is also a successful pilot, because the operator learned something cheaply.
The defining feature of the pilot stage is that the voice agent is not load-bearing. There is always a human on the other side of the curtain. The drive-thru attendant is wearing a headset. The phone has a fallback to a hostess. The kitchen knows that if the order ticket looks wrong, they should call back. The agent is being measured, not relied upon.
The trap of the pilot stage is that the metrics it produces — accuracy on a constrained menu, response latency in a quiet test environment, customer satisfaction in a self-selected pilot cohort — are not predictive of stage two performance. A pilot can run at 95 percent accuracy on a 40-item menu and collapse to 70 percent when the LTO drops and the menu doubles in a week. A pilot can run at sub-second latency on a wired ethernet connection and stutter at three seconds on a franchise’s actual 4G failover. A pilot can have a 4.7-star CSAT in a Tuesday lunch test window and a 2.9-star CSAT on a Friday night with a line of cars and a tired manager.
The pilot stage is not where you evaluate a vendor’s technology. It is where you evaluate a vendor’s instrumentation. What does their pilot dashboard show you? Can you see escalation reasons? Can you see the audio of failed calls? Can you see which menu items get misheard most often, and how that distribution shifts week over week? A vendor whose pilot tooling is good is a vendor who has run a lot of pilots and learned what operators need to see. A vendor whose pilot tooling is a CSV emailed weekly has not.
Most of the voice AI vendors raising money in January 2025 are being valued on stage one data. This is the first place the gap between pricing and deployment hides.
Stage 2: Drive-Thru
The drive-thru stage has, at present, one publicly visible reference deployment that operators across the industry are using as their benchmark. It is Wendy’s FreshAI, and the number that travels with it is 86 percent. That is the accuracy figure Wendy’s published on its corporate blog (https://www.wendys.com/blog/wendysr-square-deal-blog/transforming-ordering-experience-wendys-freshai-update) describing the pilot phase of its drive-thru voice agent. Every vendor I have spoken to this month has either quoted this number, claimed to beat it, or argued about how it is calculated.
I want to be careful here, because Wendy’s number is a pilot-phase number — which means by the framework I am laying out, it is technically stage one data. But the deployment context is unambiguously stage two: real lanes, real customers, real menu, real business hours, at a publicly traded national chain. The agent is load-bearing for the lanes it is in. Wendy’s has been transparent about the operational realities — escalation paths, menu coverage, ongoing tuning — in a way that most of their peers have not.
The reason the drive-thru stage matters as its own category is that the failure modes are unique to the lane. Wind noise. Engine idle. A customer leaning out of a passenger window. An order modification shouted from the back seat by a child. A menu with limited-time offers that change every six weeks. A staff that needs to be able to override the agent without breaking the order flow. None of these failure modes exist in a phone deployment, and none of them are well-modeled by pilot test environments.
The drive-thru stage is also where the economics of voice AI become interesting for the first time. A drive-thru lane handles, on a busy day, several hundred orders. If the voice agent reduces average handle time by ten seconds and improves upsell capture by even a small percentage, the math at chain scale becomes material. This is the stage where vendors can charge per-lane subscription fees that justify themselves on labor and throughput, and it is the stage that the largest QSR brands are most actively piloting. McDonald’s, Carl’s Jr., Hardee’s, Checkers, Del Taco, and others have all run drive-thru voice AI pilots of various depths over the past three years. Most have ended or paused. Wendy’s is the most visible ongoing one.
The drive-thru stage is the stage at which voice AI has the most pressure to deliver real numbers, because the deployment is observable. A customer in line can hear the agent. A franchisee can see the order accuracy report. A regional manager can watch the lane time. There is nowhere for a bad deployment to hide.
Stage 3: Phone-Channel
The phone-channel stage is, on paper, the simplest of the four. The voice agent answers the phone. It books reservations, takes takeout orders, answers questions about hours and location, handles modifications and cancellations, and routes the calls it cannot handle to a human. The deployment surface is small — one phone line per restaurant — and the integrations are well-defined: a reservations system, a POS, a customer database.
The reason it matters is the baseline. PolyAI’s customer data, cited in their press release with OpenTable last fall (https://www.prnewswire.com/news-releases/polyai-partners-with-opentable-to-offer-enterprise-restaurants-and-diners-reservation-support-using-voice-ai-302254773.html), is that restaurants miss “between 30 and 60 percent of phone calls.” Take a moment with that range. If you operate a restaurant that does any meaningful reservation or takeout volume, somewhere between three and six out of every ten people who try to reach you on the phone do not. They hang up. They call a competitor. They book somewhere else. They give up.
A phone-channel voice agent that simply answers those calls — without booking a single additional reservation, without taking a single additional takeout order, without resolving a single complaint — recovers real revenue. This is the cleanest ROI story in voice AI, and it is the one PolyAI has been most disciplined about telling. Their Series D, raised last year, was priced against this story. The CMSWire writeup of that round (https://www.cmswire.com/customer-experience/polyai-raises-86m-series-d-for-enterprise-voice-ai/) frames it as enterprise voice AI for customer service, and the restaurant case is a clean application of the enterprise pattern: replace the missed call.
The phone-channel stage is where the operator-vendor interface gets complicated, though, because the integrations matter more than the voice quality. A phone agent that sounds great and cannot push a reservation into OpenTable is a worse product than a phone agent that sounds mediocre and writes correctly to the reservation book. A phone agent that takes a takeout order and cannot fire it to the POS is a phone agent that has created a manual data-entry job for the host. The voice is the surface. The integrations are the product.
This is the stage at which the January fundraise wave gets most interesting, because ElevenLabs, the company that just raised at 37 times ARR, is a voice infrastructure company. They make the voice. They do not, themselves, make the phone-channel product. The phone-channel product is built by application-layer companies — PolyAI, SoundHound, Bland, Replicant, and a long tail of restaurant-specific specialists — that may or may not use ElevenLabs voices underneath. The infrastructure layer and the application layer are being valued together in this wave, and that is a category error the curve helps clarify.
Stage 4: CRM
The CRM stage is, today, vendor whitespace. I have not seen a single voice AI deployment in the restaurant industry that operates at this level, and I have looked. By “this level” I mean: the voice agent recognizes the returning guest, pulls their preferences and history, makes a contextualized decision about what to offer or how to route, and writes the outcome of the conversation back to the customer record in a way that influences future interactions across channels.
What makes this stage hard is not the voice. The voice is the easy part. What makes it hard is the customer record. Most restaurants do not have a unified customer record. They have a reservation system that knows reservation history, a POS that knows order history, a loyalty program that knows visit frequency, an email marketing tool that knows engagement, and a review platform that knows complaints. These systems do not talk to each other in any structured way. The hostess at a stage three restaurant cannot pull up a guest’s last three visits, last complaint, dietary restriction, and preferred server in a single screen. Neither, today, can the voice agent.
A stage four voice agent presumes a guest data layer that, for the vast majority of restaurants, does not exist. Building that layer is the actual work. The voice agent is the affordance on top.
I am dwelling on this because the January fundraise wave is pricing into stage four. The valuations being assigned to voice AI vendors right now do not make sense on stage one or stage two deployment data. They start to make sense on stage three deployment scale, if you believe phone-channel will roll out fast across the enterprise mid-market. They make sense on stage four, if you believe the voice agent eventually becomes the entry point to the customer record. The valuations are pricing in a future where the voice agent is not a channel but a customer relationship layer. The deployments, today, are pricing in a present where the voice agent is a pilot.
That gap is the entire investment thesis of the category. It is also the entire risk.
What the January fundraise wave priced in
Let me put the numbers down in one place. Bland AI closed forty million dollars on Wednesday this week. ElevenLabs closed one hundred and eighty million dollars yesterday at a 3.3 billion dollar valuation, which TechCrunch (https://techcrunch.com/2025/01/30/elevenlabs-raises-180-million-in-series-c-funding-at-3-3-billion-valuation/) noted is “37 times ARR.” CB Insights, in their market research note on voice AI (https://www.cbinsights.com/research/voice-ai-market-opportunities), reports that voice AI equity funding hit “$2.1B in 2024” with “nearly $500M” raised in Q1 of 2025 alone — and we are still a day away from February.
These numbers price the curve at the top end. A 37x ARR multiple is not a multiple you assign to a company at stage one or stage two. It is a multiple you assign to a company you believe will own the substrate of a category that is itself going to grow many times over. The investors writing these checks are not betting on Wendy’s FreshAI hitting 90 percent accuracy. They are betting on every restaurant phone line in America being answered by an AI in five years, and on the voice infrastructure underneath being concentrated in two or three platforms.
That bet may be correct. I am not in the business of disputing it. I am in the business of telling operators how to read the bet, and the curve is the reading guide.
If you are an operator, here is what the January fundraise wave actually tells you. It tells you that voice infrastructure is going to keep getting cheaper and better, fast. It tells you that application-layer vendors are going to have well-funded competition for the next 24 to 36 months, which means feature velocity will be high and pricing will be aggressive. It tells you that the vendors selling stage three (phone-channel) products today will be under enormous pressure to claim stage four (CRM) capabilities tomorrow, whether or not the underlying customer data layer supports those claims. It tells you that the most dangerous purchase decision an operator can make right now is to buy stage four pricing on a stage two deployment.
Where the gap between pricing and deployment hides the risk
I want to spend a section on this, because it is the part I most want operators to internalize.
The risk in voice AI procurement is not that the technology does not work. It does, well enough, at stage two and stage three. The risk is that the contracts being signed today are stage four contracts — multi-year, per-location, with platform-level lock-in — for deployments that are operationally stage one or stage two. The vendor’s incentive is to sell the contract on the future-state pitch. The operator’s incentive is to test the present-state product. The two incentives only align if the contract is structured to follow the deployment up the curve, not to assume it.
A stage-aware contract has several features. It prices based on the stage in deployment, not the stage in the pitch. It includes explicit success criteria for moving from one stage to the next — accuracy thresholds, escalation rate ceilings, integration completeness, customer record write-back. It includes an exit ramp at every stage. It does not assume that a successful pilot means an inevitable rollout. It treats each stage as a separate procurement decision with its own evidence requirements.
Most of the contracts I have seen passed under the table at industry events this month are not stage-aware. They are platform contracts written at stage four pricing for operators who have not yet completed a stage two deployment. The vendor’s logic is that the platform contract de-risks the rollout. The operator’s reality is that the platform contract front-loads the cost of a rollout that may or may not happen, on a curve the operator has not yet decided to climb.
The principle here is simple. Buy the stage you are in. Reserve the right to buy the next stage when you get there. Do not let the fundraise wave convince you that stage four pricing is a discount on stage four delivery, because stage four delivery does not yet exist.
How operators should pick a vendor against the curve
Here is the framework operationalized. When a voice AI vendor walks into your office, ask four questions in order.
First, what stage are your reference deployments at? Not your roadmap. Your deployments. If the vendor cannot name a stage three deployment in production at a comparable operator, they are selling you stage one or stage two. That is fine — but you should be paying stage one or stage two pricing, and signing a stage one or stage two contract. The vendor’s answer to this question is the most important data point in the meeting. Vendors who answer it cleanly are vendors who understand their own deployment depth. Vendors who answer it vaguely are vendors who are hoping you will not.
Second, what stage are you being priced for? A vendor who is asking for a multi-year platform contract with seven-figure annual commitments is pricing you at stage three or stage four. A vendor whose deployments are at stage one or stage two needs to defend the gap between their deployment stage and their pricing stage. That defense is usually some version of “this is the future.” That may be true. But the operator is paying today, and the operator is bearing the integration cost today, and the operator deserves to know what the vendor’s pricing assumes about delivery they have not yet shipped.
Third, what is the escalation rate, and what is the cost of an escalation? At stage two, every escalation is a labor event. At stage three, every escalation is a labor event plus a customer-experience event plus a data event. Vendors who can tell you their escalation rate at deployment, and who can tell you the cost structure of escalations in their customer base, are vendors who understand the economics of the curve. Vendors who quote you only their headline accuracy number — the 86 percent, or whatever number they have chosen to put on the deck — are quoting the easy metric, not the operational one.
Fourth, what does your customer record integration look like? This is the stage four question, and it is the question that most cleanly separates vendors who are pricing into the future from vendors who can deliver it. Ask which CRMs, reservation systems, POS systems, and loyalty platforms they have shipped production integrations for. Ask whether the integration is read-only or write-back. Ask whether the integration is real-time or batched. The vendor who is closest to stage four has the most concrete answer to this question. The vendor who is selling stage four on a stage two deployment will pivot to roadmap language. That pivot is the answer.
If a vendor can answer these four questions in order, with specifics, they have earned the next conversation. If they cannot, they have not. This is the principle the framework is built to enforce.
What 2025 will and won’t resolve
I want to close with what I expect this year to settle and what I expect it not to.
I expect 2025 to settle the question of whether voice AI can hold a stage two drive-thru deployment at scale. Wendy’s will publish more data. McDonald’s will say more about its plans. At least two of the major chains will either commit publicly to a rollout or quietly retreat, and either of those outcomes is information. The 86 percent number will either become a floor or become a footnote.
I expect 2025 to settle, partially, the question of phone-channel adoption at the enterprise mid-market. PolyAI, SoundHound, Bland, and a handful of others will publish customer counts, and we will get a clearer picture of whether the missed-call recovery thesis converts at the per-location economics the vendors are pitching. The 30-to-60-percent missed-call number will either be confirmed as the baseline operators are willing to spend against, or it will turn out that operators have rationalized those missed calls in ways the vendor pitches do not account for.
I do not expect 2025 to resolve stage four. The customer data layer at most restaurants is too far from ready, and the vendor incentives are too pointed away from the unglamorous work of building it. The vendors who are quietly investing in the data layer this year will be the vendors I am writing about in 2027. The vendors who are pricing as if the data layer already exists will, by then, either have acquired the companies that built it or have quietly repriced.
I do not expect the regulatory environment to resolve, either, but it will move. The EU AI Act’s general-purpose provisions take effect in two days, on Sunday. That is not directly a restaurant voice AI event — the high-risk classifications most relevant to AI in food service have longer phase-ins — but it is a signal about where the regulatory conversation is going to go, and operators with European exposure should expect their vendors to have answers about audit trails, voice cloning consent, and data residency that those vendors did not have to provide six months ago. I will write about this separately when the implications for our category get clearer.
The framework will travel. That is what frameworks are for. I expect to extend it through the year in our subsequent voice-agent coverage, and to refine it further in the essays we develop later in the year toward the broader margin work this column is building. The Voice-Agent Maturity Curve is the first piece of doctrine I am putting on the board for 2025. The rest of the year’s Mise essays will refer back to it, build on it, and where the evidence demands, revise it.
What I’m holding, at the end of this last day of January, is a curve with four stages, a fundraise wave that has priced the top end, and a deployment landscape that mostly sits at the bottom two. The gap between those is where every interesting decision an operator makes this year will happen. It is also where every interesting story I write this year will live.
Buy the stage you are in. Reserve the right to buy the next stage when you get there. Read the curve before you sign the contract.
— Eitan writes the Mise column. Tips: [email protected].
The Voice Agent Maturity Curve
mise
·12 min read
The Four Margins of a Restaurant
mise
·14 min read
The AI Premium in Hospitality M&A: Broker Story or Real Number?
the bottom line
·9 min read
What the DoorDash/SevenRooms Deal Actually Buys
the bottom line
·11 min read
Related posts
mise
·20 min read
The Voice Agent Maturity Curve
mise
·26 min read
The Four Tables: Why the reservation system is the most contested square foot in hospitality
mise
·22 min read