The 2025 Voice-AI Buyer's Guide for Restaurant Operators

QSR drive-thru speaker post and digital menu board with a customer car arriving in the lane during daylight.

If you're choosing a drive-thru or phone voice-AI vendor at the NRA Show next week, the right question isn't 'accuracy' — it's 'who owns the upsell rev share?' Here is the field.

My desk this morning looks like a trade-show booth crashed into it. Seven vendor pitch decks, four of them mailed in literal envelopes with literal lanyards taped to the cover sheet, three more sitting in a Dropbox folder labeled “NRA 2025 — Voice AI.” A printed map of McCormick Place. Two coffees. A sticky note from our editor that just says “rubric, please.”

The NRA Show opens in Chicago on May 17, ten days from now, and every voice-AI vendor in the QSR ecosystem is about to spend a week telling operators that their drive-thru is broken and only this particular speaker box can fix it. I have spent the last three weeks on calls with sales teams, product leads, and — more usefully — operators who already signed contracts and are now living with them. This is the field guide I wish I’d had when I started.

Vibe Check verdict, up top, because I owe you that before three thousand words of nuance: SoundHound, Hi Auto, Presto, and Kea each have a clean operator value proposition, but only two of the four are pricing in a way operators should accept. ConverseNow and the Yum × NVIDIA NIM-based stack are the two stories that reset the buy-vs-build math for chains over 200 units, and they reset it in opposite directions. If you walk McCormick Place asking “what’s your accuracy number?” you’ll get four hours of theater and a contract you’ll regret. The question that actually matters is: who owns the upsell rev share, and on what base? Everything else flows from there.

The rubric

Four dimensions. Each scored 1–5, where 3 is “fine, ship it” and 5 is “operator-defensible at scale.”

Build quality — does the system actually work in a real lane, with real ambient noise, real menu modifiers, and a real handoff to a human when it falls over? This is the only place an accuracy claim lives, and even there I weight handoff behavior heavier than the raw number.

Operator impact — labor minutes recovered per shift, ticket lift, throughput change at the order point. Measured against the operator’s prior baseline, not the vendor’s case-study chain.

Pricing transparency — flat monthly per lane, per-order, percentage of upsell, or hybrid? Is the floor visible before signing? Can a multi-unit franchisee model it in a spreadsheet without a sales engineer on Zoom?

Upsell economics — and this is the one most operators are about to get hosed on. Who owns the revenue share on AI-suggested adds? Is it on net new attach, or on every modified ticket? Does the vendor get paid when the upsell is declined? Does the meter run when the lane is closed?

A 5 in build quality with a 2 in upsell economics is a vendor you should walk away from. A 4-4-4-4 is a vendor you should sign. The arithmetic is unforgiving once you multiply by units.

SoundHound’s Polaris, decoded

SoundHound is the loudest voice-AI vendor in restaurants right now, and they have earned some of it. Their Polaris foundation model — speech-to-meaning, not the speech-to-text-then-LLM pipeline most of the field still runs — is a genuinely different architecture, and the operators I talked to who deployed it at White Castle, Krystal, and Jersey Mike’s locations describe a noticeable step change in how the system handles modifier chains. “Two number threes, one with no pickle, the other with extra pickle, and a Diet Coke instead of the regular one on the second combo” is the kind of utterance that snaps the older pipeline-style systems. Polaris generally holds.

Build quality, 4 of 5. The handoff behavior is the cleanest in the field — when the model is uncertain, it routes to a human crew member with the partial order intact, so the customer doesn’t have to repeat themselves. That’s the single biggest predictor of customer satisfaction in a voice-AI lane and most vendors still get it wrong.

Operator impact is where I have questions. SoundHound’s case studies lean heavily on labor redeployment, which is real — a crew member who isn’t taking orders is making sandwiches — but the throughput claims pitched in private decks are more aggressive than what operators report back. One Jersey Mike’s franchisee in the southeast told me their order-time number moved by about eleven seconds, which is meaningful but not the thirty-second number on the slide. Operator impact, 3 of 5, with the caveat that the number depends entirely on what your prior baseline was.

Pricing transparency is the problem. SoundHound’s standard contract is a hybrid: a per-lane monthly fee that’s reasonable in isolation, plus an upsell revenue share calculated on a base that includes anything the model suggested, even when the customer declined. The specific percentages are confidential, but the structure is the issue, not the headline number. If your crew was already trained to suggest a cookie with every sandwich and the AI now does it, you’re paying on attach behavior that pre-dates the contract. Pricing transparency, 2 of 5. Upsell economics, 2 of 5. Negotiate the base. Walk if you can’t move it.

Hi Auto’s two-anchor strategy

Hi Auto is the quietest serious player in the room, and they’re playing a different game. They have the Checkers/Rally’s relationship and they have parts of the Wingstop deployment, and what’s interesting about both anchors is how different the deployment shapes are. Checkers is a high-volume, high-noise, modifier-light drive-thru — the model needs to be fast and robust to ambient noise, but the menu surface is constrained. Wingstop’s phone-order channel is the opposite — relatively quiet input, but a menu with combinatorially large modifier options once you start counting flavors, sauces, and sides.

Running both anchors gives Hi Auto a training-data footprint that’s genuinely useful, and you can hear it in how the model handles the long-tail utterances at either end. Build quality, 4 of 5. They’re not as flashy as SoundHound on the speech-to-meaning architecture story, but the operator experience is steady and the failure modes are predictable, which is what you actually want in a lane.

Operator impact, 4 of 5. The Checkers operators I talked to describe a real labor recapture — they redeployed an order-taker to the make line and saw the make-line throughput improve enough to absorb the redeployment without adding cost. That’s the deployment shape that actually pencils out. Wingstop’s phone-order team describes Hi Auto’s product as essentially eliminating call abandonment during the Friday rush, which for a wings-on-Friday concept is the entire ballgame.

Pricing transparency, 4 of 5. Hi Auto’s standard deal is closer to flat per lane (or per phone line) with a smaller, capped upsell component on net-new attach only. That’s the shape operators should be pushing every vendor toward. Upsell economics, 4 of 5 — the cap matters, and the “net-new attach only” framing means the vendor doesn’t get paid when the model suggests a cookie to a customer who was going to order a cookie anyway. This is the right structure. If you’re shopping voice AI at NRA, ask every other vendor to match Hi Auto’s commercial shape, and watch their faces.

Callout: Build quality

The thing nobody puts on their slide is failure-mode behavior. Every voice-AI model in this field will misunderstand a customer somewhere between five and twenty percent of the time depending on what you count. What matters is what happens next. Hi Auto and SoundHound both route the partial order to a human crew member with context preserved. Two of the other vendors in this field reset the conversation, which means the customer repeats the order, which is a worse experience than not having voice AI in the first place. Ask for a live demo of the failure mode, not the happy path.

Presto’s reset

Presto has been through a real public restructuring, and there’s a lot of noise about them that isn’t useful for an operator making a 2025 purchase decision. The facts that matter: Presto’s voice AI product is in real lanes at Carl’s Jr., Hardee’s, and Del Taco, the deployments are smaller-footprint than a year ago, and the product team has narrowed focus to the lane experience after exiting broader hospitality bets.

Build quality, 3 of 5. Honest middle of the road. The model handles the standard utterance set well and falls over on the long tail more than SoundHound or Hi Auto, but the operators I talked to describe a product that is genuinely usable in a high-volume QSR environment and where the support relationship is responsive. The new commercial leadership has been clear in their pitch that they’re not trying to be the most accurate vendor in the field — they’re trying to be the vendor who’s easiest to deploy and operate.

Operator impact, 3 of 5. Real but unspectacular. The reason it’s not lower is that Presto has been deployed long enough at Carl’s Jr. that there’s a body of operator data to talk to, which is more than most of the field can say.

Pricing transparency is where Presto has materially improved. The post-restructuring contracts I’ve seen referenced are simpler than the pre-restructuring contracts — closer to a flat per-lane subscription with a defined upsell share on a defined base. Pricing transparency, 4 of 5. Upsell economics, 3 of 5. I’d negotiate the upsell base aggressively but the starting position is more operator-friendly than it was eighteen months ago. If your alternative is signing the SoundHound hybrid as written, Presto is a reasonable counter-bid to bring to the SoundHound negotiation, and if SoundHound won’t move, Presto is a defensible second choice.

Kea’s self-serve play

Kea is the vendor most operators have either never heard of or have very strong opinions about, and the reason is that Kea is not really competing for the drive-thru lane at all. They’re competing for the phone-order channel, primarily in pizza and Italian-American casual concepts, and their product is wrapped as a kind of self-serve platform that a multi-unit operator can deploy without a full enterprise integration project. Long John Silver’s is the chain-level proof point most often cited, but the more interesting story is the long tail of two-to-twenty unit independents and small franchises who are running Kea on the phone channel because their alternative is a third-party answering service or, worse, just letting calls go to voicemail during the rush.

Build quality, 3 of 5. The model is fine. It’s not pushing the speech-to-meaning frontier, but for the phone-order use case at a pizza concept, the menu surface is constrained enough that the build-quality bar is lower than the drive-thru bar. The product works.

Operator impact, 4 of 5. This is the surprise of the field. For the right concept — high call volume, constrained menu, phone-orders-as-revenue-channel — Kea materially changes the unit economics. One independent operator I talked to is putting Kea on three locations because the math is that obvious: their phone abandonment rate during dinner rush dropped from somewhere in the high teens to low single digits, and the recovered orders pay for the platform several times over.

Pricing transparency, 5 of 5. Kea publishes its pricing. It’s on the website. You can model it in a spreadsheet before you talk to a sales rep. This is the bar every vendor should be held to and almost none of them are. Upsell economics, 4 of 5 — Kea’s upsell share is reasonable and the base is well-defined. The reason it’s not a 5 is that the model itself is less aggressive on upsell than the drive-thru vendors, which is good for customer experience but means the upsell line item in the operator’s P&L is smaller. Tradeoff, not a flaw.

Callout: Operator impact

The single most useful question I learned to ask operators in the last three weeks: “What did your crew do with the time the AI gave back?” The operators who had a clear answer — redeployed to make line, redeployed to runners, redeployed to lobby cleanliness — saw real P&L impact. The operators who didn’t have a clear answer saw the AI absorb labor minutes that just disappeared into general overhead. The product is the same. The deployment plan is the difference.

The Yum × NVIDIA shadow

This is the story that should be reframing every conversation at NRA, and I’m not yet convinced it is.

In March, Yum Brands announced a voice-AI partnership with NVIDIA built on NVIDIA’s NIM microservices, with the deployment planned to reach up to 500 Taco Bell, KFC, Pizza Hut, and Habit Burger locations in 2025 (Restaurant Dive). The contractual and architectural shape of that deal is the part operators outside the Yum system should be paying attention to, because it changes the buy-vs-build math for any chain over roughly 200 units.

What the Yum × NVIDIA deal does, in effect, is package the NIM-based voice-AI stack in a way that a sufficiently sophisticated operator could, with the right systems-integrator partner, replicate without going through one of the vendors I’ve spent the last four sections describing. Not easily. Not cheaply, and probably not at fewer than several hundred units. But the option exists in 2025 in a way it did not exist in 2024, and that changes the negotiating posture of every chain large enough to consider it.

The other anchor data point is Wendy’s. Wendy’s FreshAI is reporting an 86% no-intervention rate on its drive-thru voice AI deployment and is scaling to 500-600 restaurants by year-end 2025 (per the company’s Square Deal blog and confirmed via Restaurant Dive’s reporting and Restaurant Business). FreshAI is built on Google Cloud’s stack with Wendy’s-side engineering on top. It’s the other proof point that large chains are not waiting for a single vendor to solve voice AI for them — they’re building the integration capability internally and treating the speech model as a component.

What this means for an operator walking the NRA Show floor: if your chain is large enough to be in the Yum or Wendy’s reference class, you are not the customer the vendors at NRA are pitching for, and you should be having a different set of conversations with NVIDIA, Google, and a systems integrator. If your chain is smaller than that — and most operators reading this are — then the question is which of the vendors above gives you the cleanest path to a working lane without the build-it-yourself overhead.

The shadow the Yum deal casts on the rest of the field is real and it should be making every vendor in this article nervous about their long-term pricing power. The vendor that responds to the Yum × NVIDIA pressure by lowering their upsell share is the vendor operators should reward. The vendor that responds by raising the per-lane fee to compensate is the vendor operators should walk away from.

ConverseNow and the agentic phone-order play

I’m grouping ConverseNow at the end because they’re the vendor whose category boundary has moved the most in the last twelve months. ConverseNow started as a phone-order voice-AI vendor — Domino’s was an early anchor, though the relationship there has evolved — and has expanded into drive-thru and into more agentic patterns where the AI handles not just the order capture but the order completion and the upsell sequencing across multiple turns.

Build quality, 4 of 5. The multi-turn handling is genuinely good. The model knows when to push the upsell and when to back off, which is harder than it sounds and is the part that separates the systems customers describe as “natural” from the ones they describe as “robotic and annoying.”

Operator impact, 4 of 5. The phone-order deployments at Domino’s-class concepts have a body of operator data behind them that’s deeper than anyone else’s outside of Wendy’s, and the throughput claims have held up in operator interviews better than the other vendors’ have.

Pricing transparency, 3 of 5. ConverseNow’s contracts are negotiated, not published, and the structure varies by chain size. Upsell economics, 3 of 5 — the upsell share is real but the base is more defensible than SoundHound’s. Operators report that the negotiation is genuinely negotiable, which is more than most vendors in this field can say.

The reason ConverseNow matters at NRA specifically is that they’re the vendor most likely to walk operators through a credible buy-vs-build conversation. They are, in effect, the build-partner option that’s been productized — which means the conversation with their sales team feels more like a systems-integrator conversation than a SaaS pitch. For chains in the 100–500 unit range, that’s exactly the conversation that should be happening, and it’s the conversation the more SaaS-shaped vendors are structurally less able to have.

What to buy, what to walk

The rubric, totaled:

  • SoundHound: 4 / 3 / 2 / 2 = 11. Best build quality in the field. Worst pricing structure. Sign only if you can move the upsell base.
  • Hi Auto: 4 / 4 / 4 / 4 = 16. The quiet right answer. The commercial shape every other vendor should be benchmarked against.
  • Presto: 3 / 3 / 4 / 3 = 13. Reasonable middle. Strong counter-bid leverage in a SoundHound negotiation.
  • Kea: 3 / 4 / 5 / 4 = 16. The right answer for phone-channel deployments at independents and small franchises. Published pricing is a competitive moat in this field.
  • ConverseNow: 4 / 4 / 3 / 3 = 14. The right answer for 100–500 unit chains who want a build-shaped partner without the build cost.
  • Yum × NVIDIA shadow: not a vendor you buy, but the macro that reshapes everyone else’s pricing power.

Callout: Watch-list

Three things I’m watching between now and the show floor:

One. Toast reports Q1 tomorrow (May 8). The POS-side AI product line — order-suggestion at the digital channel, outbound on the marketing side — is going to be priced into the same operator P&L as the voice-AI line, and the way Toast frames the bundle on the earnings call will tell us whether the POS layer is about to compete with the voice-AI layer on upsell economics. (More on this in our later Vibe Check on the POS-side AI.)

Two. Cava reports Q1 on May 15, two days before NRA opens. The Mediterranean fast-casual category is the most interesting non-burger, non-pizza test bed for voice AI in 2025, and the way Cava talks about digital channel mix on the call will signal whether the category is a real opportunity for the vendors above or whether the channel mix favors the discovery layer.

Three. Olo also reports Q1 tomorrow (May 8). Olo sits at the integration layer between the brand digital stack and the third-party channels, and the way Olo talks about voice AI on the call will tell us whether they’re positioning as a neutral integration layer or as a vendor in their own right. The former is good for operators. The latter is a complication.

Vibe Check verdict

If you walk McCormick Place asking which voice-AI vendor has the best accuracy, you will be sold a product on a metric that doesn’t determine your unit economics. Ask instead who owns the upsell rev share, on what base, with what cap, and whether the meter runs when the lane is closed. Hi Auto and Kea are pricing in a way you should sign. Presto is pricing in a way you should counter-bid with. SoundHound is pricing in a way you should walk away from unless they move on the upsell base. ConverseNow is the right shape for chains who want a build-partner relationship without building.

And the shadow of the Yum × NVIDIA deal — and Wendy’s FreshAI’s 500–600 unit FY25 footprint — is the macro that should be making every vendor in this field renegotiate their own pricing power. As our later coverage of the discovery layer’s voice agents argues, the upstream channel for these orders is also shifting under all of these vendors’ feet, and the operator who signs a five-year contract in May 2025 is signing into a market that is going to look materially different by 2027. As our later piece on hospitality enterprise deployment will get into, the larger-chain calculus on owned-stack vs. vendor-stack is moving the same direction in the lodging category for the same reasons. The contracts you sign this month should have term and exit clauses that reflect that.

I’ll be on the floor at McCormick Place May 17 through 20. If you’re an operator who’s negotiated one of these contracts and the structure surprised you — for better or worse — I want to hear about it.

— Sofia runs Vibe Check. Tips: [email protected].

Featured More

The Voice Agent Maturity Curve

mise

·

12 min read

The Four Margins of a Restaurant

mise

·

14 min read

The AI Premium in Hospitality M&A: Broker Story or Real Number?

the bottom line

·

9 min read

What the DoorDash/SevenRooms Deal Actually Buys

the bottom line

·

11 min read

Browse all 494 posts

Related posts

Desk Review: Lightspeed Restaurant, the Quiet Half of the Duopoly

vibe check

·

16 min read

Desk Review: Lightspeed Restaurant, the Quiet Half of the Duopoly

Desk Review: OpenTable's 'System of Record' — what restaurants are actually agreeing to on April 16

vibe check

·

18 min read

Desk Review: OpenTable's 'System of Record' — what restaurants are actually agreeing to on April 16

Desk Review: Toast Drive-Thru — the bundle, the moat, and the 15-unit floor

vibe check

·

18 min read

Desk Review: Toast Drive-Thru — the bundle, the moat, and the 15-unit floor