The Drive-Thru Voice Stack in 2025: A Buyer's Framework After Vox, Loman, Presto, SoundHound
Operators keep asking me which voice-AI vendor to pick for the drive-thru. The question is wrong. Voice AI isn't one product — it's four discrete decisions, each with different incumbents, different failure modes, and different ROI math. Here's how to buy it.
A regional QSR franchisee — fifty-something stores across two states, the kind of operator who built her business one rooftop at a time — called me last Tuesday with a question I’ve been getting in some form every week since the spring. She’d just sat through back-to-back pitches from three voice-AI vendors and a fourth had emailed that morning. “I need you to tell me which one to buy,” she said, and then before I could answer: “And don’t say it depends. I don’t have time for it depends.”
It depends.
Not because the technology is too new to evaluate or because the category is too fluid to call. Both of those things are true and neither is the reason. The reason is that operators keep walking into these meetings asking the wrong question. They ask which voice AI, as if there’s a single product with a single feature list and a single ROI calculation. There isn’t. What’s sitting under the umbrella label “drive-thru voice AI” in September 2025 is actually four different products solving four different problems, each with a different incumbent set, a different failure profile, and — most importantly — a different unit economics story. Buy them as one thing and you’ll over-pay, under-deploy, or both. Buy them as four things and the picks get obvious fast.
The contrarian thesis I want operators to internalize before the next pitch deck lands: stop evaluating drive-thru voice as one product. It’s four discrete decisions — order-taking, employee assist, payments, brand voice — and each one has different incumbents, different procurement timelines, and different right answers. The vendor consolidation pitch (“we do all four”) is a sales motion, not a buying framework. You may end up buying from one vendor across all four. But you should make four decisions, not one.
The market got crowded fast, and that’s the tell
The clearest external signal that this is a four-product category is who showed up to compete. Restaurant Business in August framed the voice-AI market as “getting crowded” — and the names they listed don’t look like competitors to each other. Presto and Hi Auto and ConverseNow and SoundHound, the mature stack, all started in order-taking but have drifted into different adjacencies. The 2023-and-after entrants — Vox, Loman, Incept, Palona, Maple, Revmo — didn’t show up to clone Presto. They showed up to attack pieces of the workflow Presto wasn’t built around.
The capital tells the same story. PYMNTS reported in late August that Vox closed an $8.7M round to bring voice AI to restaurant drive-thrus, and the framing across the trade press treated it as one more entry in a category that, per CB Insights data PYMNTS cited, saw voice-AI startups raise $2.1B globally in 2024, roughly 8× the prior year. Restaurant Technology News covered the same Vox raise with a slightly different emphasis — the word “autonomous” doing a lot of work in the headline. And Crunchbase’s August roundup made the broader market dynamics visible: capital is flooding into voice AI generally, restaurants are one of the most-named verticals, and the early-stage entrants are explicitly not trying to be Presto.
If this were one product, the late entrants would be Presto clones with a 10% pricing discount and a slightly better demo. They aren’t. Vox is pitching full autonomy on the order-taking side. Loman is pitching employee-facing assist. Palona is pitching brand voice. The market is segmenting itself in real time, and the segmentation is the framework operators should use to buy.
Decision one: order-taking
This is the product that gets all the headlines and most of the deck slides. It’s also the one with the longest deployment history, the most measurable performance benchmark, and the most established vendor set. When a Wendy’s location is “running voice AI at the drive-thru,” what they mean — almost always — is that the order-taking conversation is being handled by an AI agent rather than a crew member with a headset.
The incumbents here are real. Hi Auto powers a meaningful share of the active Wendy’s FreshAI rollouts. ConverseNow has been in market with Domino’s, Wingstop, and a long tail of chains for years. SoundHound owns the Hyundai and broader auto-OEM ambient voice surface and has translated that into restaurant deployments at multiple QSR chains. Presto has had a turbulent year publicly but remains the operating layer at thousands of Carl’s Jr. and Del Taco lanes.
The newer entrants — Vox most pointedly — are trying to beat the incumbents on a single dimension: autonomy rate, the share of orders the agent completes without human escalation. Vox’s pitch, in PYMNTS’ framing, is “fully autonomous.” Whether that lands at the rates the marketing implies is the central question, and it’s the one operators should drill on. Ask vendors for autonomy rate broken out three ways: simple orders (one item, no modifications), modal orders (your highest-volume two-to-three item combinations), and tail orders (custom builds, allergen swaps, group orders with six-plus items). The vendor that quotes one number for “autonomy rate” without that breakdown is hiding the tail performance, and the tail is where labor cost actually leaks back into the lane.
The other thing to drill on: speed-of-service contribution. Voice AI’s value at order-taking isn’t just labor reduction; it’s a steadier order-taking time across dayparts. The crew member at 9 PM who is also covering fryers, also bagging orders, also fielding lobby questions is the bottleneck the agent is replacing. Ask vendors for the variance reduction in order-taking time, not just the average. A vendor that improves your mean OE-time by ten seconds while leaving the P95 unchanged hasn’t actually changed your throughput; you’re still gated by the worst moments.
The active rollout footprint as of late summer is informative on its own. Taco Bell is deep into a multi-vendor voice-AI deployment. Wendy’s, per CEO Kirk Tanner’s February 2025 earnings remarks, was targeting 500-600 stores by year-end on FreshAI and reporting an 80 basis-point profit margin lift at company-operated restaurants in the pilot footprint (company-reported figure — treat with appropriate caveats). Bojangles, White Castle, and Taco John’s are expanding. The list of national chains with zero voice deployment is shrinking fast.
What it means for operators: order-taking is the most mature of the four decisions and the one with the most data to underwrite. If you’re a multi-unit franchisee evaluating order-taking voice as a 2026 line item, the right approach is to insist on a five-to-ten store pilot with the autonomy-rate breakdown above, a clear escalation-to-human protocol, and contractual visibility into how the agent handles your top fifty SKUs by mix. The vendors that resist that level of scrutiny aren’t ready; the ones that welcome it are the ones whose autonomy rates hold up under inspection.
For the broader category map of which agents are mature enough to deploy and which are still pre-pilot, a forthcoming framework piece on voice agent maturity walks through the diligence criteria in more depth — the short version is that the order-taking layer is the only one of the four decisions that has a stable, defensible benchmark you can hold vendors to today.
Decision two: employee assist
This is the decision most operators don’t realize is a separate buying motion until they’re three months into an order-taking deployment and the morale problem starts.
Here’s the failure mode: you deploy voice AI on order-taking, the agent handles eighty percent of orders, the twenty percent that escalate to your crew are now disproportionately the hard orders — the eight-person group, the customer with three modifications, the angry guest who refused to talk to the agent. Your crew’s job got harder, not easier, and they didn’t get paid more. Six months in, your turnover at the lane spikes and your GM is calling corporate asking why.
Employee assist is the product that solves this. It’s not a customer-facing agent; it’s an internal one. It sits in the crew member’s ear or on the lane-side tablet and does for the human what the order-taking agent does for the customer: it pulls up the modifier menu, suggests the right SKU when the customer describes the burger by attributes rather than name, pre-stages the upsell, surfaces the loyalty status. Loman is the most visible entrant explicitly pitching this. Several of the order-taking incumbents — SoundHound notably — have rolled employee assist into the same conversational surface, which is convenient but conflates two procurement decisions.
The reason it should be a separate decision: the buyer is different. Order-taking is bought by ops and finance because it shows up on the P&L as labor. Employee assist is bought by HR and store leadership because it shows up in turnover and ramp time. The ROI math is different — order-taking is dollars per labor hour saved, employee assist is dollars per turnover event avoided and weeks shaved off the four-to-six week ramp curve for a new crew member. The procurement cycle is different. The integration profile is different — employee assist has to read your training materials, your handbook, your menu changes; it doesn’t have to interface with the speaker post.
The thing that surprises operators in employee-assist pilots is how much of the value lives in the first thirty days of a new hire. Restaurants run at 100%+ annual turnover; if employee assist takes a crew member from breakeven productivity in week six to week three, that’s the entire ROI on its own, before you count anything the agent does for tenured staff. Ask the vendor for ramp-curve data, not just steady-state productivity lift.
What it means for operators: don’t buy employee assist as a bolt-on from your order-taking vendor without first scoping it independently. Get a real pilot in front of your training and HR organization. The right answer may end up being the same vendor — bundling is real and the integration savings are real — but you should make the decision from a position of having shopped it.
Decision three: payments
This is the decision operators most under-weight, and the one where the wrong answer creates the most downstream pain.
When a voice agent takes an order, it eventually has to take a payment. There are roughly three ways this happens today. One: the order is taken by voice, the customer pulls up to the cashier, and a human handles payment — the status quo, no voice in the payment loop at all. Two: the order is taken by voice, the total is read back, the customer pulls up, and an unattended payment device — a card reader, a contactless terminal — handles the payment. Three: the voice agent handles payment in the conversation itself, capturing the card details verbally or directing the customer to a phone-based payment flow.
Each path has different vendors, different PCI exposure, and different fraud profiles. Option one is the most common today and the least interesting; the voice vendor is selling you order-taking, not payments. Option two is what most current voice-AI deployments are pairing with — the voice agent hands off to the existing card terminal, and the payment vendor is whoever your POS provider integrates with. Option three is where the early entrants — some of Vox’s positioning around “fully autonomous” implies an answer here — are trying to push the category. Option three is also where the PCI and fraud exposure get real.
The mistake I see operators make is treating payments as a feature of their voice vendor rather than a procurement decision in its own right. Your voice vendor isn’t your payment processor; if they’re pitching themselves as one, you should ask very specifically how the card-not-present transaction is being captured, which acquirer they’re routing through, what their fraud-loss attribution looks like, and how chargebacks are handled. Voice-captured payments are still a small share of QSR drive-thru volume; the operational playbook for handling disputes on a voice-captured order isn’t standardized.
The simpler version of decision three for most operators in 2025: assume payments stays on your existing POS/payments stack, treat any “we do payments too” pitch from a voice vendor as out-of-scope for the initial pilot, and revisit in 2026 when the regulatory and fraud picture is clearer. The exception is if your voice vendor is offering meaningful unit economics improvements that are only available with their payments integration — in which case you should be skeptical and ask why.
Decision four: brand voice
This is the decision operators ignore until a customer recording goes viral and corporate gets defensive.
Every voice agent has a voice — a specific TTS output, a specific phrasing pattern, a specific way of handling humor, a specific failure mode when the customer says something the agent didn’t expect. Until 2024, this was largely a vendor default; you got whatever voice the platform shipped with. Starting in late 2024 and accelerating through 2025, brands started to push back on that, and a small set of vendors — Palona most explicitly — are pitching brand-voice as a separable product layer. The pitch: your voice agent should sound like your brand, not like a generic SaaS TTS voice.
The reason this matters more than it sounds: voice is a brand surface in a way text never was. A McDonald’s voice agent and a Wendy’s voice agent that both use the same vanilla TTS sound identical to a customer. That’s an extraction of brand equity that brands didn’t notice they were giving up when they signed the first generation of voice contracts. The brands that have caught this — Wendy’s is one — are starting to insist on differentiated voice as part of the procurement.
What’s available in the market in September 2025 is a spectrum. At one end, vendors who let you pick from a handful of stock voices and adjust nothing else. In the middle, vendors who will tune phrasing patterns and add brand-specific greeting/farewell sequences. At the far end, vendors offering custom-trained voices — actor-recorded or synthetic — that are unique to your brand and contractually exclusive. The price tier is steep at the far end and the engineering integration is non-trivial. For most independent operators and mid-market franchisees, the middle tier is the right place to be. For national chains with brand teams that take voice seriously, the far end is where the conversation is heading.
What it means for operators: in your RFP, include a voice-customization section. Not because you’ll necessarily exercise the most-customized option, but because the vendor’s answer tells you whether they’re a voice platform or a voice product. Platforms let you tune; products give you a default. Both are valid; you should know which you’re buying.
How the four decisions actually compose
The trap I want operators to avoid is the “we do all four” sales pitch from a single vendor. It’s not that bundling is wrong — there are real integration benefits, particularly between order-taking and employee assist where the data model is shared. It’s that bundling collapses four discrete diligence processes into one, and the weakest decision in the bundle drags down the strongest.
Here’s the composition I’d run if I were buying for a fifty-store franchisee in September 2025:
Start with order-taking. Pilot two vendors against each other on five-to-ten stores for ninety days, measured on autonomy rate (broken out by order complexity), P95 order-taking time, and crew override frequency. Pick a winner. The winner is now your primary voice vendor for purposes of integration economics — they get the inside track on the other three decisions but don’t get them by default.
Run employee assist as a separate evaluation. The right vendor may be your order-taking winner; it may also be Loman or one of the newer entrants. Decide on ramp-curve impact, not just steady-state productivity. Make sure the buyer is HR/store ops, not the same person who bought order-taking.
Defer payments. Keep your existing POS/payments stack. Re-evaluate in 2026 unless your order-taking vendor is offering specific unit-economics gains tied to a payments integration; if they are, treat that as a separate negotiation with its own diligence.
Take brand voice seriously, but right-size it to your scale. If you’re under a hundred stores, the middle tier — customized phrasing, branded greetings, stock voice — is almost certainly the right answer. If you’re a national chain, get your brand team in the room before you sign.
The reason this framework matters more in 2025 than it did in 2023: the cost of the bundling mistake compounds. A locked-in three-year contract with a vendor that’s good at order-taking and mediocre at employee assist will make your next two procurement cycles harder. The vendor knows this and will push for bundle deals with the longest available terms. The right operator response is to negotiate term length down, hold the order-taking decision to the highest standard, and keep the other three decisions live.
What McDonald’s and the larger players have already taught the rest of the category
The larger chains have been making these four decisions, in some form, for two years now. McDonald’s experience — the well-publicized end of the IBM partnership and the subsequent reset of their voice strategy — taught the industry that buying voice as one product from one vendor on a multi-year exclusive was the wrong shape. An upcoming piece on McDonald’s AI drive-thru goes deeper on that arc, but the short version is that the chains that have moved fastest in 2025 are the ones who structured their procurement to allow vendor swaps at each of the four layers independently.
Wendy’s is the cleanest example of operating the framework in practice. The order-taking layer (FreshAI, built on the Google Cloud relationship) is one decision. The employee-facing tools are a separate procurement track. Payments stayed on the existing stack. Brand voice — listen to a Wendy’s voice-AI order — is intentionally differentiated. Kirk Tanner’s February earnings remarks on the 80 basis-point margin lift at company-operated stores reflect the order-taking layer specifically; the other three contribute in different ways and aren’t typically broken out in earnings.
Taco Bell, similarly, has run a multi-vendor approach rather than a single-supplier approach, even within the order-taking layer. This is the right shape. It’s slower and more operationally expensive than picking one vendor and going, but the option value of being able to swap is worth more than the integration savings of being locked in.
Mark interpretation
The vendors are going to keep consolidating their pitches. Expect every voice vendor by Q1 2026 to claim they do all four — order-taking, employee assist, payments, brand voice — under one platform. The pitch will get more aggressive as the capital cycle continues. The Vox raise this August won’t be the last; my read on the Crunchbase market data is that we’ll see three to five more meaningfully-funded entrants in the next two quarters, each making the bundled pitch.
Operators who buy the bundle pitch will sign three-year contracts in 2026 that they spend 2027 regretting. The bundle math doesn’t pencil out at this stage of the category: the integration savings are real but small, the lock-in cost is real and large, and the right composition of vendors today is almost certainly going to look different two years from now.
The operators who buy this category the right way are the ones who treat each of the four decisions on its own terms — with its own pilot, its own buyer, its own success metric, its own contract term. It’s more work. It also produces a stack that you can evolve as the category evolves, which is the only kind of stack worth owning in a market growing this fast.
If you’re sitting in front of a voice-AI pitch deck this fall and the deck has one slide listing all four decisions as features of the same product, the right question to ask isn’t which features are best. The right question is: which of these four would you sell me standalone, and what’s the contract term if I only buy that one? The vendor’s answer tells you everything you need to know about whether they’re a real partner or just a sales motion.
— Priya covers operators for TableTransfers. Tips: [email protected].
The Voice Agent Maturity Curve
mise
·12 min read
The Four Margins of a Restaurant
mise
·14 min read
The AI Premium in Hospitality M&A: Broker Story or Real Number?
the bottom line
·9 min read
What the DoorDash/SevenRooms Deal Actually Buys
the bottom line
·11 min read