The Voice-AI Operator Bench: What Bojangles, White Castle, Taco John's, and Jersey Mike's Are Actually Doing With AI Ordering

A regional sandwich-shop counter at the lunch rush, headset on the host stand, the phone line lit.

Voice AI in December 2025 is less Wendy's-everywhere and more a quiet regional bench. The Jersey Mike's deployment is phone-orders for fifty stores, not the drive-thru — and most operator wins this year look like that one when you read the press release line by line.

I spent a Tuesday before Thanksgiving at a Jersey Mike’s in a mid-sized Northeast suburb watching the lunch rush hit the counter and the phone simultaneously, and the most interesting thing in the room was not the line. It was the second headset, sitting next to the POS, that the shift lead had not picked up once in twenty-three minutes. Phone orders were coming in — three of them in the window I watched — and they were not going to a person. They were going to SoundHound. Pickups, not the drive-thru. Sub orders, not a queue of cars. The shift lead had a clipboard with timed pickup windows on it, and the voice agent had given her four of them. She told me, after the bagger left, that her old job during the rush was to take down four sandwiches per call, repeat them back, ring them, and shout them to the line. Her new job during the rush was to look at the printer.

Here is the thesis I want to put down before the year-end pieces frame the year as the year voice AI broke through: voice AI’s December 2025 reality is not Wendy’s-everywhere. It is a quiet regional expansion across brands you have stopped reading about — Bojangles, White Castle, Taco John’s, Jersey Mike’s — and the deployment scope, in almost every case, is narrower than the vendor marketing implies. The Jersey Mike’s number every voice-AI deck cites is real. It is fifty locations. It is also, per the SoundHound release that produced the number, phone orders for pickup — not the drive-thru, which Jersey Mike’s largely does not run.

That distinction is the article. Operators reading vendor decks in 2026 budgeting season are going to be quoted Jersey Mike’s, Wendy’s, and White Castle in the same slide, and the slide is going to imply that voice AI is a drive-thru-ordering category. It is not — yet. It is a portfolio of point-solutions sitting at three different parts of the order: the phone, the drive-thru lane, and the kiosk. The brands that are live on each are not the same brands. The vendors that are live on each are not the same vendors. And the operator economics are not the same economics.

What the SoundHound release actually said

Go back to SoundHound’s January 24, 2024 release on the Jersey Mike’s deployment — the source for every “Jersey Mike’s is on voice AI” line you have read since — and read the scope sentence carefully. The release said the deployment was initially live at 50 locations and that the voice technology allowed customers to place phone orders for pick-up. The CIO, Scott Scherer, confirmed the phone-ordering scope on the record at the time. The release did not say drive-thru. It did not say in-store ordering. It did not say kiosk. It said phone orders, for pickup, at fifty stores.

This matters because the press cycle since has collapsed the scope. The Restaurant Business piece on the crowded voice-AI market — which is the most-cited single article in operator decks I have seen this fall — lists Jersey Mike’s, Wendy’s, White Castle, and a half-dozen other deployments in the same paragraph. The framing implies a category. The deployments are not the same product. Wendy’s FreshAI is at the speaker post. White Castle’s SoundHound deployment is at the speaker post. Jersey Mike’s is at the phone. The economics of those three are different, the customer experience risk is different, and the failure mode is different.

The phone-order use case is the cleanest one to ship and the least-discussed. There is no ambient kitchen noise. The customer is on a known device. The order is asynchronous — the customer is not waiting on a car behind them. The failure mode is a re-route to a human, which the customer expects from a phone tree anyway. Most importantly: the brand does not own a drive-thru lane, so phone is the whole off-premise order rail, and shipping voice AI on it is shipping voice AI on the off-premise channel that matters to them. Jersey Mike’s, in other words, did not deploy a partial solution. They deployed the right solution for the channel they have. The Wendy’s narrative does not transfer.

The regional bench, named

The brands I keep reading about in December 2025 are not the brands that were on the deck a year ago. The Wendy’s, McDonald’s, Taco Bell axis is the one the trade press covers because the unit counts are large and the parent companies are public. That coverage produces the impression that voice AI is a top-ten-QSR story. It is not, in December. It is a regional-bench story.

The bench I have been tracking, by way of the Restaurant Business voice-AI market piece and the operator conversations that ladder back to it:

  • Bojangles — expanding voice AI at the speaker post in selected Southeast markets. The Bojangles deployment is the one most likely to look like a Wendy’s-style FreshAI deployment from the outside, because it is a chicken-and-biscuit brand with a real drive-thru lane and breakfast-into-lunch dayparts that produce the kind of order-flow regularity voice models train against.
  • White Castle — the SoundHound deployment is the one with the longest operating runway, and the one whose hand-off-to-human rate is most-cited in vendor decks. Read the cited number against the deployment scope: a smaller footprint with a tight menu produces a hand-off rate that does not generalize to a brand with a larger or more LTO-heavy menu.
  • Taco John’s — voice AI expansion across the franchise footprint, the most franchise-heavy of the four named here, which is the one I would watch hardest for the operating-cost reality. Franchisees pay the per-order fee; corporate writes the contract. The economics for the franchisee are the economics that matter, and they are not the economics quoted in the parent-brand press release.
  • Jersey Mike’s — the fifty-store phone-order deployment, not a drive-thru deployment. Mark this as the most-misread item on every deck I have seen this fall.

These four brands have less in common than the voice-AI category implies they do. The category is doing the work of grouping them; the deployments are doing different things. If you are an operator reading a vendor deck this December, the right question is not “is voice AI working” — it is “which of the four channels (drive-thru speaker, phone, kiosk, app callback) is the vendor pitching me, and is that the channel that matters in my unit economics?”

Taco Bell, Yum, and the Nvidia partnership read

The biggest piece of December voice-AI news, by megaphone size, is the Taco Bell and Yum Brands partnership with Nvidia. The framing in the press around it is the one I want to flag: the partnership is positioned as a top-line endorsement of voice AI in the QSR drive-thru. The substance, if you read it carefully, is a compute-and-modeling partnership at the parent-brand level. It is not, in the December 2025 release window, a per-store deployment count. It is a forward-looking commitment with infrastructure attached.

That distinction is the distinction operators should care about. A compute partnership is a 2026-2027 story. A fifty-store live phone-order deployment is a now story. The trade press does not consistently separate the two. Vendor decks aggregate them. Operators get the impression that the QSR voice-AI category is further along than it is — because the most-impressive press releases are forward-looking infrastructure announcements, and the most-real production deployments are smaller, channel-specific, and quieter.

I am going to be careful here: the Taco Bell-Nvidia partnership is real, the compute scale is real, and the brand’s stated direction on voice AI is consistent. This is my read, not the company’s: the gap between announcement and per-store live deployment is the gap that closes in 2026, not in December 2025. The drive-thru voice-AI category is not ready to deploy at Taco Bell’s full footprint scale this calendar year, and the announcement is not claiming that it is. The trade press, in the aggregate, is leaving the reader with that impression anyway.

The vendor side: who got funded, and for what

The vendor stack matters here because the funding pattern in the back half of 2025 tells you which channel each vendor is betting on. Per BusinessWire’s August 27, 2025 release, Vox AI raised an $8.7M seed round, bringing total funding to $10M, on a positioning of autonomous voice AI for QSR drive-thrus. The Restaurant Technology News coverage of the same round cites Vox’s claimed 17x ROI on labor savings and 90+-language support. Both numbers are vendor-stated. Both are the kind of numbers an operator should ask to see in writing, by channel, before they sign a contract.

Loman AI’s $3.5M seed earlier this year is a smaller round but the more interesting one for the phone-order use case, because Loman is positioned closer to the Jersey Mike’s-style deployment — phone ordering, off-premise, asynchronous — than the drive-thru speaker post. The funding amount tracks the channel: drive-thru voice AI is the more complex engineering problem, the more capital-intensive sales motion, and the more competitive vendor field. Phone-order voice AI is the cleaner shipping path. The funding rounds are sized accordingly.

The PYMNTS piece on the voice AI funding surge names the 8x year-over-year funding number for the broader category. That number is real and the category is well-funded — but the number is not specific to QSR. It is the whole voice-AI vertical, which includes healthcare scheduling, insurance intake, and contact-center deflection. The QSR slice of it is meaningful but smaller than the headline implies. Operators reading “voice AI funding surges 8x” should not infer that there is 8x more QSR voice-AI capital deployed than a year ago. They should infer that there is more — meaningfully more — and that the vendor density at trade shows in 2026 is going to be higher than it was at the start of 2025.

The TenOneTen partner Eric Pakravan line that has been quoted everywhere this fall — “AI wasn’t ready. It is now.” — is the line that captures the venture-side conviction. It is also a line about model capability, not about per-store deployment readiness. Both can be true. Models can be ready and per-store operations can still take eighteen months to scale. The capital is pricing the former. Operators have to scope the latter.

The four-channel framework

The most useful frame I have built for myself reading vendor decks this fall is the four-channel one. Voice AI in QSR is not one product. It is, at minimum, four:

  1. Drive-thru speaker post. The Wendy’s FreshAI, White Castle SoundHound, Bojangles deployments. The hardest engineering problem — ambient noise, accent variation, menu-size scaling, hand-off mechanics under time pressure. The largest TAM if it works at scale. The longest tail to debug.
  2. Phone order for pickup. The Jersey Mike’s SoundHound deployment. The cleanest engineering problem. The use case I would deploy first if I ran a brand with strong off-premise pickup volume and no drive-thru.
  3. Kiosk / in-store voice. Less-covered in 2025, more-piloted than the press suggests. The channel where the customer is already in front of a screen — voice becomes a secondary modality rather than the primary one. The economics depend on whether voice meaningfully changes throughput at the kiosk vs. touch.
  4. Outbound / callback / loyalty. The channel almost no one talks about. AI placing or returning calls — reservation reminders, order status, loyalty engagement. The Loman-style vendors play here adjacent to phone-order. The lowest-risk channel because the customer’s expectations are lowest.

If you are an operator scoping voice AI in 2026, the right starting point is the channel that matters in your unit economics. The Jersey Mike’s choice was rational because phone is their off-premise channel. The Wendy’s choice was rational because the drive-thru is theirs. A regional pizza brand should probably look at phone-order first, drive-thru last. A coffee brand with a heavy drive-thru should look the other way. A fast-casual without a drive-thru should look at kiosk before drive-thru regardless of how much the drive-thru pitch dominates the trade-show floor.

This frame is also the frame the McDonald’s drive-thru story made me build in the first place. I am working on a forthcoming May piece on McDonald’s AI drive-thru that goes deeper into the speaker-post engineering problem (/blog/posts/the-mcdonalds-ai-drive-thru-from-apprente-to-google-a-five-year-case-study); for the channel-readiness comparison across vendors, an upcoming framework piece on voice-agent maturity (/blog/posts/the-voice-agent-maturity-curve) is the better next read.

Reading the deployment metrics operators are actually quoted

Here is the part that does not make it into the trade press, because it is mostly off-the-record: the metrics vendor reps quote in operator meetings have a structure, and the structure rewards careful reading.

The first number is almost always order-completion rate — the percentage of orders the AI takes start-to-finish without a human hand-off. This number is heavily a function of menu complexity. White Castle’s menu is short and SKU-stable, which produces a strong completion rate that does not generalize to a brand with thirty LTOs a year. The number is real for the brand it was measured at; it is misleading if you assume it transfers to yours.

The second number is handle time. The average duration of an AI-handled order. Operators read this as a labor-savings proxy, which it sometimes is and sometimes is not. Handle time is the bottleneck during a peak-rush minute, when the line is queued; it is irrelevant in a trough minute, when the staff are on standby. The savings calculation has to be queue-weighted, not average-weighted. Most decks present average-weighted.

The third number is guest-satisfaction proxy — usually a survey-based or complaint-based metric. This one is the most variable across brands because it is the most demographic-dependent. A regional brand with a guest base familiar with voice agents reads differently from a brand whose guest base associates voice automation with frustration. The number is real where measured; I would not extrapolate it to your demographic without testing.

The fourth, the one no vendor opens with: escalation rate to human. The inverse of order-completion in some decks, distinct in others — specifically, the rate at which the AI surfaces a question or confusion to a human-monitored channel. This is the operating-cost number that matters most, because it is the staffing assumption that drives the labor savings calculation. If escalation is twelve percent and you staffed for five, the labor savings is fiction.

I am going to stop the metric-by-metric here and name the practical instruction: ask for all four, by daypart, for the brand and footprint scope closest to yours. Vendors will resist the daypart cut. Insist. The peak-rush minute is the only minute that matters for the labor case, and the average smooths it out of view.

What the Vox AI 17x ROI number actually claims

I want to come back to the Vox AI 17x ROI claim because it is the number most likely to anchor 2026 vendor pitches and the most worth picking apart at the source. The framing in the BusinessWire and Restaurant Technology News coverage is that Vox’s autonomous voice AI produces a 17x return on labor savings, which is the kind of headline number that ends up in an investor deck and an operator pitch in the same week.

Read carefully: the 17x is a vendor-stated multiple, on labor savings, against an assumed labor cost baseline that the published coverage does not break out. The number could be defensible at one daypart and dilutive at another. It could be net of vendor fees and it could be gross. It could be against full-time-equivalent labor and it could be against the marginal hour. The published coverage does not say.

This is not a knock on Vox specifically. It is a knock on the press-release-to-pitch-deck-to-operator-decision pipeline. The 17x is the kind of number that will be quoted in a slide titled “voice AI ROI proof points” in 2026, alongside the Jersey Mike’s fifty-store number, alongside the Wendy’s FreshAI completion rate, alongside the White Castle hand-off rate — and the slide will conflate four different vendors at four different channels at four different brand scales. The honest answer for an operator scoping a 2026 pilot is that all four of those numbers are real where measured and none of them transfers to your footprint without a structured pilot.

The 90+ languages claim is the same shape. It is real as a model-capability statement. It is meaningful for a brand with a multilingual guest base. It is irrelevant for a brand whose guest base is overwhelmingly one language. The slide does not make the distinction; the operator has to.

What I would do in 2026 if I ran a regional brand

I want to close on the operator instruction set, because the regional-bench framing is only useful if it produces a different decision. If I ran a regional brand with a hundred to a thousand units and I was scoping voice AI for 2026, here is the order I would work in.

Start with the channel audit. Map your off-premise revenue. If phone-order pickup is more than fifteen percent of off-premise, the Jersey Mike’s-style deployment is the first one to scope, regardless of whether you have a drive-thru. If drive-thru is your largest channel by ticket count and your menu is under sixty SKUs, the White Castle / Bojangles-style speaker-post deployment is the one to scope second. If you are kiosk-heavy, the kiosk voice modality is third. Outbound is fourth.

Insist on the peak-rush metric cut. Every number a vendor quotes should come with the daypart it was measured at. Ask specifically for the Friday-Saturday dinner rush, or the Monday-morning breakfast push, depending on which is the peak for your footprint. Reject deck pages that present averages.

Read the escalation-rate disclosure carefully. If it is not in the deck, ask for it. If the vendor cannot produce it for a brand of comparable menu complexity to yours, treat the labor savings number as un-grounded.

Pilot at unit-count scale, not at the smallest footprint the vendor offers. The voice-AI failure modes that hurt unit economics — peak-rush escalation, menu-LTO drift, accent and demographic variance — surface at twenty units, not at five. The two-store pilot will produce a clean number; the twenty-store pilot will produce the real number.

And finally: ask which channel the vendor would not recommend for your footprint. The serious vendor will name a channel. The vendor that tells you their product fits all four is selling a category, not a product, and the category is not real yet. The Jersey Mike’s deployment is real. The Wendy’s deployment is real. The Taco Bell-Nvidia partnership is real. They are not the same product, they are not at the same scale, and they are not buying the same outcome. The operator who reads them that way is the operator who lands a 2026 voice-AI program that works.

The voice-AI bench in December 2025 is a thinner roster than the trade-show floor implies and a deeper roster than the McDonald’s-Wendy’s-Taco Bell-only narrative implies. The middle of the league is where the actual deployments are. Jersey Mike’s at fifty phone-order stores is the model. Read the press release one more time. The headline is the brand name. The substance is the scope.

— Priya covers operators for TableTransfers. Tips: [email protected].

Featured More

The Voice Agent Maturity Curve

mise

·

12 min read

The Four Margins of a Restaurant

mise

·

14 min read

The AI Premium in Hospitality M&A: Broker Story or Real Number?

the bottom line

·

9 min read

What the DoorDash/SevenRooms Deal Actually Buys

the bottom line

·

11 min read

Browse all 494 posts

Related posts

Darden's quiet AI strategy is buy-vs-build done right

the operator

·

19 min read

Darden's quiet AI strategy is buy-vs-build done right

Sweetgreen's Infinite Kitchen, in Public View: A Case Study

the operator

·

15 min read

Sweetgreen's Infinite Kitchen, in Public View: A Case Study

Sweetgreen's plan after selling the robot — the Sweet Growth Transformation reset

the operator

·

19 min read

Sweetgreen's plan after selling the robot — the Sweet Growth Transformation reset