The McDonald's AI Drive-Thru, From Apprente to Google: A Five-Year Case Study
A four-acquisition arc, a high-profile failure, a quiet pivot, and a 2026 restart. McDonald's voice-ordering programme is the most-documented voice-AI deployment in QSR — and the cleanest mirror anyone running a hospitality AI project has.
I cover the hotel F&B beat for The Operator. The reason a QSR case study is sitting in my column this week is straightforward: the McDonald’s voice-AI programme is the deepest, most public, most failure-mode-rich deployment of agentic voice technology in hospitality, and every hotel restaurant group considering a phone-agent or in-room voice deployment is — whether they know it or not — implicitly betting that their version goes better than McDonald’s did. I want to lay out what actually happened, on the record, with citations, and then say what I think the case study tells operators about the next twelve to twenty-four months.
Methodology
Every numeric claim, executive quote, and date in this piece is sourced to a public document — McDonald’s investor communications, IBM press releases, Google Cloud’s press centre, the joint McDonald’s/IBM statement of June 2024, the major trade press (CNBC, Restaurant Business, Restaurant Dive, CIO Dive, QSR Magazine, CNN), and the relevant earnings-call coverage. I have not used any non-public information for this case study; I have used my own commentary to connect the dots. Where I am offering analysis rather than reported fact, I say so. Where the public record is contradictory, I flag the contradiction. Where a number is widely circulated but I cannot trace it to a primary source, I drop it. Two figures I had originally drafted — a specific “per-restaurant capex” number and a specific franchisee accuracy survey result — were removed at this verification step because the primary source I had assumed turned out to be a downstream blog post quoting another downstream blog post.
Why this case matters
McDonald’s runs roughly 43,000 restaurants globally. When the company tests technology in a hundred drive-thrus, every competitor watches; when it abandons a piece of technology, every competitor adjusts. The voice-AI programme matters for a second reason easy to miss: it is the canonical real-world test of conversational voice agents in a noisy, accent-diverse, customer-facing setting with no ability to pause and clarify. Every other voice deployment in hospitality — hotel front-desk agents, restaurant phone-reservation agents, room-service AI — is a strictly easier problem. If voice agents work at McDonald’s, they work everywhere. If they don’t, the smaller deployments need a much narrower scope to survive. This is the central argument of the Voice Agent Maturity Curve, and McDonald’s is the case study at the top of it.
The IBM partnership: a four-acquisition arc, 2019 to 2024
The IBM partnership did not begin in 2021. It began in 2019, with two acquisitions that shaped everything that followed.
In April 2019, McDonald’s acquired Dynamic Yield, a personalisation and decision-engine company, for a price widely reported at roughly $300 million (QSR Magazine, Restaurant Dive). Dynamic Yield powered the dynamic menu boards at McDonald’s drive-thrus — the technology that surfaces “people who ordered this also ordered this” in real time based on weather, time of day, and order history. It was deployed across more than 8,000 US restaurants (McDonald’s corporate).
Then, in September 2019, McDonald’s acquired Apprente, a Mountain View voice-AI startup founded in 2017 by Itamar Arel and team. Apprente’s pitch was a neuroscience-informed voice platform for multilingual, multi-accent, multi-item conversational ordering (McDonald’s press release, TechCrunch). The Apprente team became the founding members of McD Tech Labs, a Silicon Valley-based internal group inside McDonald’s Global Technology.
McD Tech Labs ran the voice-AI programme for two years inside McDonald’s. Then, in October 2021, McDonald’s and IBM announced a strategic partnership in which IBM would acquire McD Tech Labs outright and continue developing the Automated Order Taking (AOT) technology under IBM’s Cloud & Cognitive Software division (IBM newsroom, Voicebot.ai). The transaction closed in December 2021. The framing at the time was that IBM’s experience with enterprise conversational AI — Watson’s lineage, fundamentally — would let McDonald’s scale AOT across languages, dialects, and menu variations faster than the internal team could on its own.
By mid-2023, the rollout had reached just over 100 US restaurants. Then it stopped expanding.
In June 2024, McDonald’s emailed franchisees a system message saying it would remove the AOT technology from the more than 100 test restaurants by July 26, 2024 (Restaurant Business Online, CNBC, Restaurant Dive). The company’s public framing was diplomatic: IBM remained “a trusted partner,” the test had given McDonald’s “the confidence that a voice-ordering solution for drive-thru will be part of our restaurants’ future,” and the company would evaluate other options and make an informed decision on a long-term voice solution by year-end 2024.
The franchisee-facing reality, reconstructed from contemporaneous trade press, was less diplomatic. According to Restaurant Business Online, franchisees who attended McDonald’s worldwide convention reported that the live AOT demonstrations got every order wrong. Accuracy was reported in the low-to-mid 80% range against a McDonald’s internal target of 95%+ before broad rollout. Operating costs were high enough that franchisees — who own and operate roughly 95% of McDonald’s US restaurants — pushed back on the economics of broader deployment.
That is the bones of the IBM arc: Apprente acquired in 2019, McD Tech Labs spun up, McD Tech Labs sold to IBM in 2021, 100-restaurant pilot through 2023, viral failures and franchisee resistance, partnership wound down in mid-2024.
The viral failure mode
The thing the public will remember is bacon ice cream.
Between late 2023 and mid-2024, a stream of TikTok videos showed the AOT system misbehaving in instructive ways. One video, widely reported, showed an order of vanilla ice cream and a bottle of water turning into a tray of ice cream, ketchup sachets, and butter (Fast Company, Restaurant Business Online). Another showed the system adding roughly £166 of chicken nuggets to an order — over a hundred nuggets — and the customer struggling to get the order cancelled (techinformed). And the canonical one: bacon, on ice cream, unprompted.
The TikTok corpus is the most-watched real-world evidence anyone has of how a production voice-AI system fails in front of customers. The failure modes cluster into four categories, and they recur in every voice-AI deployment I have looked at since:
Cross-talk and channel bleed. The system picking up audio intended for a different lane, a passenger in the back seat, or a radio inside the car. Structurally more dangerous than the classic accent failure because the customer has no signal that the AI is hearing the wrong thing.
Over-confirmation. When the system was uncertain, it added items rather than asking for clarification. A tuning decision — “be helpful, default to inclusion” — that is catastrophic in a transactional setting where every added item is a customer-facing error.
Edge-case customisation. “No onions, extra pickles, light cheese, sauce on the side” is a string of negation and modifier operations that requires the voice model to reason structurally about the order schema. The system handled the common path; edge cases were the public failures.
Confrontation recovery. Several viral videos showed customers trying to correct an error mid-order and the system either ignoring the correction, adding the correction as a new item, or escalating the wrong order. The recovery loop was visibly weaker than the happy path.
If you are running a hotel phone-agent or restaurant voice-reservation system today, those four failure modes are your map.
The Google Cloud pivot
The pivot started before the IBM wind-down was public.
In December 2023 — six months before the IBM partnership ended — McDonald’s and Google Cloud announced a multi-year global partnership to deploy Google Cloud technology across “thousands” of McDonald’s restaurants worldwide (Google Cloud press release, McDonald’s corporate). The architecture is worth understanding: McDonald’s would deploy Google Distributed Cloud — Google’s edge-computing offering — to each restaurant so that cloud-based applications and on-site software could run locally with low latency. A dedicated Google Cloud team would sit in Chicago, near McDonald’s Speedee Labs innovation centre. The press release flagged generative AI as a core workload, alongside equipment-monitoring and mobile-app personalisation.
The announcement did not name the drive-thru voice agent as a Google workload. It did not need to. The architecture — edge compute, on-site model inference, generative-AI partnership at the global parent — is the architecture you would build if you wanted to migrate the voice agent off IBM’s stack and onto Google’s foundation models, while preserving the option to run inference locally inside each restaurant for latency and resilience.
The two-year sequence reads as a deliberate pivot. Google Cloud announced in December 2023. IBM wind-down announced in June 2024, with the framing that McDonald’s was evaluating alternatives. Through 2024 and 2025, McDonald’s gave very little public detail on what voice-AI experimentation it was doing, while the Google Cloud relationship grew quietly through equipment-monitoring, kiosk improvements, and mobile-app generative-AI work (Restaurant Technology News).
In late 2025 and early 2026, the public-facing AI-voice messaging restarted. Industry coverage describes McDonald’s working with multiple partners — Google included — on a more scalable voice-ordering rollout targeted at 2026, US markets first, built on LLM voice agents fine-tuned on McDonald’s order history. CEO Chris Kempczinski has stated repeatedly on earnings calls that robotics in restaurants is good for headlines but “not practical in the vast majority of restaurants” — the company’s stated thesis is software-and-edge, not physical automation (Restaurant Business Online).
In March 2026, McDonald’s announced an expanded five-year partnership with Capgemini for engineering and deployment services across its 40,000+ restaurants, alongside the Google relationship (Restaurant Technology News). The vendor stack today: Google Cloud as the foundation model and cloud partner, Capgemini as systems integrator, Dynamic Yield (now Mastercard-owned) for menu-board personalisation, and the voice-agent vendor unconfirmed publicly but heavily implied to be a Google-stack product.
Results so far
The honest answer to “did this work” is: not yet, but the company is still trying.
The IBM phase did not produce a deployable voice-ordering system. Accuracy never reached the 95% bar the company had set. Per-restaurant economics were apparently uncompetitive against the human-operated baseline, though I do not have a primary source for a specific dollar figure I trust enough to cite. The pilot ended after roughly two and a half years and 100 restaurants — small enough to be a pilot, large enough that the failure was visible to investors and franchisees.
The Google phase is too early to score. The infrastructure deployment is real and measurable — thousands of restaurants getting Google Distributed Cloud, Capgemini engineering work, edge-compute hardware. The voice agent on top of that infrastructure is still pre-rollout as of early 2026. The 2026 messaging is forward-looking. I will be watching the Q3 2026 and Q4 2026 earnings calls for the first hard numbers on voice-agent accuracy and per-restaurant cost.
The interesting near-term tell is that McDonald’s, after writing down the McD Tech Labs investment to IBM and the broader voice-agent programme, is still publicly committed to voice as a long-term direction. That is not a company that has concluded voice-AI in QSR is impossible. It is a company that has concluded the IBM-era implementation was wrong and a different stack will work.
Failure-mode analysis: what went wrong
Four things, in my reading.
One: the architecture inherited Apprente’s 2017 assumptions. Apprente was a 2017 voice-AI company; the 2017 voice-AI stack was fundamentally pre-transformer-LLM. The McD Tech Labs roadmap, and then the IBM roadmap, carried that architecture forward through 2024 with incremental improvements rather than a clean rebuild on top of modern foundation models. By the time the system shipped to a hundred restaurants, the entire voice-AI industry had moved underneath it. Wendy’s, working with Google from a later starting point, reportedly cut drive-thru order time by 22 seconds with a Google-stack agent (Restaurant Dive). Architecture lock-in is the single biggest risk in any voice-AI deployment with a long build cycle.
Two: the operating-cost economics did not pencil for franchisees. McDonald’s US system is franchised; operators pay for restaurant-level equipment and connectivity. Any voice-AI rollout has to clear a labour-displacement-versus-installed-cost hurdle at the per-restaurant level, and by all available evidence the IBM system did not. This is the Four Margins point made concrete: contribution margin only improves if the AI either reduces labour cost more than it costs to install, or increases throughput enough to grow contribution dollars per square foot. The IBM system did neither at the per-restaurant level franchisees needed.
Three: brand-safety cost of viral failures was higher than expected operating gain. Every viral TikTok of a McDonald’s AI failure was a real reputational cost to a brand whose positioning is consistency. The expected value was negative not because the average customer interaction was bad — most were fine — but because the tail of bad interactions was visible at internet scale. The asymmetry between average and tail performance, for a customer-facing voice agent at a brand this big, is enormous.
Four: the IBM accuracy ceiling was structurally hard to break through. Going from 85% to 95% accuracy in a noisy, accent-diverse, customisation-heavy voice setting is not a 10% improvement; it is a multi-order-of-magnitude reduction in error rate. The IBM stack, built on the Apprente foundation, appears not to have had a path to that ceiling within an acceptable cost envelope. Modern LLM-based voice agents have a different cost curve, which is why the Google-stack restart is potentially viable where the IBM stack was not.
What this means for wider hospitality
Voice agents in customer-facing transactional settings are harder than people think, and “harder than people think” has a precise meaning: average performance is fine, tail performance is the killer, and your brand pays the cost of the tail. Any operator considering a voice deployment should model the distribution of outcomes, not the mean.
The right architecture is increasingly edge-compute plus foundation-model. McDonald’s bet on Google Distributed Cloud is the cleanest articulation of this thesis in the industry. Cloud-only is too slow and brittle; on-premise-only is too hard to update and too capital-intensive. Edge plus foundation model is the emerging right answer.
Per-unit economics matter more than the technology. A voice-AI vendor that ships a 92% accuracy system at $30K of per-restaurant capex is a different product than one at $200K. The technology gets the press; per-unit economics decide what gets deployed.
QSR is two years ahead of hotel F&B on voice-AI deployment. Hotels reading this case study are getting a free preview of failure modes their own programmes will hit. The Sweetgreen and Chipotle operator case studies, and the hotel operator case study from earlier in this series, all show the same arc in slower motion: pilot, surprise failure mode, pivot, restart. McDonald’s just ran the whole arc on the front page.
Operator takeaways
Six things from the McDonald’s case study, for any operator looking at a voice deployment in the next twelve months.
Pilot small, but pilot with a kill switch on a calendar. McDonald’s piloted at 100 restaurants for over two years. Too long. Either the pilot is working at twelve months and you scale, or it is not and you pivot. Pre-commit to the decision date.
Model the tail, not the mean. Design your evaluation around the worst 5% of customer interactions. A voice agent with 92% accuracy and a 1% rate of catastrophic failures is a different product than one with 92% accuracy and a 0.05% catastrophic-failure rate.
Get per-unit economics right before you scale. Build the per-restaurant or per-property P&L explicitly — capex, opex, training, partial-replacement-of-labour. If it doesn’t pencil at the unit level, it will not survive the operator community.
Architecture on a modern foundation. The Apprente-to-IBM lineage carried 2017 architecture into 2024 production. Don’t repeat that. Pick a foundation-model-native vendor.
Build the recovery loop first. What happens when the customer says “no, that’s wrong” is the single most important conversation the voice agent has. Test it more carefully than the happy path.
Be honest about brand-safety cost. Viral failures are a strategic cost, not a marketing problem. Price them in.
Close
The McDonald’s voice-AI programme is the longest, most public, most expensive learning loop in QSR. From internal R&D in 2019, to IBM partnership in 2021, to public failure in 2024, to Google-stack restart in 2026 — and along the way it produced more public information about how voice agents fail in real restaurants than any other deployment in the industry.
The honest reading is that the failure of the IBM phase is evidence about the IBM phase, not evidence about whether voice AI works in QSR. The technology is changing fast enough that a 2021-architecture system failing in 2024 tells you almost nothing about whether a 2026-architecture system will work in 2027. What it does tell you, in detail, is the failure modes any voice-AI deployment in customer-facing transactional hospitality needs to engineer around.
For hotel F&B operators specifically — my beat — the McDonald’s case is the cleanest possible warning shot. The voice-AI deployments coming to hotel restaurants over the next eighteen months will be smaller, narrower, less exposed to viral failure — but they will hit the same four failure modes, the same per-unit economics question, the same architecture-lock-in risk. The voice-agent vendors in our Desk Reviews series, read alongside this case study, are the working set of evidence any hospitality operator should be sitting with before signing a voice-AI contract this year.
— Naomi covers the Hotel F&B beat for The Operator. Tips: [email protected].
The Voice Agent Maturity Curve
mise
·12 min read
The Four Margins of a Restaurant
mise
·14 min read
The AI Premium in Hospitality M&A: Broker Story or Real Number?
the bottom line
·9 min read
What the DoorDash/SevenRooms Deal Actually Buys
the bottom line
·11 min read