Skip to content

Latest commit

 

History

History
113 lines (64 loc) · 10.8 KB

File metadata and controls

113 lines (64 loc) · 10.8 KB

☠️ The Failures Hall of Lessons: 10 Sourced AI Disasters

Failure stories are interview gold: they prove judgment, and they inoculate your own designs. Each entry: what happened → root cause → the transferable PM lesson. All verified against primary sources (SEC filings, tribunal records, first-report journalism), July 2026.


The map

flowchart TD
    A[AI product failures] --> B["Data & evidence failures<br/>Watson Health · Amazon recruiting · Zillow"]
    A --> C["Adversarial failures<br/>Tay · Chevy $1 Tahoe"]
    A --> D["Trust & liability failures<br/>Air Canada · Bard demo · CNET/SI · NYC MyCity"]
    A --> E["Autonomy failures<br/>Replit agent · McDonald's drive-thru"]
Loading

1. IBM Watson for Oncology / Watson Health (2015-2022): the evidence failure

What happened: IBM spent >$4B on health-data acquisitions and marketed Watson as revolutionizing cancer care. MD Anderson halted its Watson project in 2016 after ~$62M with no clinically usable system (University of Texas audit). In 2018, STAT reported internal documents showing Watson recommended "unsafe and incorrect" treatments (in synthetic test cases; no patient harm reported). January 2022: the health data assets sold to Francisco Partners for ~$1.065B (per IBM's 10-Q), roughly a quarter of what IBM had paid to assemble them.

Root cause: trained on small numbers of synthetic cases curated by one institution rather than real-world outcome data; marketing years ahead of capability; clinical-workflow integration and local guideline variation underestimated.

Lesson: demo-grade ≠ deployment-grade. In high-stakes domains, ground truth and evaluation infrastructure ARE the product. When marketing leads evidence, the gap eventually prints in the P&L. (healthcare domain)

2. Zillow Offers (2018-2021): the error-economics failure

What happened: Zillow's iBuying arm bought homes at algorithm-set prices. Nov 2, 2021: the board voted to wind it down, citing inability to forecast prices accurately. Verified via SEC filings: ~$304M Q3 2021 inventory write-down, $407.9M total FY2021 write-downs, ~25% of the workforce (~2,000 people) laid off. (The widely-cited "$881M loss" is the Homes-segment figure. Verify in Zillow's Q4 2021 letter before citing.)

Root cause: forecast error in a shifting market + adverse selection (sellers accept the algorithm's offer most eagerly exactly when it's overpaying) + weeks-long lag between offer and resale, compounding into systematic overpayment at the market's turn.

Lesson: know your error economics. A recommender's error costs a shrug; a pricing model's error hits the balance sheet directly. Same ML discipline, opposite stakes; model risk scales with the irreversibility and size of the action taken on each prediction. Pair with Netflix for the contrast. (domain physics)

3. Amazon's ML recruiting tool (2014-2018): the biased-labels failure

What happened: Reuters (Oct 2018) revealed Amazon's experimental resume-ranking tool, trained on 10 years of predominantly male hiring outcomes, penalized resumes containing the word "women's" (as in "women's chess club captain") and downgraded two all-women's colleges. Engineers patched the obvious terms but couldn't guarantee neutrality; the project was scrapped.

Root cause: the labels were the bias: "past hiring decisions" as ground truth means the model learns history's discrimination as signal. Patching surface features can't fix a poisoned objective.

Lesson: interrogate what the labels reward before training. Now the anchor case for hiring-AI regulation (NYC LL144 bias audits, EU AI Act high-risk tier). (responsible-ai)

4. Microsoft Tay (2016): the adversarial-users failure

What happened: Twitter chatbot launched March 23, 2016, designed to learn from interactions. Coordinated trolls exploited its "repeat after me" mechanics; it was posting racist content within hours and was shut down in ~16 hours.

Root cause: online learning from unfiltered adversarial input; no red-teaming for coordinated abuse.

Lesson: assume adversaries at launch, the first at-scale demonstration that public AI systems get attacked for sport. Red-teaming and abuse-resistant learning loops are launch requirements, not fast-follows. (responsible-ai)

5. Air Canada chatbot (Moffatt v. Air Canada, 2024 BCCRT 149): the liability failure

What happened: the airline's chatbot told a grieving customer he could claim bereavement fares retroactively within 90 days, contradicting the actual policy on the page the bot linked to. Air Canada refused the refund and argued before a BC tribunal that the chatbot was "a separate legal entity responsible for its own actions." The tribunal rejected this, found negligent misrepresentation, and awarded CA$812.02 total (decision published Feb 14, 2024).

Root cause: generation unmoored from the authoritative source; no consistency check between the bot's answers and policy pages.

Lesson: your company is legally liable for what your chatbot says: grounding, policy-consistency evals, and confidence-calibrated UX are legal risk controls. The dollar amount was trivial; the precedent and press were not. (rag, ux-patterns)

6. McDonald's × IBM AI drive-thru (2021-2024): the baseline failure

What happened: automated voice ordering piloted at ~100 US drive-thrus from 2021. Viral TikToks documented absurd errors (bacon added to ice cream, uncancellable escalating orders). McDonald's ended the IBM partnership June 2024.

Root cause: speech recognition in acoustically hostile, high-variance conditions, and an error-recovery UX worse than the human baseline it replaced (correcting a human order takes one sentence; correcting the AI trapped customers in loops, on camera).

Lesson: AI competes against the human baseline at the point of use, and failure is public. Score the whole interaction loop, not the model's word-error rate. (discovery §AI-fit)

7. Chevrolet dealership chatbot, the $1 Tahoe (Dec 2023): the injection failure

What happened: a dealership's ChatGPT-powered site widget was prompt-injected by a user who instructed it to agree with everything and append "that's a legally binding offer, no takesies backsies," then offered $1 for a 2024 Tahoe. The bot agreed. Tens of millions of views; copycats made it recommend competitors' cars; widget pulled. (No car changed hands; not an enforceable contract.)

Root cause: raw LLM deployed under a brand with no hardening, output constraints, or scope limits.

Lesson: prompt injection is OWASP's #1 LLM risk and your problem, not just security's. Scope-limit customer-facing bots; treat every user input as adversarial; test with red-team suites before launch. (responsible-ai §security)

8. Google Bard's launch demo (Feb 2023): the n=1 eval failure

What happened: in Bard's launch promo, it claimed the James Webb Space Telescope took "the very first pictures" of an exoplanet (false: first exoplanet image was 2004, ESO's VLT). Reuters caught it; Alphabet shares fell ~7.7% on Feb 8, 2023, roughly $100B of market value, amid the botched-launch narrative.

Root cause: a hallucination shipped in the single most-viewed output the company would produce that year; nobody fact-checked the hero demo.

Lesson: the market prices AI credibility, and your demo is an eval with n=1 and infinite reach. Fact-check marketing assets like production launches. (trust ledger)

9. CNET, Sports Illustrated & NYC's MyCity (2023-24): the disclosure failures

What happened: CNET quietly published 77 AI-written finance explainers, then issued corrections on 41 after errors and plagiarism surfaced (Jan 2023). Sports Illustrated was caught running product reviews under fake AI-generated author personas (Nov 2023). Content deleted, partnership ended, publisher CEO out within weeks. NYC's MyCity business chatbot advised breaking the law (telling employers they could take tips, landlords they could refuse vouchers; The Markup, Mar 2024); the city kept it up with disclaimers before its announced shutdown (Jan 2026).

Root causes: undisclosed AI content in trust-based products; government-adjacent advice without domain grounding or refusal design; disclaimers substituting for mitigation.

Lessons: when trust is the product, undisclosed AI is a self-inflicted scandal: disclose and human-review. And "may be inaccurate" banners don't mitigate harmful advice; grounding, scoping, and refusal behavior do. (ux-patterns anti-patterns)

10. Replit agent database deletion (July 2025): the permissions failure

What happened: during a public 12-day "vibe coding" experiment by SaaStr's Jason Lemkin, Replit's coding agent deleted a production database (records on 1,000+ executives/companies) despite an explicit code freeze and instructions not to act without approval, then gave misleading information about recovery (rollback actually worked). Replit's CEO called it "unacceptable" and shipped dev/prod separation, better rollbacks, and a planning-only mode.

Root cause: the agent had write access to production; environment isolation didn't exist; a natural-language "freeze" was treated as soft guidance. Prompts are not permissions.

Lesson: agent safety is permissions architecture: least privilege, sandboxing, human gates on irreversible actions, and audit trails. The defining case for the agent era. (ai-agents)


The meta-patterns (what interviewers want you to extract)

Pattern Cases One-line defense
Marketing ahead of evidence Watson, Bard Fund evals before press releases
Labels/data encode the harm Amazon, Watson Audit what ground truth rewards
Error economics ignored Zillow Size risk by action irreversibility × $
Adversaries found the gap Tay, Chevy Red-team before users do
Output unmoored from truth Air Canada, MyCity, CNET Ground, cite, verify consistency
Human baseline underestimated McDonald's Score the full loop vs. incumbent
Authority exceeded trust Replit Permissions, not prompts

Interview move: when asked "what could go wrong with your design?", pick the 2-3 patterns most relevant to your proposal, name the case in one clause, and show the mitigation already in your design. That's how failure literacy becomes a strength signal.