
A model can win a gold medal at the International Math Olympiad and still misread the hands on an analog clock roughly half the time.
That's not a hypothetical. It's a real, documented result from today's top AI systems, and it captures something most businesses learn the hard way: being impressive in a demo and being reliable in production are two different things, and the gap between them is where many AI projects quietly die.
The demo goes great. The pilot gets funded. Then it hits real customers, real edge cases, and real messy data, and it starts getting things wrong in ways nobody tested for.
That's not just a model problem. It's a systems problem, and it's almost always solvable once you know what you're actually building toward.
Why prototypes fail
Gartner's research is blunt about this: at least 30 percent of generative AI projects get abandoned after the proof of concept, and the reasons are consistent: poor data quality, weak risk controls, costs that balloon past what anyone budgeted for, and a business case nobody could quite pin down.¹
That pattern shows up at a much larger scale in research from MIT's Project NANDA, which looked at hundreds of enterprise AI initiatives and found that 95 percent produced no measurable return, while only a small fraction actually delivered real value.² The common thread wasn't the model. It was that these systems were built to work once, in a demo, rather than engineered to keep working under real conditions.
A prototype has to impress someone for ten minutes. A reliable system has to be right consistently for months, including on inputs nobody thought to test.
The frontier is jagged, and that's the whole problem
Stanford's most recent AI Index put a name to something many people building with AI already suspected: today's models are what researchers call a "jagged frontier." They're superhuman at some things and surprisingly weak at others, sometimes within the same task.³
The clock reading example is almost funny until you translate it into a business context. AI agents tested on real computer based tasks, the kind involved in actual office work, still fail roughly one in three attempts.³ In a benchmark measuring factual accuracy, hallucination rates across today's leading models ranged from 22 percent to 94 percent depending on how the question was framed.³ Documented AI failures and incidents in the real world rose to 362 in 2025, up from 233 the year before.³
None of that means the technology isn't good enough to build on. It means capability alone was never the thing standing between a business and a reliable AI system.
Engineering was.
What reliability actually requires
Strip away the jargon, and a reliable AI system comes down to four things.
It's grounded in real information, not guesses.
A model with no access to your actual pricing, policies, and history will answer confidently and be wrong just as confidently. Reliable systems pull from real business data before they generate a response, instead of filling gaps from general knowledge.
It's tested before it ever reaches a customer.
Not once, at launch, but continuously against realistic scenarios, including the messy and unusual ones the demo never covered. If nobody is checking how the system handles edge cases, the first time it happens will be in front of a customer.
It's watched after it goes live.
Models drift. Data changes. What worked in month one can quietly start failing by month four, and if nobody's monitoring for that, the business is the one that finds out, usually from a customer instead of a dashboard.
Monitoring is what turns "it worked when we built it" into "it's still working."
It lives inside the actual workflow.
A reliable system that nobody uses because it sits in a separate tab isn't reliable. It's irrelevant. The most dependable AI systems are wired directly into the tools people already work in, so there is no extra step where context gets lost or adoption breaks down.
In practice, building this requires more than connecting an AI model to an application. Production systems need evaluation pipelines, retrieval layers, permission controls, observability, and feedback loops that allow the system to improve over time. The model is only one component. Reliability comes from the surrounding architecture that controls what information the system receives, what actions it can take, and how performance is measured after deployment.
Miss any one of these and the system might still work in the demo. It won't hold up in the business.
Where this is heading
The uncomfortable truth in Stanford's latest research is that the industry's own tracking of reliability is behind its tracking of raw capability. Nearly every major AI lab publishes benchmark scores for how capable their models are. Far fewer publish anything close to the same rigor for how often those models are safe, accurate, or trustworthy in practice.³
Businesses that wait for the industry to solve this for them will be waiting a while.
That gap is exactly why reliability is becoming the actual competitive line, not model access. Two companies can license the same model. Only one of them will have built the data pipeline, testing discipline, and monitoring required to make it dependable once real customers start using it.
As the basic model layer becomes increasingly accessible, that engineering discipline is what will separate AI systems businesses can trust from the ones that quietly get switched off after the first bad answer.
Our take
At Jurisa, we treat reliability as an engineering problem, not a hope.
Before we care what a system can do in a demo, we care whether it's grounded in the right data, tested against the scenarios that actually happen in your business, monitored once it's live, and built into the workflow people already use.
That's the difference between an AI project that gets a round of applause in a meeting and one that's still quietly doing its job a year later, unnoticed, because it never gave anyone a reason to turn it off.
Key takeaways
• Capability and reliability are not the same thing. Today's most advanced models are excellent at some tasks and surprisingly unreliable at others.
• At least 30 percent of generative AI projects get abandoned after the proof of concept, usually over data quality, cost, or unclear business value.
• A reliable system needs four things: real business data, ongoing testing, live monitoring, and integration into the actual workflow.
• Most AI failures happen after launch, not before, because nobody was watching for drift.
• As models converge in capability, engineering discipline, not model access, becomes the real competitive advantage.
Sources
- Gartner, Inc. (2024, July 29). Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept by End of 2025. Gartner. https://www.gartner.com/en/newsroom/press-releases/2024-07-29-gartner-predicts-30-percent-of-generative-ai-projects-will-be-abandoned-after-proof-of-concept-by-end-of-2025
- Challapally, A., Pease, C., Raskar, R., & Chari, P. (2025, July). The GenAI Divide: State of AI in Business 2025. MIT Project NANDA. https://www.aigl.blog/state-of-ai-in-business-2025/
- Stanford Institute for Human Centered Artificial Intelligence (HAI). (2026, April). The 2026 AI Index Report: Technical Performance and Responsible AI chapters. https://hai.stanford.edu/ai-index/2026-ai-index-report