Uber’s “Silent Failure”: Why Your Data Needs Observability, Not Just Monitoring
It wasn’t a hacker attack. It wasn’t a server meltdown.
It was a simple calculation error.
For a period of time, Uber’s internal accounting system was deducting taxes from driver earnings before taking their commission, instead of after. It was a minor logic error buried deep in a data pipeline.
The cost? Uber reportedly underpaid drivers by millions of dollars in total (approx. $900 per driver in some instances) and had to issue massive repayments.
But the scariest part wasn’t the money. It was the silence.
The pipelines didn’t crash. The dashboards didn’t turn red. The database reported “Success.” For months, the data was technically “flowing,” but it was logically wrong.
This is the “Silent Failure”—the single greatest threat to a modern enterprise. And before they built their world-class data platform, it was Uber’s recurring nightmare.
The 100-Petabyte Blind Spot
At its peak confusion, Uber sat on over 100 Petabytes of data with no unified way to see if it was healthy. They faced two specific, revenue-draining problems that will sound familiar to any Malaysian executive:
- The “Surge” Blackout: Uber’s pricing algorithm depends on real-time supply and demand data. If a data pipeline lagged by just 5 minutes, the system would think there were “zero” cars available or “zero” demand. The result? The surge multiplier would default to 1.0. Drivers wouldn’t log on, riders couldn’t get cars, and Uber lost massive revenue—all because a pipeline was “stale” but not “broken.”
- The “Zombie” Data: Teams were hoarding data “just in case.” With no accountability, storage costs skyrocketed, and analysts wasted 30% of their time just figuring out which dataset was the “real” one.
Uber realized they couldn’t just manage data anymore. They had to observe it.
The Turnaround: Inside “DataCentral”
To stop the bleeding, Uber didn’t just buy a tool; they mobilized an army of Silicon Valley’s best engineers to build a custom platform called DataCentral.
They moved from “monitoring” (is the server up?) to “Observability” (is the data accurate?). Here is the “Secret Sauce” of how they fixed it:

Figure 1: Clio
1. “Clio”: The Fingerprint of Health
Uber built a tool called Clio that “fingerprinted” every single query. It didn’t just check if a report ran; it checked how it ran.
- Before: “The report generated successfully.” (But it contained 0 rows).
- After: “This report usually contains 10,000 rows. Today it has 50. ALERT.“
- The Lesson: Anomaly detection is more valuable than simple status checks.

Figure 2: Contactless in Action
2. “Contactless”: Automated Root Cause Analysis
When a data job failed, engineers used to spend hours digging through logs. Uber built Contactless, an AI-driven system that analyzed error logs and instantly told the engineer: “This failed because the Finance Team changed a column name in the source database.”
- The Lesson: You cannot afford to spend 4 hours diagnosing a 5-minute fix.
3. “Chargeback”: The Culture Shock
This was the most controversial fix. Uber started “billing” internal teams for the data they stored. If Marketing wanted to keep 5 years of raw logs, they saw the cost on their monthly P&L.
- The Result: Teams instantly deleted Petabytes of junk data. The “Single Source of Truth” emerged because no one could afford to keep the lies.
The Reality Check: You Are Not Uber
Uber solved this problem with a massive budget and a dedicated team of 50+ Data Reliability Engineers.
Most companies in Malaysia do not have that luxury.
You cannot pause your business for two years to build a proprietary “DataCentral.” But you also cannot afford the $900-per-driver mistake. You are likely facing the same “Silent Failures” right now—inventory numbers that don’t match the warehouse, or financial reports that require three days of manual Excel glue to fix.
You need Uber’s results without Uber’s R&D budget.
The Solution: Lestar.ai
At Mandrill Tech, we built Lestar.ai to be the “DataCentral” for the enterprise that wants to move fast.
We took the complex engineering concepts from giants like Uber—Observability, Anomaly Detection, and Unified Lineage—and packaged them into a platform that deploys in days, not years.
- Instead of “Clio”: Use Lestar CEO 360. It tracks your financial and operational metrics in real-time. If revenue dips or a metric flatlines (the “Surge Blackout” scenario), you get an alert instantly—not at the end of the month.
- Instead of “Contactless”: Use our Generative AI Chatbot. You don’t need to parse logs. You can simply ask your data: “Why is the Q3 margin report showing a discrepancy?” Lestar’s AI traces the lineage and highlights the anomaly for you.
- Instead of “Chargeback”: Use Lestar’s Centralized Repository. By unifying your data into one “Single Source of Truth,” you naturally eliminate the expensive, confused data silos that plagued Uber.
Don’t Wait for the $900 Mistake
Uber learned the hard way that Data Observability isn’t an IT upgrade; it’s an insurance policy for your revenue.
Your business runs on data. If you can’t see it, you can’t trust it. And if you can’t trust it, you can’t lead.
Ready to turn the lights on? Stop fearing silent failures. Let’s build your Single Source of Truth today with Lestar.
Source: Technical details and engineering insights for this article were adapted from Uber’s official case study, DataCentral: Uber’s Observability and Chargeback Platform.



