For most of computing history, performance was an engineering problem you could reason about at your desk. Profile the code, count the database calls, size the hardware. If a system was slow, the cause was usually something you could hold in your head: an unindexed query, a synchronous call that should have been async, a loop that ran one order of magnitude too many times. You could draw the system on a whiteboard, and the whiteboard would basically be correct.
That era is over. In cloud-native systems, performance has become an emergent property. It arises from the interaction of many moving parts, most of which you don’t directly control, and that shift changes how we have to engineer for it. The core argument is simple to state and hard to internalize: in a modern-day distributed system, you cannot optimize what you cannot observe.
Performance is no longer about your code
Consider what actually determines the speed of a cloud-native application today. It’s spread across dozens of loosely coupled microservices. Those services share network fabric and physical infrastructure with other tenants on a Kubernetes cluster. They’re scheduled into containers by an orchestrator making placement decisions you never see, filtering and scoring nodes behind the scenes before binding a workload to whichever one wins, and they scale up and down automatically in response to demand that arrives in unpredictable bursts.
In that environment, application behavior is driven less by algorithmic efficiency than by distributed orchestration, virtualization overhead, and workload variability. A minor shift in traffic distribution or container scheduling can introduce a bottleneck that’s genuinely difficult to diagnose after the fact, because the conditions that caused it may no longer exist by the time you go looking. Inter-service calls that look trivial in isolation — a few milliseconds here, a few there — accumulate into measurable delay across a long service chain, the same way a handful of small taxes compound into a large bill.
And in a multi-tenant cloud, a competing workload can degrade your responsiveness in ways you can’t reproduce on demand. This is well known enough in cloud architecture circles to have its own name: the noisy neighbor problem, where one tenant’s activity on shared infrastructure quietly degrades another’s performance even though nothing in your own stack changed. The performance you get is the sum of all these interactions, not just a property of any single component you wrote.
Why traditional monitoring falls short
The instinct is to throw monitoring at the problem. But there’s a meaningful difference between monitoring cloud-native systems vs. actually exposing where fault could lie. And that is exactly where that difference matters. Cindy Sridharan made a case earlier, describing traditional monitoring as failure-centric, built around a list of known ways things break, whereas observability is meant to support open-ended debugging of failures nobody predicted.
That distinction matters in practice because averages are exactly the wrong lens for distributed systems. Average latency can sit directly on top of a system that is intermittently unbearable for a meaningful slice of users, and because averages smooth over the short-term, however, it is the sharp events that hurt the most. This isn’t a new observation; in the now-classic paper on large-scale systems, Google engineers Jeff Dean and Luiz André Barroso described how tail latency at scale becomes the dominant experience for users of any sufficiently large distributed system, because a request that fans out to hundreds of components only needs one slow component to become a slow request overall. As a system grows, the odds that at least one component is having a bad moment approach certainty. Averages don’t see that. Observability, done well, means being able to ask why a specific request was slow and reconstruct the answer from execution traces, logs, and fine-grained runtime telemetry across every service that request touched. The goal isn’t a prettier dashboard. It’s the ability to follow one request through the system and see exactly where the time went.
Designing for runtime, not just design-time
The deeper lesson is about when we assess performance, not just how. For years, performance testing was a design-time activity: run a load test before release, and if the numbers look good, ship. Cloud-native systems break that assumption, because so much of what affects performance — scheduling decisions, autoscaling behavior, neighbor interference, resource contention under real traffic shapes — only exists at runtime and only under real workloads. Offline benchmarks against fixed, synthetic traffic patterns simply don’t capture it, because production traffic doesn’t arrive in the tidy, evenly distributed way a load-testing script assumes.
So, performance assurance has to move closer to production: continuous telemetry from running services, compared against known-good baselines, so that degradation is caught as it emerges rather than discovered after users feel it and complain. There is a growing body of research pushing in exactly this direction, on observability-driven and adaptive autoscaling approaches that use runtime signals rather than fixed thresholds to decide when and how to scale. The results across that research vary by system and workload, and the specific percentage improvements aren’t quoted as a single universal number. What is consistent is the underlying finding: systems that scale and adapt based on what’s actually happening at runtime tend to outperform ones that rely on static, pre-configured rules, particularly on tail latency, which is the metric static rules are worst at protecting. The specific numbers in any one study matter less than what produced the improvement: measuring the system as it actually runs, not as it was assumed it would run when someone configured the autoscaler months earlier.
The takeaway
Cloud-native architecture gave us tremendous flexibility, and it took away our ability to reason about performance. The right response isn’t to fight that trade. It’s to accept it and instrument accordingly.
Treat observability as a design requirement, not an afterthought bolted on after an incident. Measure at the percentiles, not the averages, because the averages are hiding the exact failures that matter most to real users. And assess performance where it actually happens: at runtime, under real load, continuously, rather than only in a pre-release load test that can’t anticipate what production traffic will actually look like.
In systems this complex, visibility is the prerequisite for everything else. You can’t optimize what you can’t observe, and you can’t observe it unless you designed it, from the start, to be seen.