What It Is
If p99 latency is 10 milliseconds, 99 of every 100 requests finish in 10 milliseconds or less, while the slowest one percent take longer. It sits alongside p50 (the median) and p95. Percentiles describe the distribution of response times more truthfully than a simple average.
Key Points
- Beats averages: averages can hide serious problems.
- Tail latency: often dominates the experience in distributed systems.
- Common causes: garbage collection pauses, lock contention, retries, and overloaded queues.
- Improvements: efficient memory use, timeouts, and load testing.
Why It Matters
A fast average can still mean poor performance for a small group of users. In systems where one user action triggers many internal calls, the slowest calls often decide the overall experience, so tail latency deserves its own targets.
How ClearLeaff Applies It
We design core systems in Rust to sustain sub-10ms p99 latency under real production load, tracking percentile metrics continuously.