Metric & Benchmark

P95 Latency

95th-percentile request latency — 5% of requests are slower than this.

Definition

P95 latency captures the slow tail that defines user-perceived performance. A high P95 reveals overload, GC pauses or cold-start issues invisible to the median. Many SLOs are set on P95 or P99 rather than P50.

Common use cases

  • SLO definition
  • Performance budgets
  • Alerting

Related terms

    P95 Latency — AI Glossary | Railwail