Metric & Benchmark
aka median latency

P50 Latency

Median per-request latency — half of requests are faster, half slower.

Definition

P50 (50th-percentile) latency is the median response time of an inference service. It describes a typical user's experience but hides tail behaviour. Pair it with P95/P99 to catch the slow requests that drive perceived slowness.

Common use cases

  • SLA monitoring
  • Capacity planning

Related terms

    P50 Latency — AI Glossary | Railwail