Metric & Benchmark
aka median latency
P50 Latency
Median per-request latency — half of requests are faster, half slower.
Definition
P50 (50th-percentile) latency is the median response time of an inference service. It describes a typical user's experience but hides tail behaviour. Pair it with P95/P99 to catch the slow requests that drive perceived slowness.
Common use cases
- SLA monitoring
- Capacity planning