The middle of a sample
The median divides sorted observations into two halves. It describes the middle of a sample but does not say how bad a slow request can be. A complete presentation also includes the observation window, number of samples and a definition of which outcomes entered the calculation.
The tail needs evidence
A p95 describes the slower end under a stated quantile method. With few observations, it can depend almost entirely on one request. We withhold p95 until the configured minimum sample count is met. That threshold is a display rule, not a guarantee of statistical precision.
Do not hide failures
A timeout has no completed latency to sort, but still matters to reliability. Report failure counts separately instead of treating failures as zero latency or removing them silently. Preserve invalid records and exclusion reasons so an aggregate can be reconstructed and reviewed later.
Ask three questions
Do the model settings match? Do the windows overlap? Do both groups have enough observations? An apparently faster model may have a more favorable sample. The initial composite score remains inactive while baseline and uncertainty policy are evaluated; no percentage should imply precision the data cannot support.