gRPC Risk Scores Explanation

Executive Summary: Server Risk Score

The Server Risk Score provides management with a single, easy-to-compare measure of how reliably each server is handling customer orders. It combines order failures, server availability, and response speed using the most recent seven days of monitoring data.

A lower Risk Score is better. Servers are ranked from the lowest Risk Score to the highest.

What the Risk Score Measures

Service Targets

Measurement Target
Order error rate 1% or less
Timeout rate 1% or less
P95 order latency 3 seconds or less
Median order latency 2 seconds or less

How to Interpret the Score

A score of approximately 100 means the server's combined performance is at the target levels shown above. A score below 100 indicates better overall performance, while a score above 100 indicates increasing operational risk.

The score is not capped, so a severely unhealthy server may have a score significantly greater than 100.

Statistical Reliability

Because actual order tests occur less frequently than availability checks, the calculation applies a conservative statistical adjustment to the observed order error rate. This prevents a server from appearing completely risk-free simply because no failures happened to occur in a limited sample.

As more successful observations are collected, this uncertainty adjustment becomes smaller.

Measurement Window

All measurements use a rolling seven-day period containing up to:

The score is recalculated regularly. This allows the rankings to reflect recent performance while the seven-day window prevents isolated events from causing excessive short-term movement.

Management Use

The Risk Score is an operational decision aid, not the probability that a server will fail. It is designed to answer a practical management question:

Which servers currently offer the best combination of order reliability, availability, and response speed?