NOFire.ai
Resources/SRE reference/The four golden signals

The four golden signals

NOFire AI

What are the four golden signals and what should they alert on?

Latency, traffic, errors and saturation. The four exist to keep monitoring focused on what users experience rather than on what a machine is doing, and the discipline they impose is that a page should come from one of the four rather than from any metric that happens to be collectable.

At a glance

SignalWhat it measuresTypical formWhat it catches that the others do not
LatencyHow long a request takesp50, p95, p99, split by success and failureDegradation while everything still technically works
TrafficHow much demand the service is underRequests per second, or a domain equivalentA drop that means an upstream is broken, not that you are healthy
ErrorsThe rate of failed requestsProportion of non-success responsesExplicit failure, including failures the caller sees but you log as fine
SaturationHow close a constrained resource is to its limitQueue depth, memory headroom, connection pool useThe failure that is coming in ten minutes
UtilisationHow busy a resource is right nowCPU or disk percentageLittle on its own, and it is the most over-alerted metric in most estates

How it works

The four golden signals are a filter applied to a problem of abundance. A modern estate emits far more metrics than anyone can watch, and the instinct is to alert on the ones that are easy to collect, which is how teams end up paged for CPU utilisation at 3am on a service that is serving every request perfectly.

Latency, traffic, errors and saturation are chosen because each corresponds to something a user could notice, directly or imminently. Latency and errors are experienced now. Traffic is the context that makes the other two interpretable, since a doubling of errors matters differently at ten requests a second than at ten thousand. Saturation is the leading indicator, the one that tells you a failure is arriving rather than that it has arrived.

Saturation is also the one most often misread. It is not utilisation. A CPU at 90% may be entirely fine, while a connection pool at 90% with a queue forming behind it is minutes from failing. The useful saturation signal is the one measuring the most constrained resource for that particular service, which differs per service and requires knowing what actually limits it.

Two related methods narrow the same ground for different subjects. RED covers rate, errors and duration and fits request-driven services. USE covers utilisation, saturation and errors and fits resources such as disks, pools and queues. The golden signals span both, which is why they read as a superset rather than a rival framework.

The discipline the four impose is subtractive. They are a rule about what deserves to wake somebody rather than a list of things to add.

How it is measured

Measure latency at percentiles, never as a mean, and separate successful from failed requests before looking at either. Fast failures drag an average down, so a service returning errors quickly can appear to improve as it degrades. This single mistake accounts for a lot of monitoring that looked healthy through an outage.

Measure errors as a proportion rather than a count. Absolute error counts scale with traffic, so a fixed threshold pages during every busy period and stays quiet during a quiet-hours outage.

Choose the saturation metric per service by identifying what actually constrains it. Thread pools, connection pools, queue depth, memory headroom and disk throughput are the usual candidates. Alerting on generic CPU across everything is the default that the golden signals exist to replace.

Alert on symptoms from these four, and leave the rest as diagnostic context. A page should mean users are affected or imminently will be. Everything else belongs on a dashboard you consult during an investigation rather than in a rotation. How alert triage works and how it is measured covers what happens to the signals that do fire, and how service level objectives are set covers turning latency and error signals into a target with a consequence attached.

Common misconceptions

That you should alert on all four for every service. The four are candidates, not a checklist. Many services have no meaningful saturation signal worth paging on, and adding one to satisfy the framework produces noise.

That saturation means utilisation. They come apart exactly when it matters. A queue with headroom on a busy CPU is fine, a full queue on an idle CPU is an outage in progress.

That golden signals replace observability. They tell you something is wrong. They do not tell you why, and a team with four perfect signals and no ability to investigate has swapped one problem for another. The difference between observability and monitoring covers that boundary.

That a traffic drop is good news. Traffic falling to zero is the classic silent outage: no errors, excellent latency, nobody being served. Alerting on unexpected absence of traffic catches failures that error-rate alerting cannot see.

That more signals mean better coverage. Past a certain point, additional alerts reduce the attention paid to all of them. The four are a reduction, and treating them as an addition misses the argument entirely.

Frequently asked questions

What are the four golden signals?
Latency, how long requests take. Traffic, how much demand there is. Errors, the rate of failed requests. Saturation, how close a resource is to its limit. Together they describe a service from the outside in.
How do golden signals relate to RED and USE?
RED covers rate, errors and duration for request-driven services. USE covers utilisation, saturation and errors for resources. The golden signals span both, which is why they read as a superset rather than a competing method.
Should latency be measured as an average?
No. Averages hide the tail, which is where the users who are suffering live. Measure at percentiles, typically p50, p95 and p99, and alert on the higher ones.
Should successful and failed request latency be measured together?
No, and mixing them is a common trap. Fast failures pull the average down and can make a degrading service look like it is improving. Separate the two before reading either.

What to take from this

Use them to decide what pages you, not what you graph. Most estates collect far more than four signals and would be better off alerting on fewer.

Go deeper: how alert triage is measured

Back to SRE reference