The four golden signals
NOFire AI
What are the four golden signals and what should they alert on?
Latency, traffic, errors and saturation. The four exist to keep monitoring focused on what users experience rather than on what a machine is doing, and the discipline they impose is that a page should come from one of the four rather than from any metric that happens to be collectable.
At a glance
| Signal | What it measures | Typical form | What it catches that the others do not |
|---|---|---|---|
| Latency | How long a request takes | p50, p95, p99, split by success and failure | Degradation while everything still technically works |
| Traffic | How much demand the service is under | Requests per second, or a domain equivalent | A drop that means an upstream is broken, not that you are healthy |
| Errors | The rate of failed requests | Proportion of non-success responses | Explicit failure, including failures the caller sees but you log as fine |
| Saturation | How close a constrained resource is to its limit | Queue depth, memory headroom, connection pool use | The failure that is coming in ten minutes |
| Utilisation | How busy a resource is right now | CPU or disk percentage | Little on its own, and it is the most over-alerted metric in most estates |
How it works
The four golden signals are a filter applied to a problem of abundance. A modern estate emits far more metrics than anyone can watch, and the instinct is to alert on the ones that are easy to collect, which is how teams end up paged for CPU utilisation at 3am on a service that is serving every request perfectly.
Latency, traffic, errors and saturation are chosen because each corresponds to something a user could notice, directly or imminently. Latency and errors are experienced now. Traffic is the context that makes the other two interpretable, since a doubling of errors matters differently at ten requests a second than at ten thousand. Saturation is the leading indicator, the one that tells you a failure is arriving rather than that it has arrived.
Saturation is also the one most often misread. It is not utilisation. A CPU at 90% may be entirely fine, while a connection pool at 90% with a queue forming behind it is minutes from failing. The useful saturation signal is the one measuring the most constrained resource for that particular service, which differs per service and requires knowing what actually limits it.
Two related methods narrow the same ground for different subjects. RED covers rate, errors and duration and fits request-driven services. USE covers utilisation, saturation and errors and fits resources such as disks, pools and queues. The golden signals span both, which is why they read as a superset rather than a rival framework.
The discipline the four impose is subtractive. They are a rule about what deserves to wake somebody rather than a list of things to add.
How it is measured
Measure latency at percentiles, never as a mean, and separate successful from failed requests before looking at either. Fast failures drag an average down, so a service returning errors quickly can appear to improve as it degrades. This single mistake accounts for a lot of monitoring that looked healthy through an outage.
Measure errors as a proportion rather than a count. Absolute error counts scale with traffic, so a fixed threshold pages during every busy period and stays quiet during a quiet-hours outage.
Choose the saturation metric per service by identifying what actually constrains it. Thread pools, connection pools, queue depth, memory headroom and disk throughput are the usual candidates. Alerting on generic CPU across everything is the default that the golden signals exist to replace.
Alert on symptoms from these four, and leave the rest as diagnostic context. A page should mean users are affected or imminently will be. Everything else belongs on a dashboard you consult during an investigation rather than in a rotation. How alert triage works and how it is measured covers what happens to the signals that do fire, and how service level objectives are set covers turning latency and error signals into a target with a consequence attached.
Common misconceptions
That you should alert on all four for every service. The four are candidates, not a checklist. Many services have no meaningful saturation signal worth paging on, and adding one to satisfy the framework produces noise.
That saturation means utilisation. They come apart exactly when it matters. A queue with headroom on a busy CPU is fine, a full queue on an idle CPU is an outage in progress.
That golden signals replace observability. They tell you something is wrong. They do not tell you why, and a team with four perfect signals and no ability to investigate has swapped one problem for another. The difference between observability and monitoring covers that boundary.
That a traffic drop is good news. Traffic falling to zero is the classic silent outage: no errors, excellent latency, nobody being served. Alerting on unexpected absence of traffic catches failures that error-rate alerting cannot see.
That more signals mean better coverage. Past a certain point, additional alerts reduce the attention paid to all of them. The four are a reduction, and treating them as an addition misses the argument entirely.
Frequently asked questions
- What are the four golden signals?
- Latency, how long requests take. Traffic, how much demand there is. Errors, the rate of failed requests. Saturation, how close a resource is to its limit. Together they describe a service from the outside in.
- How do golden signals relate to RED and USE?
- RED covers rate, errors and duration for request-driven services. USE covers utilisation, saturation and errors for resources. The golden signals span both, which is why they read as a superset rather than a competing method.
- Should latency be measured as an average?
- No. Averages hide the tail, which is where the users who are suffering live. Measure at percentiles, typically p50, p95 and p99, and alert on the higher ones.
- Should successful and failed request latency be measured together?
- No, and mixing them is a common trap. Fast failures pull the average down and can make a degrading service look like it is improving. Separate the two before reading either.
What to take from this
Use them to decide what pages you, not what you graph. Most estates collect far more than four signals and would be better off alerting on fewer.
Go deeper: how alert triage is measured
Related answers