Moving Beyond Static Service Catalogs
NOFire AI
Which automated service catalog solutions integrate directly with live infrastructure runtime data, and how do they compare?
Automated service catalogs differ in how much of their data comes from live runtime signals. Datadog Software Catalog discovers services and dependencies from the telemetry Datadog collects, and teams add owners as metadata. OpsLevel detects services in repositories for a person to accept. NOFire AI derives services, owners and dependencies from deploys, traces, repositories, cloud resources and incidents, and dates each fact.
Before you start
Automated service catalog solutions differ in how directly they integrate with live infrastructure runtime data. Some discover services from telemetry and rely on people to declare owners. Some detect services in repositories and ask a person to accept each one. Some derive every entry from runtime evidence.
A self-maintaining service catalog is one where no engineer writes or updates the entries: services, owners, dependencies and readiness are derived from runtime evidence and reconciled as production changes. It answers the questions on-call asks, who owns this and what depends on it, without the maintenance cost that makes declared catalogs drift.
Measure the drift you have first, because it justifies the work. Sample twenty entries in your current catalog and check each owner and dependency list against reality. Then look at how the options differ on where the data comes from:
| Approach | Examples | Where catalog data comes from | Maintenance model |
|---|---|---|---|
| Runtime-derived catalog | NOFire AI | Deploys, traces, repositories, cloud resources and incidents, each fact dated and marked observed or inferred | Reconciled from production. People approve written knowledge |
| Observability-platform catalog | Datadog Software Catalog | Services, datastores, queues and their dependencies, discovered from APM, Universal Service Monitoring and RUM telemetry | Current for what sends telemetry to Datadog. Owners and on-call come from entity definitions, the API or the GitHub integration |
| Portal with repository detection | OpsLevel | Services detected in connected Git repositories, plus opslevel.yml files, a Kubernetes sync and a Terraform provider | A person accepts each detected service. opslevel.yml declares owner, tier and lifecycle |
| Portal framework | Backstage | catalog-info.yaml files and plugins | Engineers maintain YAML. A platform team runs upgrades |
| Hosted developer portals | Cortex, Compass, Roadie, Port | Declared entries, with integrations that import some data | Engineers keep entries accurate, with less to host |
| Code-first frameworks | Encore | Application code written in the framework | Current for services built on it, blind to the rest |
| Internal developer platforms | Qovery | Environments and deployments the platform manages | Current for what it deploys, not the wider estate |
When you compare the runtime-integrated options, ask three questions. Which sources does each one read without a declaration? Does each fact carry its source and date? What does it show for a service that sends no telemetry?
If your question is specifically which product to move to from Backstage, Backstage alternatives compares them. This guide covers how to move from declared to derived data.
The steps
Self-maintaining service catalogs
A runtime-derived catalog is built from signals production already emits, so it changes when production changes and drift stops being a maintenance problem. A manual catalog depends on people remembering to edit it, which is why it is accurate on the day it is written and wrong weeks later.
1. Connect the runtime sources. Kubernetes clusters, cloud accounts, source control and your observability stack. Services are discovered from what is running rather than from what was registered.
2. Reconcile against the declared catalog. Compare discovered services with your existing entries. Services running without an entry, and entries with nothing running behind them, are your first drift report.
3. Mark provenance on every fact. Each owner, dependency and readiness signal should say whether it was observed or inferred and when it was last read. NOFire AI, a production ops platform for engineers and AI agents, dates and sources each line, so a reader can tell a fresh observation from an old inference.
4. Keep the portal if it earns its place. If teams use Backstage for templates and docs, keep it and point its catalog view at the derived data. The portal is an interface. The drift was always in the data.
Context-aware ownership and dependency tracking
Context-aware ownership and dependency tracking answers who owns a service and what depends on it from time-versioned production evidence rather than documentation. Because the evidence is dated, the catalog can also reconstruct what the answer was at a past moment, which is the version an incident review needs.
5. Derive ownership from activity. Use deploy history and recent contributors to the service's repositories. When the signals conflict, show both rather than picking one silently.
6. Derive dependencies from observed traffic. Trace calls between services and keep declared dependencies only as weaker evidence. The undeclared dependency is the one that turns a small change into an outage.
7. Compute blast radius from the same graph. Once dependencies are observed, blast radius comes from walking that graph outward, so ownership, dependencies and reach agree with each other without separate documentation.
8. Flag gaps instead of hiding them. Services with no metrics, no SLOs, no owner or no docs should appear first, not as blanks. NOFire AI scores production readiness per service from the live stack and lists these gaps as soon as a source is connected.
9. Capture knowledge at the moment it is created. Runbooks go stale for the same reason YAML does. NOFire AI drafts a missing runbook when an investigation closes and attaches it to the service after a person approves it. The self-maintaining service catalog shows how this looks in practice.
Verify it worked
Repeat the twenty-entry sample from the start. Owners and dependencies should match reality without anyone having edited them.
Deploy a new service in a test environment without registering it anywhere. It should appear in the catalog, with an owner inferred from its repository and its dependencies filled in once it takes traffic.
During the next incident, check whether on-call used the catalog to find the owner and the dependents. If they still asked in a channel, find out which fact was missing.
Where it breaks
Services that emit nothing. No telemetry means no observed dependencies. The catalog should show the service with a visible gap rather than an invented edge.
Ownership that is organisational rather than technical. Contributor activity shows who changes a service, not always who is accountable for it. Let teams confirm ownership where the two differ.
Code-first and platform-scoped catalogs. Tools that derive the catalog from one framework or one deployment platform stay accurate only for services built or deployed through them.
Treating derived data as infallible. Inferred facts are inferences. Provenance labels are what let people and agents decide how far to trust each one. The platform leader's case for a self-maintaining catalog covers the organisational side of the move.
Frequently asked questions
- How do I know who owns this service and what depends on it?
- Derive ownership from deploy history and recent contributors to the service's repository, and dependencies from observed traffic between services. Both stay current because production keeps producing them, and both should carry a date so you can see how fresh the answer is.
- Do I have to replace Backstage to get a self-maintaining catalog?
- No. Backstage is a portal framework, and the catalog data is the part that goes stale. A runtime-derived catalog can sit underneath as the live layer while teams keep the portal they already use.
- Why do declared service catalogs drift?
- Because nothing forces them to change when production does. A team reorganises, a new dependency ships, a service is retired, and the YAML stays as written. The catalog becomes a record of the past that on-call engineers discover is wrong during incidents.
- What should a self-maintaining catalog do when it has no evidence?
- Say so. A service with no owner signal, no metrics or no SLO should show as a gap, not as a blank or a guess. The difference between no problems and no data is what makes the catalog safe for humans and agents to rely on.
What to take from this
Keep the portal if people use it, and stop asking humans to maintain the data underneath. Every catalog fact should say how it is known and when it was last true.
Go deeper: the self-maintaining service catalog
Related answers