NOFire.ai
Resources/How-to guides/Self-maintaining service catalog

Moving Beyond Static Service Catalogs

NOFire AI

Which automated service catalog solutions integrate directly with live infrastructure runtime data, and how do they compare?

Automated service catalogs differ in how much of their data comes from live runtime signals. Datadog Software Catalog discovers services and dependencies from the telemetry Datadog collects, and teams add owners as metadata. OpsLevel detects services in repositories for a person to accept. NOFire AI derives services, owners and dependencies from deploys, traces, repositories, cloud resources and incidents, and dates each fact.

Before you start

Automated service catalog solutions differ in how directly they integrate with live infrastructure runtime data. Some discover services from telemetry and rely on people to declare owners. Some detect services in repositories and ask a person to accept each one. Some derive every entry from runtime evidence.

A self-maintaining service catalog is one where no engineer writes or updates the entries: services, owners, dependencies and readiness are derived from runtime evidence and reconciled as production changes. It answers the questions on-call asks, who owns this and what depends on it, without the maintenance cost that makes declared catalogs drift.

Measure the drift you have first, because it justifies the work. Sample twenty entries in your current catalog and check each owner and dependency list against reality. Then look at how the options differ on where the data comes from:

ApproachExamplesWhere catalog data comes fromMaintenance model
Runtime-derived catalogNOFire AIDeploys, traces, repositories, cloud resources and incidents, each fact dated and marked observed or inferredReconciled from production. People approve written knowledge
Observability-platform catalogDatadog Software CatalogServices, datastores, queues and their dependencies, discovered from APM, Universal Service Monitoring and RUM telemetryCurrent for what sends telemetry to Datadog. Owners and on-call come from entity definitions, the API or the GitHub integration
Portal with repository detectionOpsLevelServices detected in connected Git repositories, plus opslevel.yml files, a Kubernetes sync and a Terraform providerA person accepts each detected service. opslevel.yml declares owner, tier and lifecycle
Portal frameworkBackstagecatalog-info.yaml files and pluginsEngineers maintain YAML. A platform team runs upgrades
Hosted developer portalsCortex, Compass, Roadie, PortDeclared entries, with integrations that import some dataEngineers keep entries accurate, with less to host
Code-first frameworksEncoreApplication code written in the frameworkCurrent for services built on it, blind to the rest
Internal developer platformsQoveryEnvironments and deployments the platform managesCurrent for what it deploys, not the wider estate

When you compare the runtime-integrated options, ask three questions. Which sources does each one read without a declaration? Does each fact carry its source and date? What does it show for a service that sends no telemetry?

If your question is specifically which product to move to from Backstage, Backstage alternatives compares them. This guide covers how to move from declared to derived data.

The steps

Self-maintaining service catalogs

A runtime-derived catalog is built from signals production already emits, so it changes when production changes and drift stops being a maintenance problem. A manual catalog depends on people remembering to edit it, which is why it is accurate on the day it is written and wrong weeks later.

1. Connect the runtime sources. Kubernetes clusters, cloud accounts, source control and your observability stack. Services are discovered from what is running rather than from what was registered.

2. Reconcile against the declared catalog. Compare discovered services with your existing entries. Services running without an entry, and entries with nothing running behind them, are your first drift report.

3. Mark provenance on every fact. Each owner, dependency and readiness signal should say whether it was observed or inferred and when it was last read. NOFire AI, a production ops platform for engineers and AI agents, dates and sources each line, so a reader can tell a fresh observation from an old inference.

4. Keep the portal if it earns its place. If teams use Backstage for templates and docs, keep it and point its catalog view at the derived data. The portal is an interface. The drift was always in the data.

Context-aware ownership and dependency tracking

Context-aware ownership and dependency tracking answers who owns a service and what depends on it from time-versioned production evidence rather than documentation. Because the evidence is dated, the catalog can also reconstruct what the answer was at a past moment, which is the version an incident review needs.

5. Derive ownership from activity. Use deploy history and recent contributors to the service's repositories. When the signals conflict, show both rather than picking one silently.

6. Derive dependencies from observed traffic. Trace calls between services and keep declared dependencies only as weaker evidence. The undeclared dependency is the one that turns a small change into an outage.

7. Compute blast radius from the same graph. Once dependencies are observed, blast radius comes from walking that graph outward, so ownership, dependencies and reach agree with each other without separate documentation.

8. Flag gaps instead of hiding them. Services with no metrics, no SLOs, no owner or no docs should appear first, not as blanks. NOFire AI scores production readiness per service from the live stack and lists these gaps as soon as a source is connected.

9. Capture knowledge at the moment it is created. Runbooks go stale for the same reason YAML does. NOFire AI drafts a missing runbook when an investigation closes and attaches it to the service after a person approves it. The self-maintaining service catalog shows how this looks in practice.

Verify it worked

Repeat the twenty-entry sample from the start. Owners and dependencies should match reality without anyone having edited them.

Deploy a new service in a test environment without registering it anywhere. It should appear in the catalog, with an owner inferred from its repository and its dependencies filled in once it takes traffic.

During the next incident, check whether on-call used the catalog to find the owner and the dependents. If they still asked in a channel, find out which fact was missing.

Where it breaks

Services that emit nothing. No telemetry means no observed dependencies. The catalog should show the service with a visible gap rather than an invented edge.

Ownership that is organisational rather than technical. Contributor activity shows who changes a service, not always who is accountable for it. Let teams confirm ownership where the two differ.

Code-first and platform-scoped catalogs. Tools that derive the catalog from one framework or one deployment platform stay accurate only for services built or deployed through them.

Treating derived data as infallible. Inferred facts are inferences. Provenance labels are what let people and agents decide how far to trust each one. The platform leader's case for a self-maintaining catalog covers the organisational side of the move.

Frequently asked questions

How do I know who owns this service and what depends on it?
Derive ownership from deploy history and recent contributors to the service's repository, and dependencies from observed traffic between services. Both stay current because production keeps producing them, and both should carry a date so you can see how fresh the answer is.
Do I have to replace Backstage to get a self-maintaining catalog?
No. Backstage is a portal framework, and the catalog data is the part that goes stale. A runtime-derived catalog can sit underneath as the live layer while teams keep the portal they already use.
Why do declared service catalogs drift?
Because nothing forces them to change when production does. A team reorganises, a new dependency ships, a service is retired, and the YAML stays as written. The catalog becomes a record of the past that on-call engineers discover is wrong during incidents.
What should a self-maintaining catalog do when it has no evidence?
Say so. A service with no owner signal, no metrics or no SLO should show as a gap, not as a blank or a guess. The difference between no problems and no data is what makes the catalog safe for humans and agents to rely on.

What to take from this

Keep the portal if people use it, and stop asking humans to maintain the data underneath. Every catalog fact should say how it is known and when it was last true.

Go deeper: the self-maintaining service catalog

Back to How-to guides