The change that passed review on a two year old map

A change went to the board with a clean impact statement. One application server, one window, two dependent services, no customer impact expected. The window opened, the server went down, and a claims intake service on nobody's list stopped taking work for four hours. The dependency map behind that statement had not been touched in twenty-six months.

Almost every CMDB is partly wrong, and almost everyone near one knows it. That is not the interesting part. The damage shows up in what people do next. Once a change manager has been burned by a stale map, they stop reading it and ring application owners directly, and a ten minute review becomes an afternoon of calls.

Decay speeds up from there. A CMDB people consult gets corrected, because someone notices a wrong relationship at the moment it costs them something. A CMDB nobody opens collects errors quietly, and by the time anyone returns the gap has grown far enough that fixing it becomes a project instead of a correction.

Completeness only matters for the classes you decide about

Completeness asks whether the things that exist in your estate are represented at all. Total completeness is the wrong target. Nobody needs every switch port and every installed font, and a programme that chases them runs out of budget before it reaches the classes anyone consults.

Completeness for the classes you make decisions about is achievable. Write the decisions down first. Which services does this change touch. Who gets called at two in the morning when this database is unreachable. Which hosts run an operating system nobody patches. In most ServiceNow ITSM practices those questions name fewer classes than people expect.

A narrow set can also be measured against something outside the CMDB, which is the only measurement worth much. Compare the server class against what the virtualisation platform reports, or the application list against what finance pays for. A count drawn from the CMDB alone tells you how many records exist.

Correctness depends on who fills the attribute

Correctness asks whether the attributes on a record describe the thing as it stands today. The answer depends almost entirely on one property of each attribute, which is whether a machine populates it or a person does.

Attributes fed by discovery or by a platform integration drift for a while and get corrected on the next run. Attributes a person typed once during onboarding start rotting the day they were entered. Support group, business criticality, environment and owning application are the usual offenders, and they are the attributes people lean on hardest.

So decide, attribute by attribute, whether a machine can fill it. If it can, take the manual field away or make it read only, so nobody fights the discovery source. If it cannot, give it a named owner and a review interval. An attribute with neither an automatic source nor an owner should come off the form, because a confident wrong value does more harm than a blank one.

Relationships carry the value and take the decay

A configuration item with no relationships answers no interesting question. It confirms that a server exists, which nobody needed to ask. The questions people bring to a CMDB are about what depends on what, and those answers sit in the relationships rather than the attributes.

Dependency mapping is the most valuable part and the least maintained. Application to service, service to database, database to host, host to location. Each link takes effort to establish, and none of that effort repeats when the estate changes underneath. A team that mapped for six months in year one and stopped has a map of year one.

How much you can automate varies by release and by which discovery sources are in use, so check what your instance can infer before keeping a relationship by hand. Traffic and agent based discovery propose lower level dependencies well enough. The step up to a business service stays a human judgement, so it needs an owner and a review point, usually a change approval board that moves and notices when the map stops matching reality.

Duplicates begin with the identification rules

Duplicates do more damage than any other health problem, and they start the same way. Two sources describe one server. One reports it by short host name, the other by fully qualified name. Nothing tells the platform these are the same machine, so you hold two records, and every decision taken against either is partly wrong.

The mechanism that decides is the set of identifying attributes and the reconciliation rules built on them. For each class, identification says which attributes, in which order, make an incoming record the same thing as one you already hold. Serial number first, then name plus domain, then something weaker. Reconciliation then decides which source may write which attribute.

Those rules are the single most consequential configuration in this area, and they usually sit untouched at whatever the platform delivered, because nobody reads them until duplicates surface in a report. Changing them later is expensive. Records created under the old rules have to be merged by hand or by script, and merging loses history.

The test is deliberately mechanical. Take a class fed by two sources, pull twenty records, and check whether each physical thing appears exactly once. Then hunt the softer duplicates, where one source truncates a name and the other does not. Count them, read the identification rules, and work out which attribute would have caught them.

Decide which source wins for which class

Without a written precedence decision, the last writer wins. Two integrations both claim the operating system attribute on the server class, they run on different schedules, and the value swings between them all week. Nobody raises it, because nobody watches a field that keeps changing to a plausible value.

Precedence belongs to architecture rather than to whoever configured the most recent integration. Deciding that the virtualisation platform owns host and memory, that the software agent owns installed software and that a human process owns business criticality shapes what every consumer can rely on. Write it down per class and per attribute, with a date and a name on it.

The same argument turns up whenever two applications share one CMDB. We have covered running CSM and ITSM against a single CMDB and the fight over who owns the CMDB, and both land on precedence. A class with two owners and no precedence rule has no owner.

Measure health against a decision, not a score

An overall health score is easy to publish and hard to act on. A number that moved from 68 to 71 tells nobody what to do on Monday. Measure against the decisions instead. What share of changes to production services arrive with a confirmed dependency list. How many records in the classes you page on have an owner who still works here.

Staleness deserves a place among the signals. A record whose attributes have not moved in four months is suspect however correct it looks, because either the thing stopped changing or the source stopped reporting. Report last updated age by class and by source, and investigate the source first, since one silent feed leaves hundreds of records stale.

Retirement is the half most programmes skip. A CMDB that never removes anything fills with things that no longer exist, and those records resurface in impact analysis and in licence counts. Decommissioning needs a step that retires the record, and a thing missing from several consecutive discovery runs should raise a task. How many runs depends on your schedule and coverage.

Pick the two classes your next major change depends on and write this down for each before the board meets. An afternoon of it costs less than another quarter of impact statements nobody believes, and the deadline moved closer once agents started reading the CMDB and acting on what they found.

  • The decisions this class is used for, and the people who make them.
  • Which source is authoritative for each attribute anyone reads, and which sources may only create records.
  • The identifying attributes, in the order the identification rules apply them.
  • The relationship types that have to be present before the class is useful, and who confirms them.
  • The age at which a record here stops being trusted without a check.
  • What triggers retirement, who may retire a record, and where the retired ones go.