Most infrastructure change happens overnight, when an outage would affect the fewest users. The cost is that the work runs when the people who know the system best are tired, vendor support is on its night roster and the engineer in the data hall may be working from a document written a fortnight earlier. A window succeeds or fails mostly on what happens before it opens and after it closes.
Why the work moves overnight
Contractual maintenance windows, trading hours and batch schedules usually fix the slot before the technical plan is written. That order matters. The window is a fixed box, and the work has to fit inside it with room left to undo it.
What must be ready before the window opens
Anything that can be done in daylight should be. Before the change starts, confirm:
- Site access approved for the named engineer, with the correct ID and any escort booked.
- Replacement parts on site, checked against the part number rather than the description, with firmware versions noted.
- Cables labelled at both ends and target ports recorded on the rack elevation.
- Configuration backups stored somewhere reachable even if the device being changed is the one that fails.
- A bridge call with the remote engineer, the change owner and an escalation contact for each vendor involved.
The common failure is an engineer arriving to find the rack does not match the drawing. A photo survey or pre-visit a few days earlier costs little and catches the mismatch in daylight. For work in a colocation hall, data centre smart hands support booked with the runbook attached, not just a ticket reference, gives the engineer time to read it.
Abort criteria and the rollback budget
Every plan needs a go/no-go time written down before work begins. Take the end of the window and subtract the measured rollback time plus a margin for the unexpected. That is the latest point to decide whether to continue. It belongs in the plan, not in arithmetic done at three in the morning.
Abort criteria need to be specific enough that nobody argues at that hour. “Service not restored” is too vague. “Storage paths not all active on both controllers by the go/no-go time” can be checked and reported. Rollback time should come from a rehearsal or the last time the same change was undone, not a guess. Restoring configuration, refitting the old part and waiting for an array to rebuild are the steps that tend to run long.
The handover that makes the morning uneventful
The change is not finished when the engineer leaves site. Before the first users log in, the day team needs a short written handover: what was changed, what was tested, what was not, which alarms were acknowledged and why, and who to call. Record the serial numbers of anything removed and where the old parts went, because a faulty part in a returns queue looks identical to a good spare.
The most useful line in any handover names what to watch. If a new switch should show a specific uplink at full speed, or a backup job should finish by mid-morning, say so. The day team can then confirm success in minutes rather than learning about a problem from a user ticket.
