Executive Summary
- Paul Tanner, Technical Support Director at Opengear, spills the beans on what really happens when critical infrastructure fails at 4 am.
- Before the fault can be properly diagnosed, the engineer on duty must establish what remains reachable, which management paths are still available and whether there is a dependable route back into the affected site.
- For data centre operators, the broader lesson is that resilience depends on more than just maintaining power, cooling, and connectivity during normal operations.
At 4am, an alert shows that a remote data centre or edge facility has disappeared from the management environment, even though the site itself may still be operating. The switches are not responding through the usual route, management access is unavailable and the on-call engineer initially suspects an ISP problem.
Further investigation shows that the upstream service is working normally. Instead, a firmware task carried out during the maintenance window has unexpectedly disrupted the management path. Valuable time is spent examining the wrong area before the team identifies what has actually happened.
That first judgement can shape the entire recovery effort. Before the fault can be properly diagnosed, the engineer on duty must establish what remains reachable, which management paths are still available and whether there is a dependable route back into the affected site.
In those first minutes, the priority is to establish the facts. Engineers need to know which systems continue to respond, whether the problem is confined to one location or affects a broader section of the data centre and edge infrastructure estate, and whether the production network alone has failed or management connectivity has disappeared too. Answering those questions reduces the risk of treating the most visible symptom as the underlying cause.
Assumptions are especially risky during an outage. A configuration that has worked reliably for years may begin to fail because something has changed elsewhere in the environment. A device may look like the problem even when the real issue sits upstream. A provider outage may not yet have been acknowledged by the provider. In every case, the engineer needs enough access and enough information to properly investigate and work through the issue.
Listen before diagnosing
Communication is key to the technical response. During a live incident, network and data centre operations teams often have to deal with pressure from several parts of the organisation at once. Senior leaders expect regular updates, while service owners, site managers and facilities teams need clarity about the operational impact.
Engineers still need time to investigate without being repeatedly pulled away from the work. A clear process helps keep those pressures under control by setting out what is known, what is under investigation and what will happen next.
The more difficult scenario is when the network has failed and the team has lost its standard way back in. Power, cooling and the equipment itself may remain operational, but if the primary management path is unavailable, engineers can still be locked out of the infrastructure they need to recover. At that stage, the outage is about more than just the original failure, it also depends on whether the team can access the systems needed to fix it.
This is where out-of-band management becomes especially important for organisations operating distributed data centres, colocation environments, edge facilities and network infrastructure. A separate management plane, often supported by secure cellular connectivity, provides operations teams with an independent, secure route into critical infrastructure.
Across distributed environments, that independent path may determine whether recovery can be completed remotely or requires an engineer to travel to the facility. Without it the team might have to depend on remote-hands support, wait for confirmation from a provider or dispatch a specialist to the affected facility. With secure out-of-band access, engineers can reach routers, switches and firewalls even when the main network is offline. That gives them a practical point of entry without forcing risky workarounds during an already pressured incident.
That point of entry must also be a secure one. When normal access disappears, teams may feel pressure to restore connectivity through whatever route is fastest. Temporary routes and improvised access methods may solve one immediate problem while creating another. Network teams need a recovery method that restores operational control without weakening the environment in the process.
Certain incidents will, of course, inevitably take more time to resolve. A device might lose its connection to the network management plane while a firmware update is taking place. A remote facility could become unreachable because of an upstream provider issue that has not yet been reported. These situations cannot always be solved through a standard checklist, because the visible failure may only be one part of a much wider problem.
For engineers responsible for data centre and network infrastructure, the key requirements are reliable access and a clear route for escalation when more specialist input is needed. Response time remains vital, particularly in the early stage of a major outage, but speed only helps if it leads the team towards the right decision.
Fast answers still require human judgement
That need for speed explains why AI-assisted tools are increasingly used for ticket triage and knowledge searches across network support. When applied effectively, they can direct an engineer to the appropriate area of investigation within minutes rather than hours.
The practical limitation is confidence without accuracy – an AI system can produce a well-formed, plausible-sounding answer that simply does not reflect the device configuration or specific type of failure involved. During an outage, that distinction matters because an incorrect answer delivered with confidence is still incorrect.
That is why AI is most useful when it is treated as a way to speed up the search rather than replace technical diagnosis. It can operate much like autocomplete, identifying useful documentation, suggesting a relevant procedure and limiting the time spent searching for information. The engineer must still compare the output with the actual environment before taking action. On its own, the technology cannot be treated as an authoritative source.
The danger is that a rapid response can feel like the right one, especially when the pressure is on. A suggested answer may sound entirely reasonable, while failing to match the actual configuration or failure pattern. Validation is therefore essential. The engineer must retain a clear understanding of the infrastructure and be able to judge whether the suggested path makes sense.
For data centre operators, the wider lesson is that resilience depends on more than maintaining power, cooling and connectivity during normal operations. Teams must also be able to regain visibility and control when the primary management path disappears. Patterns identified during overnight incidents should inform infrastructure design, remote-access planning and escalation procedures across the wider estate.
Secure out-of-band management gives engineers an independent route back into critical systems, while experience enables them to interpret what they find and select the correct response. Together, reliable access, clear communication and informed human judgement help data centre teams restore control when normal operations suddenly disappear.



