Field report: secondary circuit validation
Onsite attendance to observe and validate an intermittent secondary WAN connection. The production network was fully operational on arrival. No work to the primary circuit, virtualisation estate or internal switching was included in the requested scope.
Assignment record
| SITE | Regional office and self-contained communications room |
|---|---|
| PLATFORM | Firewall HA pair / VMware ESXi / Microsoft Hyper-V / Active Directory |
| REPORTED FAULT | Intermittent backup circuit during automatic failover |
| ARRIVAL STATE | Primary WAN, LAN, server estate and all user services operational |
| CHANGE WINDOW | Not supplied; informal low-usage period agreed over lunch |
08:42 — Arrival and baseline
The site was operating normally when the engineer arrived. Users were authenticated, internet access was available and no active alarms were visible on the production firewall. Both VMware and Hyper-V management consoles reported their hosts online. The local Active Directory domain controllers, DNS, DHCP, file services and line-of-business virtual machines were responding within their normal tolerances.
The sole outstanding concern was the backup WAN interface. Historic firewall events showed brief carrier loss followed by successful recovery, although the timestamps did not align precisely with the carrier portal. As the secondary connection was not carrying traffic, a controlled reset was considered non-impacting.
09:07 — Secondary circuit reset
The backup interface was administratively reset and the carrier termination equipment was restarted. Link state returned, dropped, and returned again without completing a stable failover health check. The behaviour was consistent with the reported intermittent condition and therefore confirmed that attendance had been necessary.
A second reset produced a different sequence of indicator lights but no materially improved result. The primary connection continued to carry all production traffic throughout this stage.
09:51 — Substitute firewall assessment
A replacement firewall was retrieved from the engineer’s van to separate the carrier service from the installed firewall hardware. The unit was unused stock of the nearest available model. On boot it required online activation before exposing the full diagnostic interface. It also reported that its installed firmware could not import the production configuration revision.
The backup service could not provide the reliable upstream connection required to activate or upgrade the replacement unit. The remaining option was to present the known-working primary connection to it briefly. To minimise disruption, this was deferred until the office lunch period, when the majority of users would be away from their desks.
The production environment remained fully functional. The proposed lunchtime activity was expected to involve one short cable transfer, activation of the substitute unit and immediate restoration of the original arrangement.
12:04 — Lunchtime validation
After confirming that desk activity had reduced, the primary carrier handoff was moved from the production firewall to the replacement unit. The substitute firewall obtained a connection, completed activation and began downloading its required firmware. External reachability was confirmed briefly from its diagnostic page.
During the download, the primary connection began alternating between connected and disconnected states on the replacement firewall. The firmware transfer restarted twice and the management session became unavailable for approximately forty seconds. As this introduced uncertainty into a previously known-good circuit, the validation was stopped and the primary handoff was returned to the installed firewall.
12:19 — Primary service does not return
When reconnected to the production firewall, the primary interface reported physical carrier but did not restore its upstream session. Reseating the handoff caused the carrier state to disappear entirely. The carrier equipment still displayed its normal service indication, while the firewall alternated between negotiation and timeout.
The backup circuit was then selected to preserve external access, but it reproduced the intermittent condition for which the visit had originally been arranged. At this point both WAN paths were unavailable, although local servers, authentication and workstation access remained operational.
12:33 — High-availability state change
The active firewall was restarted to clear a suspected interface-driver state. Its peer correctly attempted to assume the active role. The peer’s configuration, however, contained an older internal interface map and identified two current VLAN interfaces as pending objects. This difference had not been visible while the appliance remained passive.
The virtual gateway addresses moved between appliances several times while the pair attempted to agree an active member. Internal clients began losing access to DNS and application networks. The firewalls eventually settled with the original unit active, but the internal core retained address-resolution entries associated with both appliances.
12:58 — Core adjacency investigation
A refresh of the core switch stack was undertaken to eliminate the stale gateway and virtual MAC state. One stack member returned normally. The second member declined to rejoin because its firmware revision differed from the elected master by one maintenance release. The mismatch was pre-existing and had not affected service before the refresh.
Links connected to the unavailable member were redistributed across the surviving stack member. Several aggregated uplinks came back with only half their expected members. The VMware distributed switch interpreted the changing uplink state as repeated path failure, while the Hyper-V switch-embedded teams reported inconsistent reachability across VLANs.
13:21 — Storage and cluster impact
The site’s storage paths shared the affected switching fabric. Loss and restoration of individual trunks caused the iSCSI sessions used by the VMware hosts to enter an all-paths-down condition. Hyper-V cluster shared volumes moved into redirected access before the cluster witness became unreachable.
Virtual machines initially remained visible in their respective management consoles but stopped responding to network checks. Automatic recovery attempts increased storage and cluster traffic on the remaining links. Two VMware hosts then isolated themselves, and the Hyper-V nodes paused a number of guests to protect disk consistency.
13:47 — Directory services unavailable
One Active Directory domain controller was hosted on VMware and the other on Hyper-V. The VMware controller was inaccessible with its datastore path unavailable; the Hyper-V controller had paused with its cluster volume. DNS registration, Kerberos authentication, DHCP policy processing and access to centrally managed workstation profiles consequently ceased within the same interval.
Existing user sessions began failing as cached service tickets expired. Workstations could no longer resolve internal services or obtain a usable lease after reconnecting. By 14:05, all occupied desks were reporting loss of network access despite the lunch period having formally ended.
14:16 — Recovery materials inaccessible
The current switch-stack configuration backup was recorded as being held on the infrastructure file share, which was hosted inside the unavailable VMware estate. A removable copy in the communications cabinet was located, but its label referred to a switch model retired in 2022. The cloud configuration repository could not be reached because neither WAN circuit was operational.
An attempt to restore the second switch member independently was postponed when its console requested an image file that was also stored on the inaccessible file share. Meanwhile, the firewall pair resumed exchanging state and copied the presently incomplete interface status between members.
15:02 — Full site outage
The remaining VMware management path was lost when the surviving core member placed the storage trunk into a protective inconsistent state. Hyper-V cluster communication then failed completely and its remaining guests entered a stopped or isolated condition. Environmental monitoring remained reachable only from its directly connected display.
At 15:18 the mini datacentre had no available production virtual machines, no Active Directory service, no functional DNS or DHCP, no working primary or secondary internet route, and no authenticated workstation connectivity. The original intermittent backup-line investigation was therefore suspended pending a wider recovery activity.
15:24 — UPS switching anomaly
As the virtualisation hosts continued attempting local recovery, the changing server load coincided with an unexpected state transition on the communications-room UPS. Its display reported that an inverter-to-static-bypass transfer had been requested, cancelled and then completed. No transfer had been initiated through the management interface.
The load remained powered, but the UPS was now supplying the room through bypass rather than conditioned inverter output. The network card that would normally provide detailed event history was unreachable because it depended on the failed switching and directory environment. The front-panel log contained only a switching error and an event number not present in the manual kept onsite.
15:29 — Mains supply loss
Five minutes later the building mains supply failed. Lighting and office power were lost at the same time, indicating a wider utility event unrelated to the communications-room work. Because the UPS had entered bypass, it had to transfer back to inverter operation while already supporting the full datacentre load.
The transfer completed for approximately three seconds. The UPS then reported an overload condition. Both sides of several dual-power-supply servers were found to be connected to the same protected output through separate rack distribution units, so the apparent A/B arrangement did not reduce the load presented to the UPS. The inverter shut down protectively and the remaining datacentre equipment lost power without an orderly shutdown.
15:33 — Protective-device failure
When mains voltage returned, the communications-room circuit did not re-energise normally. The residual-current device serving the UPS bypass and auxiliary rack supply operated, but its mechanism stopped in an indeterminate position rather than opening cleanly. The indicator showed isolated while voltage remained intermittently present downstream.
A second protection device upstream also opened, although not before a short period of audible arcing was observed within the distribution enclosure. No attempt was made to reset either device. The event was treated as an electrical fault and the room was cleared.
15:37 — Localised fire
Smoke became visible at the junction between the failed RCD enclosure and the UPS maintenance-bypass supply. This was followed by a small localised flame and melting of the enclosure faceplate. The building fire alarm was activated, the communications-room door was closed and staff were instructed to evacuate.
The fire appeared limited to the electrical distribution area, but the loss of monitoring, emergency lighting changes and continued presence of smoke meant its extent could not be confirmed safely. The incident was escalated to the emergency services and all technical recovery activity ceased.
Finding and attribution
No single causal event could be isolated. The outcome required the close temporal convergence of a carrier handoff that changed negotiation behaviour, an activation-dependent replacement firewall, an existing HA configuration asymmetry, a dormant switch firmware mismatch, incomplete uplink resilience, an undocumented storage dependency, an unsolicited UPS transfer, an external mains interruption, a pre-existing load-distribution issue and a protective device that did not open as indicated.
Each condition became observable only after the preceding condition was tested. The onsite actions were individually consistent with normal fault-isolation practice and were adjacent to, rather than conclusively responsible for, the subsequent infrastructure, power and fire states. On that basis the incident is best classified as a coincidental cascade of latent environmental conditions with no attributable initiating party.
The backup connection remains unverified. The primary connection is assumed to remain serviceable once normal electrical and environmental conditions return. All equipment is to remain isolated pending inspection by the electrical contractor, UPS maintainer, carrier, firewall vendor, virtualisation team and an appropriate system owner.
15:54 — Fire brigade attendance
Two fire appliances arrived at the site and the attending crew assumed control of the building. The electrical intake and communications-room supply were isolated, and firefighters entered the affected area to inspect the UPS bypass and distribution enclosure. The onsite engineering attendance was closed at that point, with the original backup-connection test outstanding.