Field report: colocation infrastructure transfer
Planned weekend relocation of a replicated production estate from a client server room into two prepared colocation racks in West London. The programme accepted several hours of overnight downtime in preference to a live infrastructure migration.
Assignment record
| SOURCE | Client head-office server room / on-premises production estate |
|---|---|
| DESTINATION | West London datacentre / two private colocation racks |
| PLATFORM | Application, database and storage tiers across three replicated node sets |
| NETWORK CHANGE | New public addressing / prepared switching and firewall configuration |
| OUTAGE WINDOW | Saturday 22:00 to Sunday 06:00 |
| TRANSPORT CONTROL | Two routes / node set C separated from sets A and B |
Monday — Existing-estate discovery
The client operated its principal business applications from a server room within its head office. The estate had grown in place over several years and contained three complete replication sets, each comprising database, application and storage nodes. Writes were copied between sets continuously, allowing service to continue after an individual node or set failure.
The client had chosen colocation to improve power, cooling, carrier diversity and physical security without replacing the functioning server hardware. A live migration was considered unnecessarily complex because the business could tolerate several hours of overnight service loss. The approved approach was therefore a controlled shutdown, physical transfer and restart before Sunday sunrise.
Half Assed engineers inventoried chassis serials, rail types, switch ports, VLANs, addresses, firewall translations, service dependencies and startup order. Existing diagrams were incomplete, but configuration exports and cable tracing established the active topology.
Tuesday — Historic configuration confirmation
Several network decisions could not be reconciled from the current records. The original installing engineer had left the client company several years earlier, so contact was arranged to confirm the intended replication VLAN, storage heartbeat links and a pair of firewall rules whose comments referred only to temporary migration work.
The former engineer remembered that the storage nodes used dedicated point-to-point links in addition to the documented switched network. They also confirmed that one apparently redundant public address was still used by an external partner callback. Both details were added to the transfer plan and tested against the live estate.
A definitive port schedule was produced from the survey rather than relying on the historic drawing. Every physical connection received a source port, destination port, cable identifier, service, VLAN and target-rack position.
Wednesday — Address and cutover plan
The datacentre allocation required new public and management addresses. Internal application and replication ranges could remain unchanged. New addresses were added to the server interfaces as secondary configuration while the old routes remained active, allowing each system to arrive with its destination values already present.
Firewall objects, NAT rules, management access and monitoring sources were prepared for both address sets. DNS time-to-live values were reduced, carrier reverse records were scheduled and external partners received the new allow-list values. The old public rules would remain available until destination validation completed.
The startup plan placed storage and replication first, followed by databases, directory services, application workers and public endpoints. No server was to accept production writes until all three node sets reported a consistent replication position.
Thursday — Destination preparation
Two numbered racks were built at the West London datacentre. New managed switches, firewalls, serial-console equipment and power-distribution units were installed and labelled. Rack elevations mirrored the source grouping sufficiently that cable lengths and airflow direction could be retained.
The switches were configured with production, replication, storage, management and out-of-band networks. Firewall policy was loaded from the approved destination configuration. Cross-connects, public transit, DNS resolution and management VPN access were tested using temporary equipment.
Every intended server position carried its matching chassis number, power feeds, switch ports and patch leads. The destination team completed a dry startup sequence against the rack schedule before the client hardware left its existing room.
Friday, 18:30 — Physical preparation
With the office still operating, engineers began tagging cables and labelling server faces, rails, power supplies and switch ports. Each chassis received a large transport number visible from the front, rear and packaging. Matching numbers were fixed to its destination rails and rack position.
Power feeds were traced separately so redundant supplies could be restored to distinct distribution units. Fibre transceivers and short specialist cables were removed from the general patching stock and assigned to individual servers. Photographs recorded every rack immediately before the change.
The three replicated node sets were marked A, B and C. Sets A and B would travel in the equipment van. Set C would travel in an engineer’s car by a different route so a single transport incident could not destroy every current copy of the infrastructure.
Saturday — Final service checks
Saturday work concentrated on replication health, data consistency and the shutdown runbook. Background maintenance was stopped, queues were drained and application writes were reduced ahead of the outage. All three node sets reported a matching transaction position and storage checks completed without error.
The client reviewed the planned unavailability and confirmed that service could stop at 22:00. The accepted business impact was the loss of online access overnight rather than the technical and operational risk of attempting a live storage and address migration.
Datacentre engineers rechecked power capacity, rail positions, switch configuration and remote access. The transport vehicles carried blankets, antistatic covers, sealed chassis boxes, rail kits and the numbered manifest.
Three complete replica sets were online and consistent. Destination racks, switching, firewalls, carrier services and addresses had passed readiness checks. The transfer plan deliberately separated one complete replica set from the other two during transport.
22:00 — Planned shutdown
The public service was placed into maintenance mode and new sessions were refused. Final application queues drained, databases checkpointed and replication reached a common confirmed position. Monitoring and scheduled jobs were stopped before application services were shut down in reverse dependency order.
Database and storage nodes were powered down only after all replication journals closed cleanly. Switch-port state and final shutdown times were added to the manifest. The old firewalls remained powered long enough to confirm that no production traffic continued.
The outage began within the approved window. No unexpected service remained active and no server required a forced power-off.
22:34 — De-racking and packing
Engineers disconnected cables in numbered groups, removed servers from their rails and checked each chassis against the transport manifest. Rails, bezels and specialist interconnects travelled with the relevant node. Drive carriers remained locked in position.
Replica sets A and B were loaded at opposite sides of the van’s restrained cargo area with padding between chassis. Set C was loaded into the rear of the designated engineer’s estate car and secured independently. The remaining application and network equipment travelled in the van.
A final sweep found no unmanifested drive, removable medium or live chassis. The source room was photographed empty except for retired cabling and the old edge equipment.
Sunday, 00:12 — Vehicles depart
The vehicles left the client courtyard together. At the exit, the van turned left and the engineer carrying replica set C turned right. Both planned routes converged only near the West London datacentre.
The office coordination contact recorded both departures and expected arrival intervals. The vehicles were not intended to remain together, stop at the same location or follow each other through the final approach.
00:31 — Dual-carriageway diversion
The van encountered a closed section of dual carriageway where overnight roadworks had introduced a diversion. Signs directed traffic through an industrial area before rejoining the principal route. The diversion had not appeared in the afternoon traffic check.
The van crew attempted to notify the engineer carrying set C, but calls went directly to an unavailable message and text delivery was not confirmed. They initially treated this as a temporary mobile-network issue and informed the office coordinator that the separate engineer could not be reached.
The office replied that the engineer had made brief contact from another telephone to report a vehicle breakdown and lack of mobile signal. A replacement car had been dispatched from the Half Assed office. The van continued along the signed diversion.
00:44 — Obstruction on industrial street
The diversion entered an unlit industrial street away from the dual carriageway. Partway along it, the van headlights revealed an unlit car stopped partly in the running lane. The van driver swerved to avoid a direct collision and grazed the stopped car’s wing mirror.
The car was the vehicle carrying replica set C. Its engineer was waiting nearby for the replacement vehicle and had placed a warning triangle behind the car, although the bend and poor lighting left limited approach visibility. The server chassis in the rear remained secured and showed no obvious movement.
The engineer reported that the car had lost electrical power and stopped. Their mobile telephone had discharged unexpectedly; they believed the same electrical fault responsible for the breakdown had also prevented the in-car charger from operating. No reliable contact remained until the replacement arrived.
00:53 — Van continues to bridge
The van crew confirmed that the engineer was safe but elected not to wait beside the stranded car. Keeping sets A and B in the same location as set C would defeat the separation control if another vehicle entered the unlit road.
The van continued toward the far end of the industrial street intending to reach the datacentre by an alternative junction. It then encountered an unannounced bridge closure with barriers across the entire road. There was no route around the works.
The driver turned the van and began returning along the same street toward the broken-down car and the diversion entrance.
01:01 — Replacement vehicle arrives
From the return direction, the van crew saw another car approaching the obstruction. Knowing that the unlit broken-down vehicle remained farther along the road, the van stopped well short and illuminated the area without attempting to pass.
The oncoming car swerved when it finally saw the obstruction and braked heavily. It stopped inches from the van’s front bumper. The driver was the replacement engineer dispatched from the Half Assed office.
The crew explained the bridge closure and escorted the replacement car back to the stranded vehicle. Replica set C was transferred from the failed car into the replacement car using its original padding and restraints. The original engineer took the replacement car and the node set; the replacement driver remained with the failed vehicle to meet the requested tow truck.
01:16 — Second separation attempt
The van and the replacement car set off again. At the bottom of the industrial street, one would turn left and the other right to restore route separation. Before either reached the junction, headlights from a large oncoming vehicle appeared along the single-width section.
The replacement car tucked tightly against the side of the road. The van reversed through an open gate into an industrial yard so the larger vehicle could pass. As it approached, its beacons and recovery equipment identified it as the tow truck requested for the original breakdown.
While both transport vehicles waited, the van crew received a call from the replacement driver attending the failed car. They reported that the car had unexpectedly restarted and that the tow was no longer required. The message arrived too late to turn the approaching recovery vehicle around.
01:18 — Multi-vehicle collision
During the call, the previously broken-down car appeared at speed from the far end of the street. Its driver attempted to pass the oncoming tow truck, swerved and struck the rear of the stationary replacement car carrying replica set C.
The impact drove the failed car’s engine and front structure into the replacement car’s load area. The restrained chassis forming replica set C took the direct line of the intrusion. Drive cages, rails and server enclosures were compressed between the engine block and the forward body structure.
At the same time, the tow-truck driver swerved into the industrial yard to avoid the oncoming car. The recovery vehicle struck the front of the stationary equipment van. The van was forced backward against the yard boundary while sets A and B remained inside its cargo area.
01:20 — Evacuation and ignition
The occupants sustained minor impact injuries but remained conscious and were able to leave their vehicles. The van’s front structure, fuel system and electrical installation had been damaged by the tow-truck impact. Liquid began spreading beneath the engine bay.
All personnel moved away from the vehicles and contacted the emergency services. Within minutes, the fuel leak ignited near the van’s damaged front. Fire travelled into the engine compartment and beneath the cargo floor.
Attempts to reach the rear doors were abandoned as heat increased along the side panels. Replica sets A and B, the remaining application servers and most transport documentation stayed inside the van.
01:31 — Fire brigade attendance
Fire appliances and police arrived at both ends of the industrial street. Firefighters treated the van as a vehicle fire with an unknown electrical and battery load in its cargo. The tow truck and both cars were isolated while casualties were assessed.
The van fire extended through the cargo area before being brought under control. Water and foam entered through the opened side and rear sections. Server chassis were exposed to flame, smoke, firefighting water and debris from the distorted vehicle body.
Police closed the road, documented the collision positions and retained the vehicles for recovery. The planned datacentre cutover was formally abandoned. The destination racks remained powered but empty.
Following day — Replica set C assessment
The replacement car’s rear body had collapsed around the transported node set. The direct impact from the other car’s engine block had driven through the protective packing and server chassis. Storage cages were flattened, boards fractured and drive carriers ejected or folded into adjoining metalwork.
Individual SAS disks were recovered from the wreckage, but several had broken castings or displaced spindle assemblies. Others were bent sufficiently that platters contacted their housings. No complete server, RAID membership set or replication journal survived the impact in readable condition.
The damage was mechanical rather than electronic alone. Laboratory imaging could not establish a complete replica from the partial disks, and missing sectors included current database and storage metadata.
Following day — Replica sets A and B assessment
The two replica sets removed from the van were charred and saturated. Chassis nearest the front of the cargo area had lost bezels, cabling and portions of their drive backplanes. Those farther rearward contained standing water mixed with foam, soot and vehicle residue.
Heat indicators and reflowed solder showed that both storage groups had exceeded normal equipment limits before firefighting. Several disk housings had warped, while others admitted water through breather paths after rapid cooling. Controllers and RAID cache modules were destroyed or electrically shorted.
Drives were catalogued by transport number and examined without ordinary power. Attempts to construct complete membership groups failed because every set contained unreadable or mechanically damaged members beyond its redundancy tolerance. Replication could not compensate: all three replicas had been offline and were physically destroyed during the same transport interval.
Data-loss conclusion
| REPLICA SET A | Fire, heat and firefighting-water damage / no complete readable storage group |
|---|---|
| REPLICA SET B | Fire, impact and water damage / RAID membership beyond recovery tolerance |
| REPLICA SET C | Obliterated by direct engine-block intrusion into replacement vehicle load area |
| APPLICATION SERVERS | Charred and water-soaked in equipment van |
| DESTINATION ESTATE | Racks, switching and firewalls ready but contained no production data |
| INDEPENDENT BACKUP | No current recoverable copy outside the three transported replica sets |
The three-node replication architecture protected the running service against an ordinary hardware failure, but the transfer placed every current replica on the road during the same outage. Route separation prevented the sets from initially travelling together; the diversion, breakdown, bridge closure and recovery response subsequently brought all transport vehicles onto the same industrial street.
Set C was destroyed by the direct collision between the restarted car and its replacement. Sets A and B were destroyed by the tow-truck impact, fuel fire and firefighting operation involving the van. No complete replica, backup or reconstructable storage membership survived.
The colocation racks remained fully configured and ready to receive equipment. There was no data capable of being installed into them. The incident was closed as a total loss of the production infrastructure and all current data.
The source estate was shut down cleanly and the destination estate passed every readiness check. All three current replication sets were nevertheless destroyed before arrival: one by direct engine-block impact and two by vehicle fire and water exposure. With no separate recoverable backup, the complete production dataset was irretrievably lost.