Four hours of downtime Then everything restarted.
A Lausanne studio’s file server fails after ten years. Four hours later, the editors are working again.
- Client
- Pain Fromage Studio Video production
- Location
- Lausanne French-speaking Switzerland
- Scope
- File server Rushes and live projects
- Solution
- Daily replication Immutable off-site copy
Contents
Pain Fromage Studio, video production in Lausanne.
A team that shoots, edits and delivers on its own, ten years of client projects, terabytes of footage in circulation. The studio holds its delivery dates, and expects the same of its IT support in Lausanne. The whole production chain runs through one central file server, and through the copy made of it.
Visit Pain Fromage StudioA ten-year-old server that does not restart.
The NAS had always held up, to the point where no one noticed it any more. One morning, it did not come back. The machine would eventually restart, much later, but you do not entrust a production to a server that can no longer be trusted.
Without the file server, editing stops and deliveries wait. Four hours separated the failure from the resumption; twelve days later, everything was back to normal.
Three choices set the length of the outage.
The NAS did not fail out of nowhere. Each of these choices made sense when it was made; together, they decided what a failure would cost on the day it came.
A server that had not failed once
Ten years of service without incident. A machine that reliable becomes invisible, and an invisible machine is on no replacement plan.
A weekly replication
The remote copy left once or twice a week. Up to seven days of work could disappear with the server.
A link sized for overnight
Remote access had been sized for nightly synchronisations, not for an editing team working over it in the middle of the day.
Failover to the remote copy, then replacement.
Switch over to the remote replica
More than fifty kilometres from the site, a replication server held a complete, recent copy of the files. We opened the remote access already in place to the team’s machines and sent out the new credentials and the server’s address: the editors got their projects back.
Four hours after the failure, the team was working again, directly on the replica turned temporary production server. In the process, we tuned the link to use the lines’ full bandwidth: working remotely had to be smooth, not just possible.
A setup designed before the failure
The distance was no accident. More than fifty kilometres separated the original from its copy, so that a local disaster could not reach both sites at once. An immutable copy, impossible to alter even with administrator access, completed the setup.
A recovery plan comes down to two measures: how much work a failure can erase (the RPO) and how long it takes to get going again (the RTO). Both were known in advance.
Replication cadence
Replace, without rushing
A new NAS was ordered the same day. The emergency failover was holding production; the replacement could be done properly, without improvisation.
Restore, then return to normal
When the hardware arrived, we restored all the data from the replication server to the new NAS, then reconnected the workstations to their usual environment.
The remote replica went back to its role as the safety copy. The setup that had just saved the production stayed in place, reinforced.
Four hours down, a setup reinforced.
Editing resumed
4 h
Four hours after the failure, the team was working again on the remote replica, now the interim production server.
Timeline of the intervention, in real conditions.
Between two off-site copies
24 h
The replication ran once or twice a week, leaving up to seven days of work exposed. It now runs every day to a site more than fifty kilometres away, in an immutable format.
The failure cost four hours. Without a complete, recent copy hosted more than fifty kilometres away, it would have cost the projects in progress and the clients’ trust. That copy made all the difference.
Production restarts, the failover is written down.
For the studio
The risk, before
A central server down, editors unable to open their projects, and delivery dates that do not move.
The benefit, since
Four hours of interruption, then a team working again on the remote copy while the new hardware arrived.
For the team
The risk, before
Up to seven days of work exposed between two copies, a standby link too slow to work over, and a failover procedure that lived only in people’s heads.
The benefit, since
One copy a day, more than fifty kilometres away, immutable; full-speed standby access that opens to the whole team in one click; and a written procedure, played once in real conditions.
Testimonial
The day our NAS failed, we did not have to improvise. Four hours later the team was working again: production had resumed. Our backup did not only protect our files: it kept the company running.
What the failure revealed.
- The failover revealed that the remote link, sized for nightly synchronisations, was throttling day-to-day work. We tuned it during the intervention, then replaced it with a faster technology.
- The failover procedure lived only in people’s heads. It is now written down, right up to the click that opens the emergency access to the whole team.
The setup that saved the production stayed in place, reinforced. A full copy leaves every day for a site more than fifty kilometres away, in a format even an administrator account cannot alter. And the scenario has been played once, for real.
Your recovery plan
Start by testing your recovery.
What is copied, where, and how the business gets back to work: three questions whose answer can be checked rather than assumed. No commitment.
Book an audit