The broken state of archive in backup

Ancient archive of dusted papers in a forgotten storage place.

Let’s say you have a front-end estate of 500 TB of data that you’re protecting. And let’s also say that you’re doing what I’d describe as fairly industry-typical retention times for those backups, viz.:

  • Daily backups kept for 31 days
  • Monthly backups kept for 7 years (84 months)

What does that look like, over time? Well, for the purposes of a more realistic sizing, we’ll assume that you have a 10% linear annual growth rate. (Previously I’d say that most companies over-estimate when they default to saying they have a 10% annual growth rate, but with AI models and data sets befuddling the waters, this is now a number which is very much up for grabs.)

At the end of 7 years, the bulk of data held by the backup system is archival data: your ‘final’ monthly backup, in month 84 will be 974.36 TB (assuming our first backup was taking after amortised monthly growth occurred in month 1). Even if that’s a full backup of 974.36 TB + 30 incremental backups (and these days, synthetic fulls definitely change how this is perceived), it’s a pittance compared to the cumulative build-up of those monthly backups: 59,961.47 TB – almost 60 PB.

So, after 84 months, the daily backups account for just 1.6% of the protected data footprint; the remaining 98.4% of the backup footprint is your archive copies.

Those archive copies aren’t being kept for shits and giggles: there’s very likely a regulatory compliance requirement coming into play – a legally enforceable records-retention directive that drives the need to be able to recover old data for tax purposes, financial reasons, etc.

Let me ask you this: are you keeping those copies to tick a box, or with the intention of actually recovering them should a regulator/etc. require you to? Because if you’re keeping those copies for the latter rather than the former, you may be kidding yourself.

There’s a problem in the data protection industry (regardless of whether we call it backup, data protection, or cyber resilience) when it comes to archive, and it can all be summarised by a single word: obsolescence. While obsolescence can be a problem for short-term retention backup (e.g., what happens for recoverability of your older backups when you decide to switch from Oracle on AIX to Oracle on Exadata, which creates an Endian switch in your data), it’s just a subset of the problem we see with archival backup. The obsolescence factor really comes into its own when you’re looking to recover backups taken 3, 5, 7+ years ago for regulatory compliance purposes. You know – those situations where if you’re unable to recover the data, the company may very well face serious fines or legal action.

In data protection we have the concept of cascading failures – data loss (irretrievable data loss, as opposed to “I now need to go recover the data”) shouldn’t ever come from a single failure point: you should always have your backup, your backup should always have a secondary copy, that secondary copy should always be in another location, and so on. Cyber attacks very much play on introducing cascading failures – at their simplest, you’ve had your production and backup copies attacked. You now need to go through more complex rebuild processes as you have to scrub your environment, isolate safe data and restore it – after reconstructing what you’re going to restore to.

But archival restore is where cascading failures will really come into the fore. The proverbial excrement hitting the spinning blades sort of thing. People often think I’m joking when I tell the story of a customer who told me once (over 15 years ago) that the first step in their recovery from long-term backup procedures was “Go check eBay for the compatible tape drive”, but I don’t joke around when it comes to recovery and it really was their documented process.

So what could change between now and when you need to restore from a long-term archival backup for regulatory purposes?

  • Platform (for physical systems)
  • Virtualisation platform (e.g., VMware to anything-that-is-not-VMware)
  • Location (on-premises vs cloud)
  • Byte ordering (big endian vs little endian)
  • Operating system version (practically guaranteed except where it’s dead tech that’s been declared irreplaceable)
  • Operating system type (e.g., Solaris to Linux)
  • Application (e.g., Oracle to PostgreSQL)

The TL;DR of that list is “everything”, but there are some really important examples of change that could happen I left off the list above. These are:

  • Your people
  • Your processes
  • Your business priorities

These kinds of changes will monumentally impact how you can achieve recovery of data from archival backups. Together they represent business velocity (what’s important to the business and what it is focused on), and institutional knowledge.

But, as Apple love to occasionally drop in their product releases, there’s one more thing. And in the spirit of Apple, I’m saving possibly the biggest thing (in this case though, the biggest headache) for the end. What else can change between teh time the archival backup was taken and the time you need to restore?

  • The backup software – not just versions, the vendor technology stack
  • The backup hardware

And if you think all the other problems are difficult to deal with – wait until you hit the final two.

It’s fair to say that backup vendors don’t have much inventive to solve this particular problem. The argument invariably is: if you want to recover from your long-term backups, stay with our product, the one you used to trigger those long-term backups. And I can understand that. Except even staying with the vendor doesn’t give you that guarantee. Will it still support recovering some version N database in a version M operating system when you’re up to N+7 and M+4 and N/M support was dropped 3 years ago?

So. Here we are.

The problem is multi-faceted. Leaving aside everything else I’d mentioned (OS versions, platform, application, etc.) and just focusing on the backup software/hardware platform, we can summarise the challenge as:

  1. The format of the backups.
  2. The catalogue for the backups.
  3. The medium the backups are written to.

For (3), it’s arguably less of a problem if you’re writing to disk as disk. That means deduplication storage isn’t necessarily an impediment here, so long as you can get a filesystem view. (Though it’s important to note – deduplication is going to help you with the backup storage footprint, but not much else.)

But (2) and (1) are going to be real headaches. Now, some backup products give you a little edge towards this by writing their backups in fully open format – ArcServe, I’m told, can write in cpio or tar format (but not if you’re backing up Windows), and NetBackup writes at least some types of backups in tar format – albeit slightly modified. AMANDA definitely uses open formats, and while Bacula writes its own format, I understand it comes with tooling you can use to extract backups regardless of whether you’ve got a catalogue.

Oh boy. The catalogue.

In the first, triple episode of Stargate Universe, where the heroes find themselves trapped on a starship (the Destiny) hundreds, if not thousands of galaxies away from home, with few provisions, there’s a funny moment where they have a random collection of provisions in sealed boxes they managed to grab before escaping to Destiny. They have to inventory each box manually because they’re barcoded and they didn’t bring the barcode scanner/database with them. That’s the catalogue problem in a nutshell (albeit with fewer stargates).

Congratulations, you’ve got 40 PB of backups in product X but you’ve moved on to product Y and no longer have access to the catalogue that wrote the backups. Very small needle: meet a very large paddock full of haystacks. Let me frame this properly: you may have a regulatory compliance requirement to restore a single, 40 KB document from 40 PB of vendor-format backups. Do you know which backup it comes from?

In lieu of better options, I’ve seen businesses generate CSV files or otherwise plain text dumps listing the content of every backup taken by their system to provide some level of catalogue independence should they change backup products later. It’s honestly not a bad idea – particularly if you’re not keeping granular indices for the full retention time (e.g., you might keep those archival filesystem backups for 84 months, but only keep file-by-file restore indices for the first 12 months) – though keep in mind, these dumps will build up over time and will need appropriately organised storage that is also protected. Doing so gives the immediate advantage of a simple, searchable register of protected content. What’s more, a relatively simple Perl/Python/etc. script could be used to look for desired keywords, file names, etc., in parallel (hint: if you create the script, you don’t have to use AI tokens every time you want to run a search). Then all you need to do is execute a recovery.

Honestly, my preferred option here would be that every backup product support an automated and manual export facility. I’m going to upset the keep-the-government-out-of-business-folks here and suggest that it’s time lawmakers who decide on regulatory compliance retention need to also start thinking of formats and recoverability mandates. (If you think this is inappropriate overreach, you haven’t been paying attention to how governments issue security mandates.) I’ll describe the ideal process via a manual export, but for automated, consider it just happening by default once every monthly backup reaches say, 12 months of age:

  • I select a backup for manual export and tell the product to export it
  • I select destination media (e.g., disk, tape or object)
  • The backup is read and inline-converted* to an OS-native format (e.g., tar), and written to the selected destination media
  • Both before and after the backup, metadata content is written – essentially, a plain-text catalogue of all the content of the backup, when the backup was originally written, what system it came from, etc. Essentially – everything you need to not only find content, but prove it’s an unmodified backup from a particular point in time.

* Inline-conversion, I recognise, is problematic, particularly when we hit applications. For filesystems, it should be straight-forward. But ideally, virtual machines should be converted to a standardised format like QCOW2, relational databases should be converted to SQL dumps, or in the worst case, independent data files that can be simply attached to an equivalent database, etc. That’s the start of the first list of problems creeping into the ‘export’ process, so we can say they’re two related, but somewhat independent problems to solve.

As I mentioned, automated conversion should be a logical extrapolation of the above. We may want to keep the backups in the vendor-native format for a while for faster, easier recovery (e.g., 80% of your recoveries might come from the last 31 days, 10% from the first 12 months, and the remaining 10% from true long-term backups). But at some point, the system should be able to either move or at least copy the backups into an exported format where having access to the original backup product just isn’t necessary.

This is something I’m particularly passionate about. If you’re a vendor/company who wants to solve this problem for its users – talk to me: I want to work with you.

—
Note: This blog is entirely human-powered. Any errors are shamefully accepted by the human writer rather than blamed on LLM-hallucinations.

Cover image attribution: From BigStock Photo.

Share: Linkedin
Author: Preston de Guise