Shortened links are a digital preservation and web archiving nightmare. You can imagine how they need to work:
Create a unique short code for a given (target) URL (like a hash, but far far shorter)
Pair the short code with your URL in a database.
Create a redirect rule on the URL shortening server from the new source URL to the target URL.
Send the shortened link to the caller, e.g. shortURL.com/123badf00d
In-perpetuity: continue to pay for your domain; maintain the database; look after redirect rules during server migrations; ensure duplicate short-codes are not created.
There aren’t many rewards in a discipline that is about taking the long term view but occasionally something comes up that you can take some pride in.
Last month, Ed Summers put out a call on Mastodon: digipres.club where he was wrestling with a CD-R format that was difficult to recognize. The disks likely held precious data belonging to his late brother.
Much of the search area had already been examined and narrowed down by folks in the community, including Misty de Meo, Roxi Ruuska, Ethan Gates, and Johan van der Knijff who all contributed suggestions and analysis..
Ed was able to share a copy of one of his disk images, and I had some time that I could dedicate to taking a look as well.
The situation might be familiar to others: a digital file that isn’t recognized by the major file format identification tools, and yet, because of its context, you know it is something that might be important.
I have different experiences with these types of files, sometimes they are valuable (and you want to look after them), sometimes they are not (and it can still benefit you to get rid of them). The process of finding this out often follows a similar path.
In this instance the files turned out to be incredibly valuable and I wanted to elaborate on the path of discovery. Even though it really isn’t very sophisticated, I hope it will be helpful to those with unidentified digital records who might find the task of identifying them quite daunting.
We might not have a second life, but what if I told you there was a second internet? Not the deep web, but another web that we engage with nearly every day?
Think about it, that QR code you scanned for more information? That payment link you followed on your electricity bill? The website you’re told to visit at the end of a television ad?
The antipodes of the internet are these terminal endpoints, material and not necessarily material objects that represent the end of the freely navigable web — the QR code on a concert poster is the web printed onto the physical world. There is every chance it will be scanned and followed by someone from a mobile device, but it’s a transient object, something that will exist for a short amount of time, and then disappear into the palimpsest of the poster board or wall it was pasted on until it eventually disappears.
This is part of the materiality of the internet that has long fascinated me. Perhaps it comes from being a student of material culture, but if we look around, we see the Internet everywhere!