I’m not sure you and the author are talking about the same thing. He mentions the fact you can’t even read the news from 10 years ago, the content has simply disappeared. No amount of YouTube videos can replace that.
The problem is not “everything, everywhere” or a lack of filters but the extreme commercialization of all content available, closed networks, the short life of URLs…
The irony here is that Youtube videos from ten years ago are still alive and well. As Youtube makes a much better places for publishing and archiving content than the rest of the Web. With Youtube you don't have to worry about URLs changing or domain names expiring or anything like that. You just publish your video once, get a unique video-id and don't have to worry about anything else. Google's monopoly worked here in our favor for once and they have been reasonably good in not breaking old content (not perfectly, as video annotation got crippled pretty badly).
I think that's where the rest of the Web fell short. The Web has no concept of "publishing". There is no ISBN when you write a blog, no library were you could look up that ISBN. It's all just a file on a server or an entry in a database, that will get mangled and lost in the coming years. Worse yet, the article itself isn't even accessible from the Web, it's mixed together with a user-interface, ads and other stuff or spread across multiple pages. All this makes it quite tricky to keep old content readable and archived for the future.
This also leads to a weird situation that a lot of publications are still avoiding the Web after 20+ years. They publish as ePub or PDFs instead, which you can somewhat access from the Web, but really aren't well integrated into it. But it's by far the easiest way to ensure that a text document published today will remain readable a decade down the line.
Unlisted youtube videos from 10 years ago are all gone... Google made the decision to delete every video 'shared by URL' because of the possibility that the URL generation algorithm had leaked. It was legally less risky to delete all the content than to risk leaking all the content to the open web.
IMO, they made the wrong call - it would have been better for the internet as a whole to notify all users that "We have a new 'share by link' option, and no longer consider the old links private. Please update all old links and then click this button to disable the old URL. If you don't click the button, videos shared by URL may be discovered by others in the future."
Unfortunately when I look through my old Favourites playlist, which at some point reached the limit of 5000 videos, I can see how many of those videos are now private, deleted, or blocked in my country. The worst part is that in many cases I can't even recover the video title, so I have no idea what has been lost. A possible solution would be to store the titles separately, but I didn't think about this while I added stuff to my collection.
I do agree though that the unique, unchanging URL is a huge boon, when I look at the situation in my browser's bookmarks for comparison.
I made our wedding playlist on YouTube, since I pay for premium, and about 30% is now gone after a couple years. It would be nice if YT at least gave some indication of things they do not plan to delete, i.e. it's an official source not an upload subject to the whims of DMCA.
>The worst part is that in many cases I can't even recover the video title, so I have no idea what has been lost.
I use two methods to recover deleted video titles:
1) google videoId part of the URL
2) look up cleaned up (with removed playlist-related parameters, just "youtube.com/watch?v=[videoId]" form) video URL in web.archive.org - sometimes it even has an archived video itself!
There's plenty of videos on youtube that have been removed. Every few years when I look through my list of liked videos, a couple more are gone forever. Granted, this is likely by the creator themselves, but that doesn't matter when the purpose is archival.
On a playlist of about 27 thousand videos (me and my cohorts links to each other over ~8 years) I observe about 10% missing, its a lot!
Most of these are not information but amusing, random, short videos, and of course some of it is account deletions/disabling - copyright strikes - and people taking down their own stuff, but still, a lot!
When they delete videos, they don't even list the title of the video anymore, they remove the thumbnail and everything, so you won't even know that something you added months ago is no longer available.
> The irony here is that Youtube videos from ten years ago are still alive and well.
They're not, though.
Because Youtube re-encodes videos every couple of years, with new "better" lossy compression algorithms. And each time the videos get successively worse.
Watching a 2008 Youtube video will not only look grainy because it's 360p, but it'll look actively *worse* than it did in 2008 because of all the lossy compression that was applied to it over the decades.
YouTube has definitely kept the original videos for as long as I can remember, so any transcodes you see are only one generation after the original (plus maybe one more for if they didn't launch with this feature)
And yes the new codecs really are better. Same quality at lower bitrate.
I know the norm when I read an article more than a year or two old that has tons of YouTube embeds is for most or all of them to show a "video gone" error. Sometimes even ones that are weeks old are like that.
It seems to be at least as bad as the rest of the web as far as data/link rot goes.
As a sibling comment noted, I'm pretty sure Youtube keeps the original video files, so generational loss should not be a huge problem.
The problem is that Youtube is bitstarving 240p and 360p streams. That makes sense when a video is available at a higher resolution, and the 240p version is for people stuck on dial up or whatever. But in cases where 240p is the highest resolution available, Youtube should provide a high quality 240p stream!
> The irony here is that Youtube videos from ten years ago are still alive and well.
Tons of them aren't. I've run into many linked from wikipedia footnotes that no longer exist, particularly digitized film from the 20th century. A ton of wikipedia pages still cite old films, documentary footage, linking to videos which were uploaded by Jeff Quitney, who was banned from youtube a few years ago because some of the old films he uploaded contained material that ran counter to modern values (I think the one that eventually got him banned was an old christian film warning children about homosexual predators.) When they banned him, they took down a ton of completely innocuous videos because a tiny minority were offensive.
> There is no ISBN when you write a blog, no library were you could look up that ISBN.
An ISBN is a string of characters, much like a URI/URL, and offers no more "protection" for long term access and the latter. Books get mangled and lost too; their only benefit is that it is harder to mangle and lose them, and there is likely to be more than one of them.
> The irony here is that Youtube videos from ten years ago are still alive and well.
Do videos survive account deletion ?
In particular, GDPR has provisions about deleting account info after years of inactivity, and Youtube is apparently not an exception (https://qr.ae/pv9o9j)
So except the chunk of accounts that will stay active for the years to come, a big part of youtube videos should be disappearing progressively.
I think you forget or may not have grown up with the microfiche.
The reality of the internet is that everyone has a voice and things will only be archived if someone gives a damn to archive them. And that's fine. Some information deserves to be transient. Hell, we've survived millennia without this level of information storage. Does every single YouTube video, Reddit post, and Flickr photo really need to live forever? No. Would it be nice? Sure.
That's fine for stuff that was printed in newspapers, but you (like the topmost commenter) seem to be replying to something different from what the article is talking about.
Paper is great, because it doesn't just evaporate when you look away. Although it does degrade, it's a slow process—slow enough that you can notice that it's happening and think to yourself, "Gee, I maybe ought to do something about this." It's not the same for transient digital media. There are bonafide news items and other digital content that are now no longer accessible because they were digital-first but the business incentives were so misaligned and/or their legacy has been so mismanaged that, perversely, it's easier to access to the content of a 50-year old news article than it is for others that are 5–15 years old. People can always trawl through their parents' and grandparents' belongings and come across the only known surviving copy of something and donate it to a library or sell the collection in a yard sale or eBay. (Whether they know it's the only one known to exist or not isn't a precondition.) That's not just less likely with the Web, but it's drastically less likely. No one's gleaning much from the unevicted entries in someone's browser cache.
One of the functions of public libraries is to archive the news. One could browse weekly papers going back decades, either as hard copies or microfilm. There's a good chance major city libraries still do that, or have digital scans.
BBC used to be an extensive and complicated site with non-news articles, study guides and curated collections. All of that was destroyed with a revamp. You can still read the news articles, but that's all that's left.
Not really. There are vast archives of newspaper articles accessible through a web interface. You just need library access. The free, open web is mostly just a spam ocean, but if you make an effort to access the services that catalog useful information, it’s still very useful. The Google web is shit, though.
There are scans of old newspapers going back ages, and current ones also. Stop using the web, it sucks ass. Use ProQuest and other similar databases. Web searches are wastes of time for most things. Google wants to train you to believe research is not a skill and that you can get usually get good information from the open web. Neither of those points are true.
You may not read this it has been a couple of days, but thought I should reply.
I am not sure why you replied to my comment l, with talk about not using the web, google, etc.
After all, I specifically said that trusting web bawed properties is an issue.
And you seem to have missed a key part of my comment, the state of paper based newspapers today.
They're dying.
Huge, multi-millon dollar, national newspaper corps, which used to have Saturday newspapers 2 or 3cm thick, when I was a kid, are now a mere 1/2 cm if lucky.
Many local newspapers are gone. Just gone.
The ones which remain, are barely standing on their feet.
So my comment, re archival, is that old methods of scanning phycial newspapers are not viable, and getting worse yearly.
Yes. A conversation about the utility of the internet as an archive of news material presupposes access to the internet in some form. But in the Western world at least, notwithstanding certain American locales where right-wing activists defund libraries which carry wrongthink pro LGBT, trans or anti-racist literature, access to a public library is not uncommon even in rural or poor locales.
The "All content available" statement is a bit much. I'd go so far as to say there are a good portion of sites that have commercialized information. But I can easily access other forms of information not commercialized by avoiding mainstream views.
The problem is not “everything, everywhere” or a lack of filters but the extreme commercialization of all content available, closed networks, the short life of URLs…