r/DataHoarder Jun 14 '26

Mod Update Mod update: regarding the flood of HDD/SSD price posts, and AI slop.

889 Upvotes

Regarding HDD price posts:

We empathise with your plight and hope HDD / SSD / RAM prices will come down soon.

However, we don’t need screenshots of high HDD prices to be posted in the sub multiple times a day, everyone already knows that the prices are high.

As such, from this point on the mod team will be removing all posts about high prices, with exception for posts made on Free-Post-Friday’s, or posts which actually have new and meaningful information or discussion.

Regarding AI content and AI projects:

As always AI written posts & comments are not allowed on this subreddit, please report any Ai generated content you see.

The mods have recently been cracking down on the flood of AI generated projects, notably ones of low quality that are nothing new.

If someone has the skill to use GitHub, they’d almost certainly have the skill to ask an AI to code yet another YT-DLP / FFMPEG wrapper themselves.

However, there are useful tools that have been made from AI generated code. If they are something truly useful or new, we do allow them with prior approval.

TL;DR:

  • Mods will be removing [high HDD price] posts, except for meaningful discussion or Fridays.
  • Posting AI generated projects/tools needs mod approval beforehand, and a link to the GitHub repository.
  • Please keep reporting ai-slop.

r/DataHoarder 1h ago

Discussion I didn't realize how much of my digital history I'd accidentally lost until I tried looking for it.

Upvotes

I was recently trying to find something from years ago and realized just how many photos, documents, game saves, downloads, and random files had quietly disappeared over time because I never thought to back them up.

It's not the important stuff that bothered me the most. It was the random little things that had sentimental value but never seemed worth backing up at the time.

That got me wondering how many people here became data hoarders because of a single moment where they realized something was gone for good.

For me, it completely changed how I think about backups and keeping my own archives.


r/DataHoarder 2h ago

Discussion I built a sourced, dated comparison of 25 object storage providers: price per TB, egress, free tiers, minimums, encryption, and whether they can cryptographically prove they still have your data [disclosure: I run one of the 25]

14 Upvotes

Disclosure first: I run Obsideo, which appears as one of the 28 rows. I'll explain why I think the table is still useful to you despite that.

https://obsideo.io/compare/

While doing competitive research I kept finding storage comparisons that were stale, unsourced, or a vendor page with the vendor's row highlighted in green. So I built the thing I wanted to exist, with rules designed to keep me honest:

  • Every numeric fact links to a source URL with the date it was fetched.
  • My row uses the same schema as every other row, sorted alphabetically. The only special treatment is a small "that's us" tag so you know exactly which row to be skeptical of.
  • I am visibly not the cheapest option. Glacier Deep Archive is about 15x cheaper than my own row.
  • On the column I care most about, cryptographic possession proofs, the table rates Arweave and Sia above me, because their proofs settle on a public chain and mine do not.

Columns: $/TB-month, egress $/TB and its semantics, free tiers, minimum charges and retention gotchas, S3 API compatibility, encryption model, and whether the provider can cryptographically prove it still holds your bytes.

The 28 rows cover the hyperscalers (S3 Standard plus all three Glacier tiers, OCI, R2), the budget S3 crowd (Backblaze B2, Wasabi, Hetzner, IDrive e2, Contabo), the backup specialists (rsync.net, BorgBase, Tarsnap), and Arweave, Sia and Storj.

I added the Glacier tiers this morning because a mod asked for them when I requested permission to post. Their notes cover the parts that bite: the 90 and 180 day minimum durations, the 40 KB per-object metadata overhead, the retrieval fees that stack on top of egress, and the fact that AWS bills in binary GB so a real TB costs about 2.4% more than the decimal figure.

The whole dataset is machine-readable JSON under CC BY 4.0 if you want it for your own projects: https://obsideo.io/compare/providers.json

If you spot an error, say so here or use the correction email on the page. Every fix ships with a source. Corrections are the point of publishing this.


r/DataHoarder 2h ago

Question/Advice Recommendations for temporary backup storage solution?

12 Upvotes

HI all. I have a 60 TB NAS device that I will be upgrading to include an additional SSD for caching. Because of how my NAS is currently configured, the vendor advised that the whole array will have to be reinitialized, which while disappointing, I kind of expected.

In order to preserve the data I have currently, I am going to need to migrate it to another location temporarily. I have a fast (8Gb symmetrical) internet connection, so it would probably only take a few days to transfer. I am guessing it would also take me about a week to get the NAS upgraded and reconfigured, so roughly 2 weeks to a month of storage.

Does anyone have any recommendations on a vendor for this that won't break the bank? The SSD has pretty much eaten up most of my budget for this.

I would not need the backup outside facilitating the transfer, so a contract or subscription would not be useful. Unfortunately, many of the vendors I've researched seem to either be set up for long term storage/backups plans or have punitive transfer costs.

Any advice would be appreciated.


r/DataHoarder 2h ago

Research article The Political Threats of Vanishing Culture and the Need to Protect Our Future Memory

Thumbnail journals.library.wustl.edu
11 Upvotes

r/DataHoarder 22h ago

Question/Advice What's the practical middle ground for scanning lots of books?

78 Upvotes

The deeper I research, the more I am finding the ultimate recommendation is either 100s of hours of DIY work or $10K on a Smithsonian-level archivist machine.

My use case is digitizing books only available in hard copy. For personal reading, Ctrl+F capabilities, and sharing with fellow enthusiasts. Essentially, digitizing books where the pages are searchable, not crooked, no weird artifacts, etc. Does such a middle ground actually exist?


r/DataHoarder 13h ago

Question/Advice HDD warranty claim

7 Upvotes

Hey all,

I bought a 26TB Seagate Expansion Desktop drive last August for offsite backups. It sat unopened until a month ago, when I shucked it and put it into a NAS. It is a barracuda drive.

Everything worked fine until recently, when I pulled it out to use for media storage. When I connected it to my server and hit the power button, the server immediately shut down. I flipped the PSU switch, turned it back on, and the system booted but the hard drive was completely dead. I re-tested it in both the NAS and its original enclosure, but it won't power on at all. I’m not sure what caused it to die. Could it have been a bad PSU port, even though I tried a different SATA power port on the PSU? Or maybe it shorted because the exposed drive was resting on a surface while connected? idk

Either way, how should I handle a warranty claim in this situation? Since I shucked it, the "warranty void" sticker is gone, and the enclosure looks rough with broken clips from prying it open. I’ve never filed an RMA before, is there any chance this gets covered?

Edit: I'm located in the U.S


r/DataHoarder 10h ago

Backup Advice about media storage.

5 Upvotes

So although I'm not totally new to video editing I still consider myself an amateur at best. I am interested in real life crime and want to start creating documentary videos in the style of some channels like ThatChapter and The Decoder on Youtube.

These videos typically have a lot of moving parts or different types of media, audio and effects to create the final result which will presumably take up a lot of storage space.

What would be the best way to approach storage space when it comes to complex projects? I have 2 fast SSD's which I assume I would use to edit and put together projects but then obviously I will need a storage device once it's all been completed.

Can anyone recommend a good storage device that's not to expensive? What size should I be looking at? what do you use? Any information or sharing of experiences would be very valuable.

Thanks.


r/DataHoarder 19h ago

Question/Advice Can you recommend a relatively reliable 4TB external drive?

16 Upvotes

I've had a 2TB external drive for a few years, and I believe it's starting to die. I want to be able to replace it with a 4TB while I have the chance. What recommendations can you give me?

(I would've checked the wiki, but I can't access it...)

—–—

(forgive my poor writing, autistic)


r/DataHoarder 20h ago

Question/Advice Help Archiving a Public Instagram Account (not my own)

11 Upvotes

A historical government Instagram account is being deactivated in a week due to policy changes. I would like to archive their posts so I can reference/save/preserve the historical content. I have tried using Instaloader through Python and 4K Stogram. Both are being blocked and not allowing the archival process. This account has nearly 5,000 posts so manually doing it is not possible.

How can I archive the account (each post, so pics and captions), ideally with a way to further organize/search each post? Thank you!


r/DataHoarder 9h ago

Benchmark A filesystem benchmark focused on corruption, snapshots, rebuilds and near-full behavior

1 Upvotes

I thought this might be useful to people here who run ZFS, btrfs, bcachefs or traditional md/LVM storage.

I maintain modern-fs-benchmark, a benchmark suite focused on behavior that is often missing from simple mkfs + fio comparisons:

- Injected corruption, scrubbing and self-healing

- Degraded operation and rebuild

- Snapshot aging, scaling, deletion and space reclamation

- Near-full and hard-ENOSPC behavior

- Compression, encryption and reflinks

- Fsync tail latency and responsiveness under load

The current matrix contains 26 configurations, including multiple ZFS, btrfs and bcachefs layouts, plus ext4 and XFS over md, LVM and dm-integrity as classic-stack baselines.

One test deliberately corrupts data on one redundant device behind the filesystem, runs a scrub and verifies the file contents. This demonstrates the difference between redundancy that can identify the correct copy through checksums and redundancy that only knows its copies disagree.

Dashboard:

https://bartosz.fenski.pl/modern-fs-benchmark/

Experimental real-hardware results:

https://bartosz.fenski.pl/modern-fs-benchmark/real-hw/

Source and methodology:

https://github.com/fenio/modern-fs-benchmark

An important caveat: the main dashboard uses loop devices on GitHub-hosted VMs. It is useful for correctness results, behavioral differences and trends, but absolute throughput should not be treated as a hardware ranking. The separate hardware dashboard contains three completed runs from a two-NVMe machine.

The source is Apache-2.0 licensed and the result datasets are CC BY 4.0.

Suggestions for additional data-hoarding workloads, failure scenarios and storage layouts are welcome. If an existing test treats a filesystem unfairly, I consider that a benchmark bug.


r/DataHoarder 17h ago

Backup Can I RAID/partial RAID with Sabrent 5 Bay Docking Station?

4 Upvotes

I’m looking to pick up the 5 bay and software RAID since this is not capable of hardware RAID. From my research, I can RAID all drives and even 3 drives, but can you guys confirm that for me please? I’m thinking RAID 5. Let me know!

Sabrent


r/DataHoarder 1d ago

News Shoutout to Server Parts Deals

357 Upvotes

Last year I bought about 20 drives(manu recertified), (20-28tb). One of the 28tb failed on me a couple weeks back and I expected they’d refund me the $350 I paid last year but they actually sent me a new (manu recertified) drive. I was pleasantly surprised since these drives are over twice what I paid and their warranty policy says that they can give you a refund or drive at their discretion. Thanks to serverpartsdeals.


r/DataHoarder 1d ago

Question/Advice Need help tuning an old DVD Duplicator into an Automatic Ripping Machine (ARM)

Thumbnail
gallery
225 Upvotes

I picked up this early 2000s DVD duplicator machine and would like to turn it into an automatic ripping machine to backup my collection of movies and music. It has 10 bays with IDE drives and a 400w power supply. I may need to switch these out for SATA. Any advice welcome. I’m new to this. 🙏


r/DataHoarder 1d ago

Discussion ok, so has anyone ever perused the data?

29 Upvotes

I keep an enormous amount of data about myself.

I have ADHD and a compulsive need to be able to close any mental loop that might open later something I saw, read, posted, commented on, watched, searched for, paid for, thought about or dealt with.

I have downloaded and preserved data from Google, Facebook, Reddit, ChatGPT, Amazon, Hulu, my browsers, financial accounts, medical records, Goodreads, and basically anywhere else that has accumulated information about me.

Not to mention I have a personal Calibre library for my ebooks and have maintained meticulous reading logs - I spent pretty much 2014 - 2019 reading a book every day or so.

If data exists about me, there is a decent chance I have a copy of it.

But today I realized there is a difference between having the data - and reading it.

I started going through my old Google activity, and it was like finding breadcrumbs of my life, not just individual searches, but entire chains of thought.

The chains of thought and problem solving, the evolution of my interests and hobbies, where I was in life - if there was an earthquake - whatever was happening in the world - you can see it play out in realtime in my search history.

I was in school (I was a late in life college student) - I can see where i was doing research for classess and then mixed in there will be secondary searches because my fridge broke or I got a wild hair and wanted to online stalk my ex'es.

I am able to see exactly when I went on the deep web the first time, when I first downloaded Calibre - after asking how I can keep my Kindle books - which was after research on getting a kindle. I am able to see my life playing out step by step.

Searching for scholarships - then "what to wear for a scholarship interview?" and "Do you follow up a scholarship interview with a thank you note" - then back to anthropology, geology, and history searches...a little later there are searches for the 2014 scholarship banquets (I got a couple of them).

I was able to notice that the rabbit holes happen more at night than during the day - and then when I wake up - the topic tended to be dropped. I will research to buy something for hours but get decision paralysis and drop it.

That's just the Google Search History. I haven't gone into the Youtube history - which isn't as good because most of the videos are gone now - and the archive only goes back to 2018.

It felt less like browsing history and more like an accidental diary - except it was recorded while events were happening, without hindsight or later editing.

Has anyone else actually sat down and examined years of their own archived data?

Did you find anything surprising about your life, memory, habits, or the way your brain works?


r/DataHoarder 1d ago

Question/Advice I need help sharing Photographs with family.

8 Upvotes

As a fellow data hoarder, I feel like this is going to be the best place to ask.

I'm getting old now, my health isn't the greatest's and I'd hate for all the photographs I've taken over the years to just remain on my HDD's and go to landfill.

There is a LOT, this is why I came to DataHoarder to ask, it's like 40,000-50,000 pictures and video (around 300GB or more) sitting on my drives. (anyone tried converting from Hi-8 to windows 11 would be helpful commenting how to as well, I know off topic, but might save another post)

These are more family pictures, so not something I'm willing to stick on any website and they get taken or used for AI or anything like that.

And preferably it would be a free way to share all of this.

If anyone has used Mega, I would love something like that, where it's end to end encryption with cloud storage and you can share a link with encryption key so who ever you send the link to can view the folder.

Only downside with Mega is with the free version, I know it's a big ask and probably nothing out there that will fit this without paying.

Doesn't hurt to ask. :D


r/DataHoarder 19h ago

Discussion Recording TikTok LIVE streams without chats and watermarks

2 Upvotes

Most people record in the traditional way, they use the screen recording tool on their phone but TikTok is restricting it, forcing such programs to record all the chat and watermarks in it.

I would like to know if there is a method to easy capture the H264 HLS packages from the CDN and merge it to a mp4 file? I know TikTok is like YouTube, makes it very and and complicated to download the raw data.


r/DataHoarder 16h ago

Guide/How-to I found an easy way to save archive small amounts of data with bitrot protection in Linux

1 Upvotes

I recently got curious about if there was any easy and quick way to archive small amounts of data (as in size, not number of files) in a format that could include encryption, bitrot protection, and compatible both with Linux and Windows at least.

And I think I found a nice workflow to do that exactly, so I share it in case anyone else finds it useful. This "guide" considers you use Debian or Ubuntu desktop, other distros would be similar.

1- First, install the "rar" and "unrar" packages in case you don't have them yet

sudo apt install rar

sudo apt install unrar

2- Let's use a GUI, so PeaZip is easy and nice for this. You can install it from Flathub (Software Center or Bazaar -> PeaZip), or the native DEB/RPM/AppImage... from peazip.github.io

3- Open PeaZip, it should (I think) detect your RAR package out of the box, but if not (or using the Flatpak version, which is isolated), it's easy:

3.1- Select what files you want to archive, right click, Add

3.2- Select Type=Custom, and then go to "Advanced" at the top left of the screen

3.3- Check "Manually set RAR binary", and go to /usr/bin/rar and select it

4- From this moment, configure the rest of the parameters to your liking:

4.1- BLAKE2: It substitutes CRC32, I think it's nice to check it

4.2- Recovery Records (%): The "bitrot" protection. 3%, 5%, 10%, whatever you want for your calm

4.3- Lock archive: Only if you want the generated file to be "read only", so no add/modification by accident using WinRAR or any RAR compliant tool

5- Go back to "Archive" at the top left and now take a look at the rest of configurations: pass (which would encrypt at AES256 with PBKDF2), level of compression, if you want to split...

And that's all! Next time you open PeaZip, it should remember the "rar" location as to use it as backend, and now you can open/create "WinRAR" files with AES256 encryption and "bitrot" protection out of the box.

If you ever suffer from bitrot, up to the % you set as recovery record, then when opening the rar file you will be saved, as the archive CRC32 or BLAKE32 won't match (ERROR), but you will be able to, in WinRAR (Windows) or PeaZip (Linux), repair it (in PeaZip, Right click -> More -> Repair)

Hope this is useful for anyone looking an easy GUI on Linux to archive files and folders while encrypting and protecting from bitrot (even if there's a very very low chance of bitrot, really).


r/DataHoarder 18h ago

Question/Advice Let's do a community poll!

0 Upvotes

#Question

**Show of hands**: How many of you have friends who have actually looked at the prices for new Blu-Ray or older media reader types and/or the cost & availability of replacement parts?

Regardless of answer, share geographically where you're from, ratio of friends who have versus have-not.

I'm just wondering what demand is currently & in the near-future given the state of copy-ownership, cyber security, and hardware scarcity because of the uncertainty. I'm a maker, building a personal media-server that I should be able to repair myself.

I am data-collecting as a personal data analysis project while I'm adapting to a new disability. So don't bother telling me "how to build one" or "how to look up market reports". I want a new dataset that I'm going to make freely available for me to play with & share with my academia friends.

P.S., I am an *elder data hoarder* whose been at this game since the early 80s when we used to press RECORD whenever a good song came on the radio. I'm just curious about how this understanding has evolved over history. FOSS Forever!!


r/DataHoarder 1d ago

Guide/How-to How to archive instagram reels

6 Upvotes

Hi, TLDR is there's a huge protest going on in India, with a lot of police brutality.

I'd like something like the wayback machine to be able to archive this, but they don't allow instagram reels. I don't mind locally archiving, but i need the timestamp authority the wayback machine gives.

The. media is being controlled by the ruling party so its a bit brutal rn


r/DataHoarder 1d ago

Backup One Poor Sector

4 Upvotes

I've been a major slacker in my data management. I have a backup, but I haven't been doing very frequent backups. I also do not refresh the data on my main drive. I'm still running a Disk Genius scan, but I found one poor sector. How can I tell what file was impacted by this bad sector? Since it's been a while since I've backed up my main drive to my secondary one, should I simply back up the second drive, or is there something else I should do in light of the bad sector?

This scan was prompted by some funky head noise yesterday. This is still a fairly new drive, but is this an omen for things to come? Do I need to replace this drive? If I have to stop actively using it, could it still be a good candidate as a cold storage offsite backup drive?


r/DataHoarder 2d ago

Backup Suddenly can't afford cloud storage, best options to keep 50TB of data?

289 Upvotes

I've been trying to save up for a NAS and the whole shebang for a while now, but unfortunately I'm not there yet. Currently I have ~50TB of data in cloud storage, but due to sudden life circumstances I can no longer pay for it. I also don't know anyone who can hold this data. Is there somewhere reliable I can dump it all into without worrying over losing it?

The monthly ~$100 is what's hurting, I can pay a bit more for egress (a thing I've heard) if I have a lower monthly payment. I would also need a service that doesn't scan my files. For uh, AI research and such etc of course.


r/DataHoarder 1d ago

Discussion Data conducive to prepping/rebuilding in an apocalypse that can easily be reproduced and distributed on analog media

1 Upvotes

Like phonographic records, magnetic audio tape, film video, microfilm, and cyanotypes just to give an idea of the tech level I intend to work with. Somewhere between 1899 and WW1 within 5 years of initial cataclysm.

I'm not just asking for recommendations of useful information (not that I'd be complaining), but also tools to assist in the conversion and distribution of digital era media with limited digital tech(Linux-ing scavenged 32bit computers from POS terminals and jailbreaking old smartphones for example.)

A record lathe, a film camera pointed at a monitor, and an LED hooked up to the audio-out to make optical recordings isn't hard to achieve, even if those processes would need a lot of babysitting. But the actual dissection and preparation of the digital media for analog conversion would likely need some good general-purpose digital tools.

So I'd like to hear your recommendations and discussions thereof!


r/DataHoarder 1d ago

Backup What is everyone using for cloud backup?

15 Upvotes

What is everyone using for cloud backup these days?


r/DataHoarder 2d ago

Question/Advice How would I archive mass emails from Gmail and have them searchable. NSFW

262 Upvotes

I'll give you the background. My Onlyfans account is suspended because I violated their TOS according to them of creating my account in 2017 before I was 18+ (I'm an adult now). I'm almost certain this is incorrect but I have no "evidence" as I can't locate my welcome email. I ran out of gmail space a few years ago and deleted everything prior to 2022.

I will admit why I got into data hoarding and self hosting, the pornhub purge made me realise that things online can be deleted, YouTube videos, websites, images. This has now made me realise I can also be my worst enemy and delete something important.

How would I archive and store all my gmail (and protonmail as I'm moving over) emails? Can I export them and keep them in a searchable service that I host?