# Unlocking Decentralized Storage: Which Solution Holds the Key? Source: https://x.com/BramasPaul/status/1905624795980398647 ## Summary The article compares the cost of decentralized storage with centralized options and argues that price should not be the only factor in choosing a provider. It cites a Coingecko research paper claiming decentralized storage is cheaper than centralized storage by around 78.6% on average, while other figures in the piece argue that Arweave can cost more than some centralized options. The article compares Arweave at about $7/GB one-time, Filecoin with Lighthouse at about $4-5/GB one-time, and Walrus at $0.25/GB per month, currently subsidized to $0.05/month. It recommends first weighing dynamic storage, data permanency, read needs and replica count, and it ends by promoting KYVE as a solution for verifying stored data. ## Article The promise of decentralized storage is captivating: data sovereignty, enhanced security, and resistance to censorship. But alongside these benefits comes a critical question for developers and users alike – what is the true cost of storing data across those solutions? While decentralization offers a compelling alternative to traditional cloud storage, the price landscape is complex and varies significantly between solutions. The True Cost of Decentralized Data Storage: A Comparative Analysis The expansion of decentralized storage is undeniable, with more and more solutions launching every year and many already on the market. Data sovereignty, enhanced security, and resistance to censorship promise to paint a picture of a future where users, not corporations, control their data. But beneath the surface of this technological revolution lies a crucial question: what is the real cost of storing data on these emerging platforms? A Coingecko research paper claimed that “Decentralized storage is cheaper than centralized storage by around 78.6% on average”. While decentralized storage offers a compelling alternative to the familiar pricing tiers of cloud giants like Amazon and Google, the landscape is far more nuanced and complex than it looks, and the pricing might not be the only factor you should look for. This article delves into the current pricing models of some leading decentralized storage solutions, uncovering the factors that influence the cost and providing a framework for making informed decisions depending on your needs and usage. On the other hand, figures of data in the industry, such as @gakonst, raised the question if web3 storage solution can do cheaper than S3. Some builders and researchers say that it’s not possible. For instance, the @0xSmit answer about The Wayback machine data size serves as an example and exposes the prices of different solutions, including Arweave and how it’s more expensive compared to some centralized storage solutions. Some other answers such as the one from @b1ackd0g (Co-founder and CTO of Mysten Labs) bring a more nuanced perspective that it’s not only about pricing but “provenance, the usability of the data, immutability and push the point that a decentralized storage solution done right is a dev platform, not a commodity”. Beyond the Centralized Cloud: Understanding Decentralized Storage Before we dive into the numbers, it's essential to understand the fundamental differences between traditional cloud storage and its decentralized counterpart. Instead of relying on massive data centers owned by a single company, decentralized storage distributes data across a network of independent nodes. Your files are encrypted (depending on the solution used and the user choice), fragmented, and spread across multiple nodes, to ensure immutability and security. This architecture offers inherent advantages: Resilience: There is no single point of failure. Even if one node goes offline, your data remains accessible thanks to redundancy. Security: Data breaches are significantly more difficult when there's no central target. Partial Censorship Resistance: No central authority can easily deny access or delete your data, but a node can refuse to store a specific type of data that would not follow jurisdiction prerogatives, such as pedo-pornographic or hateful content. Comparing Costs: A Look at Four Key Players The table below provides a detailed breakdown of pricing and features for Arweave, Filecoin (with Lighthouse), ETHStorage, and Walrus. Here's a brief overview of each, highlighting their key differentiators: Arweave: The "set it and forget it" option. You pay once for permanent storage (guaranteed for at least 200 years). Ideal for long-term archiving. The upfront cost is higher, but there are no recurring fees. Filecoin (with Lighthouse): A dynamic marketplace where prices fluctuate. Lighthouse makes it easier to store permanently on Filecoin and allows upfront payment for a certain period of time. ETHStorage: Tightly integrated with the Ethereum ecosystem. The price is tied to ETH, so it's a good fit for projects already in the Ethereum ecosystem. However, the long-term cost model and retrieval fees need further clarification. Walrus: Currently the most affordable option due to significant subsidies. This makes it attractive for testing and development, but the long-term sustainability of this pricing model is a key question to consider. 💾 Decentralized Storage Showdown 💾 Which solution fits your needs? Let's break it down: 📦 @ArweaveEco 💰 ~$7/GB (one-time, min. 200 years) 🔄 High redundancy (block weave) 🚀 Free access, no bandwidth limits 🗂 20 replicas 📦@Filecoin (@LighthouseWeb3) 💰 ~$4-5/GB (one-time, min. 25 years) 🔄 High redundancy (replication & erasure coding) 🚀 Free access up to a limit, then $0.1/GB 🗂 At least 3 replicas 📦 @EthStorage 💰 <4.43 ETH/GB (one-time, indefinite) 🔄 High redundancy (Ethereum L2) 🚀 Likely free access, unclear limits 🗂 5000 replicas 📦 @WalrusProtocol 💰 $0.25/GB per month (currently subsidized at $0.05/month, max 2 years) 🔄 High redundancy (erasure coding) 🚀 Free access for now, high bandwidth 🗂 5 replicas Notes: Prices are indicative and can fluctuate significantly based on token prices, network conditions, and storage provider settings. "Indefinite" storage means it is intended to last for the lifetime of the network, subject to certain minimum guarantees (where applicable). More solutions exist and should be looked into, including @StorJ, @Jackal_Protocol, Sia, BNBGreenfield, @Codex_storage, @irys_xyz, and Cascade Pastel, but if we would have included all of them, the article would have ended up being a full research paper. Beyond the Price per Gigabyte: What Really Matters? Mentioning the Coingecko Storage provider report, we can see that Filecoin was the cheapest option in 2023 and that, if you compare all the decentralized storage solutions, they are cheaper than the centralized one, but how come? Centralized storage's higher costs reflect the substantial capital and operational expenses of maintaining dedicated infrastructure. Decentralized storage, by contrast, utilizes globally distributed, surplus computing power as Bitcoin miners use surplus electricity. This cost advantage is further amplified by market structure: an oligopolistic centralized market versus a competitive, open, decentralized one. Pricing is an important factor in decision-making but shouldn’t be the only one, and a few questions should be answered during the decision-making process in our opinion in this order: Dynamic Storage: Do I need dynamic storage or not? Data Permanency: For how long do I want the data to be stored? Read-Optimization: Does the data need to be accessed as a public good, and how many users will need to read the data? Data Security: How many replicas of my data do I want? Answering those four simple questions will already help you understand which provider you might use and give you a path to investigate then, of course, the pricing will come in the end and can impact your decision between two providers, but shouldn’t be the only decision-making fact otherwise you might end up paying for a solution that doesn’t fit your needs. For example, if you need a car because you work in construction, and you have a choice between a Fiat Panda or a Fiat Ducato: both of them are cars and can help you; also, one is cheaper than the other; however, you can fit way less in your Panda than in your Ducato. In the end, if you go for the cheaper option–meaning the Fiat Panda–you will indeed save money at first but having a product that won’t fit your long term needs and will possibly end up changing anyway to a Ducato. Prices were converted from € to $ following the pricing of each model in Germany. Even worse, it’s to overpay in the long run. It’s not because you pay more at once that you pay more in the end, and sometimes, paying more doesn’t make any sense if you are overpaying for a service (maybe you don’t need the picture of your grandma to be stored for 200 years). Pricing over the years, comparison of solutions Let’s look at some numbers and charts to illustrate all of that, few points to consider: The pricing for each solution is at the date T of the article's writing, not of the publication. Pricing can be different. These graphics have been created with public information found in documentation and online resources. For our example, the total data size in the beginning is 40TB, with a growth of 72 TB per year. With that, an average size per year is calculated and used to calculate the pricing for Walrus and S3 for each year because a monthly payment is required to keep the data alive that has been stored. The pricing used for Walrus considers the subsidies. For our example, the data calculation for Arweave and Filecoin using Lighthouse is as follows: the first-year total data (112TB) is used for the first-year pricing, and only the growth of 72TB per year needs to be stored after that. For Lighthouse starting at Year 26, we consider that the data archiving for Year 1 needs to be paid again to avoid loss of the data. We have added S3 without reading in our graphic to show the price difference with a centralized storage solution and S3 with 10 external users reading 100% of the data if we want to take a public-good approach to the data stored. The cost in $ per TB downloaded for S3 used for the calculation is $85. In the graphic above, we can see that Arweave and Filecoin using Lighthouse for a period of 5 years of storage are not the cheapest solutions, and that an approach with Walrus would make more sense. Data storage would be cheaper on S3, but that wouldn’t come with “full” access to or readability of the data. S3 pricing can be found here. If we had a full download of the stored data, which doesn’t go in the T&C of AWS for free, the cost would be much higher. If we zoom out and consider that we need a piece of data stored for 30 years and are still producing data every month, then Filecoin with Lighthouse and Arweave becomes more interesting than S3 or Walrus. If we continue to zoom out, the pricing for Arweave will stay the same while other pricing will continue to increase. Conclusion: Depending on your needs and how long you want to store your data, a higher upfront cost may make sense, but just considering the selection of a decentralized storage solution without understanding your needs and the payment mechanism is nonsensical. The overall cost can be better understood with the graphics we have provided. In the end, it's all about the user's needs. Storing blockchain historical data or books requires long-term, immutable solutions while storing pictures of your holidays requires a shorter, non-immutable solution (no one wants to keep the pictures with their ex permanently on their drive, right?). Storing is cool, but if you can’t retrieve what’s the point? Another important point when storing your data is not only the cost of storage but how easy, and sometimes even possible, it is to retrieve that data. Many projects, and let’s be honest, this is mostly a web3 issue rather than a web2 one, are launching with brilliant ideas in theory but forget the practical aspect of it. It’s like storing your private key for your Bitcoin, now worth $800M, on a hard drive but throwing that disk away in the landfill. The data is indeed secure and stored somewhere but can’t be accessed. And trust us, following the James Howell story, you don’t want that to happen to your Bitcoin or your data. The missing piece of storing solutions for Public goods storage Wherever you decide to store your data, none of those solutions offer trustless pre-validation between the data source and the data storage, which will then allow the end user to have access to verifiable but verified data (not to be confused with provenance). KYVE is the missing piece that enables this feature and, on top of that, provides open-source tooling for data users to leverage that data in a streamlined way. KYVE is built as an agnostic public good solution capable of validating any deterministic data onto any storage solution, centralized or decentralized, bringing data validation to the next era. KYVE storage cost takes into account our compression methods, reducing the data size stored on Arweave but not reducing its usability same method could be put in place before storing the data on any provider by anyone and, by extension, reducing the cost for all providers but would require some extra work to do this compression correctly and more to build the methods to retrieve those data. ## Comments **macbudkowski.eth**: Interesting comparison. I'm wondering about the limitations of these 30 year time horizon analysis. The biggest question is "Will these projects still be around in 2055?". I imagine the probability of these projects staying online is different for Arweave, Filecoin, Walrus and Amazon S3. Also the storage costs go down YoY, and this trajectory is even at the core of the Arweave thesis. Will we produce so much more data that it will outpace these decreasing costs? BTW could you explain a bit more how does this KYVE work? I don't get this part > Wherever you decide to store your data, none of those solutions offer trustless pre-validation between the data source and the data storage, which will then allow the end user to have access to verifiable but verified data (not to be confused with provenance) **paulbramas.eth**: That's a good question and something we can't predict with certainty, but we can assume that most of these projects will evolve and become more competitive over the years, driven by new technological innovations. The more we learn about proper data management, the better we should be able to develop efficient data compression systems. That's the theory, at least. However, when you see some blockchains requiring the storage of exabytes of data due to a lack of understanding of the costs involved, it's clear that there is still a long way to go before people prioritize proper data management. (For example, storing such data on S3 would cost around $500M per year.) At KYVE, we focus on proper and efficient data management, including pre-trustless validation, optimized data compression, and effective data retrieval.