# The Masterplan: Decentralizing AI Source: https://mercuryprotocol.substack.com/p/the-masterplan ## Summary Mercury Protocol argues that decentralizing access to data and computing power is the way to reduce the risks of advanced AI controlled by a few corporations. The article says that Phase 1 will open large amounts of real-life data to anyone who needs it, with data owners compensated and privacy protected through cryptography and secure hardware. Phase 2 will connect existing operator nodes in a decentralized P2P network, so data buyers can distribute training across many GPUs, with the number of operators capped only by how many are online. ## Article The development of AI is progressing at breakneck speeds. On the one hand, this is positive, because sufficiently advanced AI could be the stepping stone we need to cure diseases, reverse aging, and get to a post-scarcity society. On the other hand, however, there are some grave risks we must pay attention to. This is why our mission at Mercury Protocol is twofold. It is to help humanity build advanced AI that brings us closer to these goals, while also mitigating the risks that come with advanced AI. This post sheds some light on what we are doing to achieve this. First, let’s look at some of the risks. I don’t want to go down the AGI alignment and control route, however severe the risks might be there. Instead, I want to focus on risks associated with artificial intelligence that is already advanced but can’t be called AGI yet. Namely, models we have now (like ChatGPT) or models that -hopefully- we will soon have (i.e. models using genome data for life extension, etc.). Currently, one of the biggest risks is that a handful of corporations have absolute control over these advanced models. “Centralized AI is inherently disaligned AI.” - Balaji Srinivasan OpenAI, which started as a non-profit is now at least partially for-profit, and ChatGPT is kept inside a walled garden. All we can do is trust them and hope for the best. Other big companies heavily invested in AI like Google and Meta are even more closed off. I have nothing against these companies. They are being led and built by some amazing people. Still, if I imagine a world where a few big companies driven by profit have total control over the technology that can disrupt every single facet of what we currently know as modern civilization, I find it scary. So if you agree that only a select few having access to this technology is bad, what alternatives are there? I see two: Everyone has access to it No one has access to it I believe option two would be as bad as the current situation. These technologies have an unimaginably vast array of things to offer to humanity, so they obviously have to exist. So that leaves us with option one as the only logical way forward. Another reason why I strongly side with this option is my faith in human ingenuity and creativity. The best way to move forward is to tap into the brain power of our species as a whole and enable people to contribute meaningfully to AI. Now the question is, how could we do this? Let’s begin by looking at the two things that are essential to advanced AI. They are: Access to huge amounts of data Access to huge amounts of computing power These enable human ingenuity, which in turn enables us to develop advanced AI. The Maslow Pyramid of Advanced AI This brings us to our masterplan, which has two phases. Phase 1 will open up access to large amounts of real-life data for anyone who needs it. Phase 2 will open up access to large amounts of computing power to anyone who needs it. If you wonder why enabling access to data comes before enabling access to computing power, just look at the Maslow Pyramid of Advanced AI above. The most basic need is for data. You can have thousands of GPUs if you have no or bad quality data to train it on. And if you really need computing power, you can just order some GPUs online (granted, it’s likely to cost a fortune). However, if you need data you can’t go online and order it. At least for now. “We don’t have better algorithms, we just have more data.” - Peter Novig, Google’s Chief Scientist Our long-term vision is that every single device on the planet that generates data is integrated with the Mercury Protocol. They automatically earn passive income on behalf of their owners. In Phase 1, our work is focused on bringing this about. Mercury is a decentralized data marketplace, and it has three main actors. Data owners - they own and control the data, data buyers - they are data scientists, AI startups, researchers, etc. who need the data, and operators - they have the computing power and run Mercury software to execute the work on the data. A small side note: these roles can overlap and the system can still maintain security and privacy. For example, an actor can be both a data buyer and an operator. This means it can buy data and then run the operator software on its own infrastructure to train the AI model. Phase 1: Access to huge amounts of data This piece does not attempt to do a deep dive into our protocol architecture, you can refer to our whitepaper for that. What’s essential to understand is that we make it super easy for data sellers to collect and share data, and we use cryptography and secure hardware technologies to make sure that both the model and the data are kept private and secure all the time. So far, only big tech had access to these private and valuable datasets, and they seldom compensate the data owners in any way, while also handling concerns over privacy and security negligently, or not at all. The time has arrived for better alternatives. Mercury will pave the way for them. Once you have quality real-life data, you need enough computing power to train a model. Buying GPUs is a prohibitive upfront investment for many startups. You can use AWS but they also work with insane margins, and they’re centralized, so you’re relinquishing control over your infrastructure to a third party. Phase 2 will be about opening up access to as much computing power as needed. In this phase, we connect the already existing operator nodes in a decentralized P2P network. The data buyer sends off an untrained AI algorithm to the Mercury Protocol and specifies how much computing power it needs (i.e. number of GPUs) or among how many operators the network should distribute the job. This enables startups and data scientists to tap into the power of thousands of GPUs distributed across the world. Phase 2: Access to huge amounts of computing power The upper limit on this is the number of operator nodes online in the network, but anyone with off-the-shelf hardware can spin up an operator node. The fact that an operator is running secure hardware is always verified on-chain. The protocol distributes the work among n available operators, each does 1/n-th of the training, then the results are aggregated into a final model. The data buyers can buy data directly on the Mercury Protocol, upload an existing dataset in encrypted form to Filecoin (operators will pull the data from Filecoin), or send a web scraping script to the operators (operators will run the script to get the data). The technologies that have a huge impact on civilization are those that enable human ingenuity and open up new dimensions for creation. Think of the PC, the iPhone, or the Ethereum Virtual Machine. They introduced paradigm shifts in how individuals and businesses across the world bring products into existence. They are all enablers on a planetary level. If Mercury manages to open up access to massive amounts of data and massive amounts of computing power, it can be the technology that enables AI to reach its full potential.