Secondary Data And The "Full Cup" Problem

Words
1088
Reading
5 min
Listen
Play
9y

There is this story of a Zen student, who was ardently seeking wisdom and liberation. After many years of practice and pilgrimage, he learns about a very special monk, who was living alone, on top of a mountain. This monk had extraordinary powers and tremendous insight. Our student decides he is going to visit that monk and get enlightened.

After a very difficult trip, in which he nearly dies, one sunny morning he finds himself on top of that mountain, in front of a small hut. Inside, there was a simple man, with a gentle face.

-- Are you the famous Zen monk who knows the path to liberation? asks our student.
-- Are you thirsty? asks back the old man.
-- Not specifically for tea, answers the student. I want to find liberation, can you teach me?
-- Well, let me fix you a cup of tea, says the old man.

Thinking that the old man may have gone a bit senile, the student sits down and patiently wait for the tea, hoping to find his next opportunity to immerse in the Zen monk's wisdom. The old man puts a cup on the table, takes a tea pot and starts pouring. He pours until the cup is full and then continue pouring.

The student jumps up and yell:

-- Can't you see you're pouring on me, old man?
-- Well, says the monk, continuing to pour, your mind is just like this cup: it's full but yet it yearns for more wisdom. You cannot teach a mind which already thinks it knows the path.

The "Full" Cup Problem

We live in the golden era of the blockchain. We're just scratching the surface. And by that I mean we're literally at a very shallow level of our blockchain usage.

Because, beyond all the governance mechanisms, beyond all minting algorithms, there is a very sensitive part of the whole story which is overlooked, simply because we're at the very beginning of this. This part is blockchain storage.

If you want to run a Bitcoin node, you already need generous storage. If you want to run a Steemit node, you need at least 100GB of SSD storage, just to be sure you're accessing all the data you need, at a speed comparable to one of a traditional relational database. There is so much info piling up, and the paradox is, the more popular the blockchain technology is, the bigger the amount of data grows.

With every new transaction we're adding to an already thick layer of transactions and we bank on the immutability of this data. But soon enough we'll reach a point where our storage will outperform our data retrieval capabilities. In other words, we'll be able to store a lot of data, but retrieving that data will be a big problem. We will have a "full cup" blockchain.

It may seem that this is not going to happen any time soon, but the speed of growth of these strange things called blockchains is tremendous. When I first joined Steemit, a year ago, all you needed was a mere 16GB of HDD to store the data. Now you need almost 10x more space. And not only the space is the problem, because HDDs tend to be cheap these days, but the speed of access is what creates some problems, especially on blockchains with higher TPS parameters (TPS = Transactions Per Second). If you imagine a blockchain with a few terabytes of data and then a transaction between a very early account (one at the "beginning" of the storage) and a late account (one at the "end" of the storage), all needing to happen in less than 3 seconds, including network latency, then you start pondering.

Secondary Data

I personally see no other solution to this than downgrading some of the data we use. Right now every piece of information is considered necessary for a transaction to be stored in the blockchain, so the need for a "secondary", less important level may seem inadequate.

But we already use the concept of "secondary data" in many areas of our lives. We make our data secondary based on a few paradigms.

One of them is usefulness. If something is not deemed useful, or even "solicited" or called for, we put it in a secondary layer. The most prominent example is email spam. Try to live your life without your email spam filters for a day, and you'll se what happens. You'll be overwhelmed.

But there are other paradigms we use, and some of them are enforced at the legal level. In Europe, for instance, we have a law based on "the right to be forgotten" stance, which gives you the right to "erase" your public data from public information sources, like Google searches. If you think something is threatening your privacy, you can just ask Google to eliminate it. This is another example of "secondary data".

I think the most important challenge of the whole blockchain industry in the next 4-6 years is this "full cup" problem and how to define and manage the amount of "secondary data" for each use case.

I can see situations in which entire chunks of blockchains are saved under one identifier and made available only on specific requests. I can also see a lot of AI assistance here, meaning we will rely on AI to learn how and how much of this "secondary data" is accessed and what makes a piece of data suitable to be considered "main" or "secondary".

I don't think generation algorithms or governance methodologies are the most important direction for the next 4-6 years, because, in my humble opinion, we are already saturated here.

But if the problem of effectively store and retrieve the immutable data that the blockchain generates is not solved, then the entire network will simply halt.


I'm a serial entrepreneur, blogger and ultrarunner. You can find me mainly on my blog at Dragos Roua where I write about productivity, business, relationships and running. Here on Steemit you may stay updated by following me dragosroua@dragosroua.


Dragos Roua


You can also vote for me as witness here:
https://steemit.com/~witnesses


If you're new to Steemit, you may find these articles relevant (that's also part of my witness activity to support new members of the platform):

Secondary Data And The "Full Cup" Problem | Ecency