As usual, I've ended up with a degree of mash and chaos in my head, which happens when I try to make sense of some idea, and then I straddle to learn further about its context and all the things that surround it. And at this point, I'm so flooded with new information that, in the end, it comprises an incomprehensible heap, where I cannot figure out how particular things are connected to each other, or to arrange all those facts and concepts into a neat composition and system.
The initial concept is to merge blockchain tech and AI tech for the benefit of AI developers and with an eye on all the possible advantages of implementing blockchain wherever it's possible. Maybe, the problem is that I still don't envision in my mind the scope of possible blockchain applications, the way the evangelists of this tech see it. For me, it includes several handy features: An architecture, in which all data is stored as a chain of locks, sequentially attached one upon another, where it's impossible to change anything in this construction after it's done. This has both advantages and disadvantages, as I see it. The advantages are: The data is secure and cannot be tampered with. Therefore everybody can be guaranteed that any data extracted from this database is in its initial form, in which it had been submitted. Plus, in accordance with blockchain initial purpose (transaction ledger), every record includes the signature, identifying the source. It might help, if necessary, to confirm ownership and authorship of data, if there are issues, related to copyrights. Which theoretically makes it viable to use such structure to store information which ownership should be protected. In fact, one of the applications can be a database for patents, scientific and creative work, requiring copyright protection.
Still, I don't see how blockchain can prevent anybody from stealing the information it contains. For example, somebody would decide to pool into it some valuable blueprints of some new breakthrough and secret tech. Since it's an open and decentralized database, which, simply according to its protocol, gets copied by the multiple computers just in order to fulfill the protocol. In such an open environment it's not clear how to protect the information from just being extracted and published somewhere else. Well, unless it's encrypted. Yes, for example, if I tried to find a way to create a common information pool for a number of scientific organizations willing to share information with each other, this architecture would do. The sensitive information can be encrypted using private keys owned by those scientific institutions. They would share those keys with each other, thus giving each other access to information, they think, they are willing to share. The things to worry about here are possible leaks of private keys.
Another problem is that information on blockchain really cannot be altered after it's been posted there. For example, if we take some blueprint or, say, novel, which is in the process of editing when it's constantly changed and updated. It means each small update would lead to adding a new record to this database. Plus all the records would need to have a common identifier, showing that those are the revisions of the same thing. I guess, there is no problem with that, apart from that it can generate a really huge amount of information, most of which would be garbage. And, due to blockchain specifics, no garbage (or data that's not relevant anymore for instance) can be simply purged from the database
It leads to yet another problem: The blockchain database is supposed to be light, not weighed down by massive loads of data. It's essential since the protocol implies that the whole database is replicated among many computers (nodes), many of which might not have a lot of disk space or high network bandwidth. The way it was initially planned, with blockchain being a ledger for monetary transactions, this ledger would never reach a really significant size. I think I remember reading some calculations in the original White Paper, predicting that even decades from now Bitcoin blockchain would still be small enough for this replication protocol to work without a hitch. But when people begin to talk about storing huge amount of data on the blockchain, it becomes unclear, how it's supposed to be decentralized. Because decentralization in blockchain virtually means that ALL the information is get copied to many computers (nodes) on the network. And it happens literally every time a new butch of transactions is added to the chain. The new version of the chain with all new signatures and hashes is replicated by all the full-nodes. ( computers receiving and storing the full copy of the ledger, hence the name) The full-nodes check the validity of new changes and, after reaching a consensus among themselves whether this new version is valid, either approve or reject the updated ledger.
Well, the massive amounts of data just don't fit in this picture. First of all, there is a limitation on block size, which prevents storing anything large on the blockchain. Even if we increase this limit significantly, maybe ten-fold or hundred-fold, it'll still be an obstacle for trying to push into this database something really big. (What if I want to store some movies in HD format there) If we eradicate the block size limit, many mechanisms fundamental to classic blockchain architecture, such as mining, won't work anymore. So it'll be a different kind of distributed ledger architecture. What's remarkable, on Ethereum blockchain the storage space is quite expensive, and it's not accidental. Since Ethereum is less strict, regarding the things (code and data) that can be stored on it, the high cost of storage space is a safeguard against swamping the blockchain with massive amounts of garbage.
And the most important problem, there'll be no way to keep this architecture decentralized. If blockchain database is measured in Terabytes and Petabytes, the very protocol that makes this system decentralized won't work. Such amount of information simply cannot be replicated by multiple nodes. (many of them just average workstations with average hard drives and Internet bandwidth.) It will only be possible to store it on similarly massive piles of hardware in the huge data centers with the breeze of powerful cooling systems in their aisles.
So, well, we have a project, aiming to create a marketplace of Big Data for AI developers. And the marketplace of AI algorithms for owners of Big Data, who need those algorithms to put their data to some use.
There are other questions that make me wonder: Why the owners of data would be willing to share it? And are there really the valuable data sources not owned by big companies like Google or Facebook (that have their AI algorithms as well)? Plus, how the platform would guarantee that this valuable data won't be leaked or stolen? How would it guarantee that the data provided by some sources isn't rubbish? How is it going to store such huge amounts of information, considering the limitations of blockchain architecture discussed above? And how is it going to price that data, considering that likely it's assumed that the buyer doesn't see the data until he's bought it? And how to prevent the further proliferation of data after it's purchased and left the marketplace, considering that it might be exclusive or sensitive, or the seller would like to get the maximum advantage, selling it to as many buyers as possible? Basically, this concept of sharing in the open space available for everybody doesn't align well with the concept of data, which is currently considered as valuable as gold, and therefore is held in safe storages behind multiple locks and firewalls in an encrypted form.
I envision what currently might be used as data sources for training AI algorithms. Those are either data sources publicly available from the beginning, which might have been the case with the training of Chess and Go neural networks. If this is the case, there is no problem for anybody just to access those sources and train their own algorithms. Or those are highly sensitive, secret, classified, commercially valuable, expensive (pick your choice) sets of data, which public unavailability is defined by the character of that information. Consider, for example, police databases and court cases, that can be used to train the AI-based judicial and criminal expert systems. I don't think that kind of information can be under any circumstances be shared on the open marketplace. Unless such kind of sensitive information is shared with affiliated organizations, developing AI algorithms, which those institutions - owners of sensitive information trust. I think; everything there is going to be conducted via the closed private channels of communication anyway. Neither of that will have anything to do with the open Big Data market.
And, by the way, it raises a different question: How in such an open Big Data marketplace anybody can be sure that they are not dealing with the malicious actors. For example, under the pretense of buying data for purposes of training AI, they can buy data for some nefarious goals, like blackmail, terrorist activities, cybercrime, etc.
As far as I understand from the quite vague White Paper of the project, data won't be transferred to the project's network; rather their blockchain would provide access and APIs to respective data sources. Which is just as well, since this at least eliminates a mind-blowing question of how those Terabytes and Petabytes are going to be flung across the Internet, and how blockchain is going to handle this. (The answer is, no way.) Well, even not taking into consideration the fact, that blockchain is not designed to store the huge amounts of data, I doubt if this project would be able to finance the data centers capable of storing such amount of information. Even if they'll conduct a successful ICO. And there is another question, why blockchain? Like if this is a concept of selling the information sources, is there any reason to store all the proceedings of such system on the blockchain. It can be any kind of database. Plus openness of blockchain architecture is not particularly convenient for this kind of project. Namely, at least part of the deals might need secrecy. Those are the situations when the fact of the deal is revealing enough information in itself. On blockchain, there are no secrets though.
Although, if I look at this project from a positive, dreamy, and naive vantage point, it looks kinda cool. The AI scientists need the huge amounts of data, which are stored somewhere behind closed walls and are not available to anybody just because of the absence of marketplace where data owners can trade it. And data owners cannot do anything with that data because they don't have necessary tools - AI algorithms and neural networks, (why did they accumulate all that data in the first place then?) that they, in turn, would be willing to buy on such marketplace. So it's sort of a place where AI researchers and Big Data owners meet each other.
The thing that makes me doubt the whole idea is: This Big Data trading business is really not that simple. It includes many complex issues, including security, information sensitivity, political and ethical aspects. It might include a complicated web of agreements and amendments, regarding how this data can be used, distributed and stored. All this kind of stuff. As I see it the Big Data sale is a unique and subtle process, accompanied by heaps of legal paperwork, produced by brilliant lawyers. Basically, because Big Data is not peanuts. Well, at least if we talk about valuable data. If the data can be shared with anybody within a framework of common rules of an open auction, without any subtleties, with regard to the data specifics, it's not likely that it is some valuable data. Most likely it is rubbish. Plus, I still cannot wrap my head around the idea of people and companies that have accumulated huge amounts of valuable data but never tried to sell it, although, they cannot use it due to the lack of respective algorithms and neural networks. (Why wouldn't they buy those things or at least outsource this task to the researchers who have those things) Plus, I have a feeling that the worth of those datasets might be too significant for them to be traded in such environment of an open auction.
Ok, there are some things to ponder, while I make a pause and let all the information settle in my head.
First, data exchanges and markets exist. It's not like they are particularly successful at the moment, but the fact is, that a number of companies collect various data with the intention to sell it, and there are existing platforms, facilitating those sales. An example is DEX that also has a relation to company X, which prospects we are trying to assess. DEX is not new on the market, and after facing all the problems related to the centralized way of data, trading hopes that blockchain aka distributed ledger platform would allow to solve them. Mostly the problems with trust among sellers or buyers, and the perception of data owners that they lose control of their data. Something like that. (Probably I'll need once again to reiterate those points in my head to keep them at hand)
Company X is in the cooperation with one major AI foundation, promoting the development of AI, plus helping and encouraging independent developers. So the power of AI wouldn't be in the hands of a handful of major big players like Google, Facebook, and Amazon. Who, in addition, control a significant part of data essential both for marketing decisions and for training AI algorithms that would eventually be able to do a lot of important stuff, including making those strategic marketing decisions.
The concentration of Big Data in the hands of a few major players is Evil since it puts up the barriers preventing new IT startups and companies from development. Plus it allows big IT companies to analyze the situation so well that they can eliminate the potential rivals on the early stages of their development. (Facebook buying WhatsApp case in point) So to prevent those powerful from getting omnipotent, it's a socially important goal to find a way to make the big data arrays available to independent and fledgling startups. (How to wrestle that data from the hands of big boys is another question)
There is an existing implementation of a hybrid database solution, combining the best properties of blockchain architecture and the industrial level databases, designed to store actual data. The thing is called BigchainDB. (Or something like that; I need to check) In other words, it's actually possible to store huge, industrial-level size amounts of information on blockchain after all. (I won't get into technical details about this project, there is no time to study it. What I can tell with certainty, though. It doesn't use proof-of-work since it doesn't have its built-in currency. I might assume that it doesn't require full-scale replication for reaching the consensus either. So most likely it supports only lightweight replication of headers similar to how it happens with "light nodes" on Bitcoin blockchain. And it doesn't include mining. There are obvious implications for the reliability of the security model of such system, but I'm definitely not going to dive into it at the moment. Suffice to say, that the architect of the project is pretty sure that it's going to work, and this architecture is secure enough.
The X project cooperates with that BigchainDB (or whatever it's called) project and bases its core architecture on it. So it can actually store the big amounts of information on the blockchain. (Or should we rather call it a distributed ledger) Although, the security model of the system is not as good as that of Bitcoin blockchain.
Blockchain technology can actually solve a number of existing problems, accompanying the use of massive amounts of data. Like the reliability of data sources (signatures ftw), whether the data is in its initial state and hadn't been tampered with, the verifiable trails, allowing to check everything that happened to the data in the process, all updates and stuff. without the need to trust some central actor, maintaining those audits (they are self-verifiable)
The constantly developing AI brings us to the dangerous territory when it can conquer the humanity, take control of all the resources and take away all the jobs. It's a path to extinction, literally.
It's essentially important to implement UBI
So, once upon a time...
The X project is created by the founders of DEX and BigchainDB with an aim to address the problems of open marketplaces for Big Data. There are existing data marketplaces with their common problems, such as: There are difficulties of revealing the provenance of data. For data providers, it's necessary to trust buyers, for example, that they are going to use received data in accordance with terms and conditions of an agreement. Like they are not going to share this information all over the Internet, especially if those are some medical records, provided for training of some AI-based medical expert system. So DEX is some project that worked in this area of data marketplaces, and after facing all those difficulties got hooked on the idea of a decentralized marketplace, based on the blockchain tech. (a specialized version combining the features of blockchain and traditional databases - BigchainDB ) Ok, here we got to describe this field of big data marketplaces a bit, how it's riddled with problems, and how the data providers feel that they lose control of their data when they post it on those marketplaces. Sure, something has to be done about it.
So, speaking of expert systems, there is another talking point in relation to project X. Project X pitch itself as an ultimate solution for AI developers, who suffer from lack of sufficient data necessary to train their systems. While that data might be accumulated in databases of various companies that cannot use it because they lack respective algorithms. Or, maybe, simply that data and whatever they can extract from it is useless for them, who knows. In any case, data is accumulated in silos and cannot be put to use because there are no marketplaces where it can be sold to those who could actually use it. I mean those developers of AI algorithms. Or rather such marketplaces exist, but they don't satisfy the requirements such data trade poses. Because data can be sensitive (like medical records), data can be only provided for some specific use; data can only be provided on a condition that the buyer can guarantee that it won't be leaked afterward. So those are the problems that current data marketplaces cannot solve. And the data providers feel that they lose control of their data.
Here it would probably be cool to mention all the positive implications of rapid AI development. Like it can create art, music, puzzle out complex problems, inundating humanity since its inception, decipher human genome, etc. Also, it can decimate the humanity, take over the Earth, appropriate all the resources, steal all the jobs, in other words, effectively bring the human civilization to the untimely end. But this is a negative way to look at the singularity.
Speaking of singularity, the company X is collaborating with the Singularity project to bring together AI algorithms and data from blockchain powered marketplaces, capable of making those AI algorithms powerful. The point is, the Singularity project trusts that the company X can deliver that data. Here it might make sense to talk a bit about the problems and prospects of AI systems, and how vital it is to have the huge datasets to train sufficiently advanced AI, applicable in real life. From image and sound recognition to self-driving cars, smart IoT, stuff like that.
Another talking point is how big and evil corporations are monopolizing both Big Data and AI development, increasing their power at the expense of small and fledgling startups and inventors, withering in their shadow.
Another talking point is what's BigchainDB, and how it combines the advantages of traditional databases, designed for storing industrial size amounts of data, and blockchain architecture, making such database decentralized and tamper proof.
So this is all about the introductory paragraph: Massive amounts of data hoarded behind firewalls, unable to find its way to the customer. Problems of centralized data marketplaces, and the subtitles of data trade. Problems of AI algorithms that need adequate data sources to reach the level when their output is precise enough for them to find the real-life application. The evil of data and AI being monopolized by a few big corporations. Prospects of AI. BigchainDB and how cool it combines blockchain and the features of normal databases with their SQL queries and adequate performance.
So, speaking of the project X itself. First, it's built on BigchainDB architecture and, as a result, is capable of storing significant amounts of information. It uses the token system to incentives its participants. It doesn't use "heavy" mining with proof-of-work confirming the transactions. Rather all the information is confirmed by the special members (nodes), who get rewards by maintaining the whole platform. So the platform will include the participants with different roles: Data suppliers and Data consumers. (Which is self-explanatory) Data mashers, who do some magical manipulations with data, making it more valuable. Like sorting, filtering, labeling, normalizing, converting it to adequate format, whatever. Data keepers, who perform all those maintenance activities in the system, allowing it to function without a hitch, similar to miners in the Bitcoin blockchain. And the last group that somewhat represent authorities, like they uphold the order in the system and responsible for preventing the anarchy. Something like that.
So there are several interesting things. The system doesn't store data; rather it provides APIs for the existing data sources. It stores the information, describing the sources, confirming its status as intellectual property, etc. Also, it somehow guarantees that the data, which the consumer who bought the information receives, is the same information he has paid for, like it hasn't been altered, and it hasn't been replaced with something different. (Probably using hashes, checksums, and other stuff like that)
Moreover, the platform itself is not a data marketplace. Data marketplaces are built on top of that platform, either using a supplied pattern, helping to design marketplace to fit the standards and interfaces of the platform, or not. Nonetheless, the system of data pricing is built into the system core. So the data providers and data consumers operate on the level of data marketplaces, and data marketplaces use the core functionality to price its data and to confirm the data quality and validity (more about that below)
So how the system determines that the data supplier doesn't supply garbage? There is another interesting element of the protocol, related to that. The supplier stakes tokens on the dataset he provides, and the nodes (probably keepers) vote whether they think the data is garbage or not. They also stake tokens on the outcome, and, after the vote, those who's voted against the majority lose their stakes. So the quality of data is confirmed by this form of proof-of-stake protocol. So, it's assumed, that such system will guarantee that there'll be no garbage in the system.
So, eventually, the company X was created with a support of DEX and BigchainDB, combining their knowledge about data marketplaces, AI, Big Data, and how to adapt blockchain architecture for storing Big Data. Plus, project X is collaborating with the various projects, related to data licensing and Intellectual Property (IP), IoT, various applications of AI, in other words, the whole constellation of crazy people.
Trying to jumpstart the brain. Trying to clear the fog and cotton from my head, while the deadline is flashing alarm signals behind me, while I feel powerless to do anything about that. Because the thoughts just float around the insides of my skull not being able to stick together in solid construction. And the flashing alarm is the worst distracting factor, causing my thoughts to fly away like a bunch of frightened birds, leaving my head absolutely empty with the scraps and scratches of ideas and words, drifting aimlessly around. And I jump frantically, trying to catch some of them and glue them together and organize them into something coherent, but they dodge me or dissipate in my hands. Then they coalesce somewhere else, and I continue my efforts to not to slide into sleepiness and numbness when the words follow the thoughts and disappear somewhere in the black hole of fatigue.
The scenery consists of angry wind and dusk. There is some source of energy inside, sputtering heat like a broken nuclear reactor, but I cannot channel it, and it gets wasted. The heat materializes in the form of words and images. The only thing, I cannot direct it into accomplishing a single meaningful thing I need to accomplish. Maybe, I thought about it too much and burned out. Maybe, all those misgivings and missed deadline produced this mental paralysis, like I've convinced myself that it's so important, that now it turned into some anti-magnet pushing me away every time I try to approach it. Or maybe I cannot produce the phrases that I would consider good enough, convincing enough.
It's still all about data and data exchanges, and how AI needs data, and how data exchanges can provide that data, but they fail to do that eventually. And how banal all of this sounds after so many repetitions. I need something new, something beautiful like a lightning bolt or a trailer for a thriller. Something that would just sound cool. But mostly I come up with reiterations of the same talking points, and after so many repetitions it feels almost painfully commonplace. There should be some strong introduction. Maybe with some examples.
Like "In the world of AI, scientists at some point realized that the quantity of data is far more important than the quality of the models. And starting from that point, everybody started chasing data, trying to catch it, badger it from the dens and lairs and hiding places."
Like "We all strive to wrestle the data from the greedy bony hands of big corporations holding it, trying to create a super-intelligence that will enslave us in the end. And to do that we need those free and decentralized marketplaces where people will have the liberty to share, to give and take. And the generation of young and fledgling AI scientists will emerge creating a flurry of smart algorithms that would finally parade through the crumbling alleys among deteriorating humanity" Something like that.
Maybe, I need to find a way to convert the dry technical details (incomprehensible for anybody not completely tuned in, including the persons under the alcohol influence or those who just have recently woken up) into fiery and inspiring marketing slogans.
"The Company X, using the ingenuity of decentralization and the power of masses, wrestles the Big Data from the hands of the few chosen and gives it back to the people. Adding breakthrough features of democratized control over the network by a group of honest blokes, motivated by the material rewards and their stakes that they are unwilling to lose. And the smooth flow of data, getting filtered, processed, converted, and improved in the process. And how the company Y observed for years the troubles and frustrations of data markets until one day those guys met and after a fair amount of camaraderie they came up with a new cool solution. And how the company Z saw the potential of the blockchain but regretted that it didn't have cool features of normal databases, namely an ability to store and manipulate data, SQL queries, and stuff like that. Then they all came together, and the folks from an AI foundation joined, and they made a sorta really cool thing."
Maybe I should look at all this as some symbol, a geometric shape with vertices. The vertices are the key points: AI, Big Data, Big Data trade and marketplaces, blockchain and databases, Intellectual Property and licensing. Then I can draw multiple lines, connecting those vertices, and build a narrative along those lines.
Another problem is the cotton numbness of my brain and the fact that the thoughts currently tend to escape and fly away. Maybe I can keep this geometric shape in my head like an anchor, all the time, and it will prevent me from floundering. There is a flow of energy, which is strong like a stream of liquid metal, but I don't have the mold to shape it, so it just splatters around and gets wasted.