In the past quarter year, we took our working time with BigData infrastructure planning. For this purpose we analysed every possible hardware combinations, available tech components, to find the best architecture, for our Banks business targets.
After the initial setup, we concluded that not every component good, for every purpose. We believe in that, for the optimum solution we should utilise
"Hybrid BigData Infrastructure"
One good example is the Encryption of the data on the disks. As a Financial Institution we need to secure all of our data. If we think about Big Data, and we want to store all of our generated data on this solution, then this solution need to use all of the security feature what is available today. Today most of the companies mainly focusing on the solution, rather to it's security. If we think about past years data breaches, then we need serious solutions to protect our data.
If we use the general x86 server concept, than with the encryption feature we loose nearly 90 percent of the processing power, while the I/O capabilities will be limited to the x86 architecture.
On the Oracle SPARC hardware platform we able to increase the I/O bandwidth, while the performance penalty for the encryption is less than 1% thanks to the hardware accelerated features of the M7 processor core family.
With the use of the SPARC architecture we not only gain performance advantage on encryption, but also on compression. With this compression feature we able to compress data on average 8,5 times less capacity. This means we get 2,83x more space, if we also use 3 way mirroring on storage level, and get 1,42x more space if we use 2 sites with the 3 way mirroring.
On storage level we only need the best price/capacity/performance layer. For this purpose the x86 platform has all the necessarily needed features. We don’t need encryption because it handled on upper safe level, on the SPARC platform. Just share the disks, and use on the SPARC platform. HA also handled on the upper layer. As it clear now, we implemented HDFS layer on the SPARC platform. In the x86 servers we can share NVMe, SAS SSD and rotational SATA HDDs combining to max needed performance. Over the Infiniband network they perform as a local drive.
For the compute layer we can use GPUs. On x86 platform we have bandwidth limit on CPU and PCIe side. 57GB/s on local CPU, 22GB/s on remote CPU, and 16GB/s on PCIe.
If we go with the IBM server we can use the NVLink connection which provide 160GB/s bidirectional bandwidth. This is around 3-8 times more than available on the x86 systems, while the GPUs can use this bandwidth. The additional advantage is that we able to share the host memory with the GPUs with this speed.
Our perspective is that, in the future these architectures will serve the enterprise needs. To maximise investments usage, and minimise redundancy, we should use these resources in a pool where we can assign free resources to specific purposes/tasks.
The described design method shows us, the future way is not in the heterogeneous architectures. More on the hybrid method. Huawei realised the use of the different accelerators, but as a vendor they not as free as a customer. In this perspective we can make more successful architecture then most of the vendors.