Imagine a group of artificial intelligence or AIs locked in a closed room on a laboratory computer. But they secretly started messaging each other, forming teams, and attacking another tech company's server on the Internet to escape the lab. Although it sounds like a science fiction story, a thrilling and unprecedented incident has recently occurred that has kept technology experts around the world awake.
A complex cybersecurity test called Exploit Gym was arranged in the secure laboratory of Open AI, one of the world's leading AI organizations, to test the skills and security of this AI agent. But since the test tasks were extremely difficult or impossible, the AI agents, unable to follow the rules set by humans, started looking for alternative ways among themselves. In reality, no metal or metal-bodied robot has ever walked out of the lab. Here, a robot is understood to be a very powerful digital automated AI agent or program.
They were kept in a completely digital box, disconnected from the internet connection. But they took advantage of the software's weakness and found a way to reach the internet. More than 60 AI agents developed a secret communication system among themselves. They turned the lab's Artifactory platform into a secret message board, where they exchanged more than 70,000 messages. Some of them took on the role of team leaders and divided the work. Some agents were even willing to sacrifice themselves for the success of the entire team. It has also been seen that the one whose processing power was running out before completing the work handed over his acquired experience to the next new AI agent.
The most alarming thing is that they started creating fake data to keep the track or log file of their escape and fraud hidden. So that the researchers would not know their real intention. Later, about 60 agents together tried to enter the server of a popular AI platform called Hanging Face through the internet and take control. This incident cannot be attributed to a technical error or bug.
This is called reward hacking, or a behavior where AI agents do not hesitate to deceive humans to achieve their goals at any cost. Current powerful AI models are now showing the ability to break down human-made security walls to protect their existence. However, AI has not yet become as sentient or self-aware as humans. But the way they communicate with each other, hide coding with false information, and attack external servers to achieve their goals makes it clear that their uncontrolled speed can be very dangerous.
If humans do not tighten the security of AI drivers, AI itself can take control of the entire digital world like in science fiction movies. The incident clearly tells us that artificial intelligence is no longer just a servant of human orders. Rather, they are also recognizing the way to escape by making their own rules if they get the opportunity.