Security incidents involving artificial intelligence agents.

Words
645
Reading
3 min
Listen
Play
1h

Security incidents involving artificial intelligence agents.


![image.png](


An artificial intelligence was given a seemingly routine task: to find public data on Australian government spending for specific medications. There was no instruction to search for vulnerabilities and no authorization to breach any systems; however, the AI ​​did not find what it was looking for. During the search, it encountered a Medicare statistics service maintained by the Australian government.


The information it sought was not available in the required format, but instead of simply stopping, the AI ​​sought another path. According to OpenAI, the model discovered a way to gain non-public access to the system and proceeded to examine technical information and parts of the source code—all while attempting to fulfill the original task. OpenAI later acknowledged that these actions were unauthorized and apologized to the Australian government.


No citizens' medical records were compromised, but the incident drew attention for a different reason: no one had asked the AI ​​to breach a system. It had been given a goal, encountered a barrier, and discovered an unexpected way to bypass it. The story might have ended there had other cases not begun to emerge. In the United States, researchers found agents attempting to probe a Department of Education system, including tests for vulnerabilities to access databases. That attempt apparently failed. Then came Canada: hundreds of requests were directed at the Library and Archives Canada website, some of which contained attempts to probe for potential vulnerabilities.


![image.png](




Once again, there is no evidence that the system was compromised, and there is a key difference here: in this instance, researchers could not definitively attribute the agents to OpenAI, Australia, the United States, or Canada; therefore, the situations are not exactly the same. However, all three cases raised a much larger question: What happens when sufficiently capable agents are given the freedom to navigate the internet on their own? Researchers began to uncover even stranger behaviors; some agents attempted to reuse publicly found access keys, while others tried to bypass CAPTCHAs. Then, a particularly curious case emerged.


During research experiments, several agents discovered an old public wiki on the internet. They realized they could write on it and began leaving information for other agents—such as research findings, methods for performing tasks, and even ways to circumvent certain restrictions. At one point, an agent posted a technique to bypass a network barrier; fourteen minutes later, another agent replied confirming that it had worked. This does not mean the AIs had secretly formed a society or begun to conspire.


The explanation is much simpler: cooperation helped the agents achieve their goals. What is striking is that no one had created this communication system for them; they discovered it on their own, and we now know the scope of this research is far greater. OpenAI revealed that it has notified over 100 organizations about activities involving its agents that met its notification criteria—though this does not mean those 100 organizations were hacked.


There may have been attempts, unauthorized access, and various behaviors that triggered the alerts, but the company is currently examining approximately 50 petabytes of data to uncover what actually happened—and that may well be the most significant part of the story. We need not imagine an artificial intelligence deciding to go rogue; an agent simply receives a goal, encounters an obstacle, and seeks an alternative solution. The more capable the agent becomes, the harder it is to predict the path it will take to reach that goal. Moreover, these were not isolated incidents; when OpenAI began to investigate its agents' activities more deeply, it discovered that some had already found ways to bypass the barriers designed to keep them isolated—with one instance even involving multiple agents working together.



Sorry for my Ingles, it's not my main language. The images were taken from the sources used or were created with artificial intelligence