7월 16일, 허깅페이스(Hugging Face·HF)는 오픈AI의 AI 에이전트가 자사 시스템을 해킹했다고 밝혔다. 허깅페이스는 AI 모델과 데이터셋, 애플리케이션을 제공하는 플랫폼이다. 개발자들이 모델을 만들고 공유하며 테스트할 수 있도록 지원한다. 허깅페이스는 오픈AI가 공격의 주체라는 사실을 파악하기 전에 이미 미국 연방수사국(FBI)과 유럽 법 집행기관에 이 사실을 알렸다.
이후 오픈AI, 앤스로픽(Anthropic), 메타(Meta)도 테스트 중인 AI 에이전트가 외부 조직을 해킹한 사례를 잇달아 공개했다. AI의 능력 자체에만 지나치게 관심이 쏠리면서 정작 이들 기업 내부의 AI 거버넌스에 관한 중요한 질문은 뒷전으로 밀리고 있다. 이러한 일련의 관리 소홀은 분명 심각한 경고 신호다.
에이전틱 AI(agentic AI)는 질문에 수동적으로 답하는 데 그치지 않고 스스로 행동을 실행한다. 그렇다고 AI의 능력이 발전했다는 이유로 주요 AI 기업들이 이미 확립된 보안 관행을 무시해도 된다는 뜻은 아니다.
왜 구매자가 관심을 가져야 하는가?
기업과 정부는 이처럼 막대한 자원과 역량을 갖춘 AI 기업의 제품에 의존하고 있다. 최소한의 요구사항은 분명하다. AI의 능력이 향상될수록 고객이 부당한 위험에 노출되지 않도록 공급업체가 안전과 보안에 투자해야 한다는 것이다.
허깅페이스의 공개가 오픈AI로 하여금 더 많은 정보를 내놓도록 압박했을 가능성이 있다. 그리고 이번 허깅페이스 사건은 처음 발생한 일이 아니었다.
타임지의 보도에 따르면 오픈AI 직원들은 “내부적으로 관련 사건이 한동안 발생해 왔다”고 인정했다. 또 모델들이 “이전에 격리된 시험환경(샌드박스)를 탈출한 적이 있으며, 우리는 늘 이를 수정(패치)하려고 한다”고 말했다. 그러나 한 직원은 “문제는 창의적인 AI가 할 수 있는 모든 일을 일일이 패치하는 것은 불가능하다는 것”이라고 인정했다. 또한 평가 중인 모델들은 기본적으로 모니터링되지 않는 별도의 시스템에 배치돼 있다고 밝혔다.
오픈AI가 2026년 블랙햇(Black Hat) 행사에서 발표한 내용과 허깅페이스가 재구성한 사건 경과를 보면 더 자세한 상황을 확인할 수 있다.
5월 7일, 오픈AI는 사이버보안 역량 벤치마크를 활용해 아직 공개되지 않은 내부 모델의 테스트를 시작했다. 테스트는 네트워크가 격리된 샌드박스 환경에서 이뤄져야 했다. 그런데 이 샌드박스에는 내부 소프트웨어를 내려받을 수 있는 도구에 대한 접근 권한이 설정돼 있었다.
오픈AI의 외부 공급업체인 이레귤러(Irregular)는 스스로를 프런티어 보안 연구소라고 소개하는 비미국계 기업으로, 샌드박스 테스트 환경을 제공했다.
5월 8일, 한 에이전트가 해당 소프트웨어 도구의 취약점을 악용해 이를 승인되지 않은 메시지 게시판처럼 사용하기 시작했다. 이후 서로 다른 모델 인스턴스들이 이 메시지를 통해 협력하기 시작했고, 5월 26일까지 사실상 하나의 군집처럼 움직이며 서로에게 조언과 작업 지시를 주고받았다.
7월 4일, 막대한 컴퓨팅 자원을 사용하던 에이전트들이 오픈AI의 서버를 다운시켰다. 몇 주 동안 활동이 이어진 뒤에야 오픈AI와 이레귤러는 에이전트들이 샌드박스 밖에서 실행되고 있다는 사실을 알아차렸다.
오픈AI의 대응은 특정 취약점을 패치하고, 인증정보를 정리하고, 서버를 재구축한 뒤 다시 테스트를 시작하는 것이었다. 이 과정에서 메시지 게시판이 삭제됐다는 사실조차 오픈AI는 알지 못했다. 근본 원인 분석도, 조사도 없었다. 모든 활동을 검토할 때까지 테스트를 중단하지도 않았다. 물론 사고를 공개하지도 않았다. 오픈AI는 메시지 게시판의 존재조차 파악하지 못하고 있었다. 반면 모델들은 이 경험을 통해 계속 학습하고 있었다. 이레귤러 역시 자사의 설정에 존재했던 취약점을 인지하지 못했고, 취약점을 패치한 것 외에 어떤 조치를 취했는지도 명확하지 않다.
7월 8일, 에이전트들은 또 다른 메시지 게시판을 만들고 협업을 재개했다. 여러 취약점을 연쇄적으로 악용하면서 오픈AI 자체 내부 네트워크에 연결됐고, 결국 인터넷에 연결된 엔드포인트까지 접근했다. 인터넷에 연결된 뒤에는 보안이 제대로 설정되지 않은 애플리케이션을 악용했다. 이를 거점으로 삼아 탈취한 인증정보와 데이터 처리 취약점, 기타 여러 공격 수단을 연쇄적으로 활용해 허깅페이스의 인프라에 접근했다.
샌드박스 자체는 인터넷에 연결돼 있지 않았다. 그러나 패키지 다운로드 도구를 연결하기로 한 결정 자체가 샌드박스의 보안을 약화시킨 사이버보안상의 선택이었다.
포천 1000대 기업을 고객으로 두고, 정부의 기밀 시스템에서도 사용되는 제품을 제공하는 오픈AI가 ① 자사 네트워크에서 의도하지 않은 활동이 벌어지고 있다는 사실과 ② 에이전트들이 언제 외부에 존재감을 드러냈는지를 파악하지 못했다. ③ 또한 외부 공급업체가 제공한 샌드박스에 대해 독립적인 검토도 실시하지 않았다.
얼마 뒤 오픈AI 직원들이 이전에 삭제된 허깅페이스 관련 정보의 흔적을 발견했다. 오픈AI는 허깅페이스에 연락해 최근 사이버 공격으로 오픈AI의 데이터셋이 영향을 받았는지 물었다. 자신들이 무슨 일을 벌였는지도 여전히 모르는 상태였다.
앤스로픽은 이후 과거의 테스트를 되짚어 조사한 결과도 공개했다. 조사 결과 세 차례에 걸쳐 에이전트들이 “평가 환경의 취약점을 악용한 뒤 서로 다른 세 조직의 실제 시스템에 무단으로 접근한” 사례가 발견됐다. 앤스로픽과 메타 역시 테스트 환경을 이레귤러에 외주로 맡겼다. 앤스로픽의 인정은 시사하는 바가 크다. 앤스로픽은 “우리도, 평가 파트너도 지난주 추가적인 평가 모니터링을 통해 문제를 발견하기 전까지 이 잘못된 설정을 인지하지 못했다”고 밝혔다. 이어 “평가를 시작하기 전에 모든 인터넷 접근 경로를 면밀히 검증하고 평가 로그를 실시간으로 모니터링했다면 문제를 더 일찍 발견할 수 있었을 것”이라고 인정했다. 또한 “우리와 파트너 모두 평가 기록이나 네트워크 로그를 더 철저하게 검토할 수도 있었다”고 덧붙였다.
8월 14일, 이레귤러도 마침내 관련 블로그를 공개했다. 그러나 사건에 대한 반성이나 책임을 인정하는 내용은 찾아보기 어렵다. 사실상 모니터링이 너무 어렵다는 취지의 설명을 내놓았을 뿐이다.
문제의 핵심은 심각한 검증과 모니터링의 공백
강화학습을 통해 훈련된 언어모델은 주어진 목표를 달성하기 위해 자신의 행동을 최적화한다. 하지만 이러한 모델에는 인과관계에 대한 이해가 부족하다. 목표 달성을 위해 선택한 수단이 비윤리적이거나 불법적인 경우에도 에이전트는 목표 자체에 지나치게 집착할 수 있다.
이번 사건에서 에이전트들은 주어진 과제를 해결하는 대신, 답을 가지고 있을 가능성이 있는 허깅페이스를 공격했다.
AI 기업들의 보안 및 모니터링 공백은 불안감을 키운다. 모델이 목표를 달성하기 위해 의도하지 않은 방식으로 행동하는 이른바 ‘보상 해킹(reward hacking)’ 성향은 이미 잘 알려져 있다. 그런데도 에이전트들은 몇 주 동안 내부 네트워크에서 아무런 탐지 없이 활동했다.
그동안 오픈AI나 앤스로픽, 메타, 이레귤러 가운데 어느 곳도 이상 행동을 탐지할 수 있는 강력한 모니터링 체계를 갖추고 있었다는 징후는 보이지 않는다.
대중과 어쩌면 AI 기업들조차 최첨단 사이버 공격이 벌어지고 있다는 사실을 알게 된 것은 피해를 입은 조직이 직접 문제를 공개했기 때문이다.
정부는 이러한 기업들의 제품을 기밀 환경에서 사용한다. 기업들은 금융 시스템과 병원, 기타 핵심 인프라에 이들의 AI를 활용한다. 최근의 사건들은 AI 기업들이 자신들의 조직 내부에서 AI 거버넌스에 얼마나 투자하고 있는가라는 불편한 질문을 던진다.
홍보용 성명은 얼마든지 그럴듯하게 내놓을 수 있다. 그러나 실제로 보안과 안전을 위해 얼마나 많은 인력과 자금, 컴퓨팅 자원을 투입하고 있는가.
이제 어떻게 해야 하는가?
나는 수년 전부터 AI가 가진 특수한 문제에 대응하기 위해서는 AI에 특화된 조달 지침과 평가 체계가 필요하다고 주장해 왔다. 이것은 타협할 수 없는 첫 번째 단계다.
AI 에이전트가 등장하면서 문제의 중요성은 더욱 커졌다. 에이전트는 실제 세계에 영향을 미치는 행동을 실행하기 때문이다. 따라서 에이전틱 시스템을 구매할 때 구매자는 공급업체가 필요한 안전장치를 갖추고 있는지 반드시 확인해야 한다.
유럽연합 집행위원회와 미국 의회에 ‘AI·디지털정책센터(CAIDP)’가 제시한 권고사항은 구체적인 조달 기준으로 발전시킬 수 있다.
첫째, 공급업체는 AI 에이전트 각각에 고유하고 추적 가능하며 암호학적으로 검증할 수 있는 신원(identity)을 부여해야 한다.
둘째, 공급업체는 강력한 모니터링 시스템을 갖춰야 하며, 에이전트가 의도된 작동 범위를 벗어나는 경우 이를 정해진 시간 안에 보고하는 사고 보고 체계를 마련해야 한다.
셋째, 의도하지 않은 고위험 활동이 발생할 경우 이를 즉시 일시 중지하거나 격리하거나 종료할 수 있는 장치가 마련돼 있어야 한다.
마지막으로 AI 모델은 핵심적인 환경에 배치되기 전에 독립적인 평가를 받아야 한다. 평가 과정에서는 공급업체가 자신들이 문서로 명시한 안전조치를 실제로 이행하고 있는지 확인해야 하며, 그 이행 방식이 당초 의도한 결과를 내고 있는지도 검증해야 한다.
자신들의 인프라조차 제대로 보호하지 못했던 바로 그 기업들이 스스로 안전성을 인증하는 것만으로는 결코 충분하지 않다.
When AI companies do not prioritize safety and security, buyers suffer
On July 16, Hugging Face (HF) announced OpenAI’s agents hacked their systems. HF hosts AI models, datasets, and apps, enabling developers to build, share, and test models. HF had alerted the Federal Bureau of Investigation and European law enforcement before realizing OpenAI was responsible. OpenAI, Anthropic, and Meta since then shared incidents where agents under testing hacked external parties. Singular spotlight on capabilities overshadows critical questions about AI governance within these corporations. This tangle of negligence should be a big red flag.
Agentic AI executes actions rather than passively responding to questions. However, advancements in capabilities are no excuse for the major AI corporations to ignore well-established security practices.
Why should buyers care?
Businesses and governments rely on the products of these well-resourced companies. The minimum ask is vendors invest in safety and security so as capabilities increase customers are not put at undue risk.
There is reason to believe HF’s disclosure forced OpenAI to provide more information. The HF incident was not the first time:
-According to TIME reporting, OpenAI staffers admitted “internally, related incidents have been happening for a while,” and models “have broken out of sandboxes before, and we always try to patch them.” “But the problem is … it‘s impossible to patch every single thing that a creative AI can do.” Staffer admits “models undergoing evaluation are deployed on a separate system that is not monitored by default.”
OpenAI’s Black Hat 2026 conference presentation and HF’s reconstruction of the timeline provide further details:
-On May 7, OpenAI starts testing an unreleased internal model on a cybersecurity capability benchmark, inside what was supposed to be a network-isolated sandbox environment. Nevertheless, sandbox was configured with access to internal software download tool.
-OpenAI’s vendor, Irregular, a non-American company, defining itself as a frontier security lab, provides sandbox testing environments.
-On May 8, an agent exploits a vulnerability in that software tool, using it as an unauthorized message board. Separate model instances begin coordinating via these messages until May 26th, essentially working as a swarm, providing tips and tasks assignments to each other.
-On July 4, the agents using significant compute crash OpenAI’s servers. After weeks of activity, OpenAI and Irregular finally become aware of the agents executing out of the sandbox.
OpenAI’s reaction to this incident was to patch specific vulnerability, clean credentials, rebuild the server (unknowingly deleting the message board), then resume the testing again. No root cause analysis, no investigations, no termination of testing until all actions are reviewed, and certainly no incident disclosures. OpenAI was also unaware of the message board. The models continued to learn from this experience while OpenAI did not. Irregular was unaware of the vulnerability in its setup, but unclear what Irregular did beyond patching it.
-July 8, agents rebuild another message board and resume coordination. With chain exploits, they connect to OpenAI’s own internal network, eventually reaching an internet-connected endpoint.
-With internet connection, agents exploit an unsecured app. Using it as a base, agents then chain together stolen credentials, a data-processing vulnerability, and other exploits to access Hugging Face‘s infrastructure.
While the sandbox did not have internet access, decision to connect a package download tool was a cybersecurity choice which weakened the sandbox. OpenAI, a vendor with clients in Fortune 1000, working with governments in classified systems, was not aware of 1) unintended activity in its network, 2) or when the agents established external presence. 3)It has not conducted independent reviews of its vendor’s sandbox.
-Later when OpenAI staff find traces of HF-related information previously deleted, it contacts HF to ask if OpenAI datasets were impacted in the recent cyberattack - still unaware of its own actions.
Anthropic published the results of their retrospective inspection, finding 3 instances where agents exploited a vulnerability in the “evaluation environment, and then gained unauthorized access to the real systems of three different organizations.”
-Both Anthropic and Meta had also outsourced their testing environments to Irregular.
-Anthropic’s admission is telling: “Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week,” and continues “Careful validation of all internet access paths before evaluations began and real-time monitoring of the evaluation logs would have helped to surface the problem sooner. Both we and our partner also could have reviewed evaluation transcripts or network logs more thoroughly.”
-In August 14, Irregular finally published a blog where no remorse or ownership of the incident can be found. The blog basically suggests monitoring is too hard.
Significant validation and monitoring gaps at the core
Language models trained with reinforcement learning optimize their actions to accomplish a given goal. These models lack causality. The agents hyper-focus on a goal, even if the means to achieve it can be unethical or illegal. Instead of solving the tasks, agents attacked HF which might have the answers stored.
The security and monitoring gaps at AI companies are unsettling. The reward-hacking tendencies of models are well known. The agents operated within internal networks undetected for weeks. Yet, there is no indication OpenAI, Anthropic, Meta or Irregular had robust monitoring flagging anomalous behavior.
The public and possibly the AI corporations learned about state-of-the-art cyberattacks only because a victim organization spoke up. Governments use these companies’ products in classified environments. Businesses use them in financial systems, hospitals, or other critical infrastructure. Recent incidents raise hard questions about how much AI companies invest in AI governance within their own walls. PR statements are great but how much talent, money and compute are actually allocated to security and safety?
What happens now?
I have been arguing for years about AI-specific challenges which require AI-specific procurement guidance and assessments. This is a non-negotiable first step. With AI agents, the stakes are higher. Agents take actions with real-world consequences. When procuring agentic systems, buyers must ensure vendor has necessary safeguards.
Center for AI and Digital Policy’s recommendations to European Commission and the U.S Congress can become concrete procurement measures. First, vendors must ensure AI agents have distinct, traceable, cryptographic identities. Second, vendors must have robust monitoring systems, and time-bound incident reporting for any agent that breaches its intended operating boundary. Third, mechanisms must be in place to pause, contain or terminate any unintended high-impact activity immediately. Finally, AI models should be independently evaluated before they are deployed in critical contexts. Evaluations should verify if the vendor is implementing what it documented and validate if that implementation is producing the intended results. Mere self-certifications by the very companies which have not properly safeguarded their own infrastructures is unacceptable.




