BTC
ETH
HTX
SOL
BNB
View Market
简中
繁中
English
日本語
한국어
ภาษาไทย
Tiếng Việt

GCSA Agent Achieves 91.3% Score on CyberGym, Ranking Among World's Leading AI Cybersecurity Agents

星球君的朋友们
Odaily资深作者
2026-08-29 11:45
This article is about 1835 words, reading the full article takes about 3 minutes
GCSA Agent demonstrates autonomous vulnerability analysis and PoC generation capabilities in globally challenging real-world vulnerability benchmark tests.
AI Summary
Expand
  • Key Takeaway: The Global Cybersecurity Alliance (GCSA) announced that its AI security agent achieved a 91.3% success rate on CyberGym's real-world vulnerability evaluation, entering the leading tier, marking a shift in cybersecurity capabilities from underlying large language models to agentic workflows with full autonomous verification capabilities.
  • Key Elements:
    1. Developed by the University of California, Berkeley, CyberGym includes 1,507 real-world vulnerability test cases covering 188 large-scale software projects, requiring AI agents to autonomously complete analysis, localization, PoC construction, and execution verification within unpatched codebases.
    2. The GCSA Agent operates on Grok 4.5 and Grok 4.6 models, achieving a 91.3% success rate through an agentic security workflow. A task is only deemed successful if the PoC triggers the vulnerability on the vulnerable version and cannot be reproduced on the patched version.
    3. The core capability shift lies in AI needing to enter real execution environments, autonomously completing a full security research loop encompassing vulnerability description comprehension, code retrieval, attack surface identification, hypothesis validation, and PoC iteration.
    4. CyberGym's open experiments reveal that AI agents have already discovered multiple previously unknown zero-day vulnerabilities and incompletely patched security fixes, demonstrating the potential for autonomous vulnerability analysis to transition toward genuine discovery capabilities.
    5. GCSA's goal is moving beyond benchmark testing toward building AI Security Agents, driving their participation in the complete security lifecycle of vulnerability discovery, analysis, validation, and remediation, using a collaborative model to enhance attack path analysis and execution-level verification efficiency across large codebases.

Hong Kong, August 29, 2026 — The Global Cybersecurity Alliance (GCSA) announced today that the GCSA Agent has achieved a 91.3% success rate on the CyberGym benchmark, entering CyberGym's "Leading Systems Above 90%" tier.

CyberGym (https://www.cybergym.io/cybergym) is a large-scale, real-world cybersecurity evaluation framework developed by a research team at the University of California, Berkeley. It comprises 1,507 historical real-world vulnerability test cases covering 188 large-scale software projects, designed to assess the practical capabilities of AI agents in real-world vulnerability analysis scenarios.

Unlike traditional AI benchmarks that primarily evaluate code comprehension, knowledge Q&A, or static analysis capabilities, CyberGym requires AI agents to directly confront real vulnerable code environments.

In its core Level 1 test, an AI agent is provided only with a vulnerability description and an unfixed codebase. The agent must autonomously complete code analysis, vulnerability identification, attack path reasoning, PoC construction, and execution validation. A task is only deemed successful if the generated PoC successfully triggers the target vulnerability in the vulnerable version while failing to reproduce it in the patched version.

Therefore, what CyberGym measures is not merely whether an AI "understands code," but whether the AI can truly complete the entire process from security analysis to vulnerability reproduction and validation.

From Large Models to Security Agents

In this CyberGym test, the GCSA Agent ran on Grok 4.5 and Grok 4.6 models, ultimately achieving a 91.3% success rate.

This result also reflects a significant shift taking place in the AI cybersecurity field:

The underlying large model alone is no longer the sole determinant of ultimate security capability.

Real-world vulnerability research typically requires the continuous completion of multiple stages, including vulnerability description comprehension, large-scale code retrieval, attack surface identification, vulnerability hypothesis formulation, test input generation, program execution, feedback analysis, and iterative PoC refinement.

The GCSA Agent has built an agentic security workflow around this entire process.

Its goal is not simply to leverage large language models for code analysis, but to enable AI to enter real execution environments, autonomously form hypotheses around security issues, gather runtime evidence, execute tests, and ultimately validate security findings with reproducible results.

This CyberGym test provides a quantifiable external benchmark for this capability.

Real-World Vulnerability Research Capabilities

The core value of CyberGym lies in narrowing the gap between traditional AI testing and real-world cybersecurity research.

Its evaluation environment restores the pre-patch state of software project code. AI agents may need to autonomously pinpoint issues within large codebases containing thousands of files and millions of lines of code, ultimately generating PoCs that can genuinely trigger the target vulnerabilities.

More notably, further research on CyberGym has shown that this agentic security capability is not limited to reproducing known vulnerabilities.

In open-ended vulnerability research experiments, AI agents have already discovered multiple previously unknown zero-day vulnerabilities and security patches that were not fully remediated in history, demonstrating the potential for autonomous vulnerability analysis technology to migrate toward real-world vulnerability discovery capabilities.

For GCSA, this represents an even more important development direction.

Benchmark scores are not the end goal. GCSA's objective is to further build AI Security Agents capable of serving real-world cybersecurity scenarios, progressively participating in the complete security lifecycle of vulnerability discovery, analysis, validation, and subsequent remediation.

Building AI-Native Cybersecurity Capabilities

As artificial intelligence accelerates software development, AI is also transforming the methods of vulnerability research and cyber offense and defense.

Facing software systems of ever-increasing scale and complexity, the next generation of cybersecurity frameworks will increasingly rely on collaboration between human security experts and autonomous AI agents.

AI Security Agents are expected to help security teams:

* Discover software vulnerabilities with real exploit value earlier;

* Automatically analyze complex attack paths within large codebases;

* Automatically generate PoCs for execution-level vulnerability validation;

* Reduce false positives in traditional security detection through real runtime results;

* Accelerate vulnerability assessment, validation, and remediation efficiency;

* Expand the scale of software and systems that professional security teams can cover.

The GCSA Agent's 91.3% score on CyberGym represents a significant milestone in GCSA's efforts to build AI-native cybersecurity capabilities.

Looking ahead, GCSA will continue to advance autonomous vulnerability analysis, AI Security Agents, and intelligent cybersecurity technology research, further translating cutting-edge AI capabilities into real-world security strength, providing technical support for building a more secure, trusted, and resilient digital environment.

Source: GCSA Global Cybersecurity Alliance

Official Website: www.gcsa.org

Safety
technology
AI
Welcome to Join Odaily Official Community