SentinelOne has introduced what it describes as the first long-horizon reverse-engineering benchmark for frontier AI models. This benchmark uses the Fast16 malware, a 2005 Windows malware designed to interfere with LS-DYNA engineering software, as its test case. Fast16 is believed to have been used in Iran's nuclear weapons development program, similar to Stuxnet.
SentinelLabs researchers tested several leading AI models, including OpenAI's GPT-5.5 and GPT-5.6 Sol, Z.ai’s GLM-5.2, and Anthropic’s Opus 4.x. The benchmark evaluates a model's ability to sustain a trustworthy investigation across eight escalating stages, particularly when new evidence contradicts earlier conclusions. This approach differs from scoring models on isolated tasks.
OpenAI's GPT-5.6 Sol was the only tested model that successfully completed all eight stages of the benchmark, achieving this in three separate runs with different reasoning-effort settings. Other models, such as GPT-5.5, GLM-5.2, and Opus 4.7 and 4.8, performed solid local analysis but failed to progress through all stages. GPT-5.5 stalled at the initial stage, while Opus models prematurely declared work finished.
SentinelLabs attributes GPT-5.6 Sol's success to its "project-scale recovery" capability. This refers to a model's ability to retract disproven conclusions, trace downstream dependencies, correct the root cause, and integrate that correction throughout the entire investigation, rather than just patching immediate errors. This capability is crucial for complex, multi-stage analytical tasks.
Despite GPT-5.6 Sol's strong performance, SentinelLabs researchers concluded that human oversight remains essential. Even the best AI runs made semantic errors, accepted weak quality controls, and claimed readiness prematurely. The researchers suggest that the best current use for these AI models is as a supervised investigative agency, with human analysts defining objectives, identifying blind spots, and retaining final publication authority.
✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →
One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.
One email a day. Unsubscribe in one click, any time.
Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.
▶ Play today's briefNew every morning, and the back catalogue is archived by date.
SentinelOne developed the first long-horizon reverse-engineering benchmark for frontier AI models, using the Fast16 malware as a test case. OpenAI's GPT-5.6 Sol was the only model to complete all eight stages of the benchmark, demonstrating superior "project-scale recovery" compared to other models. This benchmark highlights the current capabilities and limitations of AI in complex, multi-stage investigations, emphasizing the continued need for human oversight.