DeepSeek V4.1 Flash Scores Perfect on AI Hacking Benchmark for $4.65
Security firm reports that DeepSeek V4.1 Flash achieved code execution on all 11 vulnerable targets in its AI hacking benchmark, with all four fixed targets remaining secure, at a total cost of $4.65. A detailed review confirmed the results, finding six solutions followed the planned attack path and five successful routes not distinguished by the original scoring system. The review highlighted the model's strong hacking ability and identified areas where the benchmark needed stricter checks.
Coverage timeline
Hacker Newstalhof8
DeepSeek V4.1 Flash produced an extraordinary result in our AI hacking benchmark. It gained code execution on all 11 vulnerable targets, while all four fixed targets remained secure. The accepted runs cost only $4.65. A perfect score at that price deserves a detailed review. We looked into every command, request, and successful attack. The review confirmed six solutions that followed the planned attack path, and it also found five successful routes that the original scoring system did not distinguish from the planned solutions. The result gave us two useful insights. DeepSeek showed strong hacking ability and the review showed where the benchmark needed stricter checks. ## **A large attack run for less than five dollars** DeepSeek worked inside isolated copies of Grafana, Jenkins, and Nextcloud. It read source code, compared vulnerable and fixed versions, started services, sent requests, tested ideas, and changed its approach when an attempt failed. Across the full benchmark, the model