-
Evaluating Ox Alpha on Hack The Box Challenges
A new model called Ox Alpha landed on OpenRouter last night, and you can use it for free through Monday, August 24. Today, I tested it on my HTB-Challenger BenchmarkThe HTB-Challenger Benchmark evaluates LLMs’ ability to find and exploit security vulnerabilities. It tests models against selected Hack The Box challenges of varying difficulty and measures their performance. For more information, visit the HTB-Challenger Benchmark page. , and hear me out: stop what you’re doing, cancel your weekend plans, and test it! If you need to cancel all your meetings, your child’s birthday party, or your romantic weekend for two, do it. It’s worth it! Just don’t tell my wife I said that 😉
But really, do it.
-
Evaluating GLM 5.3 on Hack The Box Challenges
The GLM 5.3 model was released just last week. The interesting fact is that it’s just a retrained version of GLM 5.2, which doesn’t give you much hope for good results. But Z.ai’s marketing gave us sentences like: “As we scaled post-training, cyber capability developed faster than we expected.” and graphs showing better results in CyberGym than both Mythos 5 and GPT-5.6 Sol.
Also, the last model I tested from Z.ai was GLM 5.1, and it was excellent. So I thought: “OK, maybe, just maybe, this time the marketing isn’t overhyped. Maybe I really can get a new open-weight champion…” Well, spoiler alert: I didn’t.
-
What Does the HTB‑Challenger Benchmark Actually Measure?
I’m sure many of you have come to this blog, checked the HTB-Challenger BenchmarkThe HTB-Challenger Benchmark evaluates LLMs’ ability to find and exploit security vulnerabilities. It tests models against selected Hack The Box challenges of varying difficulty and measures their performance. For more information, visit the HTB-Challenger Benchmark page. results for your favorite model and wondered why they differ so much from official benchmarks or your own experience. “DeepSeek V4 Flash is the best model I have ever used. How could this moron put it at the bottom of his benchmark?!?” I hear you shouting.
And fair enough. If you use the model through Cursor, Claude Code, OpenCode, or another modern application, I agree that my results may have little to do with your experience. But if you’re wondering how the model would perform in your own pentesting or security testing harness, I think you should look at them carefully.
And because I realized that I had done a really poor job of explaining what my benchmark actually measures, I put together this post to clarify it.
-
Evaluating DeepSeek V4 Pro 0813 on Hack The Box Challenges
The release of DeepSeek V4 Pro back in April got a lot of attention - so DeepSeek decided to release it again! The new version is called DeepSeek V4 Pro 0813, and it is not exactly clear what changed. Reading its release page, you can almost feel the marketing team’s desperation to put at least something there. The few concrete details are that it uses the same core V4 Pro architecture with approximately 1.6T total and 49B active parameters, adds a DSpark speculative-decoding module (whatever that is), and supports a 1M-token context.
DeepSeek reports large improvements across several agent benchmarks, although those are the developer’s own results. I had finished testing the previous Pro version, now called DeepSeek V4 Pro 0423, shortly before the new version appeared (sic!). I therefore ran 0813 through my HTB-Challenger BenchmarkThe HTB-Challenger Benchmark evaluates LLMs’ ability to find and exploit security vulnerabilities. It tests models against selected Hack The Box challenges of varying difficulty and measures their performance. For more information, visit the HTB-Challenger Benchmark page. to see whether it performed better on cybersecurity tasks.
-
Evaluating Hy3 on Hack The Box Challenges
Tencent’s Hy3 belongs to the category of “interesting LLM models you may never have heard of.” It appeared to be quite popular on OpenRouter.ai at the start of the summer, when a free version was available, and it seemed to have earned a good reputation. So I thought it might be worth including in the HTB-Challenger BenchmarkThe HTB-Challenger Benchmark evaluates LLMs’ ability to find and exploit security vulnerabilities. It tests models against selected Hack The Box challenges of varying difficulty and measures their performance. For more information, visit the HTB-Challenger Benchmark page. .
-
Evaluating GPT‑5.6 on Hack The Box Challenges
Until now, I had tested only the GPT-5.6 Luna model because the two more advanced GPT-5.6 models, Terra and Sol, rejected my test prompts, flagging them as a “possible cybersecurity risk.” Fortunately, this issue turned out to have an easy solution, so I could finally spend a couple of days testing the whole GPT-5.6 family for my HTB-Challenger BenchmarkThe HTB-Challenger Benchmark evaluates LLMs’ ability to find and exploit security vulnerabilities. It tests models against selected Hack The Box challenges of varying difficulty and measures their performance. For more information, visit the HTB-Challenger Benchmark page. .
-
Evaluating Muse Spark 1.2 on Hack The Box Challenges
I have to admit something: I have opinions. That’s not exactly a character flaw, but mixing opinions with facts and hard data is rarely a good idea. Sometimes, though, keeping the two separate is difficult.
Like now.
Meta is back, this time with Muse Spark 1.2, an ambitious new model. Unfortunately, I have strong opinions about Meta and its business model, and I would rather not risk letting those opinions color my usual discussion of the test results. So this time, I’ve decided not to offer any commentary. Here are the hard numbers from my testing. I’ll leave the interpretation to you.
-
Evaluating Grok 4.6 on Hack The Box Challenges
When I tested an older version of Grok this spring, my only comment was blunt: “grok-4.20 was useless.” So when Grok 4.5 was released a month ago and started receiving very positive reviews, I was quite surprised.
Today, Grok 4.6 was released and I’ve finally had a chance to test it against HTB challenges, and the team behind it deserves an apology. I’m not sure what the folks at xAI did, but the performance increase over the older version is jaw-dropping.
-
Evaluating DeepSeek V4 Pro on Hack The Box Challenges
DeepSeek V4 Pro was the best (and most expensive) LLM I tested with Strix a few months ago. After the heartbreaking results of its smaller sibling, DeepSeek V4 Flash 0731, in my HTB-Challenger tests, I was curious to see how the Pro version would handle the new challenges.
-
Evaluating DeepSeek V4 Flash 0731 on Hack The Box Challenges
I tested the previous version of DeepSeek V4 Flash, now called DeepSeek V4 Flash 0423, with Strix back in April, and I absolutely fell in love with it. It delivered great results at a very low price and became my go-to model for most tasks that didn’t require the capabilities of frontier models.
That’s why I was looking forward to testing the latest, improved version on my HTB-Challenger BenchmarkThe HTB-Challenger Benchmark evaluates LLMs’ ability to find and exploit security vulnerabilities. It tests models against selected Hack The Box challenges of varying difficulty and measures their performance. For more information, visit the HTB-Challenger Benchmark page. . But the results… Oh dear, oh dear, oh dear. Where do I even start?
-
Evaluating Qwen3.8 Max on Hack The Box Challenges
Qwen3.8 Max was released recently, but it didn’t create much buzz. I was wondering if it was simply overshadowed by Kimi K3, which was released just a few days earlier, or if its performance wasn’t strong enough to raise any eyebrows.
-
Evaluating Kimi K3 on Hack The Box Challenges
Kimi K3 landed with a big splash just three weeks ago, and the initial reviewers seemed to agree on one thing: it’s very good, but it’s also quite expensive. So here I am, the late reviewer, with my own numbers. Let’s see if they line up with the prevailing opinion.
-
Evaluating GPT‑5.6 Luna Pro on Hack The Box Challenges
What the heck is GPT-5.6 Luna Pro, I hear you asking. That’s a very good question! To answer it, let me quote its description on OpenRouter.ai: “GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with
reasoning.modeset toprofor higher-quality responses on complex tasks.”Okay, so how much better, and how much more expensive, is it compared with GPT-5.6 Luna, I hear you asking now. And that’s exactly what I can tell you.
-
Evaluating GPT‑5.6 Luna on Hack The Box Challenges
My first thought was: let’s kick off the HTB-Challenger BenchmarkThe HTB-Challenger Benchmark evaluates LLMs’ ability to find and exploit security vulnerabilities. It tests models against selected Hack The Box challenges of varying difficulty and measures their performance. For more information, visit the HTB-Challenger Benchmark page. with some state-of-the-art models. Let’s see what these Fables, Opuses, Terras and Sols can do with offensive security challenges. As it turned out, not much - sooner or later, all of them refused to continue with the HTB challenge-solving workflow, with messages like “This content was flagged for possible cybersecurity risk.” So I was left with just GPT-5.6 Luna. And it was not so bad after all.
Update (17 August 2026): I have since solved the refusal issue with GPT-5.6 Terra and Sol by joining OpenAI’s Trusted Access for Cyber program. You can read about the solution and see my test results for the full GPT-5.6 family.
-
Introducing HTB‑Challenger: Benchmarking LLM models on Hack The Box Challenges
Choosing an LLM model for offensive security work is difficult. Public benchmarks can help compare model capabilities, but they rarely show how much it costs to complete a task from start to finish. The price per million tokens does not tell the whole story: different models may require vastly different numbers of tokens and tool calls before they solve a problem - or fail to solve it.
On top of that, as Winston Churchill allegedly said, “The only statistics you can trust are those you falsified yourself.”
I previously spent a lot of time evaluating new LLMs on offensive security tasks with Strix. I am still enthusiastic about Strix, and it remains my go-to tool for penetration testing, but it is no longer the best fit for the repeated model comparisons I want to run.
Strix performs autonomous security testing against complex applications, so a single run can require many steps, tool calls, and tokens. That makes repeated testing expensive. Strix is also evolving rapidly and improving with each version. Results produced with different versions are therefore not directly comparable: changes in the agent could be mistaken for changes in model performance. A fair comparison would require pinning one Strix version and rerunning every model against it.
I therefore revived an older idea and turned it into my new testing approach: the HTB-Challenger BenchmarkThe HTB-Challenger Benchmark evaluates LLMs’ ability to find and exploit security vulnerabilities. It tests models against selected Hack The Box challenges of varying difficulty and measures their performance. For more information, visit the HTB-Challenger Benchmark page. .
-
How much better is Strix 1.0? Results from a small rerun
It has been a while since my last blog post, but I definitely haven’t stopped playing with AI-assisted security tools. Summer is here, and my schedule is finally a little less crowded, so I’m planning to write a few blog posts about what I’ve learned lately.
This post should be short and to the point. After a few months in its v0.x phase, Strix recently reached adulthood with the release of version 1.0. Since I ran quite a few tests with Strix v0.8.3 in the past, I was curious whether the improvements in Strix v1.0.4 justified the major-version bump.
-
Deepseek V4 with Strix: a quick test
Deepseek released V4 yesterday in two variants. V4 Pro has 1.6T total parameters with 49B active, while V4 Flash is the smaller, faster, cheaper sibling with 284B total and 13B active. In the release notes, the company claims that V4 Pro rivals the world’s best closed-source models in reasoning and leads all open models on agentic coding benchmarks. It also says that V4 Flash comes surprisingly close despite its smaller size and can match Pro on simpler agentic tasks.
Those are exactly the kinds of claims worth testing, so I wanted to see how both models hold up in my local lab setup and how they compare with the other models I have tested recently.
-
Kimi K2.6 with Strix: a quick test
The Kimi K2.6 was released just yesterday, and looking at the benchmarks quoted in the release blog post, one could easily get the impression that it is the best model ever released. So I decided to do a quick test.
-
Agentic AI pentesting with Strix: results from 18 LLM models
Over the last couple of months, I spent close to a hundred hours testing an autonomous AI pentesting tool called Strix with 18 different LLM models. My goal was to evaluate which LLM model performed best with the tool in this lab setup and what that might say about autonomous AI pentesting more generally.
After a few dead ends and a lot of discarded results (I summarised that earlier failed testing in my How not to test LLM models post), I finally arrived at a methodology that I think produces meaningful practical benchmark of real Strix usage under my specific provider, tier, pricing, and rate-limit constraints.
This post contains the results of my testing and a few observations.
-
How not to test LLM models
In the Czech Republic, we have a whole lore built around a fictitious character called Jára Cimrman. He was partially a genius (one of the greatest playwrights, composers, teachers, travellers, inventors, detectives, gynecologists and sportsmen, among many other things) but mostly a loser (“… while running away from one furious tribe, he missed the North Pole by just seven meters, thus almost becoming the first human to reach the North Pole.”) One of his strongest skills was finding dead ends. He found many ways in which things should NOT be done and helped humanity many times by being able to authoritatively say: “This isn’t the way to do it, my friends!”
After spending several days trying to compare the performance of different LLM models, I’m sure Jára would be very proud of me.
-
How GPT‑5.4 performed with Strix ‑ and why it fell short
GPT-5.4 was released just yesterday and because I’m currently testing the strix autonomous AI tool for web penetration testing, the temptation to compare it with other LLM models was too strong to resist. As I already spoiled in the title, the results were pretty bad. But there could be a good explanation for this.
-
LLM model statistics from my Strix testing
In my previous post I summarized a few impressions from my strix testing (TL;DR I was impressed).
Since then, I have collected some hard data and summarized it on this page. I still haven’t run enough tests to be able to objectively compare different models, but I believe that page is not a bad starting point when selecting an LLM model for your own testing.
Beyond the numbers, here are some short personal observations for each model.
-
Strix ‑ First impressions
We’ve all heard it: penetration testers are over. Their job will soon be done by agentic AI frameworks that can find the same (or even more elusive) vulnerabilities for a fraction of their bloody money - and since they don’t need to sleep, eat, or have a work-life balance, they can run 24/7.
And you, Red Teamers, are next.
Ok, doomers, you got my attention. I decided to look at one of these rising AI penetration testing superstars, strix, and be generous enough to share my random thoughts with you. If you plan to test this tool yourself, check the APPENDIX: Practical tips for Strix testing section at the end of this post - I think I can save you some time and money.
Here’s the TL;DR for those of you who don’t have enough time or patience to read my whole rant:
- After this test, am I scared to death and looking for a plumbing job? No, not yet.
- Am I impressed? Yes, I am. Actually, thinking about it, I’m very impressed.