Rendered at 13:57:11 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
BrawnyBadger53 4 hours ago [-]
I can only assume this whole post is meant to be an ad for the cyber frost model? The charts being unreadable such that only cyber frost is identifiable, benchmarks being chosen to mostly support it, and the model being only 2 days old all makes me rather suspect.
girvo 5 hours ago [-]
I’ve been playing with Qwen 3.8 Flash Next uncensored (using the Heretic v2 method) for security exploration, and have been quite impressed, so I’m not surprised to see it near or at the top here. But also it’s a far more powerful base model, so that shouldn’t be too surprising either.
prettyblocks 28 minutes ago [-]
The abliterated qwen3.8 has been really good too
bede 4 hours ago [-]
Please use a categorical colour palette when visualising data like these
Incipient 1 hours ago [-]
Has anyone tried these security models for finding bugs from the outside vs say Fable reviewing code for bugs on the inside?
flipping_beacon 5 hours ago [-]
Would have been better with something else other than the gradient colour scheme
b112 5 hours ago [-]
Indeed... the graphs are basically just annoying to look at, as the colours are impossible to use as unique identifiers. Can't understand the decision on that.
I also can't read the actual numbers for the lighter shades of green, there's not enough contrast.
This may seem unimportant, but some of these models are tiny, others huge. If you have a specific task and the #1 model is a 180B MoE, you maybe can't run that. So you may want to look for smaller models, with high scores.
Lastly, the order of the models on the graph, isn't the order of the models discussed. Which makes the graph less inline with everything.
With this degree of disconnect on how to present even the most basic concepts, I question the author's underlying logic in how they test even.
I also can't read the actual numbers for the lighter shades of green, there's not enough contrast.
This may seem unimportant, but some of these models are tiny, others huge. If you have a specific task and the #1 model is a 180B MoE, you maybe can't run that. So you may want to look for smaller models, with high scores.
Lastly, the order of the models on the graph, isn't the order of the models discussed. Which makes the graph less inline with everything.
With this degree of disconnect on how to present even the most basic concepts, I question the author's underlying logic in how they test even.
Just wow
Also, the graduated colour scheme works only on the first plot, it's misleading on the others.