Comparison

Open-source AI pentest tools, compared

A factual, feature-by-feature look at the open source AI pentest tools people evaluate together — and where Darkmoon is a genuine alternative to Strix, XBOW and PentestGPT. Verifiable facts only, with honest credit where competitors lead.

This is Darkmoon's own comparison, compiled from public sources on 2026-09-16. It is not a third-party endorsement and not an analyst citation — no external study lists Darkmoon. Star counts change over time; treat them as a signal, not a live figure. We do not claim any competitor cheats — we simply state each tool's public conditions.

Feature matrix

Tool by tool

ToolPublic tractionLicense & hostingModel locationProof & fix loopCoverageThird-party signal
Darkmoonus940★ (ASCIT31/Dark-Moon)GPL-3.0 · self-hostedLocal (Ollama / llama.cpp) + Privacy GatewayProof-of-exploit per finding + remediation PR loop (42/57 demonstrated)Web, API, AD/identity, K8s, cloud, CI/CD, DB, IoT/firmware, LLM1 verified press hit (Help Net Security); open reproducible benchmark
Strix≈62,900★ (usestrix/strix)Apache-2.0 · cloud modelCloud LLMPoC exploit + one-click autofix → fix PR + retestWeb-application focusedGitHub Trending sprint; security-press coverage; strong analyst mindshare
XBOWCommercial (closed cloud)Proprietary · SaaSCloud (hosted)Autonomous exploitation; 1,060 vulns / 22 CVEs reportedWeb / bug-bounty targets#1 US HackerOne leaderboard (Jun 2025); mainstream tech press
PentAGI≈24,600★ (vxcontrol/pentagi)Open source · self-hosted DockerCloud LLMFully autonomous; no published remediation benchmarkGeneral autonomous pentestLarge GitHub following; self-host community
PentestGPT≈15,500★ (GreyDGL/PentestGPT)Open sourceCloud LLMInteractive assistant (guides a human; not an autonomous executor)General guidanceUSENIX Security academic paper; 2023 first-mover
HexStrike AI≈11,900★ (0x4m4/hexstrike-ai)Open sourceBring-your-own LLM via MCPMCP server exposing 150+ tools to any LLMTool-runner breadth (any tool you wire in)Security press after real-world abuse on Citrix CVEs; MCP ecosystem
VulnHuntr≈2,800★ (protectai/vulnhuntr)Open sourceCloud LLMStatic analysis that finds real 0-days (no exploit execution)Python source code (SAST-style)Dark Reading, The Register, The Hacker News; real disclosed CVEs
ShannonCommercial / research (Keygraph)ProprietaryCloud, white-box (reads your source)Reports 96.15% (100/104) on the XBOW white-box validation setBenchmark-set targetsHeadline benchmark score
AISLECommercialProprietaryNot independently confirmedPositioned as autonomous find-and-fix (public specifics unverified)Not independently confirmedListed for completeness; we could not verify specifics
Being fair

Where competitors lead today

  • Strix~67× the stars, GitHub-Trending reach, press mindshare, and it owns the exact proof + fix-PR + retest message.
  • XBOWA third-party-verified result (top of a public HackerOne leaderboard) that no self-published number can match.
  • PentAGIFar more stars and self-host visibility.
  • PentestGPTAcademic legitimacy (a USENIX paper) and first-mover recognition.
  • HexStrikeMCP-native distribution and organic news pickup.
  • VulnHuntrEditorial reach far beyond its star count, on the back of real disclosed CVEs.
  • ShannonA higher headline benchmark number (white-box).
The wedge

What Darkmoon does that the leaders do not

Local-LLM + Privacy Gateway

Darkmoon can run entirely on a local model, and its Privacy Gateway means the model never sees your real IPs, hosts or credentials. Most high-star competitors call a cloud model.

Breadth in one open tool

Web, API, Active Directory & identity, Kubernetes, cloud, CI/CD, databases, IoT/firmware and LLM endpoints — several competitors specialise in one surface.

Honest, reproducible proof

The 42/57 remediation dossier and the black-box offensive benchmark publish their limits and invite you to reproduce or contest them.

The wedge in depth: self-hosted & sovereign environments.

Reproduce the numbers yourself

Darkmoon is open source. Clone it, point it at a lab you own, and read every line — the benchmarks invite you to contest them.