The defining feature of the autonomous and agentic penetration testing market is not a technology. It is price opacity. Almost nobody publishes a number. Pricing lives behind sales calls, per-asset spreads, and scoping conversations, which makes it nearly impossible for a buyer to know what fair looks like. This report opens the black box: what the market actually charges, how it charges, and what buyers actually pay, with every figure labeled as a public list price or a third-party estimate and dated to 2026.
Opacity is not an accident of a young market. It is a strategy, and it works because three mechanics reinforce each other. Understanding them is the first step to reading a quote for what it is.
Most vendors in the category publish no price at all. The path to a figure runs through a demo, a discovery call, and a quote tailored to what the seller learns about your budget on the way. When no one lists, no one can be anchored against, and every buyer negotiates alone.
Where pricing is per IP or per asset, the unit rate compresses steeply with volume. Public UK G-Cloud figures for one leading vendor run from roughly £40 per IP at the low end to about £2.40 per IP at very high volume, a spread near sixteen times. The headline unit price tells you almost nothing without the volume attached.
Even a known unit rate is multiplied by a scope the vendor helps define: which assets count, which modules are included, how many engagements or credits, what a “seat” is. The final invoice is a function of variables set in a conversation, not on a page, so two buyers of the same product routinely pay very different totals.
The table below compiles the pricing signals available across the category as of 2026. Public list prices are figures a vendor publishes openly. Third-party estimates are reported or inferred figures, not vendor list prices, and should be read as directional. Nothing here is a quote, and all figures are point-in-time.
| Vendor / Product | Price signal | Basis | Source |
|---|---|---|---|
| Horizon3.ai NodeZero | ~£40 → ~£2.40 / IP·yr | Per active IP per year, sales-led. UK G-Cloud list runs ~£40/IP (≤2,500 IPs) down to ~£2.40/IP at very high volume, an ~16× spread. | Public list UK G-Cloud |
| Horizon3.ai NodeZero | ~$18.6k / yr | Typical annual deal size. Company reported ~$250M funding and a >$2B valuation. | Third-party estimate |
| Pentera | ~$35k → $100k+ | Entry to upper range, priced by asset count and modules. No public list price. | Third-party estimate |
| XBOW | ~$4k | Reference figure. Company reported ~$120M funding; moved from pure self-serve autonomy toward human-in-the-loop in mid-2026. | Third-party estimate |
| Cobalt (PTaaS) | ~$30k | Median human engagement. Delivered on an 8-hour credit model. | Third-party estimate |
| Cobalt Autonomous | ~$3,500 | AI-delivered testing tier, per engagement. | Third-party estimate |
| Synack (PTaaS) | ~$105k | Median human program. | Third-party estimate |
| Synack “Sara” | ~$4,181 | AI-delivered testing, per engagement. | Third-party estimate |
| HackerOne | ~$40k | Representative engagement. | Third-party estimate |
| Bugcrowd | ~$40.5k | Representative engagement. | Third-party estimate |
| Astra | $1,999 & $5,999 / yr | Self-serve tiers, each including a human-verified pentest and compliance report. The closest low-end analog to transparent, productized testing. | Public list |
| Aikido | ~$300–600 / mo + credits | Platform subscription plus pentest credits ($500 up to $4,000), on a run-now-pay-later model. | Public list |
| Beagle | $99–359 / mo | Self-serve subscription tiers. | Public list |
| TurboPentest | $99 / $299 / $699 | One-time pricing tiers. | Public list |
| Pentest-Tools | ~$95 / $140 / $190 / mo | Self-serve subscription tiers. | Public list |
| Vonahi vPenTest | ~$2,999 | MSP / IP-block oriented pricing. | Third-party estimate |
| Hacktron | ~$350 | Dynamic, usage-based pricing. | Third-party estimate |
Figures compiled as of 2026 and are point-in-time. Public list prices are published by the vendor; third-party estimates are reported or inferred and are directional only, not quotes. Currency shown as published. Funding and valuation figures are company-reported or press-reported.
Two data points cut through the opacity because they come from vendors that sell both human and AI testing. Cobalt Autonomous lands near $3,500 and Synack’s “Sara” near $4,181, both third-party estimates. These are the same firms whose human programs are estimated near $30k and $105k respectively. In other words, the market’s own statement of what AI-delivered, validated testing is worth per engagement is roughly one-third to one-tenth of the price of the human program sitting next to it on the same price sheet. That ratio, set by the incumbents themselves, is the most honest anchor in the category.
Reported ~$40M funding, positioned at the fully-autonomous end of the spectrum. No public list price.
Reported ~$30M funding, positioned at the human-validated end. No public list price.
Reported ~$250M funding and a >$2B valuation, the category’s largest by capital, still sales-led per active IP.
Behind the numbers sit a handful of pricing models. Each answers “what do we charge for” differently, and each has a predictable failure mode for the buyer. The model matters as much as the rate.
Used by: Horizon3.ai, Vonahi, asset-based tiers at Pentera.
Pro: scales with the thing being protected, and is easy to reason about for a fixed estate. Con: unit rates swing by an order of magnitude with volume, penalize a large or elastic surface, and turn every cloud autoscale event into a budgeting question.
Used by: Cobalt (8-hour credits), Aikido (pentest credits), TurboPentest (one-time).
Pro: maps cleanly to a project or a compliance deadline, and is easy to buy in discrete amounts. Con: a credit is a proxy for value that can run out mid-need, and continuous coverage becomes a recurring purchase decision rather than a default.
Used by: platform tools with a tester-seat component.
Pro: predictable for a fixed team and familiar from other SaaS. Con: seats are unrelated to how much testing actually happens or how much surface is covered, so the price tracks headcount rather than protection.
Used by: Astra, Beagle, Pentest-Tools.
Pro: transparent, self-serve, and easy to compare, which is why the low end of the market prices this way. Con: a flat tier can mask real differences in depth and coverage, so a low monthly number is not automatically a good value without knowing what a run includes.
Used by: Aikido (platform + credits), Hacktron (dynamic).
Pro: aligns cost to actual use and lets a buyer start small and grow. Con: consumption is hard to forecast, and a bill that moves with usage is exactly the kind of variable that reintroduces the opacity a published price was supposed to remove.
Every model is a proxy for the same thing: how much testing, over how much surface, at what depth. The cleaner the value metric, the easier it is for a buyer to predict the bill and trust the vendor. Opaque metrics are where margin hides.
Cutting across vendors and models, spend clusters into three bands by buyer size. These are estimated ranges synthesized from the signals above, not vendor list prices, and they describe annual outlay for continuous or repeated testing.
Bands are estimated ranges as of 2026, synthesized from the pricing signals above. They are directional, not quotes, and individual outcomes vary widely with scope, asset count, and negotiation.
One fact reframes the whole table. Automated AI testing has very low marginal cost. Once the agent exists, running it against one more asset or one more time is inexpensive, which means margins on automated testing are very high across the category. The price a vendor charges is therefore a statement about value and willingness to pay, not a pass-through of what delivery costs.
That is not a criticism, it is how software pricing works. But it has a clear implication for buyers: since cost is not what sets the number, opacity has no cost-based justification. A vendor that can charge anything the market will bear and still make a high margin has every incentive to keep the number behind a call. The antidote is not a lower price. It is a published one.
The value metric is straightforward: verified domain by depth, plus Operator-hours, so what you pay tracks how much testing happens over how much surface, not a per-asset spread negotiated on a call. A free, self-serve Recon tier lets you see your attack surface before you spend anything, and every paid tier ships exploit-proven findings, proof rather than probability.
It is a different posture from the black box this report describes: the metric is legible, the on-ramp is self-serve, and you can start on the free Recon tier without talking to sales.
See your attack surface with the free, self-serve Recon tier, or talk to us about continuous, exploit-proven coverage that scales by verified domain and depth.