Methodology
How we evaluate 190 tools across 17 categories
The value of a directory is entirely in its judgement. Here is what that judgement consists of, so you can decide how much of it to trust.
What gets listed
A tool earns an entry if it is actively maintained, has real users outside its own company, and does something a team could reasonably build a process around. That is deliberately a low bar for inclusion and a high bar for how it gets written up.
Abandoned projects are removed rather than left with a warning. Where a tool has been acquired or its licence has changed in a way that affects buyers — and that has happened repeatedly in this market — the entry says so.
Every entry states its limitations
A directory that only lists features is a price list with extra steps. The most useful sentence in any tool review is the one that tells you where it falls down, because that is the thing you will discover in month three otherwise.
The rule is simple: if we cannot name at least two concrete limitations, we do not understand the tool well enough to recommend it. "Has a learning curve" does not count. "The proxy architecture breaks on apps with strict CSP" does.
Pricing records the metric, not just the number
The headline price on a vendor page is rarely what determines your invoice. Device clouds bill per parallel session. Visual testing bills per screenshot. Analytics bills per event or per tracked user. Test management bills per Jira user, including the ones who never open a test case.
So every entry records the starting price and, separately, what actually drives the bill. That second field is the one worth reading before a procurement conversation.
Figures are indicative and were last reviewed August 2026. Vendors change pricing frequently and enterprise pricing is negotiated. Always confirm on the vendor's own page — every entry links directly to it.
Adoption is a signal, not a score
Each tool carries an adoption rating from one to five. It reflects how widely the tool is deployed and how easy it is to hire for — not how good it is. A category standard can still be the wrong choice for you, and a niche tool can be exactly right.
We do not publish an overall quality score. Any single number that ranks Playwright against Tricentis Tosca is measuring nothing, because they are bought by different people to solve different problems under different constraints.
How the stack finder ranks
The finder treats your answers as two different kinds of input. Hard requirements — a budget ceiling, self-hosting, open source, organisation size — are filters. A tool that violates one never appears, even if it would otherwise be the obvious answer.
Everything else is weighted. Language overlap with your stack, platform coverage, whether authoring matches who will actually maintain the tests, and adoption all contribute. A tool whose primary purpose is the category in question always outranks one that covers it as a secondary feature, and the results say so explicitly when that is the case.
Each recommendation shows the reasons it scored well and the caveats that apply to your specific answers. Nothing is stored, no account is required, and the result is encoded in the URL so you can share or bookmark it.
Independence and funding
No vendor pays for placement, ordering, or a more favourable write-up, and no vendor reviews an entry before it is published.
Where affiliate arrangements exist in future they will be disclosed on the entry itself, and they will not change the ordering or the wording of a limitation. If that ever stops being true, this page will say so.
Corrections
Facts change: prices move, licences change, companies get acquired, features ship. If something here is wrong or out of date, it should be fixed rather than defended.
The most common source of error in a catalogue this size is pricing drift, followed by acquisition-driven roadmap changes. Both are flagged in entries where they are known to have happened.
The short version