Methodology
How we test, and how we label what we haven't
Most review sites imply everything was tested. We do the opposite: every verdict carries a visible evidence label saying exactly what we did, and anything we haven't verified says so. Here's the whole system.
The evidence ladder
Every review, comparison and guide states one of five evidence levels. The label sits next to the verdict, not buried in a footer.
| Label | What it means | What it doesn't |
|---|---|---|
| Hands-on tested | Used on real work, on real jobs, over a sustained period. | A guarantee it fits your specific trade or workflow. |
| Trial account used | We set up a real trial account and worked through core workflows. | Long-term reliability or support quality under load. |
| Desk assessed | We checked current UK pricing against official vendor sources, read the official documentation and terms, and weighed public user evidence. No hands-on use yet. | That we've run it on a live job. Day-to-day feel, speed and support quality are reported from user evidence, not our own use. |
| Provisional | An early view based on partial evidence. Expect this to change. | A settled recommendation. |
| Insufficient evidence | Not enough verifiable information to reach any verdict. | That the product is bad; we simply can't judge it yet. |
What a desk assessment involves
Launch verdicts are desk assessments, and we'd rather tell you that plainly than imply testing that hasn't happened. A desk assessment means we: verify current UK pricing directly against the vendor's own pricing page (and record the date); read the official documentation, plan matrices and terms; weigh public user evidence from named, linkable sources: trade forums, app-store reviews, community threads; and check whether the vendor has any commercial relationship with us, which we disclose either way.
Anything a desk assessment can't know, such as how the app feels on site, how support behaves when you're stuck, is marked unverified in the criteria table rather than guessed.
Why there are no scores yet
A number like 8.4/10 implies measurement. Publishing one off the back of reading pricing pages would be dressing research up as testing, the exact habit this site exists to call out. So until a product has at least trial-level evidence, its review carries a verdict label and criterion ratings tied to checkable facts, and no overall number.
When hands-on testing begins, scores will use the weighted framework below, the same core criteria within a category, and never finer than half-point precision. 8.5 means something; 8.73 is theatre.
| Criterion | Weight |
|---|---|
| Value for money | High |
| Setup & onboarding | High |
| Mobile usability on site | High |
| Quoting & estimating | High |
| Invoicing & payments | High |
| Scheduling | Medium |
| UK tax & compliance fit (VAT, CIS, MTD) | High |
| Integrations (Xero, Sage, QuickBooks…) | Medium |
| Support quality | Medium |
| Pricing transparency | Medium |
| Cancellation & data export | Medium |
| Sole-trader fit | Contextual |
| Growing-team fit | Contextual |
The verdict labels
A controlled set, with no invented superlatives and no “editor's choice” badges with mystery criteria.
- Recommended
- Best for
- Good fit
- Consider with caution
- Poor value
- Not recommended
- Provisional
- Insufficient evidence
Sources, dates and corrections
Every substantive claim carries a source with the date we checked it, listed in the “Evidence & sources” section of the page it appears on. Prices are re-verified when we update a page, and the date always tells you how fresh the fact is. If we got something wrong, the corrections page explains how to tell us and what we do about it.
Money never touches any of this: the editorial policy sets out what commercial partners can and can't buy, and the affiliate disclosure explains how links pay us.