The Builder's Verdict

Methodology

How we test, and how we label what we haven't

Most review sites imply everything was tested. We do the opposite: every verdict carries a visible evidence label saying exactly what we did, and anything we haven't verified says so. Here's the whole system.

The evidence ladder

Every review, comparison and guide states one of five evidence levels. The label sits next to the verdict, not buried in a footer.

Evidence levels, what each means and does not mean
LabelWhat it meansWhat it doesn't
Hands-on testedUsed on real work, on real jobs, over a sustained period.A guarantee it fits your specific trade or workflow.
Trial account usedWe set up a real trial account and worked through core workflows.Long-term reliability or support quality under load.
Desk assessedWe checked current UK pricing against official vendor sources, read the official documentation and terms, and weighed public user evidence. No hands-on use yet.That we've run it on a live job. Day-to-day feel, speed and support quality are reported from user evidence, not our own use.
ProvisionalAn early view based on partial evidence. Expect this to change.A settled recommendation.
Insufficient evidenceNot enough verifiable information to reach any verdict.That the product is bad; we simply can't judge it yet.

What a desk assessment involves

Launch verdicts are desk assessments, and we'd rather tell you that plainly than imply testing that hasn't happened. A desk assessment means we: verify current UK pricing directly against the vendor's own pricing page (and record the date); read the official documentation, plan matrices and terms; weigh public user evidence from named, linkable sources: trade forums, app-store reviews, community threads; and check whether the vendor has any commercial relationship with us, which we disclose either way.

Anything a desk assessment can't know, such as how the app feels on site, how support behaves when you're stuck, is marked unverified in the criteria table rather than guessed.

Why there are no scores yet

A number like 8.4/10 implies measurement. Publishing one off the back of reading pricing pages would be dressing research up as testing, the exact habit this site exists to call out. So until a product has at least trial-level evidence, its review carries a verdict label and criterion ratings tied to checkable facts, and no overall number.

When hands-on testing begins, scores will use the weighted framework below, the same core criteria within a category, and never finer than half-point precision. 8.5 means something; 8.73 is theatre.

Scoring criteria and weights
CriterionWeight
Value for moneyHigh
Setup & onboardingHigh
Mobile usability on siteHigh
Quoting & estimatingHigh
Invoicing & paymentsHigh
SchedulingMedium
UK tax & compliance fit (VAT, CIS, MTD)High
Integrations (Xero, Sage, QuickBooks…)Medium
Support qualityMedium
Pricing transparencyMedium
Cancellation & data exportMedium
Sole-trader fitContextual
Growing-team fitContextual

The verdict labels

A controlled set, with no invented superlatives and no “editor's choice” badges with mystery criteria.

  • Recommended
  • Best for
  • Good fit
  • Consider with caution
  • Poor value
  • Not recommended
  • Provisional
  • Insufficient evidence

Sources, dates and corrections

Every substantive claim carries a source with the date we checked it, listed in the “Evidence & sources” section of the page it appears on. Prices are re-verified when we update a page, and the date always tells you how fresh the fact is. If we got something wrong, the corrections page explains how to tell us and what we do about it.

Money never touches any of this: the editorial policy sets out what commercial partners can and can't buy, and the affiliate disclosure explains how links pay us.