Software evaluation criteria are the weighted standards you score every candidate tool against, written before you look at vendors. A workable set for most small teams: functional fit (30%), usability and adoption (20%), integrations (15%), data portability (10%), security and compliance (10%), support quality (5%), and total cost of ownership (10%). Adjust the weights to your situation, then score each tool on identical evidence.
That definition is easy to nod at and hard to do. The rest of this guide turns each criterion into something concrete — what it actually means, what to test during a trial, and how to tell a passing score from a failing one — so you end up with a decision you can explain to whoever signs the invoice.
What makes a criterion good rather than decorative?
A criterion earns its place on the scorecard when it meets three tests.
It is specific to your work, not the category. "Good automation" is decorative. "Can trigger a follow-up sequence when a deal stalls for seven days, without me writing code" is a criterion — you can pass or fail a tool on it in ten minutes.
It is testable with evidence you can gather. If the only way to score it is to trust a vendor's marketing page, it isn't a criterion; it's a hope. Prefer things you can observe in a trial, verify in documentation, or confirm with a support ticket.
It is weighted honestly. Every criterion you list will be met by some tool and missed by another. If you weight everything equally, the cheap-to-satisfy criteria drown out the ones that will actually determine whether the tool survives its first quarter. Assign weights before you see the scores — otherwise you'll unconsciously reweight to justify the tool you already like.
This is the same discipline behind our broader criteria-first framework for choosing business software; this article is the criteria layer of it, written out in examples.
Which software evaluation criteria examples should be on your scorecard?
Here is a reusable starting scorecard. The weights are a template for a small team buying a single operational tool — treat them as a first draft to argue with, not a standard.
| Criterion | Suggested weight | What it really measures | How to test it in a trial |
|---|---|---|---|
| Functional fit | 30% | Whether the tool does the three to five jobs you bought it for, well — not whether it has the longest feature list | Write your top five real tasks; complete each one in the trial account with your own data |
| Usability and adoption | 20% | Whether the people who must use it daily will actually use it | Hand the trial to the least technical teammate with no training and watch where they stall |
| Integrations | 15% | Whether it connects to the tools already in your stack, natively or through a documented API | Connect your two most important systems during the trial, not after |
| Data portability | 10% | Whether you can get your data out in a usable format if you leave | Run a full export on day one of the trial and open the file |
| Security and compliance | 10% | Whether it meets your obligations for the data you'll put in it | Request current certifications, read the sub-processor list and data-residency terms |
| Support quality | 5% | Whether help arrives when something breaks at an inconvenient hour | Open a real support ticket during the trial and time the response |
| Total cost of ownership | 10% | The full cost shape over two to three years, including growth, seats, add-ons and migration | Model your expected usage at 2x today's volume against the published pricing model |
Score each criterion 1–5 on a shortlist of three to five tools, multiply by the weight, and total. The number won't make the decision for you, but a tool that wins on total while losing badly on a 30%-weight criterion is telling you something important.
How do you weight criteria for your own situation?
Start from the template and move weight toward whatever failure would hurt most.
- Replacing a tool people already resent? Push usability and adoption up. A technically superior tool nobody opens scores zero in practice.
- Buying into a stack with heavy hand-offs? Push integrations up, and be specific about direction — a "connection" that only pushes data one way may not solve your problem. Our guide on how to build a software stack that actually connects covers what to check.
- Handling customer, health, or payment data? Security and compliance stops being a 10% criterion and becomes a gate: fail it and the tool leaves the shortlist regardless of the total.
- On a tight or unpredictable budget? Raise total cost of ownership and model the growth curve carefully — per-seat and per-contact models bill very differently as you scale, and a tool that's comfortable today can become the largest line in your stack after one good year.
- Small team, no admin? Weight setup effort and default configuration. Tools that require a specialist to configure carry a hidden staffing cost.
What are the criteria most buyers forget?
Data portability. Almost nobody tests the export until they're leaving, which is exactly the worst moment to discover it produces a format nothing else reads. Export on day one, open the file, and check whether relationships between records survived — not just the record list.
Contract and billing mechanics. Annual commitments, auto-renewal windows, notice periods, and what happens to your data after cancellation are all decision-relevant and all knowable before you sign. Ask for the terms in writing during the trial; a vendor unwilling to answer plainly has answered.
The downgrade path. Everyone models the upgrade. Ask what happens if you shrink — do seats release mid-term, or are you locked at your peak?
Admin and permissions. Who can see what, and whether you can remove a departing employee's access without breaking their automations.
Migration cost in hours, not dollars. The subscription price is visible; the two weeks someone spends cleaning and re-importing data is not. Include it in total cost of ownership, because it's usually the largest first-year number.
How do you turn criteria into a trial that proves something?
A trial is an experiment, and an experiment needs a hypothesis. Run it like this:
- Write the five real tasks first — actual work from last week, with real data, not the vendor's sample dataset.
- Give every candidate the same tasks and the same time box. Uneven effort produces uneven scores that feel like insight and aren't.
- Involve the daily user, not just the buyer. The person who will live in the tool should score usability; the person paying should score cost.
- Test one failure path deliberately — break something, undo something, ask support something — because reliability under stress differentiates tools that look identical when everything works.
- Run the export before the trial ends, while you still have an account to export from.
- Score the same day. Impressions decay into brand feelings within a week; write the numbers down while the evidence is fresh.
- Record the tie-breaker in advance. If two tools land within a few points, decide now what settles it — usually adoption risk or exit cost, rarely features.
Frequently asked questions
How many evaluation criteria should I use? Five to eight for most small-team purchases. Fewer, and you'll miss something structural like data portability. More, and the weights get so thin that no single criterion can influence the outcome, which defeats the point of weighting.
Should price be a criterion or a filter? Both, at different stages. Use budget as a hard filter to build the shortlist, then score total cost of ownership as a weighted criterion among the survivors — that way you compare cost shape, not just headline price.
How do I score criteria I can't test in a trial, like support quality? Generate evidence rather than guessing: open a genuine support ticket during the trial and record the response time and usefulness. For security, ask for current documentation and read the terms. If a vendor won't produce evidence, score it low — that's information, not a gap.
Can I reuse the same scorecard for every category? The structure yes, the weights and the functional-fit questions no. Functional fit always has to be rewritten per category — the criteria that matter for a CRM are pipeline, automation, reporting and value, as covered in our guide on how to choose a CRM for a small team, and they look nothing like the criteria for an email platform.
What if no tool scores well? Either your weights are wrong or the category doesn't solve your problem. Re-read the failing criterion: if every candidate fails the same one, you may be asking software to fix a process issue instead.
Criteria first, candidates second. Once your scorecard is written and weighted, the shortlist stops being a matter of taste — and that's the point at which a side-by-side comparison becomes useful rather than persuasive. Compare business software side by side on Nexuswoot to see how the same criteria play out across email marketing, CRM, SEO and live chat tools, with the scoring made explicit. Nexuswoot may earn a commission from some of the tools it compares; rankings follow the published criteria, not payouts.