Guide / Atlas James

Acceptance criteria: when is software
actually done?

“It works” is a start, not a finish line. Learn how to write testable acceptance criteria for features, integrations, performance, security and AI-assisted workflows throughout delivery.

By Atlas James6 min read

A feature looks finished in a demonstration. Then someone enters an unexpected postcode, the accounts system stops responding, or an AI-generated answer sounds convincing but is wrong.

“Done” needs more than a working happy path. Across the software development life cycle, acceptance criteria turn expectations into checks you can run, review and agree. They help you decide whether work is ready for its next stage, rather than leaving that decision to instinct.

Here’s how to write them, from the first conversation through to live operation.

Start with an observable result

Acceptance criteria describe the conditions a particular feature or change must satisfy. A shared definition of done covers recurring delivery standards, such as code review, relevant tests, documentation and deployment checks. You need both: one explains what this change must do; the other explains how your team delivers work responsibly.

A useful criterion identifies:

  • Context: the user, starting state, inputs and environment.
  • Behaviour: the action or event being tested.
  • Result: something observable, with a measurable threshold where needed.
  • Evidence and owner: how you’ll check it and who accepts the result.

“Users can update their details” leaves plenty open to interpretation. Try:

Given a signed-in customer with a valid UK postcode, when they save a new delivery address, then the address persists after refresh and appears on their next order. Existing orders remain unchanged.

The examples below are illustrative, not universal targets. Agree thresholds around your users, risks, budget and dependencies.

1. Discovery: agree what success means

Start with the job someone needs to complete. Identify users, permissions, exceptions and consequences of failure before discussing screens.

For each proposed capability, ask:

  • What must happen, and what must never happen?
  • Which systems or suppliers does it depend on?
  • What data will it handle?
  • Who can accept it, and what evidence will they need?

For an invoice integration, success might mean that approved invoices reach the accounts platform without duplication. For an AI-assisted support workflow, it might mean producing a draft for human review, not sending an answer automatically.

The discovery-stage acceptance gate should be explicit: named stakeholders approve the scope, exclusions, dependencies and initial criteria. Record unresolved questions with an owner and a decision date. An unknown supplier limit is an investigation task, not a promise.

2. Design: make the criteria testable

Translate the agreed outcomes into journeys, data rules and failure behaviour. Include empty states, invalid inputs, expired sessions and unavailable services.

Choose how each criterion will be tested. A calculation may suit an automated test; a keyboard journey needs interaction checks; an operational handover needs a walkthrough.

For performance, define the workload before choosing a speed target. For security, identify access boundaries and relevant threats. For AI, define permitted source material, prohibited actions and the evaluation examples.

The design gate is reached when the team can explain how it will verify each criterion, with test data and environments identified. Resolve conflicting expectations here: “instant updates” and a supplier’s overnight batch export cannot both describe the same integration.

3. Build: cover five kinds of acceptance criteria

Write criteria alongside the work, before implementation gets too far. These five areas catch different kinds of unfinished business.

Features: test outcomes and boundaries

Include successful actions, invalid inputs and permissions.

When a customer submits an address without a postcode, saving is blocked, a clear message identifies the missing field, and their other entries remain intact.

For accessibility, describe relevant interactions rather than writing “accessible”. For example, specify that the form can be completed using a keyboard, focus remains visible, and validation errors are associated with their fields. Wider accessibility requirements still need an agreed assessment scope.

Integrations: test failure and recovery

Cover field mapping, authentication failures, timeouts, duplicate messages and reconciliation.

When the same approved invoice event is delivered twice, only one invoice is created in the accounts platform, and both attempts are traceable without exposing credentials in logs.

Also specify retry limits and what happens afterwards. A supplier sandbox can support testing, but it may not reproduce production behaviour. Record that limitation and plan controlled live checks.

Performance: define the conditions

“Loads quickly” is not a test.

In the agreed staging environment, with 50 concurrent users and 20,000 representative records, 95% of order-search requests complete within two seconds during a 15-minute test, with no failed requests.

Those figures are example targets. Define what is measured, the dataset, traffic pattern and acceptable failure rate. Separate server response time from the full experience on a user’s device.

Security: test specific controls

Avoid “the system is secure”. Describe boundaries you can examine.

An authenticated customer cannot read or change another customer’s order through either the interface or a direct API request. Denied requests do not reveal that order’s contents.

Agree checks for secrets, logging, session handling and dependency risks where relevant. Passing these checks does not prove complete security or UK GDPR compliance; broader assurance and data-protection responsibilities need separate consideration.

AI-assisted workflows: test quality and control

Judge outputs against an agreed evaluation set and scoring rubric. Include incomplete documents, conflicting information and malicious instructions embedded in supplied content.

The workflow drafts answers using approved reference material and records supporting references. It cannot send replies; a named reviewer must approve each response before sending.

Set a measurable quality threshold, define unacceptable errors, and specify what happens when evidence is missing. Keep deterministic controls, such as permissions and approval requirements, separate from model-quality scores. A good average score must not excuse a critical failure.

4. Validation: collect evidence, not just opinions

Run the agreed checks against an identified build and configuration. Record pass or fail, supporting evidence and defects. For AI evaluations, also record the model, prompt, evaluation dataset and scoring method so comparisons remain meaningful.

User acceptance testing should involve people who understand the real work. Give them representative tasks rather than asking whether the software “looks right”.

The validation gate requires an explicit decision. Failed criteria should block acceptance unless an authorised owner agrees a documented exception, including its impact, mitigation and review date. Don’t quietly rewrite the target to match the result.

5. Release: distinguish accepted from ready to operate

A feature can pass testing while its release remains unready. Agree deployment checks covering monitoring, support ownership, migration verification and recovery.

For example, require a rehearsed recovery procedure in staging and confirmation that the on-call contact receives a test alert. Check whether rollback is actually possible after data changes; some releases need a forward-recovery plan instead.

Approve release only when operational checks pass and remaining risks are understood by the decision-maker.

6. Operation: keep the agreement useful

After release, check behaviour under real conditions. Monitor response times, integration backlogs and errors. For AI workflows, review output quality and approval behaviour within agreed privacy and retention boundaries.

Define triggers for intervention: a sustained backlog, worsening latency or a critical AI error should have an owner and a response. Re-run relevant acceptance checks when dependencies, models or business rules change.

A reusable acceptance record

Use this structure for each important criterion:

  • Outcome: What should the user or business achieve?
  • Scenario: Given… when… then…
  • Conditions: Data, environment, workload and dependencies.
  • Pass threshold: The observable result and permitted tolerance.
  • Evidence: Test result, evaluation report or walkthrough record.
  • Decision: Owner, status, exceptions and review date.

Start small, but make the decision clear. Good acceptance criteria don’t eliminate every risk. They make “done” an agreement you can inspect, rather than a feeling you have to defend.

If you’d like help turning a loose brief into testable delivery criteria, let’s talk.

  • Software development
  • Acceptance criteria
  • Software testing
  • Project delivery
  • AI workflows
Atlas James

Software, apps and AI workflows. Independent thinking from our studio in Kettering, backed by 15+ years of team experience.

Put an idea into practice