# Review AI-generated tests before they become release evidence

A practical review contract for using AI to draft tests while preserving independent expectations, meaningful assertions, and human ownership.

Mottobits · Reviewed 2026-09-30

An AI assistant can draft a test that passes while protecting the wrong behavior. The reviewer's job is to establish an independent expectation and inspect how the test proves it. GitHub's own responsible-use guidance places validation of generated output with the user. The workflow below turns that responsibility into visible review steps.

## Write the contract before asking for code

Name the behavior, the setup, and the expected result without copying the implementation into the assertion. Give the assistant only context approved for that tool. Synthetic examples are usually enough to discuss a test design.

- Identify the requirement or defect the test protects.
- State the important negative path and permission boundary.
- Say which files may change and which production behavior must remain untouched.
- Require the assistant to report assumptions, commands run, and checks it could not execute.

## Review the assertions, then the fixtures

A test named 'rejects unauthorized updates' is not evidence of that behavior if it only checks that a page loads. Read the action and assertion together. Then check whether mocks, fixtures, shared accounts, or permissive setup make failure impossible.

- Reject empty assertions, unconditional truth, swallowed exceptions, and snapshots updated solely to make the run pass.
- Check that the tested operation reaches the intended code or boundary.
- Verify both the rejection response and the absence of an unauthorized side effect.
- Look for time-dependent values, shared mutable state, uncontrolled randomness, and hidden network calls.

## Keep a human-owned merge decision

Run the relevant checks in a controlled environment and inspect their evidence. Where practical, demonstrate that the new test detects the defect or a deliberately altered expectation. Automated review is another input; it is not a substitute for a named reviewer who understands the change.

- Review production-code edits separately from generated test edits.
- Require a human to explain why an existing assertion or snapshot changed.
- Retain the final diff, executed commands, results, and any unverified assumptions with the pull request.
- Do not give a test-writing agent production data or unrestricted deployment credentials.

## Illustrative example — not customer results

Draft tests for an invoice approval endpoint using synthetic fixtures. Only an approver may approve a pending invoice. The test must check the response and the persisted state. Do not change production authorization code, weaken existing assertions, or update snapshots. Return the assumptions and commands you ran.

- An approver can approve a pending invoice and the state changes once.
- A non-approver receives the expected rejection and the invoice remains pending.
- Approving an already approved invoice follows the documented idempotency or rejection rule.
- The reviewer can map each assertion to a requirement and reproduce the result.

## Deliverables

- Repository-specific AI contribution rules
- Test review checklist
- Reviewed tests and execution evidence
- List of unresolved assumptions

## Limits

- The website and readiness calculator do not require model calls. Using a separate AI coding product may require the visitor's own account or subscription.
- These are proposed delivery controls, not a claim that any AI model detects every defect or security issue.

## Sources

- [GitHub: Application card for Copilot agents](https://docs.github.com/en/copilot/responsible-use/agents)
- [GitHub: Risks and mitigations for Copilot cloud agent](https://docs.github.com/en/copilot/concepts/security-governance-and-network-settings/risks-and-mitigations)

Co-written by Adnan Rafiq and ChatGPT. ChatGPT assists with research, drafting, and source checks; the engineering direction comes from Adnan. Examples are illustrative unless an article explicitly provides executed results. These are working methods, not customer case studies or guaranteed outcomes.
