# Build a Playwright harness your AI agents must pass

A client-owned automation suite for critical workflows, with CI evidence, protected test ownership, and enforced review before agent-generated changes merge.

Mottobits · Reviewed 2026-09-30

Give your AI agents a repeatable way to prove that a change preserves the business workflows that matter. Build the Playwright suite in your own repository, run it in CI, and keep the test contract and enforcement settings under accountable human ownership. The agent can propose a fix and use failure evidence to improve it; its normal permissions should not let it waive the checks, approve its own changes, or rewrite the merge rules.

## Turn critical workflows into an independent contract

Name the customer actions a change must preserve before asking an agent to implement it. For each workflow, record the role, starting data, expected outcome, and consequence of failure. A human who understands the business approves that contract.

- Cover a successful path, invalid input, and a denied action; verify that rejected requests do not change protected data.
- Choose the relevant browsers, roles, tenants, and integration boundaries. Document known defects separately from intended behavior.
- Give each critical scenario a stable identifier and an owner so reviewers can see exactly which behavior the suite protects.

## Build a suite the agent can run and humans can trust

Use isolated fixtures and a controlled test environment. Prefer role or label locators and assertions that wait for an observable result. A passing page load is not enough for a workflow that saves data, changes permissions, or triggers a business action.

- Keep tests independent, use synthetic or approved sanitized data, and isolate payments, email, and other external side effects.
- Check persisted outcomes and relevant API boundaries as well as the UI. Demonstrate that a deliberate wrong result makes the relevant assertion fail.
- Enable Playwright's forbidOnly in CI. Separately reject unapproved skips, narrowed test selection, weakened assertions, and snapshot changes; forbidOnly does not cover those cases.
- Document installation, fixture reset, local execution, and failure reproduction so the client can maintain the harness without the original implementer.

## Protect the harness and the merge path

Assign human code owners to the tests, fixtures, Playwright configuration, test-runner scripts, CI workflows, and CODEOWNERS itself. Enable required owner review; a CODEOWNERS file alone does not enforce approval. Protect the target branch with an active ruleset requiring pull requests, review of the latest changes, and the named CI check.

- Select the expected status-check source where supported, and verify that another identity cannot satisfy the check with an unrelated success status.
- Give agents only the repository and task permissions they need. Keep rule administration, bypass rights, protected-branch writes, production credentials, and approval authority outside the agent identity.
- Review bypass lists, including administrator, team, and app entries. A bypass-capable identity can undermine the intended gate.
- Limit the CI token's permissions and use isolated runners and test credentials. A check source restriction does not make an editable workflow trustworthy.

## Require evidence and test the controls

Run the agreed suite on the change being reviewed and retain the commit, environment, executed scenario IDs, results, and failure evidence. Configure Playwright traces on a failed run or retry and keep artifacts access-controlled. A stable required job should fail when expected scenarios did not run or their result is missing.

- Preserve failed-attempt evidence and track flaky tests. A passing retry must not silently remove an agreed critical scenario from review.
- Keep approval explicit for test quarantine, updated snapshots, reduced coverage, and CI changes. Do not use success-on-error settings for the required gate.
- Rehearse with a disposable pull request: a broken workflow fails CI, a deleted or skipped critical test requires intervention, and a modified harness cannot merge without its owner.
- Use the same rehearsal to confirm that the agent's real credentials cannot bypass the active rules. Record any exception and who controls it.

## Illustrative example — not customer results

Update contact validation within the agreed application files. Run the approved Playwright scenarios for authorized edits, invalid values, and read-only users. Use failure traces to diagnose problems. Do not remove checks, change fixtures to hide a failure, update snapshots, or alter CI to get a pass. If the contract needs to change, explain why and request owner review.

- An authorized edit survives reload; invalid input and a read-only user's update leave the stored record unchanged.
- The CI report identifies the exact commit and shows every agreed critical scenario ran.
- A seeded defect makes the appropriate test fail and produces reviewable diagnostic evidence.
- Changes to tests, fixtures, runner configuration, workflows, or ownership rules receive the required human approval.
- The agent's configured identity cannot merge through a failing required check or waive the review.

## Deliverables

- Critical-workflow map and approved acceptance criteria
- Client-owned Playwright tests, fixtures, and run instructions
- CI reports and traces with retention and access rules
- Protected test and workflow ownership with verified required checks
- Agent permission boundary, control rehearsal, and maintenance handover

## Limits

- No harness is unbreakable. Administrators or other authorized identities can change repository settings, and an identity with applicable bypass rights can circumvent rules. Verify permissions and enforcement in the client's actual repository.
- Passing tests only establish the behavior exercised by those tests. Missing scenarios, incorrect expectations, or compromised execution can still let a defect through; retain human review and other appropriate checks.
- GitHub protection features depend on plan and repository visibility. Confirm availability before promising a particular enforcement setup.
- The workflow and examples are illustrative delivery guidance, not claims about a completed client engagement.

## Sources

- [Playwright: Best practices](https://playwright.dev/docs/best-practices)
- [Playwright: Test configuration](https://playwright.dev/docs/test-configuration)
- [Playwright: Continuous integration](https://playwright.dev/docs/ci)
- [Playwright: Trace viewer](https://playwright.dev/docs/trace-viewer)
- [GitHub: About rulesets](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-rulesets/about-rulesets)
- [GitHub: Available rules for rulesets](https://docs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-rulesets/available-rules-for-rulesets)
- [GitHub: About code owners](https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/about-code-owners)
- [GitHub: GITHUB_TOKEN permissions](https://docs.github.com/en/actions/tutorials/authenticate-with-github_token)

Co-written by Adnan Rafiq and ChatGPT. ChatGPT assists with research, drafting, and source checks; the engineering direction comes from Adnan. Examples are illustrative unless an article explicitly provides executed results. These are working methods, not customer case studies or guaranteed outcomes.
