Skip to content
FunDev
FunDev
playwright

Playwright Agents: AI Plans, Writes, and Repairs Tests

Playwright Agents: AI Plans, Writes, and Repairs Tests
2 views
8 min read
#playwright

Writing E2E tests takes substantial resources, so from an efficiency standpoint they rarely get written. Tests break when selectors change, and the time spent repairing them makes it hard to focus on developing new features.

The Test Agents introduced in Playwright offer an interesting answer. AI explores the application, creates a test plan, generates code, and even repairs failing tests itself.


The Roles of the Three Agents

┌─────────────────────────────────────────────────────────────────┐
│                    Playwright Agents Flow                       │
│                                                                 │
│   🎭 Planner          🎭 Generator         🎭 Healer            │
│   ┌───────────┐       ┌───────────┐       ┌───────────┐        │
│   │ 앱 탐색     │  ──▶  │ 코드 생성   │  ──▶  │ 실패 수정   │        │
│   │           │       │           │       │           │        │
│   │ specs/*.md│       │ *.spec.ts │       │ 선택자 갱신  │        │
│   └───────────┘       └───────────┘       └───────────┘        │
│                                                                 │
└─────────────────────────────────────────────────────────────────┘

Planner: Explore the App and Create a Test Plan

Planner opens a browser and explores the application directly. If there is a login page, it identifies the fields and buttons, then creates a test plan in Markdown based on what it finds.

# Login Test Plan
 
## Scenario 1: Successful Login
**Steps:**
  1. Navigate to /login
  2. Enter 'funq' in the 이름 field
  3. Enter 'password' in the 비밀번호 field
  4. Click 로그인 button
 
**Expected Results:**
  - Redirect to main calendar page
  - User session is established

In effect, AI performs exploratory testing for you.

Generator: Turn the Plan into Runnable Code

Generator reads Planner's plan and produces Playwright test code that actually works.

The key is live verification. Before writing code, it opens a browser and checks whether the action actually works. It verifies that getByRole('button', { name: '로그인' }) really finds the button before writing the test.

test('should verify successful login', async ({ page }) => {
  await page.goto('/login');
  await page.getByLabel('이름').fill('funq');
  await page.getByLabel('비밀번호').fill('password');
  await page.getByRole('button', { name: '로그인' }).click();
 
  await expect(page).toHaveURL('/');
});

Healer: Automatically Repair Failing Tests

Suppose a developer changes a button's Korean “Log in” label to “Sign In.” The test fails.

Healer analyzes the **accessibility tree** at the failure point to find a semantically equivalent element. “There is no button with the Korean 'Log in' name, but there is a 'Sign In' button in the same place” → update the selector and rerun the test.

❌ Error: getByRole('button', { name: '로그인' }) - Element not found
✅ Fix: Updated to getByRole('button', { name: 'Sign In' })

Installation and Setup (Claude Code)

1. Install Playwright

npm init playwright@latest

2. Initialize the Agents

npx playwright init-agents --loop=claude

Files generated by this command:

.claude/agents/
├── playwright-test-planner.md    # Planner 에이전트 정의
├── playwright-test-generator.md  # Generator 에이전트 정의
└── playwright-test-healer.md     # Healer 에이전트 정의
specs/
└── README.md                     # 테스트 계획서 저장 위치
tests/
└── seed.spec.ts                  # 시드 테스트 (로그인 등 초기 설정)
.mcp.json                         # MCP 서버 설정

3. Add the MCP Server

claude mcp add playwright npx @playwright/mcp@latest

You can now check the agent list with /agents in Claude Code.


The Workflow in Practice

Step 1: Ask Planner for a Test Plan

Use the playwright-test-planner subagent to explore http://localhost:3000
and create a test plan for:
1. 로그인 기능
2. 캘린더 뷰 전환 (월/주/일)
3. 일정 생성

Save it as specs/calendar.md.
Use tests/seed.spec.ts as the seed.

Planner explores the app and writes detailed test scenarios to specs/calendar.md.

Step 2: Ask Generator to Generate Code

Use the playwright-test-generator subagent to generate tests
from specs/calendar.md.
Put generated tests under tests/calendar/.

Generator creates runnable test code from the plan.

Step 3: Run the Tests and Use Healer

npx playwright test

If any tests fail:

Use the playwright-test-healer subagent to fix the failing tests.
Keep the original test intent - only fix locators and waits.

A Real Application: The Family Calendar Project

I tried Playwright Agents on the calendar project.

Part of the Test Plan Created by Planner

### 4.3. Create weekly recurring event with specific weekdays
 
**File:** `tests/event-management/recurring-weekly-custom.spec.ts`
 
**Steps:**
  1. Open event creation dialog
  2. Enter title '수영 수업'
  3. Select '매주' (Weekly) from recurrence dropdown
  4. Click weekday buttons: 월, 수, 금
  5. Verify selected buttons are highlighted
  6. Save the event
 
**Expected Results:**
  - Multiple weekdays can be selected
  - Backend receives weekdays: ['MO', 'WE', 'FR']
  - Calendar shows event on correct days

It understood the Korean UI naturally and included even the complex recurring-event logic in the plan.

Code Created by Generator

test.describe('Adding New Todos', () => {
  test('Add Valid Todo', async ({ page }) => {
    const todoInput = page.getByRole('textbox',
      { name: 'What needs to be done?' });
    await todoInput.click();
    await todoInput.fill('Buy groceries');
    await todoInput.press('Enter');
 
    await expect(page.getByText('Buy groceries')).toBeVisible();
    await expect(page.getByText('1 item left')).toBeVisible();
  });
});

It prioritizes getByRole and getByText—user-facing locators. Since they do not depend on CSS selectors, they are more resilient to UI changes.


The Core Design Principle: The Accessibility Tree

What distinguishes Playwright Agents from other AI automation tools is its use of the accessibility tree rather than screenshots.

┌─────────────────────────────────────────────────────────────┐
│  Vision Model 방식              │  Accessibility Tree 방식   │
├─────────────────────────────────┼────────────────────────────┤
│  스크린샷 픽셀 분석              │  DOM 구조화된 데이터        │
│  해상도에 따라 오차 발생          │  역할(Role) + 이름(Name)    │
│  토큰 소모량 높음                │  토큰 효율적               │
│  이미지 처리 오버헤드            │  빠른 응답 속도            │
└─────────────────────────────────┴────────────────────────────┘

It conveys the page structure to AI in the same way a screen reader describes a web page to a visually impaired user. Noise such as advertisements and decorative elements is removed, leaving only semantic information such as buttons, links, and input fields.


The Importance of a Seed Test

For apps with complex authentication, seed.spec.ts is essential.

// tests/seed.spec.ts
import { test as setup, expect } from '@playwright/test';
import path from 'path';
 
const authFile = path.join(__dirname, '../playwright/.auth/user.json');
 
setup('authenticate', async ({ page }) => {
  await page.goto('/login');
  await page.getByLabel('이름').fill(process.env.E2E_USER!);
  await page.getByLabel('비밀번호').fill(process.env.E2E_PASS!);
  await page.getByRole('button', { name: '로그인' }).click();
 
  await expect(page).toHaveURL('/');
 
  // 세션 상태 저장
  await page.context().storageState({ path: authFile });
});

With this setup, Planner, Generator, and Healer all start authenticated. They can focus on testing the core features without wasting tokens analyzing the login page.


Practical Tips

1. Give Healer Clear Guidelines

Healer may delete assertions or change the test's intent to make it “pass.” Specify constraints in the prompt:

Fix the failing tests, but:
- Keep the original test intent
- Do NOT remove assertions
- Only update locators and waits

2. Regenerate the Agents When Updating Playwright

npx playwright init-agents --loop=claude

Refresh the agent definition files to incorporate tools and instructions added in new versions.

3. Manage Token Costs

Loading the MCP server alone consumes 4,000–8,000 tokens. It is a good idea to disable the server when testing is done, or use a separate session.


Limitations

It Cannot Detect Logic Changes

Healer is good at handling selector changes, but cannot distinguish cases where a test fails because the application's business logic has changed. In those cases, it marks the test with test.fixme() and requests developer intervention.

Non-Interactive Design

Healer attempts the most reasonable fix without asking the user questions. In complex situations, the applied change may differ from the intended one, so its changes must be reviewed.

Token Costs

Exploring a complex app can consume a substantial number of tokens.


Closing Thoughts

Playwright Agents is an attempt to change the way tests are written.

Previously, developers had to code every selector and step explicitly. Now, they can provide a high-level intent such as “Verify the user-registration flow.”

It is not perfect yet. Token costs, detection of complex logic changes, and the non-interactive design still need improvement. But the direction is clear: the QA engineer's role will evolve from “test scripter” to “agent architect.”

Rather than writing individual test cases, the core skills will be designing the system prompts agents follow and building robust seed tests and fixtures.


References

Leave a comment if you have any questions!

Related posts

Comments

Korean and English pages share this conversation.

Write a comment

0 / 5,000
You will need this password to edit or delete this comment.

Loading comments…