Skip to main content

QA · AI

Tester Kit: agents run the tests, people decide

From 3–4 manual tasks a day to 10–15 agent sessions running in parallel

RoleProcess designer, skill builder and trainer

Documents from this case: Tester Kit, Scenario
automated test cases with agents
1,000+
Tester team time
−70%
recurring bugs in the same module
30% → 5%
agent sessions in parallel
10–15

Client names are withheld; examples and data are simulated. The process, methods and result figures are real.

The problem

The Tester team tested 100% manually. Most members are not strong technically, so they could not write automation the traditional way. Each person finished at most 3 to 4 tasks a day, with not enough time to re-run old cases, while the Dev team was already using AI and shipping features faster than they could be tested.

Before Tester KitOld bugs came back because no one had time to re-run old cases
  • Recurring bugs
  • New bugs
  • Bugs in one moduleabout 30% were recurring bugs

Context and role

  • A Tester team of four, serving several projects in parallel.
  • Test cases were kept in spreadsheets, with many cases repeated between releases.
  • I designed the process, built the skills, trained the team and tracked adoption, alongside my BA role.
  • Constraints: testers do not write code; agents must not decide pass or fail on their own; every result must have evidence.

Step 1: Classifying the test cases

Not every case should go to an agent.

ClassificationWhich cases go to agents and which are tested by hand
ExampleGiven to agentWhat the person does
API and data checksEndpoint returns the right code, database is in the right stateYesVerifies the report
Repetitive, stable UI flowsSign in, create a record, filter a listYesVerifies the evidence
Multi-step business flows with calculationsPartial payment, then check receivablesYesVerifies each step and the figures
Look and feel, user experienceLayout, colours, messagesNoTests by hand
Third-party integration, real paymentsWebhooks, payment gatewayNoControlled manual testing
ExploratoryFinding defects outside the scriptsNoA skilled tester

Step 2: Designing Tester Kit

The kit has 8 skills, packaged so that a tester only needs to install it in the IDE to run the automated cases, with no code to write. The representative skills form a chain:

Tester KitOne job per skill, one output per job
  1. Generate test casesfrom stories, acceptance criteria, business rules
  2. Pre-run checkenvironment, sample data, test accounts
  3. Run a case seton staging, screenshots, logs
  4. Regressionre-run old cases for the module that changed
  5. Draft bugsteps to reproduce, expected, actual
  6. Collect evidencefolder to attach to the ticket
RulesApplied to every skill
  • The agent gives only a preliminary result: looks like a pass, looks like a fail, or not sure.
  • Every result must come with evidence: screenshots, logs, actual values.
  • Do not touch production data. Use test accounts only.
  • The final pass or fail decision belongs to the tester.

Step 3: Running in parallel

Each tester opens 10 to 15 IDE agent tabs (Claude Code, Codex, Antigravity), and each session runs one group of cases. The agents run first while the tester watches the summary sheet.

A tester's day
  • Tester
  • Agent
  1. TesterSplit cases into groups
  2. AgentRun in parallelone group of cases per session10–15
  3. AgentPreliminary resultswith evidence
  4. TesterVerify the evidence
  5. TesterManual testingfailed cases, cases that cannot be automated
  6. TesterQA decision
Only the failed cases and the cases that cannot be automated need to be tested by hand.

Step 4: The “looks like a pass” situation

The simulated example below shows why a person still has to verify. Case: record VND 1,000,000 against an invoice of VND 3,000,000; receivables must go down by exactly VND 1,000,000.

What the agent sees

  • The screen shows “Partially paid”.
  • The balance has gone down.
  • Conclusion: looks like a pass.

What the tester sees

  • The balance on screen did go down.
  • But the invoice record in the database has not changed status: only the interface was updated temporarily.
  • Decision: fail. The draft-bug skill creates a report with before and after screenshots and the query result.

Step 5: The tester’s new role

The tester goes from doing everything by hand to verifying the results of what AI has done.

New role
  1. Design test casesAI produces the draft
  2. Choose cases for agents
  3. Coordinate agent sessionswatch the summary sheet
  4. Verify the evidencetest failed and uncertain cases by hand
  5. Decide pass or fail
  6. Bug report with evidence
Both the Tester team and the BA team were trained to use the skill set.

Results

By internal measurement:

Time on repetitive tasks−70%
Agent sessions each tester coordinates10–15instead of 3–4 manual tasks a day
Automated test cases written1,000+plus hundreds of manual cases
Testers moved to the new process4/4
  • Before
  • After
  • Recurring bugs in the same module✓ 30% → 5%

    30%
    5%
Because old cases are re-run after every change.

On the US client project, counted from the project documentation:

Automated test cases, last 2 weeks of 09/2026901across 31 runs; 9 false passes blocked
Real tickets used to distil the kit20+in 6 days
Recurring pitfalls collected39from about 95 real failures
AI tools it runs on4Claude Code, Codex, OpenCode, Antigravity

The kit has been handed over to testers, including one on Ubuntu, with a guide for non-technical users.

What I learned

  • People who do not write code can still direct automated testing if the tool speaks their language: test cases and evidence.
  • The agent is not the one who decides pass or fail. The tester remains responsible for verifying the test results.
  • Automation is most valuable where it frees up time for the work that needs people: hard cases and exploratory testing.
  • Test case
  • Integration test
  • Regression
  • Automated testing
  • IDE agent
  • Governance

Want to talk about this case?

Contact