QA · AI
Tester Kit: agents run the tests, people decide
From 3–4 manual tasks a day to 10–15 agent sessions running in parallel
RoleProcess designer, skill builder and trainer

- automated test cases with agents
- 1,000+
- Tester team time
- −70%
- recurring bugs in the same module
- 30% → 5%
- agent sessions in parallel
- 10–15
Client names are withheld; examples and data are simulated. The process, methods and result figures are real.
The problem
The Tester team tested 100% manually. Most members are not strong technically, so they could not write automation the traditional way. Each person finished at most 3 to 4 tasks a day, with not enough time to re-run old cases, while the Dev team was already using AI and shipping features faster than they could be tested.
- Recurring bugs
- New bugs
Bugs in one moduleabout 30% were recurring bugs
Context and role
- A Tester team of four, serving several projects in parallel.
- Test cases were kept in spreadsheets, with many cases repeated between releases.
- I designed the process, built the skills, trained the team and tracked adoption, alongside my BA role.
- Constraints: testers do not write code; agents must not decide pass or fail on their own; every result must have evidence.
Step 1: Classifying the test cases
Not every case should go to an agent.
| Example | Given to agent | What the person does | |
|---|---|---|---|
| API and data checks | Endpoint returns the right code, database is in the right state | Yes | Verifies the report |
| Repetitive, stable UI flows | Sign in, create a record, filter a list | Yes | Verifies the evidence |
| Multi-step business flows with calculations | Partial payment, then check receivables | Yes | Verifies each step and the figures |
| Look and feel, user experience | Layout, colours, messages | No | Tests by hand |
| Third-party integration, real payments | Webhooks, payment gateway | No | Controlled manual testing |
| Exploratory | Finding defects outside the scripts | No | A skilled tester |
Step 2: Designing Tester Kit
The kit has 8 skills, packaged so that a tester only needs to install it in the IDE to run the automated cases, with no code to write. The representative skills form a chain:
- Generate test casesfrom stories, acceptance criteria, business rules
- Pre-run checkenvironment, sample data, test accounts
- Run a case seton staging, screenshots, logs
- Regressionre-run old cases for the module that changed
- Draft bugsteps to reproduce, expected, actual
- Collect evidencefolder to attach to the ticket
- The agent gives only a preliminary result: looks like a pass, looks like a fail, or not sure.
- Every result must come with evidence: screenshots, logs, actual values.
- Do not touch production data. Use test accounts only.
- The final pass or fail decision belongs to the tester.
Step 3: Running in parallel
Each tester opens 10 to 15 IDE agent tabs (Claude Code, Codex, Antigravity), and each session runs one group of cases. The agents run first while the tester watches the summary sheet.
- Tester
- Agent
- TesterSplit cases into groups
- AgentRun in parallelone group of cases per session10–15
- AgentPreliminary resultswith evidence
- TesterVerify the evidence
- TesterManual testingfailed cases, cases that cannot be automated
- TesterQA decision
Step 4: The “looks like a pass” situation
The simulated example below shows why a person still has to verify. Case: record VND 1,000,000 against an invoice of VND 3,000,000; receivables must go down by exactly VND 1,000,000.
What the agent sees
- The screen shows “Partially paid”.
- The balance has gone down.
- Conclusion: looks like a pass.
What the tester sees
- The balance on screen did go down.
- But the invoice record in the database has not changed status: only the interface was updated temporarily.
- Decision: fail. The draft-bug skill creates a report with before and after screenshots and the query result.
Step 5: The tester’s new role
The tester goes from doing everything by hand to verifying the results of what AI has done.
- Design test casesAI produces the draft
- Choose cases for agents
- Coordinate agent sessionswatch the summary sheet
- Verify the evidencetest failed and uncertain cases by hand
- Decide pass or fail
- Bug report with evidence
Results
By internal measurement:
- Before
- After
Recurring bugs in the same module✓ 30% → 5%
30%5%
On the US client project, counted from the project documentation:
The kit has been handed over to testers, including one on Ubuntu, with a guide for non-technical users.
What I learned
- People who do not write code can still direct automated testing if the tool speaks their language: test cases and evidence.
- The agent is not the one who decides pass or fail. The tester remains responsible for verifying the test results.
- Automation is most valuable where it frees up time for the work that needs people: hard cases and exploratory testing.
- Test case
- Integration test
- Regression
- Automated testing
- IDE agent
- Governance
