QAWAI / Product quality and documentation

Build with AI.Demand evidence.

QAWAI means QA with AI. I built it to bring evidence-based software review into the coding agent a founder already uses, with a clear method for finding defects, challenging assumptions and verifying what deserves action.

My contribution

Product lead · UX strategy, review methodology and AI-directed delivery

Project context

v1 build complete · Public release pending

Born from the difficulty of building with AI

I’m Ivan Pedai, a UX Director and AI DesOps Lead. While building Launcherry, my flagship product, I was already using Claude Code and Codex to implement and review work. Yet getting a convincing result was only the beginning: bugs, broken flows and documentation drift still needed a deliberate way to be found, understood and resolved.

I built QAWAI around that problem: QA with AI, directed by human judgment. Its review methodology connects intended behaviour, documentation, code and observable product behaviour. It defines the scope, brings different disciplines into the review and turns findings into decisions and fix jobs. The methodology’s catalog contains more than 1,100 entries, selected for the product and authorized review scope. The shipped CLI/MCP engine applies the criteria suitable for read-only software review; it does not execute the entire catalog on every audit.

My role: turn a working method into a product

I own the concept, product strategy, UX, review methodology and acceptance decisions. I am the sole developer; AI delivers the implementation under my direction. My contribution is deciding what the product should help people understand, how the workflow should behave and which evidence is sufficient to act.

The founder’s question is practical: what deserves attention before the next release? I designed a plain-language report that connects each finding to its evidence, its implications and a copy-ready fix prompt. Technical detail has to support a decision.

The first proof came from Launcherry

The July 2026 review and follow-up cycle using QAWAI’s early methodology recorded 59 findings across product quality, UX, code and documentation. By the 11 July checkpoint, 57 were fixed or resolved: 97%, rounded. Two low-priority checks remained blocked by hardware or live-model access. The initial audit also added 26 regression tests, each shown to fail before the relevant fix or guard.

One concrete example: the mobile navigation dock covered the campaign approval bar, making approval impossible on a phone. The review and fix cycle separated that product defect from test failures, corrected the layout and restored mobile checks as a delivery gate. That is the value I wanted to preserve: an observable failure, a clear decision and verification of the correction.

These are development-phase results from applying QAWAI’s review methodology to my own project, rather than customer outcomes or a count of bugs missed by a particular model. Later independent reviews still found defects; I used them to improve coverage and challenge the method itself.

Explore the bugs and their verification

The strategic pivot: meet founders inside their coding agent

The first direction was an installable desktop app. On 31 July, I chose to deliver the review engine through a command-line tool and MCP, the connection standard coding agents use. The founder was already building inside an agent; asking them to install and learn a separate application added another workflow to maintain.

The engine, evidence verification, findings ledger and HTML report carried forward. The pivot removed desktop packaging, macOS signing and notarisation, auto-update work and Electron-specific compatibility from the launch path. Development could focus on the review capability instead of completing a second application around it.

Over MCP, the existing coding agent performs the review and QAWAI checks the evidence and builds the report. This also avoids launching a second coding-agent session inside the first. I reduced the scope that needed development and ongoing support, and kept the product close to where fixes happen. I did not measure the counterfactual hours or tokens saved.

Keep the specification and the implementation in conversation

AI delivery becomes harder when the intended product, its documentation and its code disagree. QAWAI starts by clarifying intent and checking claims against the implementation. This gives specification-led development a feedback loop: a written requirement can be challenged by what the product actually does.

QAWAI’s Full Audit includes a dedicated documentation-versus-code review. A conflict requires evidence from both sides. Its reconciliation questions capture the owner’s intent as decision records, so an accepted tradeoff can be distinguished from an unintended defect. Documentation generation and ongoing maintenance remain a planned extension.

Verify the evidence before asking someone to act

QAWAI prepares a sanitized, disposable copy of the project. The selected agent reviews that copy; the verifier then checks the quoted excerpts against its files. Findings whose evidence does not reproduce are dropped before reporting.

A matching excerpt does not establish that an interpretation is correct. The report states its scope and limitations. QAWAI’s CLI/MCP review is read-only. Browser checks, executed failure tests and live-service validation remain separate steps in my development workflow. The architecture below shows where the agent’s reasoning ends and the product’s verification begins.

Apply the quality bar to the product itself

I used the same review, fix and verification discipline on QAWAI. Project records document regression checks, an installed-package test and a completed read-only bypass review. The v1 implementation is complete; a final live re-audit and public-release setup remain open.

This case demonstrates how I connect UX judgment, product strategy and AI-directed delivery: identify a problem in real work, make the method repeatable, remove unnecessary delivery scope and help the owner make an informed decision.

The method in practice / Selected findings

The bugs that make
“looks finished” dangerous.

I built QAWAI to challenge the assumptions behind a working product: what reaches the AI, what can execute, when money changes hands and whether the test environment can expose the failure. These cases come from the early QAWAI methodology applied to Launcherry, QAWAI’s self-audit and later bypass reviews.

QAWAI · 31 July 2026

P0 · Critical

The “sanitized” copy still contained secrets

The audit workspace excluded common secret files, yet copied .dev.vars, a platform-specific file used for local credentials. Sensitive values could reach the reviewing AI and be quoted into a shareable report.

Why it was easy to miss. The familiar .env rules and their tests passed. The blind spot was another platform’s naming convention, inside the very feature meant to protect the user.

Verified correction. The exclusion rules were expanded. A re-audit through the installed package confirmed that the file was absent from the copy. Three review passes had independently identified the defect.

QAWAI · 3 October 2026

P1 · High

A repository could start a command before its review

A repository’s agent configuration survived into the audit copy. The default driver could load it and start a repository-supplied command on the user’s machine, despite the read-only review promise.

Why it was easy to miss. The risk happened during agent startup, before the review itself. Pre-approving read tools also did not remove the agent’s other tools.

Verified correction. Configuration files were excluded and the driver’s available tools restricted. A real CLI probe created a marker with the old settings; with the new settings, the fixture’s command did not run and only read tools remained. No live model call was needed for that check.

QAWAI · 3 October 2026

P1 · High

Paid content arrived before payment became final

A funded licence could download a paid audit catalog, then release the reserved credits. Letting the reservation expire had the same effect: the content was delivered and the balance restored.

Why it was easy to miss. Reserve, download and refund each looked reasonable in isolation. The bypass appeared when the full sequence was challenged. A probe downloaded the pack five times while the balance stayed at three credits.

Verified correction. Delivery now commits the charge on the server. Tests failed on the old release and expiry paths, then passed after the fix. A fresh review also caught a new failure case: a lost client confirmation could stop an already-paid audit. That received its own failing regression test and correction.

Launcherry · 11 July 2026

P2 · Medium

The test database hid a launch race

Concurrent launch requests could all pass the same read-then-write guard. The lightweight test database serialized requests, hiding the overlap that a production-style database could expose.

Why it was easy to miss. A green concurrency test was not evidence of real concurrency. The review kept the race as an unconfirmed hypothesis until the right environment was available.

Verified correction. On real PostgreSQL, eight of eight concurrent requests passed the old guard. Database uniqueness constraints and conditional approval were added, with regression checks. This case is included for the testing lesson; its recorded severity was Medium.

The discipline behind the findings

QAWAI’s July self-audit accepted 53 evidence-verified findings, including 1 Critical and 11 High. All twelve were fixed, none waived; subsequent reviews continued to challenge those fixes. The initial review of Launcherry added 26 regression tests shown to fail before the relevant fix or guard.

My role is to define the product’s promises, direct the review and decide what evidence is enough to act. That includes challenging QAWAI’s own advice: one generated payment fix would have broken legitimate audit resumes, so I chose receipt-based entitlement instead.

Dated findings from my own products, with severity as recorded in their reports. The cases span different review methods and are not one combined audit or customer benchmark. Executed failure checks belong to my development workflow; QAWAI’s read-only CLI/MCP engine checks quoted file evidence. A missing finding in a later report is not, by itself, proof of a fix.

QAWAI / Product architecture

Your agent reviews.
QAWAI checks the evidence.

A high-level map of the built Full Audit workflow. The owner provides intent; the selected coding agent supplies reasoning; the local engine verifies excerpts and records the result.

Inside the founder’s existing workflowTerminal or coding agent → QAWAI CLI / MCP
  1. Prepare a copy

    Sanitized disposable workspace; source project unchanged by the audit.

  2. Clarify intent

    Owner questions and decision records; compare documentation with code.

  3. Review by discipline

    Six role passes performed by the selected agent at Full Audit depth.

  4. Verify and report

    Check excerpts against files; consolidate findings into a ledger and HTML report.

SecurityPrivacyArchitectureInterface consistencyFailure casesMisuse resilience
Owner decision → fix prompt → implementation → re-audit

Fixes happen in the owner’s development workflow. A re-audit compares reported findings; absence in a later report does not by itself prove a defect is fixed.

Code is processed by the selected AI provider, not uploaded to QAWAI’s servers. Excerpt verification establishes that the cited evidence exists; it does not certify the interpretation or complete defect coverage. This read-only product does not execute browser, load or live-service tests.

Continue exploring

Enterprise employee app

Read the next case

Clearer decisions.
Accountable AI delivery.

Discuss a design leadership or AI-enabled delivery role.