•

Drizz raises $2.7M in seed funding •

•

Featured on Forbes

•

Drizz raises $2.7M in seed funding •

•

Featured on Forbes

Logo

Schedule a demo

?
Payment succeeds, order creation fails
never written

Specs describe what should happen. Bugs live in what could.

Drizz reads your app, maps everything that could happen around each feature, and proposes what to test, with the reasoning attached. Your team approves the plan. Drizz writes the tests.

The requirement
"Users can apply a coupon at checkout."
2
what should happen
25
what could
Code accepted
Total drops 20%
Letter O typed
Expired
Already used
Wrong category
Below minimum
Hits cap
Apply twice
Removed
Swapped
Quantity changed
App killed
Network drops
12s response
Timeout
Retry doubles
Tax after discount
Free delivery lost
Partial match
Per-unit
Rounding
Stacked codes
Returning user
Free account
Guest
Session expires
27 scenarios from one line. Illustrative example.

Trusted by mobile teams at

27
scenarios from one requirement line
2
human approval gates before a test runs
0
routes approved without a real device
1
sitting to review a scope, not a day

Six ways a feature goes wrong that a spec never mentions

An experienced QA engineer runs this expansion for every feature, every release. Mostly in their head.

What should happen

2
  • SAVE20 is accepted
  • Total drops by exactly 20%
The part most suites cover.
What could happen · 25 scenarios the spec never mentions

The code is wrong

6
  • Typed SAVE2O, with a letter O
  • Expired yesterday
  • Already used on another order
  • Valid for shoes; cart has groceries
  • Cart is $49.50; minimum is $50
  • 20% off $2,000, capped at $100

The user improvises

5
  • Taps Apply twice
  • Removes it. Is the old total back?
  • Swaps in a better code
  • Goes back, changes quantity
  • Kills the app mid-checkout

The network wobbles

4
  • Drops the moment Apply is tapped
  • Responds in 12 seconds
  • Times out
  • Retry applies it once, not twice

The maths shifts

6
  • Tax applied after the discount
  • Subtotal falls below free delivery
  • Three items; code covers one
  • Quantity 5, per-unit discount
  • Rounding on $19.99
  • Two codes stacked where rules forbid

The account differs

4
  • First-order code, returning user
  • Premium-only code, free account
  • Guest, no account at all
  • Session expires mid-checkout

Four of four passed. The bug shipped anyway.

Release 4.12 · Checkout suite
100% passed
✓
Login
✓
Add product
✓
Checkout
✓
Payment success
Production · two days later

Customers charged. Orders never created. Nothing failed in CI, because nothing was checking.

The fix was easy. Knowing to test for it was the hard part.

  • The bottleneck was never running tests. It’s deciding which ones matter.
  • Scope decisions happen once, quietly, and set the ceiling on everything downstream.
  • A suite can only be as good as the list it was built from.
  • That list is usually in one person’s head.

Drizz does the expansion. You make the call.

The same coupon feature, followed through all three steps.

Step 01

It reads your app first. Not a description of it.

Drizz walks the application on a device and records what is actually there: every screen, control and route, and every way into a journey. That’s why a later “Tap Apply” points at a button that exists.

Step 02

It proposes a scope. Every line has a reason.

Drizz turns the map into a prioritised plan and explains each entry with what it observed. Your QA lead reviews it the way they’d review a colleague’s: add what Drizz couldn’t know, drop what doesn’t apply.

Step 03

It writes tests that have to prove something.

Approved scenarios become runnable tests, in plain English or as code alongside your existing suite. Every test states what counts as a pass. A test that only clicks through a flow is rejected.

App map · Checkout journey
Observed on device
Product
→
Cart
→
Checkout
→
Payment
→
Confirmed
Product
→
Buy now
→
second entry point into Checkout
Coupon code
text field, free text
Apply
button, calls the server
Remove
appears only after apply
Total
updates in place, with delivery
Step 03 · What gets written
2 checks
Expired coupon is rejected
Route verified on device
  1. Tap “Checkout”
  2. Type “EXPIRED10” in “Coupon code”
  3. Tap “Apply”
  4. Validate “This coupon has expired” is visible
  5. Validate the total is unchanged
Coupon flow walkthrough
Not written · would always pass
  1. Tap “Checkout”
  2. Type “SAVE10” in “Coupon code”
  3. Tap “Apply”
  4. Tap “Place order”

Four steps, nothing checked. This passes on every run even when the coupon is applied at the wrong rate, so it would never tell you the discount broke.

AI-written tests read well and fail on the first run

Because they’re written from a description of the app. Here’s what changes when they’re written from the app itself.

Generated from a document
Source of truth
What the spec says the app does
“Tap Apply”
Assumes an Apply button exists
Navigation
Inferred from docs or source code
Done when
It’s written
Generated by Drizz
Source of truth
What the app actually does, read on a device
“Tap Apply”
Refers to one Drizz has seen, on a screen it can reach
Navigation
Can’t be approved until a real device has walked the route
Done when
It has run on a device and been checked

“Seven tests failed. None were bugs. Each had trusted a route read from source code.”

That was one of our own runs. We didn’t ask the system to be more careful. We made it impossible to approve a route no device had walked. The principle: never check work against the same source that produced it.

Nothing enters your suite because a model was confident

Two human gates, placed where judgement matters: is this the right set of things to test, and are these the right tests.

Drizz
Maps the app
Drizz
Proposes scope
Your team
Approves scope
Drizz
Writes tests
Your team
Approves cases
Drizz
Runs on devices
Proposed scope · Checkout → Apply coupon
In review · QA lead
Scenario
Priority
Why it’s in scope
✓
Valid code updates the total in place
High
Observed
Total recalculates without a reload
✓
Expired or used code: error shown, total unchanged
High
Observed
An error state renders beneath the field
✓
Subtotal drops below free delivery after discount
High
Observed
The delivery fee line changes with the subtotal
✓
Quantity changed in cart after applying
High
Route
Checkout → cart → checkout is reachable
✓
Checkout entered through Buy now
Medium
Route
A second entry point, different cart state
✓
Network drops mid-apply
Medium
Observed
Apply calls the server; no offline state found
+
Apply coupon on a low-end Android device
High
Added by QA lead
Most of our checkout traffic is on older Android
–
Guest user applies a coupon
—
Dropped
Guest checkout is disabled in this market
7 approved · 1 added · 1 dropped · Illustrative example
Approve scope

Questions to ask any AI test-generation tool

1

Does it read the app, or a description of the app?

Ask what the tool saw before it wrote the step. If the answer is a spec or the source, its navigation is inferred.

2

Can it show why each scenario is in scope?

A list without reasons can’t be reviewed, only accepted.

3

Where are the human gates?

Ask which decisions a person makes, and what happens if they change one.

4

Has a device walked the route?

A step that depends on a route no device has reached is a guess with good grammar.

5

What does the tool reject?

If it never rejects its own output, it has no standard. A walkthrough that asserts nothing shouldn’t pass.

6

How long does reviewing a scope take?

Measure it in a real session. Reading generated steps line by line isn’t a review.

7

What happens when the feature changes?

Ask whether the scope is rebuilt from scratch or updated against what actually changed.

What Drizz decides, and what stays with your team

What Drizz decides

  • Expands a requirement into the scenarios around it
  • Attaches a reason to each one, from what it observed on the device
  • Prioritises, and writes the approved set as runnable tests
  • Refuses to write a test that cannot fail

What stays with your team

  • Judgement about risk. Drizz doesn’t know most of your users are on older phones, or that one feature carries most of your revenue.
  • Exploratory testing. Finding the bug nobody thought to look for is still a human skill.
  • Catching everything. No testing approach does. The aim is fewer blind spots and a scope you can defend.

What connects to the scope decision next

The scope decision is separate from everything downstream, so each new input connects to it without rebuilding the rest. These are not available today.

In development

Specs, designs and PRs as inputs

Acceptance criteria, documented edge cases and designed states such as empty, error and offline, feeding scope before the build lands.

In development

Change-impact analysis

For a pull request: the shared components it touched, and the journeys that quietly depend on them.

In development

Regression selection

The subset of your existing suite relevant to a change, with the reason each test was included.

In development

Learning from past defects

Areas that have failed before carry more weight in the scope when related code changes.

Take something useful with you.

Benchmarks, calculators and checklists you can use whether or not you pick Drizz.

UI-TapBench benchmark

94.51% tap accuracy against frontier models. 570 screenshots, 20 apps, open dataset.

Download the report →

Appium cost calculator

What your current suite costs and what Vision AI saves. Four inputs, no signup.

Run the numbers →

Mobile app testing checklist

Pre-release checks across functional, UI, performance, security, device and network.

Get the checklist →

Mobile testing tools compared

11 tools compared on real-device coverage, CI fit and maintenance burden.

Compare tools →

Questions? We’ve got answers.

What is test planning in software testing?

Test planning is deciding what to test before tests are written: which scenarios matter for a feature, why each is in scope, and which can be left out. It sits between a requirement and a test suite, and it determines coverage more than execution does.

How do you decide what to test for a feature?

Start from what should happen, then expand into what could: wrong input, repeated actions, network failure, calculation edges, and different account states. A single requirement line commonly expands into twenty or more scenarios.

Why do bugs ship when all tests pass?

Because a passing suite only proves the tests that were written. If a scenario was never scoped, nothing checks it, and CI stays green while the defect reaches production.

Can AI decide what to test?

AI can propose a scope and explain its reasoning. It cannot weigh business risk, which is why Drizz puts the plan in front of your QA lead to approve, edit or reject before any test is written.

How is Drizz different from AI that generates test cases from a spec?

Tests generated from a document assume the app matches the document. Drizz walks the app on a real device first, so every step refers to a control it has seen on a screen it can reach.

What stops Drizz writing tests that always pass?

A test that only taps through a flow without asserting anything is rejected. Every approved test states what counts as a pass.

Who approves the test plan?

Your team, at two points: once on the proposed scope, and again on the written test cases. Nothing enters the suite automatically.

Does test planning replace exploratory testing?

No. Exploratory testing finds the problems nobody thought to look for. Test planning covers the scenarios that can be anticipated, so exploratory time isn’t spent on them.

Find out what could happen in your app.

In a first session, point Drizz at your app and review the journey map and proposed scope it produces.

Ship your next iOS build with the suite already green.

Upload an .ipa, describe a flow, watch it run on a real iPhone. Twenty minutes, no framework setup.