•
Drizz raises $2.7M in seed funding •
•
Featured on Forbes
•
Drizz raises $2.7M in seed funding •
•
Featured on Forbes

Nothing hinders the deployment of mobile applications like flaky tests.
Your mobile tests execute much slower than the traditional unit tests, so you end up waiting for a long time to get your build results.
Also, multiple iterations of the same test sometimes provide very different results. What makes this worse is that the flaky test doesn't even notify the user that a false failure has just occurred, thus being a total waste of time.
Finally, iOS and Android builds generate lots of unpredictable concerns with regard to asynchronous application states, several devices of varying types, diverse operating systems, different network conditions, and shared CI infrastructure. In such a case, a single test can pass 10 times and then fail on the 11th run for absolutely no reason.
The worst thing to do is let the failure be and try to run a build again to make the pipeline green. The frequency of the test suite passing or the number of builds you run after failing to make the build pass should have absolutely no relevance in this case.
Eventually, false alarms will hide legitimate bugs, cause build pipeline trust to erode and moreover, you’ll fail to notice the real regressions that occur in your codebase.
You need to trust the results of a failed test enough to look for a reason. Because flaky tests ultimately lead to bad engineering practices, developers start to ignore failures, and also raise the number of retries, whilst QA teams waste a lot of time checking and investigating false alarms. The only way to mitigate this is to enable trust in your test suite.
Flakiness has an identifiable cause. This guide breaks down the 4 core drivers of test flakiness on iOS and Android, namely: timing, locators, device variability, and environment drift. Along with this, we will also discuss solutions that can be used across different frameworks in the context of CI/CD.
Mobile automation has a lot of factors working against it that increase the chance of failing tests. These factors include the asynchronous nature of application state, the behavior of the native OS, various form factors with their own performance characteristics, the quality of the network, and the shared infrastructure of CI. The assumption that these factors will always behave the same is a mistake because if they do not, the test may fail intermittently.
The goal is not simply to make a failing run go green. It is to make the test deterministic. A test should produce the exact same result under the same exact conditions each and every time, given the same app build and the data used.
Flakiness is intermittent, non-deterministic behavior where the same test passes and fails under equivalent conditions without an intentional app or test-code change.
Automated mobile tests are subjected to test flakiness.
An example could be a login test that tries to find or tap a button before the login screen has finished rendering, is recorded as a failure, and later, is recorded as a success with no change in the test.
Flakiness should not be confused with all production issues that occur intermittently. A test which fluctuates between a failure and success could indicate a bug that is unique to a device and could be a race condition or an issue in the backend.
Flakiness is not an indicator that the failure is inconsequential. If a test case fails intermittently, that test case should be examined closely.
On iOS and Android, flakiness commonly comes from 4 areas:
The sections that follow describe how to locate and resolve these common underlying issues in a CI/CD pipeline.
An app may render a control that has not yet been prepared to receive input or start a transition before data, animations, and business logic have finished. An app may seem to work on a fast device when we add a fixed delay, but the delay may fail to work under heavy usage or prolong the execution of the app on a faster device.
Avoid fixed sleeps
Commands such as Thread.sleep() and time.sleep() make tests slower and less reliable. Setting a 3 second wait for a screen that loads in 1 second is a waste of time. Unfortunately, it can still fail when a slower runner needs 4 seconds.
Use conditional waits
It is a better practice to use a bounded, condition-based wait that is tied to an observable application state such as when a control is visible and enabled, an expected text appears, a loading indicator disappears, or a known network and UI state has reached.
Appium
new WebDriverWait(driver, Duration.ofSeconds(10))
.until(ExpectedConditions.visibilityOfElementLocated(By.id("login_button")));
Visibility isn't enough for a tap. Ensure the element is enabled, unobscured, and ready for the action you would like to perform on the element.
Espresso
Using Espresso (UI testing framework created by Google for native Android applications), you can run automated tests to simulate user actions such as clicking buttons, typing text, etc on a physical device or emulator. It waits for certain framework operations to complete before proceeding. For custom asynchronous operations, you need to register an IdlingResource or expose a reliable app state that the test can observe before taking action.
XCUITest
Using XCUITest (Apple's framework for automating UI tests for iOS applications), you can allow tests to interact with and verify UI elements such as buttons, text fields, screens, and other controls. The command waitForExistence(timeout:) confirms that an element exists in the UI hierarchy. However, it doesn't guarantee that the element is visible or tappable. You need to pair it with appropriate checks, such as exists, isHittable, and expected screen content.
Maestro and similar tools
While auto-waiting does work to eliminate the need to model the expected UI state, it is better to use an explicit wait based on a condition that must be met before a transition, API response or animation.
Control animations in test environments
If you're using Android devices exclusively for testing, it is best to reduce or disable system animations wherever required. For example:
adb shell settings put global window_animation_scale 0
However, you'll need to apply the equivalent configuration deliberately on iOS simulators and test devices. As a practice, always retain a small set of release checks with realistic animations when animation behaviour itself is a part of what you want to validate.
On r/Playwright, a developer who rewrote a flaky suite summarized it cleanly: "The real lesson: flakiness is usually a waiting problem, not a selector problem." In same thread, another commenter added a practical warning for SPAs: "waitForResponse is right call but SPAs will still bite you when response lands before component finishes rendering." The fix is combining network-level waits (API response received) with DOM-level checks (element is visible and interactive). And on topic of hardcoded timeouts, one reply put it bluntly: "You don't actually win until you go back and delete page.waitForTimeout calls someone added in a panic six months ago, because those are ones that mask real bugs."
A study analyzing 201 fixes across 51 Apache projects found that 45% of flaky test fixes addressed async timing issues. A 2026 benchmark report confirmed same range across mobile and web.
While selector-based tests work in some cases, it is important that you figure out stable methods of locating UI controls. A test becomes fragile if it is based on the hierarchy, details of the presentation, or text that changes often and is user facing. QA Wolf's analysis of production test suite failures found that DOM changes and brittle selectors account for about 28% of test failures. Not majority, but consistent.
Use stable identifiers
You can use accessibility identifiers or framework supported test IDs for controls that the test must be operable. Accessibility identifiers are the most reliable when they are deliberately assigned, unique, stable and are maintained during UI changes. The best practice is to give clear naming conventions to these identifiers and keep them stable across routine UI refactors.
But refrain from using accessibility labels as test-only IDs. This is because these labels may be improved for screen readers or localized for different regions, rendering them useless in most cases. Try to keep your user-facing accessibility labels separate from the stable accessibility identifiers used for testing.
Locator guidance
For React Native and cross-platform applications, you must confirm how test IDs map to native automation and ensure native driver behaviors are not assumed. Different frameworks can behave and respond to different attributes.
Make identifiers part of delivery
Slightly increase the scrutiny level: make sure your code changes a testable control, and within your pull request, your identifier and affected tests have also been updated. This would help elevate automation hooks to the level of a UI contract.
On r/Playwright, a commenter gave most practical advice for any flaky test: "Find out why test is failing. That will help you figure out solution." It sounds obvious, but most teams skip investigation and jump to retries. Retries mask root cause. On r/PracticalTesting, a post made case that "Flaky tests need owners, not just retries." Assigning ownership to specific team members forces investigation instead of suppression.
An action can be valid on one device but invalid on another. This is the reason why tests pass on one device and fail on another. These types of variations include screen size and density, OS permission flows, manufacturer customizations, and hardware capabilities. On top of that, rendering speed varies, especially on Android devices.
Some examples are as follows:
Not every failure that is device specific is a test issue. If an important action is not available on a supported screen that is smaller, then that can be a real product defect.
On r/QualityAssurance, a tester recommended pragmatic first step: "Disable specific tests until you have time to fix and stabilize them." Quarantining device-dependent flakes keeps your CI signal clean while you investigate.
On r/devops, another commenter pointed to test pyramid as structural fix: "The textbook solution is to have majority tests as unit test, maybe 20% of tests should be integration tests and lastly perhaps 5-10% system level tests." Fewer E2E tests means fewer opportunities for device-specific flakes to block your pipeline.
Tests often result in passing local runs and failing Continuous Integration (CI) runs. CI runners may have less CPU, memory, or system resources, be limited to older versions of dependencies, run a higher number of concurrent tests, or access more unpredictable or slower services.
Reproducible CI
On r/Playwright, a commenter described CI specific pattern: "This comes up all time and usually is a mix of an infrastructure (bottlenecks that get revealed when you increase workers in your CI environment) and poor test code (writing tests that are not parallel friendly or brittle)."
On r/devops, another pointed to shared state: "With 'flaky' tests you will likely have some of following - global state being used between tests that are not being accounted for correctly, such as a global logger, a global tracing provider, etc." Both are CI-specific problems that don't show up when running tests locally one at a time.
The 3 root causes (timing, devices, environments) exist in every testing framework. However, 1 category (selector fragility) is structural. It exists because selector-based frameworks couple your tests to your app's internal element identifiers.
Drizz's Vision AI removes that coupling.
How it works:
What this fixes:
What this doesn't fix:
The numbers: Teams using Vision AI report 90%+ reduction in flaky test failures. The reduction comes primarily from eliminating selector fragility and reducing timing flakes through adaptive waits. The remaining flakiness is device and environment related, which requires infrastructure fixes, not framework changes.
For teams where selector maintenance eats sprint time, removing selector layer is highest-leverage fix. For teams where timing or environment drift is dominant problem, fixes in sections 1 and 4 of this guide apply regardless of which framework you use.
The pattern across every Reddit thread on flaky tests is same: teams retry instead of investigating, patch instead of fixing, and add sleep commands instead of understanding wait. The four root causes listed above are what investigation should target.
1. Determine the cause of failure. Identify if failure can be reproduced by building the application, using the same test data, environment, and device configuration.
2. Log evidence relating to the failure, such as video, application logs, traces, and network data.
3. Assign the failure to an owner. The owner is responsible for the investigation, coordination of the fix, and stabilization of the test by deleting or quarantining it.
4. The failure should be addressed by removing sleeps, elaborating on the tests using state-based synchronization, improving locators, and addressing test data isolation, among other things.
5. Quarantine should be temporary. Quarantined tests need to include a ticket, owner, reason, and deadline. They cannot just go off to never return from release coverage.
6. Measure the trend. Separately track intermittent failures and confirmed product regressions, as well as retry rates and time-to-resolution.
One type of visual or AI-assisted testing is to help reduce the coupling of implementation-level selectors by locating an element based on its appearance, label, and context rather than on a selector. This is useful when the UI is frequently refactored and results in a costly case of maintaining locators.
This is not a general solution for mobile flakiness. Visual methods still require reliable synchronization, test data, a variety of supported devices, and assertions that distinguish the correct target from targets that look similar. They also require validation for localization, dynamic content and various layouts of the application.
When teams consider using Drizz's Vision AI testing, evaluate the trade-offs it would bring to your suite in terms of the time required to maintain locators and to resolve false failures, an increase in execution time, and potential missed regressions.Take the reported improvement numbers from the vendor as a frame of reference to check for yourself, do not take them at face value and run a controlled experiment instead.
The first step is to replace the fixed sleeps in the tests that consistently produce flaky results with bounded waits on UI or application states. Additionally, isolate shared test data, and ensure you record evidence for each test failure.
Another approach is that you may quarantine known flaky tests so they don't block deploys. Gather execution data over 1-2 sprints. Then fix or delete them. This stops bleeding while you address root causes.
As a part of continuous integration, relevant tests can be automatically run multiple times to overcome flaky infrastructure. However, this should not be the end of the investigation and remediation efforts. Document the failure and the result from the multiple attempts to automatically run the test in order to determine possible causes.
Yes. For most mobile testing frameworks, accessibility identifiers are easier and better options than XPath.
No. Real devices exhibit behavior from the OS and hardware that may not be exhibited by emulators. Real devices also have issues with the performance and management of the device. Use both emulators and real devices to form your testing strategy.
No. Visual testing helps reduce locator instability, however it does not completely eliminate flaky results due to other factors such as variable devices, state, or performance issues.
About 28%, based on QA Wolf's analysis of production test suite failures. The larger category is async timing at roughly 45%.
No. Use explicit, condition-based waits. Sleep commands waste time when app is fast and cause failures when app is slow.
Re-run it on same commit without code changes. Passes on retry = flaky. Fails consistently = real regression. Track both categories separately.
Quarantine known flaky tests so they don't block deploys. Gather execution data over 1-2 sprints. Then fix or delete them. This stops bleeding while you address root causes.
Yes. Drizz uses adaptive wait logic that detects screen state before executing next step, instead of static timers. It doesn't eliminate all timing issues, but it handles common cases without explicit wait commands.
Related Content:
Self-healing test automation | Why test automation fails | Self-healing test automation tools | Book a demo