Quick Answer
Self-healing test automation detects when a test breaks due to UI changes and repairs it automatically — without manual intervention. On mobile, four architectures power self-healing: selector fallback (tries alternative locators), multi-locator fingerprinting (scores candidates against a composite profile), NLP-to-selector re-mapping (re-derives locators from natural language intent), and vision-based (identifies elements visually with no selectors at all). The right choice depends on how aggressively your mobile UI evolves.
Automated tests begin to fail when the buttons move about or when their identifiers change. If a self-healing tool unhelpfully directs a test to click an incorrect button and reports that the test passed, then you have a harder problem to spot.
Self-healing automated tests attempt to fix broken tests due to changes in the UI. The value of each solution is generally based on how well it identifies the desired element and how it confirms that the repair is valid. Just clicking another button is insufficient. The test has to carry out the required action and validate the output.
Repair mechanisms differ among solutions. Some tools attempt to use backup selectors, compare labels and IDs, or even use AI to comprehend the context of the test step and identify elements on the rendered UI. Repair mechanisms for self-healing automated tests handle different types of issues and introduce different risks. For example, using the nearby “Save draft” button to perform the action in place of the “Submit” button may change the value of the test.
The best way to assess a tool is to see what it does where it fails. When it refuses a repair, see what evidence it provides for it including the target, replacement, reason and checks. A test that passes is insufficient evidence of the correct repair.
For mobile teams, consider repairs that cross your Android and iOS applications including the Flutter or React Native interfaces as needed. Changes to the UI, label, or accessibility metadata should not re-direct a mobile test to a different action or hide a bug.
Self-healing automation test barely alters for a new UI modification.
In automated UI testing, self-healing helps a test find or interact with an element after the app changes, as long as the expected behavior stays the same.
Let’s say a checkout button goes through a new name. As long as the button remains and does the same function, healing finds the button, and the test can continue.
Most implementations of self-healing repair consist of detecting a failed step, identifying a new target, and executing the step. Some tools save updated references. Others resolve targets during later runs or require review before saving the changes.
Auto-waiting is different. While self-healing tests locate an element and perform the step, auto-waiting steps in to wait for an element to be available. For example, if a button takes a long time to load, an auto-waiting test would utilize synchronization to wait for the button to load, whereas a self-healing test would utilize a new identifier locator.
Self-healing has become the most overloaded term in test automation. A Selenium plugin that tries a backup XPath calls itself self-healing. An AI platform that re-identifies elements from screenshots calls itself self-healing. A low-code tool that re-records broken steps calls itself self-healing.
They are not doing the same thing.
The confusion costs mobile teams real money. If you pick a "self-healing" tool whose healing mechanism doesn't match your failure pattern, your tests still break, you just paid more for the privilege.
4 approaches to self-healing
Locating strategies often define self-healing. Products sometimes combine these methods, so it is often necessary to evaluate the feature and method of execution.
1. Selector Fallback
If the primary selector fails, alternative selectors are attempted. For example, if the resource ID no longer matches, the accessibility ID can still be used to identify the button. The Selector Fallback method is effective when at least one of the stored references is correct. The other references can be broad and ambiguous.
What it handles well: Minor changes to selectors. For example, when a CSS class is changed by a developer or when a primary element ID is changed, but another stable attribute remains on the element (i.e. data-testid). If there is stable text in the element that can be identified, the fallback would cover that as well.
What breaks it: A major redesign to the user interface where a broad set of attributes are changed. If the element is dynamic and no stable attribute exists, a fallback is futile. The test then fails with no recovery path.
Tools that use it: Katalon, Selenium with the Healenium plugin, Ranorex.
Mobile considerations: Compared to web applications, mobile applications offer fewer stable selectors. Element trees in some cross-platform frameworks are unstable and can change with each build. A strategy that works on a web application may not work on a native mobile application.
2. Multi-locator fingerprinting
Multi-locator fingerprinting involves profiling elements using a variety of signals such as text, attributes, surrounding elements, and hierarchical position, and combining it with position. On failure, it checks the replacement element against this profile.
This approach works best with element changes that affect text. While the absence of other signals may make element identification unambiguous, changes in a lot of flexibility in element text can make identification ambiguous. This ensures that repairs can happen even when other signals change. Moving a delivery address field to a different location still allows us to match against its label.
The highest score doesn’t always guarantee correctness. If the delivery address field is removed, a different address field in the same area may get the highest score. Analyze tool's rejection threshold and behavior in case of a tie among candidates.
It can handle: Moderate changes to UI, for example, moving an element to a different container and changing the element's class, but keeping text and general position similar. This approach provides the system with more signals to work with, as compared to single-locator fallback.
Breaks when: Complete redesign of page, element's text and position changes. The fingerprint degrades because all the signals it was trained on change. It is still DOM dependent and needs a valid and parseable element tree to score candidates against.
Tools that use it: BrowserStack (Low-Code Automation), Mabl, ACCELQ, TestGrid.
The mobile problem: Android and iOS expose different element trees. Flutter and React Native apps don't necessarily produce the same tree for widgets as their native counterparts. As a result, automation created using a fingerprint built from the element tree of one platform doesn't necessarily work the same for the other.
3. Natural-language interpretation
Use case: A test step says “Tap Log In,” and the tool has to interpret that intent against the current interface.
In an NLP-to-Selector implementation, the system figures out which element is to be selected based on the instruction and available element data. Thus, it is effective for situations where the intent remains the same, but the system underneath changes.
Plain-English authoring doesn't explain the mechanism. Depending on the solution/product, it might use a variety of methods to accomplish the same task (e.g. a combination of selectors and/or screenshots, etc.). Confirm the underlying method for custom controls that are not part of the exposed UI hierarchy.
What it handles well: UI modifications that change the actual mechanism of a button (i.e. btn-login becomes auth-submit), but not the semantics (i.e. button text remains Log In). In that case, the NLP layer is able to re-map to the new mechanism based on the text. Because tests are usually written at a high level of abstraction, the re-mapping provides more context than falling back to raw element locators.
What still breaks it: The execution layer relies on the use of selectors. Tools that provide NLP-to-selector functionality add an intelligent translation step; however, the final action still resolves to clicking on an element via its DOM coordinates. If the element is rendered as a custom canvas or the accessibility tree is incomplete, no selector strategy can locate it. If the native UI tree does not expose it, the healing does not occur.
Tools that use it: testRigor, Testsigma, Functionize.
The mobile problem: The selector dependency does not go away, it is just hidden behind a natural language layer. With mobile, where the accessibility tree is frequently incomplete and/or inconsistent, the NLP layer may interpret the intent correctly, but may still not be able to find a valid selector to execute against. For more detail and a point-by-point comparison of NLP-to-selector and NLP-to-vision architectures, check out the natural language test automation explainer.
4. Vision-based execution
To execute a test case, vision-based tools capture a screenshot and identify a target. Because of selector-free execution, an element ID rename does not create a locator that has to be fixed.
This is important. Using visual targets to resolve a priori on each run is resilience by design, before even starting a healing process.
Using Drizz, you can create mobile tests via English language commands and resolve targets visually on both Android and iOS. For teams creating platform locators, the biggest selling point is probably less selector work for mobile interfaces and reusable intent.
Execution still needs to be verified. There are repeating buttons, overlays, changed wording, and new flows, which increase ambiguity. Measure the inference time and test the instructions on both platforms.
What it handles well: This framework handles changing selectors, changing DOM structure, and platform-specific element tree differences well. It also adapts well to cross-platform framework migrations and redesigns where the visual intent is preserved. Because it never relies on selectors, it is immune to the entire class of failures that other approaches heal around.
What requires attention: There is a latency penalty with vision-based execution. Compared to direct selector-based execution, it takes longer. The system may need to disambiguate between visually similar elements, for example, two identical “Submit” buttons. Drizz handles this through understanding surrounding elements and test flow. Compared to the deterministic speed of selector-based tools, this approach offers a different kind of tradeoff.
Products that use it: Drizz (mobile full vision-based execution), Applitools (visual validation layer, not execution), Quash (vision-based execution, partial).
Mobile edge: Compared to the web, mobile is more visual by nature. Users interact with pixels, not DOM nodes. Cross-platform mobile DOM doesn't exist. Tools that work at the visual layer, work equivalently for all mobile platforms (iOS, Android, Flutter, React Native, Mobile Web) without platform specific adapters or element tree parsers. Learn more about how Vision AI powers mobile testing.
Why mobile needs a distinct evaluation
Android and iOS each have their own native UI and accessibility components. Some apps also include WebViews or frameworks that display browser-based DOM elements. Because of this, support of mobile browsers does not automatically support native app screens.
For Flutter and React Native apps, the level of support varies depending on how the app has been built, how the components of the app pass information to accessibility support, how well the automation tool works with the framework. Using frameworks does not remove the need for reliable selectors.
Think about accessibility throughout the development and testing processes. Visible elements are affected in different ways by onscreen keyboards, dialogues for permissions, scrollable lists, larger font sizes, and different languages. How users navigate your application and testing tools are affected as well.
Having a visual test pass does not mean the app meets accessibility standards. Let’s say a developer removes an accessibility label. Even if the screenshot test passes (because, from their perspective, nothing visually changes), however, change could break the experience for screen-reader users.
What self-healing can’t fix
Broken application behavior
Pro-actively failing a test when the outcome of a payment is rejected, the payment total is incorrect, or the order is missing, is necessary. Just finding the ‘Pay’ button does not fix any broken payment functionality.
For a checkout test, ensure the order, the payment, and the payment amount can be verified if the test environment allows it. Reaching a confirmation screen does not greatly assist the validation of the order.
Assertions that should have been made (but were not)
Healing does not create a definition of success. Failing to verify the data on a page that should have been opened is an example of a missing test.
The tool should not lower the expected balance from ₹500 to ₹400 if that balance was displayed.
Changed Requirements and Unstable Environments
A review is needed to update the test case for an approved redesign of the password login flow with an OTP flow. There are also critical issues related to infrastructure failures, expired credentials, and unavailable test data, which also need to be addressed.
A retry may restore execution in situations involving transient failures, but it does not eliminate the root cause.
Why Mobile Makes Self-Healing Harder
Most self-healing content focuses on web testing. Mobile introduces specific architectural challenges that make several healing approaches less effective:
No universal DOM on mobile. Web apps have a standardized Document Object Model. Mobile apps have platform-specific UI trees (Android's View hierarchy, iOS's accessibility tree) that differ in structure, naming, and completeness. Cross-platform frameworks (Flutter, React Native, Kotlin Multiplatform) add another layer of abstraction that generates different element representations per platform.
Accessibility tree gaps. On web, accessibility attributes are well-standardized (ARIA roles, labels). On mobile, accessibility IDs are optional, inconsistently implemented, and frequently change between app versions. Many healing strategies depend on these attributes as fallback signals.
Dynamic rendering. Mobile apps frequently use lazy loading, virtualized lists, animated transitions, and platform-specific rendering optimizations that make element identification timing-sensitive in ways web apps are not.
Smaller screen, denser UI. Mobile screens pack more interactive elements into less space. Fingerprinting and scoring approaches have more ambiguous candidates to distinguish between.
Platform fragmentation. The same app on a Pixel 8 running Android 15 and an iPhone 16 Pro running iOS 19 may render the same screen with different element hierarchies, different accessibility labels, and different timing characteristics. A healing strategy that works on one may fail on the other.
These factors explain why tools that heal effectively on web do not automatically heal effectively on mobile. When evaluating self-healing for mobile specifically, prioritize approaches that are architecturally independent of platform-specific element trees.
Comparison: 4 Self-Healing Approaches Side by Side
| Criteria |
Selector Fallback |
Multi-Locator Fingerprint |
NLP-to-Selector |
Vision-Based |
| How it heals |
Tries ranked backup locators (ID → XPath → text → CSS) |
Scores all candidates against a composite element profile |
Re-derives selectors from natural language test intent |
AI identifies elements from the rendered screen visually |
| Selector dependency |
Full |
Full |
Hidden |
None |
| Minor UI changes |
✅ Strong |
✅ Strong |
✅ Strong |
✅ Strong |
| Major redesigns |
❌ Breaks |
⚠️ Degrades |
⚠️ Degrades |
✅ Resilient |
| Mobile native fit |
Weak — few stable selectors |
Moderate — needs element tree |
Moderate — still selector-bound |
Strong — platform-independent |
| Cross-platform (Flutter, RN) |
❌ Inconsistent trees |
⚠️ Fingerprint may not transfer |
⚠️ Selector re-map per platform |
✅ Same visual layer everywhere |
| Healing speed |
Fast — instant fallback |
Fast — local scoring |
Moderate — NLP re-processing |
Moderate — AI inference per step |
| Example tools |
Katalon, Healenium, Ranorex |
BrowserStack, Mabl, ACCELQ, TestGrid |
testRigor, Testsigma, Functionize |
Drizz, Applitools (visual layer), Quash |
How to evaluate self-healing claims: 5 questions to ask
When a vendor says "self-healing," ask these questions before signing anything:
1. How does healing resolve the element? Selector fallback? Fingerprint scoring? NLP re-mapping? Vision? The answer tells you what failure modes the tool can and cannot handle.
2. What happens when healing fails? Does the test hard-fail? Does it flag for review? Does it retry with a different strategy? A tool that fails silently when healing cannot resolve an element is worse than a tool that fails loudly.
3. Does healing produce an audit trail? Can you see what was healed, what the old and new references were, and the confidence level of the match? Invisible healing can mask real bugs — a healed test that now clicks the wrong element passes incorrectly.
4. How does healing perform on mobile specifically? Ask for a demo on a native mobile app, not a responsive web page viewed on a mobile viewport. The DOM-dependent approaches that work on web may struggle on actual native mobile screens.
5. What is the latency cost? Selector fallback is nearly instant. Vision-based resolution adds AI inference time per step. Understand the tradeoff between resilience and speed for your CI/CD pipeline requirements.
How to assess self-healing capabilities of your application
Build a controlled set of changes
Pick some representative user flows, such as logging in, checking out, or editing the user's profile. Create passing baseline runs, then build and test with each change introduced one at a time.
| Controlled change |
Expected result |
| Button ID renamed |
Verify same outcome as before |
| Move field into different container |
Enter the data |
| Add second button |
Disambiguate target |
| Remove required step |
Ensure failure to act |
| Show incorrect price |
Fail price assertion |
| Show indicators to load slowly |
Watch for within the time |
Include failures deliberately. A tool’s refusal to repair a genuine defect matters as much as its recovery from a harmless UI change.
Correctness and recovery should be measured in parallel
These should be measured separately:
- Precision of repair: Successfully repaired defects divided by all reported repairs
- Detection of defects: deliberately introduced defects flagged or failed, divided by all introduced defects
- Effort to review: Time used to evaluate repairs and accept them
- Runtime overhead: execution time with healing compared with the baseline
For instance, if a tool claims to have repaired 20 instances successfully and a review shows 2 wrong target selections, then repair precision is 90%. Reporting only 20 repaired instances hides the errors. The numbers are for illustrative purposes and should not be taken as product benchmarks.
Perform the process multiple times on target devices and builds. Compare first sets of runs to subsequent runs and check if the product caches resolutions.
Take a look at evidence and control saved changes
Collect enough info to reconstruct each repair: the initial step, selected target, screenshots or changes of locators with the assertion.
Confidence scores should be treated as signals of confidence. Just because a score is high doesn't mean that a particular action fulfilled your business requirement.
Drizz's self healing steps show passed (healed) to indicate when a step was repaired and it differentiates it from an ordinary pass. This helps QA teams to easily find test runs that require further inspection to ensure that the fix is reliable and persists across future tests.
Drizz records what was changed and why, and provides a way to review or undo an incorrect repair. This is especially useful for destructive or financial actions because tests should fail safely when a target is ambiguous instead of allowing the automation to execute the wrong action.
Select the approach that most closely aligns with your failure history
Review recent failures and assess whether the repairs actually preserve what the test was intended to validate. The goal is not simply to make a test pass again, but to restore it with the least possible review and runtime effort. Acceptance criteria should reward repairs that preserve test intent, provide clear evidence of what was tested, and continue to reliably detect real application defects.
This is where Drizz helps teams move beyond just maintaining test suites. Using AI-driven repairs to preserve the intent of the tests, Drizz reduces the effort to review and maintain tests as the application changes. This allows teams to address the expanding of mobile test coverage and reduces the time spent addressing broken test automation.
Frequently Asked Questions
What is self-healing test automation?
Self-healing test automation means that the system is able to recognize and repair test breaks due to application changes. Instead of merely failing on a broken test step due to an element ID change or an application screen change, the system is able to locate the element using other means and repair the test step.
How does self-healing work on mobile apps?
On mobile, self-healing has unique challenges because native apps on Android and iOS have no standardized Document Object Model (DOM) like there is on the web. As a result, some self-healing frameworks use a variety of approaches from a selector fallback to a multi-locator fingerprinting to NLP (natural language processing) to re-map selectors to new locators, to vision-based techniques (i.e. identifying an element visually, without selectors). Because vision-based techniques are independent of platform-specific element trees, they are purposely built for self-healing on mobile.
What is the difference between self-healing and auto-waiting?
Auto-waiting (like in the Playwright and Maestro frameworks) retries a locator until the element is found, or a timeout is hit. This addresses timing issues (slow loading elements). Self-healing handles structural issues (elements change identity, or are in a different location). Auto-waiting is asking “is the element there yet?” Self-healing is asking “where did the element go?”
Can self-healing tests mask real bugs?
Yes, since self-healing tests can resolve to the wrong element. This is most problematic with approaches that do not have semantic understanding of the test step. Vision based and NLP based approaches reduce this risk, since they evaluate candidates against the intended element action, and not element attribute similarity. Audit logs are always good to check to ensure test resolution is not masking real bugs.
Which self-healing approach is best for mobile testing?
If the mobile app is a stable native app with well-defined attributes, fingerprinting of multiple elements should provide reliable self-healing with minimal time delay. If the UI changes often, the app is built on a cross-platform framework, or it has an incomplete accessibility tree, then a vision-based approach would be best.
Does self-healing work with Appium?
Appium does not have self-healing out of the box. There are ways of integrating self-healing to an Appium pipeline. Katalon, for example, uses Appium with locator fallback strategies. Some teams integrate Healenium with Selenium/Appium. A vision-based solution, like Drizz, replaces the Appium execution layer, and therefore eliminates the selector dependency that Appium-based healing has to work around.
How much time does self healing save?
Industry data shows QA teams spend 30-60% of their testing effort maintaining tests rather than writing new tests. The majority of test maintenance is locator repair. The self-healing concept is to let a QA engineer write a test once and then rarely have to modify that test to represent a change to the application. Teams using self-healing tools typically report a 50-80% reduction in the time spent maintaining tests, depending on the frequency of their application’s UI changes.
Is self-healing only useful for UI tests?
The primary focus of self-healing is for UI tests, as locator brittleness is predominantly a problem with the UI. The same concept can be applied to API tests (self-healing against a change of an endpoint or a change in the schema) and to database tests (self-healing against a rename of a column or a change in structure). The latter two are in the early stages of development. For mobile teams, self-healing of the UI provides the greatest impact, as mobile UI changes are the main source of test flakiness.