•
Drizz raises $2.7M in seed funding •
•
Featured on Forbes
•
Drizz raises $2.7M in seed funding •
•
Featured on Forbes

It is not practical to test your application on thousands of mobile devices. The goal isn’t figuring out which devices you can test the most. The goal should be understanding device failure scenarios that would expose the greatest risk.
A Samsung device that makes up a large percentage of your user base might be more important than testing 10 other less common devices.
A version of Android that was released a while back might only make up a small percentage of your app traffic, but if it has issues with permissions, performance, or even compatibility, it could still be a big risk.
A device with very little memory could reveal crashes or even slow loading that other, more high-end devices would never experience. Screens with limited traffic might be important after you release a new feature or even make some changes to the page design.
By far your biggest mistake would be building a device matrix with just some of the more popular mobile devices and considering that your coverage is complete. You should not choose devices because of the global market share, just because your QA team has them, or just because they are the latest and greatest, or just because they are the only ones that your testing service has available. You should understand your user base and the areas that are the biggest risks to your business and understand the impact that a failure could have.
Device coverage is a technology resource allocation decision. You should be prioritizing what your user base needs the most and what poses the greatest threat to your product. Your matrix might end up only having 10 devices, or it could go as high as 20, or maybe even more.
It should encompass your users, application behavior, and operating systems. It should not be a copied list made from another company’s testing strategy. This paper outlines how to analyze user traffic, OS distribution, form factor, and bugs to create a mobile device matrix to determine a representational sample of devices to test within the validity of the business risk. This approach reduces the cost and time of test matrix maintenance while achieving meaningful coverage.
Fragmentation is a difficulty with mobile device testing. Android suffers from more than just a high number of device models. There are a plethora of modifications made by device manufacturers to the OS. Included in this modification are the background process rules, permission flows, various keyboards, browsers, and battery saving features. How a screen behaves consistently on a Google Pixel may be different from how a screen behaves on a Samsung Galaxy or Xiaomi Redmi device.
Look for differences in:
iOS suffers less from fragmentation, but there are still important considerations with iPhone versions, OS versions, screen sizes, refresh rates, and hardware capabilities. iPad adds another dimension with various layouts, orientations, and different multitasking states. The combinations explode. 10 devices with 3 versions of OS already yield 30 combinations with many different sizes, orientations, and foldable, hardware dependent features.
The answer is not maximum coverage. It is targeted coverage.
The way most teams conclude on a device matrix is through one of these three methods.
The traffic-weighted methodology begins with the devices your customers use. Extract device and OS data from Firebase Analytics, Mixpanel, Amplitude, app store reporting, crash services, or internal telemetry. Arrange device and OS combinations by either sessions or active users. The combinations with the most traffic are your first test matrix group.
Assume your analytics report the following:
A traffic weighted list would give you more overlap with your customer base than the top used devices list. This is a good method to use with stable products as it aligns the product and engineering teams with a strong rationale for device testing, In addition, it does not work well with the long tail, as there would be one percent devices with potential significant rendering and payment issues. In those cases, you should implement a fixed risk-based methodology.
A developer on r/androiddev shared: "I segment by Android version and OEM pretty aggressively now." That's traffic weighted selection in practice, even without formal name.
This testing methodology involves selecting devices based on market reports and/or regional data for devices with the highest market share. Although this methodology is employed primarily by teams with little customer data (i.e. new apps), it can also be used by products that plan to launch in a new country, region, or market.
A typical market-share matrix incorporates the following:
Global market share may not always align with your audience. An app specifically for the U.S. will require a different matrix than an app predominately used in Southeast Asia. Use regional data if available.
Risk-based testing focuses on device characteristics associated with failure-related issues observed in your app.
Prioritize devices that include:
Risk-based testing relies heavily on the maturity of your product and your bug data. If failure in Xiaomi’s low-memory devices is observed in a significant amount of your crash data, these devices warrant inclusion in the primary matrix even with less market share, compared to flagship devices.
For most teams, user traffic (or some proxy) defines the bulk of the matrix. Market share can fill in product data gaps where necessary. Adding risk-based changes includes devices that carry unique importance on a technical or business level.
A simple scoring model for your device prioritization works well:
Device Priority = user traffic + technical risk + business impact a
Score each category from 1 to 5 individually, then total the scores for each device, and place devices with high traffic and moderate technical risk in every significant test run. A device with low traffic and high payment failure rate should be placed in nightly or release testing. A device with low traffic and no historical issues should be placed in the long-tail rotation.
There is no correct answer, but most teams start with 10-20 device and OS combinations. Consider using testing tiers instead of running every test everywhere.
For every change in the code run a small group on each of those devices, which should be the highest traffic iOS and Android devices, an older supported OS, and at least one of the lower-memory Android devices. This tier should help provide feedback and catch issues before they make it to shared branches.
Next, add other manufacturers, increase screen size, add different versions of OS, add known problematic devices and testables, and tablets.
Totally fill in the risk matrix before each release. Include all of the devices that involve cameras, biometrics, GPS, NFC, push notifications, and any hardware that involves background services.
Long-tail devices should be rotated between release candidates in the test budget is constrained and all devices cannot be included.
They assist teams in validating the construction of fundamental navigation, form submission, API integration, element layout, and commonly recurring regression tests. While they do not fully recreate physical hardware, these should be avoided and real hardware be used for:
A good approach is to leverage emulators/simulators for more CI runs and real devices for high-risk regression and release testing. It is possible to buy real devices, utilize an internal device lab, or use real devices provided by a cloud testing service.
Cloud services provide access to a much wider variety of devices running different versions of different operating systems and reduce the cost of ownership of hardware. Before you make any decision, you should check their parallel execution, how fast you can get devices, which regions they serve, how integrated their offering is with your CI, and which devices they offer.
Android fragmentation:
iOS fragmentation (yes, it exists):
The math problem: if you test 10 Android devices x 3 OS versions x 5 screen sizes, that's 150 configurations. Add 5 iPhones x 2 OS versions, and you're at 160. At 5 minutes per test flow, running 50 test cases across 160 configurations takes 667 hours. That's a full time person for 4 months, running tests and nothing else.
You can't test everything. The question is what to test.
Developers on r/androiddev feel this daily. One noted that "OEM skin issues are absolute worst. One UI alone has caused me more bugs than any other single variable in Android development."
Another was more pragmatic: "Unless you're dealing with some hardware specific stuff, it will work fine on like 95% of devices, if you test on emulators." That remaining 5% is where device specific bugs live, and where real device testing earns its cost.
Drizz facilitates teams using analytics data in designing active testing strategies. Instead of teams having to select devices from a static list, it integrates usage data and prioritizes the device and OS combinations used by the greatest number of customers.
A good approach might run the top group, covering approximately 80 to 90 percent of sessions, on a continuous basis for each major CI build. The next group runs on a weekly or nightly basis, while even less used devices run on a rotational basis for each major release.
Drizz also employs Vision AI in order to identify elements rendered to the screen, rather than relying on platform specific selectors.
This solution allows one to test a flow across devices such as a Galaxy, Pixel, and iPhone without a separate selector for each device's screen hierarchy. This approach also helps reduce the amount of work necessary to write and maintain tests.
This solution, however, does not remove the need for selecting devices. Teams still need to add important hardware, low-memory devices, supported older OSs, and configurations that were used to reproduce bugs from the past. Manual exceptions also need to be added to address traffic-weighted selection.
Your app works with foldable devices, along with NFC, LiDAR, biometric authentication or similar product-related features that use hardware. These elements might need their own testing even when there is low interface demand.
This matrix becomes outdated if a team builds it and never looks at it again. To stay current, review the matrix at least once per quarter or if you:
Keep track of device specific failures as a separate entity from general application failures. Log the device model, OS version, manufacturer skin, screen size, app version, and the test case failure.
In the end, this data helps you eliminate unnecessary choices. For instance, a lot of users might be generating memory pressure failures as opposed to failures related to screen size, and a lot of old Samsung devices are producing app failures because of the old flow implementation for Android permissions.
Before you can collect fully detailed analytics, use a balanced group:
Replace with real user data as soon as possible.
The best device matrix is not the matrix with the most devices. It is the matrix related to your users, your product risks, and your release process.
Focus on the most used devices and cover configurations of previous failures, then add incremental tiers of new devices. This provides your team with a useful testing framework without a lot of cost or time on devices that are largely irrelevant to your product.
Try Android app testing, on the devices that actually break it with Drizz.
15-20 device/OS combinations cover 80-90% of most user bases. Start with your top 10, then expand based on crash data.
Yes. 24,000+ models, 7+ OS versions, and OEM skins that change rendering. A test passing on Pixel can fail on Samsung because One UI changes animation timing.
Emulators for development iteration. Real devices for regression and release testing, where rendering, memory pressure, and OEM behavior matter.
Drizz's free trial includes 300 Drizz Tokens (about 30 test runs) on real devices. For ongoing testing, Solo is $49/month for 5,000 tokens (about 500 runs), and the Team plan is custom priced.
It ranks device/OS combinations by session count from your analytics tool. Top tier (80-90% of traffic) runs on every CI build. Lower tiers run less frequently.
Yes. Foldable testing is complex because same app renders in folded, unfolded, and tabletop modes. Vision AI handles different screen states by reading whatever is rendered.