Your network can be healthy and your subscriber still unhappy.
Radio KPIs describe the pipe. Churn is decided inside the apps people open. Most operators can report accessibility and throughput to four decimal places, and still cannot say whether Instagram loaded in Cebu at 8pm — or whether a rival did it faster.
Speed is not experience, and the two diverge more every year
100 Mbps does not guarantee Netflix starts in two seconds. Throughput measures capacity, not whether a reel loaded. And almost nobody measures that: national QoS programmes report speed, coverage and availability, so no published dataset says who delivers the better Instagram. The argument is settled today by a speed-test badge — and the first operator with app-layer evidence gets to reframe it.
iOS is largely unmeasured
Most field tools need a rooted Android. iPhones are 40–55% of the high-ARPU base and route CDNs, ramp ABR and handle 4G→5G differently. Android-only is not incomplete — it is skewed.
Only a handful of apps get tested
Scripting a journey has meant a developer and a device farm, so teams test three apps and extrapolate — while subscribers judge you on twenty.
Indoor is where the traffic is
Most mobile data is consumed indoors, and passive DAS cannot measure the signal it delivers. A mall or a campus runs unobserved until a complaint arrives.
Four tools, four datasets, no common view
Drive test, app monitoring, benchmarking and probes from four vendors, each with its own KPI definitions. The numbers never reconcile, so cross-cutting questions go unanswered.
Three conversations that go badly without app-layer data
“ERAB drops are down 89%” means nothing outside the radio team. “YouTube starts 1.8 seconds faster on the sites we touched” means everything. AT&T México measured the same sites, same scripts, same days of week, before and after — and reported both layers side by side.
Internal test data cannot carry an external claim. A Southeast Asian operator took the other route: eight apps, five cities, 3,000+ sessions per city, both operators from the same kit at the same moment. The claim went to air and survived advertising-standards review.
At FIFA 2022, Ooredoo ran app-layer QoE live in their SOC across eight stadiums, iOS and Android, competitor SIM in the same kit. Every issue came with a specific cause — an overloaded CDN node, a cell under signal pressure, a BGP misroute — and each fix was confirmed at the next fixture.
Eight ways to collect. One engine. One portal.
Everything here runs one engine — 5GMARK — into one KPI schema. A crowdsourced hex, a walk-test session on the third floor of a mall and a fixed probe in an airport report throughput, latency and video start the same way. That is why one contract replaces four.
What are you actually trying to find out?
Pick the question closest to yours. Each one routes to the offering built for it — you do not have to read all eight.
App & Network Experience Monitoring
Measure the same sites with the same scripts before and after a change, and report the radio delta and the app delta side by side. The AI reporting agent produces the per-site, per-app, per-KPI comparison within a day — including the KPIs that went the wrong way.
Event & Venue Monitoring
Kits across stadiums, fan zones and transit hubs, running scripted app journeys on iOS and Android before, during and after each event, with a competitor SIM in the same box. Results stream into your SOC while the event is running, so fixes land between fixtures rather than after the tournament.
Device & iOS RF Benchmarking
Run the same journey on several handsets at the same cell in the same minute, and rank them on radio, throughput and app experience — including RSRP, RSRQ and SINR on unmodified iPhones. Operators use it three ways: device acceptance before a launch, lab validation that a tuned network behaves across the whole device range, and estate variance in the field.
Field Monitoring with 5GMARK Pro
A retail agent, a technician or a partner engineer installs an app on an ordinary handset and runs a full test protocol — radio, network, web, video and OTT app KPIs — with no rooting, no dongle and no bound device. The data lands processed in the portal, so nobody has to post-process log files.
Indoor DAS Monitoring with fixed probes
A passive probe on a wall or pillar, with up to three SIMs, measuring like a handset every minute — all operators, Wi-Fi and cellular, twenty-four hours a day. It is how you turn “a user complained” into “sector 7 degraded at 14:20 and recovered at 15:05”.
Indoor Walk Test
Walk the floor with a phone against a collection form you define, and get a per-site, per-zone picture of coverage and app experience — including on iOS. Repeat it after a change and the two walks compare directly, because the KPI definitions never move.
Crowdsource Data
We already run 5GMARK, a consumer measurement app with its own user base. In most markets that means a large body of geo-tagged, operator-identified QoE records has already been collected — across every operator, indoor and outdoor. It is the fastest thing here to switch on, and it shows you where to put everything else.
EdgeSDK & Quick Test
Drop the measurement engine into the app you already ship. A Quick Test button gives a subscriber a 90-second read of their own connection — and gives you a continuously refreshed national dataset from real handsets on real plans, collected under the consent terms you set.
Collection is plural. Everything after it is singular.
Eight agents — app-test kits, event kits, device labs, field phones, fixed probes, walk-test handsets, crowdsourced devices and your own app. Pick one or pick all eight.
All of them run the 5GMARK engine. Same test protocols, same KPI names, same units, same conditions for a valid sample.
Radio metrics are captured before, during and after every app journey — so a slow Instagram load can be traced to the cell, the handover or the CDN node that caused it.
The portal, the AI reporting agent and the raw export. One dataset for engineering, commercial and regulatory reporting — and it is yours, exported nightly.
Numbers that reconcile
- A crowdsourced hex and a probe reading can sit in the same chart without a caveat
- A walk test in March compares to a walk test in September, because nothing in the definition moved
- A regulatory submission and an internal engineering report draw on the same records
Where each agent is weak
- Crowdsourced data follows where people are, not a quota matrix — use it to find the problem, not to prove an SLA
- A fixed probe is precise at one address and blind everywhere else
- Packet capture is available on Android and on the probe, not on iOS
- Voice MOS needs additional hardware and sits on the roadmap
“ERAB drops are down 89%” means nothing. “YouTube starts 1.8s faster” means everything.
Scripted app journeys run alongside radio capture, so every app KPI carries the conditions it was measured in. Run it before and after a change on the same sites with the same scripts and you can state what the optimisation bought, per site and per app, in a report that lands within a day.
Radio context captured before, during and after every journey
Most tools measure the app or the radio. Measured separately, you get a slow timestamp and no explanation. Here the radio layer is sampled through the whole journey — so when Instagram takes four seconds we can say whether RSRP was degraded at that instant, whether a handover was in progress, and whether Android on the same kit did the same thing.
What the radio looked like before the app opened
Without a baseline, every anomaly is arguable.
The journey executes as a real user would
YouTube search then play. Netflix browse then stream. Instagram scroll then upload. WhatsApp call setup. Journeys, not synthetic file transfers — and radio metrics keep recording throughout.
What happened to the radio afterwards
A stall during a 4G→5G transition is a different problem from a stall on a stable cell.
Which server actually served the app
Decomposed per session and per CDN node, with raw PCAP export on Android and on the probe. A surprising number of “network” problems live here.
The comparison writes itself
The reporting agent correlates pre and post app KPIs with the changes applied between them, and produces the full comparison within 24 hours. No analyst assembles it.
AT&T México — an LTE optimisation, measured at the app layer
Five-day same-day-of-week windows. Same kits, same locations, same scripts. Samsung Galaxy S25, LTE Android.
The dashboards this offering ships with
Forty-plus apps ready, and a new one takes minutes
Scripts are generated by the AI scripting agent from a plain-language description of the journey — no developer, no device farm, no sprint.
31 shown.
Eighty thousand people arrive at once. The RAN dashboard stays green.
Stadium cells are dimensioned for a normal day. On a match day everyone streams and posts at once, and your brand is judged on whether a story uploads. Mozark puts scripted app measurement inside the venue — iOS and Android, competitor SIM in the same box — live in your operations centre while the event runs.
Kits where the crowd is, run entirely from a desk
A kit is a compact enclosure holding real handsets — typically two Android and two iOS — with power, connectivity and remote control. Once installed, nobody goes back to it. That matters: the zones generating the most complaints at a major event are often the hardest to walk into, and a security-restricted VIP area cannot be investigated by sending an engineer.
Stadiums and concourses
One or more kits per venue, latched to the serving cell, running through the event window and the hours either side.
Portable units
Carried through the crowd on match day, so measurement follows the density rather than where it was expected.
Fan zones, transit and airport
The experience is not confined to the bowl. Arrival halls, metro interchanges and fan parks carry their own load profile and their own reputational risk.
Pre, during and post — and a competitor in the same box
Illustrative results below are from the FIFA World Cup Qatar 2022 tournament summary across all stadiums and fan zones.
Share of samples scoring above 80 on the Application Quality Index — above 80 is a good or very good experience. Both operators measured from the same kit, at the same location, at the same moment.
| App | Ooredoo | Competitor | Gap |
|---|
Every finding came with a cause, and none of them needed new hardware
The most useful result of the tournament was not a score. It was four specific, actionable root causes that RAN dashboards could not have surfaced.
City-level averages looked healthy. Excess load times traced to a single cell under signal pressure at peak — invisible in any aggregated view.
High session variance flagged an inconsistency in ABR and CDN routing. Where routing was optimal, play start and home load beat the competitor.
The CDN was hosted internally, yet the primary node was overloaded: TLS handshake up around a third and time-to-first-byte up 44% against the secondary node. Every TikTok KPI was inflated by one node.
Same app, same network, three outcomes across three venues — one routed to a local CDN and optimal, one routed via a distant region with three times the TLS time, one on mixed anycast and unpredictable.
Packet-level analysis, per session
TLS handshake, TTFB and TCP setup decomposed per session and per CDN node, with the serving endpoint named. Raw PCAP export for every session.
The loop closed in hours
Every change applied between matches — scheduler tuning, CDN policy, handover thresholds — was confirmed in app KPIs at the next fixture. Action to evidence in hours, not weeks.
The comparison held up
Same device, same kit, same script, both SIMs at the same moment. That is what turns a competitive result into something marketing can use.
The references for this offering
“We have a strategic commitment to finding, establishing and developing partnerships with world-leading technology providers. Working with Mozark’s innovative solution enabled us to deliver a vastly enhanced fan experience at the world’s greatest sporting event.”
Radio metrics on an unmodified iPhone — and every handset ranked on the same cell.
Run the same journey on several handsets at the same cell in the same minute and the differences are the device, not the network. On Android that has always been possible. Doing it on a stock iPhone — RSRP, RSRQ and SINR during an app test, no jailbreak, no enterprise profile — is what finally brings 40–55% of the premium base into the dataset.
An Android-only programme is not incomplete — it is skewed
If iPhones behaved like the rest of the base, leaving them out would cost coverage but not accuracy. They do not. iOS handles 4G→5G differently, Netflix ramps ABR harder, and in several markets Instagram picks CDN endpoints by OS. Measured on both platforms at one location, the same network produces different conclusions — and the missing platform carries the highest ARPU.
The high-ARPU half is invisible
iPhones dominate the premium segment in most markets Mozark operates in. Android-only testing means the subscribers with the greatest revenue impact have never appeared in a QoE data point.
The platforms genuinely differ
CDN routing, ABR behaviour and network-transition handling all diverge by OS. Benchmarks built on one platform are systematically shifted, not merely narrower.
App KPIs without RF are timestamps
A slow load with no radio context can be logged but not diagnosed. That is the gap that made iOS app testing interesting but not actionable.
Same cell, same second, same script — so the handset is the only variable
Illustrative output from a live deployment. Click a column heading to sort.
| Device | App QoE score | DL throughput | Median RSRP | Video start | Rating |
|---|---|---|---|---|---|
| iPhone 15 | 94.7 | 721 Mbps | −88 dBm | 1.3 s | Excellent |
| Galaxy A14 5G | 91.2 | 48.3 Mbps | −89 dBm | 1.6 s | Excellent |
| Pixel 7 | 89.1 | 52.8 Mbps | −90 dBm | 1.7 s | Excellent |
| iPhone SE | 82.3 | 34.7 Mbps | −92 dBm | 2.1 s | Good |
| Redmi 12 | 74.6 | 28.4 Mbps | −94 dBm | 2.6 s | Good |
| OnePlus Nord | 67.8 | 18.2 Mbps | −96 dBm | 3.4 s | Average |
| Realme Narzo | 58.2 | 11.4 Mbps | −97 dBm | 4.8 s | Poor |
Where to point care and capex
A device family under-performing at the same cell as its peers is a device or firmware problem — not a site to spend money on. It is also a specific message for the care script and the retail proposition.
A neutral field comparison
Lab and field results diverge. This is your handset against its competitive set on live networks, radio context recorded, measured by a third party with a documented methodology.
The minimum operating threshold
The exact network conditions at which each app degrades, per device. That threshold turns a coverage target into a subscriber-experience target.
Three standing jobs, not one benchmarking exercise
The device comparison above is the output. These are the recurring programmes that consume it.
Device acceptance on every new release
Every handset an operator puts in its channel gets tested against the device it replaces and against the range it sits in. Same cell, same script, same minute — so a launch candidate that under-performs its predecessor on throughput, attach time or app experience is caught before it reaches a store shelf, not after the returns start.
Lab validation across the device estate
A tuned network is only tuned for the devices you checked. Run the same protocol across a representative device set in the lab and you can confirm a parameter change behaves the same on a flagship, a mid-tier Android and an older iPhone — or find the one chipset family it regressed. This is the repeat-purchase use case: it runs every optimisation cycle, not once a year.
Estate variance, continuously
Lab results tell you what a device can do; the crowdsource and field layers tell you what your installed base is actually getting. When one device family under-performs its peers at the same cells across a market, that is a firmware or device problem — and a very specific message for care, retail and the vendor conversation.
A tier-1 MENA operator switched iOS on, and found a different network
A mature QoE programme, reasonable confidence in it, every insight from Android. Adding iOS on the same kits at the same locations produced three findings the old programme could not have — one of which moved the operator’s position in a regional benchmark table.
Same kit, same moment: iOS Instagram resolved to a more distant CDN endpoint than Android — roughly three times the load time for a large share of premium subscribers. Nothing in the old programme could have surfaced it, because it had no iOS in it.
Visible only with simultaneous dual-OS testing and radio event logging. It changed how the operator sequenced its 5G handover work — the population most affected was the one it least wanted to affect.
Because radio context was captured during the call attempt, failures could be attributed to specific high-utilisation cells rather than to the app. That list went into the next optimisation cycle.
The same platform, read from the other side of the table
Field evidence for claims a lab cannot support
The measurement is device-centric by construction: a handset, at a cell, running a journey, with the radio recorded. No rooting, so devices under evaluation stay in shipping configuration — and iOS can be included in a comparative study, which is normally the hardest part.
Field testing that does not need a field engineer.
A retail agent, a technician or a partner installs an app on an ordinary handset and runs a full protocol — radio, network, web, video and OTT app journeys. No rooting, no dongle, no bound device. Results arrive in the portal already processed, so nobody spends a week on log files.
The reason field data is scarce is that field testing is hard to staff
Legacy tools need a rooted handset, a specific model, a laptop, a hardware-bound licence and someone who can read a scanner trace. That is expensive, so campaigns are quarterly and the data is always a little stale. Remove those constraints and a field programme becomes something else: continuous sampling by whoever is already out there.
Any modern handset, from the store or an MDM push. Nothing flashed, nothing rooted — the device stays under warranty and inside your compliance policy.
The tester logs in and picks the protocol you configured, or scans a QR at the site. The licence follows the login, so it moves between phones and people.
Press start and walk, drive or stand. Nothing to interpret on screen, nothing to record by hand.
Records upload, validate and appear on the map within minutes. There is no post-processing step.
Pick what a field test should contain
Each block is switched on or off centrally. The tester never chooses — they just press start.
Radio & cellular context
Captured continuously through the whole protocol, not once at the start — so any later result can be explained by the conditions it ran in. On iOS this is the 5GMARK Pro capability: RSRP, RSRQ and SINR on an unmodified iPhone.
Throughput, latency, jitter, loss
Configurable file sizes and durations against a Mozark test server or one inside your own network. Idle and loaded latency are both recorded, because the difference between them is where bufferbloat shows up.
Web browsing & page quality
Any URL you nominate — your own portal, a payment gateway, a government service, a news site. Google Lighthouse metrics are collected alongside raw load time, so a slow page can be attributed to the network or to the page itself.
Video streaming quality
HLS and MPEG-DASH sessions measured end to end, with ITU-T P.1203 video MOS. This is the KPI set most field tools do not carry, and it is the one that tracks closest to what subscribers complain about.
OTT app journeys
Real journeys inside real apps — open, search, scroll, play, upload — scripted by the AI scripting agent rather than by a developer. Forty-plus apps are ready on day one and a new one takes minutes, not a sprint.
Protocol & path testing
For when the question is where the problem sits rather than whether there is one. TWAMP isolates the last mile from the core and the transit; traceroute exposes the hop that changed.
What changes when the tool stops being specialist
Assessment reflects the publicly documented capabilities of each category, not a named product.
| Legacy drive-test stack | 5GMARK Pro | |
|---|---|---|
| Device requirement | Specific rooted Android models | Any off-the-shelf iOS or Android handset |
| Radio metrics on iOS | Not available | RSRP · RSRQ · SINR, unmodified device |
| Who can run a test | Trained RF engineer | Any staff member or partner |
| Licence model | Bound to a device or a dongle | Follows the login, moves between phones |
| OTT app journeys | A handful, developer-scripted | 40+ ready, new ones in minutes |
| Video MOS (ITU-T P.1203) | Separate tool | In the same protocol |
| Data processing | Export logs, post-process, build report | Processed on arrival, in the portal |
| Commercial model | CapEx-heavy, per rig | SaaS, per user |
Where non-specialist field testing already runs
A passive DAS cannot measure itself. This can.
DAS and small-cell systems inside malls, airports and enterprise buildings have no native way to know what signal they deliver. The EdgeX probe sits on a wall or pillar, receives like a handset, carries up to three SIMs and measures every minute — all operators, cellular and Wi-Fi, with no site visit and nothing transmitted.
Two tiers, one platform
EdgeX Pro — the in-building workhorse
Intel Elkhart Lake embedded SoC with 5G/4G across four M.2 slots and four micro-SIMs, Wi-Fi 7 with MLO, a 10G SFP+ port and integrated GNSS. Fanless, screen-less and battery-less — which is why it generates almost no heat even at full load and runs for months without intervention. Rated −20°C to +60°C at 95% relative humidity, so ceiling and pillar mounts in tropical venues are within spec.
EdgeX Lite — cost-efficient rollout
Raspberry Pi 5 based, fanless, PoE+ powered, with dual-band Wi-Fi 5 and cellular via a USB modem where it is needed. Where the requirement is Wi-Fi and fixed-line panels rather than multi-operator cellular, Lite delivers the same engine and the same KPI schema at a materially lower unit cost — which matters when the deployment is measured in hundreds of sites.
Rotation is what makes the comparison fair
The scheduler cycles the SIMs — ten seconds each by default, adjustable remotely — so all three operators are measured from the same antenna position within the same minute. Location, hardware and environment are held constant, which is the only way an operator comparison survives scrutiny.
Drive-test depth, delivered passively and continuously
Signal layer
- RSRP · RSRQ · RSSI · SINR
- Cell ID · TAC · LAC · MCC/MNC
- Technology and band in use
- Sensitivity down to about −100 dBm
Performance layer
- Download and upload, configurable size
- Latency, jitter, packet loss
- DNS lookup and TCP handshake
- TWAMP two-way active measurement
Where the problem sits
- Hop-by-hop traceroute latency
- Loss per node, route-change detection
- Domestic versus international split
- PCAP deep inspection, optional
What the user meets
- Any URL or API — reachability and response
- TTFB · FCP · LCP · Speed Index
- Payment portal and enterprise app checks
- Wi-Fi 6/7 alongside cellular, same device
One dashboard, every probe, every site
Screens from a live twelve-probe fixed deployment. Every module carries a raw-data export.
How many probes does a venue need?
Placement is driven by business need, not only by cell sectors. Wherever people transact digitally, something should be measuring.
Minimum one probe per sector, so every coverage zone has a baseline.
Payment counters, food courts, gates, checkouts, transport interchanges — anywhere an SLA is promised.
A venue of this size is a single-site deployment: one dashboard, one commissioning window, remote configuration from that point on.
Under two hours per site
Mount the back plate, attach the probe, connect PoE or DC, insert SIMs, power on. Mozark completes commissioning, onboarding and test configuration remotely. No RF engineer needs to be on site.
Zero site visits after that
Remote reboot, remote reconfiguration of every parameter, remote SIM-rotation changes, and staged OTA firmware updates that roll back automatically if they fail. Only a physical power cut needs a person.
Postpaid, strongly recommended
Prepaid SIMs interrupt service when they exhaust. Mozark can manage procurement and fold the data cost into the service fee transparently — the total is the same either way.
Where indoor probes already run
A probe is precise at one address and blind everywhere else
That is the point of it, not a criticism. A probe gives an unarguable continuous record where it sits, and nothing about the rest of the estate. Pair it with a walk campaign to decide where probes go, and crowdsourced data to see the market around them.
Walk the sites first
See the market around it
EdgeX Lite and Pro data sheets
Most of the traffic is indoors. Most of the measurement is not.
A drive-test route stops at the kerb. Walk testing takes the same engine inside — mall, terminal, hospital, campus — with every measurement tagged against a collection form you define centrally. So twenty people walking eighty sites produce one dataset rather than twenty spreadsheets.
Define the form, assign the sites, walk them, read the rollup
Set up the campaign: the collection form, the test protocol, the site list and who walks which site by when. Zones are usually the ones that matter commercially — entrances, payment counters, food courts, gates, wards, lecture halls.
The tester walks the site with an ordinary handset, confirming the form at each zone. GPS is unreliable indoors, so the form entry is what anchors the measurement to a place — and the app will not close a session with it incomplete.
Radio conditions, throughput, web, video and app KPIs are recorded continuously and stamped with the form values and the serving cell — so a weak zone can be traced to the antenna or the DAS sector feeding it.
Rank zones worst to best, roll up by venue, city and region, and diff the campaign against the previous pass. A DAS change that improved the third floor and hurt the basement shows up immediately.
The hard part is not the walk. It is twenty walks agreeing with each other.
One engineer in one mall produces good data by accident. Twenty people across eighty sites produce a spreadsheet nobody can use — one tagged the level “L2”, another “2F”, a third left it blank. So the form is configurable and mandatory: define the fields once, centrally, and the app will not close a session until they are filled.
The form is yours, not ours
- Free text, dropdown, numeric, date, photo or barcode fields
- Required or optional, with validation on the values
- Pre-filled from the assigned site, so the tester confirms rather than types
- Different forms per campaign — an acceptance test asks different questions from a complaint visit
What a venue walk usually captures
Every measurement in the session inherits these values, so the data arrives already grouped the way you want to read it.
Field engineers, partner contractors, retail staff and a benchmarking agency can all collect into the same campaign. Because the form is fixed, their results are directly comparable — and you can see who collected what.
Eighty sites walked over three weeks roll up by venue, by city, by region and by operator without anyone building a pivot table. The worst zones sort to the top on their own.
Re-run the same campaign after a DAS change and the two passes align field for field. A change that fixed the third floor and hurt the basement is visible immediately rather than argued about.
Sites are assigned to testers with a due date. The portal shows what has been walked, what is outstanding and which sessions failed validation and need redoing — before the campaign is closed.
A walk test answers a question. A probe answers it forever.
These two are designed to be bought together. A walk test is precise, cheap and repeatable — but it describes one hour on one day. A probe describes every hour of every day at one address. Most venues need the first to find out where to put the second.
The question has a date on it
- Accepting a new DAS or small-cell build
- Investigating a specific complaint or a specific zone
- Running a multi-site sweep with several teams to one form
- Checking a change landed the way it was meant to
- Surveying a venue you do not yet monitor
- Benchmarking a competitor’s in-building performance
The question never goes away
- An SLA has to be evidenced continuously
- Degradation needs to be caught before a complaint
- The zone is security-restricted and hard to walk
- Multi-operator comparison must run all day, every day
Your market and your rivals — before you deploy anything.
5GMARK has been in consumer hands since 2012, and every test it has run is geo-tagged, operator-identified and classified indoor or outdoor. That history becomes your competitive baseline immediately — your network and every rival, same handsets, same engine, nothing to install.
Someone else already did the fieldwork
Real users, real handsets, real plans — including indoors, where a drive-test route never goes. A 17-step model classifies each test indoor or outdoor at 95% accuracy, and anti-bias controls strip manipulation and device skew before reporting. A picture of the market not derived from your own counters.
Reach nothing else can match
- City, district and route-level benchmarking against every competitor, from day one
- Shows where to send a kit, a walk or a probe — so the expensive tools go to the right places
- An independent second source when a rival, a regulator or a journalist disputes a claim
- Wi-Fi tests reveal home broadband performance nationwide, not only at panel addresses
And we will say so in the room
- Sampling follows population density, not a quota matrix
- Plan and device are inferred from the handset, not verified from a bill
- It will not carry an SLA claim on its own — pair it with a probe or a kit for that
- Use it to find the problem; use a kit, a walk or a probe to prove and fix it
The data, in the portal, on day one
Screens from live national deployments. Every view has a Download Data button — the aggregates are not the deliverable, the records are.
You get the records, not just the pictures
Every screen above sits on top of a record set you can pull. Nightly export of raw records to storage you control, plus API access and CSV or shapefile download from the portal itself — timestamped, geo-tagged, operator-identified, with the confidence score attached.
Straight into your BI or planning tool
The 5GMARK API serves the same records to whatever you already use — a data warehouse, a planning tool, a BI dashboard. Nothing here is a black box that only renders inside our portal.
Yours from day one
Mozark holds no residual rights over your tenancy’s records, and there is no aggregated-only licensing tier. The raw data is yours from the first export, not at contract end.
Two minutes on the crowdsource layer
How the collection works, what the anti-bias controls remove, and how the same records roll up from street level to a national league table.
Who runs this at scale
Crowdsourced data finds the problem. Something else proves it.
The honest use of this is as a scout: which cities, districts and competitors deserve attention, faster and cheaper than anything else. A number that carries an SLA or an advertising claim comes from a kit, a walk or a probe under controlled conditions — and because the KPI definitions are shared, the two readings sit in one chart without a caveat.
Send a field campaign
Walk it, then monitor it
Put a probe there
Put the measurement engine inside the app you already ship.
Your self-care app is already on millions of handsets. EdgeSDK turns it into a measurement fleet: a Quick Test button gives the subscriber a 90-second read of their own connection, and gives you a national dataset from real devices on real plans — your brand, your consent terms.
Three places the engine can live
A Quick Test screen subscribers run themselves, plus optional background sampling. The care agent sees the last test when the customer calls.
Rather not touch the main app? A separate one under your brand on both stores. Same engine, same portal, faster to launch.
The engine compiles into CPE firmware. Always on, wired, no user action — the closest thing to a probe at every address you already own.
An OEM, content partner or enterprise customer can carry it too. Records land in your tenancy with the source tagged.
Ninety seconds, and the subscriber understands the answer
The subscriber taps one button
No settings, no server address, no choice of test file size. The screen shows a dial and a running commentary in plain language. Everything technical happens underneath.
Radio conditions and raw network performance
The radio context is captured first, so every later result can be explained by the conditions it ran in. On Android this includes packet-level capture where you enable it; on iOS the radio metrics come through 5GMARK Pro without any device modification.
Web page load and video start
This is the part the subscriber recognises. A page that renders and a video that starts are the two things they judge you on, and both are measured against the same definitions used by every other agent on this platform.
One score, and a reason
The subscriber sees a number and a sentence — “video start is slow, coverage indoors is the likely cause”. Your care team sees the same record with every underlying KPI attached.
Signed, encrypted, uploaded
Records are validated and confidence-scored on receipt, then land in the same portal as every other agent. Raw export to your own storage runs nightly — the data is yours from day one, not at contract end.
A national dataset that refreshes itself
The last test, on the agent’s screen
When a subscriber calls, the agent opens the most recent Quick Test rather than asking them to describe the problem. Repeat calls about the same address group together.
Complaint clusters become work orders
Where several subscribers in one grid square report the same symptom, that is a site to look at — evidenced, not anecdotal.
Your own numbers, refreshed daily
A dataset you control, from your own subscribers, on your own plans — rather than a quarterly report bought from a benchmarking house.
It runs quietly under the consent terms you set
Network performance metrics only — no browsing history, no files, no credentials, no message content. Sampling frequency, background behaviour and consent wording are yours to configure, and a subscriber can withdraw at any time. SOC 2 certified, GDPR compliant, ITU-T SG12 associate; raw records export nightly to your storage.
Eight collection agents. One set of definitions.
The least visible part of the platform, and the one that decides whether any of it is useful. Same test protocol, same units, same conditions for a valid sample — whether the record came from a crowdsourced handset, a field walk or a probe. Without that you have four datasets. With it, one.
The reason four tools never reconcile
Vendors measure throughput over different durations, against different servers, with different definitions of a valid sample. Latency loaded or idle. Video start with or without ads. None of those choices is wrong alone — but four of them in one organisation produce four numbers for the same network. One engine removes the argument rather than adjudicating it.
Agents become interchangeable
Start with crowdsourced data, add walk tests where it looks weak, put a probe where it stays weak. The series continues rather than restarting.
Time series survive
A measurement from two years ago compares to one from today, because the definition did not move underneath it.
One dataset, three audiences
Engineering, commercial and regulatory reporting draw on the same records — so the numbers in a board pack and in a regulatory return cannot disagree.
Which agent collects what
Where a cell is empty it is a platform limitation we state openly, not a gap we hide. Nothing here is redefined per agent.
| KPI family | App & event kits | Device lab | Field handset | Fixed probe | Crowdsource & SDK |
|---|---|---|---|---|---|
| Throughput & consistency | ✓ | ✓ | ✓ | ✓ | ✓ |
| Latency, jitter & loss | ✓ | ✓ | ✓ | ✓ | ✓ |
| Latency under load (bufferbloat) | ✓ | ✓ | ✓ | ✓ | ~ |
| DNS & TCP connect | ✓ | ✓ | ✓ | ✓ | ✓ |
| Web browsing / Core Web Vitals | ✓ | ✓ | ✓ | ✓ | ✓ |
| Video streaming QoE | ✓ | ✓ | ✓ | ✓ | ✓ |
| OTT video MOS (ITU-T P.1203) | ✓ | ✓ | ✓ | ✓ | ~ |
| App journeys, 40+ apps | ✓ | ✓ | ✓ | ~ URL/API only | ✗ |
| RF / cellular radio — Android | ✓ | ✓ | ✓ | ✓ | ✓ |
| RF / cellular radio — iOS | ✓ | ✓ | ✓ | ✓ | ~ |
| Wi-Fi link metrics | ✓ | ✓ | ✓ | ✓ | ✓ |
| TWAMP segment isolation | ~ | ~ | ~ | ✓ | ✗ |
| Traceroute & path analysis | ✓ | ✓ | ✓ | ✓ | ~ |
| CDN endpoint & TLS decomposition | ✓ | ✓ | ✓ | ✓ | ✗ |
| Packet capture (PCAP) | ✓ Android | ✓ Android | ~ Android | ✓ | ✗ |
| Voice & call quality | ~ | ~ | ~ | ~ | ~ |
| Availability & uptime | ✓ | ~ | ✗ | ✓ | ✗ |
✓ available today · ~ partial, conditional or on roadmap · ✗ not possible on that agent
Every KPI family, filterable
Filter by which agent collects it and which layer it sits in. Platform support is stated honestly, including where something is not available.
15 families shown.
Netflix · YouTube · Instagram · TikTok · Snapchat
Rather than screenshots, open the product.
One login, one master filter, and modules different people use for different reasons. The radio team lives in one place, the commercial team in another, both on the same records. Where a question has no module, the reporting agent answers it and builds the chart.
Pick a role, then click through it the way that person would
Every module carries the question it exists to answer. Eight modules shown; the screens are from live deployments.
Click the screen to enlarge it. Arrow keys move between modules.
Three places the platform does the work instead of you
App scripting agent
Describe an app journey in plain language and the agent generates the test script — element selectors, actions and assertions — with you approving each step. TikTok, Instagram, a custom fintech app or your own self-care app are onboarded the same way, in minutes rather than a development sprint.
Reporting agent
Ask natural-language questions across the QoE dataset and get summaries, anomaly highlights and trend analysis with the charts attached. It is also what produces the automatic before-and-after comparison after an optimisation window closes.
Real-time device monitoring
Live health status of every test device, kit and probe — so you know that data collection has stopped before the gap shows up as a suspiciously good week in a report.
Operators, venues and regulators — in thirty-plus countries.
Deliberately mixed. Regulator work is here because it is the hardest audience to satisfy on methodology; venue and enterprise work because the indoor problem is the same whoever owns the building.
Who we work with
Operators, venues, transport, retail and regulators. Named where the customer allows it.
Ooredoo · AT&T · Globe · Bouygues · BTC
Event monitoring at two World Cups, optimisation validation, app benchmarking, fixed broadband and an embedded SDK — plus two tier-1 programmes under NDA in MENA and Southeast Asia.
Marina Bay Sands · SNCF · Carrefour
80 fixed probes across a hotel and casino complex, a decade of rail connectivity measurement, and 3,000+ retail sites self-monitoring.
FCC · TRAI · ARCEP · TDRA · IBPT
Mozark built the FCC BDC Speed Test App and TRAI’s MySpeed, MyCall and DND apps, and supplies over 80% of the data on France’s public QoE portal.
European Commission · Nokia · a top-tier OEM
Crowdsourced measurement across 27 member states, and raw crowdsource data supplied to a leading device manufacturer under a global frame agreement.
Certified, patented and audited
The methodology has already been reviewed by the hardest audiences there are.
Deployments, grouped by what you would buy.
Pick an offering and see who runs it and what it produced. Numbers are as measured — including the ones that went the wrong way.
Twenty-one programmes across eighteen organisations
MÉXICO
Optimisation proved in app KPIs, not radio KPIs
The problem
28,373 ERAB drops in the baseline window, and no way to translate that into subscriber impact for commercial stakeholders.
What we did
Kits latched to serving cells across CDMX. Five-day pre and post windows, same scripts, same days of week. Sites ranked by app-impact score, not signal.
Result
X latency 127.5 → 49 ms. Instagram home load −36.6%. YouTube rebuffering eliminated. PRB utilisation 31.3% → 16.9%. 19 of 21 sites improved.
OPERATOR
App benchmarking turned into an advertising claim that survived review
The problem
A rival held the speed-test award. Internal test data cannot carry an external claim.
What we did
Eight apps that drive loyalty in that market, both operators from the same kit, four weeks, methodology documented for external review.
Result
WhatsApp call completion 93.8% vs 81.5%. YouTube play start 1.6s vs 2.3s. Claims went to air and survived advertising-standards review unmodified.
Globe Telecom & DICT — app failures traced to network root causes
The problem
App complaints and network counters could not be connected, for either the operator or the department.
What we did
App testing kits benchmarking 40+ applications across competing operators in the Philippine market.
Result
Per-app, per-operator comparison with root causes attributable to specific cells and CDN paths.
USA
Host-city measurement run on 5GMARK Pro as a field test app
The problem
Event-scale measurement across US host venues, at a cadence and a headcount no specialist drive-test rig could staff.
What we did
5GMARK Pro deployed as a field test app on off-the-shelf handsets — radio, network, web, video and OTT app KPIs from a single protocol, run by field teams rather than RF engineers.
Result
Event measurement without an event-sized specialist team, on the same KPI schema as every other Mozark agent.
Live app QoE in the SOC across 56 matches
The problem
80,000 concurrent users against cells dimensioned for a normal day. VIP zones were security-restricted — no engineer could investigate. Android-only monitoring would have missed half the premium base.
What we did
30 kits and 60 handsets across 17 locations in 8 stadiums, plus 5 hotspot locations — fully remote. Dual-SIM, dual-OS from the same box. Six apps, pre / peak / recovery windows.
Result
Four root causes found and fixed between fixtures: an overloaded CDN node, a cell under signal pressure, a BGP/anycast misroute, an ABR routing inconsistency.
“Working with Mozark’s innovative solution enabled us to deliver a vastly enhanced fan experience at the world’s greatest sporting event.”
TIER-1
Radio metrics during app tests on unmodified iPhones
The problem
A mature QoE programme in which every insight came from Android. The highest-ARPU, highest-churn-risk subscribers had never appeared in a data point.
What we did
Two iOS and two Android per kit, same journeys, same cell, same minute — with RSRP, RSRQ and SINR captured on stock iPhones, no jailbreak, no enterprise profile.
Result
iOS Instagram resolved to a more distant CDN node — 3× the load time, invisible before. A 0.8–1.2s app-continuity gap on 4G→5G that Android handled better. WhatsApp failures correlated to specific high-utilisation cells.
TDRA and ARCEP — national handset benchmarking on both platforms
The problem
Spectrum and QoS policy needed device-representative evidence, not Android-only samples.
What we did
Smartphone benchmarking across iOS and Android at national scale, measuring voice and data performance.
Result
A methodology audited by two regulators — which is why the same comparison holds up in a commercial argument.
PROGRAMMES
Device acceptance before launch, and lab validation after a network change
The problem
A launch candidate is signed off on lab radio figures, and a tuned network is validated on whichever handsets were to hand. Neither says how the device range behaves on the live network.
How it is used
Acceptance: the candidate against the device it replaces and its price band, same cell, same minute. Validation: the same protocol across a representative device set before and after a parameter change.
What it catches
A handset that under-performs its predecessor before it reaches a shelf, and the one chipset family a change regressed while every other device improved.
USA
5GMARK Pro as the field test app for FIFA World Cup 2026
The problem
Field coverage at event scale, with more sites than a specialist team could reach.
What we did
The full test protocol on ordinary handsets, assigned centrally and run by field staff.
Result
Event-scale field measurement without an event-sized specialist team.
3,000+ stores self-monitoring, no specialist on site
The problem
Network quality where customers transact, across a store estate no field team could cover.
What we did
Self-monitoring kits deployed across hypermarkets and retail locations, run by store staff rather than RF engineers.
Result
Continuous coverage of the estate. Similar deployments now run with La Poste and Geoptis across the French postal network.
Globe Telecom & DICT — operator benchmarking in the field
The problem
Benchmarking competing operators needed field coverage that specialist rigs could not sustain.
What we did
App testing kits and off-the-shelf handsets running the same protocol across the market.
Result
Comparable results from non-specialist collection, for both an operator and a government department.
TDRA and ARCEP — field campaigns at national scale
The problem
National QoS reporting required field measurement no single team could staff continuously.
What we did
Hybrid crowdsource and field campaigns since 2015, on standard consumer handsets.
Result
More than 80% of the data on France’s public QoE portal comes from Mozark collection.
Marina Bay Sands — 80 fixed probes, two operators, one dashboard
The problem
A passive in-building system across a hotel and casino complex with no way to measure what it delivered.
What we did
An 80-location fixed-probe network monitoring StarHub and Singtel simultaneously from the same antenna positions.
Result
24/7 KPI dashboards and SLA validation for the venue operator, with per-operator comparison at the same moment.
SNCF — in-train and trackside connectivity
The problem
Connectivity along a moving high-speed corridor that no fixed survey describes.
What we did
Fixed probes across rail infrastructure measuring in-train and trackside performance, alongside national QoS campaigns.
Result
A decade of continuous measurement feeding both the operator and the regulator.
Bouygues Telecom — residential and indoor coverage at scale
The problem
Residential QoS and indoor coverage evidence for SLA reporting and planning.
What we did
Probe-based fixed broadband and indoor monitoring at operator-grade data confidence.
Result
SLA reporting and network planning running on measured data rather than modelled coverage.
PT
Commercial venues across three regulatory environments
The problem
Venue operators and enterprise clients needed hotspot performance evidence across borders.
What we did
Fixed probe deployments with centralised dashboards, commissioned and configured remotely.
Result
One dashboard across three markets — no separate login or contract per site.
Crowdsourced data to identify served and underserved areas across the EU
The problem
Comparing connectivity across 27 member states, each with its own national methodology.
What we did
Crowdsourced measurement on one engine and one KPI schema across all 27 countries, classifying served and underserved areas from real-device records.
Result
Performance genuinely comparable between countries — because the same test, the same definitions and the same validity rules ran everywhere.
OEM
Crowdsourced raw data supplied under a global frame agreement
The problem
A device maker needed field connectivity data across markets, in a form its own teams could analyse.
What we did
Crowdsourced records delivered as raw data under a global frame agreement, rather than as a licensed dashboard.
Result
An ongoing global supply relationship with one of the largest handset manufacturers.
ARCEP — crowdsourced data underwriting a public national portal
The problem
Four operators, four methodologies, no comparable dataset ARCEP could publish and defend.
What we did
Continuous QoE collection on all four operators, iOS and Android, with anti-bias controls before reporting.
Result
More than 80% of the public monreseaumobile.com data comes from Mozark collection.
BOTSWANA
The measurement SDK running inside the operator’s own app
The problem
Reaching subscribers with measurement without shipping another app or another device.
What we did
EdgeSDK integrated into BTC’s existing customer app — their brand, their consent terms, our engine.
Result
Subscriber-side measurement from real handsets on real plans, feeding the same portal as every other agent.
TRAI MySpeed, MyCall and MyDND — one SDK behind three apps
The problem
Three consumer apps and a web tester had to produce consistent, reproducible results at national scale.
What we did
A single ITU-T compliant measurement SDK behind all of them, with device integrity checks, geo-verification and anomaly detection.
Result
Millions of tests monthly with no third-party data dependency, plus India’s first browser-based national speed test.
FCC — the SDK inside a federally accredited consumer app
The problem
Every commercial measurement tool failed federal security accreditation.
What we did
The measurement engine inside a FedRAMP-aligned consumer app, with remote configuration of servers, protocols and parameters — no store resubmission.
Result
The first federally compliant consumer broadband app at national scale, with a closed complaint-to-resolution audit loop.
What is different is not one feature. It is that these were separate purchases.
Drive test from one vendor, app monitoring from another, benchmarking from a third, probes from a fourth — four KPI definitions, four portals, four contracts. The change here is architectural rather than clever: one engine, one schema, one portal, and no capital case for every new question.
Six things few suppliers do together
Radio metrics on unmodified iOS
RSRP, RSRQ and SINR captured during app tests on stock iPhones — no jailbreak, no enterprise profile, no device modification. This is unusual, and it is what brings 40–55% of the premium base into the dataset.
Off-the-shelf handsets
No rooting, no specific model, no dongle, no licence bound to hardware. That is what lets non-specialists run field tests, and it is why coverage can be continuous rather than quarterly.
App journeys at breadth
Forty-plus apps ready on day one and a new one scripted in minutes by the AI agent — including your own self-care app, your own streaming service and a partner’s app.
Crowdsource and probe on one engine
The same measurement code runs in a consumer app, in your app via the SDK, on a field handset, in an event kit and in fixed hardware. Few suppliers cover both ends of that range, and fewer still with one schema.
Remote by default
Commissioning, configuration, SIM rotation, firmware and reboots are all remote. Only a physical power cut needs a person on site — which is what makes restricted zones measurable at all.
You own the raw data
Nightly export of every raw record to storage you control, with timestamps, agent IDs and confidence scores. No aggregated-only black box, and no residual rights held back until contract end.
The three routes an operator can take
Assessment reflects the publicly documented capabilities of each category, not a named vendor.
| Capability | Legacy drive-test platforms | Crowdsource-only vendors | Mozark |
|---|---|---|---|
| Radio metrics on iOS | ✗ | ✗ | ✓ |
| Off-the-shelf, unrooted devices | ✗ | ✓ | ✓ |
| App journey breadth (40+ apps) | ~ | ✗ | ✓ |
| Your own telco app monitored | ✗ | ✗ | ✓ |
| OTT video MOS (ITU-T P.1203) | ~ | ~ | ✓ |
| Indoor DAS / small-cell monitoring | ✗ | ✗ | ✓ |
| Three operators from one fixed device | ✗ | ✗ | ✓ |
| 24/7 fixed monitoring on the same platform | ~ | ✗ | ✓ |
| Event and venue deployment, fully remote | ~ | ✗ | ✓ |
| Fully autonomous, no field expert per test | ✗ | ✓ | ✓ |
| AI script generation for new apps | ✗ | ✗ | ✓ |
| AI reporting and self-service dashboards | ~ | ~ | ✓ |
| SaaS commercial model, no CapEx | ✗ | ✓ | ✓ |
| Raw data ownership, nightly export | ~ | ✗ | ✓ |
| On-premise hosting where policy requires it | ✓ | ✗ | ✓ |
What is coming, stated separately from what is live
Everything described elsewhere on this page is in production today. These four are not, and we would rather say so here than in a footnote.
Voice speech quality
MOS scoring for voice calls — objective speech quality measurement across operators. Requires additional hardware alongside the handset.
Indoor floor mapping
Automated floor-plan overlays for indoor testing, extending offering 06 from a ranked zone list to a rendered heat map on your own plan.
Advanced AI insights
Trend forecasting and automated root-cause analysis in natural language, beyond the current summarisation and anomaly highlighting.
5G NR KPI expansion
Extended NR metrics including NR-RSRP, SSB beam tracking and clearer SA/NSA differentiation.
Specifications, without a document trail.
The hardware sheets and the API summary are embedded inside this document — no server, no login, no link that expires. The KPI framework itself is browsable in the page rather than attached as a PDF.
The KPI framework is in the page, not a PDF
One definition per KPI, applied to every collection agent — browsable and filterable rather than attached, so it can never be a stale document in someone’s inbox.
What is measured
Radio, Wi-Fi, network performance, web, video streaming, OTT video MOS, app journeys, transaction journeys, PCAP, protocol and path, location context, probe availability, alerting, provenance and voice.
Where each one runs
The coverage matrix states which agent collects which family — including where a KPI is not possible on a platform.
How it is recorded
The explorer gives the metric set, the platform support and the conditions under which a sample counts.
EdgeX probe data sheets
Available on request
These are either commercially sensitive or written per engagement, so they are shared directly rather than published.
Pricing and engagement models
SaaS licence bands per offering, kit and probe pricing, and a three-year total cost view. Scoped to the deployment size you have in mind rather than a list price.
Test methodology and field-level KPI list
Test protocols, the field-level KPI list, sample validity conditions, statistical governance and panel sizing — the document a competitor’s technical team or an advertising standards body would review.
SOC 2 report and DPA
The current SOC 2 report, the data processing agreement, hosting options and the data-residency position for your market.
Start with the smallest thing that proves the point.
None of this needs to begin with a programme. Each starting point below answers a real question in weeks, and produces data you keep whether or not anything follows.
Pick the one that matches the argument you are currently losing
A crowdsourced read of your market
We show you what the existing 5GMARK dataset already says about your network and each of your competitors, city by city. Nothing to deploy, nothing to sign beyond an NDA. If the picture matches your internal view, that is useful; if it does not, that is more useful.
A 30-day pilot on one city or one venue
A small number of kits or probes in a place where you already suspect a problem — a mall, a stadium, a corridor, a set of cells you are about to optimise. Full portal access, full raw data export, a defined question to answer.
A head-to-head app benchmark
Your network against a named competitor on the apps that drive loyalty in your market, measured from the same kits at the same moments, with the methodology documented to a standard that survives external review.
Talk to us
Sanket Jain
Sales, business development and delivery — Connectivity
To scope anything properly we need three things
- The market or markets, and whether competitors are in scope
- The question you want answered — a claim to defend, a change to validate, a venue to monitor, a device set to compare
- Any constraint we should design around: data residency, procurement route, existing supplier contracts