The operator's estate, more than twenty-seven thousand devices, generated a continuous flood of fault and performance data over three different collection mechanisms. Nothing monitored those devices centrally in real time, so what was happening inside the network was fragmented across collection silos and the operations centre saw it late, partially, or not at all.
Twenty-seven thousand devices, and no single place to see their fault and performance state as it happened.
Traps and metrics, event streams and polled collection each lived in its own pipeline, with no correlation across them.
Carrier-scale telemetry arrived faster than any team could read it, and meaningful signals drowned in the stream.
Root-causing a fault meant hopping between element managers and raw counters, device by device.
Where views existed they showed symptoms, and getting an actual live value still meant logging into the device.
Device faults, server health and service state were monitored in separate tools with separate truths.
x101 was deployed as the central intelligence layer for the operations centre: ingesting all three collection paths at carrier scale, processing fault and performance streams in near real time, and putting investigation on the platform rather than the engineer. Probable causes are suggested with evidence, approved procedures run automatically, incidents are raised with context, and engineers ask for live values in plain language instead of hunting through consoles.
Traps and metrics, event streams and polled collection from 27,000+ devices land in one pipeline, normalised and correlated on arrival.
Events and performance indicators are handled as they land, so degradation surfaces within moments rather than at the next reporting cycle.
Each fault is investigated against topology, history and correlated signals, and the probable cause is presented with its evidence. The engineer decides.
Engineers ask for current values and states in plain language and get answers drawn from the live estate, alongside the central dashboards.
Approved procedures run automatically for recognised patterns, incidents are raised with context attached, and every action flows through a reversible, audited record.
The same layer watches the underlying infrastructure and the services riding on the network, so cause and impact are visible in one place.
The operations centre moved from fragmented, delayed device monitoring to one near-real-time picture of twenty-seven thousand devices, with investigation, live values and action available in the same place the fault appears.
Measured in production within six months of go-live.
Running as governed procedures rather than manual work.
Three ingestion paths correlated into a single pane.
Raised automatically with probable cause and context attached.
| Dimension | Before | After · with x101 |
|---|---|---|
| Device visibility | No central real-time view across the device estate | 27,000+ devices in one near-real-time fault and performance view |
| Data ingestion | Three collection paths in separate, uncorrelated silos | One pipeline, normalised and correlated on arrival |
| Investigation | Manual, device-by-device root-causing across consoles | Probable cause suggested with evidence attached, engineer in command |
| Access to truth | Log into the device to read an actual value | Live values and states answered in plain language |
| Response | Ad-hoc manual procedures, untracked | 85% of routine operations automated as governed procedures, 43% MTTR reductionBoth measured in production within six months of go-live |
| Coverage | Network, infrastructure and services in separate tools | One governed layer across devices, infrastructure and services |
At twenty-seven thousand devices the question is never whether the data exists, it always does. The question is whether anyone can see it in time, trust what it means, and act on it safely.
Solution summary · x101 carrier-scale network observability deployment
Nothing here was built for one customer. Each capability below is standard platform behaviour, applied to this problem.
About this case study. The customer's identity is withheld at their request and is referred to throughout as a Tier-1 telecom operator in India. The MTTR reduction and routine-operations automation figures are measured production results from this deployment; the remaining figures are indicative, and actual results vary with estate scope, telemetry profile and rollout phase.