01

The challenge

The operator's estate, more than twenty-seven thousand devices, generated a continuous flood of fault and performance data over three different collection mechanisms. Nothing monitored those devices centrally in real time, so what was happening inside the network was fragmented across collection silos and the operations centre saw it late, partially, or not at all.

  • No central real-time view

    Twenty-seven thousand devices, and no single place to see their fault and performance state as it happened.

  • Three ingestion worlds

    Traps and metrics, event streams and polled collection each lived in its own pipeline, with no correlation across them.

  • Volume beyond human triage

    Carrier-scale telemetry arrived faster than any team could read it, and meaningful signals drowned in the stream.

  • Investigation as archaeology

    Root-causing a fault meant hopping between element managers and raw counters, device by device.

  • Dashboards without answers

    Where views existed they showed symptoms, and getting an actual live value still meant logging into the device.

  • Network, infrastructure and services apart

    Device faults, server health and service state were monitored in separate tools with separate truths.

02

What x101 does

x101 was deployed as the central intelligence layer for the operations centre: ingesting all three collection paths at carrier scale, processing fault and performance streams in near real time, and putting investigation on the platform rather than the engineer. Probable causes are suggested with evidence, approved procedures run automatically, incidents are raised with context, and engineers ask for live values in plain language instead of hunting through consoles.

Step 01

Three paths, one pipeline

Traps and metrics, event streams and polled collection from 27,000+ devices land in one pipeline, normalised and correlated on arrival.

Step 02

Processed as it arrives

Events and performance indicators are handled as they land, so degradation surfaces within moments rather than at the next reporting cycle.

Step 03

Investigated, not just displayed

Each fault is investigated against topology, history and correlated signals, and the probable cause is presented with its evidence. The engineer decides.

Step 04

Live values, asked for in words

Engineers ask for current values and states in plain language and get answers drawn from the live estate, alongside the central dashboards.

Step 05

Known patterns, handled

Approved procedures run automatically for recognised patterns, incidents are raised with context attached, and every action flows through a reversible, audited record.

Step 06

Network, infrastructure and services together

The same layer watches the underlying infrastructure and the services riding on the network, so cause and impact are visible in one place.

Governed autonomy, deliberately bounded. x101 suggests root causes rather than acting on its own judgment, automation is limited to predefined and approved procedures, and every incident, ticket and action carries a full audit trail and a rollback path. Autonomy is earned scope by scope, never assumed.

03

The impact

The operations centre moved from fragmented, delayed device monitoring to one near-real-time picture of twenty-seven thousand devices, with investigation, live values and action available in the same place the fault appears.

43%

Lower MTTR

Measured in production within six months of go-live.

85%

Routine ops automated

Running as governed procedures rather than manual work.

27,000+

Devices in one view

Three ingestion paths correlated into a single pane.

~60%

Incidents auto-created

Raised automatically with probable cause and context attached.

Carrier-scale network operations before and after x101
DimensionBeforeAfter · with x101
Device visibilityNo central real-time view across the device estate27,000+ devices in one near-real-time fault and performance view
Data ingestionThree collection paths in separate, uncorrelated silosOne pipeline, normalised and correlated on arrival
InvestigationManual, device-by-device root-causing across consolesProbable cause suggested with evidence attached, engineer in command
Access to truthLog into the device to read an actual valueLive values and states answered in plain language
ResponseAd-hoc manual procedures, untracked85% of routine operations automated as governed procedures, 43% MTTR reductionBoth measured in production within six months of go-live
CoverageNetwork, infrastructure and services in separate toolsOne governed layer across devices, infrastructure and services

At twenty-seven thousand devices the question is never whether the data exists, it always does. The question is whether anyone can see it in time, trust what it means, and act on it safely.

Solution summary · x101 carrier-scale network observability deployment

04

The parts of x101 this uses

Nothing here was built for one customer. Each capability below is standard platform behaviour, applied to this problem.