01

The challenge

The operator runs one of India's largest radio access networks: hundreds of thousands of sites and millions of cells, generating around 12.5 million alarms every day. Fault management was still anchored to a reactive, rule-based, alarm-centric model that had stopped scaling with the network.

  • Alarm floods and fatigue

    Static, hardcoded correlation rules missed complex multi-domain patterns, and real incidents were buried under symptomatic noise.

  • Slow, expert-dependent root cause work

    Finding the true cause meant several specialists manually interpreting logs, performance data, topology and history, which stretched resolution times.

  • Resolutions that did not hold

    Manually matching an incident to a playbook led to mismatched steps, repeat tickets and incidents that bounced between teams.

  • Reactive by design

    Work started only after a hard failure or a threshold breach. There was no forecasting of degradation and no way to act before customers felt it.

  • Every step needed a person

    Triage, diagnostic validation and recovery were all manual, which set the size and cost of the operations floor.

  • Data that cannot leave the building

    Network, topology and customer-impact data had to stay inside approved environments, which ruled out every public AI service.

02

What x101 does

One platform took fault management from reactive and manual to proactive and closed-loop. Six coordinated capabilities turn raw operational signals into a single actionable incident, an evidenced root cause and an approved action, with human approval and governance applied at every stage.

Step 01

One source of truth

Alarms, performance data, tickets, inventory and topology are consolidated and normalised into a single operational picture, with topology stitched together across domains.

Step 02

Cause separated from symptom

x101 reasons across the whole picture to tell the primary fault from its symptoms, grounded in the operator's standards and validated against live topology. The result is one clean incident carrying its confidence score and its citations.

Step 03

Early warning

Statistical and time-series models surface degradation signatures before they become service-affecting alarms. This path is pure machine learning, with no language model in the prediction.

Step 04

The write-up, drafted for you

x101 produces the operator-grade reason for outage, the root cause analysis and the plan of action on approved templates, then enriches and closes the ticket behind it.

Step 05

The right people told

The right message reaches the right stakeholder over email and SMS, with response timers, escalation paths and delivery tracking handled by the platform.

Step 06

Looking ahead

Forecasting models flag capacity exhaustion, congestion and service degradation before they arrive, so maintenance is scheduled instead of scrambled.

Governance runs across every stage. Each request carries the permissions of the person asking, confidence thresholds decide what x101 may do on its own, high-impact actions wait for human approval, and every decision is written to an immutable, replayable record.

03

Deployed where the data lives

The defining constraint was sovereignty: operational, topology and customer-impact data could not leave the operator's approved environment. x101 was deployed entirely on the operator's own premises, platform and AI models together, so the intelligence runs where the data already sits.

The platform

Inside the operator's data centre

  • Private deployment. x101 runs inside the operator's own data centre, with no dependency on an external cloud provider.
  • Carrier-grade resilience. Redundant, multi-instance services with automatic scaling and disaster recovery, built for round-the-clock operation.
  • Data stays inside the estate. x101 reads the existing inventory, fault and performance systems in place, and nothing crosses an approved boundary.
  • Enterprise security throughout. Single sign-on with multi-factor authentication, role-based access at every level, and encryption in transit and at rest.
  • Complete audit record. Every question, retrieval, action, approval and result is logged with its own transaction identifier.
The AI models

The operator's own, on its own hardware

  • In-house inference. Language models are served on the operator's dedicated hardware, and model selection is restricted to those on-premises services.
  • No public AI service involved. No prompt, alarm, ticket or network detail is ever sent to a commercial AI endpoint.
  • Tuned to this network. On-premises tuning and retrieval over the operator's standards, equipment documentation and resolution history produce answers specific to this estate.
  • Models trained on operator data. Detection and forecasting models are trained, versioned and retrained in-house, with drift monitoring, and are never externalised.
  • Guardrails at every interaction. Input and output checks stop prompt manipulation, sensitive-data leakage, unsupported conclusions and unapproved actions before they reach a person or a device.

04

The impact

Moving from a reactive, manual model to a closed-loop one changed the economics of the operations floor: less noise, faster resolution, fewer repeat incidents and a smaller, higher-value team footprint.

~60%

Less noise

Alarms compressed into actionable incidents with the primary cause named.

~50%

Faster resolution

A shorter lifecycle from first detection through to closure.

~40%

Efficiency gain

Manual triage and ticket handling avoided across the front line.

20–40%

Fewer incidents

Prevention and faster isolation, phased in as forecasting matured.

Fault management before and after x101, across seven dimensions
DimensionBefore · reactive and manualAfter · x101, on-premises
Alarm handlingMillions of raw alarms, static rules, constant fatigueAround 60% compressed into unique actionable incidentsSymptomatic noise and flapping suppressed automatically
Root causeManual, several specialists, hours per complex incidentAutomated and explainable, with confidence and citationsGrounded in operator standards, validated against topology
Time to resolveLong detection-to-resolution lifecycleAround 50% faster across defined incident families
Resolution qualityFrequent repeat and bouncing ticketsAround 60% resolved permanently on defined types
Failure postureReactive, action only after a breachProactive, with degradation and capacity risk forecast ahead of impact
Operations effortEvery triage and recovery step needed a personAround 40% efficiency gain, low-touch under governance
Data and AI posturePublic AI services off-limits, no safe path to adoptFully on-premises, sovereign and auditable end to end

By running the platform and its AI models entirely on its own premises, the operator gained modern AI in the network operations centre without ever compromising data sovereignty, turning a compliance constraint into an advantage.

Solution summary · x101 network fault management deployment

05

The parts of x101 this uses

Nothing here was built for one customer. Each capability below is standard platform behaviour, applied to a network operations problem.