Customer Case Study · Telecom · North America
Infrastructure observability at hardware depth, insight before failure
How a Tier-1 telecom operator in North America gained its first unified view of a very large server estate, from server telemetry all the way down to HPE iLO hardware sensors, and moved from reacting to alerts to acting before anything breaks.
- Firstunified, central view of the entire server estate, previously none existedNo central view existed before
- 100%of HPE iLO endpoints connected centrally, previously disconnected islandsEvery endpoint, previously islands
- ~70%of hardware-related issues surfaced before service impactCaught before service impact
- ~40%faster MTTR on infrastructure incidentsAgainst the prior manual process
- ~50%reduction in manual health-checking effort across the estateEngineer effort returned
- ~60%of routine remediations executed under governed automationUnder governed automation
The Challenges
The operator runs a very large fleet of physical and virtual servers supporting heavy, continuous data pulling and processing workloads. There was no central solution monitoring this infrastructure at any comparable level, and the hardware layer was effectively invisible: every HPE iLO management interface was an island, reachable only one server at a time.
- No central observability
Server health lived in scattered tools and manual checks; nobody had a single picture of the estate.
- iLO data locked per server
Thousands of HPE iLO interfaces, each holding temperature, power and component health data, none of it available centrally.
- Hardware failures announced themselves
Degrading components were discovered when something broke, not before.
- No headroom for surprises
The servers ran heavy, continuous data pulling and processing around the clock, so even a small capacity or thermal drift quickly turned into a service-affecting incident.
- Manual, repetitive checking
Engineers walked server lists by hand for health verification, slow, partial, and unsustainable at estate scale.
- Alert-then-react posture
Everything downstream of a failure, nothing ahead of it: no forecasting, no preemption.
The architecture, and why hardware depth mattered
Most server monitoring stops at the operating system. That is where the metrics are easy to collect and where the tooling already exists, and it is one layer above where the failure usually starts. A fan degrading, a power supply drifting, a thermal envelope creeping upward: none of that shows in a CPU graph until it has already become an outage. Pulling the sensor layer into the same correlation as the workload layer is the whole architecture.
Governance across every stage: role-based access, approval gates on production-changing actions, full audit trails, and rollback on every automated remediation. Autonomy is graduated: the platform earned each class of preemptive action against live results before running it unattended.
What was switched on, and what was not
Preemptive action carries a failure mode that reactive action does not: acting on a prediction that was wrong. Moving a workload off a server that was never going to fail costs real capacity and real trust, so the boundary here is drawn tightly around actions that are reversible and cheap to have been wrong about.
Around 70 percent of hardware-related issues are now surfaced before service impact. That figure is about surfacing rather than about acting, and the distinction is deliberate: seeing a failure coming is the achievement here, and what to do about it stays largely with the engineers who own the capacity plan.
Runs without a human
- Continuous telemetry capture from agents across thousands of servers
- Native iLO collection of temperature, power, fan and component health, centrally
- Correlation of the hardware and workload layers into one view
- Surfacing emerging patterns such as thermal drift or a degrading component
- Routine remediations the operator has approved, under governance
Does not run without a human
- Anything that removes capacity from the estate without a reversible path back
- Scheduling a physical component replacement. The platform raises it; a person schedules it
- Acting on a prediction outside the approved remediation set, however strong the pattern
- Changing detection thresholds in a direction that reduces coverage
What happens now when hardware starts to degrade
The sequence below is what replaced walking the estate. The step worth noticing is the second one, because it is the only place a hardware sensor and a workload metric have ever been in the same sentence.
The Sentinel loop, as it runs here
Every Opstral deployment runs the same four-stage loop: Observe, Investigate, Act, Optimize. It is the methodology rather than a feature list, and the point of setting it out per deployment is that you can see which stages carried the weight in this one and which did not.
- ObserveTake in every signal the estate produces, normalised and correlated as it arrives rather than after somebody goes looking.HereServer telemetry and HPE iLO hardware sensors stream continuously, so the hardware layer and the workload layer arrive together rather than being reconciled after an outage.
- InvestigateWork the signal into a probable cause with the evidence attached, before anyone is notified.HereHardware and workload are correlated rather than merely collected, and Sentinel surfaces the degradation pattern before the threshold is crossed. A disk that is failing looks different from a disk that is full.
- ActRun the approved procedure where the blast radius allows it, or hand a named human the plan, the evidence and the rollback.HereApproved remediation runs under governance. Everything else is raised for a named engineer with the pattern and its evidence attached.
- OptimizeFeed the outcome back so the next run of the loop is better informed than the last.HereThe outcome adjusts the baseline, so the definition of normal for that class of hardware moves with the estate rather than being set once at install.
- Both layers stream continuously
Agents send server, workload and resource telemetry. The iLO interfaces, previously disconnected islands checked one server at a time, now report temperature, power, fan and component health into the same place.
PlatformObserve - Hardware and workload are correlated, not just collected
A thermal reading on its own is a number. The same reading alongside the workload that server is carrying, and the behaviour of its neighbours, is a pattern. This is the step that makes prediction possible rather than just monitoring.
PlatformInvestigate - Sentinel surfaces the pattern before the threshold
Emerging thermal drift, a component trending toward failure, a power envelope creeping. None of these cross a static threshold until late, which is why threshold monitoring finds hardware problems by outage.
PlatformInvestigate - Approved remediation runs, under governance
Where the pattern maps to an approved action and the action is reversible, it executes: workload moved, throttling applied. Around 60 percent of routine remediations now run this way.
PlatformAct - Everything else is raised for a person
A component scheduled for replacement, or any action that removes capacity without a reversible path, is raised with the evidence attached. The platform predicts; the engineer who owns the capacity plan decides.
Named engineerAct - The outcome adjusts the baseline
What happened feeds back into detection thresholds and procedures, so the estate baseline keeps improving rather than being tuned once at go-live and left to drift.
PlatformOptimize
The Impact
The estate went from invisible to instrumented, and from reactive to preemptive. Hardware stopped failing by surprise, engineers stopped walking server lists, and the heaviest data workloads gained the thermal and capacity headroom visibility they never had.
| Dimension | Before | After · unified & AI-driven |
|---|---|---|
| Estate visibility | Scattered tools, no central view at this level | One governed layer across the entire estate, server telemetry to hardware sensor |
| Hardware layer | iLO interfaces disconnected, checked one server at a time | 100% of iLO endpoints centrally connected: temperature, power, component health |
| Failure posture | React after breakage or threshold breach | ~70% of hardware issues surfaced and handled before service impact |
| Resolution speed | Slow, manual cross-checking across silos | ~40% faster MTTR on infrastructure incidents |
| Operations effort | Manual health walks across thousands of servers | ~50% less manual checking; ~60% of routine remediations automated under governance |
| Output | Raw dashboards, where they existed at all | Insights and recommendations with governed, reversible actions attached |
Connecting the hardware layer changed the physics of the operation: when you can see temperature, power and component health across every server centrally, and an AI is watching the trend lines, failures stop being surprises and start being work orders.
At a glance
- Industry
- Tier-1 Telecom Operator
- Geography
- North America
- Engagement
- Production deployment
- Scope
- Infrastructure & Hardware Observability
- Scale
- Thousands of servers, heavy data workloads
- Platform components
- Sentinel AI, server telemetry, HPE iLO hardware sensors, governed remediation
- Evidence basis
- Improvement figures are indicative, reflecting target and expected outcomes of the deployment rather than measured production results. Customer identity withheld by request.
See your estate down to the sensor
Agent-based telemetry, hardware-layer integration, and Sentinel's analysis in one governed platform, deployed inside your environment.
Customer identity withheld by request and referred to throughout as a "Tier-1 telecom operator in North America." Improvement figures are indicative, reflecting target and expected outcomes of the deployment; exact results vary by estate scope, workload profile and rollout phase.