By Doug Green
“AI goes blind at exactly the moment when you need it most.”
In this Technology Reseller News podcast, Vishal Gupta, Director of Product Management at ZPE Systems, explains why AI-driven infrastructure management needs an independent path to the devices it is expected to monitor, troubleshoot and recover.
AIOps platforms have become increasingly effective at detecting problems, correlating events and automating routine infrastructure operations. The problem, Gupta says, is that these systems often run on the same production infrastructure they manage.
When a network outage or hardware failure occurs, the AI platform can lose both its connection to the affected equipment and access to the telemetry it needs to diagnose the problem.
“That is the gap out-of-band fills,” says Gupta.
Out-of-band management provides an independent management plane that remains separate from the production network. Even when the primary infrastructure is unavailable, IT teams—and increasingly AI agents—can still reach devices through console access, examine system logs and kernel messages, and take corrective action.
Gupta compares the architecture to an airport. Aircraft use the runway for normal operations, while emergency and service vehicles have separate roads and infrastructure. If the runway becomes unavailable, the service infrastructure can still reach the aircraft.
The same principle applies to resilient IT operations. An isolated management environment should have its own connectivity, security, routing, switching, storage and compute capabilities. It may also include failover connectivity through 4G, 5G or satellite services such as Starlink.
Building Out-of-Band for a Larger Edge
ZPE Systems developed its Nodegrid Net Services Router 2U, or NSR 2U, in response to customers operating increasingly large and complex edge environments.
These environments can include branch offices, remote facilities, ships, oil rigs, cell sites and other locations outside traditional data centers. They frequently contain more devices, require greater bandwidth and have fewer trained personnel available on-site.
The NSR 2U was designed around three priorities: greater capacity, increased resiliency and support for AI workloads.
The modular platform offers 10 expansion-card slots, allowing customers to configure the system around their particular deployment. It also includes redundant, field-serviceable power supplies and fans, two NVMe storage slots with RAID support, four native 10-gigabit SFP+ ports and an increased Power over Ethernet budget.
ZPE has even addressed the possibility that the out-of-band device itself could fail. Two NSR 2U systems can be interconnected so that one system can provide remote console, power and reset control for the other—effectively providing out-of-band management for the out-of-band infrastructure.
Taking NVIDIA Jetson AI to Remote Locations
ZPE Systems has also developed an NVIDIA Jetson AI Expansion Card for the Nodegrid NSR family. The card supports NVIDIA Jetson Orin Nano and Orin NX modules, providing local AI processing within the isolated management environment.
This allows organizations to deploy AI agents close to the infrastructure and data they manage, without relying entirely on a remote cloud connection.
A key capability is remote lifecycle management. IT teams can remotely flash the Jetson operating system, deploy or update AI agents and models, access the console, and power the device on, off or into recovery mode.
Ordinarily, updating or recovering an edge AI device may require someone to travel to the location and connect directly to the hardware. ZPE’s approach is intended to reduce those truck rolls while allowing organizations to manage distributed AI infrastructure centrally.
Potential applications extend beyond AIOps. The platform can support real-time video analytics, object detection, smart recording, manufacturing quality control, sensor-data aggregation and local automation. GPIO and I2C interfaces also allow sensors measuring conditions such as temperature, vibration or voltage to feed information directly into locally running AI models.
Asking the Hard Infrastructure Questions
Gupta says much of the AI conversation remains focused on models, software and the token economy. Those areas are important, but they can obscure fundamental infrastructure questions.
Where will an AIOps platform run? Can it survive the outage it is expected to resolve? Will it still have a path to the affected equipment? Can it access sufficiently accurate data to diagnose the problem and select the right recovery action?
“If you can’t answer these questions, then there’s a gap in your AIOps strategy,” Gupta says. “No software and no model will fix it for you.”
As AI becomes more autonomous, infrastructure resilience will determine whether AI agents can move beyond identifying failures to actually recovering from them. ZPE Systems is positioning isolated out-of-band infrastructure, the NSR 2U and edge-based Jetson AI processing as the foundation for making that transition possible.
More at Enterprise Network Management Solution | ZPE Systems