Building a Network Management System That Stays Up
Building a Network Management System That Stays Up
A Network Management System has one job that everyone takes for granted: it has to be the thing that's still working when nothing else is. Operators look to the NMS precisely when the network is on fire, alarms flooding in, links flapping, a card failing in a remote site. After two decades building these systems, I've learned that designing for the bad day is the whole discipline.
FCAPS is a contract, not a checklist
Fault, Configuration, Accounting, Performance, and Security, FCAPS gets taught like a list of boxes to tick. In practice it's a contract with the operations team. Fault management has to surface the *right* alarm, not every alarm. Performance management has to tell you something is degrading before the customer does. Configuration has to make device onboarding repeatable across thousands of network elements without a human babysitting each one.
When I review an NMS design, I'm really asking one question for each FCAPS area: what does the night-shift engineer actually need at 3 a.m., and does this give it to them without making them think?
Alarm correlation is where you earn trust
A raw alarm stream is noise. One fiber cut can light up hundreds of downstream alarms, and an operator drowning in red learns to ignore the system entirely. Good alarm correlation collapses that storm into a single root cause, "this card is down, here's everything it took with it."
Getting correlation right is unglamorous work: topology awareness, suppression rules, time windows that don't fire too early or too late. But it's the difference between an NMS people rely on and one they mute. The goal isn't to show more; it's to show less, and have the less be true.
Onboarding devices is the unsexy moat
Every NMS demo looks great with one shiny device. The real test is the tenth vendor, the firmware that drifts from the YANG model, the legacy box that only speaks SNMP while the new gear talks NETCONF. Device onboarding that scales, model-driven, protocol-flexible, resilient to the inevitable deviations, is what separates a product from a prototype.
I've spent a lot of my career in that seam between SNMP, NETCONF, and RESTCONF. It's tedious, and it's exactly the work that makes everything above it possible.
Design for the failure, present the calm
The systems I'm proudest of look boring in steady state and stay coherent under stress. High availability and scalability aren't features you bolt on at the end, they're decisions you make in the first architecture review and defend in every one after. Build for the storm, and the sunny days take care of themselves.

Narendra skipped presentations and built real AI products.
Narendra Billakanti was part of the April 2026 cohort at Curious PM, alongside 18 other talented participants.
