NOC Performance: 4 metrics before the incident even starts.
Alert triage decides everything downstream, and it runs completely unmeasured.
Most NOC performance advice rides along with the incident process and its metrics. What is missed is the work before the incident, especially the work that never becomes one. That gap means detection, one of the NOC's most critical functions, goes untracked, and it leaves both performance and root cause open to speculation.
These are the alert metrics, the measures of the work before an incident exists, named here the way root9 reports them. They outline a team's strengths and weaknesses instead of averaging everyone into an incident number.
Time to Detect. From arrival to first human touch. This shows how quickly the team is able to react and evaluate any incoming alert, real or not. Insights: Disruption patterns here can expose distractions, overload, or individual performance issues. Time to Decide. From arrival to the real-or-not call, the moment the alert is cleared, becomes an incident, or pages someone. Judgment speed is the NOC's craft, and no standard framework has a name for it. Insights: Sliced by service or component, it shows where training or documentation is thin. Time to Incident. From arrival to a tracked incident, the whole span before the incident record exists, counted only for the alerts that become one. Insights: This is the one you can already approximate today, because the incident looks back at the alert that started it. Alert-to-incident ratio. What arrived versus what became an incident. Insights: This is the workload nobody sees, and the number that finally shows what the room absorbs on a bad week.
One of these can look wrong on the page. Time to Decide sometimes reads slower than Time to Incident, which seems impossible until you see that the two count different populations. Time to Incident is computed on the alerts that became incidents. Time to Decide is computed on every alert the room made a call on, the ones that turned out to be nothing included.
That gap is the whole argument in one place. The alerts that became incidents were the ones that got attention, and a handful of them is a small and flattering sample. If you only track incident metrics, you are seeing that fraction and nothing else. The work that never became an incident is where the real capacity of the room shows up, and it is the half nobody has been able to see.
It starts too late. MTTR, as most shops run it, starts when an incident is opened. The noticing, sifting, correlating, and judging that produced the incident all happened before that timestamp. Your NOC can get twice as fast at the part it does alone and the dashboard will not move. It only counts the survivors. The six alerts that became incidents get measured. The four hundred the NOC sifted, suppressed, or caught early never enter any metric anywhere. The NOC's largest work product is an invisible denominator.
Take the good case. A monitor fires, the alert becomes an incident, and there is a record of how it was handled. Alert time, who opened the incident, incident time. That satisfies the RCA and gives you a timeline. But when handling breaks down, an alert missed, an escalation that came late, that record cannot isolate where the failure actually happened, because everything around the failure is missing.
How many other alerts were in flight at the time?. How did this operator perform on the last 20 or 200 alerts?. Does this alert have high signal quality, or a history of noise?. Were any changes made to the alert or its monitor recently?. How many other operators were actively working alerts?. Where does this alert land? A shared channel, an inbox, a portal?
"Joe took 22 minutes to open an incident" is not enough to identify an area of improvement. Maybe Joe was carrying forty alerts mid-storm. Maybe this alert has cried wolf nine times this month. Maybe it landed in a portal nobody had open. There is no other span in an incident's history where a blind spot like this would be tolerated. Around alert handling, it has always been the norm.
Every measure above, and every question on that list, needs the alert layer to be a system of record, a timestamp for arrival, for eyes, for the decision, for the handoff, kept per signal. In most shops the alert layer is a pile of inboxes and channels, so the data to compute them does not exist. This work is not measured because nobody cares. It is unmeasured because everything upstream from the incident has traditionally been hard to capture.
In house. This is shaped differently in every shop, so there is no single answer. You start by centralizing and normalizing the alerts, then track and report every action taken on them. Off the shelf. root9 started at exactly this blind spot, and SigOps is the tool that came out of it. Setup is an email forwarding rule or a webhook, pointing the alerts at a shared team board.
root9 built SigOps this way, the board is the system of record for alerts, every arrival, every decision, every handoff captured as the room works, and Triage Health reads the alert metrics straight off it, Time to Detect, Time to Decide and Time to Incident, with MTTR alongside in its correct place. The context lives there too, what else was in flight, how a signal usually behaves, what changed recently, so the next RCA starts from a record instead of a recollection. The work that never becomes an incident finally earns credit, because for the first time it is countable.
What metrics should a NOC track? In addition to the ones most commonly listed, MTTR, MTTA, SLA compliance, a NOC should track the alert metrics, the measures of the work that happens before an incident exists. Time to Detect, from arrival to first human touch. Time to Decide, from arrival to the real-or-not call. Time to Incident, from arrival to a tracked incident. And the alert-to-incident ratio, what arrived versus what became an incident. Those are the ones you will not find on every other page, and they measure the team rather than the process.
Does MTTR measure NOC performance? No. MTTR, as usually implemented, is a shared measurement of the incident or impact time for all involved. The NOC plays a primary role, but it is best used as a process metric rather than a team metric. It is also a questioned one on its own terms, the VOID research project analyzed roughly 2,000 incident reports across 660 organizations and concluded incident durations are too skewed for a mean to describe reliability. The alert metrics, which track how a team handles alerts regardless of their disposition, are a better measure of NOC performance and will support improving the MTTR.
Why is it so hard to find root cause when an alert is missed? Because the context around the miss was never recorded. A name and a timestamp, who opened the incident and when, says nothing without what else was in flight, how that alert usually behaves, whether it changed recently, who else was working, and where it landed. Those answers only exist if the alert layer keeps state and history. In most shops it does not, so the RCA stops at speculation about a person instead of a picture of the moment.
How do you improve NOC performance? Start by measuring the right work. Baseline the alert metrics, Time to Detect, Time to Decide, Time to Incident and the alert-to-incident ratio, even roughly, even by hand for a week. The fixes they point to are mostly structural, centralize and normalize the alerts, then track and report every action taken on them, whether you build that in house or run it on a tool that already keeps the record. Re-measure after each change, frameworks alone change nothing the NOC does.