Detection engineeringElasticSOC October 2026

Cutting SOC triage from 3,000 hours a month to under 400

At a managed SOC covering 14 customer environments, our analysts were logging about 3,000 hours a month on alert triage. A month of tuning later, that number was under 400. Coverage did not shrink to get there. This is how I did it.

Measure hours, not alerts

We were seeing roughly 800 alerts a day, about 90% of them from Elastic and the rest from Splunk. Every alert became a ticket, and analysts logged the time they spent on each one.

That time tracking turned out to be the most useful data we had. Alert count tells you which rules are loud. Hours tell you which rules are expensive. They are not the same list. A rule that fires 200 times a day but closes in ten seconds costs less than one that fires twelve times and takes twenty minutes each, because every one of those twenty minutes is an analyst's time.

At around 24,000 alerts a month, 3,000 hours works out to about seven and a half minutes per alert. No analyst is doing real investigation in seven and a half minutes, 800 times a day. Most of that time was spent confirming that something expected had happened again.

Find where the hours go

I pulled the time logged per ticket, grouped it by the rule that generated the ticket, and sorted by total hours. Then I worked from the top down.

The top of the list was not surprising. Anyone who had worked a shift could have guessed most of it:

What these had in common was the disposition. Analysts closed them as "typically seen." The investigation always ended the same way, and nobody learned anything from it.

One question for every rule

For each rule near the top, I asked: if this fires, what does the analyst do differently from what they did last time?

If the answer was "nothing," the rule needed to change. The kind of change depended on why the answer was nothing:

Why it was noise What I changed
Fires on a known, benign source Allowlist that specific source, scoped as narrowly as possible
Fires on activity that is only interesting at volume Raise the threshold, or alert on the pattern instead of each event
Fires many times for one underlying event Group duplicates into a single alert
Analyst spends the time looking up context Add the context to the alert: account type, asset owner, whether the source is a known scanner
Useful to see, but never needs a response Move it to a report or dashboard
Nobody could say what it was for Retire it

Every option in that table got used. The last one least often, and only after checking that the rule was not the sole coverage for anything.

An example: failed logins from expected users

This is a reconstruction of the pattern, not the production rule.

The original alert fired when an account crossed a small number of failed logins in a short window. It made no distinction between a user who mistyped a password three times, a service account with a stale password, and an attacker working through a list of accounts.

The tuned version split that into separate questions:

The detection got more specific without getting blinder. The noisy version and the tuned version see the same events. Only the tuned version asks a question an analyst can answer.

Not cutting real coverage

Reducing alert volume is easy if you are willing to stop detecting things. The work was making sure we were not.

What changed

Triage time went from about 3,000 hours a month to under 400, an 87% reduction, over roughly a month of work.

The number was the least interesting result. With fewer tickets, the alerts that remained got a real look, and response times on them improved. Analysts had time to work on projects instead of clearing a queue. Customers noticed too, and their concerns went down.

If you are starting this yourself

Tuning is not a one-time project. New customers, new software and new scanners bring new noise. What made the difference was having a measure everyone could see, and a habit of checking it.

← All writing