Cutting SOC triage from 3,000 hours a month to under 400
At a managed SOC covering 14 customer environments, our analysts were logging about 3,000 hours a month on alert triage. A month of tuning later, that number was under 400. Coverage did not shrink to get there. This is how I did it.
Measure hours, not alerts
We were seeing roughly 800 alerts a day, about 90% of them from Elastic and the rest from Splunk. Every alert became a ticket, and analysts logged the time they spent on each one.
That time tracking turned out to be the most useful data we had. Alert count tells you which rules are loud. Hours tell you which rules are expensive. They are not the same list. A rule that fires 200 times a day but closes in ten seconds costs less than one that fires twelve times and takes twenty minutes each, because every one of those twenty minutes is an analyst's time.
At around 24,000 alerts a month, 3,000 hours works out to about seven and a half minutes per alert. No analyst is doing real investigation in seven and a half minutes, 800 times a day. Most of that time was spent confirming that something expected had happened again.
Find where the hours go
I pulled the time logged per ticket, grouped it by the rule that generated the ticket, and sorted by total hours. Then I worked from the top down.
The top of the list was not surprising. Anyone who had worked a shift could have guessed most of it:
- Failed logins from expected users. People mistyping passwords, and the same handful of accounts failing every morning.
- Scripts and service accounts running with stale credentials. A scheduled task with an old password fails on every run, every day, until someone fixes it.
- Vulnerability scans. Authorized scanners doing exactly what they are supposed to do, which looks a lot like reconnaissance.
What these had in common was the disposition. Analysts closed them as "typically seen." The investigation always ended the same way, and nobody learned anything from it.
One question for every rule
For each rule near the top, I asked: if this fires, what does the analyst do differently from what they did last time?
If the answer was "nothing," the rule needed to change. The kind of change depended on why the answer was nothing:
| Why it was noise | What I changed |
|---|---|
| Fires on a known, benign source | Allowlist that specific source, scoped as narrowly as possible |
| Fires on activity that is only interesting at volume | Raise the threshold, or alert on the pattern instead of each event |
| Fires many times for one underlying event | Group duplicates into a single alert |
| Analyst spends the time looking up context | Add the context to the alert: account type, asset owner, whether the source is a known scanner |
| Useful to see, but never needs a response | Move it to a report or dashboard |
| Nobody could say what it was for | Retire it |
Every option in that table got used. The last one least often, and only after checking that the rule was not the sole coverage for anything.
An example: failed logins from expected users
This is a reconstruction of the pattern, not the production rule.
The original alert fired when an account crossed a small number of failed logins in a short window. It made no distinction between a user who mistyped a password three times, a service account with a stale password, and an attacker working through a list of accounts.
The tuned version split that into separate questions:
- Failures followed by a success from an unusual source stayed a high-priority alert. That is the case that matters.
- Failures across many accounts from one source became its own alert, for password spraying.
- Repeated failures from a known service account stopped generating triage tickets. Instead, they went to the account owner as a fix. A service account failing every morning is a hygiene problem, and suppressing it forever just hides it.
- Failures from authorized vulnerability scanners were excluded by scanner address, not by account.
- Every remaining alert carried the account type and asset owner, so the analyst could tell at a glance whether it was a person, a service or a scanner.
The detection got more specific without getting blinder. The noisy version and the tuned version see the same events. Only the tuned version asks a question an analyst can answer.
Not cutting real coverage
Reducing alert volume is easy if you are willing to stop detecting things. The work was making sure we were not.
- Every change was mapped against MITRE ATT&CK. Before a rule was tuned or retired, I checked which techniques it covered and whether anything else covered them. If the tuned rule was the only coverage, it had to keep detecting that technique.
- Every change was peer reviewed. A coworker reviewed each change as a GitLab merge request, and pipeline checks had to pass before anything was deployed.
- Data was never dropped, only alerts. Moving a rule to a dashboard or a report kept the events searchable. If a quiet alert turned out to matter during an investigation, the evidence was still there.
What changed
Triage time went from about 3,000 hours a month to under 400, an 87% reduction, over roughly a month of work.
The number was the least interesting result. With fewer tickets, the alerts that remained got a real look, and response times on them improved. Analysts had time to work on projects instead of clearing a queue. Customers noticed too, and their concerns went down.
If you are starting this yourself
- Start with time spent, not alert count. If your ticketing system does not track it, start tracking it before you tune anything.
- Work the top of the list first. A few rules almost always account for most of the cost.
- Ask what the analyst would do differently. If the answer is nothing, the rule is not finished.
- Fix the source when you can. A stale service account password is a ticket for its owner, not a permanent exception.
- Drop alerts, never data. And have someone else review every change.
Tuning is not a one-time project. New customers, new software and new scanners bring new noise. What made the difference was having a measure everyone could see, and a habit of checking it.