When a log source goes quiet: baselining ingestion per source
Every detection on this site assumes the logs are arriving. When a source stops, nothing alerts, and a quiet SIEM looks exactly like a safe environment. This is how I monitor log sources so that a gap gets noticed in an hour instead of during the next investigation.
The first version: "no logs received"
At an earlier SOC, our check was the obvious one: alert when a source sent nothing for a set period.
It caught dead sources, and it was noisy. Plenty of logs are intermittent by nature. Software installation logs on a host might be empty for days and then produce a burst. Every one of those quiet periods raised an alert, and every alert was closed as normal.
The noise was not the worst part. Analysts learned that a quiet source was usually fine, so "that's how that log source usually is" became the default answer. That is the habit that lets a real outage sit unnoticed.
The question that matters
"Did any logs arrive?" is the wrong question. The useful one is: did this source send roughly what it usually sends at this time?
That needs a baseline for each source. At a FedRAMP cloud SOC with several customer environments, I built that baseline in Splunk for every index in every environment, and for each AWS sourcetype separately: CloudTrail, IAM and VPC Flow Logs. Each customer environment had its own alert and dashboard panels.
Two comparisons, not one
Each source's last 24 hours is compared with two things:
- The same day in previous weeks. Log volume follows the work week. Monday is not Saturday, and comparing a Sunday with a weekday average flags every weekend as an outage.
- The average day over the last 14 days. This catches gradual changes that a same-day comparison misses, and gives a steadier baseline when a single week was unusual.
A source is flagged only when it is outside the expected range on both comparisons. Volume that is normal for a Sunday but low for the fortnight is not an outage. Volume that is low on both is.
The query
This is a reconstruction of the logic rather than the production search, which I no longer have. The structure is the same: 24-hour windows going back 14 days, a same-weekday average, a 14-day average and standard deviation, and a minimum tolerance of 10% of the average.
| tstats count where index=<customer_indexes> earliest=-15d@h latest=@h
by index, sourcetype, _time span=1h
| eval days_ago = floor((relative_time(now(), "@h") - _time) / 86400)
| where days_ago <= 14
| stats sum(count) as events by index, sourcetype, days_ago
| eval key = index . "|" . sourcetype
| xyseries key days_ago events
| fillnull value=0
| untable key days_ago events
| eval days_ago = tonumber(days_ago)
| stats sum(eval(if(days_ago=0, events, 0))) as current
avg(eval(if(days_ago>0, events, null()))) as daily_avg
stdev(eval(if(days_ago>0, events, null()))) as daily_stdev
avg(eval(if(days_ago=7 OR days_ago=14, events, null()))) as same_day_avg
by key
| eval band = max(2 * daily_stdev, 0.10 * daily_avg)
| eval status = case(current < same_day_avg - band AND current < daily_avg - band, "low",
current > same_day_avg + band AND current > daily_avg + band, "high",
true(), "ok")
| eval status = if(status = "low" AND current = 0, "silent", status)
| where status != "ok"
| rex field=key "^(?<index>[^|]+)\|(?<sourcetype>.+)$"
| table index, sourcetype, status, current, same_day_avg, daily_avg, daily_stdev
A few details are doing more work than they look:
- Windows are rolling 24 hours, not calendar days, so the alert can run every hour and always compare like with like.
- The
xyseries,fillnull,untablestep fills in zeros for windows where a source sent nothing. Without it, a source that went completely silent would have no row for the current window and would drop out of the results, which is exactly the failure this search exists to catch. The zeros also keep the average honest for intermittent sources. - The tolerance band is two standard deviations, with a floor of 10% of the average. A noisy source gets a wide band. A very steady source would otherwise have a band near zero and alert on trivial changes.
- Intermittent sources mostly tune themselves. Software install logs have a large standard deviation, so a quiet day sits inside the band. That was the main source of noise in the first version.
Too many logs is a finding too
The search flags high volume as well as low. A source sending several times its usual volume usually means something changed: a logging level left on debug, a duplicated input, a misconfigured forwarder, or an application in an error loop. Occasionally it is the activity you are looking for. Either way someone should know, and in Splunk it also shows up on the license bill.
Alert and dashboard
The alert ran every hour and fired when any source in an environment crossed its threshold. The dashboard showed the same sources over 14 days, so an analyst could see the shape of the drop or spike before deciding what to do:
| tstats count where index=<customer_indexes> earliest=-14d@d by sourcetype, _time span=1d
| timechart span=1d sum(count) by sourcetype
The pairing mattered. The alert says something is off. The dashboard shows whether it is a cliff, a slow decline or a single bad day.
What it caught
It found real problems: hosts that were down, logging that had been misconfigured, and ingestion issues on the customer side that nobody had noticed.
In a FedRAMP environment, a logging gap is not a curiosity. NIST SP 800-53 control AU-5 requires a response to audit logging failures, and a monitored source that stops sending is treated as an outage that needs to be resolved right away. An hourly check with a clear status meant those gaps were found within the hour, not whenever someone happened to look.
What changed for the analysts
Compared with the "no logs received" alert, it was much quieter, and the alerts it did raise were easier to read: a real issue or a quiet log, with the numbers to show which.
The bigger change was the habit. When an alert means a source is outside its own normal range, "that's how that log source usually is" stops being an answer, because the search has already checked how it usually is.
If you build one
- Baseline each source separately. One threshold across all sources is the "no logs received" alert with extra steps.
- Compare with the same weekday as well as the overall average, and require both before alerting.
- Make sure silent sources stay in the results. Test it by searching a time range where a source really did stop.
- Alert on high volume as well as low.
- Expect new sources to be noisy for their first two weeks, until they have a baseline. Holidays will also show up as low. Both are easier to explain than an outage nobody noticed.
For the same check written against CloudTrail and Entra ID, see Five CloudTrail detections to deploy first and Five Entra ID detections to deploy first.