SplunkDetection engineeringFedRAMP October 2026

When a log source goes quiet: baselining ingestion per source

Every detection on this site assumes the logs are arriving. When a source stops, nothing alerts, and a quiet SIEM looks exactly like a safe environment. This is how I monitor log sources so that a gap gets noticed in an hour instead of during the next investigation.

The first version: "no logs received"

At an earlier SOC, our check was the obvious one: alert when a source sent nothing for a set period.

It caught dead sources, and it was noisy. Plenty of logs are intermittent by nature. Software installation logs on a host might be empty for days and then produce a burst. Every one of those quiet periods raised an alert, and every alert was closed as normal.

The noise was not the worst part. Analysts learned that a quiet source was usually fine, so "that's how that log source usually is" became the default answer. That is the habit that lets a real outage sit unnoticed.

The question that matters

"Did any logs arrive?" is the wrong question. The useful one is: did this source send roughly what it usually sends at this time?

That needs a baseline for each source. At a FedRAMP cloud SOC with several customer environments, I built that baseline in Splunk for every index in every environment, and for each AWS sourcetype separately: CloudTrail, IAM and VPC Flow Logs. Each customer environment had its own alert and dashboard panels.

Two comparisons, not one

Each source's last 24 hours is compared with two things:

A source is flagged only when it is outside the expected range on both comparisons. Volume that is normal for a Sunday but low for the fortnight is not an outage. Volume that is low on both is.

The query

This is a reconstruction of the logic rather than the production search, which I no longer have. The structure is the same: 24-hour windows going back 14 days, a same-weekday average, a 14-day average and standard deviation, and a minimum tolerance of 10% of the average.

| tstats count where index=<customer_indexes> earliest=-15d@h latest=@h
    by index, sourcetype, _time span=1h
| eval days_ago = floor((relative_time(now(), "@h") - _time) / 86400)
| where days_ago <= 14
| stats sum(count) as events by index, sourcetype, days_ago
| eval key = index . "|" . sourcetype
| xyseries key days_ago events
| fillnull value=0
| untable key days_ago events
| eval days_ago = tonumber(days_ago)
| stats sum(eval(if(days_ago=0, events, 0))) as current
        avg(eval(if(days_ago>0, events, null()))) as daily_avg
        stdev(eval(if(days_ago>0, events, null()))) as daily_stdev
        avg(eval(if(days_ago=7 OR days_ago=14, events, null()))) as same_day_avg
    by key
| eval band = max(2 * daily_stdev, 0.10 * daily_avg)
| eval status = case(current < same_day_avg - band AND current < daily_avg - band, "low",
                     current > same_day_avg + band AND current > daily_avg + band, "high",
                     true(), "ok")
| eval status = if(status = "low" AND current = 0, "silent", status)
| where status != "ok"
| rex field=key "^(?<index>[^|]+)\|(?<sourcetype>.+)$"
| table index, sourcetype, status, current, same_day_avg, daily_avg, daily_stdev

A few details are doing more work than they look:

Too many logs is a finding too

The search flags high volume as well as low. A source sending several times its usual volume usually means something changed: a logging level left on debug, a duplicated input, a misconfigured forwarder, or an application in an error loop. Occasionally it is the activity you are looking for. Either way someone should know, and in Splunk it also shows up on the license bill.

Alert and dashboard

The alert ran every hour and fired when any source in an environment crossed its threshold. The dashboard showed the same sources over 14 days, so an analyst could see the shape of the drop or spike before deciding what to do:

| tstats count where index=<customer_indexes> earliest=-14d@d by sourcetype, _time span=1d
| timechart span=1d sum(count) by sourcetype

The pairing mattered. The alert says something is off. The dashboard shows whether it is a cliff, a slow decline or a single bad day.

What it caught

It found real problems: hosts that were down, logging that had been misconfigured, and ingestion issues on the customer side that nobody had noticed.

In a FedRAMP environment, a logging gap is not a curiosity. NIST SP 800-53 control AU-5 requires a response to audit logging failures, and a monitored source that stops sending is treated as an outage that needs to be resolved right away. An hourly check with a clear status meant those gaps were found within the hour, not whenever someone happened to look.

What changed for the analysts

Compared with the "no logs received" alert, it was much quieter, and the alerts it did raise were easier to read: a real issue or a quiet log, with the numbers to show which.

The bigger change was the habit. When an alert means a source is outside its own normal range, "that's how that log source usually is" stops being an answer, because the search has already checked how it usually is.

If you build one

For the same check written against CloudTrail and Entra ID, see Five CloudTrail detections to deploy first and Five Entra ID detections to deploy first.

← All writing