← back to theory
Blue Team

Building a SOC — Detection to Response

“Assume breach” is only a strategy if you can actually see the breach. The Security Operations Centre is the function that turns an estate's torrent of telemetry into a small number of investigated, contained incidents. It is people + process + a data pipeline, run continuously — and its quality is mostly the quality of its detection content, not the brand on the SIEM.

This page walks the SOC end to end, from a raw event on an endpoint to a contained incident and the lesson that improves the next detection. It is where the C2 and evasion tradecraft elsewhere on this site is meant to be caught.

The detection-to-response pipeline

  SOURCES          PIPELINE             BRAIN              ACTION
  -------          --------             -----              ------
  Endpoint/EDR \
  Network/NDR   \   collect +           SIEM (hot)          Tier 1 triage
  Identity      +-> normalise  ----->   + detections  --->  Tier 2 investigate
  Cloud         /   (Cribl/              mapped to ATT&CK    Tier 3 hunt / IR
  App/DB       /    forwarders)          UEBA + threat intel     |
                     |                       |                   v
                     v                       +------------->  SOAR playbook:
                 Data lake (cold)                            enrich -> contain
                 retention + search                          (isolate host,
                                                              disable account)
                                                                   |
                                                                   v
                                           NIST IR: contain -> eradicate ->
                                           recover -> lessons learned

1. Telemetry — what gets collected

SourceWhat it tells youExamples
EndpointEDR, process/file/registry, Sysmonprocess trees, injection, persistence
NetworkNGFW, proxy, DNS, NDR, flowbeaconing, lateral movement, exfil
IdentityAD/Entra sign-ins, PAM, auth logsanomalous logons, kerberoasting, impossible travel
CloudCloudTrail, Azure Activity, GuardDutyIAM changes, resource abuse
App / dataWAF, DAM, app logsinjection, sensitive-data access
The hardest, least glamorous SOC problem is visibility: you cannot detect what you do not collect. A detection idea is worthless if the data source isn't onboarded — see Windows telemetry & logging for what to turn on and why.

2. The log pipeline & tiering

At enterprise scale this is hundreds of thousands of events per second and tens of terabytes per day, so you cannot keep everything hot in a premium SIEM. The pipeline tiers the data:

  • Hot tier — high-value, detection-relevant logs in the SIEM for real-time correlation, short retention.
  • Cold / data-lake tier — the bulk routed cheaply (Cribl, object storage) for compliance retention and search-on-demand.
  • Normalisation at ingest — structured, relevant data into the SIEM, not raw firehose; this is where most SIEM cost is won or lost.

3. The brain — SIEM, detections, UEBA, threat intel

  • SIEM (Splunk ES, Microsoft Sentinel, Google SecOps, Elastic) correlates everything and raises alerts.
  • Detection content is the heart — written and tuned as code, mapped to MITRE ATT&CK (its own page).
  • UEBA baselines normal and flags the anomalous (a service account suddenly interactive) — catching credential abuse that signatures miss.
  • Threat intelligence enriches alerts with context (is this IP/hash/TTP known-bad, who uses it) and drives proactive blocking and hunting.

4. SOAR — automation is mandatory, not optional

With default rule-sets producing roughly 70% false positives at around 15 minutes of triage each, you cannot hire your way out of the volume. SOAR (Palo Alto XSOAR, Splunk SOAR, Tines, Torq) runs playbooks that auto-enrich an alert, auto-contain (isolate host, disable account), and auto-close known noise — so analysts spend time on judgement, not toil.

5. The analyst workflow

TierRoleTypical SLA / focus
Tier 1Triage — validate & prioritise alerts, close known FPs, escalate real onesAcknowledge P1 in minutes; high throughput
Tier 2Investigation — scope across tools, pivot through telemetry, contain, eradicateHours; owns the incident
Tier 3Threat hunting & IR — hypothesis-driven hunts, forensics, major-incident leadProactive; deepest skill

A follow-the-sun model (regional SOCs hand off around the clock) gives 24×7 coverage; an MDR/MSSP augments for surge and after-hours. You staff for the median and burst to the partner for the peak.

6. Incident response lifecycle

Confirmed incidents follow the NIST SP 800-61 lifecycle:

  1. Preparation — tooling, runbooks, retainers, and rehearsals before the event.
  2. Detection & analysis — confirm, scope, and classify.
  3. Containment — stop the spread (isolate, disable), short- and long-term.
  4. Eradication — remove the foothold, persistence, and root cause.
  5. Recovery — restore cleanly, monitor for re-infection.
  6. Lessons learned — feed the gaps back into detections and controls (the loop that makes the SOC better).

7. The metrics that matter

MetricWhat it measures
MTTDMean time to detect — how long an attacker dwells unseen
MTTRMean time to respond/contain — how fast you stop the bleed
Dwell timeTotal intruder-present time; the headline risk number
Detection coverage% of relevant ATT&CK techniques with a tested detection
False-positive rateAlert quality; drives analyst burnout and automation need
Beware vanity metrics. “10 million events blocked” says nothing; MTTD/MTTR, dwell time, and ATT&CK coverage are the leading indicators of whether the SOC actually reduces risk.

Red team ↔ blue team

The SOC is the function that catches what the offensive pages do: C2 beaconing shows up as periodic egress to a rare destination; evasion is an arms race against the EDR and logging sources above; DCSync is directory-replication from a non-DC. Every attack technique has a detection opportunity — building those is detection engineering.

Key takeaways

  • The SOC is a pipeline: sources → tiered logging → SIEM/detection/UEBA/TI → SOAR → tiered analysts → IR → lessons learned.
  • Visibility is the hard part — you can't detect what you don't collect.
  • Automation (SOAR) is mandatory at scale, not a luxury.
  • Measure MTTD/MTTR, dwell time, and ATT&CK coverage — not blocked-event counts.

Related reading