“Assume breach” is only a strategy if you can actually see the breach. The Security Operations Centre is the function that turns an estate's torrent of telemetry into a small number of investigated, contained incidents. It is people + process + a data pipeline, run continuously — and its quality is mostly the quality of its detection content, not the brand on the SIEM.
This page walks the SOC end to end, from a raw event on an endpoint to a contained incident and the lesson that improves the next detection. It is where the C2 and evasion tradecraft elsewhere on this site is meant to be caught.
The detection-to-response pipeline
SOURCES PIPELINE BRAIN ACTION
------- -------- ----- ------
Endpoint/EDR \
Network/NDR \ collect + SIEM (hot) Tier 1 triage
Identity +-> normalise -----> + detections ---> Tier 2 investigate
Cloud / (Cribl/ mapped to ATT&CK Tier 3 hunt / IR
App/DB / forwarders) UEBA + threat intel |
| | v
v +-------------> SOAR playbook:
Data lake (cold) enrich -> contain
retention + search (isolate host,
disable account)
|
v
NIST IR: contain -> eradicate ->
recover -> lessons learned
1. Telemetry — what gets collected
| Source | What it tells you | Examples |
|---|---|---|
| Endpoint | EDR, process/file/registry, Sysmon | process trees, injection, persistence |
| Network | NGFW, proxy, DNS, NDR, flow | beaconing, lateral movement, exfil |
| Identity | AD/Entra sign-ins, PAM, auth logs | anomalous logons, kerberoasting, impossible travel |
| Cloud | CloudTrail, Azure Activity, GuardDuty | IAM changes, resource abuse |
| App / data | WAF, DAM, app logs | injection, sensitive-data access |
The hardest, least glamorous SOC problem is visibility: you cannot detect what you do not collect. A detection idea is worthless if the data source isn't onboarded — see Windows telemetry & logging for what to turn on and why.
2. The log pipeline & tiering
At enterprise scale this is hundreds of thousands of events per second and tens of terabytes per day, so you cannot keep everything hot in a premium SIEM. The pipeline tiers the data:
- Hot tier — high-value, detection-relevant logs in the SIEM for real-time correlation, short retention.
- Cold / data-lake tier — the bulk routed cheaply (Cribl, object storage) for compliance retention and search-on-demand.
- Normalisation at ingest — structured, relevant data into the SIEM, not raw firehose; this is where most SIEM cost is won or lost.
3. The brain — SIEM, detections, UEBA, threat intel
- SIEM (Splunk ES, Microsoft Sentinel, Google SecOps, Elastic) correlates everything and raises alerts.
- Detection content is the heart — written and tuned as code, mapped to MITRE ATT&CK (its own page).
- UEBA baselines normal and flags the anomalous (a service account suddenly interactive) — catching credential abuse that signatures miss.
- Threat intelligence enriches alerts with context (is this IP/hash/TTP known-bad, who uses it) and drives proactive blocking and hunting.
4. SOAR — automation is mandatory, not optional
With default rule-sets producing roughly 70% false positives at around 15 minutes of triage each, you cannot hire your way out of the volume. SOAR (Palo Alto XSOAR, Splunk SOAR, Tines, Torq) runs playbooks that auto-enrich an alert, auto-contain (isolate host, disable account), and auto-close known noise — so analysts spend time on judgement, not toil.
5. The analyst workflow
| Tier | Role | Typical SLA / focus |
|---|---|---|
| Tier 1 | Triage — validate & prioritise alerts, close known FPs, escalate real ones | Acknowledge P1 in minutes; high throughput |
| Tier 2 | Investigation — scope across tools, pivot through telemetry, contain, eradicate | Hours; owns the incident |
| Tier 3 | Threat hunting & IR — hypothesis-driven hunts, forensics, major-incident lead | Proactive; deepest skill |
A follow-the-sun model (regional SOCs hand off around the clock) gives 24×7 coverage; an MDR/MSSP augments for surge and after-hours. You staff for the median and burst to the partner for the peak.
6. Incident response lifecycle
Confirmed incidents follow the NIST SP 800-61 lifecycle:
- Preparation — tooling, runbooks, retainers, and rehearsals before the event.
- Detection & analysis — confirm, scope, and classify.
- Containment — stop the spread (isolate, disable), short- and long-term.
- Eradication — remove the foothold, persistence, and root cause.
- Recovery — restore cleanly, monitor for re-infection.
- Lessons learned — feed the gaps back into detections and controls (the loop that makes the SOC better).
7. The metrics that matter
| Metric | What it measures |
|---|---|
| MTTD | Mean time to detect — how long an attacker dwells unseen |
| MTTR | Mean time to respond/contain — how fast you stop the bleed |
| Dwell time | Total intruder-present time; the headline risk number |
| Detection coverage | % of relevant ATT&CK techniques with a tested detection |
| False-positive rate | Alert quality; drives analyst burnout and automation need |
Beware vanity metrics. “10 million events blocked” says nothing; MTTD/MTTR, dwell time, and ATT&CK coverage are the leading indicators of whether the SOC actually reduces risk.
Red team ↔ blue team
The SOC is the function that catches what the offensive pages do: C2 beaconing shows up as periodic egress to a rare destination; evasion is an arms race against the EDR and logging sources above; DCSync is directory-replication from a non-DC. Every attack technique has a detection opportunity — building those is detection engineering.
Key takeaways
- The SOC is a pipeline: sources → tiered logging → SIEM/detection/UEBA/TI → SOAR → tiered analysts → IR → lessons learned.
- Visibility is the hard part — you can't detect what you don't collect.
- Automation (SOAR) is mandatory at scale, not a luxury.
- Measure MTTD/MTTR, dwell time, and ATT&CK coverage — not blocked-event counts.