This is a general guide to how a very large, high-stakes organisation is actually defended — not a list of products, but the reasoning: the scenario and its constraints, the reference models the design hangs on, and the decisions (above all the networking and identity decisions) that have to be made when 100,000+ people, their devices, and the regulated personal data they handle across many countries are in scope. It describes common industry practice, not any particular deployment.
To keep it concrete, everything below is designed for one fictional organisation, Meridian Group. The numbers are illustrative but realistic — they're the order of magnitude that changes what the right answer is. At this scale, the hard part is rarely knowing the control exists; it is making it work across hundreds of thousands of objects, dozens of jurisdictions, and an operating budget that cannot buy its way out of every risk.
1. The brief: the environment we're defending
| Dimension | Meridian Group (illustrative) |
|---|---|
| Workforce | ~120,000 employees + ~25,000 contractors/partners → >140,000 identities |
| Endpoints | ~150,000 managed (Windows, macOS, mobile) + BYOD |
| Servers / workloads | ~12,000 VMs & hosts on-prem + tens of thousands of cloud workloads & containers |
| Footprint | 4 regional data centres, Azure + AWS (multi-cloud), offices in 40+ countries |
| Data | Customer & employee PII, payment-card data (PCI), some health data, financial records, IP |
| Regulatory | GDPR (EU), India DPDP, US state privacy + HIPAA (one unit), PCI DSS, SOX, sector regulators |
| Identity estate | Hybrid: on-prem Active Directory + Microsoft Entra ID, plus SaaS federation |
| Business reality | Growth by acquisition (frequent M&A), legacy & OT in some sites, 24×7 operations, mergers add unknown estates overnight |
Those facts are the design's actual inputs. Scale (100k+) kills anything that depends on manual effort per object. Multi-jurisdiction turns “encrypt the data” into “encrypt it and keep EU data in the EU and prove it.” M&A means the estate is never fully known. Good security design makes choices that survive all three.
The non-negotiable design stance at this scale: assume breach (someone is already inside — design the interior accordingly), Zero Trust (never trust by network location; verify identity and device posture every time — per NIST SP 800-207), and defense in depth (no single control is load-bearing). Everything else is implementation detail of these three.
2. The reference models the design hangs on
A programme this size is not invented from scratch; it is assembled from established models so it is defensible to auditors, boards, and regulators, and so newcomers can read it. The ones that actually shape the drawings:
- NIST CSF 2.0 — the organising spine (Govern, Identify, Protect, Detect, Respond, Recover). Every programme area maps to a function.
- CISA Zero Trust Maturity Model v2 — the structure this page follows: five pillars — Identity, Devices, Networks, Applications & Workloads, Data — over three cross-cutting capabilities (Visibility & Analytics, Automation & Orchestration, Governance), scored across four stages (Traditional → Initial → Advanced → Optimal). It gives the roadmap a shared language of “how mature is each pillar.”
- Microsoft Enterprise Access Model — the modern privileged-access design (Control / Management / Data-&-Workload planes) that supersedes the legacy Tier 0/1/2 model and the old ESAE “Red Forest.” Central to a hybrid AD + Entra estate.
- ISO/IEC 27001 for the certifiable management system, CIS Controls v8 for a prioritised implementation checklist, and MITRE ATT&CK to measure detection coverage against real adversary behaviour.
- SABSA / TOGAF for the business-driven architecture method — every control traces up to a business requirement and down to an implementation, so spend is justifiable.
- Cloud shared-responsibility models — the dividing line of who secures what in Azure/AWS, which is itself a design artefact at this scale.
The rest of this page is organised by the five Zero-Trust pillars, then the three cross-cutting capabilities, then resilience, then the decision register and roadmap — the same order the design is usually presented.
3. Pillar 1 — Identity (the real perimeter at 140,000 accounts)
With no network perimeter to trust, identity is the primary control plane — and it is also where almost every real breach of an org this size begins. At 140,000 identities the problems are lifecycle and privilege, not login screens.
Identity fabric
- One authoritative IdP (Entra ID) fronting SSO for thousands of apps, with on-prem AD synchronised for legacy. Consolidation means one policy engine and one kill-switch — critical when you must disable a compromised account across 2,000 SaaS apps at once.
- Phishing-resistant MFA as the default, not an add-on: FIDO2 / passkeys bound to the origin so credentials can't be relayed or replayed. SMS/OTP is retired. This single move removes the majority of credential-based intrusions.
- Conditional Access as the real policy surface: allow this identity to this app only from a compliant, managed, low-risk device and expected location — otherwise step-up or block. Device compliance (from the endpoint pillar) is an input here, which is what ties identity and device together.
- Identity Governance (IGA) is the control that scale demands: automated joiner/mover/leaver from the HR system of record, access certifications, and role mining (SailPoint/Saviynt). Manual provisioning does not exist at 140k; orphaned and over-entitled accounts are the slow leak that IGA plugs.
- External identity (B2B) for the 25k contractors/partners — scoped, time-boxed, sponsored, and auto-expired, kept distinct from employee identity.
- Workload & service identity — managed identities and short-lived tokens instead of static service-account passwords; secrets in a vault (HashiCorp Vault / cloud secret stores), never in code. Stale service accounts with kerberoastable passwords are a classic way these estates fall.
Privileged access — the Enterprise Access Model
The crown jewels are the accounts that control the estate, so they get their own architecture. Meridian implements the Enterprise Access Model: assets are classified into the Control plane (identity and infrastructure control — domain controllers, Entra global admin, Azure root management groups, AD CS), the Management plane (server and workload administration), and the Data/Workload plane (the apps and data themselves). The rule that makes it work: a higher-plane credential never authenticates to a lower-plane, internet-exposed asset, because that is exactly the path that turns one phished laptop into Domain Admin.
- Privileged Access Workstations (PAWs) — dedicated, hardened, separately-managed devices for Control-plane admin. The population is small (typically 5–15 true Tier-0 admins even at this size), so PAWs are affordable for the highest tier and expand outward from there.
- PAM with just-in-time elevation (CyberArk/Entra PIM): no standing admin rights; access is requested, approved, time-boxed, and session-recorded. This collapses the window in which a stolen admin token is useful.
- Tiered/plane isolation enforced by Conditional Access and network so that administering a DC is only possible from a PAW through a jump host — directly defeating the lateral-movement and credential-theft paths that the AD attack-path map on this site is built around.
The directory itself
Design bias is single forest (a forest, not a domain, is AD's real security boundary) unless a hard requirement — a regulated or acquired entity needing true isolation — justifies another. Domains are kept few; regional child domains are used only where replication or administrative-boundary needs demand it. Sites and ~2–5 domain controllers per major location give resilience and sane logon locality; read-only DCs serve smaller/edge sites. AD and Entra are treated as the most critical asset in the company, because whoever owns the directory owns everything in it.
If budget were rationed, identity wins. Strong IdP + phishing-resistant MFA + IGA at scale + the Enterprise Access Model shrink the attack surface more than any appliance — and they are the controls an attacker emulating a real adversary will hit first.
4. Pillar 2 — Devices & endpoints (a 150,000-agent fleet)
Endpoints are where users and attackers land. At 150k devices the strategy is managed, measured, and revocable, and device posture feeds straight back into the identity pillar as a Conditional-Access signal.
- Unified endpoint management (Intune/MDM + autopilot provisioning) so every device is enrolled, configured to a hardened baseline, and reports compliance — and a non-compliant device is denied access rather than trusted.
- EDR/XDR on every endpoint and server (CrowdStrike / Defender XDR / SentinelOne): behavioural detection, one-click isolation, and telemetry feeding the SOC. Fleet management (agent health, coverage gaps) is itself an ops discipline at this size — an unmonitored 2% is 3,000 blind endpoints.
- Hardening to CIS Benchmarks / STIGs with drift detection, and application control / allowlisting (WDAC/AppLocker) on high-value systems — the strongest counter to commodity malware and the defense-evasion and living-off-the-land tradecraft attackers rely on.
- Patch & vulnerability remediation at scale via rings (pilot → broad → critical) with SLAs tied to exposure — because unpatched internet-facing services are the other great breach vector besides identity.
- Rich telemetry (Sysmon, PowerShell script-block logging, auditing) — see Windows telemetry & logging for what gets recorded and why in-memory tradecraft tries to dodge it.
- OT/IoT & unmanageable devices (manufacturing, medical, building systems) that can't take an agent are discovered and monitored passively (Claroty/Nozomi) and isolated aggressively on the network — the device you can't patch, you contain.
5. Pillar 3 — Networks (segmentation and access across 40 countries)
This is where the big networking decisions live, and at Meridian's scale they are made deliberately. The organising principle is segmentation to limit blast radius, and the depth of that segmentation is the single largest network trade-off in the whole design.
User & site access: de-perimeterised by default
- SASE as the access fabric for 145k users in 40+ countries — Secure Access Service Edge converges SWG, CASB, ZTNA and FWaaS in the cloud so a user in any office or at home gets identical policy and inspection close to them, instead of backhauling traffic to a regional DC.
- ZTNA replaces VPN. The old “on the VPN = on the LAN” model is gone; users reach only the specific applications they're entitled to, brokered by identity and device posture. No user is ever “on the flat network.”
- SD-WAN connects sites over the internet with policy and resilience, retiring expensive legacy MPLS where it can.
Data-centre & cloud: segmentation and east-west control
- Security zones: untrusted → DMZ → internal → restricted enclaves (cardholder-data environment for PCI scope reduction, PII stores, OT, and the Control plane). The most sensitive zones are the most isolated, and PCI scope is deliberately minimised by segmentation to shrink audit burden.
- Micro-segmentation (Illumio/NSX/cloud security groups) controls east-west (server-to-server) traffic so a foothold on one workload can't freely reach its neighbours. North-south has always been inspected; east-west is the modern battle and the difference between “one host owned” and “the data centre owned.”
- NGFW (Palo Alto/Fortinet) at zone boundaries — app-aware, user-aware, TLS-inspecting — and NAC/802.1X so an unmanaged device lands in quarantine, not on the server segment.
- Default-deny egress + DNS security. Outbound is allowlisted and DNS is filtered and monitored (Protective DNS), which strangles command-and-control and exfiltration. A tightly controlled egress is one of the most under-rated wins at scale.
- Cloud landing zones: a standardised, policy-guardrailed account/VNet structure (hub-and-spoke, private endpoints, no public IPs by default) so tens of thousands of cloud workloads inherit segmentation and logging rather than each team inventing its own.
- Encryption in transit everywhere, internal included, so sniffing a segment yields nothing and relay attacks lose their cleartext.
6. Pillar 4 — Applications & workloads (DevSecOps + cloud at scale)
Thousands of applications and a multi-cloud estate create new risk daily, so security shifts left (into the build) and right (into runtime), and is delivered as a platform teams consume rather than a gate they queue for.
- Secure SDLC as a platform: threat modelling, secure-coding standards, and security gates wired into CI/CD — SAST, DAST, SCA (open-source components), IaC scanning, and secrets scanning running automatically so 1,500 developers get fast feedback instead of a central bottleneck.
- Software supply chain: SBOMs, dependency pinning, signed builds, and artefact provenance — the response to a decade of supply-chain compromises.
- Runtime protection: WAF and API security (schema validation, rate limiting, bot management) in front of internet-facing apps.
- Cloud-native protection (CNAPP) unifying CSPM (misconfiguration), CWPP (workloads/containers), and cloud IAM analysis (Wiz/Prisma) — continuously, because an over-permissive cloud role is the cloud's Domain Admin, and a public storage bucket is the cloud's open file share.
- Shared-responsibility discipline: explicit ownership of which controls are the cloud provider's and which are Meridian's, documented per service — the most common cause of cloud breaches is assuming the provider has you covered when they don't.
7. Pillar 5 — Data (the thing we're actually protecting, across jurisdictions)
Ultimately the attacker wants data, and Meridian holds regulated PII in many countries, so controls wrap the data itself — the layer that still protects you after every other control is bypassed, and the layer regulators care about most.
- Classification taxonomy (public / internal / confidential / restricted) with labels (Microsoft Purview) that drive encryption, DLP, and sharing. You can't protect — or prove you protect — what you haven't labelled.
- Data residency & sovereignty: the multi-jurisdiction twist. EU personal data stays in EU regions, India DPDP and other local rules are honoured, and cross-border transfer is governed — which constrains where workloads, backups, and even SIEM analytics may run. This single requirement reshapes the whole architecture (regional data stores, regional log handling) and is a first-class design input, not an afterthought.
- Encryption at rest with keys in a KMS/HSM (ideally customer-managed keys for the most sensitive sets) and in transit with TLS; stolen disks or sniffed traffic are then useless.
- DLP & DSPM: Data Loss Prevention on email/web/endpoint to stop sensitive data leaving, and Data Security Posture Management to find where regulated data actually sprawled to (the shadow copies in a SharePoint nobody owns).
- Tokenisation / masking for the highest-sensitivity fields (card numbers, national IDs) and database activity monitoring on the crown-jewel stores — narrowing PCI scope and limiting what a compromised app can read.
8. Cross-cutting: Visibility & Analytics — the SOC at scale
Assume-breach only works if you can see the breach. At Meridian's scale the SOC is an engineering problem before it is a staffing one, because the telemetry is enormous.
The log-volume reality (why you can't just “send everything to the SIEM”)
Rough back-of-envelope: endpoints with EDR emit on the order of ~0.3 events/sec each and busy servers far more; across ~150k endpoints, ~12k servers, network, identity, and cloud, Meridian generates on the order of hundreds of thousands of events per second and tens of terabytes per day (a common planning figure is ~1,000 EPS ≈ 8.6 GB/day). You cannot afford to keep all of that hot in a premium SIEM, and you don't need to. So the pipeline is tiered:
- Hot tier — high-value, detection-relevant logs in the SIEM (Splunk/Sentinel) for real-time correlation and short retention.
- Cold / data-lake tier — the bulk of telemetry routed cheaply (via a pipeline like Cribl) to object storage for compliance retention and search-on-demand during investigations.
- Normalisation & routing at ingest so the SIEM sees structured, relevant data, not raw firehose — this is where a large part of SIEM cost is won or lost.
Turning data into response
- Detection engineering mapped to MITRE ATT&CK, written as code (Sigma rules in version control) and measured by technique coverage, not vendor dashboards. Quality of content, not brand of SIEM, defines a SOC.
- SOAR & automation are mandatory, not optional: with default rule-sets producing ~70% false positives at ~15 minutes of triage each, you cannot hire your way out — playbooks must auto-enrich, auto-isolate, and auto-close so humans spend time on judgement.
- UEBA to catch credential abuse signatures miss (impossible travel, a service account suddenly interactive).
- Deception — honey-tokens and decoy accounts (like the LDAP honeypot in one of the lab writeups): zero false positives, so a single touch is a high-confidence incident.
- Operating model: a tiered, follow-the-sun SOC across regional hubs (also helping with data-residency on where logs are analysed), proactive threat hunting, and an MDR partner for surge/after-hours — you staff for the median and burst to the partner for the peak.
9. Cross-cutting: Automation & Orchestration
Nothing at this scale is done by hand twice. Provisioning, hardening, patching, certificate issuance, access certification, and incident response are codified (infrastructure-as-code, policy-as-code, detections-as-code, SOAR playbooks). Automation is both a security control (consistency, no manual drift) and the only way the team keeps up — the ratio of assets to engineers makes manual operation a vulnerability in itself.
10. Cross-cutting: Governance, risk & the operating model
- Risk management as the steering wheel: a live risk register, scored impact/likelihood tied to the CIA triad, and board-level reporting. Finite budget goes where it buys the most risk reduction; residual risk is formally accepted by a named owner.
- Compliance mapping once, report many: a unified control framework mapped to GDPR, DPDP, PCI, HIPAA, SOX and ISO so a single control set satisfies many regimes and audits don't each reinvent evidence.
- Third-party & supply-chain risk: vendor assessments, continuous monitoring, and contractual security requirements — at this size, a large share of risk lives in suppliers and SaaS.
- M&A security: a repeatable playbook to assess, isolate, and integrate an acquired company's unknown estate without letting its compromise become yours — acquisitions are a top way large enterprises inherit breaches.
- Security organisation (target operating model): a CISO with architecture, GRC, security engineering/platform, SecOps (SOC/IR), identity, AppSec, and offensive teams — federated into business units with central policy. Clear RACI so “who owns this control” is never ambiguous.
- Offensive validation: continuous vulnerability management with risk-based prioritisation (CVSS for severity, EPSS for exploit probability, CISA KEV for known-exploited), regular penetration testing, red teaming, purple teaming, and Breach & Attack Simulation. The web pentest checklist and AD attack-path map here are the attacker's side of this loop.
11. Resilience & recovery (when, not if)
At this scale the question is not whether a serious incident happens but how fast you recover, so recoverability is a first-class design goal — and ransomware is the modelled worst case.
- Backups that survive the attacker: 3-2-1 plus immutable / offline copies that ransomware can't reach, at petabyte scale, with tested restores (an untested backup is a hypothesis).
- AD forest recovery as a specific, rehearsed plan — rebuilding the identity fabric is the critical path after a domain-wide compromise, and it is slow and error-prone if practised for the first time during the incident.
- Multi-region DR with defined RTO/RPO per service tier, failover exercised, and crown-jewel services designed to fail over cleanly.
- Incident response at scale: a defined lifecycle (prepare, detect & analyse, contain, eradicate, recover, learn), a retainer, legal/comms/regulatory- notification playbooks (GDPR's 72-hour clock, etc.), and tabletop + live exercises up to board level so the plan is muscle memory, not a document.
12. The key design decisions (decision register)
Everything above is well-understood; the craft is the decisions, because security always trades against cost, usability, performance, and operational reality. These are the big calls at Meridian's scale and the way a high-stakes design usually leans — each one is a conscious trade with governance accepting the residual risk.
| Decision | The tension | How this design leans — and why |
|---|---|---|
| Single vs. multiple AD forest | Simplicity & cost vs. hard isolation | Single forest by default (the real security boundary); a second only for a truly isolated/regulated or freshly-acquired entity |
| Segmentation depth | Blast-radius control vs. operational friction & cost | Micro-segment the crown jewels and PCI/PII enclaves; coarser elsewhere — you can't afford to micro-segment everything, so target it |
| Prevention vs. detection | Blocking breaks business; detection needs people | Both, but invest heavily in detection & response — prevention always eventually fails |
| SIEM: ingest everything vs. tier | Visibility vs. multi-million cost | Tier it — hot SIEM for detection, cheap data lake for retention/search; route at ingest |
| Central vs. regional SOC | Consistency vs. coverage & data residency | Follow-the-sun regional hubs under one central detection standard + MDR burst |
| Default-deny egress | Strong C2/exfil control vs. app-team friction | Default-deny with a managed allowlist — worth the friction |
| TLS inspection | Visibility vs. privacy, performance, breakage | Inspect at the edge with carve-outs for privacy-sensitive & certificate-pinned traffic |
| Cloud vs. on-prem | Elasticity vs. control & residency | Cloud-first with guardrailed landing zones; cloud security funded like the data centre |
| Data residency vs. central analytics | Legal localisation vs. one global view | Keep regulated data/logs regional; correlate on metadata/federated search, not by moving raw PII |
| Legacy & OT | Can't patch / can't touch vs. exposure | Isolate aggressively (segment or air-gap) and compensate with passive monitoring |
| Build vs. buy | Fit & control vs. speed & support | Buy platforms, build the detection content, integrations, and automation |
| Usability vs. lockdown | Friction drives shadow IT | Make the secure path the easy path (SSO, passkeys, self-service) rather than piling on friction |
The recurring truth: there is no absolute “secure” configuration, only an acceptable risk for this business. The value of good design is knowing exactly what is being traded and making that trade on purpose — then having the telemetry to know when the trade stops being acceptable.
13. Sequencing it: a maturity roadmap
You don't build all of this at once; you raise each Zero-Trust pillar through the CISA maturity stages, leading with the controls that cut the most risk per dollar. A defensible sequence for Meridian:
| Horizon | Focus | Representative moves |
|---|---|---|
| 0–6 months — stop the bleeding | Identity & visibility | Phishing-resistant MFA everywhere, kill legacy auth, EDR to full coverage, PAW + PIM for Tier-0, SIEM for the critical log sources, asset inventory |
| 6–18 months — raise the floor | Segmentation, IGA, DevSecOps | IGA joiner/mover/leaver, Enterprise Access Model rollout, SASE/ZTNA for users, PCI/PII enclave micro-segmentation, CI/CD security gates, CNAPP |
| 18–36 months — optimise | Automation & coverage | SOAR-driven response, detection-as-code with measured ATT&CK coverage, DSPM + data residency controls, purple-team loop, immutable backup + AD-recovery drills |
| Ongoing — sustain | Measure & adapt | Continuous validation (BAS/red team), risk-based patching, M&A integration playbook, board reporting, roadmap refresh |
Progress is tracked with metrics the board can read: MFA/EDR coverage %, mean time to detect/respond (MTTD/MTTR), % privileged access just-in-time, patch SLA attainment on KEV/EPSS-prioritised vulns, phishing-simulation failure rate, and detection coverage against ATT&CK — leading indicators of risk, not vanity counts of blocked events.
Why it matters — even from the attacker's chair
Every control here is something a red teamer will meet, and the best offensive work comes from understanding the defender's model rather than just the next exploit. Knowing why egress is filtered, why the directory is the crown jewel, why privileged access is plane-isolated, and why the PII enclave is micro-segmented tells you where the friction and tripwires are — and which single misconfiguration quietly collapses a carefully-built, hundred-thousand-user stack.
References & further reading
- NIST SP 800-207 — Zero Trust Architecture
- CISA — Zero Trust Maturity Model v2
- Microsoft — Enterprise Access Model
- NIST Cybersecurity Framework 2.0
- CIS Critical Security Controls v8
- MITRE ATT&CK