What we tested, and on what
The estate is our own: 31 managed endpoints, one Microsoft 365 tenant, Defender for Endpoint on every device and Microsoft Sentinel as the security information and event management (SIEM) platform. Every rule in it was written here, by us. Testing ran over 2 days from a managed host under a named account, with each command written to a transcript. Techniques were taken from execution, persistence, credential access, discovery and collection. Nothing in initial access: phishing people who know the test is running measures nothing.
The counts
- Alerted, with an analyst-actionable alert inside 15 minutes: 12
- Logged, with the evidence present in the SIEM platform and no rule that matched it: 6
- Silent, with no telemetry collected at all: 4
The 4 silent techniques divide in two. Two were cloud control-plane actions where the directory audit log was retained inside the tenant and never forwarded to the SIEM platform. The other 2 ran on a self-managed virtual machine built for laboratory work and never enrolled in the endpoint agent. Our asset inventory listed it as enrolled. We wrote that inventory by hand eight weeks earlier, and neither gap was a detection engineering problem. Both were inventory problems wearing a detection engineering label.
The alerts that fired for the wrong reason
Of the 12 alerts, 3 keyed on artefacts of the test rather than on the behaviour: file paths under the Atomic Red Team working directory, the word atomic in a command line, and one binary name that only that project uses. We re-ran those 3 with renamed binaries, altered arguments and a different working directory. One fired. Two did not. The honest alerted count for this estate is 10 of 22, and we wrote both of the rules that failed, for this estate, knowing what they were meant to catch.
A rule that only fires on the default filename is not a detection. It is a signature for one tool, and the tool is ours.
What we changed
- The 2 brittle rules rewritten against process lineage and handle access, using Sysmon event 10 for LSASS rather than image names
- Directory sign-in and audit logs forwarded to the SIEM platform, with retention set to 12 months
- The unenrolled machine enrolled, and the asset inventory rebuilt from the directory rather than kept by hand
- The 6 logged-not-alerted techniques worked through in order of exploitability, 4 closed so far
The limit of this exercise is the size of the estate. Thirty-one endpoints and one tenant will tell you whether a rule expresses the behaviour. They will not tell you whether it holds at 30,000, where the same logic returns a queue no analyst can work and the rule gets tuned until it is decorative. We will not know that until we run it somewhere that size, and we will publish the counts when we do. What we will commit to now is the shape of the report: techniques tested, techniques alerted, and the date each figure was measured, in place of a rule count. A coverage figure that only ever rises is a figure nobody is testing.