What the Exercise Tells You About Detection
The findings are about your systems. The timeline is about your team. Read from the defender’s side, a red team report is a much more uncomfortable document.
Most of a red team report describes what the attackers did. The part that matters more describes what you did, and it usually gets read too quickly.
Three ways detection fails
No data. The activity produced no log anywhere: an unmonitored system, a log source that was never onboarded, a network segment nothing watches. The fix is collection, and it costs money.
Data, no alert. The evidence was collected and nothing drew attention to it. This is the largest category in most engagements, and it is the cheapest to fix: usually a rule that does not exist or a rule scoped to the wrong source. It is also the most demoralising to read, because the answer was sitting in the platform the whole time.
Alert, no response. Something fired and nothing happened. Acknowledged and closed, lost in volume, or raised at 02:00 to a rota that does not operate at 02:00. It is the one failure mode here that no tool fixes, and the one that most often survives another year of spending.
A report that says "not detected" without distinguishing these three gives you a number and no diagnosis.
Some teams keep a fourth column: alerted, responded, not contained. Someone acted and the action did not stop the attack (an account disabled while the session token stayed live, a host isolated after the data had already left). Separate it from the third category where your timeline supports the distinction, because it points at runbooks and tooling instead of staffing.
Agree what counts as detected before the exercise starts
The argument at the debrief is almost always definitional. The red team says the beaconing was never caught. Your analyst says the traffic was logged and there is a ticket. Both can be true, and the classification above only works if the bar was set in advance.
Set it in the rules of engagement: detection means a human read the evidence and acted on it inside the exercise window. A log line that existed is collection. A dashboard nobody was watching is collection. Write that down while nobody's reputation depends on the answer.
What the timeline is actually measuring
Two intervals, measured on your own estate instead of borrowed from a benchmark: how long from the first attacker action to a human knowing, and how long from a human knowing to the action that stopped it. Most detection metrics you already report are proxies for those two.
For a bank the first interval has a regulatory shape. Under the RBI Directions issued on 31 July 2026, ¶182 requires cyber incidents to be reported on the DAKSH platform within six hours of detection. That six-hour clock starts where the first confirmed detection sits on your timeline, and the exercise is the only thing that measures what happens before it starts. See red teaming under the RBI Directions 2026.
The limits of "We had no alerts"
Teams frequently conclude the exercise proves their monitoring is useless. It usually proves something narrower: that the specific techniques used, against the specific systems touched, within that window, were not caught.
Two things limit what a single red team can tell you. The exercise tested one route. A different objective would have exercised different detections. And a team that got caught early may have run into bad luck as much as a strong defence.
Breadth is what purple teaming is for. A red team gives depth on one path; purple gives coverage across many.
Four things the exercise left untested
Write this list at the same time as the findings. It is what stops a board reading one result as a verdict on the whole programme.
- Techniques nobody used. The team chose a route and stopped when the objective was met. Everything it walked past is unmeasured, including things you believe you would catch.
- The hours nobody worked. An exercise run on ordinary weekdays says nothing about a four-day festival weekend, a quarter-end close or the nights under a change freeze, when the rota is thinnest and escalation is slowest.
- Detection under load. Your analysts met the exercise in a normal queue. The same alert arriving during a live outage is a different experiment.
- Everything outside the agreed scope. The subsidiary, the acquired entity still on its own domain, the plant network. Attackers do not respect the boundary you drew for commercial reasons.
Two things that distort the score
You caught the exercise, not the technique. A red team's infrastructure carries artefacts a real attacker's may not: a domain registered weeks ago, a hosting range your filtering already dislikes, a payload compiled that morning. Where the alert fired on one of those, ask whether the underlying behaviour would have fired anything at all.
Your side was primed. If word travelled beyond the white cell (a change ticket raised for the test window, an approval mail to a team lead, a service desk told to expect unusual calls), you measured a team that was watching for it. This inflates the result in the direction everyone prefers, so it rarely gets challenged in the debrief.
When your monitoring belongs to someone else
A lot of Indian monitoring is bought, not built: a managed service covering business hours, an on-call arrangement after that, and your own people holding escalation. The exercise then produces findings about two organisations, and only one of them is in the room.
Settle three things before the exercise.
- Permission. Your contract with the provider may require notice before adversarial testing crosses their platform. Finding this out mid-exercise stops the exercise.
- Who receives the timeline. A provider who first sees the undetected column inside your steering committee pack arrives defensive. A provider handed the timeline directly, with the classification attached, usually writes the missing rules.
- Who pays for the rules. Detection engineering on a third-party platform is often a chargeable change. A backlog of rule gaps you cannot fund is a finding you will read again next year.
What to do with it
Take the timeline and, for every action marked undetected, decide which of the three failure modes it was. That classification is the actual work product, and it converts a narrative into a funded plan: collection gaps go to engineering, rule gaps to the detection team, response gaps to whoever owns the rota.
Then re-run the specific techniques and confirm they now fire. That is a purple team exercise. It is cheap, and it turns the money already spent into something measurable.
| Failure mode | Evidence that proves it | Who closes it | What closed looks like |
|---|---|---|---|
| No data | No event from that host or segment anywhere in the platform for the timestamp window | Infrastructure and platform engineering | The source is onboarded and a query returns the action |
| Data, no alert | The raw event exists at the timestamp and no detection matched it | Detection engineering | A re-run of the technique produces an alert |
| Alert, no response | Alert record showing fire time, acknowledgement time and the action taken | Whoever owns the rota and the escalation path | The same alert, re-fired in a drill, reaches a human who contains it |
Map each undetected action to its MITRE ATT&CK technique before filing it. It costs an hour and it makes the row portable: next year's exercise, your platform's own coverage claims and the rule someone eventually writes all refer to the same identifier.
The constraint on closing the list is rarely ideas. It is whoever writes detection rules having the time, and a re-test to prove the rule works. Put a name and a date against each row at the debrief, while the people who can commit to both are still in the room.
What it proves outside your own team
To a board. The classification, not the breach story. Counts in three columns that you generated on your own estate, the same three counted again after remediation, and a cost attached to each. That reads as a programme and it survives the question "what did we get for the money".
To a SEBI-regulated IT Committee. The CSCRF technical clarifications of 28 August 2025 recommend that regulated entities consider deploying a range of security solutions in consultation with their IT Committee, such as threat simulation, vulnerability management and decoy systems. It is a recommendation to consider, and that consultation goes better with evidence from your own estate than with a vendor's slide. See red teaming and SEBI CSCRF.
To a bank's audit committee. ¶151 of the RBI Directions sets vulnerability assessment at least every six months and penetration testing at least every twelve, for systems that are critical and/or sit in the DMZ with a customer interface. ¶162 says red teams may be used. The case for the exercise is therefore one you make internally, and the detection timeline is the strongest part of it.
To an auditor. Every Indian regulator that accepts an audit report accepts it from a CERT-In empanelled auditor, and where an empanelled auditor is engaged, ¶159 brings CERT-In's audit policy guidelines into the supervisory relationship. We have held empanelment continuously since 2008.
When not to buy this
You already know the answer is "no data". If the systems in scope send nothing to a platform, the exercise buys an expensive restatement of a gap you can list today. Onboard the sources first, then test whether they are any use.
Nobody is on the other end. The three categories only separate if someone owns collection, someone owns rules and someone owns the rota. Where one person holds all three alongside a day job, the report has no recipient and every row lands in the same queue.
You want coverage, not depth. If the question is which techniques in your threat model you can see, a red team answers it for the handful it used. A purple team answers it across the set, and your people watch the rules get written as it runs.
Nobody gets blamed
A red team report names the moment a person did not challenge a stranger, or an analyst closed an alert. If the exercise produces disciplinary consequences it will be the last honest one you get, because everyone who might have reported something ambiguous will now stay quiet.
Say so before the exercise, in writing, from someone senior enough for it to hold. Then make the debrief a joint session with the red team and your defenders in the same room, so the timeline gets walked through together. See what a red team report contains.
About the author
Siddarth G
Practice Director — Cybersecurity
Leads Security Brigade's offensive security practice with deep expertise in vulnerability research, penetration testing, and red team operations. Ranked Top 80 globally on Bugcrowd.
Continue reading
All articles →Choosing a Red Team Provider
Every firm answers yes to every capability question. Six questions where the generic yes runs out, and what a real answer sounds like.
Red Teaming and SEBI CSCRF
What the Cyber Security and Cyber Resilience Framework asks of regulated entities, where adversarial testing sits within it, and how the tiering decides how much applies to you.
Social Engineering Assessments and the Consent They Require
Testing people is not testing systems. What can be assessed, what must be agreed first, and why individual results should almost never leave the room.