Social Engineering Assessments and the Consent They Require
Testing people is not testing systems. What can be assessed, what must be agreed first, and why individual results should almost never leave the room.
Social engineering targets your staff, not your systems. That changes what has to be agreed before it starts and what may be reported afterwards.
What the assessment covers
Email phishing is the most common. Voice pretexting, calling the service desk, is frequently the most effective route of all. Physical entry tests whether anyone challenges a stranger. Messaging platforms come up occasionally, where the same pretexts work with less scepticism because people expect email to be the attack surface.
What it does not cover is the fix. The awareness programme, the report button in the mail client, the verification script the service desk reads from: those stay yours to build. What moves the price is the number of distinct pretexts, the number of waves, the size of the target population, out-of-hours calling and whether physical entry is in scope. Lead time is the item buyers underestimate: a credible sending domain has to be registered and aged before it is used, so an exercise booked for next week runs on infrastructure prepared earlier or accepts a weaker pretext.
Who you test, and who builds the list
If you supply the target list, you get a named population that is comparable between rounds. If we build it, it comes from your website, job adverts, conference speaker pages, filings and the mail format inferable from one published address. How much of a workforce can be reconstructed from public sources is itself a finding, and usually the one that lands hardest with a board. How red teams get in covers the reconnaissance side.
The safer default is both: reconnaissance first, then the list goes to the sponsor before anything is sent, so people on long-term leave, in a live disciplinary process, or in roles where a pretext would be cruel come out of it quietly.
What has to be agreed first
Beyond the usual rules of engagement, three things specific to testing people.
Who authorises it. Employment and privacy questions arise the moment a real person's behaviour is recorded. In many organisations this needs HR and legal to have seen the plan. Some jurisdictions, and anywhere a works council exists, need more than that. That sign-off is what makes the exercise defensible afterwards.
Which pretexts are off limits. Some work extremely well and should not be used. Anything invoking redundancy, disciplinary action, a bereavement, a medical matter, or a real named individual's authority in a way that damages trust in them. An attacker would use all of them. The test still has to leave your organisation functioning afterwards. Whether synthetic or cloned audio may be used on calls belongs on the same list, settled in writing before anyone dials.
How results are reported. The most important, and covered below.
Two more get skipped and cause trouble later. What is captured and what is kept: a harvesting page can record that a password was submitted without recording the password. Decide which, decide where that data sits and when it is destroyed, and agree that anything actually entered triggers a reset. The stand-down path: one named person, reachable out of hours, who can confirm within minutes that an event is ours and can halt the exercise. Anyone testing on site carries written authorisation with that number on it and produces it the moment they are challenged. See physical penetration testing for how that plays out on the ground.
What goes wrong in practice
The gateway eats the wave. Nothing lands and you have tested mail filtering instead of people. Decide before you buy which one you are paying for. Both are legitimate. If the sender is allowlisted so mail reaches inboxes, the report has to say so, because a click rate from an allowlisted wave is not comparable to one from mail that survived your filters.
The floor gets warned. Someone spots the first message, posts it to a group chat, and the rest of the numbers collapse. That is a result, not contamination: record the time to first report and the time to a workforce-wide warning, because a company that warns itself in minutes has just demonstrated the defence you were looking for.
Your own response fires. A case opens, your sending domain goes into a takedown process, a fraud team freezes something. Agree de-confliction before the start: who is told, in what order, and what evidence is preserved before anything is torn down. For a bank this runs against a clock. Paragraph 182 of the RBI Directions of 31 July 2026 requires cyber incidents to be reported on the DAKSH platform within six hours of detection, and a named contact who can confirm inside minutes that the event is ours is what keeps that call correct.
Someone is genuinely distressed. The banned-pretext list and the stand-down contact exist for this. If an exercise harms a person, it stops being defensible whatever the sign-off said.
Individual results should stay out of the report
A phishing exercise can produce a list of who clicked. That list is almost never the useful output and is frequently harmful.
What is useful: how many, how quickly, whether anyone reported it, how long until the first report, and what the service desk did when called. Those measure the organisation.
What is harmful: names. Once staff learn that an exercise produces a list managers see, the rational response is to stop reporting anything ambiguous. The reporting rate is the single most valuable defensive metric you have. An assessment that improves click rates while destroying reporting rates has made you less safe.
The narrow exception is a person who repeatedly hands over credentials after training. That is a conversation for their line manager. Even then, the pathway should be agreed before the test, never improvised from results.
Reading the numbers honestly
A click rate without a pretext description is not comparable to anything. A generic template sent to two thousand mailboxes and a targeted message built from reconnaissance are different exercises. Quoting them as the same percentage across quarters measures the pretext, not the people.
The number to track is the reporting rate: what proportion recognised it and told someone, and how fast. That is what shortens an intrusion, and it is what a good assessment is designed to improve.
Who owns the fix
Almost none of it belongs to the security team, and an assessment whose findings are all addressed to them produces the same result next year.
| Route | What a failure shows | Who has to change something |
|---|---|---|
| Email phishing | A pretext was believed and credentials followed | Identity, for authentication a relayed password cannot satisfy; internal comms, for the reporting route |
| Voice pretexting | The service desk verified a caller on facts an attacker can gather | The service desk process owner: a callback to a number already on record, and a check that is not employee ID or a manager's name |
| Physical entry | Nobody challenged a stranger, or a door was held open | Facilities and building management. In a shared tower the reception and the guards are the landlord's, and your policy does not bind them |
| Bank-detail change | An instruction was actioned on the strength of an email | Finance, with callback verification against a number held before the request arrived |
Where it sits in a regulated programme
Social engineering is usually bought inside a red team engagement. For an Indian bank, the RBI Directions of 31 July 2026 set the cadence the rest of the programme is measured against: paragraph 151 requires vulnerability assessment at least every six months and penetration testing at least every twelve, for systems that are critical and / or sit in the DMZ with a customer interface. Red teaming appears at paragraph 162, which says red teams may be used, and the small finance bank, payments bank and credit information company Directions use the same permissive word. Red teaming under the RBI Directions 2026 takes it paragraph by paragraph.
The firm is held to a tighter standard than the technique. Paragraph 156 requires the credentials and competency of the testing firm and of the personnel assigned to be established at selection, appointment, engagement and renewal, and paragraph 159 brings CERT-In's audit policy guidelines into the supervisory relationship where a CERT-In empanelled auditor is engaged. Every Indian regulator that accepts an audit report accepts it from a CERT-In empanelled auditor, so empanelment is the first question you ask a provider. We have held it continuously since 2008. To a board, the artefact that carries weight is not a click chart: it is the service desk transcript, the timeline of what your own people did, and the authorisation trail showing the exercise was sanctioned before it ran.
When not to buy one
- Staff have nowhere to report. No button, no monitored mailbox, nobody triaging it. Build the channel, then measure it.
- Remote access still accepts a password alone. You know how the exercise ends. Spend the money on the control.
- Results are headed for appraisals, or the reporting pathway is not agreed in writing. If HR will not put its name to the plan, it is not ready to run.
- You have just had a real incident. A simulated attack weeks later reads as a trap, and the goodwill costs more than the finding.
- You want better detection. Replaying the technique with your defenders in the room gets there faster: see purple team, when it beats a red team and what the exercise tells you about detection.
For how these routes fit into a wider engagement, see how red teams get in, and for the paperwork that governs all of it, scoping a red team.
About the author
Siddarth G
Practice Director — Cybersecurity
Leads Security Brigade's offensive security practice with deep expertise in vulnerability research, penetration testing, and red team operations. Ranked Top 80 globally on Bugcrowd.
Continue reading
All articles →Choosing a Red Team Provider
Every firm answers yes to every capability question. Six questions where the generic yes runs out, and what a real answer sounds like.
Red Teaming and SEBI CSCRF
What the Cyber Security and Cyber Resilience Framework asks of regulated entities, where adversarial testing sits within it, and how the tiering decides how much applies to you.
What the Exercise Tells You About Detection
The findings are about your systems. The timeline is about your team. Read from the defender’s side, a red team report is a much more uncomfortable document.