Payload staging
TL;DR: Staged = small stub fetches the real payload at runtime (smaller dropper, more wire signal); stageless = entire beacon ships in one blob (heavier dropper, no fetch). In-proc = post-ex runs in the beacon thread (quieter, riskier); fork-and-run = post-ex spawns a sacrificial process.
What it is
Staging is the choice between shipping a small loader that pulls the implant over the network (“stager”) versus shipping the whole implant (“stageless”). Within an active implant, “fork-and-run” vs “in-proc” describes whether post-ex tooling runs in the beacon’s own process or in a freshly-spawned sacrificial one. Both decisions are detection-surface tradeoffs.
Preconditions / where it applies
- You control the payload-generation step of your C2 framework
- You’re making an explicit OPSEC choice based on what the target’s EDR / network monitoring is good at
Technique
Stager vs stageless.
- Stager: dropper is tiny (a few KB) — easier to fit in a phishing attachment, easier to obfuscate, but the runtime fetch from stage-2 hosting is visible on the wire. Classic CS HTTP stager hits
/<checksum8>. - Stageless: dropper is the whole beacon (hundreds of KB to MB) — no fetch, but the artifact itself contains all signatures. Modern OPSEC defaults to stageless because edge content inspection often flags the stager’s deterministic checksum URI.
In-proc vs fork-and-run for post-ex (Cobalt Strike model, applies broadly).
- In-proc BOF (Beacon Object File): post-ex runs as a thread inside the beacon process. No extra process create. Failure crashes the beacon. No fork-and-run cleanup.
- Fork-and-run: spawn
rundll32(or a configured sacrificial process), inject the post-ex tool, run it, exit. Detection-rich because of the spawn + injection pattern. Survives tool crashes. Default for legacymimikatzetc.
Choose:
- Use in-proc BOF when the post-ex action is fast, well-tested, and you’ve already paid the cost of having beacon in memory
- Use fork-and-run when the tool is unstable (.NET assemblies, large legacy code) or when you want to push the noise into a process that’s expected to be noisy
- Use module-stomping / process-injection variants when even the spawn is too loud — overwrite a benign DLL’s
.textwith shellcode inside an existing process
Other staging modes.
- DLL stager: drops a DLL, loaded via LOLBin (rundll32, regsvr32)
- Shellcode loader: PE → shellcode (Donut, sRDI, pe2shc) → injector → memory
- Self-contained EXE: signed via stolen cert, packed, anti-emulation
- HTML-smuggled stageless: base64-encode the stageless artifact into a JS
Blobwith MIMEoctet/streamand trigger download client-side viamsSaveOrOpenBlobor a synthesised<a>click. Onlytext/htmltraverses the proxy, so MIME/extension blocks on the perimeter never see the executable; the file materialises from ablob:URL with no fetch event in most network monitors. Splitting the base64 across multiple<script>chunks or gzip-pre-compressing it further defeats inline content scanners that timeout on large inline strings.
Detection and defence
- Stagers: deterministic URI patterns, fixed sizes, fixed cipher results — Suricata rules catch them within hours of public release
- Fork-and-run: parent (legitimate proc) spawns child (rundll32/sacrificial) → immediately injected → makes outbound — three-step rule is broadly written
- In-proc BOF: harder to detect on process tree but memory artifacts (unbacked executable) still show
- Defenders should monitor for unbacked RX/RWX memory in long-running processes (Moneta, PE-sieve, HollowsHunter)
References
- Cobalt Strike blog — staging/stageless rationale
- Outflank blog — in-proc post-ex BOF design notes
- TrustedSec blog — sacrificial processes and OPSEC guidance
- ired.team — HTML/JS file smuggling — Blob-based client-side delivery that pairs cleanly with a stageless beacon
- c2-protocol-design process-injection-techniques