Zeroization, part 2: Clear the stack and registers

In [part 1](/2026/10/06/zeroization-1/), we added a wipe and ended up with an extra stack copy that survived it. We can't just add another `secure_memzero()` for that copy: the compiler's spill slot has no name in our C program. But we know where to look. The operation used stack space and registers. If we clear those, a new spill inside the cleared region is already covered. Let's try that. We'll need to work out how much stack to clear, which registers we can touch, and how to get control back to do the cleanup. ## Decide where the operation ends First, let's decide what we're trying to erase. Suppose our application derives an ephemeral key, processes a batch of records, and then discards the key. We can put that whole sequence inside one cleanup boundary: ```text enter boundary derive ephemeral key process records retire owned key and context buffers leave boundary clear stack used by the work and its callees clear residual registers return status ``` Notice where the key derivation happens. If we derive it in the caller first, we may already have copies in that caller's frame or preserved registers. Wrapping the later encryption call won't remove those, so the derivation belongs inside too. And what comes back out? A status is easy to preserve while clearing the temporary computation state. Returning a secret key means deliberately retaining that key somewhere, and output buffers still need their own lifetime rules. We're cleaning the temporary storage used by this operation after it returns. A key the caller keeps for the next batch is outside that promise. ## Let's ask the compiler ```sh clang -O3 -fzero-call-used-regs=all -c crypto.c ``` GCC and Clang can generate the register-clearing instructions for us. The flag applies to the compilation unit; the `zero_call_used_regs("all")` attribute lets us select a function. Here's a declaration requesting it for the function's definition: ```c struct request; __attribute__((noinline, zero_call_used_regs("all"))) int secret_batch(struct request *request); ``` Now look at the end of the scalar arm64 example from part 1. After calculating the return value, we get these instructions among the added overwrites: ```asm mov x8, #0 mov x9, #0 movi.2d v0, #0000000000000000 movi.2d v1, #0000000000000000 ``` There we go. `x8` actually gets cleared this time. The result in `x0` survives, as it should. This [Godbolt comparison](https://godbolt.org/z/nc5KKe44o) shows the same source with and without register cleanup on AArch64 Clang 22.1.0. Let's try something bigger. I tested register clearing on libaegis with GCC 16.2 on arm64. The register matches disappeared, but 103 recognized secret words were still on the stack. Of course: we've asked for register cleanup. It doesn't discover spills in callees, and a spotless wrapper epilogue can leave a key schedule below it on the stack. The word `all` also needs a closer look. It refers to the compiler's supported call-used register set, not every architectural register, preserved register portion, flag, or stack slot. Options ending in `-gpr` are narrower still: they omit vector registers, which are where a lot of cryptographic state lives. Check the [attribute's contract](https://clang.llvm.org/docs/AttributeReference.html#zero-call-used-regs) and the epilogue for your target. Now let's ask GCC to handle the stack as well: ```sh gcc -O2 -fstrub=all -fzero-call-used-regs=all -c crypto.c ``` Its `strub` mechanism tracks stack use with a watermark, then clears the recorded region after the call. The compiler inserts both operations, so those unnamed spill slots can be covered too. [GCC documents those interfaces here](https://gcc.gnu.org/onlinedocs/gcc/Stack-Scrubbing.html). I tested stack scrubbing alone as well: the stack matches disappeared, but register matches remained. With both facilities enabled, the test found no matches. Better. But the scan only recognized particular byte representations in the state it sampled. It didn't prove the absence of every possible encoding or inspect every architectural extension. Before applying that command to a library, check which functions it actually covers. Mode selection has eligibility rules, and separately built callees need review. An `internal` wrapper can retain arguments or results outside its scrubbed body. An `at-calls` interface changes the function type and must agree across declarations and calls. Marking a function `callable` permits the call without promising to scrub its frame. Those distinctions are in the [`strub` contract](https://gcc.gnu.org/onlinedocs/gcc/Common-Attributes.html). ## What about our own register-wiping helper? Writing a few zeroing instructions sounds easy enough. Let's look at what happens when an ordinary function tries to clear a register it must preserve for its caller. Here's an abbreviated AArch64 sequence: ```asm str x19, [sp, #-16]! mov x19, #0 ldr x19, [sp], #16 ret ``` There's our `mov x19, #0`. And right after it... `ldr` restores the old value. Oops. The helper has undone its own wipe. The caller gets its register back, as the ABI requires, and the temporary stack save is still in memory too. So a separate helper can't freely destroy all registers and remain an ordinary ABI-compliant call. Putting the registers in an inline assembly clobber list doesn't get us out of this either. It tells the compiler to protect live values, which can mean making more copies. A boundary wrapper instead preserves the caller's original state, runs the secret work, clears that work's stack, and clears the residual register state it may discard. Any secret already present in the preserved caller state remains the caller's responsibility. Then we need the right register list for the actual ABI and instruction set. On AArch64, preserved vector-register portions need different handling from fully volatile vectors. On x86, clearing XMM state alone isn't a policy for AVX-512 registers and masks, and `vzeroupper` leaves the low 128-bit lanes intact. Windows x64 also preserves vector registers that the common Unix x86_64 convention doesn't. Don't copy a register list between targets and call it portable. ## How much stack should you wipe? Let's say we clear 4 KiB below the current stack pointer. Why 4 KiB? Adding up our local arrays won't answer that. We need to include callees, compiler spills, alignment, saved registers, and ABI areas such as the x86_64 red zone. The whole region also has to be valid to write. We could measure how deep one call goes, but recursion, variable-sized frames, callbacks, and instrumentation can change it. The deepest point reached in a test doesn't give us a maximum for every input. This is the limit of a fixed-size stack scrub, including a call to `sodium_stackzero(n)`. It clears a requested scratch region; it doesn't discover the entire preceding call tree's stack use or clear the registers for you. A compiler can track the extent or calculate a bound under appropriate restrictions. Or we can arrange to own the entire stack used by the work. ## Give the work its own stack Let's give the operation a dedicated stack. Run the sensitive sequence there, switch back, and clear the whole allocation. Now we have a region with a known start and end. The wrapper I tested has this interface: ```c long secure_call_on_stack(long (*fn)(void *), void *ctx, void *stack, size_t stack_len); ``` This is an assembly-wrapper interface, not a standard C function you can get by including a system header. The callback returns a status and receives its inputs through `ctx`. Now look at what the wrapper does after the callback returns. Here's the actual AArch64 stack-clearing excerpt: ```asm mov sp, x21 mov x9, x19 b 2f 1: stp xzr, xzr, [x9], #16 2: cmp x9, x20 b.lo 1b ``` That first `mov` is easy to skip over. `x21` holds the wrapper's original stack pointer, so we're switching back *before* overwriting the dedicated stack. `x19` holds the dedicated region's start and `x20` its aligned end; the callback preserves these metadata registers. Register cleanup and restoration of the caller's state follow this excerpt. They belong in the same reviewed wrapper; calling arbitrary C helpers afterward can create new saves or change the register contents again. I tested the wrapper around selected operations from libaegis, libsodium, and libhydrogen on arm64 and translated x86_64. After each dedicated-stack call, every byte in that allocation was zero. The residue scan also found none of its recognized secret words in the sampled registers or dead stack. This time, the cleanup covered the whole dedicated allocation. The prototype still omitted x87, SVE, and SME handling, so it demonstrates the approach on that target subset rather than providing a production wrapper for arbitrary native code. Require an aligned, mapped region with a usable size that's a multiple of the stack alignment. Check the size rather than silently accepting an unwiped tail. Add guard pages and adequate stack probing for the platform, and don't let stack growth escape the allocation. Lock the pages and exclude them from dumps if that's part of your policy, checking whether those requests succeed. There's more to switching stacks than changing `sp`. Unwinding, sanitizers, profilers, and language runtimes may keep their own stack bookkeeping, so the custom stack needs to be integrated with them. Either support exceptions, cancellation, and nonlocal exits through this boundary, or prohibit them and enforce normal callback return. A `longjmp` that skips the cleanup defeats the design. And if a signal interrupts the work, its saved register context may land on a signal stack outside the region we're about to erase. `sigaltstack()` selects a signal-handler stack; it doesn't arrange ordinary function calls on your dedicated stack. ## Pay for cleanup at the right boundary So how much did this cost? On the M5 Max, I measured about 29 ns for a 64-byte AEGIS-128L encryption before wrapping. Clearing registers and a 64 KiB dedicated stack added about 965 ns in that run. Ouch. Cleanup took much longer than encryption. Doing it once after a handshake or a batch that shares the same secret lifetime may be a reasonable tradeoff. Choose the boundary from the lifetime you need, then benchmark that complete operation. But batching delays cleanup too. If an intermediate needs to disappear before the next record, spreading its cleanup cost across the whole batch changes that promise. ## When you need a stronger argument So far, I've checked particular binaries and run tests. Can we get a stronger argument than checking again after each compiler change? A compiler designed for cryptographic assembly can calculate cleanup after register allocation and stack layout, once it knows where the temporary storage is. The Jasmin work in [High-assurance zeroization](https://eprint.iacr.org/2023/1713) develops and proves an architectural stack-zeroization property, with register and flag cleanup discussed as part of the return boundary. Its scope is the operation's temporary machine state and preserved functional result. It doesn't erase caller-owned inputs, an earlier OS context save, or a snapshot taken while the key was live. The paper's discussion of composition with timing and speculative security also goes beyond its machine-checked development. The useful difference is that the cleanup accounts for the final storage locations. We still have to define where the operation ends and which data is supposed to survive it. For ordinary C, use the available compiler facilities or a reviewed execution wrapper, and document their limits. Keep explicit wipes for owned buffers that remain outside that boundary. [Part 3](/2026/10/06/zeroization-3/) covers those lifetimes, including the extra boundary WebAssembly introduces.