Zeroization, part 2: Clearing the stack and registers

In part 1, we saw that code intended to wipe secrets can end up creating extra copies.

So what can we do?

The best mitigation is to use a compiler designed to natively address that issue.

Jasmin is a perfect example of this. Jasmin is a high-level assembler, similar to Dan Bernstein’s qhasm, that makes writing assembly code very pleasant, if only because it can do automatic register allocation.

But Jasmin does a little bit more: the compiler is able to prove that values tagged as secrets are never returned as public values, are not involved in operations that wouldn’t run in constant time, and are safe against some classes of SPECTRE attacks. And for code using pointers, it’s even able to prove memory safety.

In addition to that, and this is where things become relevant to our topic, Jasmin can automatically wipe all the values tagged as secret before returning from a function, regardless of where they live (registers or stack memory).

Jasmin is an awesome tool for generating high-performance, high-assurance code.

And for a real-world example of Jasmin source code, check out the Jasmin implementations of AEGIS, which are also the fastest implementations of AEGIS to date.

Unfortunately, writing everything in Jasmin is not always practical, so we have to find mitigations for common compiled languages.

Decide where the operation ends

Instead of trying to solve the problem at the microscopic level, we can take a higher-level approach: find clear boundaries within which secrets are used, but beyond which they’re no longer used.

Suppose our application derives an ephemeral key, processes a batch of records, and then discards the key. We can put that whole sequence inside one boundary:

enter boundary (function)
    derive ephemeral key
    process records
    retire owned key and context buffers
leave boundary
    clear stack used by the work and its callees
    clear residual registers
return status

Notice where the key derivation happens: inside the boundary.

For practical purposes, a boundary will be a function, so if we derive the ephemeral key in the caller first, we may already have copies in that caller’s frame or preserved registers.

Let’s ask the compiler

GCC and Clang have little-known compilation flags to automatically clear registers before a function returns:

clang -O3 -fzero-call-used-regs=all -c crypto.c

The flag applies to the compilation unit, but with the zero_call_used_regs("all") attribute, we can apply this selectively to relevant functions instead:

struct request;

__attribute__((noinline, zero_call_used_regs("all")))
int secret_batch(struct request *request);

Now look at the end of the leaky example from part 1, compiled with that feature enabled. After calculating the return value, we get these instructions among the added overwrites:

mov     x8, #0
mov     x9, #0
movi.2d v0, #0000000000000000
movi.2d v1, #0000000000000000

There we go. x8 actually gets cleared this time. And the result in x0 survives, as it should.

This Godbolt comparison shows the same source with and without register cleanup.

This is a great improvement, but it’s not enough. It doesn’t discover spills in callees, and a spotless wrapper epilogue can leave a key schedule below it on the stack.

Now let’s ask GCC to handle the stack as well:

gcc -O2 -fstrub=all -fzero-call-used-regs=all -c crypto.c

Its strub mechanism tracks stack use with a watermark, then clears the recorded region after the call. GCC documents those interfaces here.

The strub contract is also a useful read.

What about our own register-wiping helper?

We saw compiler-specific flags. But why not write our own implementations in portable C, maybe just with a bit of assembly? Writing a few zeroing instructions sounds easy enough.

Unfortunately, there’s a trap: ABIs. Calling conventions that typically require functions to preserve a set of registers.

As an illustration, let’s look at what happens when an ordinary function tries to clear a register it must preserve for its caller. Here’s an abbreviated AArch64 sequence:

str     x19, [sp, #-16]!
mov     x19, #0
ldr     x19, [sp], #16
ret

There’s our mov x19, #0. Code we presumably explicitly added to clear the x19 register containing a secret. And right after it… ldr restores the old value.

Oops! The helper has undone its own wipe :)

The caller gets its register back, as the ABI requires, and the temporary stack save is still in memory too.

So, a wipe_registers() helper (which itself is trickier to write than it appears, especially for software supposed to be portable) can’t freely destroy all registers and remain an ordinary ABI-compliant call.

And putting the registers in an inline assembly clobber list doesn’t get us out of this either. It would just tell the compiler to protect live values, which can mean making more copies.

A boundary wrapper instead preserves the caller’s original state, runs the secret work, clears that work’s stack, and clears the residual register state it may discard.

What about our own stack-wiping helper?

How much stack should we wipe when leaving a boundary?

This is not easy to answer: adding up our local arrays isn’t enough. We need to include callees, compiler spills, alignment, saved registers, and ABI areas such as the x86_64 red zone. Yay.

We could measure how deep one call goes, but recursion, variable-sized frames, callbacks, and instrumentation can change it.

Compilers have flags that can help track the stack usage of functions, but in practice, this is a mess of hard-coded, hard-to-maintain, unreliable limits that may become invalid as soon as a single change is made to the code or compiler.

Giving the work its own stack

A different approach: let’s give an operation involving secrets a dedicated stack. We can run the sensitive sequence there, switch back, and clear the whole allocation.

We allocate the region, so we know where it starts and where it ends. We know that wiping the entire region is going to wipe everything stored on the stack inside the boundary.

Plus, the region can be surrounded by guard pages, marked for exclusion from swap and core dumps, etc. Keeping the stack that holds secrets in a different (and randomized) memory location than the regular application stack also makes a whole class of attacks more difficult.

I wrote C and Zig (secretstack) implementations. For reference, the C implementation has a simple entry point that looks like this:

long secure_call_on_stack(long (*fn)(void *), void *ctx,
                          void *stack, size_t stack_len);

The callback returns a status and receives its inputs through ctx. And this is what the wrapper does after the callback returns:

    mov     sp, x21
    mov     x9, x19
    b       2f
1:  stp     xzr, xzr, [x9], #16
2:  cmp     x9, x20
    b.lo    1b

x21 holds the wrapper’s original stack pointer. We’re switching back before overwriting the dedicated stack. Then the dedicated stack gets wiped: x19 holds the start of the region and x20 its end.

The beauty of this approach is that it works with virtually anything. Any existing library. Even if that library didn’t itself try to wipe the secrets it receives or creates, we end up with zero copies of these secrets in memory after the function returns.

Clear the registers in addition to this, and you’re in pretty good shape.

Unfortunately, there’s more to switching stacks than changing sp. Unwinding, sanitizers, profilers, and language runtimes may keep their own stack bookkeeping, so integration can be tricky.

Either support exceptions, cancellation, and nonlocal exits through boundaries, or (way easier) prohibit them and enforce normal callback return.

And of course, ban longjmp.

Another issue is that if a signal interrupts the work, its saved register context may land on a signal stack outside the region we’re about to erase, because sigaltstack() uses a signal-handler stack.

So, one has to remain careful about what runs inside a boundary. But these constraints can usually be easily satisfied for cryptographic code.

Overhead

Of course, clearing entire regions and tons of registers has a cost.

If the boundary only encrypts or hashes a short input, it’s likely that clearing registers and a dedicated stack can take longer than the operation itself.

However, doing it once after a handshake or a batch that shares the same secret lifetime is likely to have a negligible cost.

As always in security, everything is about tradeoffs.

Stay tuned for part 3 :)