A good hygiene when processing secrets is to wipe them after use.
And projects receives countless PRs about adding calls to zeroization functions for anything that looks like a secret.
This is a low-hanging fruit, something LLMs love to report, and it feels theoretically useful. But unfortunately, things are a little bit more complicated. Blindly zeroing secrets can do more harm than good.
Turns out that adding a wipe for a secret can leave more copies of a secret than the original code. Copies that wouldn’t have existed without it, and that are still there after it returns.
Code snippets and their compiled output can be verified in this Godbolt example.
Let’s start with memset()
Classic starting point: the good old memset() call that gets optimized out:
We calculate an intermediate value, use it to produce a result, then clear it before returning.
#include <stdint.h>
#include <string.h>
uint64_t ordinary_memset(uint64_t input)
{
uint64_t intermediate = input ^ UINT64_C(0x9e3779b97f4a7c15);
uint64_t result = (intermediate >> 17) ^
(intermediate * UINT64_C(0xd6e8feb86659fd93));
memset(&intermediate, 0, sizeof intermediate);
return result;
}
Now remove the memset() and compare the generated instructions. They are identical.
The compiler was smart enough to understand that wiping intermediate is a useless operation since after the function returns, it’s never read again.
But before we try to fix this, notice something else in the assembly output: the intermediate never went into memory at all. It stayed in registers.
Let’s make the stores survive
The most common practice to address the issue: replacing the memset() with volatile byte stores:
volatile unsigned char *bytes =
(volatile unsigned char *)&intermediate;
for (size_t i = 0; i < sizeof intermediate; i++) {
bytes[i] = 0;
}
Let’s do it and observe the arm64 output (I used XCode, shipping Clang 21):
eor x8, x0, x8
strb wzr, [sp, #8]
strb wzr, [sp, #7]
strb wzr, [sp, #6]
strb wzr, [sp, #5]
strb wzr, [sp, #4]
strb wzr, [sp, #3]
strb wzr, [sp, #2]
strb wzr, [sp, #1]
Eight strb instructions. wzr supplies zero.
Perfect, we defeated the compiler’s optimization passes and actually wrote eight zero bytes to stack memory!
Wait… where’s the instruction that put our intermediate into those stack slots?
Surprise: there isn’t one.
x8 holds the intermediate. We’ve just written eight zero bytes to locations that never received the value we wanted to erase.
The function then calculates the result from x8 and returns without clearing it.
The actual secret value is still around and we just wasted CPU cycles for nothing.
What if we hide the wipe from the compiler?
Let’s try another common approach: put the wiping code in a dedicated function, in a separately compiled file, without LTO.
You know, some variant of a secure_zero() function that can be found everywhere in cryptographic libraries:
#include <stddef.h>
void opaque_wipe(void *pointer, size_t length)
{
volatile unsigned char *bytes = pointer;
while (length != 0) {
*bytes++ = 0;
length--;
}
}
The caller only sees the declaration, so it can’t inspect the implementation when compiling this call:
opaque_wipe(&intermediate, sizeof intermediate);
Cool, so let’s look at the caller’s assembly output now.
Intestering: there’s a new instruction before the call:
str x8, [sp, #8]
Duh? We’re storing the intermediate on the stack now? The original function didn’t do that!
Yep, the wipe needs an address. And since the caller can’t see the implementation, it has to allow for the callee reading the old contents at that address. So it writes the value there first.
The wipe clears that new stack copy and leaves x8 alone.
We’ve introduced an interval during which the value exists in memory, paid to clear it, and still kept the register copy.
Also, a recognized write-only operation can skip this particular store, and if LTO is enabled (quite the norm these days), our attempt to hide the implementation in another file is likely to become useless.
Now let’s compare a tag
Now let’s do something very common but a little less trivial: compute a tag (HMAC output, etc.), compare it, wipe it, return the result.
In this example, we’re going to compute the tag as seed[i] ^ 0x5a. This is completely dumb and insecure, but easy to follow through the assembly.
We want to securely compare it against an application-provided tag.
#include <stddef.h>
void opaque_wipe(void *pointer, size_t length);
int compare_and_wipe(const unsigned char seed[16],
const unsigned char candidate[16])
{
unsigned char computed[16];
for (size_t i = 0; i < 16; i++) {
computed[i] = seed[i] ^ 0x5a;
}
unsigned int different = 0;
for (size_t i = 0; i < 16; i++) {
different |= (unsigned int)(computed[i] ^ candidate[i]);
}
int equal = different == 0;
opaque_wipe(computed, 16);
return equal;
}
Compile the caller and wipe separately, without LTO:
clang -std=c11 -O3 -Wall -Wextra -Werror -S compare.c -o compare.s
clang -std=c11 -O3 -Wall -Wextra -Werror -c wipe.c -o wipe.o
In the C code, the comparison appears before the wipe.
But let’s take a closer look at the disassembled output.
sub sp, sp, #80
stp x29, x30, [sp, #64]
add x29, sp, #64
ldr q0, [x0]
movi.16b v1, #90
eor.16b v0, v0, v1
str q0, [sp, #16]
stur q0, [x29, #-24]
ldr q0, [x1]
str q0, [sp]
sub x0, x29, #24
mov w1, #16
bl _opaque_wipe
ldp q1, q0, [sp]
cmeq.16b v0, v1, v0
Wait. Why are there two stores of q0 after the XOR?
And why is the comparison instruction, cmeq, after the call to the wipe?
To understand, let’s map the stack. We’re going to call the stack pointer after allocating the frame S; the frame pointer, x29, is S + 64.
| Stack bytes | Contents | Wiped? |
|---|---|---|
S through S + 15 |
Candidate | No |
S + 16 through S + 31 |
Extra computed-tag copy | No |
S + 40 through S + 55 |
Addressed computed object |
Yes |
Ah! The compiler loaded the comparison operands before the call, but postponed the comparison itself until afterward.
To keep the computed tag available across the call, it made a second copy in a spill slot.
Our wipe clears the requested 16 bytes perfectly, no problem here. And then ldp reloads the other copy, and cmeq compares it.
But nothing clears that spill slot before the function returns!
Now the fun part: what if we remove the wipe and compile again?
Surprise: the computed tag no longer gets stored on the stack at all. Everything stays in registers.
Adding the wipe created the copy we left behind.
Note that there’s no compiler bug here. The return value is correct, the addressed object is overwritten, and C doesn’t promise the security ordering you intended between those two events.
You can compare both functions in Godbolt.
There, the addressed object is at sp + 32 rather than sp + 40 (which I got with Xcode), but the extra copy at sp + 16 still survives the wipe.
Disabling stack protection generates different code, but this is more of an accidental unreliable side effect than a fix or a good idea.
So we’re sometimes paying for stores, and sometimes for a new call and stack copy, and the secrets are still around. Pretty frustrating for a change that was supposed to improve cleanup.
Registers shouldn’t be ignored
Registers are copied to memory after every context switch. Registers are visible in core dumps, hibernation files and VM snapshots. Registers can leak through microarchitectural bugs (Zenbleed).
And a single SIMD register can hold an entire 256-bit secret key, or even a 512-bit hash. They also tend to be recycled less frequently than general-purpose registers. But traditional zeroing functions are just designed to write zeros to memory, and have no clue about actual execution flows.
Keep the useful wipes
Does this mean we should stop wiping password buffers? No: those buffers already exist in memory, and clearing them when we’re done is useful. The same goes for an allocated key schedule or a context we’re retiring.
But in the examples above, the values we cared about were also somewhere else.
So before adding a wipe to a small local, inspect whether the value already lives in memory, whether taking its address creates storage, and what stays live across the call. Check the optimized build you ship, including LTO and hardening flags.
Or stop blindly sprinkling calls to buffer wiping functions. There are more reliable techniques to wipe secrets.
We’ll see that in part 2.