Welcome to part 3 on zeroization! We added a wipe in part 1 and created an extra copy. And in part 2, we looked at some options to clear the stack and registers.
But why did we need some of those copies in the first place? Turns out that sometimes the answer is right there in the API.
Do we need that buffer?
Let’s go back to the textbook example: MAC verification.
We initialize the state of a MAC function (could be HMAC, an AEAD, etc.), update the state with a message to authenticate, finalize the state, compute the expected tag, compare it with the received tag, and return whether they match.
In the toy example below, the state is just two 128-bit blocks, and to produce a tag, we just XOR them to produce a 128-bit tag.
#include <stddef.h>
#include <stdint.h>
typedef struct {
uint8_t a[16];
uint8_t b[16];
} final_state;
__attribute__((noinline))
void extract_tag(const final_state *state, uint8_t out[16])
{
for (size_t i = 0; i < 16; i++) {
out[i] = state->a[i] ^ state->b[i];
}
}
static int equal16(const uint8_t a[16], const uint8_t b[16])
{
unsigned int different = 0;
for (size_t i = 0; i < 16; i++) {
different |= a[i] ^ b[i];
}
return different == 0;
}
int verify_tag(const final_state *state, const uint8_t candidate[16])
{
uint8_t expected[16];
extract_tag(state, expected);
return equal16(expected, candidate);
}
The noinline attribute deliberately prevents the extraction from being inlined, which is what a compiler would do in a real implementation if the extraction is attached to a substantial finalization function.
Look at expected.
Even if extract_tag() computes the tag entirely in registers, its interface requires it to write those bytes through an output pointer. And the caller has to provide storage for them.
We could wipe that storage afterward, but we’d still have required it to exist.
Let’s do something else, then: fuse the computation and the check of the tag in the same function:
int verify_fused(const final_state *state,
const uint8_t candidate[16])
{
unsigned int different = 0;
for (size_t i = 0; i < 16; i++) {
different |= state->a[i] ^ state->b[i] ^ candidate[i];
}
return different == 0;
}
There’s no output pointer for an intermediate tag here. Each extracted byte goes straight into the comparison, and the only result is a boolean.
The state may still be sensitive information, but the expected tag is not present in stack memory. There’s no hard promise that a compiler wouldn’t spill a temporary, but it is very unlikely as it would make absolutely no sense from a performance perspective.
So, simply combining finalization and verification can help mitigate leakage.
The same reasoning applies to passing small keys by pointer. If an out-of-line callee needs an address, the caller may have to store a value it already has in a register.
Passing by value can avoid that on a suitable ABI, even though there’s still no hard guarantee that the callee won’t spill it.
So don’t pick the pointer version merely because it looks easier to wipe. Check the generated code and decide who owns the value and when they’re done with it.
Can the comparison erase the tag?
Here’s another tempting idea.
We already XOR the tags to compare them, so why not overwrite the computed tag as we compare it with the candidate?
unsigned int different = 0;
for (size_t i = 0; i < length; i++) {
computed[i] ^= candidate[i];
different |= computed[i];
}
return different == 0;
Did that erase the computed tag?
Looking at the source code, it should. But compilers are sneaky and in a bunch of test cases, I got exactly the same instructions as a normal XOR/OR equality reduction.
All right, suppose it’s a caller-owned array and the changed contents remain observable.
Good news: in that case, those stores survive. So, when we have a match, the buffer holding the computed tag is all zero.
Cool, but how about when we have a mismatch? The candidate is public; the remaining bytes are not all zero, but are still in memory, so recovering the correct, computed tag is trivial:
remaining = computed_before XOR candidate
computed_before = remaining XOR candidate
A dumb security scanner grepping memory for leftover secrets would report that everything’s fine, while the secrets are trivial to recover.
So, don’t try to be creative. Keep tag extraction close to that comparison where it avoids unnecessary storage, and inspect whether the compiler actually kept them together.
WebAssembly makes this worse
In native code, the boundary we described in part 2 could clear the registers and stack region used by an operation.
But WebAssembly is a different beast. WebAssembly code doesn’t control the machine registers or the engine’s native stack. The engine decides how its values are represented there.
To the VM, a key is just an i64, a v128, or some bytes in memory.
And WebAssembly doesn’t define any ways to tell the engine to track sensitive copies and erase them. A key gets the same treatment as any other value of its type. Proposals to improve that situation have been ignored and eventually abandoned.
In a JIT-based VM, that leaves another optimizer free to introduce copies, spill values, move calculations across a wipe, or discard an unused assignment of zero. The VM only cares about observable behavior.
The application neither chooses those copies’ locations nor gets a way to clear them.
WebAssembly itself is useless without external, imported functions provided by the runtime. Think about them like syscalls in a native environment.
A WebAssembly application passing a secret to an external function implemented by the runtime has no control over it. The runtime can make as many copies as needed, and from within the WebAssembly sandbox, none of these copies is accessible and can be cleared.
Clearing a Wasm local doesn’t clear a register
Inspecting a WebAssembly module is not enough to predict whether secrets will be cleared, even if the disassembly of the module says so.
Here’s an example of a module that XORs an input with a constant, uses that intermediate in a calculation, then assigns zero to $intermediate:
(module
(func (export "local_clear") (param $input i64) (result i64)
(local $intermediate i64)
(local $result i64)
local.get $input
i64.const 0x9e3779b97f4a7c15
i64.xor
local.tee $intermediate
i64.const 17
i64.shr_u
local.get $intermediate
i64.const 0xd6e8feb86659fd93
i64.mul
i64.xor
local.set $result
i64.const 0
local.set $intermediate
local.get $result))
We were able to defeat the C/Rust/Zig and LLVM optimizations, and we got a local.set $intermediate that replaces the secret with 0. Perfect!
But here comes the second compiler, the one that will compile the WebAssembly code to native code so it can actually run.
I tested this with Wasmtime and Wasmer, using Cranelift and LLVM: both produced identical native function bodies with and without the local clear.
Let’s look at the relevant Wasmer output, with constant loads omitted:
eor x8, x1, x8
mul x9, x8, x9
eor x0, x9, x8, lsr #17
ret
x8 contains the secret. It’s not cleared, in spite of all our efforts and what the WebAssembly bytecode says.
The wipe can create a native spill too
Let’s see another example:
__attribute__((noinline))
int check_and_wipe(const uint8_t seed[16],
const uint8_t candidate[16])
{
uint8_t expected[16];
for (size_t i = 0; i < 16; i++) {
expected[i] = seed[i] ^ 0x5a;
}
int same = equal16(expected, candidate);
secure_zero(expected, sizeof expected);
return same;
}
The source compares first, then wipes.
But compiling this with LLVM 23 shows that the compiled Wasm keeps the computed tag and candidate in locals, calls the wipe, and only then finishes the comparison.
This is an acceptable thing to do for a compiler: it doesn’t change the function’s result or the bytes the wipe writes.
But what does the engine do with those locals across the call?
With Wasmtime, this is what I saw (relevant instructions only):
str q20, [x0, w4, uxtw]
stur q20, [sp, #0x10]
bl #0x380
ldur q20, [sp, #0x10]
x0 is the base of linear memory and w4 is the expected offset within it.
The first store writes the computed tag there. The second writes the same tag to the engine’s native stack.
The call at 0x380 clears the linear-memory object, and the final load brings back the untouched native copy for comparison.
The complete function returns without clearing that spill slot.
expected is cleared, but the engine still has another copy, at an address the module can’t wipe.
Even a module without memory uses memory
Let’s remove linear memory completely.
No C stack, no buffers, no imports, and no load or store instructions.
This complete module just computes x XOR (4 * (x + 1)), with arithmetic modulo 2^64:
(module
(func $triple (param $x i64) (result i64)
local.get $x
i64.const 3
i64.mul)
(func $combine (param $x i64) (result i64)
local.get $x
local.get $x
call $triple
i64.add)
(func $across (export "across") (param $x i64) (result i64)
local.get $x
local.get $x
i64.const 1
i64.add
call $combine
i64.xor))
If x were sensitive, where would it go?
Let’s try with Cranelift: across() keeps its original input in x19 while calling combine().
But combine() also wants to use x19, so it saves the caller’s value on the native stack:
str x19, [sp, #-0x10]!
mov x19, x4
bl #0
add x2, x19, x2
ldr x19, [sp], #0x10
Here, bl #0 calls triple(). The final load restores x19, but leaves the input on the native stack after both functions return.
The wasm module can’t do anything to wipe it.
A note on “volatile”
The volatile attribute is understood by source-to-wasm compilers and can keep a buffer wipe in the generated WebAssembly module.
But the volatile attribute doesn’t exist in the specification; from a WebAssembly runtime perspective, clearing memory with or without volatile makes no difference: the code is exactly the same, so it will be optimized out the same way.
By the way, in WebAssembly, memory barriers don’t do anything either. So don’t expect them to mitigate speculative execution or help mitigate side channels.
So what can we do?
Zeroing memory using compiled languages and/or virtual machines is pretty much impossible to do correctly.
My recommendations for library authors:
- Inspect the generated code. This is the source of truth, not the source code.
- Wipe secrets that will almost certainly be stored in memory, such as passwords, hex/base64-encoded keys, etc.
- Always wipe secrets present on the heap
- Use protected memory when possible. For example, the libsodium guarded memory functions are quite useful. If only to find bugs in implementations.
- For everything else, small values that are actually used to perform computations, don’t bother.
My recommendations for library users:
- Use encrypted swap and encrypted filesystems, and configure your OS and hardware correctly.
- Apply the recommendations from part 2 if you’re really worried about dangling secrets in memory and registers. Write libraries like zig-secretstack for your favorite programming language that are easy to use.