Zeroization, part 3: Design the lifetime, check the binary
We added a wipe in [part 1](/2026/10/06/zeroization-1/) and created an extra copy.
In [part 2](/2026/10/06/zeroization-2/), we worked out how to clear the stack and registers used by an operation.
Now let's work backward a little.
Why did we need some of those copies in the first place?
Sometimes the answer is right there in the API.
## Do we need that buffer?
Here's the 16-byte branch of the libaegis MAC verifier I tested, at commit `52c5cac`:
```c
uint8_t expected_mac[32];
switch (maclen) {
case 16:
implementation->state_mac_final(st_, expected_mac, maclen);
return aegis_verify_16(expected_mac, mac);
```
Generate the expected tag, compare it with the received tag, return the result.
Looks reasonable.
But look at `expected_mac`.
We only want to know whether the tags match, yet the backend function pointer requires a buffer to write the computed tag into.
Even if the backend could finish with the tag in registers, this interface asks it to produce an addressable object for a separate comparator.
We could wipe the buffer afterward, but we'd still have required it to exist.
Let's change the private interface instead:
```c
return implementation->state_mac_verify(st_, mac, maclen);
```
Now the backend can finalize its state, compare the extracted tag directly, and return just the verification result.
The public tag-generation API can stay unchanged for callers that actually need tag bytes.
I tried this in the actual library across all six standalone MAC variants and both tag lengths.
I moved the existing finalizer and comparator into one backend function, then checked the output.
The staging buffer was still there.
Why? The compiler kept part of finalization out of line.
The tag still had to pass through an output pointer; I'd moved that boundary inside the backend without removing it.
Separating the final rounds from the small tag-extraction step did remove that boundary in the inspected builds.
Finally.
Then I measured it, and the first version was slower on the wider variants.
The compiler had also changed which round helpers it inlined.
Restoring the earlier inlining increased code size.
Deleting an array from the source hadn't been enough to remove the storage.
Removing the storage hadn't been enough to improve performance.
We have to check the generated code and measure the complete operation.
The same reasoning applies to passing small keys by pointer.
If an out-of-line callee needs an address, the caller may have to store a value it already has in a register.
Passing by value can avoid that on a suitable ABI, but neither parameter style promises that the callee won't spill it.
So don't pick the pointer version merely because it looks easier to wipe.
Check the generated code and decide who owns the value and when they're done with it.
## Can the comparison erase the tag?
Here's another tempting idea.
We already XOR the tags to compare them, so let's write those differences back over the computed tag:
```c
unsigned int different = 0;
for (size_t i = 0; i < length; i++) {
computed[i] ^= candidate[i];
different |= computed[i];
}
return different == 0;
```
We've assigned to every byte.
Did that erase the old tag?
With a dead local array, I got exactly the same instructions as a normal XOR/OR equality reduction in the tested Clang builds.
The assignments bought us no extra erasure.
All right, suppose it's a caller-owned array and the changed contents remain observable.
Those stores survive.
Now look at what they leave after a mismatch:
```text
remaining = computed_before XOR candidate
computed_before = remaining XOR candidate
```
We know the received candidate.
XOR it with the remaining bytes and we get the computed tag back.
The buffer becomes zero on equality, but then the computed tag already equals the candidate we supplied.
Don't treat this as a wipe of an unknown expected tag after failed verification.
And don't infer anything about removal of the key or internal MAC state from what happens to the final tag.
Use a reviewed constant-time comparison.
Keep tag extraction close to that comparison where it avoids unnecessary storage, and inspect whether the compiler actually kept them together.
Then let the cleanup boundary handle temporary registers and spills.
## Let's clear the context in final()
The name sounds promising: surely `final()` is where we're done with the state?
Try resetting a reusable MAC context afterward.
It may deliberately retain initialized state for exactly that operation.
Unconditionally wiping it during `final()` or `verify()` can break the API, and a clone keeps another independent copy anyway.
So we need the point where the caller is actually done with the context, including cancellation and failed verification.
Document that point and put the explicit object wipe there:
```c
sodium_memzero(&state, sizeof state);
```
Now we're clearing an object we own and no longer need.
Another clone, an allocator's old copy, or a temporary made while processing it still needs separate attention.
This works alongside a surrounding execution boundary that clears the temporary computation state.
For a heap buffer, clear its owned sensitive extent before releasing it.
Don't use a shortened logical length if older bytes can remain in the unused portion, and don't write outside the allocation you own.
Avoid unnecessary growth and reallocation of secret-bearing buffers in the first place.
And keep the failure cleanup the caller relies on.
If authentication fails and we clear the plaintext output, the caller can observe those zeroes.
That's different from wiping a dead local array the program never reads again.
Streaming decryption needs particular care: finalization can't erase earlier plaintext chunks whose addresses it no longer has.
Keep unauthenticated output under the caller's control until verification succeeds, and define who discards or clears it on failure.
## Now let's try WebAssembly
Perhaps we can get a cleaner picture in WebAssembly.
But first, which memory are we clearing?
C's addressable stack objects live in linear memory.
The engine also has its own native stack and registers for running the module:
```text
C stack objects -> Wasm linear memory
Wasm locals -> engine-chosen registers or native spills
host-language copies -> separately managed storage
```
Our linear-memory wipe reaches the first row.
The other two are still there.
Let's take C out of the experiment entirely and explicitly clear a Wasm local.
This function computes the same synthetic result as part 1, then assigns zero to `$intermediate`:
```wat
(module
(func (export "local_clear") (param $input i64) (result i64)
(local $intermediate i64)
(local $result i64)
local.get $input
i64.const 0x9e3779b97f4a7c15
i64.xor
local.tee $intermediate
i64.const 17
i64.shr_u
local.get $intermediate
i64.const 0xd6e8feb86659fd93
i64.mul
i64.xor
local.set $result
i64.const 0
local.set $intermediate
local.get $result))
```
Check the Wasm file: the last assignment is there.
This time the zero hasn't disappeared during compilation to Wasm.
But there's another compiler involved.
I tested this on arm64 with Wasmtime 49.0.2 using Cranelift at optimization level 2, and with the installed Wasmer 7.4.2 LLVM build.
Both produced identical native function bodies with and without the local clear.
That Wasmer build reports a `-modified` version marker; these results describe that binary.
Let's look at the relevant Wasmer output, with constant loads omitted:
```asm
eor x8, x1, x8
mul x9, x8, x9
eor x0, x9, x8, lsr #17
ret
```
There it is again: the intermediate in `x8`, still present at `ret`.
The zero survived into the Wasm file, then disappeared when the engine compiled it to native code.
Assigning zero to an abstract Wasm local doesn't ask the engine to erase the physical register that held its old value.
That agrees with the [WebAssembly instruction semantics](https://webassembly.github.io/spec/core/exec/instructions.html#exec-local-set).
Now bring back our C tag example.
In Wasmtime's inspected SIMD128 output, the engine stores the computed tag in linear memory and also spills it to native `[sp, #0x10]`.
It wipes the linear-memory object, reloads the native spill, and finishes the comparison.
The function leaves that spill slot uncleared.
Same problem, another stack.
There's a useful distinction here, though.
Under the sandbox's intended isolation, a Wasm guest can't read that native stack.
If our observer can only read linear memory, clearing the owned region still removes bytes they could have read.
If they can inspect the whole engine process, that wipe leaves more outside its coverage.
Changing `__stack_pointer` or clearing the module's entire C stack won't switch or erase the engine stack.
For that, we need runtime cooperation with an explicit cleanup contract.
And there's still the host.
A JavaScript array that received a copy earlier won't change when we clear the Wasm buffer.
Dropping the instance doesn't provide a portable promise that its storage was overwritten either.
## What if we add a barrier?
One more familiar tool: an empty assembly barrier can help keep an object wipe in the compiled program.
```c
__asm__ __volatile__("" : : "r"(pointer) : "memory");
```
Look for its barrier instruction in the assembly.
There isn't one: the template is empty.
Its constraints can still change the surrounding stores and reloads, but they don't clear registers or provide a general speculative-execution defense.
If we need the cleanup boundary to hold against transient execution, we have more work to do on the actual processor.
A scrub loop can need protection against a transiently predicted early exit.
Later accesses can raise separate store-bypass questions.
The right mechanism depends on the architecture and execution path; appending a generic C atomic fence to every wipe doesn't establish either property.
And no end-of-operation cleanup can undo disclosure while the secret was live.
Keep constant-time code, platform mitigations, and isolation where they're required.
An earlier crash dump, snapshot, or saved OS context also has its own lifetime.
## What to check before shipping
After all these attempts, we should be suspicious of a test that only checks whether our named buffer contains zeroes.
Here's what to check before shipping:
1. **Check the owned bytes.** Verify the whole promised range, guard bytes, and supported zero-length behavior.
Cover success, failure, and retirement after reuse or cloning.
2. **Inspect the linked release build.** Follow sensitive values through the operation and its callees.
Look for address-taking, call-site spills, delayed comparisons, preserved-register saves, and what runs after the cleanup.
3. **Check the boundary's extent.** Establish the stack bound or validate the dedicated region.
Review ABI register coverage and every permitted exit path.
4. **Repeat for supported compilation paths.** Include LTO, stack protection, instrumentation, target features, and compiler upgrades.
For Wasm, inspect both the module and the chosen engines' native output, including scalar and SIMD builds.
5. **Measure the complete operation.** Include short inputs, long inputs, failure paths, throughput, and code size.
A result measured under Rosetta isn't a native x86 performance result.
A test that reads back zeroes proves something about that buffer.
A scan that finds no known secret words proves less than complete erasure, since another representation can evade it.
Use those checks to catch regressions alongside the generated-code review.
Clear the buffers we own when we're done with them.
Avoid interfaces that force pointless copies, and clear the stack and registers used for temporary computation state.
Document where that cleanup ends.
Then check the binary.
We've seen how easily a perfectly sensible source change can do something else entirely.