In 2025 we published Scoop, an optimizer for training side-channel models against masked implementations, and with it the first attack on the ASCADv2 dataset that didn’t rely on knowing the masks during profiling. The attack passed every check we ran. It was wrong: the model had not learned to read a single power trace. This post is about how that happens, why the standard metrics can’t see it, and what to check instead.
A key out of nowhere
ASCADv2 is a public set of power measurements from an AES implementation protected by affine masking and by shuffling the order in which the 16 S-box lookups run. Breaking this dataset in a non-worst-case (black box) setting remains, to this day, an open problem.
The extracted dataset has 500,000 profiling traces and 15,000 attack traces, each 15,000 samples long. Because the perceived information we managed to get is so small, we reallocated the traces to get a larger attack set out of the profiling traces: 200,000 to train on, the last 300,000 to attack with.
Our model was a plain multi-layer perceptron with one hidden layer and about 231 million parameters. Surprisingly, without any hyperparameter optimization, the loss converged. The guessing entropy dropped to 1 after roughly 250,000 attack traces: the correct key ranked first.
Then we zeroed every weight in the model’s first layer. That cuts the traces out entirely: whatever you feed in, the rest of the network sees the same constant. The guessing entropy still went to 1.
This paradox is concerning because at no point had the model seen the traces during the GE evaluation. It must have learned something about the labels: the distribution of the targeted byte in ASCADv2 isn’t uniform, and the model had memorised its shape. Nothing in the training curves or the key-rank curves flagged it.
What an attack actually scores
A profiled attack has two phases. First, on a device you control, you record traces and train a model to predict a secret-dependent intermediate value from a trace: here, one byte of the AES S-box output. Then, on the target device, you record new traces, run them through the model, and ask which key guess best explains what the model says.
In symbols: the model maps a trace to a probability vector over the possible values of the intermediate variable , meant to approximate the posterior . Each attack trace comes with a known plaintext , so every key guess predicts an intermediate value . The usual way to score a guess is to add up how much log-probability the model gave to the values that guess predicts. The paper calls this the posterior distinguisher:
The attack’s answer is the highest-scoring guess. To judge the attack, you sort the guesses and read off where the true key landed, its rank ; the guessing entropy is that rank averaged over attack sets of traces:
A GE curve that falls to 1 as grows is how papers in this field claim a key byte is broken. With 256 possible keys, a GE stuck around 128 means you’re guessing.
The prior gets a vote
Here is the intuition. A classifier trained to minimise log-loss learns two things at once: what the trace says about the value, and how common each value is in the first place. If one value shows up 70% of the time, a model that has learned nothing else will still bet on it. That’s the right behaviour for a classifier. It’s the wrong behaviour for a key-recovery score, because the score can’t tell which of the two kinds of knowledge it is spending.
Bayes’ rule makes the split explicit:
The first factor, the likelihood, is the leakage: how well value explains the trace. The second is the prior. The best score you could build from the measurements alone uses only the likelihood. That’s the maximum-likelihood distinguisher:
Now suppose the model is perfect, . Substitute Bayes’ rule into the posterior distinguisher and take logs, and it splits into three sums:
The last sum doesn’t depend on , so it can’t change the ranking. The middle one is the problem. If is uniform, every value has prior , the prior term is the same for all keys, and the posterior distinguisher ranks keys exactly like the optimal one. That’s the paper’s Lemma 1, and it’s why nobody worried: the textbook datasets are built to be uniform.
If is not uniform, the prior term varies with , and it favours the right key. For the right key, is the value that was really computed, so it follows the skewed distribution and keeps landing where the prior is high. A wrong key passes the same plaintexts through the S-box under a different key, which relabels the values: the high-probability ones get mapped somewhere else. Per trace, the right key’s prior term averages ; a wrong key’s averages minus a cross-entropy, which by Gibbs’ inequality is never above , and falls strictly below it once the relabelling moves the heavy values. So the prior term alone, summed over enough traces, picks out — no leakage needed.
This is worst exactly where the stakes are highest. Against a well-masked implementation the real leakage is faint and hides in higher-order statistics, so for the optimizer, memorising the label distribution is the easy way to lower the loss.
A model that never looked
The paper’s toy example strips this to the bone. Let be a single bit with , so bits, and let the traces be independent of : there is nothing to learn from them. The best a model can do is output the prior, which gives it a log-loss of
Against the 1 bit you’d expect from a uniform label, that loss looks like 0.12 bits of information learned. And if the attack set has the same 70/30 skew, the guessing entropy converges. A model that never looked at a trace passes both tests.
Try it
The demo below runs the same experiment on a full byte, live in your browser. The
intermediate value is ; its distribution is a truncated Gaussian centred on 127, the family
the paper uses to bias ASCADv1. --prior-skew narrows it; the tick marks sit on the paper’s
three bias levels, , and . --leakage sets how much each trace says about
(its Hamming weight plus Gaussian noise); at zero, traces are pure noise. The model is the
ideal one for this setup: it knows both the leakage and the prior. Each curve is the mean rank of
the right key over many simulated attack sets.
Static render of the default setting (100 simulated attack sets at build time). With JavaScript on, it re-runs live in your browser as you move the sliders.
It opens at the medium bias with no leakage. The posterior curve falls to rank 1 within a hundred
or so traces, from measurements that carry no information at all. The AOD curve, introduced in
the next section, stays on the random-guess line. Pull --prior-skew to 0 and the two curves
collapse onto each other at 128.5: that’s Lemma 1. Push it to the right and the posterior converges
in a handful of traces.
Now give it some leakage. Both curves fall, and the AOD one is the honest one. The gap between them is the part of the “attack” the prior paid for.
Divide the prior back out
If the posterior is likelihood times prior, the fix is to divide the prior back out before scoring. That’s the paper’s asymptotically optimal distinguisher (AOD):
Here is the empirical frequency of each value in the attack set. Dividing the posterior by the prior leaves the likelihood over , and that denominator is the same for every key. The paper shows that if the model converges to the ideal one, the expected AOD score ranks keys like : only the leakage counts.
It’s cheap. Counting value frequencies is linear in the number of traces, and the rest is the GE computation you already run: overall. In the demo, the AOD divides by the exact prior rather than counted frequencies; with a sharp prior and a few hundred traces, most values never appear in the attack set and would get a frequency of zero.
One caveat: the AOD doesn’t tell you there was a false positive. It just declines to have one. To see the problem you compare the AOD and posterior curves, or use the checks below.
On ASCADv1, where models are known to learn real leakage, the posterior and AOD guessing entropies come out similar. On our ASCADv2 model, the AOD guessing entropy never converges.
A quieter trap: averaging the same traces
The paper flags a second issue with how GE is estimated. You never have infinitely many attack sets, so you average the rank over subsets of the one attack set you recorded:
If each subset is as large as the whole attack set, every subset is the attack set, and the “average” is a single rank measured times. A single rank can hit 1 by luck when the true GE wouldn’t. That’s what happens when an attack needs all the traces you have, which was the case for the ASCADv2 attack. Empirically, the estimate starts behaving like a single rank once subsets pass about 30% of the attack set; the paper’s conservative rule of thumb is to keep them at most 10%.
Checks before training
Two checks run before you commit to a model.
A χ² test on the labels. Count how often each value occurs and test against uniform:
Compare it with the χ² distribution at degrees of freedom and a significance level such as . It costs . On ASCADv2, byte 4, the one the Scoop attack targeted, fails the test; byte 6 passes. Plot the two histograms side by side, even sorted, and you can’t tell which is which, so run the test rather than eyeballing. Its limit: it tells you there’s a risk, not whether a model will use the bias, nor whether the attack set shares it.
A null benchmark. Replace every trace, profiling and attack, with zeros and train and evaluate as usual. This is the toy example run on your real labels. If the GE still converges, the labels alone carry enough bias to break the key, and the attack set shares it. Your real model then has to clearly beat this baseline to mean anything. It costs a full training run, for epochs and parameters, and like the χ² test it says nothing about leakage, only about exposure to the bias.
Checks after training
Once a model looks successful, four checks ask whether it actually uses its input. Each needs a trained model first, which is the expensive part.
Activation probing. Run the attack set through and look at the hidden activations. A model that has baked the prior into its weights reacts the same way to every trace. In our ASCADv2 model, only a few hidden units ever fired, and always the same ones, whatever the input. With 16 shuffled S-box lookups in the trace you’d expect much more variety. Cost: one forward pass, . Suggestive, not conclusive.
Gradient visualisation. Take the gradient of the loss with respect to the input trace; its peaks show which samples the prediction depends on. On ASCADv1 the peaks line up with where an ANOVA says the masking shares leak. On ASCADv2 it’s noise. The catch: confirming the peaks needs an expert, and, on a masked target, access to the shares. It’s also prone to confirmation bias, and a valid model can have a noisy gradient. A clean match is good evidence; a noisy one proves nothing.
Model ablation. Set the input layer’s weights to zero, , and recompute the GE. If it still converges, the predictions come from the network’s own parameters and not from the traces: a false positive, no argument. If it falls back to random guessing, the model needed its input. That doesn’t make it a good model, only an honest one. This is the check that caught ASCADv2, and it costs one inference plus a GE, .
Prediction entropy. Compute the entropy of each output vector, . A model that outputs the prior gives nearly the same vector for every trace, so the entropy sits at one value with almost no spread. A model reading leakage is confident on clean traces and unsure on noisy ones. The two Scoop models side by side:
| ASCADv1 | 2.8269 | 0.4818 | 4.2202 |
| ASCADv2 | 3.5686 | 0.0003 | 0.0265 |
Both targets are 8-bit labels assumed uniform, so the two models should look alike. The ASCADv2 model’s predictions barely move from trace to trace, and its top guess is almost always the same value. That’s a model with one answer.
Testing the tests
ASCADv2 is so far the only public dataset known to have this problem, so the paper builds some. It resamples ASCADv1 so that its labels follow a truncated Gaussian with and , or . Those are the low, medium and high bias levels: the information a model gains from the prior alone is respectively below, about equal to, and above what the best known model extracts from the leakage. The bias goes either into the profiling set only, or into both the profiling and attack sets.
What they found:
- χ² is jumpy. It flags all three biased sets, and also the original ASCADv1 profiling set. It picks up even slight sampling imbalance.
- The null benchmark is selective. It raises a flag only when the same bias sits in both the profiling and attack sets, which is exactly when a false positive is possible.
- The training curves can give it away. At high bias, with the bias in both sets, the perceived information ends up well above that of the best known model. That can’t come from leakage.
- Every post-mortem check catches the high bias. Only ablation catches the medium one. At medium bias the model learns the prior and the leakage, and ablation is the only check that separates the two.
- The AOD changes the verdict only when it should. At low bias the posterior and AOD curves agree; at medium both degrade slightly and stay close; at high bias, with the bias in both sets, the posterior GE converges quickly and the AOD GE doesn’t. With the bias in the profiling set only, the two agree at every level: if the attack set doesn’t share the bias, the model can’t cash it in.
No single check wins. The cheap ones can’t see leakage; the informative ones need a trained model and sometimes an expert. The paper’s advice is to combine them, and to treat the AOD as the default scoring rule rather than an optional extra.
Takeaways
- A GE curve reaching 1 means the score ranked the key first. It doesn’t mean the model read the traces.
- The posterior distinguisher is only optimal when the labels are uniform. When they aren’t, it adds a prior term that favours the right key on its own.
- Score with the AOD by default. It costs almost nothing.
- Run a χ² test on your labels, and zero the input layer before you claim a break.
- Report those checks next to the GE. A reader can’t rerun them for you.