· 11 min read · companion to cascade-2026

Is it really broken?

How the way the DL-SCA community computes guessing entropy leaves the door open to false positives. And an introduction to a robust distinguisher.

In 2025 we published Scoop, an optimizer for training side-channel models against masked implementations, and with it the first attack on the ASCADv2 dataset that didn’t rely on knowing the masks during profiling. The attack passed every check we ran. It was wrong: the model had not learned to read a single power trace. This post is about how that happens, why the standard metrics can’t see it, and what to check instead.

A key out of nowhere

ASCADv2 is a public set of power measurements from an AES implementation protected by affine masking and by shuffling the order in which the 16 S-box lookups run. Breaking this dataset in a non-worst-case (black box) setting remains, to this day, an open problem.

The extracted dataset has 500,000 profiling traces and 15,000 attack traces, each 15,000 samples long. Because the perceived information we managed to get is so small, we reallocated the traces to get a larger attack set out of the profiling traces: 200,000 to train on, the last 300,000 to attack with.

Our model was a plain multi-layer perceptron with one hidden layer and about 231 million parameters. Surprisingly, without any hyperparameter optimization, the loss converged. The guessing entropy dropped to 1 after roughly 250,000 attack traces: the correct key ranked first.

Then we zeroed every weight in the model’s first layer. That cuts the traces out entirely: whatever you feed in, the rest of the network sees the same constant. The guessing entropy still went to 1.

This paradox is concerning because at no point had the model seen the traces during the GE evaluation. It must have learned something about the labels: the distribution of the targeted byte in ASCADv2 isn’t uniform, and the model had memorised its shape. Nothing in the training curves or the key-rank curves flagged it.

Zeroing the first layer. The published model's guessing entropy falls to rank 1 (dashed). With the first layer zeroed, no trace reaches the hidden layer, and the guessing entropy still falls far below random guessing (green). Curves from the paper's figures.

What an attack actually scores

A profiled attack has two phases. First, on a device you control, you record traces and train a model to predict a secret-dependent intermediate value from a trace: here, one byte of the AES S-box output. Then, on the target device, you record new traces, run them through the model, and ask which key guess best explains what the model says.

In symbols: the model F\sym{F}{F} maps a trace xi\sym{x}{\mathbf{x}_i} to a probability vector yi\sym{yi}{\mathbf{y}_i} over the possible values s∈Z\sym{zval}{s} \in \sym{Zset}{\mathcal{Z}} of the intermediate variable Z\sym{Z}{Z}, meant to approximate the posterior Pr(Z=s∣X=xi)\sym{post}{\mathrm{Pr}(Z = \sym{zval}{s} \mid \mathbf{X} = \sym{x}{\mathbf{x}_i})}. Each attack trace comes with a known plaintext pi\sym{pt}{p_i}, so every key guess k\sym{k}{k} predicts an intermediate value zi,k=f(pi,k)\sym{zik}{z_{i,k}} = \sym{f}{f}(\sym{pt}{p_i}, \sym{k}{k}). The usual way to score a guess is to add up how much log-probability the model gave to the values that guess predicts. The paper calls this the posterior distinguisher:

dSa[k]≜∑i=1Nalog⁡yi[zi,k]\sym{dpost}{\mathsf{\textbf{d}}_{S_a}[k]} \triangleq \sum_{i=1}^{\sym{Na}{N_a}} \log \sym{yi}{\mathbf{y}_i}\left[\sym{zik}{z_{i,k}}\right]

The attack’s answer is the highest-scoring guess. To judge the attack, you sort the guesses and read off where the true key k⋆\sym{kstar}{k^{\star}} landed, its rank gSa(k⋆)\sym{rank}{g_{S_a}(k^{\star})}; the guessing entropy is that rank averaged over attack sets Sa\sym{Sa}{S_a} of Na\sym{Na}{N_a} traces:

k~⋆=argmax⁡k∈KdSa[k]gSa(k⋆)≜1+∑k∈K∖{k⋆}1[dSa[k]≥dSa[k⋆]]GE(Na)≜ESa[gSa(k⋆)]\begin{aligned} \sym{khat}{\tilde{k}^{\star}} &= \operatorname*{argmax}_{\sym{k}{k} \in \sym{K}{\mathcal{K}}} \sym{dpost}{\mathsf{\textbf{d}}_{S_a}[k]} \\ \sym{rank}{g_{S_a}(k^{\star})} &\triangleq 1 + \sum_{\sym{k}{k} \in \sym{K}{\mathcal{K}} \setminus \{\sym{kstar}{k^{\star}}\}} \sym{ind}{\mathbf{1}}\left[\sym{dpost}{\mathsf{\textbf{d}}_{S_a}[k]} \geq \sym{dpost}{\mathsf{\textbf{d}}_{S_a}[k^{\star}]}\right] \\ \sym{GE}{\mathrm{GE}}(\sym{Na}{N_a}) &\triangleq \mathbb{E}_{\sym{Sa}{S_a}}\left[\sym{rank}{g_{S_a}(k^{\star})}\right] \end{aligned}

A GE curve that falls to 1 as Na\sym{Na}{N_a} grows is how papers in this field claim a key byte is broken. With 256 possible keys, a GE stuck around 128 means you’re guessing.

The prior gets a vote

Here is the intuition. A classifier trained to minimise log-loss learns two things at once: what the trace says about the value, and how common each value is in the first place. If one value shows up 70% of the time, a model that has learned nothing else will still bet on it. That’s the right behaviour for a classifier. It’s the wrong behaviour for a key-recovery score, because the score can’t tell which of the two kinds of knowledge it is spending.

Bayes’ rule makes the split explicit:

Pr(Z=s∣X=xi)=Pr(X=xi∣Z=s) Pr(Z=s)Pr(X=xi)\sym{post}{\mathrm{Pr}(Z = \sym{zval}{s} \mid \mathbf{X} = \sym{x}{\mathbf{x}_i})} = \frac{\sym{lik}{\mathrm{Pr}(\mathbf{X} = \sym{x}{\mathbf{x}_i} \mid Z = \sym{zval}{s})}\,\sym{prior}{\mathrm{Pr}(Z = \sym{zval}{s})}}{\sym{evid}{\mathrm{Pr}(\mathbf{X} = \sym{x}{\mathbf{x}_i})}}

The first factor, the likelihood, is the leakage: how well value s\sym{zval}{s} explains the trace. The second is the prior. The best score you could build from the measurements alone uses only the likelihood. That’s the maximum-likelihood distinguisher:

dSa(opt)[k]≜∏i=1NaPr(X=xi∣Z=zi,k)\sym{dopt}{\mathsf{\textbf{d}}^{(\mathrm{opt})}_{S_a}[k]} \triangleq \prod_{i=1}^{\sym{Na}{N_a}} \sym{lik}{\mathrm{Pr}\left(\mathbf{X} = \sym{x}{\mathbf{x}_i} \mid Z = \sym{zik}{z_{i,k}}\right)}

Now suppose the model is perfect, yi[s]=Pr(Z=s∣X=xi)\sym{yi}{\mathbf{y}_i}[\sym{zval}{s}] = \sym{post}{\mathrm{Pr}(Z = \sym{zval}{s} \mid \mathbf{X} = \sym{x}{\mathbf{x}_i})}. Substitute Bayes’ rule into the posterior distinguisher and take logs, and it splits into three sums:

dSa[k]=∑i=1Nalog⁡Pr(X=xi∣Z=zi,k)⏟log⁡dSa(opt)[k]+∑i=1Nalog⁡Pr(Z=zi,k)⏟prior term  −  ∑i=1Nalog⁡Pr(X=xi)⏟same for every k\begin{aligned} \sym{dpost}{\mathsf{\textbf{d}}_{S_a}[k]} = {} & \underbrace{\sum_{i=1}^{\sym{Na}{N_a}} \log \sym{lik}{\mathrm{Pr}\left(\mathbf{X} = \sym{x}{\mathbf{x}_i} \mid Z = \sym{zik}{z_{i,k}}\right)}}_{\log \sym{dopt}{\mathsf{\textbf{d}}^{(\mathrm{opt})}_{S_a}[k]}} \\ & + \underbrace{\sum_{i=1}^{\sym{Na}{N_a}} \log \sym{prior}{\mathrm{Pr}\left(Z = \sym{zik}{z_{i,k}}\right)}}_{\text{prior term}} \;-\; \underbrace{\sum_{i=1}^{\sym{Na}{N_a}} \log \sym{evid}{\mathrm{Pr}\left(\mathbf{X} = \sym{x}{\mathbf{x}_i}\right)}}_{\text{same for every } \sym{k}{k}} \end{aligned}

The last sum doesn’t depend on k\sym{k}{k}, so it can’t change the ranking. The middle one is the problem. If Z\sym{Z}{Z} is uniform, every value has prior 1/∣Z∣1/|\sym{Zset}{\mathcal{Z}}|, the prior term is the same for all keys, and the posterior distinguisher ranks keys exactly like the optimal one. That’s the paper’s Lemma 1, and it’s why nobody worried: the textbook datasets are built to be uniform.

If Z\sym{Z}{Z} is not uniform, the prior term varies with k\sym{k}{k}, and it favours the right key. For the right key, zi,k⋆\sym{zik}{z_{i,k^{\star}}} is the value that was really computed, so it follows the skewed distribution and keeps landing where the prior is high. A wrong key passes the same plaintexts through the S-box under a different key, which relabels the values: the high-probability ones get mapped somewhere else. Per trace, the right key’s prior term averages −H[Z]-\mathbb{H}[Z]; a wrong key’s averages minus a cross-entropy, which by Gibbs’ inequality is never above −H[Z]-\mathbb{H}[Z], and falls strictly below it once the relabelling moves the heavy values. So the prior term alone, summed over enough traces, picks out k⋆\sym{kstar}{k^{\star}} — no leakage needed.

This is worst exactly where the stakes are highest. Against a well-masked implementation the real leakage is faint and hides in higher-order statistics, so for the optimizer, memorising the label distribution is the easy way to lower the loss.

The prior gets a vote. Top: one trace's posterior is the likelihood times the prior; as the leakage goes to zero, the likelihood goes flat and the posterior becomes the prior. Bottom: an attack with no leakage at all. The values the right key predicts (green) pile up under the prior's peak; a wrong key's values (white) are the same ones relabelled by the S-box, and scatter. Summed over traces, the right key's prior term pulls ahead of all 255 others.

A model that never looked

The paper’s toy example strips this to the bone. Let Z\sym{Z}{Z} be a single bit with Pr(Z=0)=0.7\sym{prior}{\mathrm{Pr}(Z = 0)} = 0.7, so H[Z]≈0.88\sym{HZ}{\mathbb{H}[Z]} \approx 0.88 bits, and let the traces be independent of ZZ: there is nothing to learn from them. The best a model can do is output the prior, which gives it a log-loss of

L=−E[log⁡Pr(Z∣X)]=−∑s∈{0,1}Pr(Z=s)log⁡Pr(Z=s)=H[Z]≈0.88\begin{aligned} \sym{loss}{\mathcal{L}} &= -\mathbb{E}\left[\log \mathrm{Pr}(Z \mid \mathbf{X})\right] \\ &= -\sum_{\sym{zval}{s} \in \{0,1\}} \sym{prior}{\mathrm{Pr}(Z = \sym{zval}{s})} \log \sym{prior}{\mathrm{Pr}(Z = \sym{zval}{s})} = \sym{HZ}{\mathbb{H}[Z]} \approx 0.88 \end{aligned}

Against the 1 bit you’d expect from a uniform label, that loss looks like 0.12 bits of information learned. And if the attack set has the same 70/30 skew, the guessing entropy converges. A model that never looked at a trace passes both tests.

Try it

The demo below runs the same experiment on a full byte, live in your browser. The intermediate value is z=S(p⊕k)\sym{Z}{z} = \sym{sbox}{S}(\sym{pt}{p} \sym{xor}{\oplus} \sym{k}{k}); its distribution is a truncated Gaussian centred on 127, the family the paper uses to bias ASCADv1. --prior-skew narrows it; the tick marks sit on the paper’s three bias levels, σ=128\sym{sigb}{\sigma} = 128, 6464 and 1616. --leakage sets how much each trace says about z\sym{Z}{z} (its Hamming weight plus Gaussian noise); at zero, traces are pure noise. The model is the ideal one for this setup: it knows both the leakage and the prior. Each curve is the mean rank of the right key over many simulated attack sets.

demo://prior-skew● running in your browser

π(z), z = S(p ⊕ k) on 8 bits · σ = 64 · H(Z) = 7.82 bits

posterior distinguisherAODrandom guess (128.5)GE vs. number of attack traces (log) · mean rank over 100 attack sets

Static render of the default setting (100 simulated attack sets at build time). With JavaScript on, it re-runs live in your browser as you move the sliders.

It opens at the medium bias with no leakage. The posterior curve falls to rank 1 within a hundred or so traces, from measurements that carry no information at all. The AOD curve, introduced in the next section, stays on the random-guess line. Pull --prior-skew to 0 and the two curves collapse onto each other at 128.5: that’s Lemma 1. Push it to the right and the posterior converges in a handful of traces.

Now give it some leakage. Both curves fall, and the AOD one is the honest one. The gap between them is the part of the “attack” the prior paid for.

Where the right key lands. Each dot is one simulated attack of 30 traces with no leakage, placed at the right key's rank. With a uniform prior the dots spread evenly and the guessing entropy (dashed) sits at 128.5. As the prior narrows through the paper's three bias levels (inset), they drain onto rank 1.

Divide the prior back out

If the posterior is likelihood times prior, the fix is to divide the prior back out before scoring. That’s the paper’s asymptotically optimal distinguisher (AOD):

dSa(AOD)[k]≜∑i=1Nalog⁡yi[zi,k]Pr(Z=zi,k)\sym{daod}{\mathsf{\textbf{d}}^{(\mathrm{AOD})}_{S_a}[k]} \triangleq \sum_{i=1}^{\sym{Na}{N_a}} \log \frac{\sym{yi}{\mathbf{y}_i}\left[\sym{zik}{z_{i,k}}\right]}{\sym{prior}{\mathrm{Pr}\left(Z = \sym{zik}{z_{i,k}}\right)}}

Here Pr(Z=zi,k)\sym{prior}{\mathrm{Pr}(Z = \sym{zik}{z_{i,k}})} is the empirical frequency of each value in the attack set. Dividing the posterior by the prior leaves the likelihood over Pr(X=xi)\sym{evid}{\mathrm{Pr}(\mathbf{X} = \sym{x}{\mathbf{x}_i})}, and that denominator is the same for every key. The paper shows that if the model converges to the ideal one, the expected AOD score ranks keys like d(opt)\sym{dopt}{\mathsf{\textbf{d}}^{(\mathrm{opt})}}: only the leakage counts.

It’s cheap. Counting value frequencies is linear in the number of traces, and the rest is the GE computation you already run: O(Np+Na∣K∣)\mathcal{O}(\sym{Np}{N_p} + \sym{Na}{N_a} |\sym{K}{\mathcal{K}}|) overall. In the demo, the AOD divides by the exact prior rather than counted frequencies; with a sharp prior and a few hundred traces, most values never appear in the attack set and would get a frequency of zero.

One caveat: the AOD doesn’t tell you there was a false positive. It just declines to have one. To see the problem you compare the AOD and posterior curves, or use the checks below.

On ASCADv1, where models are known to learn real leakage, the posterior and AOD guessing entropies come out similar. On our ASCADv2 model, the AOD guessing entropy never converges.

A quieter trap: averaging the same traces

The paper flags a second issue with how GE is estimated. You never have infinitely many attack sets, so you average the rank over n\sym{nsub}{n} subsets Si′\sym{Ssub}{S'_i} of the one attack set you recorded:

GE^(Na)=1n∑Si′∈CgSi′(k⋆)\sym{GEhat}{\widehat{\mathrm{GE}}}(\sym{Na}{N_a}) = \frac{1}{\sym{nsub}{n}} \sum_{\sym{Ssub}{S'_i} \in \sym{C}{\mathcal{C}}} \sym{rank}{g_{S'_i}(k^{\star})}

If each subset is as large as the whole attack set, every subset is the attack set, and the “average” is a single rank measured n\sym{nsub}{n} times. A single rank can hit 1 by luck when the true GE wouldn’t. That’s what happens when an attack needs all the traces you have, which was the case for the ASCADv2 attack. Empirically, the estimate starts behaving like a single rank once subsets pass about 30% of the attack set; the paper’s conservative rule of thumb is to keep them at most 10%.

Checks before training

Two checks run before you commit to a model.

A χ² test on the labels. Count how often each value occurs and test against uniform:

χ2=∑k=0∣Z∣−1(Ok−Ek)2Ek,Ek=N∣Z∣\sym{chi2}{\chi^2} = \sum_{k=0}^{|\sym{Zset}{\mathcal{Z}}|-1} \frac{(\sym{Ok}{O_k} - \sym{Ek}{E_k})^2}{\sym{Ek}{E_k}}, \qquad \sym{Ek}{E_k} = \frac{\sym{N}{N}}{|\sym{Zset}{\mathcal{Z}}|}

Compare it with the χ² distribution at ∣Z∣−1|\sym{Zset}{\mathcal{Z}}| - 1 degrees of freedom and a significance level such as α=0.05\sym{sig}{\alpha} = 0.05. It costs O(Np)\mathcal{O}(\sym{Np}{N_p}). On ASCADv2, byte 4, the one the Scoop attack targeted, fails the test; byte 6 passes. Plot the two histograms side by side, even sorted, and you can’t tell which is which, so run the test rather than eyeballing. Its limit: it tells you there’s a risk, not whether a model will use the bias, nor whether the attack set shares it.

A null benchmark. Replace every trace, profiling and attack, with zeros and train and evaluate as usual. This is the toy example run on your real labels. If the GE still converges, the labels alone carry enough bias to break the key, and the attack set shares it. Your real model then has to clearly beat this baseline to mean anything. It costs a full training run, O(E⋅Np⋅∣Θ∣)\mathcal{O}(\sym{E}{E} \cdot \sym{Np}{N_p} \cdot |\sym{Theta}{\Theta}|) for E\sym{E}{E} epochs and ∣Θ∣|\sym{Theta}{\Theta}| parameters, and like the χ² test it says nothing about leakage, only about exposure to the bias.

Checks after training

Once a model looks successful, four checks ask whether it actually uses its input. Each needs a trained model first, which is the expensive part.

Activation probing. Run the attack set through and look at the hidden activations. A model that has baked the prior into its weights reacts the same way to every trace. In our ASCADv2 model, only a few hidden units ever fired, and always the same ones, whatever the input. With 16 shuffled S-box lookups in the trace you’d expect much more variety. Cost: one forward pass, O(Na∣Θ∣)\mathcal{O}(\sym{Na}{N_a} |\sym{Theta}{\Theta}|). Suggestive, not conclusive.

Gradient visualisation. Take the gradient of the loss with respect to the input trace; its peaks show which samples the prediction depends on. On ASCADv1 the peaks line up with where an ANOVA says the masking shares leak. On ASCADv2 it’s noise. The catch: confirming the peaks needs an expert, and, on a masked target, access to the shares. It’s also prone to confirmation bias, and a valid model can have a noisy gradient. A clean match is good evidence; a noisy one proves nothing.

Model ablation. Set the input layer’s weights to zero, W(1)←0\sym{W1}{W^{(1)}} \leftarrow \mathbf{0}, and recompute the GE. If it still converges, the predictions come from the network’s own parameters and not from the traces: a false positive, no argument. If it falls back to random guessing, the model needed its input. That doesn’t make it a good model, only an honest one. This is the check that caught ASCADv2, and it costs one inference plus a GE, O(Na(∣Θ∣+∣K∣))\mathcal{O}(\sym{Na}{N_a}(|\sym{Theta}{\Theta}| + |\sym{K}{\mathcal{K}}|)).

Prediction entropy. Compute the entropy of each output vector, H[yi]=−∑s∈Zyi[s]log⁡2yi[s]\sym{Hy}{\mathbb{H}[\mathbf{y}_i]} = -\sum_{\sym{zval}{s} \in \sym{Zset}{\mathcal{Z}}} \sym{yi}{\mathbf{y}_i}[\sym{zval}{s}] \log_2 \sym{yi}{\mathbf{y}_i}[\sym{zval}{s}]. A model that outputs the prior gives nearly the same vector for every trace, so the entropy sits at one value with almost no spread. A model reading leakage is confident on clean traces and unsure on noisy ones. The two Scoop models side by side:

E[H(Y)]\sym{EH}{\mathbb{E}[\mathbb{H}(Y)]}V[Y]\sym{VY}{\mathbb{V}[Y]}H(argmax⁡Y~)\sym{Harg}{\mathbb{H}(\operatorname{argmax} \tilde{Y})}
ASCADv12.82690.48184.2202
ASCADv23.56860.00030.0265

Both targets are 8-bit labels assumed uniform, so the two models should look alike. The ASCADv2 model’s predictions barely move from trace to trace, and its top guess is almost always the same value. That’s a model with one answer.

Testing the tests

ASCADv2 is so far the only public dataset known to have this problem, so the paper builds some. It resamples ASCADv1 so that its labels follow a truncated Gaussian with μ=127\sym{mub}{\mu} = 127 and σ=128\sym{sigb}{\sigma} = 128, 6464 or 1616. Those are the low, medium and high bias levels: the information a model gains from the prior alone is respectively below, about equal to, and above what the best known model extracts from the leakage. The bias goes either into the profiling set only, or into both the profiling and attack sets.

What they found:

  • χ² is jumpy. It flags all three biased sets, and also the original ASCADv1 profiling set. It picks up even slight sampling imbalance.
  • The null benchmark is selective. It raises a flag only when the same bias sits in both the profiling and attack sets, which is exactly when a false positive is possible.
  • The training curves can give it away. At high bias, with the bias in both sets, the perceived information ends up well above that of the best known model. That can’t come from leakage.
  • Every post-mortem check catches the high bias. Only ablation catches the medium one. At medium bias the model learns the prior and the leakage, and ablation is the only check that separates the two.
  • The AOD changes the verdict only when it should. At low bias the posterior and AOD curves agree; at medium both degrade slightly and stay close; at high bias, with the bias in both sets, the posterior GE converges quickly and the AOD GE doesn’t. With the bias in the profiling set only, the two agree at every level: if the attack set doesn’t share the bias, the model can’t cash it in.

No single check wins. The cheap ones can’t see leakage; the informative ones need a trained model and sometimes an expert. The paper’s advice is to combine them, and to treat the AOD as the default scoring rule rather than an optional extra.

Takeaways

  • A GE curve reaching 1 means the score ranked the key first. It doesn’t mean the model read the traces.
  • The posterior distinguisher is only optimal when the labels are uniform. When they aren’t, it adds a prior term that favours the right key on its own.
  • Score with the AOD by default. It costs almost nothing.
  • Run a χ² test on your labels, and zero the input layer before you claim a break.
  • Report those checks next to the GE. A reader can’t rerun them for you.
$ cat rousselot2026really.bib
@article{rousselot2026really,  title={Is it Really Broken? The Failure of DL-SCA Scoring Metrics under Non-Uniform Priors},  author={Rousselot, Nathan and Heydemann, Karine and Masure, Lo{\"\i}c and Migairou, Vincent and Strullu, R{\'e}mi},  journal={CASCADE},  year={2026}}