2026-10-11 · poster, CHES 2026 · interactive version: press + on a block? tour
Gradient-free Profiling for Deep Learning based Side-Channel Analysis against Masking: the Solution for the Plateau Effect? Nathan Rousselot 1,2 · Karine Heydemann1 · Loïc Masure2 · Vincent Migairou1
1 Thales, Meyreuil, France · 2 LIRMM, Univ. Montpellier, CNRS, France
TL;DR: no. A Bayesian network trained without any gradient (TAGI, Tractable Approximate Gaussian Inference) still plateaus, and its plateau grows with the masking order like Adam’s. So do all five profilers we tested: the plateau has deeper origins than gradient-based optimization.
1 Masking and the plateau effectProfiled attacks
Learn P [ Y ∣ X ] \mathbb{P}[Y \mid X] P [ Y ∣ X ] from labelled traces
Attack unseen traces
Masking of order d \sym{d}{d} d
Y = S 0 ⊕ ⋯ ⊕ S d Y = S_0 \oplus \cdots \oplus S_{\sym{d}{d}} Y = S 0 ⊕ ⋯ ⊕ S d (d + 1 \sym{d}{d}+1 d + 1 random shares)
No leakage moment of order ≤ d \leq \sym{d}{d} ≤ d depends on Y Y Y
Schematic: the training loss sits at H(Y), the loss of a blind guess, then drops sharply. The drop comes later for masking order d = 2 than d = 1, and later still for d = 3. H(Y) training time (log scale) loss schematic d=1 d=2 d=3 plateau (d=3) Plateau effect [MCLS23, RHMM25]
Losses stall at H ( Y ) \sym{entropy}{H(Y)} H ( Y ) , then drop sharply
Plateau length grows exponentially with d \sym{d}{d} d
Persists when shares are unknown during profiling
Still occurs with noiseless leakages
more $ man block-1 : Masking and the plateau effect [esc]
Neural networks trained against masked implementations stall before they learn: the loss sits at
the value of a blind guess , then drops sharply, and the stall grows exponentially with the
masking order d \sym{d}{d} d . The figure in the block is a sketch. The experiment below is real,
but small: it trains a tiny network in your browser, on the same kind of task as block 5.
● live · toy re-run in your browsernot the poster's numbers
Training loss against samples seen, log scale, for the masking orders you run. Each curve starts at H(Y) = 4 bits and falls toward its dashed floor, the best any model can do from Hamming weights; a dot marks the escape. 10² 10³ 10⁴ 10⁵ 2 3 4 H(Y) = 4 bits, a blind guess samples seen (log) loss (bits) masking order d 2 run clear
Pick an order and press run. d = 3 takes a few seconds; d = 4 rarely escapes within the demo budget.
Each order adds a random share: Y = S_0 ⊕ … ⊕ S_d, each share leaking its Hamming weight, no noise. The network is an MLP (d+1) → 32 → 32 → 16 trained with Adam, much smaller than the poster's. Hamming weights can't pin the shares down, so even a perfect model keeps some loss: the dashed floor, H(Y | leakage). Expect the same shape (a flat loss at H(Y), then a fall to the floor, later for every extra share), not the same numbers.
Run d = 0 \sym{d}{d} = 0 d = 0 , then 1 1 1 , 2 2 2 , 3 3 3 . Each extra share stretches the flat part, by a lot more each time. During
the plateau the network isn’t slowly improving: its loss sits at H ( Y ) \sym{entropy}{H(Y)} H ( Y ) . Then it leaves and falls
to its floor, the dashed line: with noise-free Hamming weights, even a perfect model can’t tell every value of the
secret apart.
2 Is the gradient to blame?Scoop [RHMM25]: the plateau is caused by gradient training
Gradient magnitudes decay exponentially with d \sym{d}{d} d
Saddle points trap the optimizer
Theory: on XOR-like targets, gradients carry almost no label information [SSS17]
Research question. If gradients cause the plateau, a learner that never computes one should avoid it.
Is gradient-free profiling the solution?
more $ man block-2 : Is the gradient to blame? [esc]
In Scoop we blamed gradient-based training for the plateau. Gradients shrink
exponentially with d \sym{d}{d} d , a saddle point sits near where training starts, and on XOR-like targets a
gradient carries almost no information about the label. Scoop changes the optimizer accordingly, and it does
shorten the plateau.
That makes a testable prediction. If the gradient causes the plateau, a learner that never computes one
shouldn’t have it. So we need a learner that trains a neural network with no gradient at all (block 3),
and a fair way to compare it with the others (block 4).
3 Bayesian learning, no gradientIn TAGI [GNA21], every weight is a Gaussian random variable.
Illustration of TAGI: a trace X enters a small network whose weights are each a Gaussian distribution, w ∼ N(μ_w, σ_w²). Step 1, forward: means and variances propagate in closed form to a predicted distribution P[Y|X]. Step 2, update: Bayes' rule, layer by layer, backpropagation style. trace X w ∼ N(μ_w, σ_w²) P[Y | X] 1 forward: propagate (μ, σ²) in closed form2 update: Bayes' rule, layer by layer↻ replay The updates use Bayesian statistics:
Δ μ θ = Cov ( θ , z ) y − μ z σ z 2 + σ v 2 \sym{dmu}{\Delta\mu_\theta} = \sym{cov}{\operatorname{Cov}}(\sym{wt}{\theta}, \sym{out}{z})\,\frac{\sym{err}{\sym{y}{y} - \sym{muz}{\mu_z}}}{\sym{unc}{\sym{sz}{\sigma_z^2} + \sym{sv}{\sigma_v^2}}} Δ μ θ = Cov ( θ , z ) σ z 2 + σ v 2 y − μ z
No gradient is ever computed, and no learning rate is used.
Bayesian layers run on GPU (CUDA, built on cuTAGI [NG25]).
more $ man block-3 : Bayesian learning, no gradient [esc]
TAGI (tractable approximate Gaussian inference, Goulet et al.) trains a neural network without
backpropagating a gradient. Every weight is a Gaussian random variable with a mean and a variance.
The forward pass pushes means and variances through the network in closed form. Learning is Bayes’
rule: condition on the label , then pass the correction back layer by layer, the way backpropagation
would, but with covariances instead of derivatives.
Δ μ θ = Cov ( θ , z ) y − μ z σ z 2 + σ v 2 \sym{dmu}{\Delta\mu_\theta} = \sym{cov}{\operatorname{Cov}}(\sym{wt}{\theta}, \sym{out}{z})\,\frac{\sym{err}{\sym{y}{y} - \sym{muz}{\mu_z}}}{\sym{unc}{\sym{sz}{\sigma_z^2} + \sym{sv}{\sigma_v^2}}} Δ μ θ = Cov ( θ , z ) σ z 2 + σ v 2 y − μ z The prediction error y − μ z \sym{err}{\sym{y}{y} - \sym{muz}{\mu_z}} y − μ z is scaled by how much the weight θ \sym{wt}{\theta} θ covaries with the output
z \sym{out}{z} z , relative to the output’s total uncertainty . Because every weight carries a
variance, the network doesn’t output one function but a distribution over functions, and it knows where it is unsure.
Below, a tiny network of this kind (one input, 32 hidden units, one output) fits a curve. The band is its
prediction ± 2σ; the thin lines are five networks drawn from its weight distributions.
● live · a Bayesian network, 1 → 32 → 1, in your browserillustration, not the poster's networks
A small Bayesian network's prediction against its input x. A shaded band shows its uncertainty, ±2σ, and thin lines show networks drawn from its weight distributions. Near the labels the band is narrow and the drawn networks agree; for x above 1.5, where no label was drawn, the band stays wide. no random labels −2 0 2 −3 −2 −1 0 1 2 3 input x mean ± 2σ 5 networks drawn from the weights labels the function they come from+1 label +10 labels one more pass reset label noise σ_v 0.15
30 labels, three passes. ±2σ at x = 0: 0.37; at x = 2.5, where no label fell: 1.89.
Before any label the band is the prior: the drawn networks disagree everywhere. Click the plot to give a label where you like. Each one is a single Bayesian update of every weight's mean and variance: the band narrows around it and the drawn networks fall into line. Nothing changes where no label has been, and the band there stays wide: the network says it doesn't know. A small σ_v trusts each label more, a large one barely moves.
Three things to notice. Each label is one closed-form update, and the step size comes out of the variances:
nothing plays the role of a learning rate. Every update also shrinks variances, so the band closes where labels
fall and stays open where none did. And the prediction is a distribution, so “I don’t know” is part of the answer.
4 Protocol: five profiling techniquesgradient-based gradient-free point estimate Bayesian closed-form updates our focus click a learner to follow it through blocks 5 and 6
Every baseline is tuned with the same budget, for a fair comparison.
Plateau: samples seen until the loss < H ( Y ) − ε < \sym{entropy}{H(Y)} - \sym{tol}{\varepsilon} < H ( Y ) − ε .
more $ man block-4 : Protocol: five profiling techniques [esc]
TAGI is one corner of a two-by-two grid: gradient-based or not, point estimate or Bayesian. The other
corners are filled by Adam and Scoop (backpropagation), an evolution strategy, and variational inference.
If the plateau came from the gradient, the two gradient-free learners would escape it.
Gradient-based Gradient-free Point estimate Adam, Scoop (backpropagation) ES (evolution strategy) Bayesian VI (variational inference) TAGI (closed-form updates)
Each one gets the same tuning budget. The plateau is measured in training samples seen until the
loss falls below H ( Y ) − ε \sym{entropy}{H(Y)} - \sym{tol}{\varepsilon} H ( Y ) − ε , with ε \sym{tol}{\varepsilon} ε = [placeholder: value of ε].
Click a learner in the grid, or in any legend, to follow it through blocks 5 and 6.
5 TAGI plateaus tooSamples until escape against masking order d, log scale. Adam and TAGI rise exponentially from 32 at d = 0 to about a million and more at d = 4; Scoop matches them at d = 0 and 1. 10¹ 10³ 10⁵ 10⁷ 0 1 2 3 4 masking order d plateau size 4.2·10⁶ 1.0·10⁶ Adam Scoop (d ≥ 2 running) TAGI one epochSetup: exhaustive dataset · 4-bit secret masked at order d (16d+1 traces) · noise-free Hamming-weight leakage · MLP (d+1) → 200 → 200 → 16 · 20 seeds
TAGI suffers from the same exponential increase (at the same rate) as gradient-based DL-SCA.
more $ man block-5 : TAGI plateaus too [esc]
The cleanest setting: a 4-bit secret masked at order d \sym{d}{d} d , every combination of share values
in the dataset (16 d + 1 16^{\sym{d}{d}+1} 1 6 d + 1 traces), noise-free Hamming-weight leakage, the same small MLP for every
learner, 20 seeds. Median samples seen before the loss breaks:
Order d \sym{d}{d} d 0 1 2 3 4 Adam 32 2,304 36,864 196,608 4,194,304 TAGI 32 1,280 53,248 327,680 1,048,576 Scoop 32 2,560 still running
TAGI never computes a gradient, and its plateau still grows exponentially with d \sym{d}{d} d , at the same rate
as Adam’s. At some orders it escapes sooner, at others later; the trend is the same. The dotted line is one
epoch, 16 d + 1 16^{\sym{d}{d}+1} 1 6 d + 1 samples: every learner needs to see the whole dataset many times over before it escapes.
A toy version of this experiment runs live in block 1.
6 Every profiling method plateausWe consider a noisy leakage of a single bit with dummy samples. The noise is additive Gaussian.
noise σ = 1.0 σ = 1.5Median samples until escape against masking order d = 0, 1, 2, for Adam, Scoop, VI, ES and TAGI, log scale. Every learner rises with d. At noise σ = 1.5 and d = 2, Scoop, VI and ES do not escape within the 2·10⁶ sample budget, and 6 of 10 TAGI runs do. 10² 10⁴ 10⁶ 0 1 2 masking order d plateau size budget: 2·10⁶ samples 0/10 escaped 6/10 Adam Scoop VI ES TAGI not escaped budget
Every method suffers from the plateau
Some sort of masking hardness
more $ man block-6 : Every profiling method plateaus [esc]
Noisy leakage: a single masked bit plus dummy samples, with additive Gaussian noise , 10
seeds, and a budget of two million samples. Median samples until escape, at d = 0 , 1 , 2 \sym{d}{d} = 0, 1, 2 d = 0 , 1 , 2 :
Learner Noise σ = 1.0 \sym{sigma}{\sigma} = 1.0 σ = 1.0 Noise σ = 1.5 \sym{sigma}{\sigma} = 1.5 σ = 1.5 Adam 128 / 576 / 3,776 128 / 1,472 / 132,480 Scoop 448 / 1,408 / 6,912 512 / 2,816 / not escaped VI 128 / 576 / 4,160 128 / 1,408 / not escaped ES 1,152 / 3,840 / 41,984 1,472 / 9,088 / not escaped TAGI 512 / 640 / 3,776 512 / 2,048 / 1,027,840
All five plateau, and all five plateau longer as d \sym{d}{d} d grows: gradient or not, Bayesian or not. At
the higher noise level and d = 2 \sym{d}{d} = 2 d = 2 , only Adam and TAGI escape within the budget, and only six of
the ten TAGI runs do. Flip the noise switch on the figure to watch every learner climb.
Every method seems to suffer from masking in the same way. This lets us conjecture that the plateau
effect comes from some sort of hardness inherent to masking, not from the way a network is trained.
7 Why does TAGI plateau?σ = 1.0 · last-layer weights · median over 10 seeds · relative to d = 0
TAGI, relative to d = 0. Size of the update: 1, 0.93, 1.01 at d = 0, 1, 2. Output–label covariance: 1, 0.53, 0.20. Size of the update 0 0.5 1 d=0 d=1 d=2 ≈1 Output–label covariance d=0 d=1 d=2 ≈0.2 relative to d = 0
Training does not stall: updates do not vanish with d \sym{d}{d} d
Output–label covariance ≈ 5× lower at d = 2 \sym{d}{d} = 2 d = 2
Updates decorrelate from the secret as d \sym{d}{d} d grows, similar to [SSS17, RHMM25]
more $ man block-7 : Why does TAGI plateau? [esc]
TAGI’s updates are easy to instrument, so we looked at what changes with d \sym{d}{d} d . Two quantities,
measured on the last layer at σ = 1.0 \sym{sigma}{\sigma} = 1.0 σ = 1.0 and expressed relative to d = 0 \sym{d}{d} = 0 d = 0 (median over 10
seeds):
Order d \sym{d}{d} d 0 1 2 Size of TAGI’s update 1 0.93 1.01 Output–label covariance, TAGI 1 0.53 0.20 Output–label covariance, Adam 1 0.51 0.20
The updates don’t shrink. TAGI keeps moving its weights as much at d = 2 \sym{d}{d} = 2 d = 2 as at d = 0 \sym{d}{d} = 0 d = 0 . What
collapses is the covariance between the network’s output and the label, about five times smaller
at d = 2 \sym{d}{d} = 2 d = 2 , and Adam shows the same drop. The learner isn’t stuck; its updates have stopped
pointing at the secret. That’s the same symptom Shalev-Shwartz et al. describe for gradients, and
that we saw in Scoop, now without any gradient involved.
In the update of block 3, that covariance is the factor Cov ( θ , z ) \sym{cov}{\operatorname{Cov}}(\sym{wt}{\theta}, \sym{out}{z}) Cov ( θ , z ) in front of the
error: the step keeps its size, but it is spent in directions that say little about Y Y Y .
8 What's next? Towards more frugal DL-SCASynthetic 1st-order masked bytes (σ = 0.5) · MLP 16–200–200–256 · 5 seeds
Against the number of profiling traces N_p, from 500 to 100,000: validation perceived information and attack traces needed, TAGI and Adam. TAGI's PI is positive from 500 traces while Adam's stays negative up to 2,000. At 2,000 profiling traces TAGI needs 5 attack traces and Adam 3,070. From 10,000 on, both are close. validation PI (bits) 0 2 4 6 1k 10k 100k attack traces to GE ≤ 1 1 10 100 1k 10k 1k 10k 100k profiling traces N_p (log) MI = 6.69 attack set TAGI Adam no attack within 10k traces--n-prof N_p = 2,000 · PI: TAGI 1.76, Adam −0.54 bits · attack traces: TAGI 5, Adam 3,070
Profiling traces needed: TAGI 500 vs. Adam 5,000
At N p = 2,000 \sym{Np}{N_p} = 2{,}000 N p = 2 , 000 : TAGI 5 vs. Adam 3,070 attack traces
Few traces: Adam is over-confident
The baseline has been tuned.
more $ man block-8 : What's next? Towards more frugal DL-SCA [esc]
One thing did come out in TAGI’s favour. On synthetic first-order masked bytes (σ = 0.5 \sym{sigma}{\sigma} = 0.5 σ = 0.5 ,
5 seeds), TAGI needs far fewer profiling traces than a tuned Adam baseline:
Profiling traces N p \sym{Np}{N_p} N p 500 1,000 2,000 5,000 10,000 PI, TAGI (bits) 0.12 0.66 1.76 3.43 5.12 PI, Adam (bits) −1.10 −0.86 −0.54 1.71 4.72 Attack traces, TAGI 35 11 5 3 2 Attack traces, Adam > 10,000 10,000 3,070 4 2
With few traces, Adam is overconfident: its perceived information is negative, so its predictions
are worse than a uniform guess. TAGI is positive from 500 traces. At 2,000 profiling traces, TAGI’s
model recovers the key in 5 attack traces and Adam’s in 3,070. With 10,000 or more, the two meet.
Drag the slider under the figure to read both panels at one N p \sym{Np}{N_p} N p .
This is synthetic data only; whether TAGI stays this frugal on real traces is the next thing to test.
References [RHMM25] N. Rousselot, K. Heydemann, L. Masure, V. Migairou. Scoop: an optimization algorithm for profiling attacks against higher-order masking. TCHES 2025(3). ePrint 2025/498 [MCLS23] L. Masure, V. Cristiani, M. Lecomte, F.-X. Standaert. Don’t learn what you already know. TCHES 2023(1).[GNA21] J.-A. Goulet, L. H. Nguyen, S. Amiri. Tractable approximate Gaussian inference for Bayesian neural networks. JMLR 2021.[SSS17] S. Shalev-Shwartz, O. Shamir, S. Shammah. Failures of gradient-based deep learning. ICML 2017.[BEG+22] B. Barak et al. Hidden progress in deep learning. NeurIPS 2022.[NG25] L. H. Nguyen, J.-A. Goulet. cuTAGI / pytagi. github.com/lhnguyen102/cuTAGI Acknowledgment This work was partially funded by the France 2030 program, managed by the French National Research Agency under grant agreement No. ANR-22-PETQ-0008 PQ-TLS.
Take-home message: the plateau effect has long days ahead of it Without any gradient, TAGI still plateaus. Masking decorrelates the leakage from the labels. The plateau effect seems intrinsic to masking and is not an optimization artifact. TAGI might be more frugal than regular DL techniques. Next: not the gradient… so where does the plateau really come from?
CHES 2026 · Antalya, Türkiye · October 11–15, 2026 Ongoing work, ideas and developments were made by humans and LLMs were not used on any research tasks.
$ cat rousselot2026poster_tagi.bib[esc]
@misc{rousselot2026poster_tagi, author = {Nathan Rousselot and Karine Heydemann and Lo{\"\i}c Masure and Vincent Migairou}, title = {Gradient-free Profiling for Deep Learning based Side-Channel Analysis against Masking: the Solution for the Plateau Effect?}, howpublished = {Poster, CHES 2026}, address = {Antalya, Turkey}, year = {2026}, month = {October}, type = {poster}, url = {https://nathan-rousselot.com/posters/ches2026-gradient-free-plateau.pdf} }