Stochastic Nonlinearities

Stochastic Nonlinearities Improve
Uncertainty Estimation

Mohammad Sarhangzadeh1, Amirreza M. Shebly1, Doruk Oner1

1NeuraVision Lab, Bilkent University

British Machine Vision Conference · BMVC 2026

8-minute paper presentation

Paper Video

An 8-minute presentation covering the motivation, stochastic activation mechanism, information preservation, experiments, and main results.

Loading paper presentation… If the player does not appear, use the YouTube link shown here.

Abstract

Reliable uncertainty estimation is important across visual recognition tasks, providing information about the confidence and reliability of model predictions. Many uncertainty estimation methods obtain predictive diversity through mechanisms that suppress intermediate features, perturb the input, or require substantial additional computation.

We reframe stochastic activation switching as a lightweight mechanism for uncertainty estimation. The method samples between nonlinear activation functions, inducing predictive diversity while keeping intermediate feature pathways active. This enables uncertainty estimation from a single trained model without the training cost of ensembles, task-specific uncertainty heads, or architectural redesign.

Across classification, semantic segmentation, and out-of-distribution detection—with convolutional and transformer-based architectures—stochastic activation switching improves uncertainty quality while preserving competitive primary-task performance.

Multiple stochastic activation passes through a single model produce predictions whose variance gives the uncertainty map.

One set of weights. Different nonlinear pathways. Predictive disagreement becomes uncertainty.

Method

Stochastic Activation Switching

Each activation unit samples between the backbone’s native activation and an alternative nonlinearity. Instead of deleting features, the method perturbs the nonlinear transformation applied to them while keeping the intermediate feature pathways available.

The same stochastic mechanism is used during training and inference. At test time, repeated activation configurations produce a stochastic activation trajectory: multiple predictions from the same learned parameters but different nonlinear pathways.

hj(l,t) = (1 − mj(l,t)) σorig(zj(l,t)) + mj(l,t) σalt(zj(l,t))
p = 0.5uniform switching
Node-leveldefault granularity
0extra trainable parameters
PyTorch StochasticActivation module that samples a switch mask and selects between the original and alternative activation functions.

A drop-in stochastic nonlinearity

The switch remains stochastic during inference. This is important: repeated activation configurations are exactly what create the predictive distribution used for uncertainty estimation.

We use the backbone’s native activation with SiLU for classification and Tanh for segmentation.

01 Train normally

Sample activation choices and optimize the standard task loss.

02 Freeze weights

At inference, keep one learned model and resample activation configurations.

03 Aggregate passes

Use the mean as prediction and empirical variance as uncertainty.

Why this stochasticity?

Preserving Structure Under Stochasticity

MC Dropout creates diversity by masking activations. Stochastic activation switching instead changes how features are transformed while keeping their pathways active.

Comparison of stochastic activation and MC-Dropout uncertainty maps on abdominal CT segmentation.
Under stronger test-time stochasticity, activation switching keeps uncertainty more localized around boundaries and prediction errors, while MC Dropout becomes more diffuse.
Diagnostic analysis

Information Preservation

Information Destruction Rate compares estimated layer input–output information under stochastic inference with a deterministic reference. Lower is better. IDR is used only for analysis—not as an additional training loss.

IDR(l) = 1 − (1/T) Σt=1T It(l) / Ifull(l)

Across Massachusetts Roads, DRIVE, CIFAR-10, and MiniImageNet, stochastic activation switching yields substantially lower IDR than MC Dropout.

Lower IDR → more estimated information retained Diagnostic only → no extra objective
Information destruction rate comparison showing lower values for the proposed method than MC-Dropout.
Experiments

Quantitative Results

Browse every reported segmentation and classification table without a long wall of rows. Choose a result family, then slide between datasets.

Qualitative results

Uncertainty Near Difficult Predictions

The uncertainty maps are obtained from the variance across five stochastic activation trajectories.

Aerial road and retinal vessel segmentation examples with input, ground truth, prediction, and uncertainty map.
From left to right: input, ground truth, prediction, and uncertainty. Higher uncertainty appears around ambiguous structures and prediction errors.
Takeaway

Predictive diversity does not have to come from masking features or training multiple models.

By resampling nonlinear activation choices inside a single trained network, stochastic activation switching produces multiple functional pathways while keeping intermediate features available. Their disagreement gives a practical uncertainty signal without an auxiliary uncertainty head or architectural redesign.

Single-model uncertaintyOne set of learned weights produces multiple stochastic predictions at inference time.
Representation-aware stochasticityThe method changes nonlinear transformations instead of explicitly removing intermediate features.
Across tasks and shiftsWe observe competitive task performance and strong uncertainty behavior across classification, segmentation, and OOD evaluation.

Current trade-off: repeated stochastic passes add inference cost, and the activation pair influences the balance between predictive performance and calibration.

BibTeX

@inproceedings{sarhangzadeh2026stochastic,
  title     = {Stochastic Nonlinearities Improve Uncertainty Estimation},
  author    = {Sarhangzadeh, Mohammad and Shebly, Amirreza M. and Oner, Doruk},
  booktitle = {British Machine Vision Conference (BMVC)},
  year      = {2026}
}
BibTeX copied