BAM! Bayesian Anything Model: a foundation model for generative computational imaging

1Université Paris Cité, CNRS, MAP5 UMR 8145, Paris, France
2Heriot-Watt University, MACS & Maxwell Institute for Mathematical Sciences, Edinburgh, UK
*Indicates Equal Contribution
DIV2K SR ×4: observation DIV2K SR ×4: BAM posterior sample ObservationBAM DIV2K | SR ×4
AFHQ Gaussian deblurring: observation AFHQ Gaussian deblurring: BAM posterior sample ObservationBAM AFHQ | Gaussian deblurring
FFHQ JPEG (Q = 10): observation FFHQ JPEG (Q = 10): BAM posterior sample ObservationBAM FFHQ | JPEG (Q = 10)
FFHQ Inpainting: observation FFHQ Inpainting: BAM posterior sample ObservationBAM FFHQ | Inpainting
Köhler Blind motion deblurring: observation Köhler Blind motion deblurring: BAM posterior sample ObservationBAM Köhler | Blind motion deblurring
LSUN Inpainting: observation LSUN Inpainting: BAM posterior sample ObservationBAM LSUN | Inpainting
LSUN SR ×4: observation LSUN SR ×4: BAM posterior sample ObservationBAM LSUN | SR ×4

One small network, many imaging problems, few steps. Drag the sliders to compare each observation (left) with a BAM posterior sample (right). A single 36M-parameter network, with the forward operator supplied at inference time, covers ×4 super-resolution (DIV2K, LSUN), Gaussian deblurring (AFHQ), inpainting (FFHQ, LSUN), non-linear JPEG restoration at Q = 10 (FFHQ) and blind motion deblurring (Köhler), all in 3 steps.

Abstract

Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models. Current practice falls into two camps. Large foundation image models are deployed as plug-and-play priors with zero-shot approximate likelihood guidance, which introduces significant bias and computational cost. Physics-aware generative models avoid this bias, but each is tied to a specific dataset, task and instrument. We introduce BAM (Bayesian Anything Model), a lightweight foundation model for few-step, physics-aware posterior sampling that generalises robustly to unseen data and tasks, zero-shot or with minimal finetuning. BAM upgrades the operator-conditioned Reconstruct Anything Model (RAM) backbone (Terris et al., 2025) into a conditional flow map, so instrument physics is specified at inference time rather than fixed during training. BAM has just 36M parameters and is pre-trained jointly on large image corpora and libraries of forward operators. A single network then draws posterior samples in a few steps, with no likelihood approximation and no guidance weights to tune. Across linear inverse problems on FFHQ, AFHQ, LSUN, DIV2K and the Köhler camera-shake benchmark, BAM outperforms in just 3 steps both specialised models and leading zero-shot methods in sample quality, at a fraction of their computational cost. BAM gives the community an accessible entry point to generative computational imaging, lowers the economic and environmental cost of training imaging models, and opens a new path for research on physics-aware Bayesian computational imaging. Code and weights will be released upon acceptance.

36M
parameters, one network for all problems
3
network evaluations per posterior sample
38/40
settings where a BAM variant has the best LPIPS
36/40
settings where a BAM variant has the best CMMD

Method

We consider imaging problems $y = A x^\star + \sigma_y w$, where the forward operator $A$ and the noise level $\sigma_y$ are known at inference time, and aim to draw samples from the posterior $p(\bm x \mid y, A, \sigma_y)$ in a handful of network evaluations. BAM is a conditional flow map: it transports a Gaussian reference $\rho_1 = \mathcal N(0, \sigma_d^2 \Id)$ to the posterior $\rho_0 = p(\bm x \mid y, A, \sigma_y)$, jumping between any two times $t \geq s$ of the probability-flow ODE. It amortises over a wide family of operators and noise levels, so a single model serves the whole family, with no retraining, or at most light finetuning, when the forward model changes.

Schematic of the BAM conditional flow map
The conditional flow map $X^\theta_{t,s}$ moves a point from time $t$ to $s \le t$ along the probability flow that carries the Gaussian reference ($t=1$) to the posterior ($t=0$), with conditioning $c = (y_\sigma, \alpha_\sigma A)$. The purple arrow is the finite jump $(s-t)\,v_\theta(x_t, t, s, c)$; the green arrow is the reverse-time tangent $-v_\theta(x_t, t, t, c)$. Paths and densities are illustrative.

Physics-aware conditioning

Feeding $(y, A, \sigma_y, t, s)$ directly to a small network is numerically fragile, as $\sigma_y$ and $(t,s)$ span several orders of magnitude. BAM therefore conditions on a rescaled measurement, defined through the stochastic interpolant

$$\bm y_\sigma = \alpha_\sigma A \bm x + \varsigma_\sigma \bm w, \qquad \varsigma_\sigma / \alpha_\sigma = \sigma,$$

so that $y$ is a realisation of $\bm y_\sigma$ at $\sigma = \sigma_y$ up to the known factor $\alpha_\sigma$. The flow map is parametrised as

$$X^\theta_{t,s}(x \mid y_\sigma, \alpha_\sigma A) = x + (s-t)\, v_\theta(x, t, s, y_\sigma, \alpha_\sigma A), \qquad 0 \le s \le t \le 1,$$

which enforces the boundary condition $X^\theta_{t,t}(x) = x$ by construction. The velocity $v_\theta$ is a RAM backbone (an operator-conditioned DRUNet with unrolled physics-aware updates) that receives the augmented measurement $[x_t\,;\, y_\sigma]$ and the stacked operator $[(1-t)\Id\,;\, \alpha_\sigma A]$ supplied at inference time; its two noise-level embeddings take $t$ and $s$. $\sigma_y$ enters only through the rescaling, which keeps the model lightweight.

Training: Lagrangian self-distillation

On the diagonal ($t = s$) the velocity is fitted to the interpolant slope, and off the diagonal ($s < t$) the endpoint derivative of the map is matched to the drift of an EMA teacher $\theta^-$ at the transported point:

$$\mathcal L_b(\theta) = \mathbb E\big\| v_\theta(\bm x_t, t, t, \bm y_\sigma, \bm A) - (\bm z - \bm x_0) \big\|_2^2, \qquad \mathcal L_{\mathrm{LSD}}(\theta) = \mathbb E\big\| \partial_s X^\theta_{t,s} - \sg\big[v_{\theta^-}(\hat{\bm x}_{t,s}, s, s, \bm y_\sigma, \bm A)\big] \big\|_2^2.$$

Images $\bm x$ are drawn from large image corpora and operators $\bm A$ from a library of forward models (super-resolution, Gaussian and motion blur, compressed sensing, inpainting, demosaicing) with $\sigma_y \in [0, 0.05]$. The measurement noise is tied to $t$ through $\sigma(t) = \sigma_{\max}\gamma t / (1-(1-\gamma)t)$, which concentrates capacity where $\bm y_\sigma$ and $\bm x_t$ carry comparable noise. The full objective adds an LPIPS term and a contrast term, active only near $s \approx 0$:

$$\mathcal L(\theta) = \mathbb E\Big[\lambda_b \ell_b + \lambda_L \ell_{\mathrm{LSD}} + e^{-4s}\big\{\lambda_p \ell_{\mathrm{LPIPS}}(\hat{\bm x}_{t,s}, \bm x_0) + \lambda_c \ell_{\mathrm{ctr}}(\hat{\bm x}_{t,s}, \bm x_0)\big\}\Big].$$

No adversarial loss is used at any stage. Starting from the public RAM checkpoint, the baseline BAM is trained for 170 hours on 8 H100 GPUs; the fine-tuned variant BAM★ continues for 16 hours on 2 H100 GPUs per target dataset, still across all problems.

Few-step posterior sampling

Each step re-noises the current iterate to level $t_k$, pairs it with a draw of the measurement interpolant at the matching level $\sigma_k = \sigma(t_k)$, and maps it back to $s = 0$ in a single network evaluation. Independent runs give independent posterior samples.

Algorithm 1: BAM posterior sampling (3 steps)
  1. Given measurement $y$, operator $A$, noise level $\sigma_y$, schedule $1 = t_1 > t_2 > t_3$, levels $\sigma_k = \sigma(t_k)$, network $v_\theta$.
  2. $\bm\varepsilon \sim \mathcal N(0, \Id)$ (held fixed across steps)
  3. For $k = 1, 2, 3$
  4. $\bm y_{\sigma_k} \gets \alpha_{\sigma_k} y + \sqrt{\varsigma_{\sigma_k}^2 - \alpha_{\sigma_k}^2 \sigma_y^2}\;\bm\varepsilon$ (add the missing measurement noise)
  5. $\bm z^{(k)} \sim \mathcal N(0, \sigma_d^2 \Id)$
  6. $\bm x_{t_k} \gets (1-t_k)\,\hat{\bm x}^{(k-1)} + t_k\, \bm z^{(k)}$ (pure noise when $t_1 = 1$)
  7. $\hat{\bm x}^{(k)} \gets \bm x_{t_k} - t_k\, v_\theta(\bm x_{t_k}, t_k, 0, \bm y_{\sigma_k}, \alpha_{\sigma_k} A)$ (one NFE: $X^\theta_{t_k,0}$)
  8. Return $\hat{\bm x}^{(3)}$

For small $\sigma_y$ the last step is taken at $\sigma_3 = \sigma_y$, where no re-noising is needed. For large $\sigma_y$, BAM sets $\sigma_2 = \sigma_y$ and takes the last step at $\sigma_3 < \sigma_y$, drawing the interpolant by a Brownian bridge between $A\hat{\bm x}^{(2)}$ and the rescaled measurement.

Results

LPIPS versus reconstruction FLOPs for BAM and baselines

State-of-the-art quality at a fraction of the cost

BAM advances the quality–cost Pareto frontier drawn by state-of-the-art methods: better perceptual quality (LPIPS) with fewer FLOPs per 512×512 image and far fewer parameters (bubble area). The same conclusion holds across metrics, inverse problems and datasets.

BAM is also robust to post-training quantisation: in INT8 under TensorRT it restores a 512×512 image in 80 ms instead of 281 ms on an NVIDIA A40 (3.5× faster), with a moderate loss in accuracy (PSNR −1.3 dB, LPIPS 0.400 → 0.427) and no fine-tuning after quantisation.

Reconstruction quality

Across four datasets, five problems and two noise levels, a BAM variant attains the best LPIPS in 38 of 40 settings and the best CMMD in 36 of 40. BAM restores sharp, realistic textures such as fur, hair and foliage while remaining consistent with the measurements. RAM, trained as an MMSE estimator, often attains higher PSNR at the cost of smoother textures, while zero-shot and latent-space methods show noise artifacts or hallucinate content inconsistent with the observation.

DIV2K and FFHQ restoration at σy = 0.05

MethodNFEs DIV2K FFHQ
PSNR↑ LPIPS↓ CMMD↓ FID↓ PSNR↑ LPIPS↓ CMMD↓ FID↓
Gaussian deblurring
BAM324.480.350.0634.429.630.250.1045.0
BAM★324.440.350.0435.630.160.250.0851.6
RAM125.630.430.4452.630.540.351.1594.7
RAM★126.000.420.3447.031.530.341.0792.4
LATINO-PRO6523.550.540.4971.923.400.461.17128.3
SILO20017.200.621.84298.824.920.380.2591.8
I2SB50––––30.240.240.0244.2
UD2M12––––29.300.290.7267.3
Super-resolution
BAM324.300.360.0839.628.840.270.1950.4
BAM★323.560.370.0941.129.370.270.1962.4
RAM125.790.430.2674.131.150.361.02107.1
RAM★125.880.420.2751.731.010.351.5293.0
LATINO-PRO6522.820.540.5995.323.180.530.90146.3
SILO20016.530.642.04308.425.040.370.2080.7
I2SB50––––26.820.391.02148.8
UD2M12––––28.160.310.3082.2
Inpainting
BAM326.610.270.0220.531.500.230.0832.1
BAM★325.660.270.0321.031.570.200.0532.4
RAM123.520.410.3066.925.550.391.53126.5
RAM★127.610.320.1835.532.400.291.1256.5
LATINO-PRO6511.250.712.01282.413.440.792.63305.0
SILO20012.920.732.91382.417.460.521.00139.4
I2SB50––––18.890.581.93174.2
UD2M12––––29.290.250.2145.54

PSNR (dB, ↑), LPIPS (↓), CMMD (↓) and FID (↓). BAM variants use three steps; BAM★ and RAM★ denote finetuning. Bold and underline mark the best and second-best values for each dataset and problem; – denotes an unavailable result.

Visual comparisons

BAM and BAM★ use three steps. Yellow boxes mark the region magnified 4× below each image. Click any image to compare it against another method with a slider.

Observation
BAM
BAM★
RAM
RAM★
SILO
TReg
GT
Gaussian deblurring σy = 0.05
DIV2K: Observation
DIV2K: BAM
DIV2K: BAM<sup>★</sup>
DIV2K: RAM
DIV2K: RAM<sup>★</sup>
DIV2K: SILO
DIV2K: TReg
DIV2K: GT
SR ×4 σy = 0.05
DIV2K: Observation
DIV2K: BAM
DIV2K: BAM<sup>★</sup>
DIV2K: RAM
DIV2K: RAM<sup>★</sup>
DIV2K: SILO
DIV2K: TReg
DIV2K: GT
Inpainting σy = 0.05
DIV2K: Observation
DIV2K: BAM
DIV2K: BAM<sup>★</sup>
DIV2K: RAM
DIV2K: RAM<sup>★</sup>
DIV2K: SILO
DIV2K: TReg
DIV2K: GT
Observation
BAM
BAM★
RAM
RAM★
SILO
UD2M
I2SB
GT
FFHQ — Gaussian deblurring, σy = 0.05
FFHQ: Observation
FFHQ: BAM
FFHQ: BAM<sup>★</sup>
FFHQ: RAM
FFHQ: RAM<sup>★</sup>
FFHQ: SILO
FFHQ: UD2M
FFHQ: I2SB
FFHQ: GT
FFHQ — SR ×4, σy = 0.05
FFHQ: Observation
FFHQ: BAM
FFHQ: BAM<sup>★</sup>
FFHQ: RAM
FFHQ: RAM<sup>★</sup>
FFHQ: SILO
FFHQ: UD2M
FFHQ: I2SB
FFHQ: GT
FFHQ — JPEG, Q = 10, pre-compression σJPEG = 0.01
FFHQ: Observation
FFHQ: BAM
FFHQ: BAM<sup>★</sup>
FFHQ: RAM
FFHQ: RAM<sup>★</sup>
FFHQ: SILO
FFHQ: UD2M
–
FFHQ: GT
Observation
BAM
BAM★
RAM
RAM★
LATINO
L-PRO
TReg
SILO
GT
AFHQ — Inpainting, σy = 0.05
AFHQ: Observation
AFHQ: BAM
AFHQ: BAM<sup>★</sup>
AFHQ: RAM
AFHQ: RAM<sup>★</sup>
AFHQ: LATINO
AFHQ: L-PRO
AFHQ: TReg
AFHQ: SILO
AFHQ: GT
AFHQ — Demosaicing, σy = 0.05
AFHQ: Observation
AFHQ: BAM
AFHQ: BAM<sup>★</sup>
AFHQ: RAM
AFHQ: RAM<sup>★</sup>
AFHQ: LATINO
AFHQ: L-PRO
AFHQ: TReg
AFHQ: SILO
AFHQ: GT
AFHQ — Compressed sensing, σy = 0.05
AFHQ: Observation
AFHQ: BAM
AFHQ: BAM<sup>★</sup>
AFHQ: RAM
AFHQ: RAM<sup>★</sup>
AFHQ: LATINO
AFHQ: L-PRO
AFHQ: TReg
–
AFHQ: GT

Full quantitative comparison

All five problems BAM is trained on, at σy ∈ {0.025, 0.05}. AFHQ and LSUN bedroom are out-of-distribution for the baseline BAM. PSNR (dB, ↑), LPIPS (↓), CMMD (↓), FID (↓); bold / underline: best / second-best value per problem, noise level and metric.

DIV2K: 64 test images. BAM and BAM★ use three steps.

MethodNFEs σy = 0.025 σy = 0.05
PSNR↑ LPIPS↓ CMMD↓ FID↓ PSNR↑ LPIPS↓ CMMD↓ FID↓
Deblurring
BAM325.010.3340.0433.5124.480.3490.0634.43
BAM★324.910.3390.0333.8024.440.3530.0435.59
RAM126.080.4010.3741.3525.630.4270.4452.55
LATINO823.290.4730.2874.6222.000.5140.39100.27
LATINO-PRO6524.070.4540.3566.0223.550.5360.4971.90
TReg20022.430.4850.3483.8921.850.5150.40102.21
SILO20017.560.6141.72287.0317.200.6241.84298.80
Super-resolution
BAM324.920.3430.0636.8024.300.3620.0839.55
BAM★324.190.3540.0636.9523.560.3730.0941.12
RAM126.490.3970.2050.9225.790.4300.2674.12
LATINO824.360.4780.3757.1522.700.5420.5697.24
LATINO-PRO6524.490.4760.3854.1122.820.5430.5995.31
TReg20022.070.4970.3976.5220.730.5650.55112.77
SILO20017.040.6311.99295.6716.530.6412.04308.43
Inpainting
BAM327.310.2240.0117.3526.610.2670.0220.46
BAM★326.560.2330.0217.7325.660.2720.0321.01
RAM123.610.3990.2961.7823.520.4120.3066.93
LATINO815.220.6921.74229.4314.450.7041.76224.50
LATINO-PRO6513.730.6781.79270.4211.250.7062.01282.41
TReg20015.150.7271.47287.3014.480.7481.67299.14
SILO20013.080.7242.85378.4212.920.7292.91382.44
Demosaicing
BAM333.370.1210.017.3830.480.1940.0212.68
BAM★332.610.1270.018.5129.850.1960.0214.02
RAM130.090.3410.2053.5628.750.3850.2670.95
LATINO818.080.5781.52199.7615.260.6421.88263.70
LATINO-PRO6518.580.5311.28163.4115.420.6101.67248.64
TReg20019.570.5290.38101.7618.280.5940.59135.35
SILO20014.920.7173.19380.7314.630.7233.00389.20
Compressed sensing
BAM330.630.1720.0112.8628.870.2210.0216.84
BAM★330.220.1810.0213.7328.360.2300.0318.48
RAM123.390.3860.3542.7423.200.4220.4749.31
LATINO816.080.6602.18250.9213.710.7002.24249.40
LATINO-PRO6516.240.6331.86248.3413.160.6792.08275.27
TReg20017.100.6591.05195.9516.270.6911.22224.85

FFHQ: 64 test images. BAM and BAM★ use three steps.

MethodNFEs σy = 0.025 σy = 0.05
PSNR↑ LPIPS↓ CMMD↓ FID↓ PSNR↑ LPIPS↓ CMMD↓ FID↓
Deblurring
BAM330.630.2320.0940.6329.630.2500.1044.98
BAM★331.170.2250.0847.3330.160.2450.0851.61
RAM131.070.3270.9483.3730.540.3511.1594.74
RAM★132.310.3160.8582.5831.530.3411.0792.38
LATINO825.080.3770.62100.2822.720.4280.85119.48
LATINO-PRO6525.890.3950.82111.2523.400.4561.17128.26
TReg20025.160.3940.82112.0724.280.4220.73117.41
SILO20025.800.3720.3892.8724.920.3750.2591.84
I2SB5031.200.2210.0136.8230.240.2400.0244.21
UD2M1229.390.2780.8365.8729.300.2880.7267.39
Super-resolution
BAM329.990.2490.1345.4828.840.2720.1950.42
BAM★330.480.2470.1356.4629.370.2680.1962.38
RAM132.490.3220.6087.4831.150.3581.02107.12
RAM★132.360.3130.8680.5531.010.3521.5293.02
LATINO825.530.4190.6699.4722.180.5101.13141.33
LATINO-PRO6526.670.4250.58107.2723.180.5290.90146.27
TReg20024.350.4470.63103.7522.280.5180.80122.72
SILO20026.120.3330.2071.7625.040.3710.2080.74
I2SB5026.200.3871.11151.2026.820.3931.02148.82
UD2M1229.630.2770.1897.6928.160.3050.3082.25
Inpainting
BAM332.900.1740.0421.3031.500.2280.0832.09
BAM★332.930.1650.0424.6931.570.2010.0532.40
RAM125.590.3991.97148.6325.550.3861.53126.45
RAM★133.360.2490.7743.2732.400.2881.1256.50
LATINO813.980.7733.25247.9413.430.7763.27248.21
LATINO-PRO6513.870.7802.53272.3413.440.7892.63304.99
TReg20014.460.7422.71224.3613.740.7562.80237.75
SILO20018.550.4970.87133.5517.460.5171.00139.38
I2SB5022.380.4911.35134.4818.890.5841.93174.26
UD2M1233.210.2160.2046.8129.290.2540.2145.54
Demosaicing
BAM336.450.1190.0310.9533.380.2020.0523.77
BAM★336.490.1140.0412.7033.940.1670.0419.25
RAM131.950.3330.57101.4830.780.3700.66110.64
RAM★137.360.1610.3418.7435.070.2260.4632.27
LATINO817.470.6672.73199.3614.160.7163.08214.16
LATINO-PRO6518.190.6351.76170.9913.120.7432.08255.00
TReg20021.150.5471.37125.3518.870.6261.78159.45
SILO20020.260.5101.43137.0021.040.4720.85118.79
Compressed sensing
BAM334.580.1510.0316.3732.410.2080.0625.52
BAM★334.910.1450.0319.1732.720.1850.0425.02
RAM124.470.4111.1775.1724.350.4611.2188.65
RAM★133.740.2621.1341.3532.720.3011.2954.27
LATINO814.550.7393.45217.9312.280.7583.59234.12
LATINO-PRO6514.730.7432.68232.9710.750.8233.04322.00
TReg20016.600.6901.95168.0715.460.7172.29197.12

AFHQ: 64 test images. BAM and BAM★ use three steps.

MethodNFEs σy = 0.025 σy = 0.05
PSNR↑ LPIPS↓ CMMD↓ FID↓ PSNR↑ LPIPS↓ CMMD↓ FID↓
Deblurring
BAM327.410.3360.2333.5226.890.3540.2635.62
BAM★326.850.2940.1823.8926.410.3110.1924.21
RAM127.850.4171.1752.1127.460.4441.4160.40
LATINO823.370.4400.4864.8121.480.4860.5071.87
LATINO-PRO6524.620.4941.4467.7722.630.5581.6092.28
TReg20023.490.4690.8976.6222.740.4970.9381.11
SILO20024.430.4200.3142.8824.050.4200.3044.82
Super-resolution
BAM327.240.3480.2934.6626.520.3710.4239.53
BAM★326.400.3080.2526.2625.680.3320.3129.22
RAM128.720.4040.8556.2927.950.4421.1767.87
LATINO823.270.4780.9270.5220.480.5601.3592.08
LATINO-PRO6525.310.4831.2963.4922.470.5771.67104.06
TReg20022.850.4930.9365.4621.090.5531.0976.40
SILO20024.600.3830.2535.2023.750.4190.2639.93
Inpainting
BAM329.320.2230.0716.0928.410.2770.1321.87
BAM★329.150.2090.0814.3327.980.2570.1219.36
RAM124.730.4191.1166.5524.680.4271.0062.86
LATINO814.200.7392.0492.7013.650.7372.0185.07
LATINO-PRO6513.300.7962.61293.1012.370.8172.93329.01
TReg20014.090.7502.33104.1713.470.7622.25100.78
SILO20020.090.5850.7164.5119.490.6311.0571.92
Demosaicing
BAM334.660.1290.037.8631.490.2150.0615.77
BAM★334.230.1250.067.5630.840.2040.1015.08
RAM130.670.3880.7362.0129.420.4380.8973.68
LATINO817.060.6502.35105.7313.910.7032.38101.36
LATINO-PRO6518.070.6141.8795.4613.830.7082.22160.66
TReg20018.830.5851.9297.0616.630.6602.35114.93
SILO20021.800.4660.3147.8421.770.4770.2847.71
Compressed sensing
BAM332.050.1730.0512.2030.100.2330.0918.23
BAM★330.970.1780.0812.0128.650.2410.1119.26
RAM123.970.4041.1238.9023.830.4461.3144.47
LATINO814.710.7002.1689.6712.480.7162.2585.35
LATINO-PRO6513.860.7692.74230.2711.160.8293.05297.84
TReg20015.220.7102.4692.9614.400.7352.5297.29

LSUN: 300 test images. BAM and BAM★ use three steps.

MethodNFEs σy = 0.025 σy = 0.05
PSNR↑ LPIPS↓ CMMD↓ FID↓ PSNR↑ LPIPS↓ CMMD↓ FID↓
Deblurring
BAM325.780.1541.8689.3525.380.1602.1590.52
BAM★326.160.1551.0249.4125.790.1601.2648.28
RAM125.970.2131.9277.5325.580.2392.2087.09
RAM★127.250.3041.8776.9026.590.3332.1186.54
COSIGN224.740.2151.0843.4024.600.2301.4344.41
UD2M1225.960.2791.6881.6123.450.3802.5598.44
CM4AI123.820.4192.29199.0423.330.4432.36213.70
Super-resolution
BAM325.670.1602.0795.9725.130.1702.38100.50
BAM★325.690.1501.0847.8025.130.1601.3849.74
RAM121.130.5252.48163.4820.830.5773.44200.71
RAM★127.090.3021.8073.4526.270.3422.2088.81
COSIGN221.730.2452.45113.8120.750.3702.90160.95
UD2M1223.500.4393.40114.5023.780.3150.6748.31
CM4AI125.780.3241.76128.7125.520.3201.57123.33
Inpainting
BAM328.130.0670.7765.9027.460.0901.1976.43
BAM★328.240.0600.4139.7427.510.0710.5538.77
RAM123.140.2242.55125.5423.300.2242.42120.08
RAM★128.640.2381.3153.2628.200.2611.5860.24
Demosaicing
BAM334.410.0280.4040.8831.550.0670.5953.57
BAM★334.660.0230.2927.6731.660.0480.3827.71
RAM130.100.1521.3169.6628.960.1821.6079.74
RAM★135.170.1140.4321.4532.910.1610.6133.69
Compressed sensing
BAM331.900.0440.4949.9330.080.0710.7560.07
BAM★332.340.0370.3229.8430.460.0530.4129.36
RAM122.650.2052.31109.2922.580.2382.58135.76
RAM★129.810.2461.6953.4029.210.2711.9162.45

Beyond natural images

Sparse-view CT. BAM extends to CT on LIDC-IDRI chest slices with 51 parallel-beam projections, a single-channel modality far from the natural images seen during pre-training. After finetuning on CT, three-step BAM samples recover fine lung vessels, and the pixelwise standard deviation over 64 draws concentrates on anatomical edges and vessels, at the same order of magnitude as the error of the posterior mean: a spatial uncertainty map at no extra training cost.

Ground truth
CT: Ground truth
Filtered backprojection
CT: Filtered backprojection
BAM sample
CT: BAM sample
BAM mean
CT: BAM mean
4×4 std. (N = 64)
CT: 4×4 std. (N = 64)
4×4 residual
CT: 4×4 residual

Left to right: ground truth, filtered backprojection, one BAM sample, posterior mean, standard deviation over 64 draws, and residual |GT − mean|. Yellow boxes and 4× strips show the same lung region in every panel.

  • Blind motion deblurring. With spatially varying kernels estimated by a pre-trained kernel predictor network, BAM outperforms both RAM and the original unrolled PnP architecture on the Köhler benchmark, although it was never trained on such operators.
  • Non-linear JPEG restoration. After finetuning, BAM restores noisy JPEG-compressed images (Q = 10), generating missing details and regularising noisy areas where zero-shot and unrolled methods generalise poorly.
  • Posterior mean. Averaging repeated BAM draws gives a Monte Carlo posterior-mean estimate whose PSNR exceeds that of the fine-tuned MMSE network RAM★ on four of five problems: samples are sharp individually yet accurate on average.

Limitations

BAM handles only additive Gaussian noise and linear or mildly non-linear forward operators. It is non-blind, so blind problems rely on an external operator estimate. It does not detect model misspecification, so under strong distribution shift it can return unreliable posteriors without warning. Finally, posterior sample quality was assessed empirically; formal guarantees for the learned conditional flow map remain open.

BibTeX

@misc{spagnoletti2026bam,
  title  = {{BAM!} {B}ayesian {A}nything {M}odel: a foundation model for generative computational imaging},
  author = {Spagnoletti, Alessio and Kemajou Mbakam, Charlesquin and Spence, Jonathan and Almansa, Andr{\'e}s and Pereyra, Marcelo},
  year   = {2026}
}