Oliver Wakeford
All projects
Four digits generated by DRAW with attention, shown at each of its ten steps. Each starts as a grey blur and sharpens into a handwritten digit.
DRAW with attention, one sample per row, steps 1 to 10.
Deep LearningShipped2026

VAE and DRAW, Compared

DRAW 93.8 against VAE 102.9 test NELBO on MNIST

Generative ModelsVariational InferenceAttentionAblation

Built with

PythonPyTorchNumPypytest

This was the M1 Deep Learning mini-project, done in a group of three. We implemented a vanilla VAE and DRAW (Gregor et al., 2015) from scratch in PyTorch, with DRAW's Gaussian filterbank attention switched on by a flag, and trained five configurations on MNIST.

More steps gave a lower loss. Holding the attention model fixed and only changing the number of steps, test NELBO (lower is better) went from 134.5 at one step to 105.0 at five. DRAW without attention finished on 93.8 against 102.9 for the VAE, though it also has four times the VAE's parameters, so not all of that gap is the recurrence.

With attention on, DRAW scored worse on MNIST. At the same 30 epochs it sat at 105.5, behind the VAE and well behind DRAW without it. Given twice the training it got to 99.7. It does that with a third of the parameters, because it only ever writes a 5×5 patch, and on a 28×28 image that saving doesn't buy much. Its samples still look crisper than the others', with cleaner stroke endings, even where the likelihood says otherwise.

It's one seed and no hyperparameter search, and we left the pixels continuous, so these numbers can't be set next to the paper's. After the course I put it on GitHub with pinned dependencies and smoke tests for both models and the filterbank, which CI runs. The whole set of runs takes about half an hour on a laptop. The six-page report has the curves and samples.

What this does not show

  • A group project of three.
  • One seed, no hyperparameter search.
  • The configurations differ in parameter count, and the best attention run had twice the training. Only the epoch-30 attention figure compares like with like.
  • MNIST only, with pixels left continuous, so the numbers aren't comparable with the DRAW paper's.
  • The table values come from the original training runs and haven't been re-run since. Training was on Apple MPS, which isn't bit-for-bit deterministic.