Variational autoencoder explained

In machine learning, a variational autoencoder (VAE) is an artificial neural network architecture introduced by Diederik P. Kingma and Max Welling.[1] It is part of the families of probabilistic graphical models and variational Bayesian methods.[2]

In addition to being seen as an autoencoder neural network architecture, variational autoencoders can also be studied within the mathematical formulation of variational Bayesian methods, connecting a neural encoder network to its decoder through a probabilistic latent space (for example, as a multivariate Gaussian distribution) that corresponds to the parameters of a variational distribution.

Thus, the encoder maps each point (such as an image) from a large complex dataset into a distribution within the latent space, rather than to a single point in that space. The decoder has the opposite function, which is to map from the latent space to the input space, again according to a distribution (although in practice, noise is rarely added during the decoding stage). By mapping a point to a distribution instead of a single point, the network can avoid overfitting the training data. Both networks are typically trained together with the usage of the reparameterization trick, although the variance of the noise model can be learned separately.

Although this type of model was initially designed for unsupervised learning,[3] [4] its effectiveness has been proven for semi-supervised learning[5] [6] and supervised learning.[7]

Overview of architecture and operation

A variational autoencoder is a generative model with a prior and noise distribution respectively. Usually such models are trained using the expectation-maximization meta-algorithm (e.g. probabilistic PCA, (spike & slab) sparse coding). Such a scheme optimizes a lower bound of the data likelihood, which is usually intractable, and in doing so requires the discovery of q-distributions, or variational posteriors. These q-distributions are normally parameterized for each individual data point in a separate optimization process. However, variational autoencoders use a neural network as an amortized approach to jointly optimize across data points. This neural network takes as input the data points themselves, and outputs parameters for the variational distribution. As it maps from a known input space to the low-dimensional latent space, it is called the encoder.

The decoder is the second neural network of this model. It is a function that maps from the latent space to the input space, e.g. as the means of the noise distribution. It is possible to use another neural network that maps to the variance, however this can be omitted for simplicity. In such a case, the variance can be optimized with gradient descent.

To optimize this model, one needs to know two terms: the "reconstruction error", and the Kullback–Leibler divergence (KL-D). Both terms are derived from the free energy expression of the probabilistic model, and therefore differ depending on the noise distribution and the assumed prior of the data. For example, a standard VAE task such as IMAGENET is typically assumed to have a gaussianly distributed noise; however, tasks such as binarized MNIST require a Bernoulli noise. The KL-D from the free energy expression maximizes the probability mass of the q-distribution that overlaps with the p-distribution, which unfortunately can result in mode-seeking behaviour. The "reconstruction" term is the remainder of the free energy expression, and requires a sampling approximation to compute its expectation value.[8]

More recent approaches replace Kullback–Leibler divergence (KL-D) with various statistical distances, see see section "Statistical distance VAE variants" below..

Formulation

From the point of view of probabilistic modeling, one wants to maximize the likelihood of the data

x

by their chosen parameterized probability distribution

p\theta(x)=p(x|\theta)

. This distribution is usually chosen to be a Gaussian

N(x|\mu,\sigma)

which is parameterized by

\mu

and

\sigma

respectively, and as a member of the exponential family it is easy to work with as a noise distribution. Simple distributions are easy enough to maximize, however distributions where a prior is assumed over the latents

z

results in intractable integrals. Let us find

p\theta(x)

via marginalizing over

z

.

p\theta(x)=\intzp\theta({x,z})dz,

where

p\theta({x,z})

represents the joint distribution under

p\theta

of the observable data

x

and its latent representation or encoding

z

. According to the chain rule, the equation can be rewritten as

p\theta(x)=\intzp\theta({x|z})p\theta(z)dz

In the vanilla variational autoencoder,

z

is usually taken to be a finite-dimensional vector of real numbers, and

p\theta({x|z})

to be a Gaussian distribution. Then

p\theta(x)

is a mixture of Gaussian distributions.

It is now possible to define the set of the relationships between the input data and its latent representation as

p\theta(z)

p\theta(x|z)

p\theta(z|x)

Unfortunately, the computation of

p\theta(z|x)

is expensive and in most cases intractable. To speed up the calculus to make it feasible, it is necessary to introduce a further function to approximate the posterior distribution as

q\phi({z|x})p\theta({z|x})

with

\phi

defined as the set of real values that parametrize

q

. This is sometimes called amortized inference, since by "investing" in finding a good

q\phi

, one can later infer

z

from

x

quickly without doing any integrals.

In this way, the problem is to find a good probabilistic autoencoder, in which the conditional likelihood distribution

p\theta(x|z)

is computed by the probabilistic decoder, and the approximated posterior distribution

q\phi(z|x)

is computed by the probabilistic encoder.

Parametrize the encoder as

E\phi

, and the decoder as

D\theta

.

Evidence lower bound (ELBO)

See main article: Evidence lower bound.

As in every deep learning problem, it is necessary to define a differentiable loss function in order to update the network weights through backpropagation.

For variational autoencoders, the idea is to jointly optimize the generative model parameters

\theta

to reduce the reconstruction error between the input and the output, and

\phi

to make

q\phi({z|x})

as close as possible to

p\theta(z|x)

. As reconstruction loss, mean squared error and cross entropy are often used.

As distance loss between the two distributions the Kullback–Leibler divergence

DKL(q\phi({z|x})\parallelp\theta({z|x}))

is a good choice to squeeze

q\phi({z|x})

under

p\theta(z|x)

.[8] [9]

The distance loss just defined is expanded as

\begin{align} DKL(q\phi({z|x})\parallelp\theta({z|x}))&=

E
z\simq\phi(|x)

\left[ln

q\phi(z|x)
p\theta(z|x)

\right]\\ &=

E
z\simq\phi(|x)

\left[ln

q\phi({z|x
)p

\theta(x)}{p\theta(x,z)}\right]\\ &=lnp\theta(x)+

E
z\simq\phi(|x)

\left[ln

q\phi({z|x
)}{p

\theta(x,z)}\right] \end{align}

Now define the evidence lower bound (ELBO):L_(x) := \mathbb E_ \left[\ln \frac{p_\theta(x, z)}{q_\phi({z| x})}\right] = \ln p_\theta(x) - D_(q_\phi\parallel p_\theta) Maximizing the ELBO\theta^*,\phi^* = \underset\operatorname \, L_(x) is equivalent to simultaneously maximizing

lnp\theta(x)

and minimizing

DKL(q\phi({z|x})\parallelp\theta({z|x}))

. That is, maximizing the log-likelihood of the observed data, and minimizing the divergence of the approximate posterior

q\phi(|x)

from the exact posterior

p\theta(|x)

.

The form given is not very convenient for maximization, but the following, equivalent form, is:L_(x) = \mathbb E_ \left[\ln p_\theta(x|z)\right] - D_(q_\phi\parallel p_\theta(\cdot)) where

lnp\theta(x|z)

is implemented as
-1
2

\|x-

2
D
2
, since that is, up to an additive constant, what

x\simlN(D\theta(z),I)

yields. That is, we model the distribution of

x

conditional on

z

to be a Gaussian distribution centered on

D\theta(z)

. The distribution of

q\phi(z|x)

and

p\theta(z)

are often also chosen to be Gaussians as

z|x\simlN(E\phi(x),

2I)
\sigma
\phi(x)
and

z\simlN(0,I)

, with which we obtain by the formula for KL divergence of Gaussians:L_(x) = -\frac 12\mathbb E_ \left[\|x - D_\theta(z)\|_2^2\right] - \frac 12 \left(N\sigma_\phi(x)^2 + \|E_\phi(x)\|_2^2 - 2N\ln\sigma_\phi(x) \right) + Const Here

N

is the dimension of

z

. For a more detailed derivation and more interpretations of ELBO and its maximization, see its main page.

Reparameterization

To efficiently search for \theta^*,\phi^* = \underset\operatorname \, L_(x) the typical method is gradient ascent.

It is straightforward to find\nabla_\theta \mathbb E_ \left[\ln \frac{p_\theta(x, z)}{q_\phi({z| x})}\right]= \mathbb E_ \left[\nabla_\theta \ln \frac{p_\theta(x, z)}{q_\phi({z| x})}\right] However, \nabla_\phi \mathbb E_ \left[\ln \frac{p_\theta(x, z)}{q_\phi({z| x})}\right] does not allow one to put the

\nabla\phi

inside the expectation, since

\phi

appears in the probability distribution itself. The reparameterization trick (also known as stochastic backpropagation[10]) bypasses this difficulty.[8] [11] [12]

The most important example is when

z\simq\phi(|x)

is normally distributed, as

lN(\mu\phi(x),\Sigma\phi(x))

.

This can be reparametrized by letting

\boldsymbol{\varepsilon}\siml{N}(0,\boldsymbol{I})

be a "standard random number generator", and construct

z

as

z=\mu\phi(x)+L\phi(x)\epsilon

. Here,

L\phi(x)

is obtained by the Cholesky decomposition:\Sigma_\phi(x) = L_\phi(x)L_\phi(x)^T Then we have\nabla_\phi \mathbb E_ \left[\ln \frac{p_\theta(x, z)}{q_\phi({z| x})}\right] = \mathbb _\left[\nabla_\phi \ln {\frac {p_{\theta }(x, \mu_\phi(x) + L_\phi(x)\epsilon)}{q_{\phi }(\mu_\phi(x) + L_\phi(x)\epsilon | x)}}\right] and so we obtained an unbiased estimator of the gradient, allowing stochastic gradient descent.

Since we reparametrized

z

, we need to find

q\phi(z|x)

. Let

q0

be the probability density function for

\epsilon

, then \ln q_\phi(z | x) = \ln q_0 (\epsilon) - \ln|\det(\partial_\epsilon z)|where

\partial\epsilonz

is the Jacobian matrix of

\epsilon

with respect to

z

. Since

z=\mu\phi(x)+L\phi(x)\epsilon

, this is \ln q_\phi(z | x) = -\frac 12 \|\epsilon\|^2 - \ln|\det L_\phi(x)| - \frac n2 \ln(2\pi)

Variations

Many variational autoencoders applications and extensions have been used to adapt the architecture to other domains and improve its performance.

\beta

-VAE is an implementation with a weighted Kullback–Leibler divergence term to automatically discover and interpret factorised latent representations. With this implementation, it is possible to force manifold disentanglement for

\beta

values greater than one. This architecture can discover disentangled latent factors without supervision.[13] [14]

The conditional VAE (CVAE), inserts label information in the latent space to force a deterministic constrained representation of the learned data.[15]

Some structures directly deal with the quality of the generated samples[16] [17] or implement more than one latent space to further improve the representation learning.

Some architectures mix VAE and generative adversarial networks to obtain hybrid models.[18] [19] [20]

Statistical distance VAE variants

After the initial work of Diederik P. Kingma and Max Welling.[21] several procedures were proposed to formulate in a more abstract way the operation of the VAE. In these approaches the loss function is composed of two parts :

x\mapstoD\theta(E\psi(x))

is as close to the identity map as possible; the sampling is done at run time from the empirical distribution

Preal

of objects available (e.g., for MNIST or IMAGENET this will be the empirical probability law of all images in the dataset). This gives the term:
E
x\simPreal

\left[\|x-D\theta(E\phi(x))\|

2\right]
2
.

Preal

is passed through the encoder

E\phi

, we recover the target distribution, denoted here

\mu(dz)

that is usually taken to be a Multivariate normal distribution. We will denote

E\phi\sharpPreal

this pushforward measure which in practice is just the empirical distribution obtained by passing all dataset objects through the encoder

E\phi

. In order to make sure that

E\phi\sharpPreal

is close to the target

\mu(dz)

, a Statistical distance

d

is invoked and the term

d\left(\mu(dz),E\phi\sharpPreal\right)2

is added to the loss.

We obtain the final formula for the loss: L_ = \mathbb_ \left[\|x - D_\theta(E_\phi(x))\|_2^2\right]+d \left(\mu(dz), E_\phi \sharp \mathbb^ \right)^2

The statistical distance

d

requires special properties, for instance is has to be posses a formula as expectation because the loss function will need to be optimized by stochastic optimization algorithms. Several distances can be chosen and this gave rise to several flavors of VAEs:

See also

Further reading

Notes and References

  1. Kingma . Diederik P. . Auto-Encoding Variational Bayes . 2022-12-10 . Welling . Max. stat.ML . 1312.6114 .
  2. Book: Lucas . Pinheiro Cinelli . Matheus . Araújo Marins . Eduardo Antônio . Barros da Silva . Sérgio . Lima Netto . 1 . Variational Methods for Machine Learning with Applications to Deep Networks . Springer . 2021 . 111–149 . Variational Autoencoder . 978-3-030-70681-4 . https://books.google.com/books?id=N5EtEAAAQBAJ&pg=PA111 . 10.1007/978-3-030-70679-1_5 . 240802776 .
  3. Dilokthanakul . Nat . Mediano . Pedro A. M. . Garnelo . Marta . Lee . Matthew C. H. . Salimbeni . Hugh . Arulkumaran . Kai . Shanahan . Murray . Deep Unsupervised Clustering with Gaussian Mixture Variational Autoencoders . 2017-01-13 . cs.LG . 1611.02648.
  4. Book: Hsu . Wei-Ning . Zhang . Yu . Glass . James . 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) . Unsupervised domain adaptation for robust speech recognition via variational autoencoder-based data augmentation . December 2017 . 16–23 . 10.1109/ASRU.2017.8268911 . 1707.06265 . 978-1-5090-4788-8 . 22681625 . https://ieeexplore.ieee.org/document/8268911.
  5. Book: Ehsan Abbasnejad . M. . Dick . Anthony . van den Hengel . Anton . Infinite Variational Autoencoder for Semi-Supervised Learning . 2017 . 5888–5897 .
  6. Xu . Weidi . Sun . Haoze . Deng . Chao . Tan . Ying . Variational Autoencoder for Semi-Supervised Text Classification . Proceedings of the AAAI Conference on Artificial Intelligence . 2017-02-12 . 31 . 1 . 10.1609/aaai.v31i1.10966 . 2060721 . en. free .
  7. Kameoka . Hirokazu . Li . Li . Inoue . Shota . Makino . Shoji . Supervised Determined Source Separation with Multichannel Variational Autoencoder . Neural Computation . 2019-09-01 . 31 . 9 . 1891–1914 . 10.1162/neco_a_01217 . 31335290 . 198168155 .
  8. Kingma . Diederik P. . Welling . Max . Auto-Encoding Variational Bayes . 2013-12-20 . stat.ML . 1312.6114.
  9. News: From Autoencoder to Beta-VAE . Lil'Log . en . 2018-08-12.
  10. Rezende . Danilo Jimenez . Mohamed . Shakir . Wierstra . Daan . 2014-06-18 . Stochastic Backpropagation and Approximate Inference in Deep Generative Models . International Conference on Machine Learning . en . PMLR . 1278–1286. 1401.4082 .
  11. Bengio. Yoshua. Courville. Aaron. Vincent. Pascal. Representation Learning: A Review and New Perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence. 2013. 35. 8. 1798–1828. 10.1109/TPAMI.2013.50. 23787338. 1939-3539. 1206.5538. 393948.
  12. Kingma. Diederik P.. Rezende. Danilo J.. Mohamed. Shakir. Welling. Max. 2014-10-31. Semi-Supervised Learning with Deep Generative Models. cs.LG. 1406.5298.
  13. Higgins. Irina. Matthey. Loic. Pal. Arka. Burgess. Christopher. Glorot. Xavier. Botvinick. Matthew. Mohamed. Shakir. Lerchner. Alexander. 2016-11-04. beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework. en. NeurIPS.
  14. Burgess. Christopher P.. Higgins. Irina. Pal. Arka. Matthey. Loic. Watters. Nick. Desjardins. Guillaume. Lerchner. Alexander. 2018-04-10. Understanding disentangling in β-VAE. stat.ML. 1804.03599.
  15. Sohn. Kihyuk. Lee. Honglak. Yan. Xinchen. 2015-01-01. Learning Structured Output Representation using Deep Conditional Generative Models. en. NeurIPS.
  16. Dai. Bin. Wipf. David. 2019-10-30. Diagnosing and Enhancing VAE Models. cs.LG. 1903.05789.
  17. Dorta. Garoe. Vicente. Sara. Agapito. Lourdes. Campbell. Neill D. F.. Simpson. Ivor. 2018-07-31. Training VAEs Under Structured Residuals. stat.ML. 1804.01050.
  18. Larsen. Anders Boesen Lindbo. Sønderby. Søren Kaae. Larochelle. Hugo. Winther. Ole. 2016-06-11. Autoencoding beyond pixels using a learned similarity metric. International Conference on Machine Learning. en. PMLR. 1558–1566. 1512.09300.
  19. Bao. Jianmin. Chen. Dong. Wen. Fang. Li. Houqiang. Hua. Gang. 2017. CVAE-GAN: Fine-Grained Image Generation Through Asymmetric Training. 2745–2754. cs.CV. 1703.10155.
  20. Gao. Rui. Hou. Xingsong. Qin. Jie. Chen. Jiaxin. Liu. Li. Zhu. Fan. Zhang. Zhao. Shao. Ling. 2020. Zero-VAE-GAN: Generating Unseen Features for Generalized and Transductive Zero-Shot Learning. IEEE Transactions on Image Processing. 29. 3665–3680. 10.1109/TIP.2020.2964429. 31940538. 2020ITIP...29.3665G. 210334032. 1941-0042.
  21. 1312.6114 . stat.ML . Diederik P. . Kingma . Max . Welling . Auto-Encoding Variational Bayes . 2022-12-10.
  22. Kolouri . Soheil . Pope . Phillip E. . Martin . Charles E. . Rohde . Gustavo K. . 2019 . Sliced Wasserstein Auto-Encoders . International Conference on Learning Representations . ICPR . International Conference on Learning Representations.
  23. Turinici . Gabriel . 2021 . Radon-Sobolev Variational Auto-Encoders . Neural Networks . 141 . 294–305 . 1911.13135 . 10.1016/j.neunet.2021.04.018 . 0893-6080 . 33933889.
  24. 1705.02239 . A. . Gretton . Y. . Li . A Polya Contagion Model for Networks . 2017 . Swersky . K. . Zemel . R. . Turner . R.. IEEE Transactions on Control of Network Systems . 5 . 4 . 1998–2010 . 10.1109/TCNS.2017.2781467 .
  25. 1711.01558 . I. . Tolstikhin . O. . Bousquet . Wasserstein Auto-Encoders . 2018 . Gelly . S. . Schölkopf . B.. stat.ML .
  26. 1901.02401 . C. . Louizos . X. . Shi . Kernelized Variational Autoencoders . 2019 . Swersky . K. . Li . Y. . Welling . M.. astro-ph.CO .