Start here ·
SVD, explained with a small picture
Singular value decomposition is a long name for a useful way of breaking an array of numbers into simpler parts. Let us start with a picture. We can understand the idea without starting with a matrix equation.
A picture is made of numbers
Zoom far enough into a digital picture and you see small squares called pixels. In a grayscale picture, each pixel has one number describing its brightness. I use 0 for black and 1 for white. Numbers between them give shades of gray. These brightness values have no physical units in this example.
The picture below is only eight pixels wide and eight pixels high. Its 64 brightness values form eight rows of numbers. That rectangular arrangement is called a matrix. SVD is a calculation we can apply to it.
Think of a rough sketch and its corrections
Suppose you draw a scene. You might begin with a rough arrangement of light and dark regions, then add corrections to make the picture closer to what you see. Something similar happens here: SVD gives us numerical patterns that we add together to recover the original picture.
I call these patterns layers in this explanation. Each layer adds to some pixels and may subtract from others. They are not transparent sheets placed over a photograph. We are adding and subtracting numbers at corresponding pixel positions.
The layers are ordered by their mathematical strength. The first makes the largest reduction in the total squared pixel error. The next makes a further reduction, and so on. A layer need not look like a familiar object. SVD does not know that this picture contains a roof or a door.
Add the layers yourself
Move the slider, or select “Play layers.” Compare the original, the rebuilt picture and the layer just added.
Loading the picture…
Original and rebuilt picture: black = 0 and white = 1, with the same scale. For the layer just added, red means add brightness, blue means subtract brightness, and white means no change. Its fixed scale is −1 to +1. The error is calculated from the numerical values, before any display adjustment.
Why are there two patterns in each layer?
Think of a curtain with a simple light-and-dark pattern. One set of numbers describes how the pattern changes from top to bottom. Another describes how it changes from left to right. Multiplying a row's number by a column's number gives the value at their intersection.
That is how one SVD layer is formed: a pattern down the rows, a pattern across the columns, and a number that sets the layer's strength. One layer has a restricted form, so it usually cannot reproduce a complicated picture by itself. Adding layers lets us describe more complicated arrangements.
This curtain comparison explains the multiplication. It does not mean that the algorithm looks for fabric, stripes or any other physical object.
The strength is like a volume control
Imagine mixing several sound tracks. A volume control changes how strongly each track contributes to the final sound. The strength assigned to an SVD layer plays a similar role. This nonnegative strength is called a singular value.
The comparison is about scaling and combining patterns. SVD does not automatically separate a recording into its instruments, and its image layers do not automatically identify physical features. A mathematically strong pattern might represent illumination, a large background region, or a feature we care about.
Keeping fewer layers is like keeping a simpler description
Suppose you describe a room to someone. A short description might give its broad arrangement. More detail makes the description closer to the actual room, but also longer. With SVD, keeping fewer layers gives a simpler numerical description of the picture.
Unlike an ordinary description, we have a precise measure of the difference: compare corresponding pixels, square their differences, and add them. For a fixed number of layers, keeping the largest singular values gives the smallest possible total squared error among matrices of that rank. Rank here is the number of independent row-and-column patterns needed to represent the matrix.
Small error does not necessarily mean that every important detail is preserved. A small pore or a thin interface may matter greatly in a material, even if it contributes little to the total pixel error. The choice of how many layers to keep depends on what we want to learn from the image.
What should you notice in this example?
Zero layers gives a black picture because we have added nothing. One layer gives a broad approximation. More layers improve the match. Four layers recover this particular picture to numerical precision. I deliberately chose a small, simple picture; four layers will not generally be enough for a microscope image.
The displayed RMSE is the root-mean-square error. It is found by averaging the squared pixel differences and taking the square root. It has the same dimensionless brightness scale as the image. A smaller value means closer numerical agreement.
The equation, if you want to see it
SVD writes the image matrix A as A = U Σ Vᵀ. A is the original array of pixel values. The columns of U contain the row-direction patterns. The rows of Vᵀ contain the column-direction patterns. The diagonal entries of Σ are the singular values, arranged from largest to smallest. All other entries of Σ are zero. The superscript T means transpose: rows and columns are exchanged.
Multiplying one column of U by its singular value and the matching row of Vᵀ gives one layer. Adding all the layers recovers A, apart from numerical rounding. In the code, I store the diagonal entries of Σ as a list named strengths rather than create a full diagonal matrix.
In this example, A and the singular values are dimensionless. U and V contain normalized, dimensionless weights. If A had physical units, its singular values would carry those units.
How does this connect to PCA?
The example above applies SVD to one image, with image rows and columns forming the matrix. In the PCA tutorial, I use a collection of images. Each whole image becomes one data row, and each pixel position becomes one feature column. I subtract the mean value of each feature before finding the directions of variation.
SVD is then a way to calculate those PCA directions. SVD itself is a matrix decomposition; PCA asks which directions explain the variation in a collection of observations. Their connection depends on centering the data and on what the matrix rows and columns represent.
There is also an important difference in the zero-component result. The SVD picture above starts from zero. In the centered PCA tutorial, reconstruction adds the mean image back, so zero principal components gives the mean image.
Next: build PCA with small pixel images
Run the small picture example in Python
Download the tested Python example. Install NumPy and Matplotlib with python -m pip install numpy matplotlib, then run python svd-picture.py. It prints the error for each layer count and displays the reconstructions. No image downloads are needed.