PCA for materials data · Part 2 of four ·
Eigenvectors and eigenvalues: begin with a stretching sheet
Before we calculate PCA, we need two unfamiliar ideas: eigenvectors and eigenvalues. We can understand their meaning using a sheet, some arrows and a simple stretching rule.
What is an arrow telling us?
An arrow has a length and a direction. We can describe it by how far its tip extends horizontally and vertically from its starting point. We call this a vector. The arrow is a picture of the vector.
For example, an arrow starting at the center and ending 2 cm to the right and 2 cm above it has a horizontal part of 2 cm and a vertical part of 2 cm. Its length is not 4 cm: using Pythagoras, its length is √(2² + 2²), approximately 2.83 cm.
We stretch the sheet horizontally
Imagine a square sheet initially 10 cm wide and 10 cm high. We stretch it to 20 cm wide while its height remains 10 cm. Its center stays fixed. This is an idealized, uniform deformation: every horizontal distance from the center doubles, and every vertical distance remains unchanged. It is a geometric example, not a prediction of how an unconstrained real sheet responds to loading.
Try the stretching rule
Blue: horizontal arrow. Green: vertical arrow. Orange: diagonal arrow. All arrows start at the marked center. The sheet outlines show its boundaries. Change the factor and observe which arrows stay along their original lines.
We examine the numbers
| Original arrow parts (cm) | New arrow parts (cm) | What happens? |
|---|---|---|
| Horizontal: (2, 0) | (4, 0) | Length doubles; direction stays horizontal. |
| Vertical: (0, 2) | (0, 2) | Length and direction stay unchanged. |
| Diagonal: (2, 2) | (4, 2) | Length changes from 2.83 to 4.47 cm; direction changes. |
The diagonal arrow's angle above the horizontal changes from 45° to approximately 26.6°. The horizontal and vertical arrows remain along their original lines. This difference is what we want to notice.
What is a transformation?
A transformation is a specified rule for changing a vector. Our rule doubles the horizontal part and leaves the vertical part unchanged. We use the letter \(T\) as a short name for this rule.
If the horizontal and vertical parts are \(x\) and \(y\), measured in centimeters, the two instructions are:
The zeros mean that the vertical part does not contribute to the new horizontal part, and the horizontal part does not contribute to the new vertical part.
Why write a matrix?
We arrange the coefficients of these two instructions in a rectangular table. This table is a matrix. Its first row gives the horizontal instruction. Its second row gives the vertical instruction.
We write a vector as a column: horizontal part first, vertical part second. Matrix multiplication applies each row's instruction to that column:
The coefficients are dimensionless here. The arrow parts on both sides are in centimeters. The matrix represents the rule we already described; it does not introduce a different operation.
What is special about the horizontal and vertical arrows?
Let us return to the factor of 2. Our horizontal arrow changes from 2 cm to 4 cm, but stays horizontal. Our vertical arrow remains 2 cm long and stays vertical. The diagonal arrow changes from parts (2, 2) cm to parts (4, 2) cm. It turns toward the horizontal.
For the horizontal arrow, applying the matrix gives exactly the same result as multiplying the whole arrow by 2. For the vertical arrow, it gives the same result as multiplying the whole arrow by 1. We cannot do this for the diagonal arrow: multiplying it by 2 would give (4, 4) cm, whereas our matrix gives (4, 2) cm.
This distinction is important. The horizontal and vertical directions have a particularly simple response to this rule. Along either direction, we need just one number to describe the change.
Why the names eigenvector and eigenvalue?
The prefix eigen comes from German and means approximately “own.” Eigenvectors are also called characteristic vectors, and eigenvalues are called characteristic values. These are alternative names for the same mathematical quantities. Georgia Tech's Interactive Linear Algebra explains the word and the definition; MathWorld also uses the name characteristic vector.
In our example, “characteristic” helps us remember that we are describing the stretching rule's own behavior. We did not declare the horizontal direction special merely because we drew a horizontal axis. It is special for this rule because the rule multiplies a horizontal arrow without turning it. If we changed the rule, its characteristic directions could change.
A nonzero arrow with this property is an eigenvector of the matrix. The number multiplying it is its eigenvalue. Our horizontal arrow has eigenvalue 2. Our vertical arrow has eigenvalue 1. The diagonal arrow is not an eigenvector of this matrix.
There is no special arrow length that we must choose. A 1 cm horizontal arrow becomes 2 cm; a 3 cm horizontal arrow becomes 6 cm. Both have the same multiplying factor, 2. We are identifying a direction and its response, rather than one particular arrow length.
We can now write the statement in symbols
We use a bold letter \(\mathbf v\) for a nonzero vector, the matrix \(T\) for our rule, and the Greek letter \(\lambda\), pronounced “lambda,” for its multiplying factor. Writing \(T\mathbf v\) means apply the matrix's rule to the vector. Writing \(\lambda\mathbf v\) means multiply every component of the vector by the same number.
In words: applying the matrix produces the same result as multiplying this vector by one number. That is the eigenvector condition. It is true for our horizontal and vertical arrows, but not for our diagonal arrow.
What if the multiplying factor is negative, zero or one?
A negative eigenvalue reverses the arrow. Its absolute value determines the length multiplier. For example, −1 reverses the arrow without changing its length. The result remains along the same line, but points the other way. An eigenvalue of zero maps the nonzero input vector to zero. An eigenvalue of one leaves the vector unchanged.
If we set the horizontal factor to 1, the sheet does not change at all. Every nonzero arrow is then an eigenvector with eigenvalue 1. There is no single preferred direction in that case.
Can we see the characteristic directions?
So far, the special directions were horizontal and vertical. They need not coincide with our drawing axes. In this second example, we use a different rule. It doubles distances along a line at 30° above horizontal and leaves distances along its perpendicular line unchanged.
Press Play. The pale dots mark the original positions, and the dark dots move toward their transformed positions. The purple arrow starts horizontally. Watch whether it remains horizontal. Then reveal the two characteristic directions and play again.
Watch first, then reveal the directions
At the end, the orange arrow is twice its original length and still points at 30°. Its eigenvalue is 2. The green arrow stays along the perpendicular line and retains its length. Its eigenvalue is 1. The purple arrow turns, so it is not an eigenvector of this rule.
All coordinates and arrow lengths in this animation are dimensionless. The slider shows a smooth progression between the original and final positions; it is not elapsed physical time or a simulation of a material deforming.
A unit vector is a vector with length one. It specifies a direction. In this animation, a unit arrow has length one in our dimensionless coordinate scale.
How do the final numbers work?
A unit arrow at 30° has horizontal and vertical parts approximately (0.866, 0.500). The rule maps it to (1.732, 1.000). Both parts have doubled, so the arrow remains along the same line. A horizontal unit arrow starts at (1, 0), but ends at approximately (1.750, 0.433). Its vertical part is now nonzero, so it has turned.
The final rule has matrix \(\begin{pmatrix}1.75&\sqrt{3}/4\\\sqrt{3}/4&1.25\end{pmatrix}\). Each row gives the rule for one new coordinate, just as in our earlier matrix example. Its eigenvector directions are 30° and 120°, and its eigenvalues are 2 and 1.
This original interactive drawing illustrates the same idea as Wikipedia's animated eigenvector examples: points on characteristic lines stay on those lines during the transformation. We use our own grid, controls and numerical example here.
A materials example: does current follow the electric field?
Our spreadsheet includes electrical conductivity. We usually introduce conductivity as one number. Some materials conduct differently in different directions, so we need more than one number to describe their response.
Explore a numerical conductivity example
Imagine a material with conductivity 6 MS/m along a horizontal principal direction and 2 MS/m along a vertical principal direction. These are invented values for a two-dimensional example, not measurements of a named material. MS/m means megasiemens per meter, or one million siemens per meter.
The electric field describes the electrical driving force per unit charge. We measure it in volts per meter, V/m. The current density describes electric current per unit area perpendicular to the flow. We measure it in amperes per square meter, A/m².
For this example, the current density's horizontal part equals 6 million times the electric field's horizontal part. Its vertical part equals 2 million times the field's vertical part, when the field is in V/m and current density is in A/m².
| Electric-field parts (V/m) | Current-density parts (A/m²) | What do we observe? |
|---|---|---|
| Horizontal: (1, 0) | (6,000,000, 0) | Both point horizontally. |
| Vertical: (0, 1) | (0, 2,000,000) | Both point vertically. |
| Diagonal: (1, 1) | (6,000,000, 2,000,000) | The field is at 45°; the current density is at approximately 18.4°. |
Along the horizontal and vertical directions, current density is parallel to the electric field. Those are eigenvector directions of the conductivity matrix. The corresponding eigenvalues are the principal conductivities: 6 and 2 MS/m. Along the diagonal direction, the current density turns toward the direction with higher conductivity.
Notice that the field and current density are different physical quantities. The conductivity eigenvalues therefore have units, S/m. They are not dimensionless length multipliers like 2 and 1 in our stretching example.
This gives us a practical reason to find eigenvectors and eigenvalues: they identify directions in which a material's response can be described by one coefficient, without mixing the perpendicular direction into the result. MIT's materials tensor-properties lecture discusses why current and electric field need not point in the same direction in an anisotropic conductor.
Another materials example: why does the plane matter?
We pull a specimen horizontally. Suppose the tensile stress is 100 MPa. Stress means force per unit area. MPa means megapascal; 1 MPa is one million newtons per square meter, or one newton per square millimeter. We assume uniform stress and no vertical or out-of-plane loading in this example.
Imagine a small plane passing through the material. We are not cutting the specimen or creating a crack. We are asking what force per unit area the material on one side exerts on the other. This vector is called traction. Its part perpendicular to the plane is the normal stress. Its part along the plane is the shear stress.
Before moving the slider, ask: if the pulling direction stays the same, will these two parts stay the same when we turn the plane? The dashed blue arrow is the plane normal, a direction perpendicular to the plane. The angle is measured from the horizontal pulling direction to this normal, not to the plane itself.
Rotate the plane inside the specimen
The arrows for traction and its two stress components use the same scale. The dashed normal indicates direction only. The specimen outline stays fixed because we are examining stress, not calculating deformation.
At 0°, the plane is vertical, perpendicular to the pull. Its normal stress is 100 MPa and its shear stress is zero. At 45°, the normal stress and the shear-stress magnitude are both 50 MPa. At 90°, the plane is horizontal. Both stresses are zero in this particular loading example. Turning the plane changes the stresses on that plane even though the applied loading has not changed.
How do we calculate the displayed stresses?
We use the symbol \(\theta\), pronounced theta, for the angle of the plane normal. We use \(\sigma_n\) for normal stress and \(|\tau|\) for the magnitude of shear stress, both in MPa. For this 100 MPa uniaxial loading:
At 45°, both sine and cosine equal approximately 0.707. Their product is 0.5, so both displayed values are 50 MPa. We report shear magnitude here; the shear arrow shows its direction on the selected side of the plane.
Where do eigenvectors enter?
The stress matrix maps a unit plane normal to the traction on that plane. For our two-dimensional example, the stress matrix has horizontal entry 100 MPa, vertical entry zero and off-diagonal entries zero. If the normal is horizontal, the traction is also horizontal. If the normal is vertical, the traction is zero. These normals are eigenvector directions of the stress matrix.
We call them principal stress directions. On a plane perpendicular to a principal direction, shear traction is zero. The eigenvalues are the principal stresses, 100 MPa and zero here. Zero is a valid eigenvalue: it means zero traction for that nonzero normal vector. We are finding special directions of the stress at a point, not directions in which the specimen must stretch.
How does this help us think about deformation and failure?
Normal tensile stress tends to separate the two sides of a plane. Shear stress tends to slide them past each other. This distinction helps us ask whether an existing crack may open or whether a crystal may deform by slip. Slip is plastic deformation produced by dislocations moving along particular crystallographic planes and directions.
For a specified slip plane and slip direction, we resolve the traction along that direction. This component is the resolved shear stress. Under the simple Schmid-law description, slip begins when its magnitude reaches the critical resolved shear stress for that slip system. The critical value depends on the material and conditions. A crystallographic slip direction need not be a principal stress direction. Our rotating plane explores possible orientations; an actual crystal permits particular slip systems.
A 50 MPa shear stress does not by itself tell us that the specimen will fail. We also need the material's resistance and the relevant mechanism. Plastic yielding is not the same as fracture. For many ductile metals, yielding is described using a combination of the principal stresses, such as the von Mises criterion. Brittle fracture also depends on flaws and crack geometry. The useful lesson is that one applied tensile stress produces different normal and shear stresses on differently oriented planes.
For the underlying stress transformations and yielding examples, see David Roylance's MIT teaching chapter on tensor transformations and chapter on yield and plastic flow.
We return to the PCA question
In the stretching example, the eigenvalue tells us the length multiplier. In the stress example, it gives a principal stress. In PCA, it gives the variance along a particular combination of features. The mathematical idea is the same, but the meaning of the matrix and its eigenvalues changes with the problem.
Our five measurement columns do not have to be the most useful axes for examining the specimens. We want to find combinations that describe the main differences between them. Eigenvectors will specify these combinations, and eigenvalues will tell us the variance associated with each one.
We must first construct the appropriate matrix from the measurements. It will be a covariance matrix, not a stretching, conductivity or stress matrix. The next article explains the mean, variance and covariance using numbers from our spreadsheet.
For another visual explanation of eigenvectors and eigenvalues, see 3Blue1Brown's lesson.
Acknowledgment and further reading
We consulted Jonathon Shlens's A Tutorial on Principal Component Analysis (2005, Version 2) while preparing this series. It explains PCA through changes of basis, covariance and the relationship with singular value decomposition. A later version is available on arXiv.