Materials and data ·

A good fit is not always a good prediction

A curve passes neatly through every measurement. Does that make it a reliable model? Sometimes it means the model has also learned the measurement noise.

Materials measurements contain variation. Specimens can differ, instruments have finite precision, and important processing variables may be missing. A flexible model can reproduce the observations without correctly describing what happens between them.

Separate the trend from the noise

For this example, I use an artificial relation between composition and hardness. Composition is the fraction of one component, from 0 to 1, and is dimensionless. Hardness is shown in HV, a conventional Vickers hardness value. The curve is invented for this demonstration; it is not data for a named alloy.

The underlying trend is H = 120 + 70x + 25 sin(2πx), where H is hardness and x is composition. Twelve observations are made by adding a fixed noise pattern to this trend. The noise control sets its scale in HV. These deviations are illustrative rather than measurements from an instrument.

Try a simple model and a flexible one

Polynomial degree controls flexibility. Degree 1 gives a straight line. A larger degree allows more curvature. Change the degree first with noise at zero, then repeat with noise present.

Fit the observations

Orange points: observations. Teal curve: underlying trend. Purple curve: fitted polynomial.

Two errors answer different questions

The training root-mean-square error, or training RMSE, compares the fitted curve with the twelve observations used to fit it. It has the same units as hardness, HV.

The trend RMSE compares the fitted curve with the known underlying trend on a dense grid of compositions. We can calculate this here because we created the trend. In an experiment, that noiseless trend is generally unknown. This demonstration therefore shows a distinction that is difficult to observe directly in real data.

A smaller training error does not guarantee a smaller prediction error. A highly flexible curve may follow the noisy observations closely while departing from the trend between them. At zero noise, extra flexibility can instead help represent genuine curvature. Complexity is useful when the data support it.

How would we check a real model?

Keep some observations out of the fitting process. Use them to evaluate predictions. If you repeatedly use the same held-out observations to choose a model, they have become part of that choice. A final independent test is useful after selecting the model.

The way we separate data matters. Multiple measurements from the same specimen should not casually be divided across training and test sets. Closely related specimens can make a model appear more transferable than it is. For a new alloy family or a new processing range, evaluate on data that represent that intended use.

Ask what the curve is allowed to do

A polynomial does not know whether a material changes phase, whether a property must remain positive, or whether the trend should be monotonic over a particular range. Those questions come from the material and the experiment.

A useful model balances mathematical flexibility, the available observations and physical knowledge. The best-looking fit is only one part of that assessment.

How the interactive example is calculated

The polynomial is fitted by least squares using a QR factorization. The calculation does not add a physics constraint. Noise remains fixed while degree changes. The continuous reference curve and its comparison grid are supplied by the synthetic example.

Read the calculation source