Population and sample standard deviation from a list of numbers.
Variance = average of squared differences from the mean. Standard deviation = √variance. Sample variance divides by (n−1) instead of n (Bessel's correction).
Variance measures how spread out a data set is by averaging the squared differences between each value and the mean ā squaring is necessary because otherwise the positive and negative differences would simply cancel out to zero for any data set, regardless of how spread out it actually is. But squaring the differences also squares the units: the variance of a set of exam scores comes out in "points squared," a unit with no intuitive real-world meaning. Standard deviation solves exactly this problem by taking the square root of variance, converting the measure back into the original units ā which is exactly why standard deviation, not variance, is the number that gets quoted in nearly every real-world context, from exam score spreads to investment volatility.
There's also a genuine, often-overlooked distinction between population standard deviation and sample standard deviation: when working with data from an entire population, the variance calculation divides by n (the total count), but when working with a sample meant to estimate a larger population's spread, the calculation divides by nā1 instead. This adjustment (called Bessel's correction) exists because a sample's own mean is calculated from the same data being measured, which subtly reduces the apparent spread ā dividing by nā1 instead of n corrects for that downward bias and gives a more accurate estimate of the true population standard deviation. Using the wrong one of the two formulas is a common, easy-to-miss statistics mistake, since both produce a plausible-looking number and the difference is usually small for large data sets but can matter meaningfully for small ones.
Because without squaring, the positive and negative differences from the mean would always cancel out to exactly zero for any data set, regardless of how spread out the values actually are. Squaring makes every deviation positive, so the total reflects genuine spread rather than canceling itself into a meaningless number.
Because standard deviation (the square root of variance) is expressed in the same units as the original data, making it directly interpretable. Variance is in squared units (e.g., 'points squared' for exam scores), which has no intuitive real-world meaning, so standard deviation is almost always the number actually reported and discussed.
Population standard deviation divides by n (the total count) and is used when you have data for an entire population. Sample standard deviation divides by nā1 instead, a correction (Bessel's correction) that accounts for the slight downward bias introduced when a sample's own mean is used to estimate spread within that same sample.
For large data sets, the difference between dividing by n versus nā1 is usually small. For small sample sizes, the difference can be meaningful, and using the wrong formula for the situation (population data treated as a sample, or vice versa) is a common statistics mistake worth double-checking before reporting a result.