Standard deviation measures spread: how far a typical value strays from the average. Consider {4, 5, 5, 6} versus {0, 5, 5, 10} — both have a mean of 5, but the second dataset is far more scattered, and its larger standard deviation captures exactly that. In investing it quantifies volatility (a fund with high SD swings harder), in manufacturing it quantifies consistency (low SD means every part matches spec), in test scores it tells you whether the class clustered together or split into strong and weak groups, and in science it sets the error bars on every measurement.
The calculation has four steps: find the mean, square each value's distance from it (squaring penalizes large deviations and keeps everything positive), average those squared distances — that average is the variance — and take the square root to return to the original units. One subtlety: for a full population you divide by N, but for a sample you divide by N - 1 (Bessel's correction), because a sample's values cluster slightly tighter around the sample mean than around the true mean, so dividing by N would underestimate the spread. Use population when you have every value (all students in a class); use sample when your data is a subset (20 voters polled out of millions). With large datasets the N versus N-1 difference fades to nothing.