Variation, Standardization and Relative Dispersion
Range, interquartile range, absolute deviations, descriptive and sample variance, group decomposition, standard scores and coefficients of variation.
§6.1 Why a centre needs a description of spread
Location measures describe where observations are situated but cannot tell us how consistently they occupy that location. The datasets 4, 5, 6 and 1, 5, 9 both have mean five and median five. Their practical implications can be very different. A service with waiting times close to five minutes is more predictable than one whose waits vary widely around five. A manufacturing process can meet a target average while producing too many items far from that target. Variation is therefore a property to describe, not merely an inconvenience around an average.
Spread must refer to a specific variable and a specific collection of eligible units. Variation in household expenditure is not the same as variation in expenditure per household member. Variation in individual scores is not the same as variation in school means. Aggregation can make values look more stable because within-group differences have been removed from view. Whenever a report gives a standard deviation or range, identify what each observation represents and whether the values are individual measurements or summaries.
Several measures of spread serve different purposes. The range describes the distance between the extremes. The interquartile range describes the width between two central quantiles. Absolute-deviation measures use distances without squaring. Variance and standard deviation use squared deviations and support useful algebraic decompositions. Relative measures compare spread with a reference scale such as the mean. No single number preserves all features of a distribution, so a graph and a location measure should accompany a spread summary when possible.
Units help interpret these quantities. A range of waiting times measured in minutes is measured in minutes. A variance of those waiting times is measured in squared minutes, while a standard deviation returns to minutes. A coefficient of variation is dimensionless under its appropriate ratio-scale interpretation. Confusing variance with standard deviation can produce a meaningless comparison between a squared quantity and an ordinary target or threshold.
Variation can have several sources. Actual differences between people or objects may coexist with repeated-measurement error, day-to-day process variation, rounding, and changes in operating conditions. A single descriptive standard deviation mixes whatever sources are present in the recorded values. It cannot separate them automatically. A later experimental-design or measurement course uses structured repeated data and models to distinguish variance components. Here, we describe the observed spread while recording the context that generated it.
Small samples can produce unstable spread summaries. The range can change dramatically when one additional observation exceeds a previous extreme. A sample standard deviation is also affected by unusual values and depends on which units entered the sample. These statements concern sensitivity and sampling variability; they do not imply that a calculated descriptive value is arithmetically uncertain. The recorded-data summary can be exactly reproducible while its relevance to a broader population remains uncertain.
Missingness can alter spread in either direction. Excluding high or low observations narrows the recorded distribution, but missing central observations can make the remaining values appear more dispersed. A complete-case standard deviation does not automatically describe all eligible units. Retain the missing count and the selection rule, just as for a mean. If different variables use different complete-case sets, their spread comparisons may involve different people or records.
In this unit, we distinguish a finite collection's descriptive variance using denominator n from the sample variance using denominator n minus one. Both are valid defined quantities, but they answer different computational questions and have different statistical roles. Software labels such as variance are often too terse to identify the convention. The textbook and its simulations display both the denominator and the eligible count when a distinction matters.
§6.2 Range, interquartile range and robust spread
The range is the maximum observation minus the minimum. It is easy to calculate and relates directly to the full observed span. For 2, 4, 4, 7, and 13, the range is eleven minutes. It uses only two observations, so it ignores how the remaining values are arranged. The samples 2, 2, 2, 2, 13 and 2, 7, 7, 7, 13 have the same range even though their concentrations differ substantially.
The range is sensitive to extremes and to sample size. Adding observations can never reduce the range of the original collection, although it can leave it unchanged. Comparing ranges from samples of five and five thousand therefore requires caution. A larger sample has more opportunities to include extreme values. A range describes the observed span; it does not estimate an immutable population boundary unless additional knowledge supports that interpretation.
The interquartile range, abbreviated IQR, is Q3 minus Q1. This book uses type-seven quartiles unless a different convention is explicitly declared. For sorted values 1, 2, 3, 4, 5, 6, 7, and 30, Q1 is 2.75 and Q3 is 6.25, giving IQR 3.5. The range is twenty-nine. The contrast shows that an upper extreme can dominate the full span while the central quantile width stays comparatively modest.
The IQR measures a central width, not the total dispersion. Two datasets can have the same IQR but very different tail behaviour. It is also possible for the IQR to be zero when many observations take the same value, even if a smaller set of observations differs greatly. A zero IQR means the chosen quartile locations coincide under the declared convention. It does not always mean every observation is identical.
The semi-interquartile range is half the IQR and is sometimes called the quartile deviation. State the name and formula because the word deviation alone can refer to several different measures. The associated quartile coefficient, (Q3 minus Q1) divided by (Q3 plus Q1), requires a nonzero denominator and a meaningful scale for the ratio. Like other relative measures, it changes when an arbitrary constant is added, so it should not be used indiscriminately across measurement scales.
A raw median absolute deviation, abbreviated MAD in this unit, is the median of the absolute deviations from the sample median. For 2, 4, 4, 7, and 13, the median is four and the absolute deviations are 2, 0, 0, 3, and 9. Their sorted middle value is two, so the raw MAD is two minutes. Some software multiplies this by a normal-reference factor close to 1.4826. We report the raw definition unless a scaling convention is named.
The normal-reference scaling makes the measure comparable to a normal distribution's standard deviation under that model; it does not make every dataset normal. A raw MAD and a scaled MAD are different reported quantities. Software packages can use different defaults, so a reproducible report should include the definition or scaling constant. This is another example in which matching a label is insufficient for matching a calculation.
Range, IQR, and MAD can all support investigation of an unusual observation, but none determines whether it is erroneous. A large range might reflect a genuinely heterogeneous population. A small IQR with several far-away values might reflect a mixture or a heavy tail. Inspect the data source, relevant groups, and a graph before deciding what happened. Robust spread describes a central pattern while preserving the responsibility to explain the rest of the distribution.
For a positive affine transformation y equal to a plus b times x, the range, IQR, and raw MAD multiply by b. More generally, absolute spread measures multiply by the absolute value of b, with quantile order reversed when b is negative. Adding a constant does not change their widths. Nonlinear transformations can change relative spacing and therefore change these measures in more complicated ways. A comparison should keep the measurement scale explicit.
§6.3 Mean absolute deviation and the role of a reference
An absolute deviation is the distance between an observation and a chosen reference value. The mean absolute deviation about a is the sum of these distances divided by n. Its units match the original variable. For the values 2, 4, 4, 7, and 13, deviations from the mean six have absolute values 4, 2, 2, 1, and 7. Their total is sixteen, giving mean absolute deviation about the mean of 3.2 minutes.
About the median four, the absolute deviations total fourteen, giving a mean absolute deviation of 2.8. The two results differ because the reference differs. The median minimizes total absolute distance, as established in unit five, so the average absolute deviation about a median is no greater than the average about any other reference for the same dataset. This is an optimization statement, not a declaration that the median-based spread always answers the preferred scientific question.
Mean absolute deviation and median absolute deviation are easy to confuse. The first averages distances; the second takes their median. Specify both the operation on distances and the reference location. In the example, mean absolute deviation about the median is 2.8, whereas the raw median absolute deviation is two. An abbreviation without a definition can conceal this difference, especially when different software uses MAD for different quantities.
Absolute deviations avoid the cancellation that makes the average signed deviation from the mean equal to zero. They do not exaggerate large distances as strongly as squares do, but they remain influenced by extreme values when the distances are averaged. Median absolute deviation provides stronger resistance to the magnitude of a small number of extremes. The distinction follows from where the averaging occurs and can be inspected by changing one value in a small dataset.
For any common reference a, the mean absolute deviation is no greater than the root mean squared deviation from a. This follows from the fact that the square of an average of nonnegative distances is no larger than the average of their squares. It gives a numerical check, but compare quantities with the same denominator and reference. The sample standard deviation uses denominator n minus one rather than n, so explain the convention before using inequalities to validate results.
Frequency tables of exact values support exact absolute-deviation calculations by weighting each distance by its frequency. If four occurs twice, its distance contributes twice. Grouped intervals do not generally preserve these distances. Replacing values by class midpoints produces an approximation, and calculating deviations from an approximated grouped mean adds another layer of approximation. State both limitations rather than treating midpoint calculations as recovered individual data.
Absolute deviation about a target can be useful even when that target is not the data's mean or median. A production line designed to make items of length fifty millimetres may be evaluated by average absolute distance from fifty. This measures departure from a specification, combining a shift in location with spread. An average absolute deviation about the sample's own centre instead removes some location discrepancy. Name the target and the purpose so that the distinction remains visible.
A report should not use a reference selected only to make dispersion look small. If the scientific aim is agreement with a fixed target, preserve that target. If the aim is variation around the observed centre, declare the centre and its convention. Descriptive flexibility is useful only when the resulting quantity remains connected to an explicit question. This reasoning will also help interpret mean squared error and residual error in later courses.
§6.4 Descriptive variance and standard deviation
Define the sum of squared deviations Sxx as the sum of (x minus x-bar) squared. The descriptive variance of the recorded collection is Sxx divided by n. We denote it by m2 when treating it as the second central moment. Its nonnegative square root is the descriptive standard deviation sN. Both are zero exactly when all eligible observations are equal. At least one observation is required for the descriptive variance, and a one-observation collection has descriptive variance zero.
For 2, 4, 4, 7, and 13, the mean is six and squared deviations are 16, 4, 4, 1, and 49. Their sum is seventy-four. The descriptive variance is 14.8 squared minutes, and the descriptive standard deviation is the square root of 14.8, approximately 3.8471 minutes. A report using denominator five should identify it as the recorded-collection or population-denominator convention, rather than silently implying an inferential estimate.
The standard deviation is not generally the mean distance from the mean. It is the square root of an average of squared distances. Squaring gives larger deviations disproportionately greater influence, and taking the square root returns the result to the original unit. This construction enables important identities but does not make it an ordinary average absolute error. The wording matters because readers often interpret standard deviation as if every observation were that exact distance from the centre.
The descriptive variance can be calculated as the mean of squared observations minus the square of the mean. This follows by expanding squared deviations and using the definition of x-bar. The identity is useful for algebra and checking a small exact calculation. On a computer, subtracting two very large nearly equal numbers can lose significant precision. Therefore a mathematically equivalent shortcut is not always a numerically safe implementation.
When values are near a large common baseline, centre them before accumulating squares or use a stable online algorithm. Measurements 1000000001, 1000000002, and 1000000003 have descriptive variance two thirds. Computing huge raw squares and subtracting the squared mean in limited precision can give an inaccurate result. Subtracting a reference near the data keeps intermediate magnitudes small without changing the variance. The textbook simulations use centred calculations.
Under y equal to a plus b times x, variance multiplies by b squared and standard deviation multiplies by the absolute value of b. Adding a constant changes location but not spread. Multiplying lengths by one hundred to convert metres to centimetres multiplies standard deviation by one hundred and variance by ten thousand. This difference is an excellent dimensional check. A negative scaling reverses order but leaves squared distances nonnegative.
For exact frequency values, Sxx is the sum of frequency times squared deviation from the frequency-weighted mean. The descriptive denominator is total frequency. Use the actual replicated count rather than the number of table rows. If weights represent something other than frequencies, the appropriate variance definition and inferential correction can differ. A general weighted variance should not borrow the n minus one correction merely because an unweighted sample formula uses it.
Variance is especially responsive to extreme magnitudes because their contributions are squared. This can be appropriate when large errors have disproportionate consequences, but it can also make a single data-entry error dominate the result. Validate records and display the distribution. Reporting a standard deviation without noting a strongly skewed or multimodal dataset can encourage readers to imagine a symmetric spread that the data do not possess.
All finite numerical datasets have calculable descriptive moments when values are finite and eligible counts are positive. This does not imply that every corresponding theoretical population has a finite variance. Heavy-tailed probability models can lack finite second moments. That topic belongs to the probability course, but the distinction is useful here: a finite recorded summary and a theoretical expectation are different objects, even when they share familiar terminology.
§6.5 Sample variance and the n minus one denominator
The conventional sample variance s squared is Sxx divided by n minus one and requires at least two eligible observations. Its square root is the sample standard deviation s. For the five waiting times in the preceding section, sample variance is seventy-four divided by four, or 18.5, and sample standard deviation is approximately 4.3012. The descriptive variance 14.8 and sample variance 18.5 come from the same deviations but different denominators. Neither is a rounding variant of the other.
The n minus one convention has an inferential justification under independent observations from a common distribution with finite variance. The sample mean is estimated from the same observations, which constrains the deviations to sum to zero. Under those assumptions, the expected Sxx equals n minus one times the population variance. Dividing by n minus one therefore gives an unbiased estimator of that variance. This is a statement about repeated samples under a model, not a statement that one sample result equals the unknown population variance.
An algebraic explanation begins with a fixed population mean mu. The sum of squared deviations from mu equals Sxx plus n times the square of x-bar minus mu. Under the stated independent common-distribution model, the expected first sum is n times the population variance, and the expected squared difference of the sample mean from mu is the population variance divided by n. Subtracting gives expected Sxx equal to n minus one times the variance. The probability course supplies the expectation rules behind this argument.
The phrase degrees of freedom describes the constraint on deviations. If n minus one deviations from the sample mean are known, the final deviation is forced by their zero sum. This intuition explains why an estimated centre consumes one degree of freedom, but it should not be used as an automatic rule for every data structure. Dependence, unequal weighting, fitted models, and complex sampling designs require their own variance formulas and justification.
Unbiased sample variance does not mean unbiased sample standard deviation. Taking a square root is nonlinear, and the expected square root of a nonnegative estimator generally differs from the square root of its expectation. The ordinary sample standard deviation is still widely useful, but the word unbiased should be attached to the variance claim under its assumptions, not carelessly transferred to every function of that variance.
A one-observation sample has no conventional sample variance because the denominator is zero. Its descriptive variance is zero, but that gives no evidence that the broader population has no variation. Software should report the sample variance as unavailable and explain the minimum count. For a constant sample of two or more observations, sample variance is zero; standardization by that zero standard deviation is then undefined.
When the recorded values constitute an entire finite population of interest, descriptive variance using n is a natural summary. When they are treated as an independent sample for estimating a broader variance, the n minus one convention has its familiar role. Real finite-population sampling can introduce other distinctions, including a finite-population variance defined with denominator N minus one. Name the convention rather than using population and sample as vague labels disconnected from an explicit target.
The correction factor between sample and descriptive variance is n divided by n minus one. It is large for tiny samples and approaches one as n increases. That numerical closeness at large n does not remove the need to state the definition. The intended use, data dependence, and sampling process remain important even when two formulas yield nearly equal displayed decimals.
§6.6 Combining groups and separating sources of spread
An overall variance cannot generally be obtained by averaging group variances alone. Differences between group means also contribute. For disjoint groups with counts nj, means x-bar-j, and within-group sums of squares Sxx-j, the total sum of squares around the overall mean equals the sum of within-group sums of squares plus the sum of nj times the squared difference between each group mean and the overall mean. The second term is called a between-group contribution.
Derive the identity by writing each observation's overall deviation as its within-group deviation plus its group's mean deviation. When squared and summed inside a group, the cross term vanishes because within-group deviations sum to zero. The remaining terms are the group's internal sum of squares and its count times the squared mean difference. Adding across groups gives an exact descriptive decomposition, requiring no normal-distribution assumption.
Consider groups 1, 3 and 7, 9. Each group has mean two or eight and descriptive variance one. The overall mean is five. Within-group sums of squares total four. The between-group contribution is two times nine plus two times nine, or thirty-six. Total Sxx is forty, giving descriptive variance ten across four observations. Averaging the two group variances would give one and miss most of the overall variation.
This decomposition explains why aggregation can hide heterogeneity. A graph of group means retains between-group locations but suppresses within-group spread. A single overall variance mixes both components. Neither summary is wrong if labelled properly, but they answer different questions. In a school comparison, the variation of student marks within schools and variation of school means across schools should not be treated as the same quantity.
To reconstruct an overall sample variance from group sample variances, first convert each group variance back to its within-group sum of squares by multiplying by nj minus one. Add the between-group contribution and divide the final total by total count minus one. The combined mean must first be computed using counts. Group definitions and measurement scales must be compatible, and groups must be disjoint so that observations are not duplicated.
Pooled within-group variance is a different quantity from overall variance. When groups share a common variance under a specified model, an estimator often divides the sum of within-group sums of squares by the sum of group degrees of freedom. For two groups this denominator is n1 plus n2 minus two. It deliberately omits between-group mean differences. The procedure belongs to later inference, but the descriptive distinction is essential here: pooled within-group spread is not total spread.
Weighted standardization of group composition can also alter the reported overall variance. A hypothetical common composition should include both the within-group contributions and the distances of group means from its weighted overall mean. Merely reweighting group standard deviations is not sufficient. Because standard deviation is a square root, an arithmetic average of standard deviations generally does not equal the square root of an appropriately combined variance.
A subgroup decomposition is not automatically a causal explanation. If travel mode groups have different waiting-time means, the between-group sum of squares describes that partition of the recorded variation. It does not prove that changing travel mode would cause the mean difference. Confounding, selection, and measurement differences remain possible. Statistical algebra can identify contributions under a grouping without establishing why the groups differ.
§6.7 Standard scores and reference distributions
A standard score expresses an observation's distance from a chosen centre in units of a chosen standard deviation. The common sample form is z equal to x minus x-bar divided by s, with s greater than zero. It is dimensionless because numerator and denominator use the same original unit. A z score of two means two chosen standard deviations above the chosen mean. It does not by itself specify a percentile or a tail probability.
For a dataset standardized using its own sample mean and sample standard deviation, the mean of the z scores is zero and their sample variance is one. Their descriptive variance using denominator n is instead n minus one divided by n. If standardized using the descriptive standard deviation sN, the descriptive variance is one and the sample variance is n divided by n minus one. These distinctions follow directly from the denominator conventions, and they should be reflected in validation checks.
Using a fixed external reference mean and standard deviation gives a different interpretation. A student's score can be described relative to a previously defined reference cohort rather than the current classroom. The resulting classroom z scores need not have mean zero or variance one. State the source and relevance of the reference. A standard score is only as meaningful as the centre, scale, and population chosen for the comparison.
Standardizing does not make a distribution normal. It shifts and rescales the original shape without changing its essential pattern under a positive scale. A strongly skewed set of waiting times remains skewed after z transformation. Statements such as approximately ninety-five percent within two standard deviations require a suitable distributional assumption; they cannot be inferred merely because the values have been standardized.
A standardized value is undefined when the chosen standard deviation is zero. In that case all values in the reference collection may coincide, so there is no nonzero spread unit. Do not silently divide by a tiny replacement constant or report every score as zero without naming a convention. The appropriate descriptive report states the constant value and explains why standardized distances are unavailable.
Standard scores enable comparison of relative positions across variables, but those comparisons do not necessarily establish equivalence of achievement or risk. Being two standard deviations above the mean in two examinations refers to two different score distributions and possibly different populations, content, and reliability. A common numerical z does not make the underlying constructs identical. Measurement validity remains important after numerical rescaling.
Extreme observations influence both the mean and standard deviation used in their own standardized scores. A very large observation can inflate s, so a self-standardized z is not independent evidence that the observation is harmless. Robust alternatives can use a median and a declared robust spread measure, but their interpretation and scaling differ. A threshold should be chosen for a justified purpose, not because it makes the plotted observations look tidy.
When calculating standardized differences from rounded summaries, allow for rounding error. A published mean and standard deviation rounded to one decimal place may not reproduce a software z to many decimal places. Retain unrounded intermediate values where possible and report sensible precision. Verification should compare calculations within a tolerance appropriate to floating-point arithmetic rather than demand impossible exact decimal equality.
§6.8 Coefficients of variation and relative comparisons
The conventional coefficient of variation, CV, is standard deviation divided by a positive mean, often multiplied by one hundred and expressed as a percentage. It describes spread relative to the mean and is most interpretable for positive ratio-scale quantities with a meaningful zero. State whether the standard deviation uses denominator n or n minus one. For the waiting-time sample with mean six and sample standard deviation about 4.3012, sample CV is approximately 71.686 percent.
Multiplying all values by a positive constant leaves CV unchanged because both mean and standard deviation multiply by that constant. Converting a positive length dataset from metres to centimetres therefore preserves the coefficient. Adding a constant changes the mean but not the standard deviation, so CV changes. This is why an arbitrary zero point can make the comparison misleading. Celsius and Fahrenheit temperature CVs can disagree for the same physical measurements.
A mean near zero can make CV extremely large and unstable. A mean equal to zero makes the conventional ratio undefined. A negative mean does not have the ordinary positive relative-dispersion interpretation. Some formulas use the absolute mean, but that is a different convention and does not repair an inappropriate measurement scale. The simulation here restricts the displayed conventional CV to a positive mean and otherwise explains why it is unavailable.
Relative spread can reverse a comparison based on absolute spread. A process with mean one hundred and standard deviation ten has CV ten percent. Another with mean ten and standard deviation two has CV twenty percent. The first process is more variable in original units, while the second is more variable relative to its mean. Both statements can be true. Choose the comparison that corresponds to the decision rather than treating one as a universal ranking of consistency.
For counts or other quantities with a structural mean-variance relationship, CV can change partly because the mean changes. Later probability models make these relationships explicit. A lower CV is not automatically evidence of a better process or more accurate instrument. Interpret the scale, practical tolerances, and data-generating context. Relative spread is one descriptive lens, and fixed absolute specifications may remain more relevant to quality control.
A CV computed from a mean of transformed values usually has a different meaning from one computed on the original scale. For positive right-skewed data, a geometric coefficient of variation may be defined through a model on the logarithmic scale, but it should not be confused with the ordinary ratio s divided by x-bar. This book keeps the introductory definition explicit and reserves specialized transformed-scale measures for later courses.
Comparisons require compatible eligible populations. A CV of customer purchase amounts among completed transactions cannot be directly interpreted as relative dispersion of all potential customers' spending if nonbuyers were excluded. The denominator of the mean and the set used for standard deviation should match. Otherwise the ratio combines summaries of different data subsets and loses a coherent ordinary interpretation.
Report CV alongside the mean, standard deviation, units, and count. A percentage by itself hides whether the original measurements are tiny, large, concentrated, or affected by a near-zero centre. Relative summaries complement absolute summaries rather than replace them. This reporting discipline makes it possible for a reader to judge whether a proportional comparison is appropriate.
§6.9 Distribution-free bounds and the limits of rules of thumb
The arithmetic of a finite dataset gives a useful bound on how many observations can lie far from its mean. If the descriptive variance m2 is positive, the proportion of observations whose absolute deviation is at least k times the descriptive standard deviation cannot exceed one divided by k squared, for k greater than zero. Each such observation contributes at least k squared times m2 to the squared-deviation total, which is n times m2. Counting those contributions yields the bound.
For k equal to two, at most one quarter can lie at or beyond two descriptive standard deviations. Equivalently at least three quarters lie strictly within that distance under this particular boundary wording. For k equal to three, at most one ninth lie at or beyond three descriptive standard deviations. These are conservative arithmetic bounds; an actual dataset may have far fewer observations in those tails. Carefully distinguish at least, greater than, and strict versus inclusive thresholds.
This finite-data result parallels Chebyshev's inequality for a probability distribution with finite variance. The population form bounds probability rather than an observed empirical proportion. The introductory proof here concerns the recorded values and their descriptive variance. Using a sample standard deviation changes the scaling and can produce a corresponding factor involving n minus one over n. Do not mix the two denominators while quoting a bound.
The familiar approximate 68, 95, and 99.7 percent rule is associated with a normal distribution at one, two, and three standard deviations from its mean. It is not a universal law of numerical data. Symmetry alone is insufficient to establish those percentages, and standardized values need not be normal. A histogram, substantive model, and later probability methods help judge when the normal approximation is plausible.
Even a dataset that looks roughly bell-shaped can show deviations in its tails, and a small sample offers limited evidence about rare outcomes. A graphical resemblance does not justify treating a tail probability as precisely known. Use descriptive counts when describing observed data and name any probability model used for broader statements. The distinction between an observed proportion and a model probability is central to the next book.
The range supplies another descriptive bound. For a finite collection with minimum a and maximum b, its descriptive variance is at most the square of b minus a divided by four. Intuitively the largest spread within a fixed interval is achieved by placing observations at the extremes with balanced weights. A derivation uses the nonnegative product (x minus a) times (b minus x) and the resulting bound on the second moment. Equality requires an appropriate balanced endpoint configuration.
Such bounds are useful for checking a result. A descriptive variance larger than one quarter of the squared observed range indicates a computational or definition error. A result below the bound can still be wrong. Validation combines several identities, dimensional checks, count checks, and independent examples. No single inequality serves as a full correctness certificate.
Do not use broad spread bounds as promises about future cases. A finite recorded range does not establish a maximum possible future wait, and a descriptive Chebyshev count statement does not guarantee the same proportion in a changing service. Generalization requires evidence about selection, stability, and a target population. The bounds clarify what the observed arithmetic implies while preserving the limits of that implication.
§6.10 Variance exploration, verification and reporting
The variation explorer uses a five-value dataset with one adjustable high value. It displays the mean, sum of squared deviations, descriptive variance, sample variance, and both standard deviations. At the initial setting 2, 4, 4, 7, and 13, verify the sum of squares seventy-four and the two denominators five and four. Predict how a large increase in the final value will affect the squared deviations before moving the control.
Every update should recompute the mean first, because changing an observation changes the reference for all deviations. Reusing the previous mean would produce the wrong sum of squares. The simulation also displays the individual centred deviations so that the learner can inspect the calculation rather than trust a decorative chart. Their signed sum should be approximately zero, and their squared sum should match the displayed Sxx.
The ratio of sample variance to descriptive variance should equal n divided by n minus one whenever descriptive variance is positive and n is at least two. For constant data, both variances are zero and the ratio is undefined even though the multiplicative identity still holds. A test should check the identity without dividing by zero. Standardized values and CV also require their own valid denominators and clearly explained unavailable states.
Three worked questions accompany this unit. The first calculates both variance conventions and standard deviations from a small dataset. The second reconstructs total spread from two groups using within- and between-group contributions. The third examines a unit conversion and a coefficient-of-variation comparison, including why adding an arbitrary constant can destroy the intended relative interpretation. Each solution preserves units and states the denominator.
A useful report combines centre and spread with context: among eligible recorded visits, the mean wait was six minutes and the sample standard deviation was approximately 4.30 minutes, based on five observations; the recorded values ranged from two to thirteen. This particular dataset is deliberately small for teaching. A real operational conclusion would require more evidence about representativeness and stability, but the descriptive statement is reproducible as written.
For a strongly skewed distribution, add median and IQR or a clear plot rather than assuming mean plus or minus standard deviation captures the whole shape. For a categorical variable, use proportions and counts rather than applying a numerical spread formula to arbitrary codes. For a ratio-scale relative comparison, include CV only when its denominator and zero point are meaningful. The choice of spread measure remains connected to the measurement and question.
Before proceeding, distinguish mean absolute deviation from median absolute deviation, explain why sample variance uses n minus one under its stated model, and derive the variance transformation rule for a change of unit. Then reconstruct an overall sum of squares from two groups. These tasks prepare for central moments and distribution shape, where the same centred-deviation machinery is extended to third and fourth powers.
Step-by-Step Statistics Solutions
Three original questions connect calculations, definitions and interpretation. Open each solution to follow the reasoning.
Calculate the descriptive variance, sample variance, and their standard deviations for 2, 4, 4, 7, and 13. State units and explain why the sample-variance correction does not imply that one realised estimate equals a population variance.
The mean is six. Squared deviations are sixteen, four, four, one, and forty-nine, totaling seventy-four. All use the same eligible five observations.
Dividing by five gives descriptive variance 14.8; dividing by four gives sample variance 18.5. Their square roots are approximately 3.847077 and 4.301163 minutes. Variances are in squared minutes.
The n-minus-one estimator is unbiased for population variance under independent common-distribution sampling with finite variance. That is a repeated-sampling property. It does not guarantee that this particular estimate equals the unknown variance or that its square root is also unbiased.
Two groups contain values (1,3) and (7,9). Calculate the overall mean, within-group and between-group sums of squares, and both overall variance conventions. Explain why averaging the group descriptive variances is insufficient.
Group means are two and eight; the overall mean is five. Each group has two values and internal squared-deviation sum two. The combined within sum is four.
Between sum of squares is two times three squared plus two times three squared, or thirty-six. Total sum of squares is forty, giving descriptive variance ten and sample variance forty thirds.
Each group descriptive variance is one. Their average misses the thirty-six units of squared variation associated with separated group means. Total spread combines within-group variation and group-location differences, without itself providing a causal explanation.
Use the five waiting times with mean six and sample variance 18.5. Calculate the sample coefficient of variation. Convert minutes to seconds and show it stays the same. Then add four to every original value and explain why CV changes although spread does not.
Sample standard deviation is the square root of 18.5. Dividing by the positive mean six gives approximately 0.716860, or 71.6860 percent. The measurement has a meaningful zero for this relative comparison.
Seconds multiply mean and standard deviation by sixty, so their ratio is unchanged. Variance instead multiplies by thirty-six hundred. Keep these different transformation rules distinct.
Adding four gives mean ten but leaves standard deviation unchanged, producing CV approximately 0.430116. This sensitivity to addition explains why a conventional ratio-based CV is unsuitable for scales with an arbitrary zero point, such as Celsius temperatures.