Statistics & Probability Introductory major course 100% Free Open Access
Chapter 5 • Theory & Derivations

Location, Quantiles and Weighted Summaries

Arithmetic mean, median, quantiles, mode, weighted means, geometric growth factors, harmonic rate averages and robust location.

§5.1 What a typical value is supposed to describe

A measure of location compresses a collection of observations into a number intended to describe its position. The compression always loses information. Two classes can have the same average examination mark while one contains students clustered near that average and the other contains students spread across almost the entire marking scale. Two villages can have equal median household income but quite different proportions of households below a particular affordability threshold. A useful location measure answers a stated question about a stated variable; it does not replace the distribution.

Begin with the observational unit. An average household expenditure describes households only if each household receives the intended weight. An average expenditure per person asks a different question and usually requires household size. If one household of two people spends 600 currency units and another household of six people spends 900, the mean household expenditure is 750. The combined expenditure per person is 1500 divided by eight, or 187.5. Averaging the two household per-person figures, 300 and 150, instead gives 225. Both calculations are arithmetically legitimate, but they describe different weighting schemes and different target units.

The variable's measurement scale also restricts the calculation. Numerical category codes assigned to blood groups do not acquire an interpretable arithmetic average merely because software accepts them. An ordered satisfaction response supports statements about order, but an arithmetic mean additionally treats the chosen coding gaps as meaningful. Heights measured in centimetres support ordinary averaging because equal numerical differences represent equal length differences. Ratios of Celsius temperatures require special caution because the zero is conventional rather than an absence of thermal energy.

There are several reasonable meanings of typical. A retailer planning total stock may need an arithmetic mean because totals are central to the decision. A tenant asking what an ordinary renter pays may prefer the median when a few luxury properties pull the mean upward. A manufacturer may need the most common nominal size, represented by a mode, because sizes form separate categories. An investor summarizing successive proportional changes may need a geometric mean, while an equal-distance speed calculation uses a harmonic mean. These measures are connected by mathematics, but their practical roles depend on how the observations combine.

A location summary should retain its unit. A mean waiting time of six minutes is not interchangeable with a mean of six recorded clock ticks unless one tick equals one minute. Percentage summaries require similar care: an average of percentage values is generally expressed in percentage points, whereas a proportional comparison between two percentages is a relative change. A report that says the average increased by five percent must distinguish an increase of five percentage points from a five-percent increase relative to the starting average.

The mean, median, and mode need not agree. Their disagreement can reveal asymmetry, category structure, measurement rounding, or small-sample irregularity. It does not by itself prove a particular distributional shape. A small irregular sample can have a mean above its median without resembling a smooth right-skewed theoretical model. Inspect the actual values and a suitable graph before attaching a shape label. Later units develop more explicit measures of shape and explain why those also require convention and context.

Location is only one part of a responsible summary. State the number of eligible observations, the treatment of missing values, and at least one measure or display of variation. The phrase average score is incomplete when the reader cannot tell whether it refers to a mean or median, whether absentees were excluded, or whether the observations came from the entire class. These details are not decorative metadata. They determine what the reported number means and whether another person can reproduce it.

This unit develops location measures from their definitions and the operations they preserve. We will use individual values when available, frequency representations when appropriate, and explicitly approximate procedures when only grouped information remains. The resulting calculations should be explainable in ordinary language before they are implemented in software. Being able to defend the denominator and the weights is more important than memorizing the name of an average.

§5.2 The arithmetic mean and its algebra

For numerical observations x1 through xn, with n greater than zero, the arithmetic mean is their sum divided by n. We write it as x-bar. The formula gives every recorded observation equal weight. Replacing every observation by x-bar preserves the total: n times x-bar equals the sum of the original values. This identity explains why the arithmetic mean is useful for questions about combined expenditure, accumulated time, or aggregate output. It also explains why the observational unit and eligibility rules matter so much.

For the values 2, 4, 4, 7, and 13, the total is 30 and the mean is six. Saying that the mean is six does not imply that anyone actually waited six minutes. The mean is a balancing location. The deviations from it are minus four, minus two, minus two, one, and seven; their sum is zero. The positive and negative deviations cancel exactly, apart from rounding introduced by a computer or by a displayed decimal.

For any proposed location a, the sum of x minus a equals n times the difference between x-bar and a. Therefore only the arithmetic mean makes the signed deviations sum to zero. This property is useful for checking calculations, but it also warns against using the average signed deviation as a measure of spread: when deviations are measured from the mean, that average is always zero, even for a highly variable dataset. Spread requires absolute values, squares, or another method that avoids cancellation.

The mean transforms simply under a change of origin or scale. If every new value y equals a plus b times x, then the new mean is a plus b times the original mean. Adding three minutes to every waiting time adds three minutes to the mean. Converting metres to centimetres multiplies the mean by one hundred. This follows by distributing the sum over the addition and multiplication, so it holds for negative b as well as positive b. The interpretation of ordering changes when b is negative, but the algebra remains valid.

It is often convenient to calculate around a reference value c. The mean equals c plus the average of the deviations x minus c. If recorded measurements are 1002, 1004, 1004, 1007, and 1013, use c equal to 1000. Their deviations are the familiar 2, 4, 4, 7, and 13, giving a mean of 1006. This coding method reduces repetitive arithmetic. It does not change the data or grant permission to round away meaningful differences. A computational reference is separate from a scientific zero point.

A frequency table can preserve the exact mean if each row contains an exact observed value xj and its frequency fj. The mean is the sum of fj times xj divided by the sum of fj. For values two, four, seven, and thirteen with frequencies one, two, one, and one, this gives 30 divided by five. Frequencies represent replicated observations, so the denominator is the total count, not the number of distinct values. Dividing by four would mistakenly give equal weight to categories rather than equal weight to observations.

The mean of a union of disjoint groups is a weighted mean of the group means, with group sizes as weights. A group of ten students with mean 70 and a group of thirty with mean 80 have a combined mean of 77.5, not 75. Their implied totals are 700 and 2400, and the combined total 3100 is divided by forty. This reasoning requires groups that do not overlap and means measured on the same variable under compatible definitions. An average of averages without denominators cannot generally recover the overall mean.

Suppose one observation changes from x to x plus d while the sample size remains fixed. The mean changes by d divided by n. This makes sensitivity transparent. In a sample of five, changing thirteen to sixty-three raises the mean by ten. In a sample of five hundred, the same change raises it by one tenth. The effect depends on the magnitude of the change and the size of the dataset. A very large unusual value can still have a substantial effect even in a large sample, so sample size alone does not make the mean robust.

The arithmetic mean is undefined for an empty collection. It should not be set to zero as a convenience. Zero means that an actual calculated quantity equals zero; absence of eligible observations means that there is no average to calculate. Missing-value handling should therefore report an eligible count. If all observations are missing after applying the stated eligibility rule, the correct output is a clear missing or unavailable result with an explanation.

§5.3 Medians and quantiles with a declared convention

The median is a central location defined through order. Sort the n observations from smallest to largest, retaining repeated values. When n is odd, the ordinary sample median is the middle observation. When n is even, it is the arithmetic mean of the two middle observations. For 2, 4, 4, 7, and 13, the median is four. For 2, 4, 7, and 13, it is 5.5. The even-sample median may therefore be a value that was not observed, just as the arithmetic mean may be unobserved.

At least half the observations lie at or below a median and at least half lie at or above it. Ties make these proportions potentially larger than one half. This statement is more accurate than saying exactly half the observations lie below and exactly half above. In the sample 1, 1, 1, 8, 9, the median is one, but no observation lies strictly below it. The definition uses weak inequalities and counts, so it remains valid in the presence of ties.

For an ordinal variable, identify a median category through cumulative counts and order. If the two middle observations belong to different categories, the arithmetic average of their numerical codes need not denote a meaningful intermediate category. A response scale labelled poor, fair, good, and excellent does not automatically contain a category halfway between fair and good. A report can state the central categories or use a declared category rule. The numerical median formula is appropriate when the variable supports the interpolation being performed.

More generally, a quantile marks a location associated with a chosen fraction p between zero and one. The first quartile corresponds to p equal to one quarter, the median to one half, and the third quartile to three quarters. Deciles use tenths and percentiles use hundredths. Sample quantiles have several accepted definitions because finite data do not contain an observed value at every desired fraction. Different software can return different quartiles without either implementation making an arithmetic mistake.

This textbook uses the type-seven interpolated quantile for numerical examples unless explicitly stated otherwise. For sorted observations, calculate the one-based position h equal to one plus n minus one times p. Let j be the integer part of h and g its fractional part. Interpolate between the observations at positions j and j plus one with weights one minus g and g. At p equal to zero use the minimum, and at p equal to one use the maximum. If n is one, every quantile equals the single observation.

For the sorted values 1 through 8, the first-quartile position is 2.75, giving Q1 equal to 2.75. The median position is 4.5, giving 4.5, and the third-quartile position is 6.25, giving 6.25. The median-of-halves convention would instead give Q1 equal to 2.5 and Q3 equal to 6.5. Always name the convention when exact agreement across tools matters. The unit-four box-plot simulation and the later numerical explorer use type seven so that their results match the displayed calculations.

An empirical cumulative distribution function answers a related but different question: what proportion of the observed values is at or below a threshold? Its inverse can define a stepwise quantile that selects an observed value. Type-seven interpolation need not produce that same value and need not make the empirical cumulative proportion exactly equal to p. The distinction is especially visible in small samples. We use interpolation as a declared numerical convention, not as a claim that an unobserved interpolated value appeared in the dataset.

Quantiles summarize several aspects of location. A median describes the centre by rank; a high percentile describes a relatively high location relevant to service delays or exposure thresholds. Reporting only a mean waiting time may hide a long upper tail that matters to customers. Reporting a ninety-fifth percentile asks how high a threshold covers most recorded waits under the chosen finite-sample definition. It does not promise that ninety-five percent of all future waits will be below that threshold without further evidence and modelling.

Positive affine transformations preserve the ordering and transform type-seven quantiles by the same affine rule as the values. Thus converting centimetres to metres converts the median and quartiles in the expected way. A decreasing transformation reverses order, and a nonlinear increasing transformation generally does not commute with linear interpolation. For example, the median of two values after squaring them need not equal the square of their averaged median. Distinguish properties of rank order from properties of a particular interpolation rule.

Quantiles require a clear eligible dataset and consistent measurement. They are less sensitive than the mean to the magnitude of a single extreme value, but they can change when enough observations move across a rank boundary. They can also be distorted by nonresponse, selection, or systematic measurement error. Resistance to numerical extremes is a narrow mathematical property; it is not protection against every source of statistical bias.

§5.4 Modes, multimodality and categorical summaries

A mode is a value or category with the largest frequency. It applies naturally to nominal categories because identifying the most frequent label requires equality comparisons rather than arithmetic distances. If transport responses are bus, bus, walk, cycle, bus, walk, the modal category is bus. It is a category description, not a numeric average. A frequency table should accompany the mode when the difference between the leading and next category matters to the reader.

There may be more than one mode. For values 1, 1, 2, 2, and 7, both one and two have the largest frequency. A rule that silently reports only the smallest tied mode loses information and should be identified as a software convention. A dataset in which every observed value occurs once has no uniquely most frequent observed value. Calling every value a mode is mathematically possible under a literal maximum-frequency definition, but it is usually more informative to report that the data have no unique sample mode.

For finely measured numerical data, exact repetition depends strongly on rounding. Heights recorded to the nearest centimetre may have a clear most frequent recorded number, while the same heights recorded with many decimal places may all be distinct. This does not mean the underlying distribution abruptly lost its high-density region. It means the exact-value frequency mode and a smooth distributional peak are different objects. A histogram can suggest modal regions, but the apparent peak depends on bin width and bin origin.

When data are grouped into unequal-width intervals, the interval with the largest count need not be the interval with the highest density. An interval of width twenty containing twelve observations has count twelve but count density 0.6 per unit. An interval of width ten containing eight observations has count density 0.8 per unit. The second interval is taller in a correctly constructed density histogram even though it contains fewer observations. A modal-class claim should state whether it refers to total count or estimated local density.

Some introductory treatments estimate a grouped mode using a formula based on the frequencies of a modal class and its two neighbours. That procedure assumes a particular local shape and class structure; it does not reconstruct the exact original mode. This book emphasizes the information actually retained. If only interval counts are available, report a modal interval or an explicitly model-based interpolation. Do not give an interpolated number a false precision that the grouped data cannot support.

Modes can help diagnose mixtures. A frequency pattern with separate concentrations may correspond to two production settings, two demographic subgroups, or two measurement conventions. It can also arise from small-sample noise or arbitrary binning. The presence of two peaks should invite investigation into data provenance and grouping rather than immediate identification of two biological or social populations. Graphical evidence and substantive explanation should be considered together.

For a binary variable coded zero and one, the mode identifies the more frequent category, while the mean gives the proportion of ones. The median depends on the central observations and can be zero, one, or one half for an evenly split even-sized numerical sample. A median of one half is not a third observed binary category. This example shows why choosing a summary according to a variable's meaning is more reliable than treating all location measures as interchangeable options in a menu.

When reporting a categorical mode, retain the category label and enough frequency information to judge how dominant it is. A modal response supported by thirty-one percent of respondents tells a different story from one supported by ninety-one percent, even if both are labelled good. A tie should be visible. Missing responses should be counted separately or handled under a declared rule, rather than being accidentally treated as the leading substantive category because a missing code occurs frequently.

The mean-median-mode relation sometimes taught for moderately skewed smooth distributions is an empirical approximation, not a universal identity. Arbitrary finite datasets do not obey a fixed equation connecting the three. It should never be used to replace a calculable median or to manufacture an exact mode from a mean. In this course, calculate the summaries directly from the available information and describe any model-based approximation as such.

§5.5 Weighted means and defensible denominators

A weighted arithmetic mean assigns nonnegative weights wi to the observations xi and divides the sum of wi times xi by the sum of wi. The denominator must be positive. Multiplying every weight by the same positive constant leaves the result unchanged. Equal weights recover the ordinary mean. Frequency weights treat a row as representing several observations; other weights can represent population composition, exposure, precision, or a policy choice. The same algebra does not make these meanings interchangeable.

Suppose an assessment contains coursework worth thirty percent and an examination worth seventy percent. Scores of 80 and 60 give a weighted score of 66. The weights encode the assessment policy, not replicated students. A simple average of 70 would apply a different policy. State the weights as part of the definition of the score so that a reader can reproduce and evaluate the result. The phrase weighted average is insufficient without identifying what the weights mean.

For combining group means, sample-size weights preserve the aggregate total. Group means of 70 and 80 with group sizes ten and thirty give 77.5. Equal weighting gives 75 and describes the average of the two group means. This may be appropriate if the groups themselves are the target units, such as two schools receiving equal institutional weight. It is inappropriate if the target is the average score of all forty students. A denominator is a statement about whose contribution counts.

For combining proportions, use the denominators of the proportions when the goal is an overall proportion across disjoint eligible units. If three of ten observations and sixty of one hundred observations have an attribute, the combined proportion is sixty-three divided by one hundred and ten, approximately 0.5727. Averaging thirty percent and sixty percent gives forty-five percent, which instead gives the two groups equal weight. Rounded group percentages may not preserve enough information to recover the exact combined count, so retain counts where possible.

Survey weights often represent how many population units a sampled unit stands for under a design or calibration procedure. A weighted mean can improve representation when the weights are defensible, but weights do not automatically correct nonresponse, measurement error, or unmeasured selection. This introductory unit explains the descriptive calculation. Later sampling courses establish design-based estimators, uncertainty calculations, and the consequences of unequal inclusion probabilities. A weighted descriptive result should not be presented as proof that all survey biases disappeared.

Zero weights exclude contributions from the weighted total. Negative weights arise in some specialised estimation procedures, but they do not have the ordinary interpretation of frequencies or population representation. They can produce a weighted result outside the range of the observations. Our basic location explorer permits nonnegative weights and rejects a zero total weight, because its purpose is the ordinary descriptive weighted mean. Specialized signed estimators require their own definitions and justification.

A weighted mean with nonnegative weights lies between the smallest and largest observations that receive positive weight. This follows because the normalized weights sum to one and form a convex combination. If a reported ordinary weighted mean lies outside that range, inspect the weights, units, and denominator. This is a useful validation rule but not a complete proof of correctness: many incorrect calculations still fall within the observed range.

Weighted totals and weighted means should not be confused. If each sampled purchase represents ten purchases in a target population, the weighted total describes represented expenditure, while the weighted mean divides by represented purchase count. The units of the total and mean differ. Similarly, averaging rates requires deciding whether exposures, persons, institutions, or periods deserve equal weight. The choice should follow the target question and the way the underlying numerator accumulates.

Weighting can also change apparent association. Standardizing subgroup proportions with common weights removes differences in subgroup composition from a descriptive comparison, as in the Simpson example in unit three. It does not remove all causal confounding. The standardized quantities refer to a hypothetical common composition and should be labelled accordingly. A good weighted summary explains both the arithmetic and the population or policy that the weights describe.

Check weighted calculations by temporarily expanding small frequency-weighted examples into repeated values. If a value four has weight two, a frequency interpretation reproduces two fours. Compare the expanded ordinary mean with the weighted formula. This check is not appropriate for every type of weight, but it is an excellent way to detect a denominator error in a frequency table. The complete record should include the sum of weights and the number of physical rows, because those quantities may differ substantially.

§5.6 Geometric means and proportional change

For positive observations, the geometric mean is the nth root of their product. It can also be calculated by taking the arithmetic mean of their logarithms and then exponentiating. The logarithmic form is usually more numerically stable because a product of many large or small values can overflow or underflow a computer's representable range. The observations must be positive for this real logarithmic definition. A software routine should reject zero or negative inputs instead of silently applying an unrelated convention.

The geometric mean is appropriate when quantities combine multiplicatively. Suppose a quantity grows by twenty percent in one year and falls by twenty percent in the next. The growth factors are 1.2 and 0.8, whose product is 0.96. The final amount is four percent below the starting amount. The arithmetic mean of the two percentage changes is zero, but it does not preserve the two-year change. The constant annual factor that preserves the product is the square root of 0.96, approximately 0.979796.

Subtracting one from that geometric factor gives a constant annual percentage change of approximately minus 2.0204 percent. Over two years, multiplying by this same factor twice recovers the observed factor 0.96. This does not predict future investment performance. It is a retrospective equivalence between a varying sequence of factors and a constant factor producing the same endpoint. The same reasoning applies to population growth, repeated efficiency changes, or any process whose successive proportional changes multiply.

Do not take a geometric mean directly of signed percentage changes and call it a growth rate. Convert each change r into a factor one plus r, with r expressed as a decimal. A ten-percent increase becomes 1.10, and a ten-percent decrease becomes 0.90. Factors must be positive for the usual logarithmic geometric mean. A change of minus one hundred percent creates a zero endpoint, while a more negative percentage may be incompatible with the underlying nonnegative quantity. These cases require explicit substantive treatment.

Geometric means also summarize positive measurements whose differences are naturally described by ratios. Suppose concentrations are 1, 10, and 100 in a common unit. Their geometric mean is ten, whereas their arithmetic mean is thirty-seven. The geometric mean is central on a logarithmic scale: the log deviations around log ten balance. The arithmetic mean is central on the original additive scale and preserves the total concentration values. Neither is automatically correct for every decision about these measurements.

Under multiplication of every observation by a positive constant c, the geometric mean multiplies by c. Under addition of a constant, there is generally no comparable simple rule. Therefore a geometric mean should not be used on a measurement scale whose zero can be shifted arbitrarily when a ratio interpretation is intended. Ratios of measured lengths or positive monetary amounts can have a meaningful zero under the chosen definition; Celsius temperature ratios usually do not. Measurement reasoning precedes formula selection.

For nonnegative weights that sum to a positive amount, a weighted geometric mean is the exponential of the weighted arithmetic mean of log values. Frequency weights can represent repeated positive observations. Time-duration weights can describe a constant equivalent growth rate across intervals of unequal duration when the factors and rates have been defined consistently. One cannot simply weight raw annual percentage changes by an arbitrary number of months and expect every compounding problem to be solved; clarify which factors correspond to which intervals.

The arithmetic mean of positive numbers is at least their geometric mean, with equality when all positively weighted observations are equal. This inequality is a mathematical comparison, not a ranking of statistical usefulness. A geometric average will often be numerically smaller because it balances multiplicative rather than additive differences. If a report changes from arithmetic to geometric averaging, the resulting lower number should not be described as evidence that the underlying measurements improved.

Logarithms depend on their base, but a correctly back-transformed geometric mean does not. Averaging natural logarithms and applying the exponential yields the same positive result as averaging base-ten logarithms and taking ten to that power. Use one consistent base through the calculation. Units should be retained in the reported back-transformed result; a mean of log values itself is on a transformed scale and should not be labelled as if it were a concentration in the original unit.

Missing and censored positive measurements require care. A value below a detection limit is not necessarily zero and cannot automatically be replaced by an arbitrarily small positive number just to enable logarithms. Such replacement can greatly affect the geometric mean. Report the detection limit and use an appropriate method in a later specialized course. Here, the geometric mean is calculated only when the eligible positive observations are genuinely available under a declared rule.

§5.7 Harmonic means, rates and exposure

For positive observations, the harmonic mean is n divided by the sum of their reciprocals. Equivalently, take the arithmetic mean of the reciprocals and then invert it. It is useful in some rate problems because a fixed amount of work or distance takes time inversely proportional to the rate. The appropriate average follows from adding the underlying times and amounts, not from a general rule that every set of rates must use a harmonic mean.

Consider travelling sixty kilometres at thirty kilometres per hour and then sixty kilometres at sixty kilometres per hour. The first leg takes two hours and the second one hour. The overall speed is total distance one hundred and twenty kilometres divided by total time three hours, giving forty kilometres per hour. This equals the harmonic mean of thirty and sixty because the distances are equal. The arithmetic mean of forty-five does not preserve total travel time in this situation.

Now consider travelling for one hour at thirty kilometres per hour and one hour at sixty kilometres per hour. The distances are thirty and sixty kilometres, and total distance ninety divided by total time two gives forty-five kilometres per hour. Equal time exposures lead to the arithmetic mean of the speeds. The numerical rates are the same in both examples; the weighting and accumulation mechanism differ. This comparison is the safest way to remember which average a speed question requires.

For unequal distances dj at speeds vj, total time is the sum of dj divided by vj. Overall speed is the sum of distances divided by that total time. This is a distance-weighted harmonic mean of the speeds. For unequal time intervals tj, total distance is the sum of tj times vj, and the overall speed is the time-weighted arithmetic mean. The two formulas agree when applied to the same journeys with consistent distance and time records.

Other rate problems should be reduced to their underlying numerator and denominator. If a machine processes a fixed number of items at different positive rates, time per item is the reciprocal rate and equal item counts lead to a harmonic mean. If each machine runs for an equal duration, combining outputs leads to an arithmetic mean of processing rates. A rate such as cases per person-year is combined by adding cases and adding person-years, which produces an exposure-weighted arithmetic average of the individual rates.

The harmonic mean requires positive values in the ordinary rate interpretation. A zero speed over a specified positive distance means the journey cannot be completed in finite time under that speed, so inserting zero into a reciprocal formula is invalid. Negative speeds may encode direction, but signed velocity requires a different physical analysis of displacement and time. A general positive-rate harmonic mean should reject these inputs and explain the restriction.

For positive observations, the harmonic mean is at most the geometric mean, which is at most the arithmetic mean. Equal values make all three equal. A small positive observation has a large reciprocal, so the harmonic mean is particularly sensitive to small positive rates. In a task where a slow stage consumes most of the time, that sensitivity is meaningful. It is not a defect to be removed by switching averages without changing the target quantity.

Units provide a useful check. Reciprocals of speeds have units hours per kilometre; averaging them and taking the reciprocal returns kilometres per hour. Multiplying a positive speed dataset by a constant multiplies its harmonic mean by that constant. Adding a constant to every rate generally has no simple corresponding effect. The scale and denominator should therefore be explicit in an explanation of the calculation.

An average unit price can also depend on how purchases are constrained. Equal spending amounts at different prices result in quantities inversely proportional to price, so the effective total price per unit can be a harmonic mean. Equal purchased quantities lead to an arithmetic mean of prices. Unequal spending or quantities requires the corresponding weights. Calling a value average price without explaining whether expenditure or quantity is fixed can conceal a substantial change in the economic question.

Before using any rate average, write down the total quantity and total exposure separately. If those totals can be obtained, their ratio is usually the clearest answer. Recognizing that this ratio equals a named mean is helpful for generalization, but the name is secondary. This approach avoids formula selection based only on the presence of words such as speed, price, or efficiency in a question.

§5.8 Robust location and optimization

Robustness describes how a procedure responds to particular departures from its assumptions or to unusual observations. It is not an all-purpose guarantee of truth. The arithmetic mean is sensitive to an observation's magnitude because every value enters the total without a bound on its contribution. The median is insensitive to how far an already extreme value moves, provided that its rank remains beyond the middle. Both can be seriously affected by systematic sampling or measurement bias.

Consider 2, 4, 4, 7, and 13 again. The mean is six and the median four. Replacing thirteen by one hundred and thirteen raises the mean to twenty-six while the median remains four. The sample size, order of the first four observations, and middle value are unchanged. This example isolates numerical sensitivity. It does not establish that the larger observation is wrong or that a scientific total should disregard it. If the values are genuine expenditures, the higher mean accurately reflects the increased total.

The arithmetic mean minimizes the sum of squared deviations from a proposed location. To see this, write each deviation from a as the deviation from x-bar plus x-bar minus a. After squaring and summing, the cross term vanishes because deviations from x-bar sum to zero. The resulting expression is the sum of squared deviations from x-bar plus n times the square of x-bar minus a. The second term is nonnegative and is zero only at a equal to x-bar.

The median minimizes the sum of absolute deviations. Imagine moving a proposed location slightly to the right without passing an observation. Each observation to its left contributes a slightly larger absolute distance, while each observation to its right contributes a slightly smaller distance. The total decreases while more points lie to the right than to the left and increases when the balance reverses. The minimum occurs at a middle value or, for an even sample, throughout the interval between the two middle values.

The familiar even-sample median is the midpoint of that minimizing interval. Other points in the interval also minimize total absolute deviation. Thus an optimization property and a conventional reported median are related but not identical definitions. With four observations 2, 4, 7, and 13, every location from four through seven gives the same minimum sum of absolute deviations, while the conventional numerical median is 5.5. This distinction helps explain why certain median definitions can legitimately differ.

A trimmed mean removes a declared number or fraction of observations from each end of the sorted sample and averages the remaining observations. It reduces the influence of extreme values while retaining an averaging interpretation for the central portion. State the trimming rule, including how a fractional count is rounded. A ten-percent trimmed mean for a small sample may remove no observations under one rule and one from each end under another. Never hide this choice.

A winsorized mean replaces selected extreme observations by retained boundary values instead of removing their rows. It preserves the number of entries but changes the extreme numerical values. Trimming and winsorization therefore answer modified descriptive questions and should not be reported as though they were the ordinary mean of the unaltered dataset. Their usefulness depends on a justified analysis plan, especially when extreme observations are scientifically important rather than contamination.

Resistance should not be created by inspecting which transformation makes a desired conclusion look stronger. Decide the summary and any robust alternative using the measurement scale, the target question, and a transparent plan. It is often informative to report the ordinary mean and median together, along with a plot and an explanation of their difference. This lets readers see both total-sensitive and rank-based aspects without pretending that one number tells the whole story.

For descriptive reporting, robust summaries can support an investigation but should not automatically remove observations. A box-plot flag means a point lies beyond a conventionally defined fence. It is not an instruction to delete the value before calculating a mean. Investigate source records and substantive plausibility first. If an analysis excludes a verified error or an ineligible record, preserve the reason and report the eligible counts before and after the decision.

The choice between squared and absolute loss also appears in later modelling courses. Squared loss penalizes large deviations more strongly, whereas absolute loss treats each additional unit of distance equally. An introductory location problem is a useful place to understand these consequences without the complexity of a full regression model. The same principle will return when we justify least-squares regression in unit nine.

§5.9 Location from grouped data and incomplete information

Grouping replaces individual measurements by interval membership. Once the original values are unavailable, exact location measures may no longer be recoverable. A midpoint-based mean assigns every observation in a class the class midpoint and applies frequency weights. It is exact only if the total deviation from midpoint across the represented observations sums to zero. The grouped table alone usually cannot establish that condition, so describe the result as an approximation.

Suppose two observations lie in the interval from zero inclusive to ten exclusive and two in the interval from ten inclusive to twenty exclusive. The midpoint-based mean is ten. Actual values 0, 1, 10, and 11 have mean 5.5, while 8, 9, 18, and 19 have mean 13.5. Both datasets produce the same interval counts. The grouped mean approximation cannot distinguish them. This is information loss, not an error that more decimal places can repair.

If every class has known finite lower and upper boundaries, bounds for the overall mean can be formed by replacing each value with its class's lower boundary and then with its upper boundary. The true mean lies between the resulting weighted boundary means, with endpoint inclusion interpreted according to the interval definitions. For the preceding example the lower boundary mean is five and the upper boundary mean is fifteen, with the upper endpoint not attainable under the right-open intervals. Open-ended classes can prevent a finite bound in one direction.

The midpoint approximation's maximum absolute error can also be bounded when class widths are finite. Each observation differs from its midpoint by at most half its class width, subject to endpoint details. The absolute mean error is therefore no larger than the frequency-weighted average of half-widths. This conservative bound does not assume values are evenly spread within classes. A smaller actual error requires additional information about within-class positions.

Grouped quantiles often use linear interpolation inside the class containing the required cumulative position. The procedure assumes an approximately uniform cumulative increase through that class. The cumulative table identifies an interval containing a relevant rank but does not reveal the exact within-class values. An interpolated grouped median or percentile should therefore be labelled as an estimate based on that assumption, not as the same exact quantity obtained from individual observations using type-seven interpolation.

A formula for a grouped quantile typically begins with the lower class boundary, adds class width times a fraction, and calculates that fraction from the target cumulative count minus the count below the class divided by the class frequency. Different choices of target cumulative position can appear in textbooks. State the rule and retain the class boundaries. If the target lies at a boundary or the class frequency is zero, inspect the definition rather than dividing blindly.

Grouped data can still support exact statements about counts at class boundaries. If thirty out of one hundred observations fall below twenty under a declared boundary rule, that proportion is exact for the grouped table. A percentile lying somewhere between twenty and thirty is less precisely located. A useful report distinguishes known boundary counts from interpolated within-class locations. This separates retained evidence from an added model.

Missing observations create another kind of incomplete information. If the missing values have no known range, the complete-data mean may be unbounded in principle even when the observed-data mean is calculable. If all values are known to lie between finite limits, calculate worst-case bounds by putting every missing value at the lower and then the upper limit. These bounds can be wide, but they honestly represent what is and is not determined by the records.

Never substitute a grouped approximation or an observed-only mean for a complete-data quantity without saying so. A table may describe all eligible units but retain coarse values; a dataset may retain precise values but omit many eligible units. These are different limitations. A reproducible summary names the quantity actually calculated, the information used, and the assumptions required to connect it to a broader target.

When individual values are available, calculate summaries from them and use grouping for presentation. If grouped figures are the only source, preserve the original table and report uncertainty due to grouping separately from later sampling uncertainty. The teaching examples here emphasize that a numerical estimate can be internally reproducible while still depending on an approximation. Accuracy includes communicating that dependence.

§5.10 A location comparison and simulation audit

The location explorer allows the learner to alter one high observation in an otherwise fixed five-value dataset. It displays the arithmetic mean and median together with the sorted values. At the initial values 2, 4, 4, 7, and 13, predict a mean of six and median of four before running the control. Moving the final observation upward by ten should increase the mean by two while leaving the median unchanged as long as the observation remains beyond the middle ranks.

This controlled change illustrates sensitivity to magnitude, not a choice to discard the observation. The display should retain the changing observation in the data table so that its contribution remains visible. The accompanying interpretation should explain that the mean preserves the total and the median follows the middle ranks. Neither should be labelled universally more accurate. Their accuracy depends on which descriptive quantity the user intended to calculate.

At each setting, verify that five times the displayed mean equals the displayed total within a sensible rounding tolerance. Verify that the sorted middle observation equals the displayed median. Check the control's permitted range and reset behaviour. A simulation that updates the chart but leaves an old numerical summary would undermine the lesson, so chart labels and text must be regenerated from the same current data state.

Weighted averages deserve a separate hand calculation even when software is used. For group sizes ten and thirty and means seventy and eighty, recover the implied group totals and then the combined mean 77.5. Compare it with the equal-group mean 75. Explain the target of each result before declaring one appropriate. This comparison tests understanding of denominators more directly than a large exercise consisting only of repeated multiplication.

For rate calculations, reconstruct the totals. Equal sixty-kilometre legs at speeds thirty and sixty take three hours altogether and give an overall speed of forty. Equal one-hour legs at those speeds cover ninety kilometres in two hours and give forty-five. A useful interactive illustration would change the exposures and show the totals, but the numerical reasoning already supplies a complete explanation. We use simulations selectively where changing a parameter reveals behaviour, rather than attaching an animation to every formula.

For proportional change, use factors. A twenty-percent rise followed by a twenty-percent fall gives a factor of 0.96, and the equivalent two-period constant factor is its square root. Keep at least several significant digits during intermediate calculation and round the final percentage appropriately. Rounding the constant factor to 0.98 and then squaring yields 0.9604, close but not exactly equal to the original factor. A discrepancy caused by displayed rounding should be explained rather than silently ignored.

The three worked questions attached to this unit address combining group means, selecting the correct average for equal-distance travel, and comparing mean with median after an extreme-value change. Each solution names the observational unit and denominator before calculating. Additional comparisons involving geometric means and quantiles appear in the prose so that the unit covers more than the three assessed question types without overloading the reader with repetitive solution cards.

A good final location report might read: among the one hundred eligible recorded visits, the median wait was twelve minutes using type-seven interpolation, the mean was fifteen minutes, and the upper quartile was nineteen minutes; five missing wait records were excluded and reported separately. The example numbers are illustrative, but the reporting structure is substantive. It identifies the data, convention, unit, eligible count, and missingness decision.

Before proceeding, explain in words why averaging category codes is usually invalid, why an average of group means needs group sizes for a combined-person mean, and why equal-distance and equal-time speeds use different weighting. Then calculate one type-seven quartile and one proportional-change factor without a software menu. These tasks connect measurement, aggregation, order, and multiplication to the formulas. The next unit adds variation, because location alone cannot reveal how far observations spread around their centre.

THREE WORKED QUESTIONS

Step-by-Step Statistics Solutions

Three original questions connect calculations, definitions and interpretation. Open each solution to follow the reasoning.

intermediate Example 5.1: Combining people rather than groups

Ten students have mean score 70 and thirty students have mean score 80. Calculate the mean across all forty students and the equally weighted mean of the two group means. State the target of each result.

Recover totals
$$T=10(70)+30(80)=3100$$

The implied group score totals are seven hundred and two thousand four hundred. The groups are disjoint and measured on the same score scale, so their totals and counts can be added.

Use the person denominator
$$\bar x=3100/40=77.5$$

Dividing by forty gives 77.5. Group sizes are the weights because the target treats each student equally.

Name the other target
$$(70+80)/2=75$$

The equal-group average is seventy-five. It describes an average of two group means with each group given equal importance. It does not describe the equally weighted forty-person mean. Whether it is useful depends on the stated institutional question.

intermediate Example 5.2: Equal distance versus equal time

A traveller covers 60 kilometres at 30 kilometres per hour and another 60 kilometres at 60 kilometres per hour. Find the overall speed. Compare it with one hour at each speed and explain the two averaging rules.

Add distance and time
$$v=120/(60/30+60/60)=40$$

The two equal-distance legs take two hours and one hour. Total distance is one hundred and twenty kilometres and total time three hours. Overall speed is forty kilometres per hour.

Recognize the harmonic mean
$$H=2/(1/30+1/60)=40$$

Because distances are equal, this equals the harmonic mean of thirty and sixty. The arithmetic mean forty-five would not preserve total travel time for these legs.

Change the exposures explicitly
$$v_{\rm equal\ time}=(30+60)/2=45$$

One hour at each speed covers thirty plus sixty kilometres in two hours, giving forty-five. Equal time produces an arithmetic mean. The correct rule follows from accumulated distance divided by accumulated time, not from the word speed alone.

intermediate Example 5.3: Sensitivity of mean and median

The values 2, 4, 4, 7, and 13 have mean six and median four. Replace only thirteen with one hundred and thirteen. Recalculate both summaries and explain their different sensitivity without assuming the new value is an error.

Track the total change
$$\bar x_{\rm new}=(2+4+4+7+113)/5=26$$

The total rises by one hundred while count remains five. Therefore the mean rises by twenty, from six to twenty-six.

Track the ranks
$$\operatorname{median}=4$$

The sorted middle observation remains four, so the median remains four. Increasing a value already above the middle does not change that middle rank in this example.

Connect to the question

The mean accurately reflects the changed total and the median accurately reflects the unchanged middle position. Resistance to one extreme magnitude is useful for some questions, while total preservation matters for others. Check provenance separately before treating an unusual value as an error.