Statistics & Probability Introductory major course 100% Free Open Access
Chapter 3 • Theory & Derivations

Classification, Frequency Tables and Cross-Tabulation

Frequency and cumulative tables, class boundaries, histogram density, conditional proportions, grouped approximations and Simpson’s paradox.

§3.1 From individual records to a statistical classification

Classification organizes observations into categories or intervals chosen for a stated purpose. It makes a large collection easier to inspect, but it also changes the resolution at which the data are represented. A list of exact waiting times retains distinctions that a table of short, medium, and long waits removes. A useful classification preserves the distinctions needed for the question while making the main structure visible.

For a categorical variable, the categories should be defined before frequencies are interpreted. If transport responses include “bus,” “walk,” and “bus plus walk,” the classification must specify whether the last response is a separate category or contributes to two indicators. A one-way frequency table ordinarily assumes mutually exclusive categories. Multiple-response tables are valid, but their category totals should not be interpreted as a partition of respondents unless the rules make them one.

For a quantitative variable, intervals can replace exact values. A waiting-time table might use zero to less than five minutes, five to less than ten, and ten to less than twenty. These intervals are mutually exclusive under a half-open convention: the lower endpoint is included and the upper endpoint is excluded. A wait of exactly five minutes therefore enters the second interval. Stating the convention prevents double counting at boundaries.

Collectively exhaustive categories include every eligible observed value. An open-ended final category can cover large values, but it limits calculations that need a midpoint or width. A category “twenty minutes or more” has no finite upper endpoint and therefore no uniquely defined midpoint. Giving it a convenient midpoint without stating an assumption creates false precision. The category can still be useful for reporting the proportion above a meaningful threshold.

The smallest unit of recorded data matters for classification. Exact event counts are integers, while a time rounded to the nearest minute represents an interval of possible underlying durations. Both might be stored as integers, but their class boundaries mean different things. A table of counts from zero through nine is an exact category for integer values. A table of rounded measurements from zero through nine may correspond to underlying intervals with half-unit boundaries, subject to the rounding rule.

Classifications should be stable enough for comparison. If a report changes the definition of “long wait” from over ten minutes to over twenty minutes between periods, the resulting percentages do not measure the same quantity. The apparent improvement can be entirely due to the boundary change. Either retain common thresholds or reconstruct comparable groups from the original data. If only grouped tables remain, exact reclassification may be impossible.

Too many classes can obscure the overall pattern by leaving most groups with very few observations. Too few can conceal important clusters or unusual values. There is no universal class count that answers every scientific question. Sample size, measurement precision, distribution shape, and meaningful thresholds should guide the choice. Numerical rules for suggesting a bin count are starting points, not substitutes for judgment.

Keep the individual records whenever feasible. A grouped table is a presentation and analysis tool, not an adequate replacement for raw observations when later analyses need exact values. Retaining the records permits alternative intervals, precise quantiles, and checks of grouped approximations. If privacy requires releasing only grouped data, explain which calculations remain exact and which require assumptions about the unobserved positions within classes.

Classification therefore connects measurement and communication. It should reveal a meaningful pattern without silently creating new information. Before reporting a table, ask whether every observation enters exactly the intended number of categories, whether boundary cases are unambiguous, and whether the reader can recover the denominator. These checks make the later numerical summaries trustworthy.

§3.2 Frequency, relative frequency, and percentage tables

An absolute frequency $f_j$ counts observations in category or class $j$. If the classes partition the usable observations, their frequencies sum to the usable sample size $n$. A relative frequency is $p_j=f_j/n$, and a percentage is $100p_j$. The relative frequencies sum to one before rounding. These identities are valuable validation checks because they connect the displayed table to the underlying record count.

A table should distinguish eligible units, usable responses, and displayed categories. Suppose 120 people were eligible, 100 returned a questionnaire, and 90 answered the transport question. A category count of 27 can be reported as 30 percent of transport respondents, 27 percent of returned questionnaires, or 22.5 percent of eligible people. Each is a different quantity. The appropriate denominator follows the question, and the table should name it.

Percentage totals can differ slightly from 100 because of rounding. Three exact proportions of one third each become 33.3, 33.3, and 33.3 percent when rounded to one decimal place, summing to 99.9 percent. That small discrepancy is not a data error. A footnote can explain it. Changing one displayed percentage merely to force a total of 100 should follow a stated rounding procedure rather than an arbitrary edit that obscures the original calculation.

Large discrepancies require investigation. If mutually exclusive category percentages sum to 130, possible explanations include multiple-response coding, duplicated units, inconsistent denominators, or calculation errors. If they sum to 80, categories may be omitted or missing responses may have been included in the denominator without being displayed. The analyst should reconcile the counts before interpreting the percentages.

A zero frequency means no observed unit belongs to the category under the collection and classification rules. It does not prove that the category is impossible in the population. An empty category can result from a small sample, limited coverage, or a rare event. A blank cell can mean a zero, unavailable information, suppressed information, or not applicable. Use distinct symbols or annotations so that readers do not infer the wrong meaning.

Tables should use a meaningful order. Nominal categories can be arranged alphabetically, by frequency, or by a subject-specific order. Ordinal categories should ordinarily retain their natural order. Quantitative intervals should be arranged by their numerical boundaries. Sorting ordinal ratings by frequency can make it harder to see whether responses lean high or low. The order is part of the communication, not just a formatting choice.

When comparing groups with different sizes, percentages often make proportions easier to compare, while counts retain information about how much data supports each percentage. A 50 percent category based on two observations is different in evidential weight from the same percentage based on two thousand observations. Display both when possible. This book's descriptive tables do not automatically attach inferential precision to every percentage; that requires a design and model.

Frequency tables can also describe event counts rather than people. A table of loans by book subject counts transactions. A person can contribute several loans. If the report later describes the percentages as “students interested in each subject,” it changes the unit without justification. The caption should specify the counted unit: records, transactions, people, sessions, or another entity.

The complete workflow is count, reconcile, divide, label, and inspect. Count using the stated categories. Reconcile totals with the collection log. Divide by the intended denominator. Label units and scope. Inspect whether missingness or category choices affect interpretation. This sequence is simple enough to apply by hand and important enough to apply even when a software package generates the table automatically.

§3.3 Class limits, boundaries, width, and midpoint

Class limits are the stated endpoints used to name a group of recorded values. Class boundaries specify the actual interval used for assignment or display. For continuously measured data, a half-open interval such as $[10,20)$ is usually unambiguous: ten is included and twenty belongs to the next interval. The width is upper boundary minus lower boundary. The midpoint is their average when both endpoints are finite.

For values rounded to the nearest whole unit, a named class of 10 through 19 can correspond approximately to underlying measurements from 9.5 to less than 19.5. The half-unit adjustment follows the stated rounding resolution. It should not be applied mechanically to every integer-valued variable. Counts are intrinsically discrete, and a hypothetical underlying continuous interval is not always the quantity of interest.

Boundary conventions are especially important when some values equal a threshold exactly. If one class is labelled “0–10” and the next “10–20,” a recorded value of ten appears eligible for both. Software may choose one convention silently, while a hand tabulation chooses another. State the inclusion rule and ensure that the implementation matches it. A clear label such as “0 to less than 10” eliminates much of this ambiguity.

Equal-width classes simplify visual comparisons and some approximate calculations. Unequal-width classes can be preferable when the data or decision thresholds require detail in one region and broader grouping elsewhere. For example, a waiting-time study might distinguish short delays carefully and combine a sparse long tail into wider classes. The unequal widths must be retained in any histogram calculation. Equal-looking bars with unequal numerical widths misrepresent the variable.

A midpoint summarizes the position of an interval, not the exact average of the observations inside it. If all values in a class lie near the lower boundary, replacing them by the midpoint overstates their contribution to the total. If they lie near the upper boundary, it understates it. Midpoint calculations are approximations unless the class representative is known to equal the class's actual mean.

An open-ended interval does not have a defined width or midpoint. A category “at least forty” cannot be drawn as an ordinary finite-width histogram bin without choosing an artificial upper boundary. A frequency bar for that category can still display its count, but it should be presented as a categorical bar rather than as a continuous density rectangle. The distinction prevents a graphical convention from implying nonexistent numerical information.

Class boundaries can be subject-specific. Regulatory limits, instrument resolution, educational thresholds, or physically meaningful ranges may be more useful than a generic numerical rule. However, boundaries chosen after inspecting the data can emphasize or hide patterns. If the report's purpose is comparison across groups, use the same boundaries or explain why different boundaries are necessary. Otherwise, visual differences can be artifacts of classification.

The width should use the same units as the variable. If duration is measured in minutes, a class from five to fifteen has width ten minutes. Frequency density then has units observations per minute, while relative-frequency density has inverse-minute units. Keeping these units visible makes it easier to understand why density heights change when units or widths change even though class probabilities remain the same.

A practical table should include interval notation, frequency, and any needed representative value or width. This makes the downstream calculation auditable. If a report provides only a named category and an average computed from midpoints, the reader cannot assess the approximation without guessing the boundaries. Numerical summaries should preserve the information needed to judge their own limitations.

§3.4 Cumulative frequencies and threshold questions

A cumulative frequency adds frequencies up to a specified boundary. For ordered classes, the cumulative count after class $j$ is $F_j=\sum_{k\leq j}f_k$. The cumulative relative frequency divides this count by the intended total. These summaries answer questions about the fraction of observations below a threshold and provide the basis for cumulative plots and grouped quantile approximations.

The threshold convention must match the class convention. With intervals $[0,5)$, $[5,10)$, and $[10,20)$, the cumulative frequency at the end of the second class counts values less than ten, not values less than or equal to ten. A recorded value of exactly ten belongs to the third class. This difference can matter when many observations are rounded to threshold values. Avoid changing “less than” to “at most” casually in prose.

For exact observations, the empirical distribution function records the fraction at or below each value. If the data are sorted, the function increases in steps at observed values. Repeated values produce larger steps. A grouped cumulative table is a coarser representation: it records totals at class boundaries without showing exactly where observations lie inside intervals. Linear interpolation between boundaries adds an assumption rather than recovering the exact empirical function.

Cumulative counts should never decrease as the upper threshold increases. Their final value should equal the total usable count for an exhaustive classification. A decrease indicates an ordering, arithmetic, or data-entry error. A final value below the total indicates omitted categories, while a value above the total indicates overlap or duplication. These checks are easy to automate and should precede interpretation of a cumulative graph.

Threshold questions often have direct practical relevance. A laboratory might ask what fraction of waits is below ten minutes, rather than which bin has the highest count. A cumulative table can answer that question exactly when ten is a class boundary. If the threshold lies inside a class, only bounds are available without additional information. The exact fraction below twelve cannot be recovered from a single class covering ten to twenty.

One can state bounds from grouped data. If 40 observations are below ten and another 20 lie from ten to less than twenty, the number below twelve lies between 40 and 60. If there are 100 observations in total, the corresponding fraction lies between 0.40 and 0.60. Assuming a uniform distribution within the class would give an interpolated value, but the assumption should be stated. The bound remains valid without it.

A reverse cumulative table counts observations at or above a threshold. It can be useful for exceedance questions, such as the fraction waiting at least twenty minutes. The relationship to the forward cumulative table depends on whether endpoints are included. For exact values, the complement of “at most twenty” is “greater than twenty,” whereas the complement of “less than twenty” is “at least twenty.” Precise endpoint language avoids off-by-one-category mistakes.

Cumulative summaries should not be used for unordered nominal categories unless the ordering itself has a justified purpose. Adding percentages for walking, bus, and bicycle in alphabetical order does not create a meaningful threshold variable. The mathematics of cumulative addition is possible, but the interpretation lacks an underlying order. A table can be arithmetically correct and conceptually unhelpful.

When communicating cumulative results, state the threshold, direction, denominator, and whether the value is exact or interpolated. For example, “Sixty percent of recorded waits were shorter than ten minutes” is clearer than “the cumulative frequency was sixty.” The former tells the reader what was counted and how to interpret it. The latter leaves the central statistical meaning implicit.

§3.5 Histograms with unequal intervals: why area must carry frequency

A histogram represents a quantitative distribution using rectangles over numerical intervals. The key principle is that rectangle area, not necessarily rectangle height, represents frequency or relative frequency. When all intervals have equal width, raw-frequency heights produce areas proportional to frequency because the common width is a constant factor. When widths differ, raw-frequency heights distort the comparison.

For a bin of width $w_j$ containing $f_j$ observations, frequency-density height is $f_j/w_j$. The rectangle's area is then $(f_j/w_j)w_j=f_j$. For a probability-density histogram, the height is $f_j/(n w_j)$, and its area is $f_j/n$. Summing the areas gives the total count or one, respectively. This area identity should be checked whenever a histogram is constructed.

Consider eight observations in an interval of width ten and twelve observations in an interval of width twenty. Frequency densities are 0.8 and 0.6 observations per unit. The wider interval contains more observations but has lower density. Drawing bars of heights eight and twelve would make their areas 80 and 240, implying an area ratio of one to three rather than the true count ratio of two to three. The difference is a graphical error, not a matter of stylistic preference.

Histogram height should not be interpreted as probability at one exact value. In a continuous model, a single point has no interval width, and density times a nonzero width gives approximate probability over an interval. A density can exceed one when the numerical scale is compressed, while total area remains one. A tall density bar is therefore not evidence that a probability exceeds 100 percent. The axis label should say density rather than probability when heights carry density units.

Changing the measurement unit changes density heights. Converting minutes to seconds multiplies bin widths by sixty and divides density heights by sixty, leaving areas unchanged. Counts and probabilities are preserved. This transformation is another useful validation check. A histogram that keeps both widths and density heights numerically unchanged after unit conversion represents a different area and therefore a different distribution.

The bin origin matters as well as width. Two histograms with the same width but shifted boundaries can make clusters look different, especially with small samples. Inspect several reasonable choices before making strong claims about the number of modes. A pattern that exists only under one carefully chosen origin deserves caution. Retaining a dot plot or raw-value display can help distinguish genuine clustering from a boundary artifact.

Sparse data do not require a histogram. A small list can be shown with a dot plot or stem-and-leaf display that retains exact values. A histogram is especially useful when many observations make individual marks hard to inspect. The analyst should choose the representation that communicates the relevant structure with the fewest hidden assumptions. More elaborate graphics are not automatically more informative.

The interactive histogram explorer in this book displays bin counts, widths, densities, and total area together. When a width changes, the height adjusts so that the count remains represented by area. Its purpose is to make the area rule visible. It does not estimate the true population density from the small teaching dataset, and changing bins does not create new observations.

To audit a histogram, verify the intervals partition the values, the widths match the numerical axis, the heights follow the stated convention, and the areas reconcile with the total. Label the variable and units. State whether the vertical axis is frequency, relative frequency, frequency density, or probability density. Those labels are mathematical information, not interchangeable descriptions of the same picture.

§3.6 Cross-tabulation and the three meanings of a percentage

A cross-tabulation classifies observations by two variables simultaneously. Rows might represent transport modes and columns might represent whether a respondent arrived late. Each cell counts units with the corresponding combination. Row totals, column totals, and the grand total are called margins because they summarize one classification after combining over the other. The same table supports joint, row-conditional, and column-conditional percentages.

Let $n_{ij}$ be the count in row $i$ and column $j$. The joint proportion is $n_{ij}/n$, where $n$ is the grand total. It answers what fraction of all observed units has that combination. The row-conditional proportion is $n_{ij}/n_{i+}$, where $n_{i+}$ is the row total. It answers how the column variable is distributed among units in that row. The column-conditional proportion divides by $n_{+j}$ and answers the reverse conditioning question.

Suppose there are 40 bus users, of whom ten are late, and 60 walkers, of whom six are late. The proportion late among bus users is 25 percent. The proportion of all respondents who are both bus users and late is ten percent. The proportion of late respondents who use the bus is ten out of sixteen, or 62.5 percent. These are all correct calculations, but they answer different questions. Confusing them is a frequent source of mistaken conclusions.

The direction of conditioning should be stated in ordinary language. “Among bus users” identifies the row denominator. “Among late respondents” identifies the column denominator. “Of all respondents” identifies the grand-total denominator. Words such as “bus users account for 62.5 percent of lateness” can be ambiguous, because they may be read as a causal contribution rather than a composition of observed late respondents. Prefer explicit descriptions of the counted units.

Conditional percentages should sum to 100 within the chosen denominator group when categories are exhaustive and mutually exclusive. Row percentages sum across each row; column percentages sum down each column. A table displaying both can become crowded, so use separate tables or clear labels. Never ask the reader to infer the denominator from where a percent sign happens to appear.

Empty margins require special treatment. If a category has no observations, its conditional percentages have a zero denominator and are undefined. Displaying zero percent in every cell of that margin falsely suggests an observed distribution with no outcomes. A dash with a note such as “no observations in this category” is more accurate. Joint proportions for its cells can be zero because the grand total remains nonzero, but that is a different calculation.

Cross-tabulation can reveal association when conditional distributions differ across groups. In the bus example, the late proportion is 25 percent among bus users and ten percent among walkers. That difference is descriptive evidence of a relationship in the observed data. It does not show that taking the bus caused lateness; route length, departure time, and selection of transport can all be relevant. The table supports a pattern, not a complete causal explanation.

A table with many categories can contain sparse cells. Collapsing categories can simplify presentation, but the collapse should have a substantive rationale and preserve the question. Combining all unusual modes into “other” can hide a group with a very different outcome pattern. Retain the original classification and describe any aggregation so that another analyst can evaluate it.

Cross-tabulation is therefore both an organizing device and a reasoning device. It forces the analyst to distinguish joint membership from conditional experience and to inspect how totals combine. The later correlation unit provides numerical measures for some forms of association. The table remains essential because it shows the actual counts behind those measures.

§3.7 Aggregation, changing composition, and Simpson's paradox

An aggregate percentage combines groups using weights determined by their sizes. If the group proportions are $p_j$ and group sizes are $n_j$, the overall proportion is $\sum n_jp_j/\sum n_j$. A simple arithmetic mean of group percentages gives each group equal weight and generally answers a different question. This distinction becomes especially important when comparing methods whose observations have different group compositions.

Consider two fictional training methods, A and B, evaluated in easy and difficult tasks. On easy tasks, A succeeds in 90 of 100 cases, while B succeeds in 19 of 20. Their success rates are 90 and 95 percent. On difficult tasks, A succeeds in one of ten cases, while B succeeds in twenty of one hundred. Their rates are ten and twenty percent. In each task category, B has the higher observed success rate.

Combining the categories gives A 91 successes in 110 cases, about 82.7 percent. B has 39 successes in 120 cases, or 32.5 percent. The aggregate comparison reverses the within-category comparison. A's cases are overwhelmingly easy, whereas B's cases are overwhelmingly difficult. The aggregate rates combine performance with composition. This is a form of Simpson's paradox: a pattern within groups can reverse when the groups are combined.

The reversal is not an arithmetic contradiction. The aggregate percentages use different weights for the two methods. A's easy-task weight is 100 out of 110, while B's is twenty out of 120. When a high-rate category receives much more weight for one method, its aggregate can exceed the other's even if its within-category rates are lower. Writing the weighted-average expression makes the mechanism explicit.

There is no universal instruction to prefer the grouped result over the aggregate result. The relevant comparison depends on the question and causal structure. If the question is what happened among all actual cases handled by each method, the aggregate describes that experience. If the question concerns performance under a common task mix, a standardized comparison can be more relevant. If the grouping variable is affected by the method, conditioning on it may introduce another interpretation problem.

Standardization applies a common set of weights to the group-specific rates. With equal weights for easy and difficult tasks, A's standardized rate is half of 90 percent plus half of ten percent, or fifty percent. B's is half of 95 percent plus half of twenty percent, or 57.5 percent. These describe a hypothetical equal task mix. They are not the actual aggregate rates and should not be reported as if they were directly observed totals.

The common weights should represent a meaningful reference population or decision scenario. Equal weights are convenient for teaching but may not reflect actual use. A policy comparison might use the target population's task distribution. Report the chosen weights and examine whether reasonable alternatives change the conclusion. A standardized figure should make its constructed nature visible.

Changing composition can also produce trends over time. An institution's overall completion rate can fall while every program's completion rate rises if enrollment shifts toward programs with lower rates. Conversely, an overall improvement can result from a shift toward easier cases without any within-group improvement. Time-series tables should therefore inspect relevant composition before attributing an aggregate trend to improved performance.

The lesson is to retain group information and denominators. Aggregation is sometimes necessary, but it should be understood as a weighted operation. A single percentage can conceal substantial differences in case mix, and a comparison of percentages can conceal different weights. Cross-tabulations provide the detail needed to identify these mechanisms before a persuasive but incomplete story is written.

§3.8 What grouped data can and cannot tell us about a mean

For exact values with frequencies, the mean $\bar{x}=\sum f_jv_j/\sum f_j$ is exact because each representative $v_j$ is an observed value repeated $f_j$ times. For interval classes, replacing every value by the class midpoint gives a grouped mean approximation. The formula looks the same, but the information supplied to it differs. An approximation should not be described as an exact reconstruction of the original mean.

Suppose two observations lie in $[0,10)$ and two in $[10,20)$. Using midpoints five and fifteen gives a mean of ten. Yet the actual observations could be 0, 1, 10, and 11, with mean 5.5, or 8, 9, 18, and 19, with mean 13.5. Both datasets produce the same grouped frequencies. The grouping has removed the within-class positions needed to distinguish their means.

Bounds can be calculated without assuming uniformity. If all observations in class $j$ lie between lower boundary $L_j$ and upper boundary $U_j$, the overall mean lies between the frequency-weighted averages of the lower and upper boundaries, with endpoint inclusion depending on the intervals. For half-open intervals, the upper bound may be a strict bound. This gives a range of possible means rather than an invented exact value.

The approximation error can also be bounded. For a finite interval of width $w_j$, a value differs from its midpoint by no more than half the width, subject to endpoint conventions. The absolute error of the overall midpoint mean is therefore bounded by the weighted average of half-widths. This is a conservative statement; errors in different classes may cancel. It shows why narrower intervals ordinarily retain more information about the mean.

An assumption of a uniform within-class distribution makes the midpoint the expected class value under that model. It does not establish that the realized class mean equals the midpoint. When many values are rounded or concentrated near a threshold, uniformity may be particularly implausible. The analyst should use known class means when available or retain the grouped result as an explicitly approximate summary.

Open-ended classes prevent simple finite bounds unless additional scientific limits are known. A final category “at least twenty minutes” contains no finite upper boundary, so a finite upper bound for the mean cannot be obtained from that table alone. A maximum operating duration or a separate recorded maximum could supply additional information, but it must be justified. Assigning a midpoint by extending the previous width is a modeling choice, not a property of the data.

Grouped approximations affect measures of variation too. Replacing all values in a class by one midpoint removes within-class variation. A variance calculated solely from those representatives cannot generally reproduce the original variance. The decomposition into within-group and between-group sums of squares in a later unit makes this loss explicit. A grouped table retains variation between classes while concealing variation within them.

The consequences depend on the intended use. A broad histogram may be adequate for showing a rough distribution shape but inadequate for estimating a high percentile near a safety threshold. A grouped mean may suffice for a rough planning calculation but not for comparing very small changes across periods. Information quality should be judged relative to the decision's required precision.

Whenever a summary uses grouped data, label it as exact, approximate, bounded, or model-based as appropriate. Give the grouping intervals and the assumptions. This language is more informative than simply adding “approximately” to every number. It tells the reader which uncertainty arises from data reduction and which arises from sampling or measurement.

§3.9 Table design, weighted records, and reproducible computation

A readable statistical table has a descriptive title, clear row and column labels, units, counts, denominators, and notes on exclusions or approximations. The title should identify the population and period without claiming a broader scope than the data support. A table called “Student transport” is less informative than “Reported main transport mode among respondents to the October pilot survey.” The latter communicates the counted group and collection context.

Place totals where they help validation. A one-way table should ordinarily include a total row, while a two-way table should show both margins. If missing values are excluded from the percentage denominator, show their count separately. If categories are mutually exclusive, the counts should reconcile with the usable sample. These checks let readers inspect the same arithmetic that the analyst used.

Weighted tables need a distinction between observed counts and weighted estimates. If each observed unit has weight $w_i$, a weighted category total is the sum of weights for units in that category. It can be non-integer and may represent an estimated population count rather than a literal number of observed records. Display the unweighted sample size as well. A weighted total of ten thousand does not mean ten thousand people were actually interviewed.

Different weights answer different questions. Frequency weights indicate repeated identical observations. Sampling weights represent selection or adjustment factors. Exposure weights can represent time or quantity at risk. Treating them as interchangeable can produce wrong summaries and wrong uncertainty calculations. A table should explain the weight's meaning and whether it has been normalized. Normalizing all weights by a common factor leaves a weighted mean unchanged but changes weighted totals.

Reproducible tabulation begins with explicit rules. Specify category mapping, interval boundaries, endpoint inclusion, missing-value treatment, and whether percentages are joint or conditional. Apply those rules using code or a transparent hand calculation. Save the table-producing procedure with the cleaned dataset. Reproducibility is especially important when updated data arrive, because manual edits can lead to inconsistent classifications across versions.

Sorting and filtering should be documented. A table displaying only the five most common categories should show that other categories were omitted or combined. A table excluding respondents with missing group labels should state the exclusion. These operations can alter the denominator and perceived distribution. A compact presentation should not silently redefine the population.

Tables can reveal privacy-sensitive information when cells are very small. A count of one in a distinctive combination of age, department, and event date can identify a person even without a name. Reporting may require aggregation or suppression. If a cell is suppressed, distinguish it from a zero and consider whether totals reveal the value indirectly. Statistical communication should retain useful information while respecting the purpose and privacy conditions of collection.

Computational validation can include assertions that frequencies are nonnegative, margins equal sums of cells, relative frequencies lie between zero and one, and cumulative values are monotone. For mutually exclusive categories, every usable record should have exactly one assigned category. For histograms, numerical widths should be positive and total rectangle area should match the intended normalization. Assertions convert important reasoning rules into repeatable checks.

An empty dataset deserves explicit handling. A percentage with a zero denominator is undefined, not zero. A histogram with no observations has no empirical distribution to normalize. Software should return an informative message rather than a visually plausible chart of zeros. A clear failure is more trustworthy than silently substituting a default that readers might mistake for evidence.

§3.10 Auditing the histogram explorer and a complete table-to-claim workflow

The histogram explorer uses a fixed synthetic set of numerical observations. The controls change boundaries or widths, and the display reports how many values belong to each bin. A density rectangle is drawn using the count divided by width, or the relative count divided by width when normalized. The total area is displayed so that the learner can verify the central mathematical identity.

Begin by identifying the smallest and largest values. Choose intervals that cover them under the stated endpoint rule. Check that each observation enters exactly one interval and that the counts sum to the dataset size. Then inspect an observation lying exactly on a boundary. The displayed allocation should agree with the lower-inclusive, upper-exclusive convention, except for any explicitly stated handling of the final upper endpoint.

Next create unequal widths. Predict which heights should change if counts are held fixed. The height of a wider interval should be lower for the same frequency because its area must remain equal to that frequency. Read the numerical density rather than relying on the apparent height of the plotted rectangle. The axes and labels should make the units and normalization visible.

Changing bins may alter the appearance of a mode or gap. That change demonstrates sensitivity to data representation, not a change in the observations themselves. If the synthetic data contain repeated values, a bin boundary can collect or separate them in ways that make small-sample patterns look dramatic. Use the raw-data list to explain what changed. A simulation should make the distinction between data and display explicit.

Test the explorer at edge settings. A zero or negative width is invalid and should not produce a rectangle. A boundary outside the data range can yield an empty bin but should not lose the other observations. A final value equal to the upper display limit should follow a documented inclusion rule. Reset should restore the original data and controls. These tests are part of interpreting the simulation, not merely software maintenance.

The complete workflow for a table starts with a question. Decide whether the question concerns category composition, a threshold, a conditional distribution, or an aggregate comparison. Choose a classification that answers it. Reconcile counts and denominators. Select a suitable table or graphic. State any information lost through grouping. Finally, write a conclusion that identifies the observed units and avoids an unsupported causal interpretation.

For example, a statement that “one quarter of bus respondents were late” should be supported by the bus row count and the late count within that row. A statement that “bus users were most of the late respondents” uses a different denominator. A claim that “bus travel causes lateness” requires evidence beyond the cross-tabulation. The progression from table to claim should preserve these distinctions at every step.

This unit's three worked questions test frequency tabulation, unequal-width histogram density, and an aggregation reversal. They require calculation and interpretation together. A correct number without a correct denominator or endpoint convention is incomplete. Conversely, a thoughtful explanation should be supported by reproducible arithmetic rather than plausible verbal intuition alone.

By the end of the unit, a reader should be able to build and audit a one-way table, a cumulative table, and a cross-tabulation; distinguish class limits from boundaries; draw a density-correct histogram; and identify what grouped data no longer reveal. These are the structural skills on which the next unit's graphical reasoning and later numerical summaries depend.

THREE WORKED QUESTIONS

Step-by-Step Statistics Solutions

Three original questions connect calculations, definitions and interpretation. Open each solution to follow the reasoning.

intermediate Example 3.1: Unequal-width histogram bins

Eight observations occupy a bin of width ten and twelve occupy a bin of width twenty. Determine count-density heights and relative-frequency-density heights. Verify the areas and explain why using raw counts as heights misrepresents the data.

Calculate density heights
$$h_1=8/10=0.8;\quad h_2=12/20=0.6$$

Count density is count divided by width. Heights are 0.8 and 0.6 observations per unit. The narrower class is taller although its count is smaller, because its observations are concentrated in less horizontal width.

Verify area and relative scale
$$0.04(10)+0.03(20)=1$$

The areas are height times width, eight and twelve, recovering counts. Dividing heights by the total twenty gives 0.04 and 0.03 per unit; their areas are 0.4 and 0.6 and sum to one.

Reject the misleading encoding

Raw heights eight and twelve would produce areas eighty and two hundred and forty, a one-to-three ratio instead of the observed two-to-three ratio. For unequal widths, histogram area must encode frequency, so a density vertical scale is necessary.

intermediate Example 3.2: Three denominators in one table

Of 40 bus users, 10 arrive late. Of 60 walkers, 6 arrive late. Calculate the proportion late among bus users, the proportion both bus and late among all 100 people, and the proportion bus among the 16 late arrivals. Explain why the answers differ.

Condition on bus use
$$P_{\rm recorded}(L\mid B)=10/40=0.25$$

The relevant eligible subset has forty people, ten late. Thus lateness among bus users is one quarter. This is a within-row conditional proportion.

Use the full-table denominator
$$P_{\rm recorded}(L\cap B)=10/100=0.10$$

The joint cell bus and late contains ten of the one hundred recorded people, so its full-table proportion is one tenth. A joint proportion is not the same as a within-bus rate.

Reverse the condition
$$P_{\rm recorded}(B\mid L)=10/16=0.625$$

Sixteen people are late and ten of them are bus users, giving 0.625. This describes travel mode among late people, not lateness among bus users. The three questions share a numerator but have denominators forty, one hundred, and sixteen.

intermediate Example 3.3: A composition reversal

In an easy stratum, option A succeeds in 90 of 100 cases and B in 19 of 20. In a hard stratum, A succeeds in 1 of 10 and B in 20 of 100. Compare each stratum, the aggregate rates, and rates standardized to equal stratum weights. Interpret the reversal descriptively.

Compare like strata

In the easy stratum the rates are 0.90 for A and 0.95 for B; in the hard stratum they are 0.10 and 0.20. B has the larger recorded success rate in both strata.

Aggregate with actual counts
$$p_A=91/110;\quad p_B=39/120$$

A has ninety-one successes in one hundred and ten cases, approximately 0.827273. B has thirty-nine in one hundred and twenty, or 0.325. A appears better overall because its cases are heavily concentrated in the easy stratum while B has many hard cases.

Use a common composition
$$p_{A,*}=0.5(0.90)+0.5(0.10)=0.50;\quad p_{B,*}=0.575$$

Equal easy and hard weights give A a standardized rate 0.50 and B 0.575. This removes the particular composition difference from that descriptive comparison. It does not establish causal superiority, because case allocation and other factors still require a justified design.