Statistical Reasoning, Populations and Evidence
Statistical questions, populations and samples, descriptive reasoning, random selection, causality, summation and reproducible evidence.
§1.1 Statistics begins with a question, not a formula
Statistics is the disciplined study of variation using data. It provides ways to describe observations, compare groups, investigate relationships, and assess what an incomplete collection of observations can tell us about a larger setting. Its usefulness comes from connecting a question to a defensible chain of measurement and reasoning. A calculator can produce a mean from any list of numbers. Statistical work asks whether the numbers belong together, what their mean represents, and which conclusions the collection permits.
Imagine that a university wants to understand how long students wait for laboratory equipment. The phrase “waiting time” is not yet a variable. Does waiting begin when a student enters the laboratory, joins a queue, requests an instrument, or finishes the previous experiment? Does it end when equipment is assigned or when the experiment actually begins? A department comparing two laboratories needs the same definition in both places. Otherwise, the apparent difference may be a difference in recording rules. The first statistical task is to make the question operational.
A well-formed investigation identifies the people or objects of interest, the quantity to be observed, the conditions under which observation occurs, and the decision the result will inform. For the laboratory question, these might be undergraduate experiment sessions during a specified semester, minutes between joining the equipment queue and receiving an instrument, ordinary teaching days rather than examination weeks, and whether another instrument would materially reduce delays. Each choice narrows the claim. A result about one semester is not automatically a result about all future semesters.
Variation is the reason a single observation rarely resolves the question. Students arrive at different times, experiments require different instruments, and equipment sometimes fails. Some variation reflects genuine differences between sessions. Some reflects random timing. Some comes from inconsistent measurement. The analyst needs methods that preserve useful variation while avoiding conclusions based on irrelevant differences. This is why statistics is more than replacing a list by one average.
Statistical reasoning works in both directions. A scientific mechanism can suggest what data to collect: a bottleneck model predicts larger delays when demand approaches capacity. Data can also suggest a question: a plot might reveal that delays peak on certain afternoons. These directions should be distinguished. If an explanation is invented after seeing an unusual pattern, the same data cannot provide a fully independent test of that explanation. A later measurement period can help assess whether the pattern repeats.
The basic workflow is question, design, measurement, organization, analysis, interpretation, and communication. These stages influence one another, but they should remain visible. Discovering a limitation during analysis may require returning to design rather than applying a more elaborate formula. Finding that the recording system excludes students who abandon the queue is a design problem. No numerical summary of completed waits can recover those missing experiences without additional information.
Statistics also involves choices about relevance. A small average delay may coexist with a small group experiencing very long waits. A policy concerned with fairness may need percentiles and group comparisons, whereas a policy concerned with total equipment use may need sums and weighted means. The “best statistic” depends on the question. A measure that answers one question well may answer another badly without any arithmetic error.
Throughout this book, a statistical result should be read as a statement with a scope, a definition, a method, and a limitation. Numbers become evidence when these pieces are supplied. When they are missing, apparent precision can conceal uncertainty. The purpose of introductory statistics is to make those pieces explicit and to develop the mathematical tools for organizing them carefully.
§1.2 Populations, samples, units, and the boundaries of a claim
A population is the complete collection of units about which an investigation seeks information. A sample is the collection actually observed or selected for observation. An observational unit is the entity to which a recorded value belongs. These definitions sound simple, but their correct use prevents several serious mistakes. The population may be students, laboratory sessions, households, plants, transactions, or moments in time. Its boundaries depend on the research question rather than on what is easiest to measure.
Suppose a campus has 8,000 enrolled students. An investigator records travel time for 200 students entering one gate between eight and nine in the morning. The nominal target population might be all enrolled students, but the observed sample represents a much narrower process: students who entered that gate during that hour and agreed to participate. Students living on campus, arriving later, using another gate, or absent that morning may have different travel experiences. Calling the 200 observations “a sample of the university” does not remove those exclusions.
The target population describes the intended claim. The accessible population describes the units the study can realistically reach. A sampling frame is an operational list or mechanism used to identify eligible units. An enrollment register can be a frame for students, but an old register may include students who have left and omit recent entrants. A list of campus internet users is not a complete frame for all students. Coverage should be checked rather than assumed from the existence of a large list.
A unit of observation can differ from a unit of analysis. In a classroom study, each student may provide an observation, while the intervention is assigned to whole classes. Students in the same class share a teacher, room, schedule, and peer environment. Treating every student as an independent replication of the intervention exaggerates the amount of independent information. Likewise, ten readings from the same thermometer are not equivalent to ten independently manufactured thermometers when the question concerns manufacturing consistency.
Time can also define units. If one person reports sleep duration on seven nights, the dataset contains seven person-night observations but only one person. It can describe that person's week. It cannot establish how much people generally sleep. A table should include identifiers that make repeated observations visible. When identifiers are removed without retaining the study structure, later analysts may be unable to distinguish variation between people from variation within a person.
For a finite population of size $N$, the measured values can be written as $x_1,x_2,\ldots,x_N$. A population quantity, such as the mean of these values, is a parameter. A function calculated from a sample, such as the mean of $n$ observed values, is a statistic. In elementary notation, the population mean is often written $\mu$ and a sample mean $\bar{x}$. The distinction concerns what the quantity refers to; it is not a judgment that one kind of average is more legitimate than another.
A census attempts to obtain information from every unit in the defined population. It eliminates sampling uncertainty only if every relevant unit is actually measured and the measurements are correct. A census can still suffer from missing units, duplicate records, incorrect responses, and inconsistent definitions. A complete database of examination marks can accurately describe those recorded marks while failing to measure understanding beyond what the examination tested.
Some populations are conceptual rather than fixed lists. A manufacturer may care about future items produced under a stable process. An environmental scientist may care about concentrations under conditions that could occur again. Here a probability model can represent the generating process. Such a model requires assumptions about stability and comparability. The analyst should not describe a changing process as one timeless population simply because all observations share a column name.
Before calculating, write a scope sentence: “This analysis describes these units, during this period, under these eligibility rules.” Then write the desired generalization separately. If the second sentence is broader, identify the design or assumptions that connect them. This habit turns hidden extrapolation into an explicit part of the investigation.
§1.3 Describing data and making inferences
Descriptive statistics organizes and summarizes the observations at hand. Examples include a frequency table, a histogram, a median waiting time, or the correlation between two measured variables. Statistical inference uses data and assumptions to draw conclusions beyond the observed collection. Examples include estimating a population mean from a probability sample or evaluating whether a pattern is compatible with a specified model. A descriptive result can be exact for the observed data even when a generalization from it is uncertain.
Consider five recorded waits of 2, 4, 4, 7, and 13 minutes. Their arithmetic mean is six minutes. There is no inferential uncertainty about that calculation if the five numbers are correct. There is substantial uncertainty about whether six minutes describes all waits during the semester. That uncertainty depends on how the sessions were selected, how variable waits are, and whether the operating conditions changed. A mean followed by many decimal places does not answer those design questions.
Inference does not transform an unrepresentative collection into a representative one. A formula for a standard error can characterize variability under a sampling model. It cannot automatically account for people who were systematically excluded. If only successful internet connections were logged, the data cannot directly measure the experience of users who failed to connect. The issue is not that inference is useless; it is that the model must correspond to the process that produced the records.
An estimate is a numerical attempt to learn an unknown quantity. An estimator is the rule that produces the estimate from possible data. The sample mean is an estimator; its value in one dataset is an estimate. This distinction becomes useful when comparing procedures. Two estimators might return similar values in one sample while having different behavior across repeated samples. Later books study bias, precision, confidence intervals, and hypothesis testing in detail. Here we establish why those concepts are needed.
Uncertainty has several sources. Sampling uncertainty comes from observing some units rather than all eligible units. Measurement uncertainty comes from imperfect instruments or definitions. Model uncertainty arises when more than one plausible representation of the process exists. Missing-data uncertainty concerns unobserved quantities. Future uncertainty concerns events that have not occurred. These sources are not interchangeable. One interval calculated under a simple sampling model rarely captures all of them.
Inference is conditional on assumptions, but that does not mean it is arbitrary. Assumptions can often be justified through design, physical knowledge, diagnostic checks, or sensitivity analysis. For example, selecting units by a documented random procedure provides stronger support for a sampling calculation than informally asking whoever is nearby. Repeating an analysis under alternative reasonable definitions can reveal whether the conclusion depends heavily on an avoidable choice.
A frequent mistake is to equate absence of a detected effect with proof that no effect exists. A small or noisy study may be unable to distinguish several possibilities. Another is to equate an observed difference with a practical improvement. A difference can be precisely measured yet too small to matter for the intended decision. Descriptive summaries should retain units so that magnitude remains interpretable. A change of 0.2 minutes and a change of 20 minutes should not be described with the same vague word “better.”
The boundary between exploration and confirmation is also important. Exploration searches for patterns, unusual observations, and plausible explanations. Confirmation assesses a question specified independently of the evidence used to evaluate it. Exploratory analysis is valuable and should be reported honestly. Problems arise when a selected surprising pattern is presented as if it were the only relationship investigated. The number and nature of attempted analyses affect how surprising a selected result really is.
An introductory report should distinguish three layers: what was observed, what was calculated, and what is inferred. For example, “The 200 respondents reported a median journey of 35 minutes” is an observed-sample statement. “This suggests substantial travel burden among respondents” is an interpretation. “All university students typically travel 35 minutes” is a broader claim requiring additional support. Keeping the layers separate makes a report more informative and easier to critique.
§1.4 Variation, signal, noise, and measurement error
Variation is difference among observations. It can reflect meaningful differences between units, fluctuations within a unit, or errors in observation. Calling all variation “noise” can erase the phenomenon being studied. Differences in household expenditure may be the central scientific interest. Differences between repeated readings of a stable reference instrument may instead reveal measurement instability. The same numerical measure of spread can play different roles in different investigations.
A useful conceptual representation is observed value equals underlying quantity plus measurement error. Written algebraically, $x_i=t_i+e_i$, where $t_i$ represents the quantity of interest and $e_i$ represents the discrepancy introduced by measurement. This is a model, not a universal identity with observable components. The underlying quantity may itself change during measurement, and some errors depend on its magnitude. Nevertheless, the representation encourages a clear distinction between the phenomenon and the recording process.
Random measurement error fluctuates across repeated observations in a way that can sometimes be modeled. Systematic measurement error consistently shifts values or changes them according to a predictable pattern. A scale that adds two kilograms to every weight can produce very consistent readings while being inaccurate. Repeating its measurements many times reduces neither the constant offset nor the mistaken calibration. Precision describes consistency; accuracy describes agreement with the intended quantity or accepted reference.
Measurement error can alter relationships as well as averages. If heights are rounded to the nearest ten centimeters, differences within a rounding interval disappear. If one variable is measured with large independent error, its observed association with another can be weaker than the association between the underlying quantities. A scatterplot that looks diffuse may therefore reflect both natural variation and recording limitations. The analyst should inspect instrument resolution before interpreting fine differences.
Resolution is the smallest distinction an instrument or recording rule can reliably show. A digital display may show three decimal places even when the instrument is accurate only to one decimal place. Display precision and measurement quality are different. If travel times are recalled approximately, reporting a mean as 37.428571 minutes suggests more precision than the data justify. Calculation software should not determine the number of meaningful reported digits.
Variability can also change with conditions. A process may be stable at low temperature and unstable at high temperature. Combining the readings into one overall variance can conceal this structure. Recording relevant conditions allows the analyst to separate groups or inspect trends. An unexplained increase in spread may be scientifically important even when the mean remains constant. Conversely, a change in the composition of the measured population can change the overall spread without changing any individual's behavior.
Repeated observations require care. Ten consecutive measurements can be correlated because the instrument warms up, the operator adapts, or the environmental condition persists. Averaging them may still produce a useful summary, but calculations that assume independent errors need additional justification. A measurement schedule that spreads observations across meaningful operating conditions may provide more information than many readings taken seconds apart.
Outliers are observations that are unusual relative to a specified comparison. An outlier can be a recording error, a legitimate rare event, a member of another population, or evidence that the model is inadequate. Its unusual position is not sufficient reason for deletion. An investigation should first verify the record, identify whether the unit met eligibility rules, and consider the scientific explanation. Removing a genuine equipment failure from a study of reliability would remove precisely the event that matters.
A good measurement protocol documents instrument calibration, units, rounding rules, timing, missing-value codes, and the procedure for correcting mistakes. It records corrections without erasing the original value. This information may look less sophisticated than a statistical model, but it is often more important. Reliable computation on unreliable measurements produces a reliably computed mistake.
§1.5 Selection, random sampling, and why sample size is not enough
Selection determines which eligible units contribute information. A probability sampling design uses a known random mechanism to select units, with selection probabilities that permit appropriate inference. A convenience sample uses units that are readily available. A voluntary-response sample includes people who choose to participate. These approaches can all yield data, but their inferential strengths differ. The sampling method should be described directly rather than hidden behind the word “survey.”
Simple random sampling without replacement selects a fixed number of distinct units from a finite population so that every subset of that size has the same probability. It is a specific design, not a synonym for casual selection. If the population contains six units and the sample size is two, there are fifteen possible unordered pairs. Each pair has probability one fifteenth under the design. Choosing the first two names after sorting alphabetically is not a simple random sample, even if the investigator has no preference for particular students.
Equal selection probability for each individual does not by itself guarantee the simple-random-sample design. Selecting one of several whole classrooms can give each student the same chance of selection while producing highly dependent group membership. A design must specify how combinations of units are selected. Cluster and stratified designs can be useful, but their analysis should reflect the design rather than borrow a formula solely because individual probabilities look equal.
Selection bias occurs when the observed collection differs systematically from the target in ways relevant to the question. An online poll about internet reliability may disproportionately attract users whose connections currently work. A library survey conducted only during weekday mornings may exclude evening users. A survey of exam preparation posted in an enthusiastic study group may overrepresent highly engaged students. The direction of bias depends on the relationship between selection and the measured variable; it should be investigated rather than assumed.
Larger samples reduce random sampling variability under suitable designs. They do not automatically reduce systematic selection bias. A million voluntary responses can describe those respondents very precisely while providing a distorted view of all eligible people. Conversely, a smaller well-designed probability sample can support a more defensible population estimate. This is not an argument for small samples. It is an argument that design quality and sample size answer different problems.
Nonresponse is a selection process after units have been sampled. If 400 students are randomly invited and only 120 respond, the invitations may be representative while the responses are not. Students with a particularly strong opinion may respond more often. Reporting the invitation method without the response rate leaves out an important part of the data-generating process. Follow-up contact, comparisons with available frame characteristics, and careful weighting can help, but they do not guarantee complete removal of bias.
Sampling weights reflect how observed units represent the target population. If one group is sampled more intensively than another, an unweighted average can overrepresent it. A weighted mean uses the weights in both numerator and denominator. However, weighting on one available characteristic cannot automatically repair differences on all unobserved characteristics. A survey adjusted to match age proportions can remain biased with respect to employment patterns if employment affects both participation and the outcome.
Random sampling should be distinguished from random assignment. Sampling selects units for observation. Assignment allocates units to treatments or conditions. A randomized experiment with volunteers can support a causal comparison for those volunteers under its design, while still having limited population generalizability. A random sample observed without an intervention can support population description without automatically supporting a causal claim. These two kinds of randomization solve different problems.
Practical sampling requires an auditable procedure. Record the frame version, eligibility rule, randomization method, replacement policy, invitations, refusals, and achieved sample. If a random number generator is used, save the code and, when appropriate, the seed. A seed improves reproducibility of selection, but reproducibility is not the same as representativeness. A perfectly reproducible convenience sample remains a convenience sample.
§1.6 Observational studies, experiments, and causal reasoning
An observational study records variables without assigning the exposure or treatment of interest. An experiment deliberately assigns a condition and observes its consequences. Both can produce useful evidence. The distinction matters because an observed relationship may reflect selection into conditions rather than an effect of the condition itself. Students who attend optional tutorials may differ from students who do not attend in motivation, prior preparation, available time, or access to resources.
Suppose tutorial attendees obtain higher examination marks. One possible explanation is that tutorials improve learning. Another is that already well-prepared students are more likely to attend. A third is that students with fewer work obligations can both attend and study more. The measured association combines these possibilities unless the design or additional assumptions separates them. Correlation describes a pattern between variables; it does not by itself identify the mechanism producing that pattern.
A confounder is a variable that creates or distorts an exposure-outcome association through its relationships with them. The exact causal role depends on the scientific model. Merely finding that a third variable is correlated with both does not establish the complete causal structure. A variable caused by the exposure can be a mediator rather than a confounder. Adjusting for it can remove part of the effect one intends to estimate. Causal reasoning must therefore precede automatic adjustment for every available variable.
Random assignment can make treatment groups comparable in expectation by breaking systematic links between assignment and pre-treatment characteristics. It does not ensure that every realized group has identical characteristics, especially in small experiments. It also does not prevent measurement error, noncompliance, attrition, interference between participants, or an inappropriate outcome definition. Randomization is powerful because it addresses a particular selection problem; it is not a universal guarantee of validity.
A control condition provides a comparison for what might have happened without the intervention. The appropriate control depends on the question. Comparing a new teaching method with no teaching can answer a different question from comparing it with the current method. Equal time, materials, and instructor attention may be necessary if the intended comparison concerns the method rather than extra resources. Design should make the competing explanations as explicit as possible.
Blinding limits the influence of knowledge about assignment on behavior or measurement. In some laboratory settings, the person measuring an outcome can be unaware of the assigned condition even when participants cannot be blinded. In teaching studies, students generally know which materials they receive, but examination marking can sometimes be blinded. The practical possibility and purpose of blinding should be described rather than treated as a box that must always be ticked.
Ethical constraints affect what can be randomized and measured. A study should not deliberately expose participants to serious harm simply because assignment would simplify statistical inference. Consent, privacy, equitable recruitment, and proportional data collection belong to study design. Ethical design is not an obstacle outside statistics; it helps determine which questions can be answered responsibly and how credible the resulting evidence is.
An outcome should be specified before results are inspected when the study intends to evaluate a particular intervention. If a team measures attendance, grades, confidence, satisfaction, and sleep, then reports only whichever result looks favorable, the final story conceals the selection process. Reporting the full set of intended outcomes and explaining changes reduces this problem. Exploratory findings can be useful when labeled as exploratory and followed by further investigation.
Causal language should match evidence. “Students who attended reported higher confidence” states an association in the observed data. “The tutorial increased confidence” states an effect of an intervention. The second statement requires a stronger connection between design and conclusion. An introductory analyst should learn to identify that difference even before studying formal causal models. Careful language protects both the reader and the scientific claim.
§1.7 Summation notation and the algebra of statistical summaries
Summation notation provides a compact language for repeated addition. If $x_1,\ldots,x_n$ are observations, then $\sum_{i=1}^{n}x_i$ means their total. The index $i$ is a placeholder identifying the observation included in each term. Replacing it by another unused index does not change the sum. What matters is the range and the expression being added. Understanding the notation allows one to read and derive statistical formulas rather than memorize their surface appearance.
A constant added once for every observation contributes $nc$ to the total. Therefore, $\sum(x_i+c)=\sum x_i+nc$. A constant multiplied by each observation can be factored out: $\sum ax_i=a\sum x_i$. These rules explain how the mean changes when a measurement scale changes. If a temperature recorded in one scale is transformed by a linear rule, its mean transforms by the same rule. The total, however, acquires the constant shift once for each observation.
The sample mean is $\bar{x}=n^{-1}\sum x_i$. Subtracting the mean from every observation creates centered values. Their sum is zero because $\sum(x_i-\bar{x})=\sum x_i-n\bar{x}=0$. This identity is exact algebra, not an approximate property of large samples. It becomes central to variance, covariance, and least-squares regression. If software produces a substantial nonzero total of centered values, either a different mean was used, some observations were excluded, or numerical/recording errors need attention.
The sum of squares and the square of a sum are different. The expression $\sum x_i^2$ adds the individually squared values. The expression $(\sum x_i)^2$ squares the total and includes cross-products between observations. For two values, the latter is $x_1^2+2x_1x_2+x_2^2$. Confusing these expressions produces incorrect variance formulas. Parentheses are part of the mathematics and should not be removed for typographical convenience.
Expanding centered squares gives a useful identity. Begin with $(x_i-\bar{x})^2=x_i^2-2\bar{x}x_i+\bar{x}^2$. Summing and using $\sum x_i=n\bar{x}$ yields $\sum(x_i-\bar{x})^2=\sum x_i^2-n\bar{x}^2$. This expression can simplify hand calculations. In computer arithmetic, subtracting two large nearly equal numbers can lose precision, so a mathematically equivalent identity is not always the most numerically stable implementation. Later units return to that distinction.
For a grouped table, a value $v_j$ with frequency $f_j$ represents repeated observations of that value. The total is $\sum_j f_jv_j$, and the number of observations is $\sum_j f_j$. The corresponding mean divides these two sums. If $v_j$ is a class midpoint rather than an exact observed value, the calculation is an approximation. The algebra of frequencies is exact; the representation of all values by midpoints may not be.
Double sums describe two-index data, such as measurements for students within classes. The expression $\sum_j\sum_i x_{ij}$ adds observations within each class and then across classes. A mean of class means is not necessarily the overall student mean. When class sizes differ, each class mean must be weighted by its size to reproduce the student-level average. This is a mathematical issue before it becomes a policy issue about whose experience receives equal weight.
Products use a different notation, $\prod_i x_i$, meaning multiplication of the terms. Products arise in geometric means and probability models. Logarithms convert a product of positive quantities into a sum of logarithms. The positivity condition matters: a formula involving logarithms cannot be applied to zero or negative values without an appropriate modification and interpretation. Statistical notation should always be read together with its domain restrictions.
When a formula is unfamiliar, translate it into a small example. Identify the index, write out three terms, keep units visible, and check a limiting case. A mean of identical values should return that value. A measure of variation should be zero when every observation is equal. Such checks do not replace a proof, but they reveal many transcription errors before those errors enter a larger analysis.
§1.8 Data-generating processes and the meaning of a statistical model
A statistical model is a simplified representation connecting possible observations to a process or set of assumptions. A model can describe a distribution of values, a relationship between variables, or a sampling procedure. It is useful when its simplifications preserve the features needed for the intended question. A model is not a declaration that the world exactly follows a mathematical equation.
For example, a model for daily equipment failures might assume a constant operating environment and comparable observation periods. If the laboratory doubles its operating hours, the number of failures per day can change even when reliability per operating hour remains the same. A model using counts without exposure information can mistake increased use for worsening quality. Identifying the data-generating process clarifies which denominators and covariates are needed.
Models also specify which observations are comparable. A concentration measured upstream and one measured downstream may have different expected values. Treating them as interchangeable observations from one distribution erases their spatial context. If the question concerns river-wide variability, combining them may be appropriate with careful design. If the question concerns pollution added between locations, the comparison must retain location. The same records can support different analyses depending on the question.
Independence is an assumption about how information in one observation relates to information in another. Measurements collected from the same household or consecutive time points often share influences. Independence should not be inferred from separate rows in a spreadsheet. A row is a storage convention. The underlying sampling and measurement process determines dependence. This distinction matters because several later inferential formulas treat independent observations as separate information contributions.
Identical distribution means that observations share a specified distributional structure. It can be plausible when units are selected under stable conditions, but implausible when different groups or periods have different mechanisms. The common phrase “independent and identically distributed” combines two assumptions, neither of which follows automatically from sample size. A study can have identically distributed but dependent observations, or independent observations with different distributions.
Statistical models should be assessed using both subject knowledge and data. A distribution that fits a histogram reasonably well may still imply impossible values outside the observed range. A regression line may approximate a relationship within the available measurements while becoming physically impossible under extrapolation. The analyst should consider the domain in which the model is intended to operate. A model can be useful locally without being a global law.
Sensitivity analysis studies how conclusions change under alternative defensible choices. An analyst might compare results with and without a verified extreme observation, under two plausible missing-data assumptions, or using alternative reasonable group boundaries. The purpose is not to choose whichever analysis produces the preferred answer. It is to reveal which conclusions are stable and which depend on uncertain decisions. Report those dependencies rather than concealing them.
A model can guide action even when it is imperfect. A simple queue summary can identify recurring high-delay periods without predicting every individual wait. A descriptive regression can summarize how two measurements vary together without identifying a causal mechanism. Usefulness should be assessed against the actual purpose. An unnecessarily complex model can make assumptions harder to inspect while adding little practical information.
Model-based statements should contain conditional language where needed. “Under the assumption of a stable measurement process” is not empty caution if stability determines the validity of the comparison. State the assumption in a way a reader can investigate. Avoid generic phrases such as “assuming all statistical conditions” that hide the actual requirements. A transparent simple analysis is easier to improve than an impressive-looking analysis with unspecified assumptions.
§1.9 Responsible reporting, reproducibility, and a first investigation
A statistical report should let a reader understand how the result was produced. This requires the question, eligibility rules, observation period, measurements, exclusions, transformations, summaries, and limitations. It does not require publishing private identifiers. Reproducibility means that the documented data and procedures can regenerate the reported results. Replication concerns whether a new investigation under comparable conditions obtains compatible findings. A reproducible error is still an error, and a single reproducible analysis is not a replication.
Keep raw records separate from cleaned data. Cleaning should be represented by documented rules or code, not unexplained manual edits. If a travel time of 350 minutes is corrected to 35 after checking the original form, record the original value, the corrected value, the reason, and who made the correction. If the value cannot be verified, preserve the uncertainty rather than silently choosing the value that improves the analysis.
An exclusion rule should follow the study question. If the target concerns ordinary laboratory sessions, a separately identified emergency maintenance day may fall outside the target. If the target concerns the actual experience of all students during the semester, excluding that day may remove relevant experience. The reason for exclusion matters more than whether the resulting graph looks cleaner. Report the number of affected records and compare conclusions when a reasonable alternative rule exists.
Units belong in tables, axes, and prose. A label such as “time” is incomplete when seconds and minutes are both possible. A percentage needs a denominator. A claim that attendance increased by ten percent is ambiguous unless it distinguishes a relative percentage change from a change of ten percentage points. From 40 percent to 50 percent is a ten-percentage-point increase and a 25 percent relative increase. Both calculations can be correct while communicating different quantities.
Privacy protection starts at data collection. Collect identifiers only when necessary, restrict access, and avoid releasing small-group combinations that can reveal individuals. Replacing names with codes does not automatically make a dataset anonymous if age, program, location, and event date identify a person. Synthetic examples in this book are deliberately invented teaching datasets; they should never be presented as measured evidence about actual students or institutions.
Authorship and review should also be described accurately. A text can be useful without pretending that a professor reviewed every sentence. Readers deserve to know what was checked and what remains uncertain. Computational verification of a numerical result is different from expert evaluation of a study design. An honest report names the type of verification rather than using a broad label such as “validated” without explanation.
For a first investigation, choose a modest question with a feasible measurement process. Compare equipment waiting times across two clearly defined sessions, summarize the lengths of a set of leaves, or describe the distribution of commute distances in a documented sample. Plan the table before collecting data. Include identifiers for repeated units, units of measurement, eligibility flags, and any relevant conditions. A small well-documented study is an excellent place to learn because its mistakes remain visible.
After collection, inspect the records before calculating summaries. Check missing values, impossible values, duplicate identifiers, inconsistent units, and whether totals agree with the sampling log. Make a simple graph. Then choose summaries that answer the question: a mean for total burden per observation, a median for a middle-ranked experience, a percentile for a high-delay threshold, and a spread measure for consistency. Explain why each quantity is relevant rather than reporting every available statistic.
Conclude at the level the design permits. A descriptive pilot can identify a plausible problem and guide a larger study without proving a general law. A comparison can reveal that further investigation is warranted without proving causation. Good statistical communication is precise about both findings and uncertainty. That precision is a strength: it tells the reader which decisions the evidence supports and which questions still require better data.
§1.10 Reading the sampling simulation and checking your reasoning
The sampling explorer accompanying this unit uses a finite artificial population so that the true population mean is known. This is a teaching advantage that real investigations rarely have. The population contains two groups with different measured values. The controls change sample size and the tendency of the selection procedure to favor one group. The display compares the observed sample mean with the mean of the full constructed population.
Begin with an equal-probability selection rule and repeat the sampling operation. Individual sample means vary. That variation is expected even when the design has no systematic preference for a group. Increasing sample size ordinarily reduces the variation under this design. It does not make every possible sample mean identical to the population mean unless the entire finite population is observed. In sampling without replacement, larger fractions of the population leave less uncertainty about the remaining units.
Next favor the higher-valued group in selection. Repeat the operation with a large sample. The sample means may become tightly concentrated around a value above the population mean. This demonstrates how an estimate can be stable and still systematically misrepresent the target. Precision and bias should be read separately. The explorer does not show that all selection biases point upward; changing which group is favored can reverse the direction.
The groups are intentionally simple. In real surveys, selection may depend on several characteristics, some unmeasured. A clear simulator cannot recover those unmeasured relationships. Its purpose is to illustrate a mechanism, not to estimate the bias of a particular real-world study. The artificial population values, sampling rule, and reset behavior are displayed so that the learner can inspect the assumptions rather than treat animation as evidence.
When testing a simulator, check values as well as movement. The mean of the full artificial population should agree with a direct calculation. If the sample includes every eligible unit under a without-replacement design, its mean should equal the population mean. The sample count should not exceed the population count. Reset should restore the documented initial settings. Invalid inputs should be rejected or handled explicitly rather than producing a plausible-looking graph from undefined quantities.
Use the simulation to ask a prediction before changing a control. For example, predict whether favoring a group will change the center of repeated sample means or mainly change their spread. Then compare the display with the prediction and explain the mechanism. This is more educational than moving sliders until a pattern looks interesting. A disagreement can reveal a misconception, a hidden assumption, or a software error.
Several reasoning checks summarize this unit. A sample is not representative merely because it is large. A census is not error-free merely because it attempts complete coverage. Separate rows are not proof of independent observations. A correlation is not proof of a causal effect. A reproducible calculation is not proof that the variables measure what the question requires. Each statement points to a design or definition that should be inspected before a numerical conclusion is trusted.
The next units develop the descriptive tools promised here. They show how measurement scales affect valid summaries, how tables preserve denominators, how graphics reveal structure, and how averages and measures of variation respond to transformations. These tools should be learned as connected ways of answering questions. Their mathematical identities are useful precisely because they can be checked, explained, and linked back to the observations that gave them meaning.
Step-by-Step Statistics Solutions
Three original questions connect calculations, definitions and interpretation. Open each solution to follow the reasoning.
A finite population consists of the four values 1, 2, 3, and 4. Calculate its mean. List every unordered sample of two distinct units, calculate each sample mean, and average those six means. Explain what this exact enumeration shows and what it does not establish for an arbitrary sampling process.
The population contains four equally weighted units. Its total is ten, so its mean is 2.5. An unordered sample without replacement cannot repeat a unit and treats the pair in either order as the same sample.
The six pairs are (1,2), (1,3), (1,4), (2,3), (2,4), and (3,4). Their means are 1.5, 2, 2.5, 2.5, 3, and 3.5. Their sum is fifteen and their equally weighted average is 2.5.
If each of those six samples is equally likely, the sampling average of the sample mean equals the finite-population mean. One realised sample can still differ from 2.5. Unequal selection probabilities or volunteer selection do not inherit this conclusion without examining their actual design. Enumeration verifies a property of this specific equal-probability mechanism.
A university has 1,000 eligible students. One hundred voluntarily answer a satisfaction form, and ninety report satisfaction. Calculate the observed proportion and response rate. Evaluate the claim that 90 percent of all eligible students are satisfied. Give a defensible description of the result.
Among the one hundred respondents, ninety have the attribute. The observed respondent proportion is 0.90. The response rate is one hundred divided by one thousand, or 0.10. These denominators refer to different questions.
Nine hundred eligible students did not respond. Voluntary participation may be related to satisfaction, so the recorded subset is not automatically representative. The counts alone do not tell us how many nonrespondents are satisfied.
A defensible statement is that ninety percent of the one hundred voluntary respondents reported satisfaction, with a ten-percent response rate among eligible students. A statement about all students requires additional evidence or justified assumptions. More precise arithmetic cannot remove the selection limitation.
For the waiting times 2, 4, 4, 7, and 13 minutes, calculate the total, mean, and signed deviations from the mean. Verify their sum and explain why their average is unsuitable as a measure of variation.
The five observations total thirty minutes, so their equally weighted arithmetic mean is six minutes. This is the recorded-data mean; it does not require any inference about a broader population.
Subtracting six gives deviations -4, -2, -2, 1, and 7. Their sum is zero, as required by the arithmetic-mean identity. The deviations retain minutes as their unit.
The signed average is zero despite the eleven-minute observed range. Positive and negative distances cancel. Absolute or squared deviations avoid that cancellation and are developed in the variation unit. Zero signed average is an algebraic check, not evidence of zero spread.