Measurement, Data Collection and Data Quality
Measurement scales, operational definitions, questionnaires, data dictionaries, missing-value bounds and defensible cleaning.
§2.1 Variables, values, and operational definitions
A variable is a characteristic that can take different values across the units or occasions being studied. A value is the recorded outcome of that characteristic for one unit under a specified measurement rule. The difference matters: “temperature” is a variable name, whereas “28 degrees Celsius measured at noon in shade” describes a value together with conditions. A useful dataset makes both the variable and its recording rule clear.
Operational definition means specifying how an abstract concept becomes an observation. If a study concerns study effort, possible measurements include hours spent with a textbook open, time spent on focused work, number of completed exercises, or attendance at organized sessions. These measurements are related but not identical. A person can spend many hours near a textbook without learning efficiently. The analyst should not quietly replace the concept with whichever measurement is easiest and then forget the substitution.
An operational definition should specify the observational unit, reference period, instrument or question, units of measurement, and rules for unusual cases. “Weekly study hours” should say whether the week is a calendar week, the previous seven days, or a typical week recalled by the respondent. It should explain whether scheduled classes are included. Without those details, two respondents can give different interpretations to the same column heading.
Variables can be qualitative or quantitative. Qualitative variables identify categories, such as laboratory room or preferred transport mode. Quantitative variables represent amounts, such as journey distance or equipment operating time. A quantitative variable can be discrete, taking separated possible values, or continuous in its conceptual measurement model. Counts of completed experiments are discrete. Duration is usually modeled as continuous, even when recorded to the nearest minute.
The recording system can make a conceptually continuous variable appear discrete. A balance that records mass to the nearest gram produces integer-valued records. That does not make mass a count of objects. Conversely, a count divided by an exposure can produce a decimal-valued rate while still originating in discrete events. Classifying a variable solely by whether its spreadsheet entries contain decimal points is unreliable.
The same concept can be recorded at different levels of detail. Age can be recorded as date of birth, completed years, or a category such as 18–24. Each version supports different operations. Categories can improve privacy and simplify reporting, but they discard within-category information. Once only an age category has been retained, a later analyst cannot recover exact ages by replacing every category with its midpoint. Midpoints are an assumption, not recovered measurements.
A derived variable is calculated from other variables. Speed can be computed from distance and duration, and a response indicator can be computed from whether a form was returned. Derived variables require definitions just as primary measurements do. A pass/fail flag should state the threshold and treatment of missing or incomplete assessments. A missing assessment should not automatically become a failure unless that classification is part of the intended institutional rule.
Units and reference periods should travel with derived quantities. A concentration might be expressed per liter, an event rate per thousand operating hours, and an expenditure per household per month. These denominators are part of the variable. If they change across rows, the values cannot be combined without conversion or a model that handles different exposures. An unlabelled numerical column is therefore not a sufficiently defined variable.
Before collection, test an operational definition on a few realistic cases. Ask whether two trained observers would record the same value. Include a person with no study time, a session that crosses midnight, an interrupted measurement, and a response outside the expected range. The disagreements reveal ambiguities that are much easier to repair before the main dataset is collected.
§2.2 Nominal scales and the danger of arbitrary numerical codes
A nominal scale distinguishes categories without imposing an order. Blood group, laboratory building, and transport mode are common examples. The meaningful information is category membership: two observations are in the same category or in different categories. A numerical code can store this information, but the numbers do not acquire quantitative meaning merely because software accepts them as numeric input.
Suppose transport is coded 1 for walking, 2 for bus, and 3 for bicycle. The mean of these codes depends on the arbitrary assignment. Exchanging the codes for walking and bicycle changes the mean even though nobody changes transport. A valid interpretation of a nominal summary should survive a relabeling of categories. Counts, proportions, and statements about category membership do survive; arithmetic averages of arbitrary codes generally do not.
A frequency is the number of observations in a category. A relative frequency divides that number by a specified total. If 30 of 100 respondents report walking, the relative frequency is 0.30 and the percentage is 30 percent. The denominator should be the number of respondents for whom the question has a usable answer, or another explicitly justified population count. Using the full dataset count when many responses are missing changes the question from “among respondents” to something else.
A mode is a category with the largest frequency. There may be more than one mode if categories tie. The modal category is not necessarily selected by a majority. If frequencies are 40, 35, and 25 percent, the first category is most common but is chosen by fewer than half of respondents. Words such as “most common,” “majority,” and “nearly everyone” should not be used interchangeably.
Some nominal variables permit multiple selections. A respondent may use both bus and walking during a journey. If the questionnaire allows several modes, category percentages can sum to more than 100 percent without an error. The unit counted is still the respondent, but the categories are not mutually exclusive. A table should say “percentage selecting each option” rather than imply that every person belongs to exactly one category.
An “other” category can be useful but should be designed carefully. It can absorb genuinely uncommon responses, or it can conceal that the provided categories were inappropriate for much of the population. If a large fraction chooses “other,” review the free-text entries and the intended classification. Reclassifying them after collection requires a documented rule so that similar responses are treated consistently.
Missing information is not automatically a substantive category. “Unknown blood group” describes knowledge about the measurement, whereas a blood-group category describes a biological classification. In some reports, unknown values should be shown explicitly as a data-quality category. In others, they should be excluded from a denominator with the exclusion documented. The choice should follow the question and should never be hidden by recoding missing values as an ordinary category.
Binary indicators form a useful special case. If a variable is deliberately defined as 1 when a condition is present and 0 otherwise, its arithmetic mean equals the proportion with the condition. This is not an exception based on arbitrary coding; it follows from the explicit meaning of the indicator. Reversing the coding changes the question to the complementary condition, and the mean becomes one minus the original proportion.
For binary indicators, always state which condition is coded one. A column named “response” could mean response received, positive response, or response approved. Without the definition, a mean of 0.72 is ambiguous. Clear coding turns a convenient algebraic result into an interpretable statistical summary. Unclear coding turns the same calculation into a potential reporting error.
§2.3 Ordinal scales, ranks, and what ordering preserves
An ordinal scale supplies a meaningful order but does not by itself specify equal distances between adjacent categories. Ratings such as poor, fair, good, and excellent are ordered. Their ordering supports statements about higher and lower categories. It does not establish that the improvement from poor to fair is numerically identical to the improvement from good to excellent. Assigning scores 1, 2, 3, and 4 adds a distance convention beyond the original ordering.
Ordinal information is preserved by strictly increasing transformations of the numerical scores. If the scores 1, 2, 3, and 4 are replaced by 1, 2, 10, and 100, the order remains unchanged. A mean can change substantially under this transformation. That sensitivity explains why the arithmetic mean of a single ordinal item needs a justification about the score scale. It is not enough to say that the software produced the answer.
Medians, ordered category frequencies, cumulative proportions, and rank-based comparisons often fit ordinal measurements naturally. The median category identifies a middle position in the ordered observations. It does not require equal numerical spacing between categories. If an even number of observations falls around two distinct categories, reporting the interval of central categories or a specified convention is preferable to inventing a category halfway between them.
Ranks replace observed values by their relative positions. A strictly increasing transformation leaves ranks unchanged when ties are handled consistently. Ranks therefore provide a way to study ordering without relying on the original distance scale. They are useful when comparing preferences, performance positions, and monotonic relationships. Rank methods do not mean that all questions about magnitude disappear; they answer questions based on order rather than original-unit distances.
Ties contain information. If several respondents choose the same rating, their equality is part of the observation. A common ranking rule assigns tied observations the average of the positions they occupy. For values 2, 2, and 5, the first two occupy positions one and two and each receive rank 1.5. The value five receives rank three. Breaking ties arbitrarily creates distinctions that were not observed and can change a rank statistic.
Ordered response categories should be balanced and understandable. A satisfaction scale that offers “very dissatisfied,” “slightly dissatisfied,” “neutral,” “satisfied,” and “extremely satisfied” mixes different verbal intensities. Respondents may not interpret the distances or even the threshold labels consistently. Cognitive testing can identify this problem. More categories do not automatically produce a more reliable measurement if respondents cannot distinguish them.
Combining several ordinal items into a numerical score introduces additional assumptions. A questionnaire total may be useful when the items measure a common construct and the scoring procedure has evidence supporting it. The total is not automatically an interval measurement simply because several item codes were added. Item content, response behavior, scoring, reliability, and validity should be considered. This introductory book does not replace a full psychometric evaluation.
The distinction should be practical rather than dogmatic. Analysts sometimes report means of rating scores when treating the scoring convention as part of the measurement model. Such reports should acknowledge the convention and examine whether alternative summaries tell the same story. Reporting category distributions alongside a mean can reveal polarization that one average conceals. A mean rating of three could arise from everyone choosing the middle category or from opposing groups choosing extremes.
Ordinal data encourage a useful question: which conclusions depend only on order, and which depend on the numerical distances assigned? If a conclusion changes when reasonable alternative scores are used, it is sensitive to the scoring convention. That sensitivity should guide interpretation. The original ordered categories should be retained so that later analysts can investigate the assumption rather than inherit an unexplained total.
§2.4 Interval and ratio measurements: differences, zeros, and transformations
An interval scale supports meaningful differences in a chosen unit, while its zero point is a convention rather than an absolute absence of the quantity. Temperature in degrees Celsius is the standard example. A difference of ten degrees has the same scale meaning wherever it occurs, but twenty degrees Celsius is not twice as hot as ten degrees Celsius. The numerical ratio depends on the chosen zero and changes when another interval scale is used.
Permissible interval-scale transformations have the form $y=a+bx$ with positive $b$. They change the origin and unit while preserving ordering and relative differences. The mean transforms to $a+b\bar{x}$, and the standard deviation transforms to $b s_x$. Ratios of raw values generally do not survive the shift $a$. Statistics should respect this distinction. A coefficient of variation based on Celsius temperature can change simply because the same temperatures are expressed in Fahrenheit.
A ratio scale has a meaningful zero and supports ratios of values measured in a common unit. Mass, length, elapsed duration, and many event counts are examples when their definitions supply a true zero. Changing units takes the form $y=bx$ with positive $b$, without an arbitrary additive shift. A duration of twenty minutes is twice a duration of ten minutes because zero minutes represents no elapsed duration under the defined measurement.
The phrase “meaningful zero” should be tied to the actual variable. A calendar year numbered zero is not an absence of time. A clock reading of zero hours is a chosen daily reference. Elapsed duration since a specified event can have a meaningful zero. The same broad concept of time can therefore appear in different measurement structures. Analysts should inspect the variable definition rather than attach one scale label to every time column.
Negative values do not automatically rule out meaningful quantitative interpretation, but they complicate summaries based on ratios or logarithms. Net financial balance can be negative even though gross expenditures are nonnegative. Temperature deviations from a reference can be negative even if an absolute temperature scale is positive. A geometric mean of raw negative values is not defined by the usual real logarithmic formula. One should not add an arbitrary constant merely to make the formula run and then treat the result as an unchanged scientific quantity.
Differences and ratios answer different questions. If one group averages 30 minutes and another 20 minutes, the absolute difference is ten minutes and the ratio is 1.5. A report might emphasize either depending on the decision. If the same comparison concerns costs near a very small baseline, a large percentage increase can correspond to a small absolute amount. Reporting both can prevent the choice of scale from exaggerating or hiding practical significance.
Percent change divides a change by a reference value. Its interpretation depends on that reference. A change from 20 to 30 is a 50 percent increase relative to 20. A change back from 30 to 20 is a one-third decrease relative to 30. Equal absolute changes in opposite directions need not have equal percentage changes because their denominators differ. Percentage calculations should therefore identify the baseline rather than state only a number followed by a percent sign.
Units should be standardized before pooling values. A list containing some journey distances in kilometers and others in meters is not one coherent numerical sample until conversion occurs. Conversion should use a documented factor and retain enough precision for the intended analysis. Rounding each converted value aggressively before calculating can accumulate error. It is usually better to calculate with adequate precision and round the final report in accordance with the measurement quality.
Scale classification is a guide to interpretation, not a substitute for subject knowledge. A variable can have the formal appearance of a ratio measurement while its operational definition is inconsistent. Reported “hours studied” may mix different activities or reference periods. Correct transformation algebra cannot fix a poorly defined quantity. Meaningful statistical calculation requires both a defensible scale and a defensible measurement procedure.
§2.5 Instruments, questionnaires, reliability, and validity
A measurement instrument can be a physical device, a questionnaire, an observation checklist, or a software logging system. Each turns a phenomenon into a record through a set of rules. The analyst should understand those rules well enough to identify their limitations. A questionnaire about laboratory satisfaction and a sensor recording equipment use both require calibration in a broad sense: the recorded values should correspond to the intended quantity.
Reliability concerns consistency under conditions where consistency is expected. A ruler that gives nearly the same length in repeated measurements can be reliable. A questionnaire item that respondents interpret differently each time may be unreliable. Reliability can be evaluated through repeated observations, agreement between observers, or relationships among items, depending on the instrument. The appropriate form of consistency follows from the measurement purpose.
Validity concerns whether the measurement supports the intended interpretation or use. An examination can reliably distinguish students who memorized a particular procedure while failing to measure conceptual understanding. A consistently biased scale can be reliable but inaccurate. Reliability is often necessary for useful measurement, but it is not sufficient evidence of validity. An instrument should be assessed against the actual construct and decision, not merely against repeatability.
Question wording can create measurement error. “How much did you enjoy our excellent new library?” embeds an evaluation in the question. “Do you agree that the new library is better and easier to use?” asks about two properties at once. A respondent who finds it better but harder to use cannot answer unambiguously. Neutral wording and one construct per question generally make interpretation clearer.
Recall periods influence responses. Asking for the exact number of study hours in the previous year demands an unrealistic level of memory. Asking about the previous day is easier but may capture an unusual day. A reference period should balance recall quality with the intended time scope. Diary methods can reduce some recall problems but introduce participation burden and may change behavior because respondents know they are being observed.
Response options should cover plausible answers without overlapping. Categories such as “0–2 hours,” “2–4 hours,” and “4–6 hours” are ambiguous at the boundaries unless a half-open convention is stated. Categories should distinguish “not applicable” from “do not know” when these meanings matter. Providing only positive satisfaction options forces dissatisfied respondents into inaccurate records. Missing or skipped answers can then reflect questionnaire design rather than disengagement.
The order of questions can affect interpretation. Asking about a recent equipment failure immediately before asking for overall satisfaction may make that incident unusually prominent. Asking sensitive questions at the beginning can discourage participation before simpler information is collected. A pilot can compare alternative ordering and identify where respondents become confused or uncomfortable. The goal is not to manipulate answers but to reduce unintended influences.
Observer training matters for checklists. If “equipment unavailable” is recorded differently by different assistants, apparent differences across sessions may reflect the observer rather than the laboratory. Provide examples, boundary cases, and a method for resolving disagreements. Agreement should be evaluated on records where observers independently apply the rule. Agreement reached only after discussion does not measure independent consistency.
A pilot study is a small-scale test of the collection procedure. It can reveal unexpected response categories, unclear timing, excessive burden, missing identifiers, and software failures. A pilot is not automatically a miniature definitive study. Its primary purpose is to improve the design. Document resulting changes, because data collected under substantially different definitions may not be directly comparable with later observations.
Measurement quality should be communicated with the numerical results. State whether values were directly measured, self-reported, recalled, inferred from logs, or computed from other fields. Mention important resolution limits and known systematic errors. Readers can then judge whether a difference is large enough to be meaningful relative to the measurement process, rather than confusing fine numerical output with fine measurement.
§2.6 Sources of data and the distinction between a record and an event
Primary data are collected specifically for the investigation at hand. Secondary data were collected for another purpose and are reused. Neither category guarantees quality. A carefully maintained administrative database may be more reliable than a rushed primary survey. Conversely, a large secondary dataset may omit the variable needed for the question or define it in a way that makes the intended comparison invalid. Relevance and collection quality must be evaluated separately.
Administrative records are often designed for transactions rather than scientific description. A library system may record loans accurately but not visits that do not produce a loan. A hospital appointment system may record bookings rather than completed appointments. A website log may record requests rather than distinct people. Before counting records, identify the event that creates a record and the rules that create additional records for the same underlying event.
Duplicate rows can arise legitimately. A student may borrow three books, generating three loan records. Counting these rows gives the number of loans, not the number of students. If the question concerns users, a stable user identifier is needed. If the question concerns borrowing intensity, multiple records per user are meaningful and should remain visible. Deleting every repeated identifier would destroy relevant information for that second question.
Data from sensors require information about sampling frequency, outages, calibration, and storage rules. A sensor that records every second during stable operation but every millisecond during an alarm does not produce a uniform sample of time. An unweighted average of all stored readings can overrepresent alarm periods. The appropriate analysis may need time weighting or a regularized observation schedule. The storage policy is therefore part of the statistical design.
Public datasets should be evaluated through their documentation. Check population coverage, variable definitions, collection dates, changes in instruments, revisions, and missing-value codes. A column with the same name across years may have changed meaning after a questionnaire revision. A sudden apparent trend can then be a discontinuity in measurement. Metadata provide the information needed to distinguish a change in the phenomenon from a change in the records.
Synthetic data are constructed rather than observed. They are useful for teaching, testing software, and exploring the consequences of assumptions. They should be labelled explicitly. A synthetic dataset can demonstrate that an aggregation reversal is mathematically possible without establishing that such a reversal occurred in a real institution. Simulation output should be reported as behavior under the constructed model, not as empirical evidence about the outside world.
Combining sources can increase usefulness but introduces matching problems. An attendance file may use student numbers, while a survey uses self-entered email addresses. A failed match can represent a typo, a missing identifier, a changed account, or a person outside the intended population. The matching process should retain the number of matched and unmatched records and the rule used. Quietly dropping unmatched records can create a selected dataset whose coverage differs from both original sources.
A data provenance record traces where each field came from and how it changed. It can include source filename, collection date, version, transformation script, and responsible person. Provenance is particularly valuable when a source publishes revisions. Saving only the current download can make an old report impossible to reproduce. Retaining the relevant version and documenting changes allows results to be checked without guessing which file was used.
Source evaluation ends with a suitability judgment. Ask whether the records measure the desired quantity, cover the desired units and period, support the required denominator, and include enough information to handle known biases. A dataset can be perfectly suitable for one investigation and unsuitable for another. Its size, reputation, or convenient format should not replace that judgment.
§2.7 A data dictionary and an analysis-ready table
A data dictionary explains each variable in a dataset. At minimum, it should include the variable name, meaning, observational unit, data type, measurement unit, allowed values, missing-value codes, source, and any derivation rule. A well-designed dictionary lets another analyst distinguish a genuine zero from an unavailable measurement and an ordered category from an arbitrary code. It is part of the dataset rather than optional decoration.
An analysis-ready table usually puts one observational unit per row and one variable per column, with stable identifiers and consistent types. This arrangement makes many analyses easier, but it must reflect the actual study structure. Repeated measurements may require one row per unit and occasion, not one row per person with an ever-expanding set of columns. The table should make the key identifying a record explicit.
For a laboratory-wait study, a possible key is session identifier plus student identifier plus request number. A student can request different instruments within one session, so student identifier alone is insufficient. Variables might include request time, assignment time, instrument type, eligibility flag, wait duration, and reason for missing duration. The key can be checked for duplicates before any summary is calculated.
Use separate fields for separate concepts. A cell containing “12 minutes, instrument broken” mixes a numerical measurement with a status note. It is difficult to analyze and easy to misinterpret. Store duration in one field and the status in another. A field containing “less than 1 minute” is a censored or interval-valued measurement, not an ordinary number. Its treatment should be specified rather than silently converted to zero.
Avoid embedding presentation features in the raw data. Merged spreadsheet cells, subtotals mixed with observations, blank rows indicating groups, and colors carrying undocumented meanings can all obscure the table's structure. A displayed report may use these features, but the analysis dataset should contain explicit variables for the information. A group label should be repeated in a field or linked through a documented key, not inferred from the color of a cell.
Dates and times require formats and time-zone rules. A date written 03/04 can mean different calendar dates in different locales. A duration obtained by subtracting clock times can be wrong when the event crosses midnight or when an offset changes. Store unambiguous timestamps and define the local reporting date separately where needed. Do not interpret a formatted display as a complete time representation.
Numeric types should be chosen according to meaning. Identifiers that contain leading zeros should usually be treated as labels, not quantities. An instrument code “0017” should not become “17” if the coding system distinguishes them. Long identifiers can also be rounded by spreadsheet software when stored as numbers. Checking identifiers after import protects the ability to join files and identify duplicate records.
A dictionary should describe derived variables in enough detail to reproduce them. For wait duration, specify the start and end events, units, rounding, and treatment of negative or interrupted intervals. For a satisfaction category, specify the response options and coding. For a group variable, specify whether it was collected directly or assigned from another field. “Calculated automatically” is not a sufficient derivation description.
Documentation should evolve with the data. If a variable's definition changes, update the dictionary and retain the previous version. Record which observations used which rule if both remain in the dataset. A table can be technically well-formed while combining incompatible definitions. Analysis readiness therefore includes semantic consistency, not just clean column names and successful import.
§2.8 Missing values, nonresponse, and what cannot be recovered by a shortcut
A missing value indicates that the intended measurement is unavailable under the recording rule. It can arise because a question was skipped, an instrument failed, a record could not be matched, or the value was not applicable. These reasons may have different statistical implications. A missing measurement is not necessarily zero, and zero should not be used as a generic missing code for a variable where zero is a plausible value.
Complete-case analysis uses only records with all variables required for an analysis. It is convenient but can change the population represented by the data. If students with longer journeys are more likely to skip the travel-time question, the complete responses may underrepresent long journeys. The issue is not solved by reporting the number of complete cases. The relationship between missingness and the quantity of interest should be investigated.
Missing completely at random is a modeling condition under which missingness does not depend on the observed or unobserved values relevant to the analysis. Missing at random allows missingness to depend on observed information, while being conditionally unrelated to the missing values once that information is accounted for. Missing not at random describes situations where the missing values themselves remain relevant to missingness after conditioning on observed information. These terms refer to assumptions about a process, not labels determined by a missing-value count.
The word “random” in these terms can be misleading in ordinary language. A person deliberately skipping a sensitive question can sometimes be compatible with a missing-at-random model conditional on measured characteristics, but not necessarily. Conversely, an apparently accidental instrument outage can affect particular high-temperature periods and therefore be related to the missing measurement. The process should be described before a formal missingness assumption is proposed.
Replacing missing values with the observed mean preserves the mean of the completed variable by construction, but it artificially reduces variation and changes relationships with other variables. It treats uncertain unobserved values as if they were known equal to one number. Such a shortcut may be acceptable for a clearly labelled display experiment, but it is not a general inferential solution. The uncertainty introduced by missing data should not disappear merely because the spreadsheet contains no blank cells.
Simple bounds can sometimes communicate uncertainty without pretending to recover values. Suppose an event indicator is observed for 80 of 100 eligible units, and 24 observed units have the event. If nothing is known about the 20 missing indicators, the full-population event count lies between 24 and 44. The corresponding proportion lies between 24 and 44 percent. The observed-response proportion of 30 percent answers a narrower question and should not be substituted silently for that range.
The same principle applies to numerical variables with credible limits. If a questionnaire permits between zero and twenty study hours, unobserved responses can be bounded within that stated range only if the range is genuinely valid for the target measurement. Arbitrarily imposed limits can create misleading certainty. Bounds are most informative when their scientific or design basis is clear and when the analyst shows how the conclusion changes as the limits vary.
More advanced methods, including multiple imputation and model-based weighting, can address missing information under explicit assumptions. They require additional study and should not be invoked as magic repairs. An introductory analyst can still do useful work by recording reasons, displaying missingness by relevant groups, comparing complete and incomplete records on observed characteristics, and reporting sensitivity. These steps often reveal which assumption matters most.
Missingness should remain visible in reports. State the number eligible, the number observed for each variable, the number used in each calculation, and the handling rule. A table where every column has a different denominator should show those denominators. Otherwise, readers may compare percentages that describe different subsets and infer a relationship that the data do not support.
§2.9 Validation, cleaning, and a defensible correction log
Data validation checks whether records satisfy the intended structural and semantic rules. Structural checks include file readability, expected columns, valid identifiers, and unique keys. Semantic checks include plausible ranges, compatible units, logical ordering of dates, and allowed category combinations. Validation is most useful when rules come from the study design rather than from a desire to make the distribution look tidy.
A range check can identify impossible values without identifying every error. A wait duration of minus five minutes is inconsistent with ordinary event ordering and needs investigation. A duration of 500 minutes may be unusual but possible in an interrupted session. A range rule should distinguish impossible from merely unexpected values. Automatic deletion of all large values can remove important evidence about failures or unequal experiences.
Cross-field checks often reveal errors that single-column checks miss. If a person is recorded as absent but also has an equipment assignment time, the fields need reconciliation. If an end timestamp precedes a start timestamp, the event may have crossed midnight, a date may be incorrect, or the clocks may be unsynchronized. The appropriate correction depends on the original record. Changing the value to the nearest plausible number without evidence is not cleaning; it is invention.
Units can be checked through patterns but should be confirmed through provenance. A cluster of durations around 600 next to durations around 10 may suggest that some values are in seconds and others in minutes. Dividing selected values by sixty can make the distribution look smoother, but smoothness is not proof of a unit error. Verify which instrument or import procedure produced those records and document the conversion rule.
Categorical cleaning requires careful treatment of spelling, case, and meaning. “Bus,” “bus,” and “BUS” may safely map to the same category under a case-insensitive rule. “Public transport” may include bus, train, or ferry and cannot necessarily be collapsed to bus. A mapping table should distinguish formatting corrections from substantive reclassification. Preserve the original text when later review may be needed.
Duplicate detection should use the intended key and event definition. Two rows with identical values can represent two legitimate units that happened to have the same measurements. Two rows with different values but the same unique request identifier can be conflicting versions of one event. Removing duplicates based solely on identical numerical values can bias a dataset toward greater apparent diversity. Event identity, not visual similarity, should guide the decision.
A correction log records the identifier, field, original value, revised value, evidence, reason, date, and responsible person or script. Unresolved cases should be flagged rather than forced into a clean-looking result. Analysts can then reproduce both the original and corrected summaries. Comparing them indicates whether the report depends substantially on a small number of uncertain corrections.
Transformations used for analysis should also be documented. Converting units, constructing age groups, taking logarithms, winsorizing extreme values, and deriving indicators all change the representation of data. Some are useful, but none should be hidden. In particular, winsorization changes actual values; it should not be described as merely removing measurement error unless error has been established. Report the rule and its effect.
An auditable cleaning procedure improves both accuracy and collaboration. Another analyst can inspect the decisions, identify a mistaken assumption, and rerun the analysis without reconstructing undocumented edits. The objective is not a dataset with no unusual values. The objective is a dataset whose values and limitations correspond as closely as possible to the intended measurements.
§2.10 A measurement audit and the limits of numerical summaries
Consider a synthetic pilot table containing a student identifier, journey duration, transport mode, satisfaction rating, and a response timestamp. Before calculating, identify what one row represents and whether repeated rows are possible. Inspect the dictionary to determine whether duration is in minutes, whether satisfaction is ordinal, whether mode allows multiple choices, and which codes represent missing values. This audit supplies the information that the numerical formulas need.
Suppose the duration column contains 0, 15, 20, 25, and 999. A mean calculated directly from these values is 211.8. If 999 is the documented missing-value code, that calculation is wrong for the intended measurement. The observed mean after treating it as missing is fifteen minutes. If zero also means “not asked,” the observed mean becomes twenty minutes. The difference is entirely due to definitions, not competing statistical philosophies. A dictionary should resolve it.
Now suppose zero is a genuine duration for a student living in the classroom building. Removing it would introduce an error. The analyst should not decide whether a value is valid by whether it makes the average more plausible. Use the recording rule and original evidence. A correct analysis may produce an unexpected result, and an incorrect analysis may produce a reassuring one.
For satisfaction ratings, display the ordered category counts before considering a score average. If respondents are split between very dissatisfied and very satisfied, the average numerical code can resemble a neutral response that nobody selected. For transport mode, use frequencies and proportions rather than an average of mode codes. For duration, choose summaries that retain its quantitative units and inspect its distribution before deciding which central value is informative.
An audit should also ask whether the collected variable answers the intended question. Journey duration measured on a rainy examination day can accurately describe that day while being a poor proxy for a typical semester. Satisfaction measured immediately after a helpful staff interaction can be influenced by that event. These are scope issues rather than arithmetic mistakes. They belong in interpretation and may motivate a different collection schedule.
The three worked questions for this unit address complementary skills: choosing valid operations for measurement scales, understanding a percentage-point versus percentage change, and bounding a proportion with missing responses. Their solutions explain the definitions before calculating. This order is deliberate. Formula selection follows the variable and question, rather than the other way around.
No numerical simulation is required for this unit's central goals. The key task is to examine definitions, questionnaires, records, and missingness assumptions. A moving chart would not necessarily make those semantic choices clearer. The absence of a simulation is a pedagogical choice, not an unfinished feature. Later units use interactive displays where changing an input reveals a mathematical property directly.
At the end of the unit, a reader should be able to produce a short measurement specification for every variable, distinguish category codes from quantitative values, preserve missingness explicitly, and write a defensible cleaning log. These skills make subsequent tables, graphs, means, and correlations interpretable. Without them, a sophisticated calculation can answer a question that the investigator never intended to ask.
Step-by-Step Statistics Solutions
Three original questions connect calculations, definitions and interpretation. Open each solution to follow the reasoning.
Five observed responses are yes, no, yes, yes, and no. Code yes as one and no as zero and calculate the mean. Then consider coding yes as seven and no as two. Calculate that code mean and explain which quantity directly reports the yes proportion.
The zero-one entries are 1, 0, 1, 1, and 0. Their sum counts yes responses exactly, so the mean is three fifths. A binary indicator is a deliberate coding with a useful proportional interpretation.
The alternate numbers are 7, 2, 7, 7, and 2, totaling twenty-five. Their code mean is five. The transformation is two plus five times the indicator, which explains the relationship between the two averages.
The proportion yes is sixty percent. The number five is an average of chosen labels, not a sixty-percent proportion or a new response category. If the coding rule is known, the indicator proportion can be recovered, but arbitrary nominal codes should not be interpreted as measured magnitudes.
There are 100 eligible records, of which 80 have an observed binary response. Twenty-four of the observed responses are positive. Calculate the observed-only positive proportion and the lowest and highest possible complete-data proportions without assumptions about the twenty missing responses.
Among observed responses, the positive proportion is twenty-four divided by eighty, or thirty percent. This is not automatically the proportion among all one hundred eligible records.
If all twenty missing responses are negative, there are twenty-four positives among one hundred. If all are positive, there are forty-four. No narrower complete-data interval follows from these counts alone.
The complete-data proportion is bounded from twenty-four to forty-four percent, while the observed subset gives thirty percent. Additional assumptions about missingness could justify a narrower estimate, but they must be named and supported. Treating missing as negative without justification chooses the lower endpoint rather than discovering the answer.
A waiting-time export contains 0, 15, 20, 25, and 999. The dictionary states that 999 denotes missing and zero is a genuine immediate-service wait. Find the raw code mean and the correct observed-data mean. Explain the effect of also discarding the genuine zero.
Adding all five exported numbers gives 1,059 and a code mean of 211.8. This averages a sentinel with measured minutes and has no valid waiting-time interpretation.
Exclude the single missing sentinel, retain the genuine zero, and average the four observed waits. Their total is sixty, giving fifteen minutes. Report four observed and one missing record.
Removing zero as well leaves three positive waits with mean twenty. That describes a selected positive-wait subset and incorrectly changes the intended observed-data target if zero is eligible. The dictionary and eligibility rule determine treatment, not a preference for a more ordinary-looking number.