SOC 220 - October 6
Watch on YouTube →
Overview
Jeff Brassard explains probability sampling as a way to select a sample that represents a larger population, defining the sampling frame, sample size, margin of error, confidence level, and sources of bias. He compares simple random, systematic, stratified, and multistage cluster sampling, showing how each method works and when limitations such as nonresponse, periodicity, or uneven subgroups can distort results.
Key takeaways
- A probability sample gives each unit in its sampling frame a known chance of selection, but representativeness still depends on a sound frame and adequate participation.
- For populations larger than roughly 20,000, sample size changes relatively little; the U.S. General Social Survey commonly uses about 500–3,000 respondents.
- A 56% estimate with a 3% margin of error should be read as an approximate 53–59% range, commonly reported with a 95% confidence level.
- Nonresponse can undermine a carefully selected sample: Pew Research’s American Religious Landscape study had a 9% response rate, raising questions about how respondents differ from nonrespondents.
- Systematic sampling is vulnerable to periodicity: selecting every 30th case from an alternating voter/nonvoter list can yield only one group.
- Stratified sampling preserves subgroup proportions, while multistage cluster sampling builds a practical sample across dispersed regions through successive selections of regions, institutions, and individuals.
Chapters
0:00
Course Notices and the October 15 Midterm Location
- The third short assignment on surveys is due Sunday, and feedback for the second assignment has been returned.
- The October 15 midterm will take place in Education South, room 129, rather than the regular classroom.
- Jeff Brassard plans to send several reminders about the exam location.
2:06
Why Social Research Uses Probability and Nonprobability Samples
- Researchers usually study a subset of a population because collecting data from every member is too costly or impractical.
- Probability sampling uses random selection and is commonly associated with quantitative research.
- Nonprobability sampling is more common in qualitative research, but quota sampling also appears in quantitative studies.
6:11
Defining Units, Elements, and Populations
- An element or unit is one case in a study; in social research it is often a person, but it can also be a school, company, city, or nation.
- The population is the full group about which researchers seek knowledge or intend to draw conclusions.
- A survey about University of Alberta students has all U of A students as its population.
8:16
Sampling Frames, Samples, and Representativeness
- A sampling frame is the list of units eligible for selection; a set of addresses in Edmonton neighborhoods could form one.
- The sample is the subset actually selected for research, often denoted by lowercase n, while the population is commonly denoted by uppercase N.
- A representative sample approximates population characteristics such as age, race, gender, and socioeconomic class.
- Older telephone sampling could use local prefixes such as Edmonton’s 487 prefix to identify addresses in Alder Grove; mobile numbers no longer map reliably to neighborhoods.
14:24
Sample Size, Margin of Error, and Confidence Levels
- Sample-size calculators, including SurveyMonkey’s, can estimate the number of responses needed; the assignment requires a sample-size decision.
- For populations above roughly 20,000, required sample size changes relatively little; the U.S. General Social Survey typically samples about 500–3,000 people.
- A margin of error of 1–5% trades precision against cost: reducing the margin generally requires a larger sample.
- A 56% poll result with a 3% margin of error corresponds to an approximate range of 53–59%; confidence levels are commonly set at 95–99%.
- For a population of 10,000, a 5% margin of error and 95% confidence level require about 370 responses; the formula is not required on the midterm.
23:19
Probability Selection, Sampling Error, and Nonresponse
- In a probability sample, every unit in the sampling frame has a known chance of selection, although those chances need not be equal.
- Convenience sampling is nonprobability sampling because researchers recruit whoever is available and willing to participate.
- Sampling error is the difference between sample characteristics and population characteristics; even well-designed probability samples cannot eliminate it entirely.
- Selecting 100 Mill Woods households and receiving 60 replies produces 40 nonresponses, a 40% nonresponse rate.
- Pew Research’s American Religious Landscape study had a reported 9% response rate, illustrating how nonresponse has become a major challenge.
30:33
Censuses and Three Sources of Sampling Bias
- A census collects data from every element in a population; governments can make census participation compulsory, but censuses are rare in other research settings.
- Nonrandom selection, a poorly constructed sampling frame, and nonresponse are three principal sources of sampling bias.
- Choosing only higher-income Edmonton neighborhoods would overrepresent middle- and upper-income residents.
- People who answer a survey may differ systematically from those who do not, so a 9% respondent group may not reflect the other 91%.
35:15
Visualizing Sampling Error with Reality-TV Viewers
- A study of reality-TV viewing and romantic expectations would ideally include comparable numbers of viewers and nonviewers for a clear comparison.
- A small imbalance—such as one extra nonviewer—creates less sampling error than a large imbalance between the groups.
- Statistical reweighting can sometimes correct imbalances, but it typically requires a sample of at least about 1,000 cases.
- Severe imbalance can make results too biased to use effectively, so the main goal is to minimize sampling error.
38:02
Simple Random Sampling and Random Number Generators
- Simple random sampling gives every unit in the sampling frame exactly the same probability of selection.
- Researchers number every unit, choose a target sample size, and use a random-number table or computer program to select cases.
- In an example with 80,000 Mill Woods homes and a target sample of 400, a generator selects 400 numbered addresses.
- The method can produce a strong representative sample with adequate size, but it depends on having a complete, usable sampling frame.
- Modern mobility, unlisted numbers, and low response to phone or mail surveys make the method harder to carry out.
43:38
Sampling Ratios and Systematic Sampling Intervals
- The sampling ratio is sample size divided by population size; 1,000 cases from 100,000 people equals 1%.
- Systematic sampling selects units at a fixed interval rather than drawing a separate random number for every case.
- With an interval of 30, researchers choose a random start from 1 to 30 and then select every 30th unit.
- For a small frame with an interval of three and a start at unit three, the selected cases are 3, 6, and 9.
47:14
Periodicity: When an Ordered Frame Skews Systematic Samples
- Periodicity occurs when the sampling frame has a repeating pattern that aligns with the systematic interval.
- If a list alternates voters and nonvoters, selecting every 30th case from an even-numbered start could select only nonvoters.
- An alternating male–female list could similarly produce a sample dominated by one gender.
- Randomizing the frame’s order before selection can reduce this risk; assignment methods must also be feasible given the available frame.
51:12
Stratified Random Sampling for Proportional Subgroups
- Stratified random sampling divides a population into subgroups and samples from each so the final sample reflects their population proportions.
- For a university with 9,000 students across five faculties, a stratified sample of 450 can be allocated to match each faculty’s share.
- A simple random sample of the same size may overrepresent natural sciences or social sciences and undersample humanities, business, or engineering.
- This approach is useful when subgroups—such as faculties—need reliable representation in the sample.
55:24
Multistage Cluster Sampling Across Canadian Universities
- Multistage cluster sampling suits large, geographically dispersed populations when a complete list of individuals is unavailable.
- To study Canadian university students, researchers could first cluster universities by region: British Columbia, the Prairies, Ontario, Quebec, and the Maritimes.
- They could then select subregions, such as the Lower Mainland in B.C. or Alberta within the Prairies, followed by particular universities.
- Additional stages could select faculties and then students, allowing broad geographic coverage without first assembling one national student list.
1:04:00
Balancing Cluster Sizes, Complexity, and Sampling Proportions
- Cluster sizes differ: the University of Toronto has many more students than smaller universities such as Simon Fraser or the University of Victoria.
- Researchers must set appropriate proportions at each stage so that large provinces such as Ontario are not represented like much smaller populations.
- Every added cluster or subgroup increases the work needed to estimate proportions and creates more opportunities to over- or undersample a region.
- Jeff Brassard ends the session after 21 slides, planning to finish probability sampling on Thursday and offer exam questions-and-answers time before the October 15 midterm.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Jeff Brassard.