ECON 371 Fall 2026 Class Recording 9/3
Watch on YouTube →
Overview
Professor Lantis reviews Stata-based standardization and sampling distributions, then develops hypothesis testing using test statistics, critical values, and p-values. Simulations of NBA player heights show how sample-mean variance falls as sample size increases and why samples of roughly 30 or more support a normal approximation; an IQ example then demonstrates rejection decisions at different confidence levels.
Key takeaways
- Standardizing a variable with (value − mean) / standard deviation gives it a mean near zero and standard deviation near one, but does not by itself make the distribution normal.
- The variance of the sampling distribution of a sample mean is the population variance divided by sample size; with NBA height variance 12.6, a sample size of 10 implies a mean variance of about 1.26.
- The Central Limit Theorem explains why sample means become approximately normal for sufficiently large samples, commonly around n = 30 or more, even when the original data are not normal.
- A two-tailed test at 95% confidence uses alpha = 0.05 and critical values of approximately ±1.96; reject the null when the absolute test statistic exceeds the critical value.
- A p-value is calculated under the assumption that the null hypothesis is true and measures the probability of observing a result at least as extreme as the sample result; it is not the probability that the null itself is true.
- For the IQ example, a test statistic of approximately -1.77 permits rejection of a hypothesized mean of 105 at the 10% significance level, but not at the 5% or 1% level.
Chapters
0:00
Course Logistics and a Correction to the Grouped Covariance Example
- Professor Lantis cautions that missing class will make it difficult to keep up as ECON 371 moves beyond statistics review.
- He corrects the violent-crime grouped-data example: the two values should be 400 and 300.
- The first quiz is postponed until after the following Tuesday class and will be due Friday at the end of the day.
3:00
Using Stata to Plot Data and Standardize Variables
- Stata do-files use a file path to locate data; keeping course files together avoids repeatedly changing that path.
- A scatter plot of state cigarette tax and cigarette cost per pack shows a positive relationship.
- A z-score is calculated as a value minus its mean, divided by its standard deviation; Stata can create it manually or with the built-in `egen` standardization function.
10:00
Checking Z-Scores and Reading Standard Normal Probabilities
- Summarizing a standardized variable should produce a mean near zero and a standard deviation near one.
- Standardizing changes a variable's scale, but does not necessarily make its distribution normal.
- A z-score of 1 means an observation is one standard deviation above the mean; standard normal tables give probabilities for z-scores.
15:00
Why Small Samples Use Student’s t Distributions
- With small samples, sample statistics vary more, so Student’s t distributions have heavier tails than the standard normal distribution.
- The t distribution is symmetric, and its shape depends on sample size; as sample size grows, it becomes closer to the standard normal.
- The sample mean is calculated by summing sample observations and dividing by sample size; Professor Lantis distinguishes the sample mean notation, y-bar, from the population mean, mu.
19:00
How Sample Size Changes the Distribution of Sample Means
- Across repeated random samples, the expected value of the sample mean equals the population mean.
- The variance of the sampling distribution of the mean is the population variance divided by sample size.
- Larger samples reduce the spread of possible sample means, making an individual sample mean more informative about the population mean.
25:00
Simulating 1,000 Samples of NBA Player Heights in Stata
- Professor Lantis uses NBA player heights as population data, with a mean near 79.1 inches and variance near 12.6.
- A Stata program repeatedly samples players, calculates each sample mean, and stores 1,000 means for sample sizes of 10, 30, and 100.
- The simulated means stay near 79.1, while their variance decreases as sample size rises; for n = 10, the theoretical variance is about 12.6/10 = 1.26.
30:00
The Central Limit Theorem and Small-Sample Uncertainty
- A histogram of the simulated sample means is approximately bell-shaped, illustrating the Central Limit Theorem.
- For sample sizes around 30 or greater, the distribution of sample means is often approximated by a normal distribution, even if the original data are not normal.
- When population variance is unknown and samples are small, inference relies on Student’s t distribution; smaller samples have heavier tails and make extreme standardized values more likely.
34:00
Null and Alternative Hypotheses for Testing a Population Mean
- A hypothesis test starts with a null hypothesis, H₀, representing the assumed population value, and an alternative hypothesis, H₁ or Hₐ, representing the claim being tested.
- For a two-sided test of whether IU students’ average IQ differs from 100, the null is that the mean equals 100 and the alternative is that it does not.
- One-sided tests instead ask whether the population value differs in a specified direction; the choice depends on the research question.
38:30
Test Statistics Measure Distance from the Assumed Mean
- A test statistic expresses how many standard errors the sample mean lies from the mean assumed by the null hypothesis.
- Its numerator is the sample mean minus the hypothesized mean; its denominator is the standard error of the sample mean.
- When population variance is unknown, the sample variance estimates it, and the standard error is based on the square root of estimated variance divided by sample size.
43:00
Two-Tailed Critical Values and Rejection Regions
- For a two-tailed test, the total significance level alpha is split between the two tails of the distribution.
- At 95% confidence, alpha is 0.05, with 0.025 in each tail and standard normal critical values of approximately -1.96 and 1.96.
- The rejection region consists of test statistics beyond the critical values; for a two-tailed test, compare the absolute test statistic with the positive critical value.
48:00
Using P-Values to Make the Same Decision as Critical Values
- A p-value is the probability, assuming the null hypothesis is true, of obtaining a test statistic at least as extreme as the one observed.
- Reject the null when the p-value is below alpha; this decision matches checking whether the test statistic falls in the critical-value rejection region.
- The common confidence levels of 90%, 95%, and 99% correspond to significance levels alpha = 0.10, 0.05, and 0.01.
55:00
Interpreting Rejection, Failure to Reject, and Significance
- Rejecting a null hypothesis does not establish the sample mean as the true population mean; it rules out the hypothesized value at the selected significance level.
- Failing to reject means the sample did not provide strong enough evidence against the null, not that the null has been proven true.
- Professor Lantis defines significance level as 100% minus confidence level and outlines the test sequence: state hypotheses, choose alpha, calculate the statistic and p-value, then reject or fail to reject.
1:05:00
In-Class IQ Test Example and Quiz Submission
- The class example tests a hypothesized mean IQ of 105 using a sample mean of 102, sample size 50, and stated population variance 144.
- The test statistic is approximately -1.77; it falls beyond the 90% two-tailed critical value of about 1.65, but not beyond the 95% value of 1.96.
- The result supports rejection at the 10% significance level, but not at 5% or 1%; students access the Canvas quiz with the in-class code 2.13.
1:10:00
Upcoming Confidence Intervals, Quiz Deadline, and Course Support
- The next Tuesday class will cover confidence intervals and additional Stata work; confidence intervals are closely related to hypothesis tests.
- Professor Lantis says the quiz will be posted after Tuesday’s class and due Friday at the end of the day; practice problems are already available.
- Students with connectivity or software problems are advised to use eduroam or campus IT support, and Professor Lantis will post updated office-hour information on Canvas.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Professor Lantis.