ECON E371 Class Recording 9/8
Watch on YouTube →
Overview
Professor Lantis teaches hypothesis testing and confidence intervals, connecting test statistics, p-values, critical values, and sample size to decisions about population means. The class then applies those ideas in Stata to county unemployment and 2020 Trump vote-share data, while reviewing the Friday problem set and quiz requirements and previewing linear regression.
Key takeaways
- For samples larger than 30, the course approximates the sampling distribution with the standard normal; a 95% two-sided test uses cutoffs of approximately ±1.96.
- The p-value and critical-value approaches are equivalent: reject the null when the p-value is below alpha or the test statistic lies beyond the critical cutoff.
- A 95% confidence interval and a two-tailed test at alpha 0.05 produce the same decision about a hypothesized population mean: reject when the value falls outside the interval.
- Higher confidence widens a confidence interval and makes rejection harder, while a larger sample reduces the standard error, narrows the interval, and can make rejection easier.
- In the Stata exercise, counties classified as having low unemployment showed higher average 2020 Trump vote share than the comparison group, motivating a formal difference-in-means test rather than establishing causation.
- The first problem set and first quiz are due Friday; the quiz uses LockDown Browser and covers material through confidence intervals.
Chapters
0:00
Course Files, Stata Preparation, and Friday Deadlines
- Professor Lantis says updated Zoom recordings are being moved to Kaltura, accessible through Canvas, while he troubleshoots compatibility issues.
- The statistics-review do-file repeats the prior code; students should practice opening data, saving do-files, and working in Stata before the first problem set.
- The first problem set and quiz are due by end of day Friday; the quiz covers material through confidence intervals and requires LockDown Browser.
- Office hours are shifted to 10 a.m.–noon the next day, with virtual office hours Friday morning for last-minute questions.
3:30
Null and Alternative Hypotheses for Population Means
- A hypothesis test asks whether a sample mean is plausible if the population mean equals an assumed value, such as 100.
- For a one-sided test of whether the mean is greater than a value, the null represents the opposite direction; two-sided tests ask whether the mean differs in either direction.
- The sample mean varies across samples according to sample size and the underlying data’s variation.
7:30
How Sample Size Changes the Test Statistic
- The test statistic standardizes the difference between the observed sample mean and the hypothesized population mean by the standard deviation of sample means.
- The standard deviation of sample means shrinks as sample size increases, so the same difference from the hypothesized mean produces a larger-magnitude test statistic.
- A larger test statistic means the observed sample mean is farther from the null value in standard-deviation units, which can count as stronger evidence against the null.
12:00
Critical Values and the Standard Normal Approximation
- A critical value marks a cutoff: test statistics farther into the relevant tail fall in the rejection region.
- For the course’s samples larger than 30, Professor Lantis uses standard-normal critical values as an approximation to the Student t distribution.
- A right-tail critical value of about 1.65 leaves 0.05 probability in the upper tail; left-tail cutoffs have the corresponding negative sign.
17:24
P-Values, Alpha, and the Rejection Decision
- A p-value is the probability, assuming the null is true, of observing the test statistic found or one more extreme in the relevant direction.
- Alpha is the chosen rejection threshold; common levels are 0.10, 0.05, and 0.01, corresponding to 90%, 95%, and 99% confidence.
- The critical-value and p-value methods give the same decision: reject when the test statistic enters the rejection region or when p-value is below alpha.
21:42
Four Steps for a Hypothesis Test
- Specify the null and alternative hypotheses, choosing a left-tail, right-tail, or two-tail test to match the question.
- Choose alpha, calculate the test statistic from the sample, and compare it with the critical value.
- Alternatively, compare the p-value with alpha; a p-value below alpha means the result is sufficiently unlikely under the null to reject it.
23:23
Two-Tailed Tests Split Alpha Across Both Ends
- A two-tailed test treats equally distant results above and below the null mean as evidence against the null—for example, sample means of 110 and 90 when testing a mean of 100.
- The total alpha is divided between the two tails, so each tail contains alpha divided by two and a one-tail p-value must be doubled.
- For the same overall alpha, two-tailed critical values are farther from zero than one-tailed values because each tail must contain less probability.
28:14
Type I and Type II Errors in Hypothesis Testing
- A Type I error is rejecting a true null hypothesis; its probability is alpha, the significance level chosen before the test.
- A Type II error is failing to reject a false null—for example, treating a sample mean near 101 as support for a null of 100 when the true mean is 105.
- Professor Lantis says the course will emphasize Type I errors and alpha rather than calculating Type II error probabilities.
31:04
Building a 95% Confidence Interval Around a Sample Mean
- A 95% confidence interval leaves 5% outside its bounds, split into 2.5% in each tail for a two-sided interval.
- The standard-normal cutoff is approximately 1.96, so the interval extends 1.96 standard errors below and above the sample mean.
- The standard error of the sample mean is based on the sample variance divided by sample size, with the square root taken to obtain the standard deviation.
36:55
Why Confidence Intervals Match Two-Tailed Tests
- The interval is centered on the observed sample mean because the unknown population mean cannot serve as its center.
- For a 95% interval, a hypothesized population mean outside the bounds is rejected by the corresponding two-tailed test at alpha 0.05.
- If the hypothesized mean lies inside the interval, the test fails to reject it; the sample mean itself is always at the center of its interval.
44:27
Margin of Error, Confidence Level, and Sample Size
- The margin of error is the amount added to and subtracted from the sample mean to form the interval’s upper and lower bounds.
- Increasing confidence from 95% to 99% requires a more extreme critical value, widening the interval and making rejection harder.
- Increasing sample size lowers the standard error and narrows the interval at a fixed confidence level, making it easier to rule out values outside the narrower range.
47:46
Extending the Test to Differences Between Two Means
- The same hypothesis-testing and confidence-interval procedures apply when the statistic is the difference between two sample means.
- The variance of the difference is formed from the variances of the two sample statistics, so its test statistic uses the standard error of that difference.
- After calculating the appropriate statistic, the critical-value comparison or p-value-versus-alpha decision works as before.
49:19
Practice Quiz: Confidence Levels, P-Values, and Intervals
- Students open the in-class quiz using access code 52.16 and practice interpreting a two-tailed test of whether a population mean equals 50.
- Rejecting at 95% implies rejection at 90%, but does not establish rejection at 99%; the 99% critical value is more demanding.
- A two-tailed rejection does not reveal whether a one-tailed test would reject unless the direction of the sample result is known.
- For a p-value of 0.075, reject at the 90% level but fail to reject at 95% and 99%; the corresponding higher-confidence intervals include the null value.
1:00:45
Importing Election Data and Comparing Unemployment Groups in Stata
- For an Excel file, use Stata’s File > Import rather than File > Open, select the option treating the first row as variable names, and set the global path to the folder containing the data.
- The election dataset includes county unemployment and 2020 Trump vote share; Professor Lantis demonstrates renaming variables, plotting unemployment with `hist`, and generating a high-unemployment indicator equal to 1 above the mean.
- A second low-unemployment indicator illustrates why comparison operators matter: values exactly equal to the mean can be coded as 0 in both groups if the inequalities are not complementary.
- Conditional histograms and `bysort` summaries show that counties with low unemployment had a higher average Trump vote share in this initial comparison; the analysis is descriptive and leads into testing a difference in means.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Professor Lantis.