ECON 371 Fall 2026 Recorded Lecture 9/1
Watch on YouTube →
Overview
Professor Lantis connects covariance, correlation, conditional distributions, and the normal distribution to show how economists describe relationships and compare data. The lecture also walks through Stata setup and file management, then demonstrates using `egen` and `gen` to calculate summary statistics and standardize a variable into z-scores; students are asked to prepare for a quiz and begin the first problem set.
Key takeaways
- Covariance summarizes whether two variables’ deviations from their means tend to have the same or opposite signs; dividing by both standard deviations makes the resulting correlation unitless.
- For grouped data, joint-category probabilities weight each product of deviations, so categories with more observations contribute more to covariance.
- A normal distribution’s mean determines its center and its variance determines its spread; approximately 68%, 95%, and 99% of values fall within one, two, and three standard deviations of the mean.
- Standardizing an observation with (value − mean) / standard deviation expresses its distance from the mean in standard-deviation units and yields a variable with mean near 0 and variance near 1.
- In Stata, students should save executable commands in a do-file, use their own folder path, and distinguish `egen` built-in summary functions from `gen` arithmetic expressions.
- A positive association between medical-marijuana legalization and above-cutoff violent crime in the lecture’s one-year state example does not demonstrate causation.
Chapters
0:00
Getting Stata Running on Student Devices
- Stata can run through IU Anywhere, but that option depends on Wi-Fi; downloading the program through IU’s technology website is an alternative.
- Students installing Stata locally should extract the downloaded archive and sign in with IU credentials.
- For saved work, Professor Lantis recommends using IU OneDrive rather than a personal OneDrive account and opening Stata before class.
4:07
Office-Hour Troubleshooting, Course Files, and Quiz Schedule
- Students still unable to open Stata are directed to Professor Lantis’s Wednesday 8:00–11:00 office hours or Zoom office hours for screen-sharing help.
- Week 1 and Week 2 files are on Canvas and should be downloaded before class for use in Stata exercises.
- Practice problems cover statistics from the prior week and current week; Quiz 1 is planned after Thursday’s class and due Monday the 7th.
8:04
Covariance Measures How Two Variables Move Together
- For each observation, covariance multiplies the deviation of class size from its mean by the deviation of math score from its mean.
- A negative product, such as below-average class size paired with above-average scores, is evidence that the variables move in opposite directions.
- Covariance has the combined units of both variables, so Professor Lantis introduces correlation by dividing covariance by both variables’ standard deviations.
13:32
Calculating Covariance from Five Paired Observations
- The classroom example uses X values 3, 6, 3, 4, 8 and Y values 4, 2, 7, 8, 6.
- The means are 4.8 for X and 5.4 for Y; each covariance term is the product of an observation’s deviations from those means.
- For this observation-by-observation example, the five products are added and divided by the number of observations.
16:43
Using Joint Probabilities for Grouped-Data Covariance
- A state-level example groups observations by medical-marijuana legalization and whether violent crime is above or below a cutoff.
- Each cell’s probability represents the share of states in that category; for example, the chance of selecting a state with legalization and above-cutoff crime is about 0.22.
- Grouped covariance weights each pair of deviations by the probability of that category, rather than treating the four categories as equally likely.
22:41
Converting Covariance into Correlation
- Correlation is calculated by dividing covariance by the standard deviation of each variable, which removes their measurement units.
- A positive correlation in the state example means legalized states are more likely to fall above the selected violent-crime cutoff in that dataset.
- Professor Lantis cautions that the example uses one year of data and does not establish that legalization causes crime to rise.
26:33
Conditional Means and Independence
- The expected math score can be calculated for the full population or conditional on education, such as having a college degree versus a high-school degree.
- If the conditional means are the same across groups, the other variable does not change the mean in this example.
- Independence implies zero covariance; in real data, a covariance close to zero is more plausible than an exactly zero value.
31:06
How Conditioning Can Change Salary Distributions
- Professor Lantis compares starting-salary distributions for students who take ECON 371 and those who do not.
- The groups may differ in both their average salaries and the shapes of their distributions, not just in one summary statistic.
- A long right tail from a few very high salaries can pull the mean above the median, indicating positive skew.
35:31
Normal Distributions: Symmetry, Mean, and Variance
- A normal distribution is symmetric, with its mean, median, and mode at the same point.
- Changing the mean shifts the distribution along the horizontal axis; it does not by itself change the distribution’s spread.
- Higher variance spreads probability farther from the mean and flattens the curve, while lower variance concentrates probability near the mean.
39:25
Comparing Height Means and Variances by Sex
- The height example places the male mean around 69–70 inches and the female mean around 64 inches.
- A wider, flatter distribution indicates greater variance; the displayed male height distribution is more spread out.
- A taller peak indicates more probability concentrated near the mean, which corresponds to lower variance.
42:46
Normal-Curve Areas and the 68–95–99 Rule
- The total area under a probability distribution is 1, representing the probabilities of all possible values.
- For a normal distribution, about 68% of observations lie within one standard deviation of the mean and about 95% within two.
- About 99% lie within three standard deviations; observations farther away are rare and may be treated as outliers.
47:28
Standardizing Observations into Z-Scores
- A z-score is calculated for each observation as the original value minus the mean, divided by the standard deviation.
- Standardization expresses a value in standard-deviation units and produces a variable centered at 0 with standard deviation 1.
- Professor Lantis introduces the formula before shifting to its implementation in Stata.
49:42
Stata Windows, Do-Files, and Data Paths
- The three useful Stata windows are the main program, the do-file editor for saved code, and the data browser for inspecting observations.
- Students should replace Professor Lantis’s OneDrive folder path with their own folder path, while keeping Canvas-provided filenames unchanged.
- Importing a CSV through Stata’s interface can reveal the required folder path and generate the corresponding import command.
55:02
Importing CSV Data and Saving Reusable Stata Code
- The state cigarette-tax example starts as a CSV; Stata’s Import Delimited workflow loads it, and File > Save can store it as a DTA dataset.
- The Results window shows commands and output but is not where students save their work; the do-file is the record of code to submit.
- Students should start the first problem set early, confirm they can open its datasets, and save their do-file; the assignment is expected after class and likely due the following Friday.
1:05:02
Using egen, gen, and Summary Statistics to Standardize Data
- `egen` with built-in functions creates variables for statistics such as the mean and standard deviation of packs per capita; `sum, detail` also reports summary statistics.
- The example creates a standardized packs-per-capita variable by subtracting its mean and dividing by its standard deviation.
- The resulting standardized variable should have a mean near 0 and a variance near 1; Professor Lantis checks those values with Stata’s summary output.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Professor Lantis.