STAT 638 (Fall 2026), Lecture 4
Watch on YouTube →
Overview
Samiran Sinha reviews STAT 638 course logistics and develops Bayesian point estimation, prior sensitivity, posterior prediction, and Jeffreys priors, using Beta–Bernoulli and clinical-trial examples. He then introduces a Poisson model for blemish counts, calculating a probability under a rate of 0.8 blemishes per square centimeter and setting up likelihood-based inference for an unknown rate.
Key takeaways
- Under squared-error loss, the Bayes estimate is the posterior mean; under absolute-error loss it is the posterior median, and under 0–1 loss it is the posterior mode.
- For a Beta(A, B) prior and Y successes in N Bernoulli trials, the posterior mean combines the prior mean and sample proportion, with the data's influence increasing as N grows.
- In the HRT example, 44 side effects among 100 participants produce a sample proportion of 0.44, but a concentrated Beta(40, 20) prior raises the posterior mean to about 0.525.
- A posterior predictive distribution averages future-outcome probabilities over the parameter's posterior; for a Bernoulli outcome, its success probability equals the posterior mean of θ.
- A uniform prior on θ does not remain uniform for odds ψ = θ/(1 − θ); Jeffreys' rule instead gives the binomial success probability a Beta(½, ½) prior.
- For a Poisson rate of 0.8 blemishes per square centimeter, an 11-square-centimeter area has expected count 8.8 and a 0.8716 probability of at least six blemishes.
Chapters
0:00
STAT 638 Schedule Correction and Course Updates
- Samiran Sinha corrects the course schedule: December 2 will be the last class, rather than December 3; the syllabus change is awaiting approval.
- The revised final-class date also affects the project presentation schedule.
- Sinha notes that the course materials for Chapters 5, 6, and 7 will be added later.
1:45
Project Groups, Self-Study, and Reproducible Materials
- Project descriptions and group assignments are posted in the course modules; most groups have three students, and one has two.
- Some groups span both course sections because each section has an odd number of students; students should contact group members and begin work promptly.
- The graduate course emphasizes independent learning, with few open-book quizzes and example code designed to be reproducible; questions can go on the discussion board.
6:33
Bayes Estimates as Decisions Under Posterior Loss
- A Bayes point estimate minimizes posterior expected loss, which averages a decision's loss over the posterior distribution of the parameter.
- Squared-error loss selects the posterior mean; absolute-error loss selects the posterior median; 0–1 loss selects the posterior mode.
- Sinha recommends the posterior mean for roughly symmetric distributions and the median when the posterior is skewed.
10:10
Beta–Bernoulli Posterior Mean and Prior-Data Weighting
- With a Beta(A, B) prior and Y successes in N Bernoulli trials, the posterior is Beta(A + Y, B + N − Y).
- The posterior mean, (A + Y)/(A + B + N), is a weighted average of the prior mean A/(A + B) and the sample proportion Y/N, which is also the maximum-likelihood estimate.
- As N grows, the data weight increases and the prior's influence diminishes; with small samples, prior parameters can materially affect the estimate.
14:29
HRT Side-Effects Example Shows Prior Sensitivity
- In a hormone replacement therapy study, 44 of 100 participants experienced side effects, giving a sample proportion of 0.44.
- With Beta(A, B) prior, the posterior mean is (44 + A)/(100 + A + B); Beta(1, 1) yields about 0.441, while Beta(1, 2) yields about 0.436.
- More concentrated priors can shift the estimate further: Beta(20, 40) gives about 0.425, while Beta(40, 20) gives about 0.525.
20:30
Posterior Predictive Distributions for Future Outcomes
- The posterior predictive distribution for a future observation integrates its likelihood over the parameter's posterior distribution, incorporating observed data and parameter uncertainty.
- For a future Bernoulli outcome, the predictive probability of success is the posterior mean of the success probability; the probability of failure is one minus that mean.
- In the HRT example with a Beta(1, 1) prior, the predictive success probability is about 0.441; a plug-in prediction using the posterior mean matches here, but need not generally.
27:30
Why Uniform Priors Depend on Parameterization
- A uniform prior on a success probability θ does not remain uniform after transforming to the odds ψ = θ/(1 − θ).
- Using the change-of-variables rule, the induced odds density is proportional to 1/(1 + ψ)², including the transformation's Jacobian.
- This example shows why a prior that appears non-informative on one parameter scale may not appear so on another.
33:55
Jeffreys Prior for the Binomial Success Probability
- Jeffreys prior is proportional to the square root of the Fisher information and is designed to be invariant under smooth parameter transformations.
- For a Binomial(N, θ) model, the Fisher information is N/[θ(1 − θ)]; ignoring the parameter-independent factor N gives a prior proportional to [θ(1 − θ)]⁻¹ᐟ².
- The resulting Jeffreys prior is Beta(½, ½), a proper distribution that places more density near 0 and 1 than the uniform Beta(1, 1) prior.
43:30
Proper and Improper Priors: Posterior Validity Matters
- A prior is proper if its density integrates to a finite value; Beta(½, ½) is proper because both shape parameters are positive.
- Jeffreys priors are not always proper, and Bayesian analyses may use an improper prior if the resulting posterior is proper.
- A proper posterior is essential for defining summaries such as means, medians, and quantiles.
48:18
Poisson Blemish Counts and the Likelihood for an Unknown Rate
- Sinha introduces the Poisson distribution for count data, with rate parameter θ > 0 and probability mass function e⁻ᶿθʸ/y!.
- For a quality target of at most 0.8 blemishes per square centimeter, an 11-square-centimeter sample has expected count 8.8; under that Poisson model, the probability of at least six blemishes is 0.8716.
- The lecture begins the inference setup for n independent, identically distributed Poisson observations, whose joint likelihood is formed by multiplying their individual probability mass functions.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Samiran Sinha.