STAT 638 (Fall 2026), Lecture 19
Watch on YouTube →
Overview
Samiran Sinha reviews STAT 638 project grading and hierarchical-model exercises, then develops a Bayesian model for comparing 30-day heart-attack mortality across five Texas hospitals with unequal sample sizes. He builds a binomial-beta hierarchy with Gamma priors on its hyperparameters, explains Gibbs sampling and Metropolis-Hastings updates, and shows how partial pooling shrinks Hospital A’s raw 20% mortality estimate to 14.5%.
Key takeaways
- In the Normal hierarchical model, conditioning on a group effect gives observation variance σ², while integrating over the random group effect gives marginal variance τ² + σ².
- For the hospital example, Yᵢ ~ Binomial(nᵢ, pᵢ) and pᵢ ~ Beta(α, β) let hospitals with small admission counts borrow information from hospitals with larger counts.
- The full conditional for each hospital probability is Beta(Yᵢ + α, nᵢ − Yᵢ + β), but α and β do not have standard full conditional distributions.
- When deriving α and β conditionals, the Beta density’s Gamma-function normalizing constants must be retained because they depend on the unknown hyperparameters.
- Random-walk Metropolis-Hastings proposes α and β on the log scale; the example’s acceptance rates were 0.4029 and 0.4172, respectively.
- Hierarchical shrinkage moved Hospital A’s estimated 30-day mortality from a raw 20% to a posterior estimate of 14.5%, while its 95% credible interval remained wider than Hospital B’s because A had far fewer patients.
Chapters
- Sinha plans to grade 17 projects and add peer assessments to reduce the influence of any one grader on students’ grades.
- Each written report is to be randomly assigned to 2–3 students, with safeguards against assigning anyone their own project.
- Students assess the written report only; Sinha evaluates the oral presentation and question-and-answer portion.
- Written projects are due by November 18 and will be placed in a shared folder.
- The model gives group parameters θ₁,…,θₘ an IID Normal(μ, τ²) distribution and observations within group j a Normal(θⱼ, σ²) distribution.
- Conditional on θⱼ, an observation has variance σ²; accounting for uncertainty in θⱼ raises the marginal variance to τ² + σ².
- Writing Yᵢⱼ = θⱼ + εᵢⱼ separates group-level variability from observation-level residual variation.
- The residual εᵢⱼ is independent of θⱼ, so their covariance is zero.
- Two observations from the same group are independent conditional on its θ parameter, so their conditional covariance is zero.
- Without conditioning on θ, observations in a shared group can covary because they share the same random group effect.
- The exercise asks students to calculate six variance and covariance quantities and compare them with their earlier answers.
- Bayes’ rule shows that μ given the group parameters θ₁,…,θₘ does not directly depend on the observations Y₁,…,Yₘ; Sinha recommends practicing Exercises 8.2 and 8.3 for the quiz and exam.
- Chapter 8 introduced hierarchical models in a Normal linear-model setting; Sinha now extends the same information-sharing idea to a Bernoulli/binomial model.
- The example compares 30-day mortality after heart attacks at five Texas hospitals; the dataset is hypothetical and was generated with ChatGPT or Claude.
- Hospital admission counts differ sharply, from 20 patients at one hospital to 500 at another, making raw mortality rates unequally precise.
- Observed mortality rates range from 9% to 20%; Hospital A’s 20% rate comes from only 20 patients.
- For hospital i, deaths Yᵢ follow Binomial(nᵢ, pᵢ), where nᵢ is admissions and pᵢ is the 30-day mortality probability.
- A Beta(α, β) distribution models the hospital-specific probabilities pᵢ, allowing mortality information to be shared across hospitals.
- Sinha places Gamma priors on α and β rather than fixing them; fixed Beta shape parameters would prevent the intended hierarchical learning.
- The model has seven unknowns: five hospital probabilities plus α and β; the slide’s Beta distribution parameter label should read α and β.
- The full conditional for each pᵢ is Beta(Yᵢ + α, nᵢ − Yᵢ + β), obtained by combining the binomial likelihood with the Beta prior.
- The conditional densities for α and β include the product of the five Beta densities and their Gamma priors.
- Because the Beta normalizing constants depend on α and β through Gamma functions, those constants cannot be dropped when deriving their conditionals.
- The α and β full conditionals are not standard distributions, so direct Gibbs draws are unavailable.
- Sinha describes the sampler as a Markov chain Monte Carlo procedure that repeatedly updates p₁,…,p₅, α, and β using their latest available values.
- The five hospital probabilities can be drawn from Beta full conditionals; α and β require Metropolis-Hastings proposals.
- Parameter-update order is flexible, and the log full-conditional densities are used to evaluate proposed hyperparameter values.
- The lecture previews random-walk Metropolis-Hastings, with fuller treatment deferred to Chapter 10.
- The example uses hospital death counts and admission totals; individual patient records are unnecessary because the death count is sufficient for the binomial model.
- Gamma-prior hyperparameters are set illustratively to 2 and 0.1, without a substantive justification; changing them across reasonable values is a sensitivity analysis.
- The code runs 50,000 MCMC iterations and stores the five p values in a iterations-by-hospitals matrix, alongside α and β.
- Initial p values use (Y + 1)/(n + 2), a smoothed estimate close to the raw sample proportion; Sinha also notes that Y/n could be used.
- Each hyperparameter proposal is generated on the log scale from a Normal random walk centered at its previous log value, then transformed back.
- The acceptance calculation combines the difference in log conditional densities with the proposal-density ratio.
- A Uniform(0,1) draw determines whether to accept the proposal; rejection leaves the parameter at its previous value.
- The proposal standard deviations are tuned by trial and error; the example uses 0.5 for both α and β.
- The reported acceptance rates are 0.4029 for α and 0.4172 for β; Sinha uses acceptance rates to assess and tune proposal scales.
- Posterior summaries include medians and the 2.5th and 97.5th percentiles as 95% credible intervals; the lecture notes that medians can be useful for skewed posteriors.
- Hospital A’s estimated mortality is 0.145 with a 0.07–0.262 credible interval, down from its raw rate of 20% through hierarchical shrinkage.
- Hospital B’s estimate is 0.102 with a 0.078–0.129 interval; the estimates reflect both cross-hospital pooling and differing sample sizes.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Samiran Sinha.