CSci574 / AI 520 Machine Learning: Classroom Session 10/06/2026
Watch on YouTube →
Overview
Derek Harter reviews CSci574/AI 520 deadlines and project expectations, then explains logistic regression for binary classification: the sigmoid maps linear scores to probabilities, and log loss gives a cost suited to those probabilities. He connects the loss to gradient descent and demonstrates how scikit-learn’s fitted probabilities define decision boundaries on Iris data, including a contour plot for two features.
Key takeaways
- Logistic regression converts a linear score into a probability with the sigmoid function, then typically predicts class 1 above 0.5 and class 0 below 0.5.
- Binary cross-entropy penalizes confident wrong predictions sharply: for a true class of 1 it uses negative log probability, and for class 0 it uses negative log of one minus that probability.
- Logistic regression’s sigmoid makes the optimization nonlinear, so it cannot use the linear-regression normal equation and instead requires gradient descent or another optimizer.
- A scikit-learn classifier’s predict_proba output reveals confidence, while predict applies a decision threshold to return class labels.
- A two-feature decision boundary can be visualized by evaluating predicted probabilities over a mesh grid and drawing the 0.5 contour; overlapping Iris classes can still produce mistakes.
- Lasso’s L1 regularization can set polynomial-regression coefficients to zero, while regularization choices should be documented and checked for underfitting or overfitting.
Chapters
0:00
Assignment 4, Project Selection, and Incremental Work Reminders
- Assignment 4 remains due Sunday at midnight; Derek Harter encouraged students to finish before the half-week break.
- Students who had not accepted the class project or updated its README could still complete those steps that day or the next to receive credit.
- Harter recommended making project progress throughout the term, including regular Git commits, rather than leaving the work until the final week.
4:44
Assignment 4: Underfitting, Overfitting, Ridge, and Lasso
- Several assignment tasks ask students to create underfit and overfit models and use ridge and lasso regression.
- Lasso uses L1 regularization, which can drive unhelpful polynomial coefficients to zero and help assess which polynomial terms matter.
- Students should complete the assignment table and describe how they manually selected regularization parameters to avoid underfitting or overfitting.
7:26
Writing Readable Jupyter Notebooks and Starting the Logistic Regression Unit
- Harter recommended using Jupyter notebooks for projects, with multiple notebooks when a larger project benefits from being divided into parts.
- A readable notebook should combine runnable code cells with explanations in Markdown cells, using headings, bullets, and tables where appropriate.
- The lecture turns from linear and polynomial regression to logistic regression as a method for classification.
10:43
Sigmoid Probabilities and the 0.5 Classification Threshold
- Binary logistic regression applies the sigmoid, or logistic, function to a linear score such as theta multiplied by the input vector.
- The sigmoid maps scores from negative to positive infinity into values between 0 and 1, interpretable as predicted probabilities.
- A probability below 0.5 is classified as class 0 and one above 0.5 as class 1; a score of zero maps to probability 0.5.
17:22
Why Logistic Regression Uses a Classification-Specific Cost
- Mean squared error, used earlier for regression, is not the standard cost for logistic regression’s probability-based predictions.
- A useful classification cost should be small when the predicted probability agrees with the true 0-or-1 label and large when the model is confidently wrong.
- The lecture introduces negative logarithms to create costs with that behavior for each of the two possible labels.
20:13
Combining the Two Label Cases into Binary Cross-Entropy
- For a true label of 1, the cost uses the negative log of the predicted probability; predictions near 1 have low cost, while predictions near 0 have high cost.
- For a true label of 0, the mirrored term uses the negative log of one minus the predicted probability.
- The two cases combine into a single expression summed over training examples, commonly called binary cross-entropy or log loss.
26:43
Logistic Regression Gradients and Why the Normal Equation Does Not Apply
- The derivative of logistic regression’s cost has a form similar to the linear-regression gradient, but it uses sigmoid predictions rather than an unconstrained linear output.
- This similarity means gradient-descent code for regression can be adapted by applying the sigmoid when calculating predictions, costs, and gradients.
- Unlike linear regression, the nonlinear sigmoid prevents solving for parameters with the normal equation, so an optimizer such as gradient descent is needed.
33:13
Iris Decision Boundaries and Updating the scikit-learn Example
- Harter demonstrates binary classification on the Iris dataset by distinguishing Virginica flowers from non-Virginica flowers.
- A logistic-regression decision boundary separates the two predicted classes; with a single feature, the example places the threshold near petal width 1.66.
- A scikit-learn output-format change caused an array-versus-scalar error in the notebook; Harter refreshed the class repository and showed that the single value must be extracted.
42:29
Using scikit-learn Predict and Predict_proba
- The scikit-learn workflow creates a LogisticRegression estimator and fits it to labeled feature data.
- The predict method returns class labels, while predict_proba provides probability estimates that reveal confidence in the classification.
- In the Iris example, petal-width values on opposite sides of the approximately 1.66 threshold receive different binary predictions.
45:49
Plotting a Two-Feature Boundary with a Mesh Grid and Contours
- For petal length and petal width, the example extracts fitted model parameters to represent the decision boundary in two dimensions.
- A mesh grid samples petal length from about 2.9 to 7 and petal width from about 0.8 to 2.7; the model predicts probabilities across the grid.
- A contour at probability 0.5 shows the classification boundary, with values on opposite sides assigned to Virginica and non-Virginica.
- Some Iris points remain misclassified because one linear boundary cannot perfectly separate the training data.
51:19
Softmax Preview, Assignment Questions, and Closing Reminders
- Harter deferred a full explanation of softmax regression, noting it can extend the probability-based approach to multiclass classification.
- He opened the floor for final questions before the next meeting and reiterated that Assignment 4 was due Sunday.
- A student who had missed the project acceptance deadline was told to complete the README and selection steps that day or the next; Harter also acknowledged a pending email.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Derek Harter.