Exams with More Learning and Less Stress with a Computer-Based Testing Facility - CS50 Tech Talk
Watch on YouTube →
Overview
Craig Zillis advocates for computer-based testing facilities to enhance student learning and reduce stress, presenting data from CS50 and the University of Illinois. His research demonstrates that more frequent, lower-stakes testing (e.g., 50-minute quizzes instead of 2-hour midterms) leads to higher scores and better retention, while second-chance testing approximates mastery learning and reduces failure rates. Zillis highlights the role of tools like PrairieLearn in enabling sophisticated, auto-gradable questions and the operational efficiency of centralized, professionalized proctoring in dedicated testing facilities.
Key takeaways
- More frequent, lower-stakes assessments (e.g., 50-minute quizzes) significantly improve student learning and retention compared to infrequent, high-stakes exams.
- Second-chance testing provides a mechanism for mastery learning within fixed semester lengths, reducing failure rates and test anxiety.
- Computer-based testing facilities, coupled with flexible question generation tools like PrairieLearn, enable secure, scalable, and efficient assessment.
- The 'learning to cheat' phenomenon, explained by fraud triangle theory, is exacerbated by unproctored assessments and increases throughout a semester.
- Centralized, professionalized proctoring and asynchronous scheduling in testing facilities streamline exam administration and accommodate student needs.
- Question generators and auto-grading in tools like PrairieLearn allow for sophisticated, reusable exam content, freeing faculty time for higher-value activities.
Chapters
- CS50 transitioned from traditional in-person proctored exams to online, open-book exams.
- The advent of ChatGPT in Fall 2022 necessitated a re-evaluation of assessment methods.
- A pilot proctored exam saw 10% of students opt for on-campus testing.
- David Malan introduces Craig Zillis, an advocate for computer-based testing facilities.
- Malan notes the perceived datedness of computer labs but acknowledges their potential for controlled technology access.
- Zillis's research focuses on applying computing and data to education.
- Zillis aims to increase student learning and reduce testing anxiety through more frequent and second-chance testing.
- GenAI makes it difficult to trust artifacts without knowing their provenance.
- Proctored exams offer trusted assessments by controlling student access and observing artifact generation.
- Students tend to procrastinate when opportunities exist; frequent testing encourages learning material as it's introduced.
- Frequent testing leads to higher aggregate studying and better exam performance.
- Research shows adding midterms significantly improves final exam scores, with diminishing returns for each additional midterm.
- For a constant amount of content, smaller, more frequent exams lead to better performance on those individual pieces.
- Results are mixed on whether smaller exams impact final exam performance.
- More frequent testing is consistently associated with lower student anxiety and higher course ratings.
- Data from a CS1 course for non-technical majors (approx. 600 students/semester) was analyzed.
- The infrequent semester had one 2-hour midterm (20%) and one 3-hour final (30%).
- The frequent semester broke the midterm into two 50-minute exams (10% each) and the final into a 50-minute midterm (10%) and a 2-hour final (20%).
- Frequent testing semester included unproctored quizzes (2%) the week before proctored exams.
- Infrequent testing semester used purely formative self-assessments (0%) every two weeks.
- Exams were constructed using question generators, ensuring identical content across semesters.
- Scores on Exam 0 (introductory) were statistically equivalent between semesters.
- Scores on subsequent proctored exams were 4-8% higher for the midterm components and ~10% higher for the final components in the frequent testing semester.
- The advantage grew over time, particularly in cumulative courses like CS1.
- Frequent testing leads to 'mass practice' (cramming) for infrequent exams.
- Practice exam generator usage is concentrated in the 48 hours before an exam.
- Frequent testing promotes more consistent studying and 'earnestness' (less brute-forcing multiple-choice options).
- Zybooks' 'earnestness' metric shows students in the frequent semester exhibit less brute-forcing behavior.
- Student studying effort is sublinear with grade percentage; smaller assessments receive disproportionately more study time.
- Breaking assessments into smaller pieces creates more '48-hour study windows,' leading to aggregate increases in studying.
- Second chance testing offers students a second attempt on exams, typically one week apart.
- A grading policy (e.g., 90% of max score, 10% of min score) incentivizes taking the first attempt seriously.
- This approach reduces failure rates, tightens grade distributions, and decreases test anxiety.
- Second chance testing allows instructors to maintain high standards by offering remediation.
- A challenging Verilog exam for a computer organization course saw 30% failure on the first attempt, but only 3% on the second.
- This system supports mastery learning within fixed-semester constraints.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, CS50.