Luis Serrano + Josh Starmer Q&A Livestream!!!
Watch on YouTube →
Overview
Josh Starmer and Luis Serrano discuss various AI and machine learning topics, including strategies for landing an AI job (portfolio building on GitHub, hands-on practice with Kaggle), the origins of StatQuest's "StatSquatch" and "Normal Saurus" mascots, and the mathematical reasoning behind Bessel's correction (n-1 denominator for sample variance). They explore real-world applications of eigenvalue decomposition in Principal Component Analysis (PCA) for denoising, the future of agents and multimodal LLMs, and the use of Transformers in omics data for gene regulation. The conversation also touches on LoRA for efficient LLM fine-tuning, XGBoost's weak learner strategy, the challenges of data contamination from LLM hallucinations, and the pros/cons of model size versus fine-tuning.
Key takeaways
- Landing an AI job requires a combination of formal learning (courses, channels like StatQuest) and practical, hands-on experience, best demonstrated through a robust GitHub portfolio.
- Bessel's correction (n-1 denominator for sample variance) can be intuitively understood by defining variance as the average squared distance between any two data points, rather than deviation from the mean.
- Principal Component Analysis (PCA), rooted in eigenvalue decomposition, is a key technique for denoising high-dimensional data by identifying and retaining the most informative components.
- LoRA (Low-Rank Adaptation) offers a significantly more efficient method for fine-tuning large language models by injecting small, trainable low-rank matrices, reducing computational cost and time.
- The increasing generation of content by LLMs poses a risk of data contamination with hallucinations, potentially degrading future training datasets; RAG, tool use, and data curation are vital countermeasures.
- Individuals with PhDs in non-AI fields are uniquely positioned to merge their domain expertise with AI skills, becoming highly valuable assets capable of driving innovation in specialized areas.
Chapters
- Josh Starmer and Luis Serrano host a live Q&A stream.
- Viewers join from diverse international locations including Chapel Hill, Toronto, India, Greece, and Switzerland.
- The stream aims to answer pre-submitted and live chat questions.
- AI jobs welcome diverse backgrounds; focus on applying existing skills to AI.
- Recommended learning resources include Coursera, Udacity, and channels like StatQuest.
- Hands-on experience is crucial, with Kaggle suggested for practice.
- Building a strong GitHub portfolio with well-commented code demonstrates initiative and collaboration skills.
- Normal Saurus originated from a friend's daughter's t-shirt design featuring a dinosaur with a normal curve.
- StatSquatch represents Josh Starmer's learning process, asking fundamental questions.
- The mascot embodies the persona of someone learning new material and asking 'dumb' questions.
- The question addresses why sample variance divides by n-1 instead of n.
- Traditional explanation involves degrees of freedom, which is not intuitively satisfying.
- An alternative perspective defines variance as the average squared distance between any two points in the dataset.
- This approach naturally leads to dividing by n*(n-1) for pairs, and then by 2 to get the variance, effectively yielding n-1 in the denominator.
- Eigenvalue decomposition is a root for Principal Component Analysis (PCA), though Singular Value Decomposition (SVD) is now more common.
- PCA is used for denoising high-dimensional datasets by reducing them to a few principal components (eigenvectors).
- This process identifies and retains the most informative variables, discarding noise.
- Agents represent the next step beyond LLMs, enabling them to perform actions (e.g., RAG for searching external data).
- Tools allow LLMs to execute tasks like sending emails or writing/compiling code.
- Multimodality integrates learning from images, sounds, and other data types beyond text.
- Agentic workflows and multimodal capabilities are seen as the future evolution of LLMs.
- Omics refers to large biological datasets (genomics, transcriptomics, proteomics).
- Transformers and LLMs are being used to generate gene sequences that can regulate gene expression.
- This research aims to control gene activation/deactivation, potentially for therapeutic purposes (e.g., preventing cells from forgetting their identity in cancer).
- BERT, an LLM, is being trained on genomic data (long sequences of DNA base pairs) to understand and generate functional sequences.
- LoRA addresses the high cost of training/fine-tuning massive LLMs (trillions of parameters).
- It leverages the observation that large models often have low-dimensional 'flat' regions.
- LoRA injects small, trainable 'low-rank' matrices into existing Transformer layers.
- This significantly reduces the number of trainable parameters, making fine-tuning more efficient and accessible.
- In XGBoost, weak learner trees can be fit on all predictors or a random subset, similar to Random Forests.
- Practically, a random subset of variables is often used per split for efficiency.
- The concept of 'sign' (+/-) for variable importance in tree models is discussed; it might relate to regression trees' impact on output value.
- The concern is that LLMs generating internet content will contaminate future training data with hallucinations.
- This could lead to a 'vanishing gradient of information,' where all online content becomes unreliable.
- Hallucinations are seen as a feature, not a bug, of LLMs; they learn language but not necessarily facts.
- Techniques like RAG, fine-tuning, and tool use are crucial for mitigating LLM unreliability.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, StatQuest with Josh Starmer.