Human Stories in AI: Fabio Urbina
Watch on YouTube →
Overview
Fabio Urbina, Associate Director at Collaborations Pharmaceuticals, leverages machine learning for early-stage drug discovery, focusing on rare and neglected diseases. His journey began with a biology degree and a computer science minor, leading to research in cell biology and computational image analysis. At Collaborations Pharmaceuticals, Fabio utilizes prototypical networks and embedding models to virtually screen vast compound libraries, significantly accelerating the identification of potential therapeutics, even with extremely small datasets (as few as 15 compounds).
Key takeaways
- Fabio Urbina's work at Collaborations Pharmaceuticals uses machine learning to accelerate drug discovery for rare diseases, a field often overlooked by traditional VC-funded biotechs.
- Prototypical networks, combined with embedding models, enable effective drug candidate prediction even with extremely small datasets (as few as 15 compounds), a significant departure from large-scale AI models.
- The core challenge in ML for drug discovery is data scarcity, requiring specialized techniques like few-shot learning and careful consideration of the model's applicability domain.
- Simple, older ML models (like SVMs, random forests) can outperform newer, complex models in data-scarce domains like drug discovery due to their robustness and efficient handling of limited data.
- Embracing discomfort and persistent learning is crucial for career changers, especially when transitioning into new fields with unfamiliar concepts and terminology.
Chapters
- Series features AI experts' career journeys.
- Fabio Urbina is Associate Director at Collaborations Pharmaceuticals.
- He combines computational tools, ML, and biology for drug discovery.
- Focuses on previously difficult-to-probe scientific problems.
- Childhood interest in science and computers.
- Bachelor's in Biology with a minor in Computer Science.
- Internship at Massachusetts General Hospital researching familial dysautonomia, a rare disease.
- Gained experience in rare diseases and early-stage drug discovery.
- PhD at UNC Chapel Hill in Stephanie Gupton's lab.
- Studied neuron cell biology and development.
- Applied computational image analysis to cell biology imaging.
- Integrated computer science and statistical analysis into experimental work.
- Decided against an academic career post-PhD.
- Sought industry internships during the last year of PhD.
- Found Collaborations Pharmaceuticals, a small biotech company (<10 people).
- Company focuses on rare and neglected diseases, funded by government grants, not VC.
- Focus on early-stage drug discovery, e.g., finding new antimalarials.
- Traditional methods: modifying existing drugs or brute-force screening.
- ML approach: train models on known drug structures (active/inactive) to predict activity of new molecules.
- Virtual screening of large compound libraries to identify promising candidates.
- Success rates vary, from 3/3 active compounds to 0/10 active.
- Achieved enrichment rates of 10-1000x better chance of finding new compounds.
- Challenge: 'Coverage' or 'Applicability Domain' – models trained on narrow chemical spaces may over-predict.
- Data scarcity is a major issue due to high experimental cost per compound.
- Drug discovery datasets are tiny (hundreds to ~100,000 data points) compared to LLMs.
- Developed a model with predictive power on a 15-compound dataset (few-shot learning).
- Utilized prototypical networks, a simplistic model designed for bias alignment.
- Model consists of an embedding model (molecule to vector) and the network itself.
- Embedding model maps similar structures to similar numerical vectors.
- Prototypical network finds the average vector (prototype) for each class (e.g., antimalarial vs. non-antimalarial).
- New compounds are classified by proximity to these prototypes.
- Model is computationally inexpensive, runnable on standard hardware, unlike large LLMs.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, StatQuest with Josh Starmer.