Mailbag: How to Bootstrap Labels for Relevant Docs in Search
Explains how to bootstrap training labels for a semantic search system using initial lexical search and user click data instead of costly human annotation.
Explains how to bootstrap training labels for a semantic search system using initial lexical search and user click data instead of costly human annotation.
Highlights ICML 2021 invited talks on applying machine learning to scientific domains like drug discovery, climate science, poverty alleviation, and neuroscience.
A talk on system design for recommendation and search systems, covering architecture and production considerations.
An in-depth technical explanation of diffusion models, a class of generative AI models that create data by reversing a noise-adding process.
A comprehensive deep learning course covering fundamentals, neural networks, computer vision, and generative models using PyTorch.
A comprehensive deep learning course overview with PyTorch tutorials, covering fundamentals, neural networks, and advanced topics like CNNs and GANs.
Analyzes the legal implications of GitHub Copilot potentially being a derivative work of GPL-licensed code used in its training.
A data scientist shares practical strategies and mindsets for influencing technical teams and driving change without formal authority.
Explores system design patterns for industrial-scale recommendation and search engines, focusing on offline/online components and retrieval/ranking stages.
A guide to essential metrics for evaluating machine learning models, including confusion matrix, accuracy, precision, recall, and more.
Explores machine learning patterns like bandits, sequential, and graph-based models for personalizing recommendations and search results.
A guide to implementing few-shot learning using the GPT-Neo language model and Hugging Face's inference API for NLP tasks.
An interview with data scientist Eugene Yan discussing his career path from psychology to Amazon, favorite ML projects, and advice for aspiring data scientists.
Introducing mltrace, an open-source lineage and tracing tool for debugging and maintaining production machine learning pipelines.
A high-level guide to tools and methods for understanding AI/ML models and their predictions, known as Explainable AI (XAI).
Explores the strategic 'metagame' of applying machine learning in industry, focusing on problem selection and business impact over pure technical knowledge.
Explores the distinction between using regression models for causal inference versus predictive inference, and the role of generalizability in prediction.
Explores how mutual information and KL divergence can be used to derive information-theoretic generalization bounds for Stochastic Gradient Descent (SGD).
An interview with AI researcher Joelle Pineau discussing her work in reinforcement learning, its applications, and advice for newcomers to the field.
A guest post sharing personal stories of imposter syndrome in tech and academia, with lessons on recognizing and managing self-doubt.