
Sign up to save your podcasts
Or


MLOps Coffee Sessions #69 with James Lamb, Building for Small Data Science Teams, co-hosted by Adam Sroka.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Abstract
In this conversation, James shares some hard-won lessons on how to effectively use technology to create applications powered by machine learning models.
James also talks about how making the "right" architecture decisions is as much about org structure and hiring plans as it is about technological features.
// Bio
James Lamb is a machine learning engineer at SpotHero, a Chicago-based parking marketplace company. He is a maintainer of LightGBM, a popular machine learning framework from Microsoft Research, and has made many contributions to other open-source data science projects, including XGBoost and prefect.
Prior to joining SpotHero, he worked on a managed Dask + Jupyter + Prefect service at Saturn Cloud and as an Industrial IoT Data Scientist at AWS and Uptake. Outside of work, he enjoys going to hip hop shows, watching the Celtics / Red Sox, and watching reality TV (he wouldn’t object to being called “Bravo Trash”).
// Relevant Links
James keeps track of conference and meetup talks he has given at https://github.com/jameslamb/talks#gallery. The audience for this podcast might be most interested in "Scaling LightGBM with Python and Dask" and "How Distributed LightGBM on Dask Works".
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, blogs, newsletters, and more: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Adam on LinkedIn: https://www.linkedin.com/in/aesroka/
Connect with James on LinkedIn: https://www.linkedin.com/in/jameslamb1/
Timestamps:
[00:00] Introduction to James Lamb
[01:11] James' background in the machine learning space
[03:24] LightGBM
[09:56] Community behind LightGBM
[13:36] Background of James in SpotHero
[20:06] Experience in Maturity Models
[22:40] Bottlenecks of tradeoffs between speed and confidence
[28:28] Tools to be excited about
[31:46] To code your own that's already out there
[36:33] Building design decisions
[39:36] Risk of the unicorn
[42:44] Cross-team empathy
[47:18] Proudest technical accomplishment and/or biggest frustration, less proud of lessons learned
[50:53] SpotHero is hiring!
[51:20] Wrap up
[51:53] Please like, subscribe, and you can leave a review!
MLOps Coffee Sessions #68 with Chris Albon, Wikimedia MLOps co-hosted by Neal Lathia.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Abstract
Chris Alban (Wikimedia ML team lead) and Neil Lithia discuss Alban's high-output drive, Wikimedia's open-source ML infrastructure, and a six-person team's role in maintaining editor-assist models like article quality prediction and mobile "Add-a-Link" features. Key topics include agile workflows for rapid model deployment via Kubeflow, repeatability through enforced policies, open-source tooling challenges (e.g., AMD GPUs), ethical governance with model cards, and community-trained models. The session highlights Wikimedia's lean, donation-funded scale and calls for contributions.
// Bio
Chris spent over a decade applying statistical learning, artificial intelligence, and software engineering to political, social, and humanitarian efforts. He is the Director of Machine Learning at the Wikimedia Foundation. Previously, Chris was the Director of Data Science at Devoted Health, Director of Data Science at the Kenyan startup BRCK, cofounded the AI startup Yonder, created the data science podcast Partially Derivative, was the Director of Data Science at the humanitarian non-profit Ushahidi, and was the director of the low-resource technology governance project at FrontlineSMS. Chris also wrote Machine Learning for Python Cookbook (O’Reilly 2018) and created Machine Learning Flashcards.
Chris earned a Ph.D. in Political Science from the University of California, Davis, researching the quantitative impact of civil wars on health care systems. He earned a B.A. from the University of Miami, where he triple majored in political science, international studies, and religious studies.
// Relevant Links
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, Feature Store, Machine Learning Monitoring, and Blogs: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Neal on LinkedIn: https://www.linkedin.com/in/nlathia/
Connect with Chris on LinkedIn: https://www.linkedin.com/in/chrisralbon/
Timestamps:
[00:00] Introduction to Chris Albon
[00:28] Do you sleep? :-)
[02:43] ML at Wikimedia
[09:27] Wikimedia workflow
[15:00] Creating a repeatable process
[19:11] Wikimedia element team size
[20:47] Wikimedia workflow and hardware
[23:56] Evaluating open source
[29:20] Lacking in ML source tooling
[33:11] Wikimedia's separate data platform
[38:14] Abstractions
[41:50] Experimentation aspect of getting models into production
[44:05] Stack of Abstraction in ML
[47:16] Chris' proudest model
[49:10] How Wikimedia works with communities
[55:24] Large language models
[1:02:16] Beautiful vision
[1:03:23] Wrap up
MLOps Coffee Sessions #67 with John Crousse, ML Stepping Stones: Challenges & Opportunities for Companies, co-hosted by Adam Sroka.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Abstract
In this coffee session, John shares his observations after working with multiple companies that were in the process of scaling up their ML capabilities.
John's observations are mostly around changes in practices, successes, failures, and bottlenecks identified when building ML products and teams from scratch. John shares a few thoughts on building long-term products vs short-term projects, on the important non-ML components, and the most common missing pieces he sees in today's ecosystem. John also elaborates on how those challenges and solutions can differ for different company sizes.
// Bio
John always liked CS/ML/AI, but it wasn't such a hot topic back then. He found opportunities to work on models in the Financial industry as a consultant from 2007 to 2017, then he went freelance to move outside of the financial industry and focus on AI/ML.
John likes to do things efficiently, and MLOps is the bottleneck, so he ended up spending more time on MLOps than on models lately.
John finished his CS degree in 2007.
// Relevant Links
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, Feature Store, Machine Learning Monitoring, and Blogs: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Adam on LinkedIn: https://www.linkedin.com/in/aesroka
Connect with John on LinkedIn: https://www.linkedin.com/in/john-crousse-31219b9
Timestamps:
[00:00] Introduction to John Crousse
[01:11] Main trends in Machine Learning
[03:07] Symptoms of a Machine Learning product
[05:05] Proper product with limited resources
[08:52] Going into production mindsets
[11:22] Bottlenecks and challenges
[14:55] Business case for Machine Learning or MLOps in small organizations
[17:04] Gathering feedback is best suited to product owners
[19:14] More substantial role
[20:11] Data factory
[24:03] Delivery patterns or tech stacks
[26:06] Bottleneck metrics
[27:28] Concept of evaluation store
[32:18] The biggest gap to bridge
[34:42] Hindrance to people's development
[35:23] "The last mile of the machine learning projects"
[36:40] MLOps assessment survey
[40:10] Who owns the product and the path to recommend
[41:34] Datamesh community
[44:41] Tips on balancing between pure autonomy
[45:58] Wrap up
MLOps Coffee Sessions #66 with Jacopo Tagliabue, Machine Learning at Reasonable Scale.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Abstract
We believe that immature data pipelines are preventing a large portion of industry practitioners from leveraging the latest research on ML: the truth is, outside of Big Tech and advanced startups, ML systems are still far from producing the promised ROI.
The good news is that times are changing: thanks to a growing ecosystem of tools and shared best practices, even small teams can be incredibly productive at a “reasonable scale”. Based on our experience as founders and researchers, we present our philosophy for modern, no-nonsense data pipelines, highlighting the advantages of a "PaaS-like" approach.
// Bio
Educated in several acronyms across the globe (UNISR, SFI, MIT), Jacopo Tagliabue was co-founder and CTO of Tooso, an A.I. company in San Francisco acquired by Coveo in 2019. Jacopo is currently the Director of AI at Coveo, shipping models to hundreds of customers and millions of users. When not busy building products, he is exploring topics at the intersection of language, reasoning, and learning: his research and industry work are often featured in the general press and premier A.I. venues. In previous lives, he managed to get a Ph.D., do sciency things for a pro basketball team, and simulate a pre-Columbian civilization.
// Relevant Links
Bigger boat repo: https://github.com/jacopotagliabue/you-dont-need-a-bigger-boat
TDS series: https://towardsdatascience.com/tagged/mlops-without-much-ops (ep 3 and a NEW open-source contribution on data ingestion coming up)
Open datasets for e-commerce and MLops experiments: https://github.com/coveooss/SIGIR-ecom-data-challenge
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, Feature Store, Machine Learning Monitoring, and Blogs: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Vishnu on LinkedIn: https://www.linkedin.com/in/vrachakonda/
Connect with Jacopo on LinkedIn: https://www.linkedin.com/in/jacopotagliabue/
Timestamps:
[00:00] Introduction to Jacopo Tagliabue
[01:35] What Reasonable Scale means
[06:40] Biggest disconnects from Reasonable Scale
[12:32] Engineers need to do and tools to use at a Reasonable Scale
[15:25] Importance of maintenance
[17:27] Bigger boat repo demonstration of Reasonable Scale
[23:09] The Four Pillars
[27:27] ETL Paradigm
[30:16] Best practices around dragons in generic decisions and comparing the new outputs and saved snapshots
[33:32] Creating a knowledge hub
[36:28] Continuation of principles
[38:06] Distributed road
[42:24] Current state-of-the-art recommender systems
[49:04] What Kovio and TUSU do in recommender systems in the world
[53:19] Stack in recommender system
[59:11] Being optimistic in the current ecosystem we're living in
[1:01:43] Wrap up
MLOps Coffee Sessions #65 with Skylar Payne, The Future of Data Science Platforms is Accessibility.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Abstract
The machine learning and data science space is blowing up -- new tools are popping up every day. While we seem to have every type of "Flow" and "Store" you could imagine, few people really understand how to glue this stuff together. Despite all the tools we have available, we still see companies failing to leverage data science effectively to drive business results.
Instead of spending time driving business results, data scientists spend their time fiddling with Kubernetes, trying to debug that Spark serialization error figuring out how to map their code into the awkward "AI Pipeline" SDK. We have an industry filled with tools built by engineers... for engineers, rather than for data scientists. It's deeply disempowering.
Meanwhile, data is still used effectively to drive decisions in many companies. Analysts have been solving very similar problems on the back of applications like Excel, Tableau, and Mode for literally decades. While there are still challenges in analytics, the MLOps space could learn something from analytics tools. Analytics tools better understand how to make their tools accessible. Analytics tools better understand the value of iterability. Analytics tools better understand that data problems are wicked problems:
- We have to iterate on the formulation and solution simultaneously
- They involve many stakeholders with different opinions
- There's no "right" answer
- The problems are never 100% solved.
If we're going to really drive the most business value from data science, we need to understand how to design our teams and tools to effectively work against such problems.
The future of data science platforms is accessibility and iterability.
// Bio
Data is a superpower, and Skylar has been passionate about applying it to solve important problems across society. For several years, Skylar worked on large-scale, personalized search and recommendation at LinkedIn -- leading teams to make step-function improvements in our machine learning systems to help people find the best-fit role. Since then, he shifted my focus to applying machine learning to mental health care to ensure the best access and quality for all. To decompress from his workaholism, Skylar loves lifting weights, writing music, and hanging out at the beach!
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, Feature Store, Machine Learning Monitoring, and Blogs: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Vishnu on LinkedIn: https://www.linkedin.com/in/vrachakonda/
Connect with Skylar on LinkedIn: https://www.linkedin.com/in/skylar-payne-766a1988/
Timestamps:
[00:00] Introduction to Skylar Payne
[00:25] Skylar's blog post overview
[00:55] Data is Wicked
[02:22] Bundling & unbundling
[05:48] ML world vs Analytics world
[08:40] Startups from various perspectives
[11:27] Setting the right building blocks
[15:05] Defining process and interfaces
[19:51] KubeFlow success stories accessibility
[21:17] Machine Learning + Data Science
[26:48] Where to spend more time?
[28:19] Privacy
[34:28] Measuring Apps Feeds
[38:46] Difficult trade-offs
[42:46] Tools improvement in workflow
[47:24] Accessibility & Iterability
MLOps Reading Group meeting on November 20, 2021
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Connect with us on LinkedIn: https://www.linkedin.com/company/mlopscommunity/
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, Feature Store, Machine Learning Monitoring, and Blogs: https://mlops.community/
MLOps Coffee Sessions #64 with Slater Victoroff, The Future of AI and ML in Process Automation.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Abstract
The Unstructured Imperative
Recent advances in AI have dramatically advanced the state of the art around unstructured data, especially in the spaces of NLP and computer vision. Despite this, the adoption of unstructured technologies has remained low. Why do you think that is? How have the dynamics changed in the last five years?
Multimodal AI
Historic AI approaches have generally been constrained to one data modality (i.e., text or image). Recently, a wide range of papers in image captioning and document understanding have emphasized the need for more sophisticated "multimodal" techniques that can fuse information from multiple modalities. What is multimodal learning, and why is it so promising? Why are we seeing such an explosion of activity? What is Indico doing in this space?
Machine Teaching
As methods of supervision become more complex and multifaceted, many researchers have begun investigating the inverse problem. That is how do we design supervision systems that more naturally follow human processes? What are some interesting trends in "the space", and where can we expect this field to go in the next few years?
// Bio
Slater Victoroff is the Founder and CTO of Indico, an enterprise AI solution for unstructured content that emphasizes document understanding.
Slater has been building machine learning solutions for startups, governments, and Fortune 100 companies for the past seven years and is a frequent speaker at AI conferences.
Indico’s framework requires 1000x less data than traditional machine learning techniques, and they regularly beat the likes of AWS, Google, Microsoft, and IBM in head-to-head bake-offs.
// Relevant Links
https://indico.io
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, Feature Store, Machine Learning Monitoring, and Blogs: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Vishnu on LinkedIn: https://www.linkedin.com/in/vrachakonda/
Connect with Slater on LinkedIn: https://www.linkedin.com/in/slatervictoroff
Timestamps:
[00:00] Introduction to Slater Victoroff
[03:52] Slater's journey into ML
[06:52] Birth of Indico
[07:04] "Unstructured data"
[09:47] Historical perspective journey of Indico
[11:13] Adoption of unstructured technologies
[16:05] Technology techniques
[21:40] Unstructured data challenges
[26:35] Changing Supervision
[28:35] Synthetic data and its role
[32:41] Human overfitting
[34:15] Intuition in productive systems
[38:21] Data flows and information flows
[40:17] Typical Indico client
[42:34] Client Accessibility
[43:08] Process Configurability
[45:20] Indico clients dealing with workflow
[47:26] "Models won't fix a broken process."
[49:19] Model understandability and explainability
[50:03] Programming data rather than programming models
[50:55] Modeling limiting factor
[51:44] "The shift is happening!"
[52:40] Machine Learning Engineering vs Software Engineering
[53:48] Advice to 2021 Machine Learning Engineers to focus on
Dmytro Dzhulgakov, PyTorch: Bridging AI Research and Production.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
Talking PyTorch is always interesting, as the Facebook ML OSS project is one of the most important parts of the machine learning tooling ecosystem. This week, we talked to Dmytro Dzhulgakov, a tech lead for PyTorch.
We started off talking about Dmytro's journey to being an engineer and tech lead at Facebook, and what his role entails. Dmytro has been at Facebook for 10+ years, so he gave some very interesting advice on how to manage a career in software engineering for the machine learning world. After that, we got deep into the present and future of PyTorch and what improvements the project is making to support MLOps workflows. PyTorch is a large project, and Dmytro shared with us the valuable lessons he learned from confronting multifaceted scaling challenges while working on PyTorch. Finally, we talked about the future of machine learning engineering, especially as it relates to how software engineers work by comparison.
// Abstract
Over the past few years, PyTorch has become the tool of choice for many AI developers, ranging from academia to industry. With the fast evolution of state-of-the-art in many AI domains, the key desired property of the software toolchain is to enable the swift transition of the latest research advances to practical applications.
In this coffee session, Dmytro discusses some of the design principles that contributed to this popularity, how PyTorch navigates inherent tension between research and production requirements, and how AI developers can leverage PyTorch and PyTorch ecosystem projects for bringing AI models to their domain.
// Bio
Dmytro Dzhulgakov is a technical lead of PyTorch at Facebook, where he focuses on the framework's core development and building the toolchain for bringing AI from research to production.
Previously, he was one of the creators of ONNX, a joint initiative aimed at making AI development more interoperable. Before that, Dmytro built several generations of large-scale machine learning infrastructure that powered products like Ads or News Feed.
// Relevant Links
https://pytorch.org/
https://pytorch.org/blog/
https://ai.facebook.com/blog/pytorch-builds-the-future-of-ai-and-machine-learning-at-facebook/
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, Feature Store, Machine Learning Monitoring, and Blogs: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Vishnu on LinkedIn: https://www.linkedin.com/in/vrachakonda/
Connect with Dmytro on LinkedIn: https://www.linkedin.com/in/dzhulgakov/
Timestamps:
[00:00] Introduction to Dmytro Dzhulgakov
[00:00] Dmytro's journey to his current position
[05:25] Interest in staying on Facebook for so long
[08:36] What PyTorch project?
[11:23] ML Infra Evolution
[16:17] PyTorch now and its future
[22:16] Balancing product development
[27:40] PyTorch's evolution in production
[37:45] Lessons learned from failures in PyTorch
[43:41] Culmination of war stories
[45:50] Seamless merging
[46:47] Future of software engineers and machine learning engineers
MLOps Coffee Sessions #63 with Dmytro Dzhulgakov, PyTorch: Bridging AI Research and Production.
MLOps Coffee Sessions #62 with Joel Grus, MLOps from Scratch.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Abstract
In this talk, Joel Grus of “I don’t like notebooks” fame shares with us his 2021 perspective on notebooks, where he thinks MLOps is now, and what his hot takes in the data space are now.
// Bio
Joel Grus is a Principal Engineer at Capital Group, where he leads a team that builds search, data, and machine learning products for the investment group. He is the author of the bestselling O'Reilly book *Data Science from Scratch*, the not-bestselling self-published book *Ten Essays on Fizz Buzz*, and the controversial JupyterCon talk "I Don't Like Notebooks." He recently moved to Texas after living in Seattle for a very long time.
// Relevant Links
Data Science from Scratch book: https://www.oreilly.com/library/view/data-science-from/9781491901410/
Data Science from Scratch, 2nd Edition book: https://www.oreilly.com/library/view/data-science-from/9781492041122/
Ten Essays on Fizz Buzz: Meditations on Python, mathematics, science, engineering, and design book: https://www.amazon.com/Ten-Essays-Fizz-Buzz-Meditations/dp/0982481829 or https://leanpub.com/fizzbuzz/
I Don't Like Notebooks talk: https://www.youtube.com/watch?v=7jiPeIFXb6U
I Don't Like Notebooks - #JupyterCon 2018 slides: https://docs.google.com/presentation/d/1n2RlMdmv1p25Xy5thJUhkKGvjtV-dkAIsUXP-AL4ffI/edit#slide=id.g362da58057_0_658
Fizz Buzz in Tensorflow: https://joelgrus.com/2016/05/23/fizz-buzz-in-tensorflow/
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, Feature Store, Machine Learning Monitoring, and Blogs: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Vishnu on LinkedIn: https://www.linkedin.com/in/vrachakonda/
Connect with Joel on LinkedIn: https://www.linkedin.com/in/joelgrus/
Timestamps:
[00:00] Introduction to Joel Grus
[01:32] Joel's background in tech
[07:47] Joel's I Don't Like Notebooks talk on Jupyter Con
[13:42] Better tooling around notebooks
[16:48] Hex
[17:20] Step function evolution
[20:41] Kinds of professionals required in Joel's organization to practice MLOps
[23:08] Evaluation process
[25:51] Sagemaker bring your own algorithm
[27:30] Flexibility of models
[31:55] Hot takes on the data science world
[34:19] Current Overall Maturity of MLOps
[37:23] Kinds of problems in NLP and search
[39:52] Finding ways to put structures
[40:50] Probabilistic nature of machine learning systems
[43:10] Data scientists catching up on writing production code
[46:33] Invaluable code review
[47:22] Common repo structure
[47:57] Reviewing codes
[49:15] Code pals
[50:36] Readability and function
[52:23] Leverage code review
[53:10] Remote work
From the publisher's feed

1,289 Listeners

286 Listeners

1,089 Listeners

622 Listeners

582 Listeners

304 Listeners

337 Listeners

203 Listeners

565 Listeners

512 Listeners

141 Listeners

102 Listeners

222 Listeners

685 Listeners

30 Listeners