
Sign up to save your podcasts
Or


Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
James Sutton is an ML Engineer focused on helping enterprises bridge the gap between what they have now and where they need to be to enable production-scale ML deployments.
----------- Connect With Us ✌️-------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn: https://www.linkedin.com/in/aponteanalytics/
Connect with James on LinkedIn: https://www.linkedin.com/in/jamessutton2/
Timestamps:
0:00 - Intro to Speaker
2:20 - Scope of the coffee session
3:10 - Background of James Sutton
8:28 - One-Shot Classifier Algorithm
12:46 - Why is it a challenge from the engineering perspective with deployment?
19:20 - How to overcome bottlenecks?
30:07 - Vision of your landscape?
34:45 - Maturity playout
38:48 - Maturity perspective of ML
41:49 - Risk of overgeneralizing system design patterns
46:10 - Reliability, Speed, Cost
46:46 - Consistency, Availability, Partition Tolerance (CAP Theorem)
47:36 - How do you go about discussing these tradeoffs with your clients?
51: 23 - How would you deal with the PII?
58:50 - Collaborative process with clients
1:00:55 - Wrap up
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
Parallel Computing with Dask and Coiled
Python makes data science and machine learning accessible to millions of people around the world. However, historically, Python hasn't handled parallel computing well, which leads to issues as researchers try to tackle problems on increasingly large datasets. Dask is an open source Python library that enables the existing Python data science stack (Numpy, Pandas, Scikit-Learn, Jupyter, ...) with parallel and distributed computing. Today, Dask has been broadly adopted by most major Python libraries and is maintained by a robust open source community across the world.
This talk discusses parallel computing generally, Dask's approach to parallelizing an existing ecosystem of software, and some of the challenges we've seen in deploying distributed systems.
Finally, we also addressed the challenges of robustly deploying distributed systems, which ends up being one of the main accessibility challenges for users today. We hope that by the end of the meetup, attendees will better understand parallel computing, have built intuition around how Dask works, and have the opportunity to play with their own Dask cluster on the Cloud.
Matthew is an open source software developer in the numeric Python ecosystem. He maintains several PyData libraries, but today focuses mostly on Dask, a library for scalable computing. Matthew worked for Anaconda Inc. for several years, then built out the Dask team at NVIDIA for RAPIDS, and most recently founded Coiled Computing to improve Python's scalability with Dask for large organizations.
Matthew has given talks at a variety of technical, academic, and industry conferences. A list of talks and keynotes is available at (https://matthewrocklin.com/talks).
Matthew holds a bachelor’s degree from UC Berkeley in physics and mathematics, and a PhD in computer science from the University of Chicago.
Check out our posts here to get more context around where we're coming from:
https://medium.com/coiled-hq/coiled-dask-for-everyone-everywhere-376f5de0eff4
https://medium.com/coiled-hq/the-unbearable-challenges-of-data-science-at-scale-83d294fa67f8
----------- Connect With Us ✌️-------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn: https://www.linkedin.com/in/aponteanalytics/
Connect with Matthew on LinkedIn: https://www.linkedin.com/in/matthew-rocklin-461b4323/
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
This time, we talked about one of the most vibrant questions for any MLOps practitioner: how to choose the right tools for your ML team, given the huge amount of open-source and proprietary MLOps tools available on the market today.
We discussed several criteria to rely on when choosing a tool, including:
- The requirements of the particular team use-cases
- The scaling capacity of the tool
- The cost of migration from a chosen tool
- The cost of teaching the team to use this tool
- The company or the community behind the tool
Apart from that, we talked about particular use-cases and discussed the trade-offs between waiting for a new release of your tool to get the missing piece of functionality, switching to another tool, and building an in-house solution.
We also touched on the topic of organising MLOps teams and practices across large companies with a lot of ML teams.
// Bio:
Jose Navarro
Jose Navarro is a Machine Learning Infrastructure Engineer making everyday cooking fun at Cookpad, where its recipe platform has more than 40 million monthly users. He holds an MSc in Machine Learning and High-Performance Computing from the University of Bristol. He is interested in Cloud Native technologies, serverless, and event-driven architecture.
Mariya Davidova
Mariya came to MLOps from a software development background. She started her career as a Java developer at JetBrains in 2011, then gradually moved to developer advocacy for JS-based APIs. In 2019, she joined Neu.ro as a platform developer advocate and then moved to the product management position.
Mariya has been obsessed with AI and ML for many years: she finished a bunch of courses, read a lot of books, and even wrote a couple of fiction stories about AI. She believes that proper tooling and decent development and operations practices are essential success components for ML projects, as well as they are for traditional SD.
----------- Connect With Us ✌️-------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn: https://www.linkedin.com/in/aponteanalytics/
Connect with Jose on LinkedIn: https://www.linkedin.com/in/jose-navarro-2a57b612/
Connect with Maria on LinkedIn: https://www.linkedin.com/in/mariya-davydova/
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
Dask
What is it?
Parallelism for analytics
What is parallelism?
Doing a lot at once by splitting tasks into smaller subtasks, which can be processed in parallel (at the same time)
Distributed work across multiple machines and then combined the results
Helpful for CPU-bound - doing a bunch of calculations on the CPU. The rate at which the process progresses is limited by the speed of the CPU
Concurrency?
Similar a but things don’t have to happen at the same time, they can happen asynchronously. They can overlap.
Shared state
Helpful to I/O bound - networking, reading from disk, etc. The rate at which a process progresses is limited by the speed of the I/O subsystem.
Multi-core vs distributed
Multi-core is a single processor with 2 or more cores that can cooperate through threads - multithreading
Distributed across multiple nodes communicating via HTTP or RPC. Why is this hard?
Python has its challenges due to GIL; other languages don't have this problem
Shared state can lead to potential race conditions, deadlocks, etc
Coordinate work across the machines
For analytics?
Calculating some statistics on a large dataset can be tricky if it can’t fit in memory
// Show Notes
Coiled Cloud: https://cloud.coiled.io/
Coiled Launch Announcement: https://medium.com/coiled-hq/coiled-dask-for-everyone-everywhere-376f5de0eff4
OSS article: https://www.forbes.com/sites/glennsolomon/2020/09/15/monetizing-open-source-business-models-that-generate-billions/#2862e47234fd
Amish barn raising: https://www.youtube.com/watch?v=y1CPO4R8o5M
MessagePassingInterface: https://en.wikipedia.org/wiki/Message_Passing_Interface
----------- Connect With Us ✌️-------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn: https://www.linkedin.com/in/aponteanalytics/
Connect with Matthew on LinkedIn: https://www.linkedin.com/in/matthew-rocklin-461b4323/
Timestamps:
0:00 - Intro to Matthew Rocklin and Hugo Bowne-Anderson
0:37 - Matthew Rocklin's Background
1:17 - Hugo Brown-Anderson's Background
3:47 - Where did that inspiration come from?
10:04 - Is there a close relationship between Best Practices and Tooling, or are these two separate things?
11:27 - Why is Data Literacy important with Coiled?
14:46 - How do you think about the balance between enabling Data Science to have a lot of powerful compute?
17:05 - Machine Learning as a space for tracking best practices experimentation
19:32 - What makes Data Science so difficult?
24:07 - How can a for-profit company complement Open Source Software (OSS)
29:40 - Amazon becoming a competitor with your own open-source technology (?)
32:50 - How do you encourage more people to contribute and ensure quality?
34:58 - Do you see Coiled operating within the DASK ecosystem?
37:30 - What is DASK?
39:19 - What should people know about parallelism?
41:28 - Why is it so hard to put things back together?
41:34 - Why does Python need a whole new tool to enable that? Or maybe some other tools as well?
44:44 - Dynamic Tasks Scheduling as being useful to Data Scientists
47:15 - Why is reliability in particular important in Data Science?
52:27 - What's in store for DASK?
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
Why was Flyte built at Lyft?
What sorts of requirements does an ML infrastructure team have at Lyft?
What problems does it solve/use cases?
Where does it fit in the ML and Data ecosystem?
What is the vision?
Who should consider using it?
Learnings as the engineering team tried to bootstrap an open-source community.
Ketan Umare is a senior staff software engineer at Lyft, responsible for the technical direction of the Machine Learning Platform, and is a founder of the Flyte project. Before Flyte, he worked on ETA, routing, and mapping infrastructure at Lyft. He is also the founder of Flink Kubernetes operator and a contributor to Spark on Kubernetes. Prior to Lyft, he was a founding member of Oracle Baremetal Cloud and led teams building Elastic Block Storage. Prior to that, he started and led multiple teams in Mapping and Transportation optimization infrastructure at Amazon. He received his Master's in Computer Science from Georgia Tech, specializing in High-performance computing, and his Bachelor's in Engineering in Computer Science from VJTI Mumbai.
Besides work, he enjoys spending time with his daughter and wife. He loves the Pacific Northwest outdoors and will try anything new.
Lyft
Pricing, Locations, Estimated Time of Arrivals (ETA), Mapping, Self-Driving (L5), etc.
What sort of scale, storage, and network bandwidth are we looking at?
Tens of thousands of workflows, hundreds of thousands of executions, millions of tasks, and tens of millions of containers!
Flyte: more than 900k workflows executed a month and more than 30+ million container executions per month
Typical flow of information?
What are the user stories you’re typically dealing with at Lyft?
How do you set it up?
On-prem, cloud, etc.
Helm installable?
Why Golang?
What problems does it solve?
Complex data dependencies? Why
Orchestrated compute on demand
Reuse and sharing
Key features
Multi-tenant, hosted, serverless
Parametrized, data lineage, and caching
Additionally, if the run invokes a task that has already been computed before, regardless of who executed it, Flyte will smartly use the cached output, saving you both time and money.
Versioning, sharing
Modular, loosely coupled
Seems like you guys recognize that the best task for the job might be hosted elsewhere, so it was important to integrate other solutions into Flyte.
Flyte extensions
Backend plugins - is it true you can create and manage k8s resources like CRDs for things like Spark, Sagemaker, BigQuery?
Drop a Star
https://flyte.org
Flyte community
----------- Connect With Us ✌️-------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn: https://www.linkedin.com/in/aponteanalytics/Connect with Ketan on LinkedIn: https://www.linkedin.com/in/ketanumare/
Round 3: Analyzing the Google paper "Continuous Delivery and Automation Pipelines in ML"
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Show Notes
Data Science Steps for ML
Data extraction: You select and integrate the relevant data from various data sources for the ML task.
Data analysis: You perform exploratory data analysis (EDA) to understand the available data for building the ML model. This process leads to the following:
Understanding the data schema and characteristics that are expected by the model.
Identifying the data preparation and feature engineering that are needed for the model.
Data preparation: The data is prepared for the ML task. This preparation involves data cleaning, where you split the data into training, validation, and test sets. You also apply data transformations and feature engineering to the model that solves the target task. The output of these steps is the data splits in the prepared format.
Model training: The data scientist implements different algorithms with the prepared data to train various ML models. In addition, you subject the implemented algorithms to hyperparameter tuning to get the best-performing ML model. The output of this step is a trained model.
Model evaluation: The model is evaluated on a holdout test set to evaluate the model quality. The output of this step is a set of metrics to assess the quality of the model.
Model validation: The model is confirmed to be adequate for deployment, and its predictive performance is better than a certain baseline.
Model serving: The validated model is deployed to a target environment to serve predictions. This deployment can be one of the following:
Microservices with a REST API to serve online predictions.
An embedded model to an edge or mobile device.
Part of a batch prediction system.
Model monitoring: The model predictive performance is monitored to potentially invoke a new iteration in the ML process.
The level of automation of these steps defines the maturity of the ML process, which reflects the velocity of training new models given new data or training new models given new implementations. The following sections describe three levels of MLOps, starting from the most common level, which involves no automation, up to automating both ML and CI/CD pipelines.
In the rest of the conversation, we talk about maturity levels 0 and 1. Next session, we will talk about Level 2.
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn: https://www.linkedin.com/in/aponteanalytics/
MLOps community meetup #36! This week, we talk to David Hershey, Solutions Engineer at Determined AI, about Moving Deep Learning from Research to Production with Determined and Kubeflow.
// Key takeaways:
What components are needed to do inference in ML
How to structure models for ML inference
How a model registry helps organize your models for easy consumption
How you can set up reusable and easy-to-upgrade inference pipelines
// Abstract:
Translating the research that goes into creating a great deep learning model into a production application is a mess without the right tools. ML models have a lot of moving pieces, and on top of that, models are constantly evolving as new data arrives or the model is tweaked. In this talk, we'll show how you can find order in that chaos by using the Determined Model Registry along with Kubeflow Pipelines.
// Bio:
David Hershey is a solutions engineer for Determined AI. David has a passion for machine learning infrastructure, in particular systems that enable data scientists to spend more time innovating and changing the world with ML. Previously, David worked at Ford Motor Company as an ML Engineer, where he led the development of Ford's ML platform. He received his MS in Computer Science from Stanford University, where he focused on Artificial Intelligence and Machine Learning.
// Relevant Links
www.determined.ai
https://github.com/determined-ai/determined
https://determined.ai/blog/production-training-pipelines-with-determined-and-kubeflow/
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn:
https://www.linkedin.com/in/david-hershey-458ab081/
Timestamps:
0:00 - Intros
4:15 - The structure of the chat
5:20 - What is DeterminedAI?
7:20 - How is DeterminedAI different than other, more standard artifact storage solutions?
9:25 - Where are the boundaries between what your tool determined AI does really well, and where it works smoothly with other things around it?
11:48 - Is Kubeflow dying?
13:54 - How do you see DeterminedAI and Kubeflow becoming more solidified?
15:55 - How does DeterminedAI interact with Kubeflow at the moment?
18:01 - What type of models are they? Is the Kubeflow metadata?
19:18 - What is a model registry, and why is it so important to have that?
23:16 - Can you give us the quick demo real fast?
30:52 - Which orchestration tool to use?
32:04 - When using Kubeflow are determined how can you deploy the model through CD tools like Jenkins?
33:40 - How is it determined to be connected to Kubeflow?
36:09 - What components do you feel are needed to do inference in machine learning? And how can we structure different models for that machine learning inference?
40:04 - Are they the same ones when we talk about ML researchers?
42:14 - How can we better be ready for when we do want to get into production?
44:59 - In this pipeline, where do you normally see people getting stopped?
47:05 - What are things that you've seen pop up that you're not necessarily thinking about in those first phases?
50:17 - What are the most underrated topics regarding deploying machine learning models in production?
52:44 - How do you see the adoption of tools such as Determined and Kubeflow by Data scientists?
54:40 - Can you explain the Determined open source components?
Second installation, David and Demetrios are reviewing the Google paper about Continuous training and automated pipelines. They dive deep into machine learning monitoring and also what exactly continuous training actually entails. Some key highlights are:
Automatically retraining and serving the models:
When to do it?
Outlier detection
Drift detection
Outlier detection:
What is it?
How you deal with it
Drift detection
Individual features may start to drift. This could be a bug, or it could be perfectly normal behavior that indicates that the world has changed, requiring the model to be retrained.
Example changes:
shifts in people’s preferences
marketing campaigns
competitor moves
the weather
the news cycle
Locations
Time
Devices (clients)
If the world you're working with is changing over time, model deployment should be treated as a continuous process. What this tells me is that you should keep the data scientists and engineers working on the model instead of immediately moving to another project.
Deeper dive into concept drift
Feature/target distributions change
An overview of concept drift applications: “.. data analysis applications, data evolve over time and must be analyzed in near real time. Patterns and relations in such data often evolve over time; thus, models built for analyzing such data quickly become obsolete over time. In machine learning and data mining, this phenomenon is referred to as concept drift.”
https://www.win.tue.nl/~mpechen/publications/pubs/CD_applications15.pdf
https://www-ai.cs.tu-dortmund.de/LEHRE/FACHPROJEKT/SS12/paper/concept-drift/tsymbal2004.pdf
Types of concept drift:
Sudden
Gradual
Google, in some way, is trying to address this concern - the world is changing, and you want your ML system to change as well, so it can avoid decreased performance but also improve over time and adapt to its environment. This sort of robustness is necessary for certain domains.
Continuous delivery and automation of pipelines (data, training, prediction service) was built with this in mind. Minimizing the commit-to-deploy interval and maximizing the velocity of software delivery and its components: maintainability, extensibility, and testability
Then the pipeline is ready, you can now run it. So you can do this continuously. After the pipeline is deployed to the production environment, it will be executed automatically and repetitively to produce a trained model that is stored in a central model registry.
This pipeline should be able to be run on a schedule or based on triggers: certain events that you have configured for your business domain - new data or drop in performance from the prod model.
The link between the model artifact and the pipeline is never severed. What pipeline trained them? What data was extracted, validated, and how was it prepared? What was the training configuration, and how was it evaluated? Etc. metrics are key here! Lineage tracking!!!
Keeping a close tie between the dev/experiment pipeline and the continuous production pipeline helps avoid inconsistencies between model artifacts produced by the pipeline and models being served - hard to debug
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn: https://www.linkedin.com/in/aponteanalytics/
Connect with Cris Sterry on LinkedIn: https://www.linkedin.com/in/chrissterry/
MLOps Meetup #34! This week, we talk to Kai Waehner about the beast that is Apache Kafka and how many different ways you can use it!
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Key takeaways:
-Kafka is much more than just messaging
-Kafka is the de facto standard for processing huge volumes of data at scale in real-time
-Kafka and Machine Learning are complementary for various use cases (including data integration, data processing, model training, model scoring, and monitoring)
// Abstract:
The combination of Apache Kafka, tiered storage, and machine learning frameworks such as TensorFlow enables you to build a scalable, reliable, but also simple infrastructure for all machine learning tasks using the Apache Kafka ecosystem and Confluent Platform. This discussion features a predictive maintenance use case within a connected car infrastructure, but the discussed components and architecture are helpful in any industry.
// Bio:
Kai Waehner is a Technology Evangelist at Confluent. He works with customers across the globe and with internal teams like engineering and marketing. Kai’s main area of expertise lies within the fields of Big Data Analytics, Machine Learning, Hybrid Cloud Architectures, Event Stream Processing, and Internet of Things. He is a regular speaker at international conferences such as Devoxx, ApacheCon, and Kafka Summit, writes articles for professional journals, and shares his experiences with new technologies on his blog: www.kai-waehner.de.
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Kai: [email protected] / @KaiWaehner / LinkedIn (https://www.linkedin.com/in/megachucky/)
________Show Notes_______
Blogpost tiered storage
https://www.confluent.io/blog/streaming-machine-learning-with-tiered-storage/
https://www.confluent.io/resources/kafka-summit-2020/apache-kafka-tiered-storage-and-tensorflow-for-streaming-machine-learning-without-a-data-lake/
Blogpost about using Kafka as a database
https://www.kai-waehner.de/blog/2020/03/12/can-apache-kafka-replace-database-acid-storage-transactions-sql-nosql-data-lake/
Example repo on github
https://github.com/kaiwaehner/hivemq-mqtt-tensorflow-kafka-realtime-iot-machine-learning-training-inference
Model serving vs embedded Kafka
https://www.confluent.io/blog/machine-learning-real-time-analytics-models-in-kafka-applications/
https://www.confluent.io/kafka-summit-san-francisco-2019/event-driven-model-serving-stream-processing-vs-rpc-with-kafka-and-tensorflow/
Istio blog post
https://www.kai-waehner.de/blog/2019/09/24/cloud-native-apache-kafka-kubernetes-envoy-istio-linkerd-service-mesh/
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
While machine learning is spreading like wildfire, very little attention has been paid to the ways that it can go wrong when moving from development to production. Even when models work perfectly, they can be attacked and/or degrade quickly if the data changes. Having a well-understood MLOps process is necessary for ML security!
Using Kubeflow, we demonstrated how the common ways machine learning workflows go wrong, and how to mitigate them using MLOps pipelines to provide reproducibility, validation, versioning/tracking, and safe/compliant deployment. We also talked about the direction for MLOps as an industry, and how we can use it to move faster, with less risk, than ever before.
David leads Open Source Machine Learning Strategy at Azure. This means he spends most of his time helping humans convince machines to be smarter. He is only moderately successful at this. Previously, he led product management for Kubernetes on behalf of Google, launched Google Kubernetes Engine, and co-founded the Kubeflow project. He has also worked at Microsoft, Amazon, and Chef and co-founded three startups. When not spending too much time in the service of electrons, he can be found on a mountain (on skis), traveling the world (via restaurants), or participating in kid activities, of which there are a lot more than he remembers than when he was that age.
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn: https://www.linkedin.com/in/aronchick/
From the publisher's feed

1,289 Listeners

286 Listeners

1,089 Listeners

622 Listeners

582 Listeners

304 Listeners

338 Listeners

204 Listeners

561 Listeners

512 Listeners

141 Listeners

102 Listeners

222 Listeners

683 Listeners

30 Listeners