
Sign up to save your podcasts
Or


**Privacy-preserving ML with Differential Privacy**
Differential privacy is without a question one of the most innovative concepts that came around in the last decades, with a variety of different applications even when it comes to Machine Learning. Many are organizations already leveraging this technology to access and make sense of their most sensitive data, but what is it? How does it work? And how can we leverage it the most?
To explain this and provide us a brief intro on Differential Privacy, I've invited Christos Dimitrakakis. Professor at University, counts already with multiple publications (more than 1000!!!) in the areas of Machine Learning, Reinforcement Learning, and Privacy.
Useful links:
Christos Dimitrakakis list of publications
Differential privacy for Bayesian inference through posterior sampling
Differential privacy use cases
Open-source differential privacy projects
Open-source project for Differential Privacy in SQL databases
MLOps community meetup #44! Last Wednesday, we talked to Savin Goyal, Tech Lead for the ML Infra team at Netflix.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Abstract:
In this conversation, Savin talked about some of the challenges encountered and choices made by the Netflix ML Infrastructure team while developing tooling for data scientists.
// Bio:
Savin is an engineer on the ML Infrastructure team at Netflix. He focuses on building generalizable infrastructure to accelerate the impact of data science at Netflix.
// Other links to check on Savin:
https://www.usenix.org/conference/opml20/presentation/cepoi
https://www.youtube.com/watch?v=lakPlz8GJcA&ab_channel=RConsortium
https://www.youtube.com/watch?v=-oMZAS9qfrE&ab_channel=AnalyticsIndiaMagazine
https://www.youtube.com/watch?v=yyWirT279tY&ab_channel=FunctionalTV
https://www.youtube.com/watch?v=QkRJ24Q0E-k&ab_channel=Matroid
----------- Connect With Us ✌️-------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Savin on LinkedIn: https://www.linkedin.com/in/savingoyal/
Timestamps:
[00:00] Background of Savin Goyal
[02:41] Breakdown of Metaflow
[05:44] In the stack, where does Metaflow stand?
[13:23] Where does Metaflow start in the Runway Project?
[15:27] What tools or storage does Netflix use for DataOps, ie, the front-end management of data sets, and how does that integrate with Metaflow? [18:56] Recommender Systems: Can you explain the other areas that you're using Machine Learning in?
[22:27] What do you feel is the hardest part of building an operating Machine Learning workflow? [28:45] 3 Pillars: Reproducibility, Scalability, Usability.
[36:05] You give so much power to people. How do you keep them from going overboard?
[37:47] Can you explain this Pillar of Usability?
[41:09] Road-based access control has been coming up a lot recently. Does Metaflow do something specific for that?
[44:49] What are some learnings that come across that you didn't have since you open-sourced when you were working at Netflix?
[48:10] What kind of trends have you been seeing? Where do you feel like the market is going?
[50:33] Have you seen some companies really interested in Metaflow? How have you been seeing them combine other tools that are out there?
Coffee Sessions #21 with Benjamin Rogojan of Seattle Data Guy, A Conversation with Seattle Data Guy
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
//Bio
Ben has spent his career focused on all forms of data. He has focused on developing algorithms to detect fraud, reduce patient readmission, and redesign insurance provider policies to help reduce the overall cost of healthcare. He has also helped develop analytics for marketing and IT operations in order to optimize limited resources such as employees and budget. Ben privately consults on data science and engineering problems, both solo as well as with a company called Acheron Analytics. He has experience working both hands-on with technical problems as well as helping leadership teams develop strategies to maximize their data.
//Other links you can check Ben on
https://www.theseattledataguy.com/mlops-vs-aiops-what-is-the-difference/#page-content
https://medium.com/@benrogojan
https://www.kdnuggets.com/2020/01/data-science-interview-study-guide.html
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn: https://www.linkedin.com/in/aponteanalytics/
Connect with Ben on LinkedIn: https://www.linkedin.com/in/benjaminrogojan/
Timestamps
[00:00] Intro to Benjamin Rogojan
[01:22] Ben's background
[03:30] What are some of your learnings/key things that jumped out at you?
[08:15] Agile and Data Science
[10:28] Likelihood of failure
[13:05] Sometimes you have to wait
[15:11] Defining your data science process
[19:55] A Layer of communication is important between the data scientists and higher-ups
[21:29] How do you navigate challenges? Are there any tools or processes you quantify to work with your clients?
[24:30] How do you show the value of your work using monitoring and observability
[27:58] How can we be better communicators?
[31:15] Have you seen other roles that really helped the jell of the team?
[33:50] What are your interests? What are you passionate about at the moment?
[34:29] Is there something new you're learning at the moment?
[37:55] Do you have a process about how you figure out even data science or ML is right for a company?
[39:33] Do you have a blog about the process you follow?
[41:24] What is one negative wisdom that you want to share with the community?
[44:35] How did you come up with the company name Seattle Data Guy?
Links mentioned in this episode:
https://medium.com/@benrogojan
https://www.cprime.com/resources/blog/agile-methodologies-how-they-fit-into-data-science-processes/
https://www.coriers.com/the-data-science-interview-study-guide/
https://medium.com/@SeattleDataGuy/from-data-scientist-to-data-leader-workshop-c6be69698af
https://towardsdatascience.com/4-must-have-skills-every-data-scientist-should-learn-8ab3f23bc325
Coffee Sessions #20 with Neal Lathia of Monzo Bank, talking about Monzo Bank - An MLOps Case Study
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
//Bio
Neal is currently the Machine Learning Lead at Monzo in London, where his team focuses on building machine learning systems that optimise the app and help the company scale. Neal's work has always focused on applications that use machine learning - this has taken him from recommender systems to urban computing and travel information systems, digital health monitoring, smartphone sensors, and banking.
//Talk Takeaways
Monzo Bank has a small but very impactful team continuously learning new things. Optimistically do their utmost to avoid “throwing problems over the wall,” and so they build systems, iterate on machine learning models, and collaborate very closely with each other and with many folks across the business.
Hopefully, all of that paints a picture of a team that aims to bring real and valuable machine learning systems to life. Monzo does not spend time trying to advance the state-of-the-art in machine learning or tweak models to absolute perfection.
//Other links you can check Neal on
Personal Website: http://nlathia.github.io/
Research: http://nlathia.github.io/research/
Press & Speaking: http://nlathia.github.io/public/
http://nlathia.github.io/2020/06/Customer-service-machine-learning.html
http://nlathia.github.io/2020/10/ML-and-rule-engines.html
http://nlathia.github.io/2020/10/Monzo-ML.html http://nlathia.github.io/2019/09/Large-NLP-in-prod.html http://nlathia.github.io/2020/07/Shadow-mode-deployments.html https://github.com/operatorai
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with David on LinkedIn: https://www.linkedin.com/in/aponteanalytics/
Connect with Neal on LinkedIn: https://www.linkedin.com/in/nlathia/
Timestamps:
[00:00] Intro to Neal Lathia
[02:48] Background of Monzo Bank
[05:06] Problems you're solving with Machine Learning at Monzo?
[08:36] Why do you think it's fairly easy to frame a lot of problems using Machine Learning?
[11:56] How do you decide on rule-based or Machine learning?
[15:33] Team Structure
[19:18] What are some challenges, like size, latency, and the like?
[21:52] How have you addressed learning skills/challenges in your team?
[26:17] Do you have something that connects your team with all the metadata you have?
[27:14] Are you also having the monitoring models in your dashboard, or is that something else?
[28:51] Why should I bring another tool that the company is not familiar with when we already have one?
[31:43] Do you feel like there will be a point in time where you need to buy a tool because one problem is taking so much of your time?
[38:30] Engineering optimization teams for machine learning?
[40:34] Take us through the idea to production?
[46:29] How do you deal with reproducibility?
[49:48] Do you have ethics people on the team?
[54:12] Why are you using GCP and AWS?
[56:09] What are these different use cases, and how do they differ?
[57:57] How do you address applications that don't work?
**The intersection between DataOps and privacy**
DataOps is considered by many as the new era of data management, a set of principles that emphasizes communication, collaboration, integration, and automation of cooperation between the different teams in an organization that have to deal with data: data engineers, data scientists to data analysts.
But is there any relation between DataOps and data privacy protection? Can organizations leverage DataOps to ensure that their data is privacy compliant?
For this episode we've invited Lars Albertsson founder of Scling and former Data Engineer at Spotify, Lars has been educating organizations on how to get value from data and engineering efficiency!
**Are Privacy Enhancing Technologies a myth**
https://medium.com/@francis_49362/differential-privacy-not-a-complete-disaster-i-guess-d0345a76a5af
Facebook and DIfferential Privacy
Opacus
Synthetic data generation
Coffee Sessions #19 with Barr Moses of Monte Carlo, Introducing Data Downtime: How to Prevent Broken Data Pipelines with Observability, co-hosted by Vishnu Rachakonda.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Bio
Barr Moses is CEO & Co-Founder of Monte Carlo, a data observability company backed by Accel and other top Silicon Valley investors. Previously, she was VP Customer Operations at customer success company Gainsight, where she helped scale the company 10x in revenue and, among other functions, built the data/analytics team. Prior to that, she was a management consultant at Bain & Company and a research assistant at the Statistics Department at Stanford. She also served in the Israeli Air Force as a commander of an intelligence data analyst unit. Barr graduated from Stanford with a B.Sc. in Mathematical and Computational Science.
// Talk Takeaways
As companies become increasingly data-driven, the technologies underlying these rich insights have grown more and more nuanced and complex. While our ability to collect, store, aggregate, and visualize this data has largely kept up with the needs of modern data teams (think: domain-oriented data meshes, cloud warehouses, data visualization tools, and data modeling solutions), the mechanics behind data quality and integrity have lagged.
To keep pace with data’s clock speed of innovation, data engineers need to invest not only in the latest modeling and analytics tools but also in technologies that can increase data accuracy and prevent broken pipelines. The solution? Data observability, the next frontier of data engineering and a pillar of the emerging Data Reliability category, and the fix for eliminating data downtime.
During this talk, listeners will learn about:
// About Monte Carlo
As businesses increasingly rely on data to drive better decision-making, it’s mission-critical that this data is accurate and reliable. Billed by Forbes as the New Relic for data teams and backed by Accel and GGV, Monte Carlo solves the costly problem of broken data through their fully automated, end-to-end data reliability platform. Data teams spend north of 30% of their time tackling data quality issues, distracting data engineers, data scientists, and data analysts from working on revenue-generating projects. Providing full coverage of your data stack – all the way from data lake and warehouse to analytics dashboard – Monte Carlo’s platform empowers companies such as Eventbrite, Compass, Vimeo, and other enterprises to trust their data, saving time and money and unlocking the potential of data.
// Other links you can check Barr on
Learn more about Monte Carlo: https://www.montecarlodata.com
What is data downtime? https://www.montecarlodata.com/the-rise-of-data-downtime/
What is data observability? https://www.montecarlodata.com/data-observability-the-next-frontier-of-data-engineering/
How data observability prevents broken data pipelines: https://www.montecarlodata.com/data-observability-how-to-prevent-your-data-pipelines-from-breaking/
MLOps community meetup #43! Last Wednesday, we talked to Nathan Benaich, General Partner at Air Street Capital, and Timothy Chen, Managing Partner at Essence VC, about The MLOps Landscape.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Abstract:
In this session, we explored the MLOps landscape through the eyes of two accomplished investors. Tim and Nathan shared with us their experience in looking at hundreds of ML and MLOps companies each year to highlight major insights they have gained. What do the ML infrastructure and tooling landscape look like at the moment? Where have they been seeing patterns emerge? What do they expect to see happen within the market in the next couple of years? What current tools out there are the most interesting to them? And last but not least, how do they go about selecting which companies to invest in?
// Bio:
Nathan Benaich is the Founder and General Partner of Air Street Capital, a venture capital firm investing in early-stage AI-first technology and life science companies. The team’s investments include Mapillary (Acq. Facebook), Graphcore, Thought Machine, Tractable, and LabGenius. Nathan is Managing Trustee of The RAAIS Foundation, a non-profit with a mission to advance education and open-source research in the common good of AI. This includes running the annual RAAIS summit and funding fellowships at OpenMined. Nathan is also co-author of the annual State of AI Report. He holds a PhD in cancer biology from the University of Cambridge and a BA from Williams College.
Timothy Chen is the Managing Partner at Essence VC, with a decade of experience leading engineering in enterprise infra and open source communities/companies.
Prior to Essence, Tim was the SVP of Engineering at Cosmos, a popular open-source blockchain SDK. Prior to Cosmos, Tim cofounded Hyperpilot with Stanford Professor Christos Kozyrakis, which later exited to Cloudera. Prior to Hyperpilot, Tim was an early employee at Mesosphere and CloudFoundry.
Tim is also active in the open-source space as an Apache member.
----------- Connect With Us ✌️-------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Nathan on LinkedIn: https://www.linkedin.com/in/nathanbenaich/
Connect with Tim on LinkedIn: https://www.linkedin.com/in/timchen
Timestamps:
0:00 - Nathan Benaich & Timothy Chen
1:36 - Tim's background
4:07 - Nathan's background
8:08 - To Nathan: What's your take on the lay of the land in the MLOps fear or space?
10:20 - To Tim: Can you give us your rundown on what you've been seeing? The greater landscape that you look at.
14:35 - To Tim: What companies right now really excite you? What are some that are doing something that has a future?
19:36 - To Nathan: What kind of companies are you looking at right now that you're doing interesting things?
22:37 - The MLOps tools mature as the companies mature.
23:45 - No tool looks exactly the same from an MLOps perspective
25:44 - Sometimes MLOps tools are not a choice by data scientists at all.
28:10 - What MLOps needs that are not being addressed by the market right now?
35:00 - What is the annotation stack?
37:28 - How do you think about it in the context of federated learning?
41.24 - Will MLOps tools eventually become idiomatic? Would that be desirable?
47:55 - How do you switch from this open-source model to the money-making model?
52:30 - Should we focus only on the open-source at first and think about monetization later? If so, are investors prepared to invest in no-revenue companies?
**AI and ethical dilemmas**
MLOps community meetup #42! Last Wednesday, we talked to Mark Craddock, Co-Founder & CTO, Global Certification and Training Ltd (GCATI), about the UN Global Platform.
Join the Community: https://go.mlops.community/YTJoinIn
Get the newsletter: https://go.mlops.community/YTNewsletter
// Abstract:
Building a global big data platform for the UN. Streaming 600,000,000+ records/day into the platform. The strategy was developed using Wardley Maps and the Platform Design Toolkit.
// Bio:
Mark contributed to the Cloud First policy for the UK Public sector and was one of the founding architects for the UK Government's G-Cloud program. Mark developed the initial CloudStore, which enabled the UK Public Sector to procure cloud services from over 2,500 suppliers. The UK Public Sector has now purchased over £6.3Bn of cloud services, with £3.6Bn from Small to Medium Enterprises in the UK.
Mark led the development of the United Nations Global Platform. A multi-cloud platform for capacity building within the national statistics offices in the use of Big Data and its integration with administrative sources, geospatial information, traditional survey, and census data.
Mark is now building a non-profit training and certification organization.
----------- Connect With Us ✌️-------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Mark on LinkedIn: https://www.linkedin.com/in/markcraddock/
Timestamps:
[0:00] - Intro to Mark Craddock
[03:35] - Mark's background
[05:05] - UN Global Platform
[05:18] - Vision: A global collaboration to harness the power of data for better lives
[05:37] - UN GWG (Big) Data Membership
[05:49] - Sustainable Development Goals
[06:21] - Using the platform
[06:30] - Approach
[06:44] - Principles
[07:29] - How big was the team that put this together?
[08:09] - Leave no one behind. Endeavor to reach the furthest behind first.
[08:24] - Platform Business Model
[10:06] - Six distinct aspects of a platform and its ecosystem
[10:46] - The platform is the only business model able to orchestrate the wide range of products and services in an ecosystem
[11:09] - Through the means of a platform organization, ecosystems are capable of providing an improbable combination of attributes
[11:55] - Platforms and business models are also one of the best organizational structures for enabling rapid evolution
[13:22] - Technology Strategy
[13:23] - Wardley Maps
[14:50] - Is this where Machine Learning tools would fit in?
[20:35] - Are you looking at how fast these are moving across to the right? How can you gauge that?
[26:57] - Is the value fluid?
[28:43] - How did you factor in the different personas?
[30:34] - How do you enable loosely coupled teams?
[35:44] - Data also moves from left to right
[42:00] - Technology Strategy Handbook
[42:20] - Achievements - July '19
[42:31] - Global Billing Intelligence
[43:15] - Privacy-Preserving Techniques Handbook
[43:26] - Cryptographic Techniques
[44:12] - Global Big Datasets
[44:55] - Big Data
[47:41] - Automatic Identification System (AIS)
[48:14] - Automatic Dependent Surveillance (ADS-B)
[48:41] - Satellite Imagery
[49:11] - Services in the platform
[49:16] - Location Analytics Service
[50:06] - Stack Sample
[50:37] - Data Sources
[51:50] - NiFi Dataflow
[52:20] - Is this how you enabled reproducibility?
[53:47] - Location Analytics Service
[55:31] - Shanghai - Flights
[55:45] - Shanghai - Cargo Ships
[56:00] - UN Global Platform
From the publisher's feed

1,289 Listeners

286 Listeners

1,089 Listeners

622 Listeners

582 Listeners

304 Listeners

338 Listeners

204 Listeners

561 Listeners

512 Listeners

141 Listeners

102 Listeners

222 Listeners

683 Listeners

30 Listeners