Data Science at Home

Data Science at Home

By Francesco GadaletaNewsTechnologyTech News
Download on the App Store
  • Favorites

    158

    Followers

  • Typical duration

    18 min

    per episode

Based on Podcast App listening data

Data Science at Home episodes

  • Compressing deep learning models: rewinding (Ep.105)

    As a continuation of the previous episode in this one I cover the topic about compressing deep learning models and explain another simple yet fantastic approach that can lead to much smaller models that still perform as good as the original one.

    Don't forget to join our Slack channel and discuss previous episodes or propose new ones.

    This episode is supported by Pryml.io

    Pryml is an enterprise-scale platform to synthesise data and deploy applications built on that data back to a production environment.

     

    References

    Comparing Rewinding and Fine-tuning in Neural Network Pruning

    https://arxiv.org/abs/2003.02389

     

    16 min
  • Compressing deep learning models: distillation (Ep.104)

    Using large deep learning models on limited hardware or edge devices is definitely prohibitive. There are methods to compress large models by orders of magnitude and maintain similar accuracy during inference.

    In this episode I explain one of the first methods: knowledge distillation

     Come join us on Slack

    Reference
    • Distilling the Knowledge in a Neural Network https://arxiv.org/abs/1503.02531
  • Knowledge Distillation and Student-Teacher Learning for Visual Intelligence: A Review and New Outlooks https://arxiv.org/abs/2004.05937
  • 23 min
  • Pandemics and the risks of collecting data (Ep. 103)

    Codiv-19 is an emergency. True. Let's just not prepare for another emergency about privacy violation when this one is over.

     

    Join our new Slack channel

     

    This episode is supported by Proton. You can check them out at protonmail.com or protonvpn.com

    21 min
  • Why average can get your predictions very wrong (ep. 102)

    Whenever people reason about probability of events, they have the tendency to consider average values between two extremes. 

    In this episode I explain why such a way of approximating is wrong and dangerous, with a numerical example.

    We are moving our community to Slack. See you there!

     

     

    15 min
  • Activate deep learning neurons faster with Dynamic RELU (ep. 101)

    In this episode I briefly explain the concept behind activation functions in deep learning. One of the most widely used activation function is the rectified linear unit (ReLU). 

    While there are several flavors of ReLU in the literature, in this episode I speak about a very interesting approach that keeps computational complexity low while improving performance quite consistently.

    This episode is supported by pryml.io. At pryml we let companies share confidential data. Visit our website.

    Don't forget to join us on discord channel to propose new episode or discuss the previous ones. 

    References

    Dynamic ReLU https://arxiv.org/abs/2003.10027

    23 min
  • WARNING!! Neural networks can memorize secrets (ep. 100)

    One of the best features of neural networks and machine learning models is to memorize patterns from training data and apply those to unseen observations. That's where the magic is. 

    However, there are scenarios in which the same machine learning models learn patterns so well such that they can disclose some of the data they have been trained on. This phenomenon goes under the name of unintended memorization and it is extremely dangerous.

    Think about a language generator that discloses the passwords, the credit card numbers and the social security numbers of the records it has been trained on. Or more generally, think about a synthetic data generator that can disclose the training data it is trying to protect. 

    In this episode I explain why unintended memorization is a real problem in machine learning. Except for differentially private training there is no other way to mitigate such a problem in realistic conditions.

    At Pryml we are very aware of this. Which is why we have been developing a synthetic data generation technology that is not affected by such an issue.

     

    This episode is supported by Harmonizely. 

    Harmonizely lets you build your own unique scheduling page based on your availability so you can start scheduling meetings in just a couple minutes.
    Get started by connecting your online calendar and configuring your meeting preferences.
    Then, start sharing your scheduling page with your invitees!

     

    References

    The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks

    https://www.usenix.org/conference/usenixsecurity19/presentation/carlini

    25 min
  • Attacks to machine learning model: inferring ownership of training data (Ep. 99)

    In this episode I explain a very effective technique that allows one to infer the membership of any record at hand to the (private) training dataset used to train the target model. The effectiveness of such technique is due to the fact that it works on black-box models of which there is no access to the data used for training, nor model parameters and hyperparameters. Such a scenario is very realistic and typical of machine learning as a service APIs. 

    This episode is supported by pryml.io, a platform I am personally working on that enables data sharing without giving up confidentiality. 

     

    As promised below is the schema of the attack explained in the episode.

     

     

    References

    Membership Inference Attacks Against Machine Learning Models

     

     

    20 min
  • Don't be naive with data anonymization (Ep. 98)

    Masking, obfuscating, stripping, shuffling. 

    All the above techniques try to do one simple thing: keeping the data private while sharing it with third parties. Unfortunately, they are not the silver bullet to confidentiality. 

    All the players in the synthetic data space rely on simplistic techniques that are not secure, might not be compliant and risky for production.
    At pryml we do things differently. 

    14 min
  • Building reproducible machine learning in production (Ep. 96)

    Building reproducible models is essential for all those scenarios in which the lead developer is collaborating with other team members. Reproducibility in machine learning shall not be an art, rather it should be achieved via a methodical approach. 

    In this episode I give a few suggestions about how to make your ML models reproducible and keep your workflow as smooth.

    Enjoy the show!


    Come visit us on our discord channel and have a chat

    15 min

About Data Science at Home

From the publisher's feed

Cutting through AI bullsh*t.
Come join the discussion on Discord!
https://discord.gg/4UNKGf3

Best of Data Science at Home

Ranked by our users in the last 21 days

More shows like Data Science at Home

On Point with Meghna Chakrabarti by WBUR

On Point with Meghna Chakrabarti

4,024 Listeners

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,245 Listeners

Nature Podcast by Springer Nature Limited

Nature Podcast

766 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

623 Listeners

Science Vs by Spotify Studios

Science Vs

12,186 Listeners

Science Friday by Science Friday and WNYC Studios

Science Friday

6,434 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

305 Listeners

The Daily by The New York Times

The Daily

111,766 Listeners

Up First from NPR by NPR

Up First from NPR

56,424 Listeners

The Atlantic Interview by The Atlantic

The Atlantic Interview

21 Listeners

Modern Wisdom by Chris Williamson

Modern Wisdom

4,095 Listeners

The Peter Attia Drive by Peter Attia, MD

The Peter Attia Drive

7,997 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

202 Listeners

Consider This from NPR by NPR

Consider This from NPR

6,371 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,872 Listeners