Data Science at Home

Data Science at Home

By Francesco GadaletaNewsTechnologyTech News
Download on the App Store
  • Favorites

    158

    Followers

  • Typical duration

    18 min

    per episode

Based on Podcast App listening data

Data Science at Home episodes

  • How to generate very large images with GANs (Ep. 76)

    Join the discussion on our Discord server

    In this episode I explain how a research group from the University of Lubeck dominated the curse of dimensionality for the generation of large medical images with GANs.

    The problem is not as trivial as it seems. Many researchers have failed in generating large images with GANs before. One interesting application of such approach is in medicine for the generation of CT and X-ray images.
    Enjoy the show!

     

    References

    Multi-scale GANs for Memory-efficient Generation of High Resolution Medical Images https://arxiv.org/abs/1907.01376

    15 min
  • [RB] Complex video analysis made easy with Videoflow (Ep. 75)

    In this episode I am with Jadiel de Armas, senior software engineer at Disney and author of Videflow, a Python framework that facilitates the quick development of complex video analysis applications and other series-processing based applications in a multiprocessing environment. 

    I have inspected the videoflow repo on Github and some of the capabilities of this framework and I must say that it’s really interesting. Jadiel is going to tell us a lot more than what you can read from Github 

     

    References

    Videflow Github official repository

    https://github.com/videoflow/videoflow

     

    31 min
  • [RB] Validate neural networks without data with Dr. Charles Martin (Ep. 74)

    In this episode, I am with Dr. Charles Martin from Calculation Consulting a machine learning and data science consulting company based in San Francisco. We speak about the nuts and bolts of deep neural networks and some impressive findings about the way they work. 

    The questions that Charles answers in the show are essentially two:

    1. Why is regularisation in deep learning seemingly quite different than regularisation in other areas on ML?

  • How can we dominate DNN in a theoretically principled way?
  •  

    References 
    • The WeightWatcher tool for predicting the accuracy of Deep Neural Networks https://github.com/CalculatedContent/WeightWatcher

  • Slack channel https://weightwatcherai.slack.com/

  • Dr. Charles Martin Blog http://calculatedcontent.com and channel https://www.youtube.com/c/calculationconsulting

  • Implicit Self-Regularization in Deep Neural Networks: Evidence from Random Matrix Theory and Implications for Learning - Charles H. Martin, Michael W. Mahoney
    45 min
  • How to cluster tabular data with Markov Clustering (Ep. 73)

    In this episode I explain how a community detection algorithm known as Markov clustering can be constructed by combining simple concepts like random walks, graphs, similarity matrix. Moreover, I highlight how one can build a similarity graph and then run a community detection algorithm on such graph to find clusters in tabular data.

    You can find a simple hands-on code snippet to play with on the Amethix Blog 

    Enjoy the show! 

     

    References

    [1] S. Fortunato, “Community detection in graphs”, Physics Reports, volume 486, issues 3-5, pages 75-174, February 2010.

    [2] Z. Yang, et al., “A Comparative Analysis of Community Detection Algorithms on Artificial Networks”, Scientific Reports volume 6, Article number: 30750 (2016)

    [3] S. Dongen, “A cluster algorithm for graphs”, Technical Report, CWI (Centre for Mathematics and Computer Science) Amsterdam, The Netherlands, 2000.

    [4] A. J. Enright, et al., “An efficient algorithm for large-scale detection of protein families”, Nucleic Acids Research, volume 30, issue 7, pages 1575-1584, 2002.

    21 min
  • Waterfall or Agile? The best methodology for AI and machine learning (Ep. 72)

    The two most widely considered software development models in modern project management are, without any doubt, the Waterfall Methodology and the Agile Methodology. In this episode I make a comparison between the two and explain what I believe is the best choice for your machine learning project.

    An interesting post to read (mentioned in the episode) is How businesses can scale Artificial Intelligence & Machine Learning https://amethix.com/how-businesses-can-scale-artificial-intelligence-machine-learning/

    15 min
  • Training neural networks faster without GPU (Ep. 71)

    Training neural networks faster usually involves the usage of powerful GPUs. In this episode I explain an interesting method from a group of researchers from Google Brain, who can train neural networks faster by squeezing the hardware to their needs and making the training pipeline more dense.

    Enjoy the show!

     
    References

    Faster Neural Network Training with Data Echoing

    https://arxiv.org/abs/1907.05550

    23 min
  • Validate neural networks without data with Dr. Charles Martin (Ep. 70)

    In this episode, I am with Dr. Charles Martin from Calculation Consulting a machine learning and data science consulting company based in San Francisco. We speak about the nuts and bolts of deep neural networks and some impressive findings about the way they work. 

    The questions that Charles answers in the show are essentially two:

    1. Why is regularisation in deep learning seemingly quite different than regularisation in other areas on ML?

  • How can we dominate DNN in a theoretically principled way?
  •  

    References 
    • The WeightWatcher tool for predicting the accuracy of Deep Neural Networks https://github.com/CalculatedContent/WeightWatcher

  • Slack channel https://weightwatcherai.slack.com/

  • Dr. Charles Martin Blog http://calculatedcontent.com and channel https://www.youtube.com/c/calculationconsulting

  • Implicit Self-Regularization in Deep Neural Networks: Evidence from Random Matrix Theory and Implications for Learning - Charles H. Martin, Michael W. Mahoney

     

    45 min
  • Complex video analysis made easy with Videoflow (Ep. 69)

    In this episode I am with Jadiel de Armas, senior software engineer at Disney and author of Videflow, a Python framework that facilitates the quick development of complex video analysis applications and other series-processing based applications in a multiprocessing environment. 

    I have inspected the videoflow repo on Github and some of the capabilities of this framework and I must say that it’s really interesting. Jadiel is going to tell us a lot more than what you can read from Github 

     

    References

    Videflow Github official repository

    https://github.com/videoflow/videoflow

     

    31 min
  • Episode 68: AI and the future of banking with Chris Skinner [RB]

    In this episode I have a wonderful conversation with Chris Skinner.

    Chris and I recently got in touch at The banking scene 2019, fintech conference recently held in Brussels. During that conference he talked as a real trouble maker - that’s how he defines himself - saying that “People are not educated with loans, credit, money” and that “Banks are failing at digital”.

    After I got my hands on his last book Digital Human, I invited him to the show to ask him a few questions about innovation, regulation and technology in finance.

    42 min
  • Episode 67: Classic Computer Science Problems in Python

    Today I am with David Kopec, author of Classic Computer Science Problems in Python, published by Manning Publications.

    His book deepens your knowledge of problem solving techniques from the realm of computer science by challenging you with interesting and realistic scenarios, exercises, and of course algorithms.

    There are examples in the major topics any data scientist should be familiar with, for example search, clustering, graphs, and much more.

    Get the book from https://www.manning.com/books/classic-computer-science-problems-in-python and use coupon code poddatascienceathome19 to get 40% discount.

     

    References

    Twitter https://twitter.com/davekopec

    GitHub https://github.com/davecom

    classicproblems.com

    29 min

About Data Science at Home

From the publisher's feed

Cutting through AI bullsh*t.
Come join the discussion on Discord!
https://discord.gg/4UNKGf3

Best of Data Science at Home

Ranked by our users in the last 21 days

More shows like Data Science at Home

On Point with Meghna Chakrabarti by WBUR

On Point with Meghna Chakrabarti

4,028 Listeners

Making Sense with Sam Harris by Sam Harris

Making Sense with Sam Harris

26,249 Listeners

Nature Podcast by Springer Nature Limited

Nature Podcast

766 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

623 Listeners

Science Vs by Spotify Studios

Science Vs

12,188 Listeners

Science Friday by Science Friday and WNYC Studios

Science Friday

6,431 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

305 Listeners

The Daily by The New York Times

The Daily

111,799 Listeners

Up First from NPR by NPR

Up First from NPR

56,447 Listeners

The Atlantic Interview by The Atlantic

The Atlantic Interview

20 Listeners

Modern Wisdom by Chris Williamson

Modern Wisdom

4,095 Listeners

The Peter Attia Drive by Peter Attia, MD

The Peter Attia Drive

8,001 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

203 Listeners

Consider This from NPR by NPR

Consider This from NPR

6,373 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,882 Listeners