DataTalks.Club

DataTalks.Club

By DataTalks.ClubTechnology
Download on the App Store

DataTalks.Club episodes

  • Redefining AI Infrastructure: Open-Source, Chips, and the Future Beyond Kubernetes – Andrey Cheptsov

    In this podcast episode, we talked with Andrey Cheptsov about ​The future of AI infrastructure.


    About the Speaker:

    Andrey Cheptsov is the founder and CEO of dstack, an open-source alternative to Kubernetes and Slurm, built to simplify the orchestration of AI infrastructure. Before dstack, Andrey worked at JetBrains for over a decade helping different teams make the best developer tools.

    During the event, the guest, Andrey Cheptsov, founder and CEO of dstack, discussed the complexities of AI infrastructure. We explore topics like the challenges of using Kubernetes for AI workloads, the need to rethink container orchestration, and the future of hybrid and cloud-only infrastructures. Andrey also shares insights into the role of on-premise and bare-metal solutions, edge computing, and federated learning.

    00:00 Andrey's Career Journey: From JetBrains to DStack

    5:00 The Motivation Behind DStack

    7:00 Challenges in Machine Learning Infrastructure

    10:00 Transitioning from Cloud to On-Prem Solutions

    14:30 Reflections on OpenAI's Evolution

    17:30 Open Source vs Proprietary Models: A Balanced Perspective

    21:01 Monolithic vs. Decentralized AI businesses

    22:05 The role of privacy and control in AI for industries like banking and healthcare

    30:00 Challenges in training large AI models: GPUs and distributed systems

    37:03 DeepSpeed's efficient training approach vs. brute force methods

    39:00 Challenges for small and medium businesses: hosting and fine-tuning models

    47:01 Managing Kubernetes challenges for AI teams

    52:00 Hybrid vs. cloud-only infrastructure

    56:03 On-premise vs. bare-metal solutions

    58:05 Exploring edge computing and its challenges


    🔗 CONNECT WITH ANDREY CHEPTSOV

    Twitter -  / andrey_cheptsov  

    Linkedin -  / andrey-cheptsov  

    GitHub - https://github.com/dstackai/dstack/

    Website - https://dstack.ai/


    🔗 CONNECT WITH DataTalksClub

    Join DataTalks.Club:⁠⁠⁠https://datatalks.club/slack.html⁠⁠⁠

    Our events:⁠⁠⁠https://datatalks.club/events.html⁠⁠⁠

    Datalike Substack -⁠⁠⁠https://datalike.substack.com/⁠⁠⁠

    LinkedIn:⁠⁠⁠  / datatalks-club  ⁠

    57 min
  • Linguistics and Fairness - Tamara Atanasoska

    In this podcast episode, we talked with Tamara Atanasoska about ​building fair AI systems.


    About the Speaker:​Tamara works on ML explainability, interpretability and fairness as Open Source Software Engineer at probable. She is a maintainer of fairlearn, contributor to scikit-learn and skops. Tamara has both computer science/ software engineering and a computational linguistics(NLP) background.During the event, the guest discussed their career journey from software engineering to open-source contributions, focusing on explainability in AI through Scikit-learn and Fairlearn. They explored fairness in AI, including challenges in credit loans, hiring, and decision-making, and emphasized the importance of tools, human judgment, and collaboration. The guest also shared their involvement with PyLadies and encouraged contributions to Fairlearn.

    00:00 Introduction to the event and the community

    01:51 Topic introduction: Linguistic fairness and socio-technical perspectives in AI

    02:37 Guest introduction: Tamara’s background and career

    03:18 Tamara’s career journey: Software engineering, music tech, and computational linguistics

    09:53 Tamara’s background in language and computer science

    14:52 Exploring fairness in AI and its impact on society

    21:20 Fairness in AI models26:21 Automating fairness analysis in models

    32:32 Balancing technical and domain expertise in decision-making

    37:13 The role of humans in the loop for fairness

    40:02 Joining Probable and working on open-source projects

    46:20 Scopes library and its integration with Hugging Face

    50:48 PyLadies and community involvement

    55:41 The ethos of Scikit-learn and Fairlearn


    🔗 CONNECT WITH TAMARA ATANASOSKA

    Linkedin - https://www.linkedin.com/in/tamaraatanasoska

    GitHub- https://github.com/TamaraAtanasoska


    🔗 CONNECT WITH DataTalksClub

    Join DataTalks.Club:⁠⁠https://datatalks.club/slack.html⁠⁠

    Our events:⁠⁠https://datatalks.club/events.html⁠⁠

    Datalike Substack -⁠⁠https://datalike.substack.com/⁠⁠

    LinkedIn:⁠⁠  / datatalks-club  


    54 min
  • Career choices, transitions and promotions in and out of tech - Agita Jaunzeme

    In this podcast episode, we talked with Agita Jaunzeme about Career choices, transitions and promotions in and out of tech.


    About the Speaker:

    Agita has designed a career spanning DevOps/DataOps engineering, management, community building, education, and facilitation. She has worked on projects across corporate, startup, open source, and non-governmental sectors. Following her passion, she founded an NGO focusing on the inclusion of expats and locals in Porto. Embodying the values of innovation, automation, and continuous learning, Agita provides practical insights on promotions, career pivots, and aligning work with passion and purpose.


    During this event, discussed their career journey, starting with their transition from art school to programming and later into DevOps, eventually taking on leadership roles. They explored the challenges of burnout and the importance of volunteering, founding an NGO to support inclusion, gender equality, and sustainability. The conversation also covered key topics like mentorship, the differences between data engineering and data science, and the dynamics of managing volunteers versus employees. Additionally, the guest shared insights on community management, developer relations, and the importance of product vision and team collaboration.

    0:00 Introduction and Welcome
    1:28 Guest Introduction: Agita’s Background and Career Highlights
    3:05 Transition to Tech: From Art School to Programming
    5:40 Exploring DevOps and Growing into Leadership Roles
    7:24 Burnout, Volunteering, and Founding an NGO
    11:00 Volunteering and Mentorship Initiatives
    14:00 Discovering Programming Skills and Early Career Challenges
    15:50 Automating Work Processes and Earning a Promotion
    19:00 Transitioning from DevOps to Volunteering and Project Management
    24:00 Managing Volunteers vs. Employees and Building Organizational Skills
    31:07 Personality traits in engineering vs. data roles
    33:14 Differences in focus between data engineers and data scientists
    36:24 Transitioning from volunteering to corporate work
    37:38 The role and responsibilities of a community manager
    39:06 Community management vs. developer relations activities
    41:01 Product vision and team collaboration
    43:35 Starting an NGO and legal processes
    46:13 NGO goals: inclusion, gender equality, and sustainability
    49:02 Community meetups and activities
    51:57 Living off-grid in a forest and sustainability
    55:02 Unemployment party and brainstorming session
    59:03 Unemployment party: the process and structure


    🔗 CONNECT WITH AGITA JAUNZEME

    Linkedin - /agita


    🔗 CONNECT WITH DataTalksClub

    Join DataTalks.Club: ⁠https://datatalks.club/slack.html⁠
    Our events: ⁠https://datatalks.club/events.html⁠
    Datalike Substack - ⁠https://datalike.substack.com/⁠
    LinkedIn: ⁠  / datatalks-club  

    56 min
  • Career advice, learning, and featuring women in ML and AI - Isabella Bicalho

    In this podcast episode, we talked with Isabella Bicalho about Career advice, learning, and featuring women in ML and AI.


    About the Speaker:

    Isabella is a Machine Learning Engineer and Data Scientist with three years of hands-on AI development experience. She draws upon her early computational research expertise to develop ML solutions. While contributing to open-source projects, she runs a newsletter dedicated to showcasing women's accomplishments in data science.


    During this event, the guest discussed her transition into machine learning, her freelance work in AI, and the growing AI scene in France. She shared insights on freelancing versus full-time work, the value of open-source contributions, and developing both technical and soft skills. The conversation also covered career advice, mentorship, and her Substack series on women in data science, emphasizing leadership, motivation, and career opportunities in tech.

    0:00 Introduction
    1:23 Background of Isabella Bicalho
    2:02 Transition to machine learning
    4:03 Study and work experience
    5:00 Living in France and language learning
    6:03 Internship experience
    8:45 Focus areas of Inria
    9:37 AI development in France
    10:37 Current freelance work
    11:03 Freelancing in machine learning
    13:31 Moving from research to freelancing
    14:03 Freelance vs. full-time data science
    17:00 Finding first freelance client
    18:00 Involvement in open-source projects
    20:17 Passion for open-source and teamwork
    23:52 Starting new projects
    25:03 Community project experience
    26:02 Teaching and learning
    29:04 Contributing to open-source projects
    32:05 Open-source tools vs. projects
    33:32 Importance of community-driven projects
    34:03 Learning resources
    36:07 Green space segmentation project
    39:02 Developing technical and soft skills
    40:31 Gaining insights from industry experts
    41:15 Understanding data science roles
    41:31 Project challenges and team dynamics
    42:05 Turnover in open-source projects
    43:05 Managing expectations in open-source work
    44:50 Mentorship in projects
    46:17 Role of AI tools in learning
    47:59 Overcoming learning challenges
    48:52 Discussion on substack
    49:01 Interview series on women in data
    50:15 Insights from women in data science
    51:20 Impactful stories from substack
    53:01 Leadership challenges in projects
    54:19 Career advice and opportunities
    56:07 Motivating others to step out of comfort zone
    57:06 Contacting for substack story sharing
    58:00 Closing remarks and connections


    🔗 CONNECT WITH ISABELLA BICALHO

    Github: github https://github.com/bellabf
    LinkedIn:   / isabella-frazeto  


    🔗 CONNECT WITH DataTalksClub

    Join DataTalks.Club: https://datatalks.club/slack.html
    Our events: https://datatalks.club/events.html
    Datalike Substack - https://datalike.substack.com/
    LinkedIn:   / datatalks-club  

    55 min
  • AI in Industry: Trust, Return on Investment and Future - Maria Sukhareva

    Reflection on an Almost Two-Year Journey of Generative AI in Industry – Maria Sukhareva

    ​About the speaker:

    ​Maria Sukhareva is a principal key expert in Artificial Intelligence in Siemens with over 15 years of experience at the forefront of generative AI technologies. Known for her keen eye for technological innovation, Maria excels at transforming cutting-edge AI research into practical, value-driven tools that address real-world needs. Her approach is both hands-on and results-focused, with a commitment to creating scalable, long-term solutions that improve communication, streamline complex processes, and empower smarter decision-making. Maria's work reflects a balanced vision, where the power of innovation is met with ethical responsibility, ensuring that her AI projects deliver impactful and production-ready outcomes.


    We talked about:

    00:00 DataTalks.Club intro

    02:13 Career journey: From linguistics to AI

    08:02 The Evolution of AI Expertise and its Future

    13:10 AI vulnerabilities: Bypassing bot restrictions

    17:00 Non-LLM classifiers as a more robust solution

    22:56 Risks of chatbot deployment: Reputational and financial

    27:13 The role of AI as a tool, not a replacement for human workers

    31:41 The role of human translators in the age of AI

    34:49 Evolution of English and its Germanic roots

    38:44 Beowulf and Old English

    39:43 Impact of the Norman occupation on English grammar

    42:34 Identifying mushrooms with AI apps and safety precautions

    45:08 Decoding ancient languages ​​like Sumerian

    49:43 The evolution of machine translation and multilingual models

    53:01 Challenges with low-resource languages ​​and inconsistent orthography

    57:28 Transition from academia to industry in AI


    Join our Slack: https://datatalks.club/slack.html

    Our events: https://datatalks.club/events.html

    53 min
  • Large Hadron Collider and Mentorship – Anastasia Karavdina

    We talked about:

    00:00 DataTalks.Club intro

    00:00 Large Hadron Collider and Mentorship

    02:35 Career overview and transition from physics to data science

    07:02 Working at the Large Hadron Collider

    09:19 How particles collide and the role of detectors

    11:03 Data analysis challenges in particle physics and data science similarities

    13:32 Team structure at the Large Hadron Collider

    20:05 Explaining the connection between particle physics and data science

    23:21 Software engineering practices in particle physics

    26:11 Challenges during interviews for data science roles

    29:30 Mentoring and offering advice to job seekers

    40:03 The STAR method and its value in interviews

    50:32 Paid vs unpaid mentorship and finding the right fit


    ​About the speaker:

    ​Anastasia is a particle physicist turned data scientist, with experience in large-scale experiments like those at the Large Hadron Collider. She also worked at Blue Yonder, scaling AI-driven solutions for global supply chain giants, and at Kaufland e-commerce, focusing on NLP and search. Anastasia is a mentor for Ml/AI, dedicated to helping her mentees achieve their goals. She is passionate about growing the next generation of data science elite in Germany: from Data Analysts up to ML Engineers.


    Join our Slack: https://datatalks .club/slack.html

    55 min
  • MLOps as a Team - Raphaël Hoogvliets

    We talked about:

    00:00 DataTalks.Club intro

    02:34 Career journey and transition into MLOps

    08:41 Dutch agriculture and its challenges

    10:36 The concept of "technical debt" in MLOps

    13:37 Trade-offs in MLOps: moving fast vs. doing things right

    14:05 Building teams and the role of coordination in MLOps

    16:58 Key roles in an MLOps team: evangelists and tech translators

    23:01 Role of the MLOps team in an organization

    25:19 How MLOps teams assist product teams

    27 :56 Standardizing practices in MLOps

    32:46 Getting feedback and creating buy-in from data scientists

    36:55 The importance of addressing pain points in MLOps

    39:06 Best practices and tools for standardizing MLOps processes

    42:31 Value of data versioning and reproducibility

    44:22 When to start thinking about data versioning

    45:10 Importance of data science experience for MLOps

    46:06 Skill mix needed in MLOps teams

    47:33 Building a diverse MLOps team

    48:18 Best practices for implementing MLOps in new teams

    49:52 Starting with CI/CD in MLOps

    51:21 Key components for a complete MLOps setup

    53:08 Role of package registries in MLOps

    54:12 Using Docker vs. packages in MLOps

    57:56 Examples of MLOps success and failure stories

    1:00:54 What MLOps is in simple terms

    1:01:58 The complexity of achieving easy deployment, monitoring, and maintenance


    Join our Slack: https://datatalks .club/slack.html

    56 min
  • Using Data to Create Liveable Cities - Rachel Lim

    We talked about:

    00:00 DataTalks.Club intro

    01:56 Using data to create livable cities
    02:52 Rachel's career journey: from geography to urban data science
    04:20 What does a transport scientist do?
    05:34 Short-term and long-term transportation planning
    06:14 Data sources for transportation planning in Singapore
    08:38 Rachel's motivation for combining geography and data science
    10:19 Urban design and its connection to geography
    13:12 Defining a livable city
    15:30 Livability of Singapore and urban planning
    18:24 Role of data science in urban and transportation planning
    20:31 Predicting travel patterns for future transportation needs
    22:02 Data collection and processing in transportation systems
    24:02 Use of real-time data for traffic management
    27:06 Incorporating generative AI into data engineering
    30:09 Data analysis for transportation policies
    33:19 Technologies used in text-to-SQL projects
    36:12 Handling large datasets and transportation data in Singapore
    42:17 Generative AI applications beyond text-to-SQL
    45:26 Publishing public data and maintaining privacy
    45:52 Recommended datasets and projects for data engineering beginners
    49:16 Recommended resources for learning urban data science


    About the speaker:

    Rachel is an urban data scientist dedicated to creating liveable cities through the innovative use of data. With a background in geography, and a masters in urban data science, she blends qualitative and quantitative analysis to tackle urban challenges. Her aim is to integrate data driven techniques with urban design to foster sustainable and equitable urban environments. 


    Links: - https://datamall.lta.gov.sg/content/datamall/en/dynamic-data.html

    00:00 DataTalks.Club intro
    01:56 Using data to create livable cities
    02:52 Rachel's career journey: from geography to urban data science
    04:20 What does a transport scientist do?
    05:34 Short-term and long-term transportation planning
    06:14 Data sources for transportation planning in Singapore
    08:38 Rachel's motivation for combining geography and data science
    10:19 Urban design and its connection to geography
    13:12 Defining a livable city
    15:30 Livability of Singapore and urban planning
    18:24 Role of data science in urban and transportation planning
    20:31 Predicting travel patterns for future transportation needs
    22:02 Data collection and processing in transportation systems
    24:02 Use of real-time data for traffic management
    27:06 Incorporating generative AI into data engineering
    30:09 Data analysis for transportation policies
    33:19 Technologies used in text-to-SQL projects
    36:12 Handling large datasets and transportation data in Singapore
    42:17 Generative AI applications beyond text-to-SQL
    45:26 Publishing public data and maintaining privacy
    45:52 Recommended datasets and projects for data engineering beginners
    49:16 Recommended resources for learning urban data science
    Join our slack: https: //datatalks.club/slack.html

    46 min
  • DataTalks.Club 4th Anniversary AMA Podcast – Alexey Grigorev and Johanna Bayer

    We talked about:

    00:00 DataTalks.Club intro

    00:00 DataTalks.Club anniversary "Ask Me Anything" event with Alexey Grigorev

    02:29 The founding of DataTalks .Club

    03:52 Alexey's transition from Java work to DataTalks.Club

    04:58 Growth and success of DataTalks.Club courses

    12:04 Motivation behind creating a free-to-learn community

    24:03 Staying updated in data science through pet projects

    26 :37 Hosting a second podcast and maintaining programming skills

    28:56 Skepticism about LLMs and their relevance

    31:53 Transitioning to DataTalks.Club and personal reflections

    33:32 Memorable moments and the first event's success

    36:19 Community building during the pandemic

    38:31 AI's impact on data analysts and future roles

    42:24 Discussion on AI in healthcare

    44:37 Age and reflections on personal milestones

    47:54 Building communities and personal connections

    49:34 Future goals for the community and courses

    51:18 Community involvement and engagement strategies

    53:46 Ideas for competitions and hackathons

    54:20 Inviting guests to the podcast

    55:29 Course updates and future workshops

    56:27 Podcast preparation and research process

    58:30 Career opportunities in data science and transitioning fields

    1:01 :10 Book recommendations and personal reading experiences


    About the speaker:

    Alexey Grigorev is the founder of DataTalks.Club.


    Join our slack: https://datatalks.club/slack.html

    54 min
  • Human-Centered AI for Disordered Speech Recognition - Katarzyna Foremniak

    We talked about:

    00:00 DataTalks.Club intro

    08:06 Background and career journey of Katarzyna

    09:06 Transition from linguistics to computational linguistics

    11:38 Merging linguistics and computer science

    15:25 Understanding phonetics and morpho-syntax

    17:28 Exploring morpho-syntax and its relation to grammar

    20:33 Connection between phonetics and speech disorders

    24:41 Improvement of voice recognition systems

    27:31 Overview of speech recognition technology

    30:24 Challenges of ASR systems with atypical speech

    30:53 Strategies for improving recognition of disordered speech

    37:07 Data augmentation for training models

    40:17 Transfer learning in speech recognition

    42:18 Challenges of collecting data for various speech disorders

    44:31 Stammering and its connection to fluency issues

    45:16 Polish consonant combinations and pronunciation challenges

    46:17 Use of Amazon Transcribe for generating podcast transcripts

    47:28 Role of language models in speech recognition

    49:19 Contextual understanding in speech recognition

    51:27 How voice recognition systems analyze utterances

    54:05 Personalization of ASR models for individuals

    56:25 Language disorders and their impact on communication

    58:00 Applications of speech recognition technology

    1:00:34 Challenges of personalized and universal models

    1:01:23 Voice recognition in automotive applications

    1:03:27 Humorous voice recognition failures in cars

    1:04:13 Closing remarks and reflections on the discussion


    About the speaker:

    Katarzyna is a computational linguist with over 10 years of experience in NLP and speech recognition. She has developed language models for automotive brands like Audi and Porsche and specializes in phonetics, morpho-syntax, and sentiment analysis.

    Kasia also teaches at the University of Warsaw and is passionate about human-centered AI and multilingual NLP.

    Join our slack: https://datatalks.club/slack.html

    49 min

About DataTalks.Club

From the publisher's feed

DataTalks.Club - the place to talk about data!

More shows like DataTalks.Club

Radiolab by WNYC Studios

Radiolab

43,801 Listeners

Hidden Brain by Hidden Brain, Shankar Vedantam

Hidden Brain

43,345 Listeners

The Knowledge Project by Shane Parrish

The Knowledge Project

2,699 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

304 Listeners

Data Engineering Podcast by Tobias Macey

Data Engineering Podcast

145 Listeners

The Real Python Podcast by Real Python

The Real Python Podcast

140 Listeners

Huberman Lab by Scicomm Media

Huberman Lab

29,187 Listeners

The Ezra Klein Show by New York Times Opinion

The Ezra Klein Show

15,915 Listeners

ReThinking by TED

ReThinking

635 Listeners

Data Career Podcast: Helping You Land a Data Analyst Job FAST by Avery Smith - Data Career Coach

Data Career Podcast: Helping You Land a Data Analyst Job FAST

163 Listeners

The Analytics Engineering Podcast by dbt Labs, Inc.

The Analytics Engineering Podcast

29 Listeners

The Tucker Carlson Show by Tucker Carlson Network

The Tucker Carlson Show

15,951 Listeners