Sample Space

Sample Space

By probablTechnology
Download on the App Store

Sample Space episodes

  • Time for some (extreme) distillation with Thomas van Dongen - founder of the Minish Lab

    Word embeddings might feel like they are a little bit out of fashion. After all, we have attention mechanisms and transformer models now, right? Well, it turns out that if you apply distillation the right way you can actually get highly performant word embeddings out. It's a technique featured by the model2vec project from the Minish lab and in this episode we talk to the founder to learn more about the technique.

    We have a Discord these days, feel free to discuss the podcast with us there! https://discord.probabl.ai

    This podcast is part of the open efforts over at probabl.

    To learn more you can check out website or reach out to us on social media.

    Website: https://probabl.ai/

    Bluesky: https://bsky.app/profile/probabl.bsky.social

    LinkedIn: https://www.linkedin.com/company/probabl

    Twitter: https://x.com/probabl_ai

    #probabl

    50 min
  • Imbalanced learn: regrets and onwards - with Guillaume Lemaitre, maintainer

    Imbalanced learn is one of the most popular scikit-learn projects out there. It has support for resampling techniques which historically have always been used for imbalanced classification use-cases. However, now that we are a few years down the line, it may be time to start rethinking the library. As it turns out, other techniques may be preferable. We talk to the maintainer, Guillaume Lemaitre, to discuss the lessons that have been learned over the last decade.

    We have a Discord these days, feel free to discuss the podcast with us there! https://discord.probabl.ai

    This podcast is part of the open efforts over at probabl. To learn more you can check out website or reach out to us on social media.

    Website: https://probabl.ai/

    Bluesky: https://bsky.app/profile/probabl.bsky.social

    LinkedIn: https://www.linkedin.com/company/probabl

    Twitter: https://x.com/probabl_ai

    55 min
  • You want to be in control of your own Copilot - with Ty Dunn, co-founder at Continue.dev

    There are many LLMs that you can use for programming these days. Some of them even go into your IDE like Cursor or Github Copilot. But what if you want to tweak these LLMs do to what you want? Instead of being stuck with the tools that a vendor gives you, the goal of Continue.dev is to allow you to customise this yourself. In this podcast we talk to Ty Dunn, co-founder of the project to learn more about this.

    If you are curious to learn more about this effort, please check out https://continue.dev. You may always want to read the manifesto over at https://amplified.dev/.

    We have a Discord these days, feel free to discuss the podcast with us there! https://discord.probabl.ai

    This podcast is part of the open efforts over at probabl. To learn more you can check out website or reach out to us on social media.

    Website: https://probabl.ai/

    Bluesky: https://bsky.app/profile/probabl.bsky.social

    LinkedIn: https://www.linkedin.com/company/probabl

    Twitter: https://x.com/probabl_ai

    1 hr 8 min
  • What it is like to maintain the scikit-learn docs - with David Arturo Amor Quiroz, scikit-learn docs maintainer

    Scikit-learn's documentation pages are celebrated. But not everyone is aware that the project actually has somebody on payroll to take care of it. In this episode we talk to Arturo about stories from the scikit-learn documentation. In particular, the docs have a recommender that few folks are aware of. People just assume that it is manually curated, but there are a few base scikit-learn tools under the hood there.

    Link to the official scikit-learn MOOC: https://inria.github.io/scikit-learn-mooc/

    We have a Discord these days, feel free to discuss the podcast with us there! https://discord.probabl.ai

    You can follow the podcast on most podcast players including apple podcasts, spotify and rss.com.

    - https://podcasts.apple.com/us/podcast/sample-space/id1739598572

    - https://open.spotify.com/show/0BnwEHuyOlHgeZfselpn1n

    - https://rss.com/podcasts/sample-space/

    This podcast is part of the open efforts over at probabl. To learn more you can check out website or reach out to us on social media.

    Website: https://probabl.ai/

    Bluesky: https://bsky.app/profile/probabl.bsky.social

    LinkedIn: https://www.linkedin.com/company/probabl

    Twitter: https://x.com/probabl_ai

    56 min
  • Sqlite can totally do embeddings now - with Alex Garcia, sqlite-vec maintainer

    Vector databases are kind of everywhere these days. There is a big pool of VC's that are pooring money into the ecosystem too. But while all of that is happening, sqlite has also gotten support for it. In this episode we talk the Alex Garcia, the maintainer of this project, and discuss how the project got created on what the future has in store.

    Sqlite-vec Github repo:

    https://github.com/asg017/sqlite-vec

    Alex Garcia blog:

    https://alexgarcia.xyz/blog/2024/sqlite-vec-hybrid-search/index.html

    Datasette discord:

    https://discord.com/invite/ktd74dm5mw

    Sqlite-vec channel on Mozilla Discord:

    https://discord.gg/Ve7WeCJFXk

    1 hr
  • How to rethink the notebook - with Akshay Agrawal, co-creator of Marimo

    Jupyter has been a great environment to explore computational ideas, but that doesn't mean that it can be the only environment for interactive coding in Python. It also comes with some downsides, which led Akshay Agrawal to create an alternative called Marimo. We discussed it in a previous livestream and figured that it was time to sit down with the creator to learn what led to the development of this exciting new too.

    You can learn more about Marimo by going to their website over at https://marimo.io

    To learn more you can check out website or reach out to us on social media.

    Website: https://probabl.ai/

    LinkedIn: https://www.linkedin.com/company/probabl

    Twitter: https://x.com/probabl_ai

    1 hr 13 min
  • You are always dealing with many tables - with Madelon Hulsebos

    When you are working on a data pipeline for ML ... you are never dealing with a single table. It always demands different tables for different reasons that all have to be mashed together in order to have something that you can learn from. But if that is the case, why do we spend so much time talking about ML pipelines that only work on a single table? Madelon Hulsebos has a Phd on the topic and so we figured that we might ask her.

    As mentioned in the podcast, here is the link to Madelon's homepage. https://www.madelonhulsebos.com/

    Some links to interesting articles from Madelon, as well as her homepage, can be found below. https://www.madelonhulsebos.com/assets/dataset_search_survey.pdf

    https://dl.acm.org/doi/pdf/10.1145/3654975

    https://dl.acm.org/doi/pdf/10.1145/3588710

    1 hr 10 min
  • How Narwhals has many end users ... that never use it directly with Marco Gorelli

    When you pip install a package you will for sure end up using it later. But often you will also install a bunch of dependencies and it is very likely that you won't directly interact with all of them. That does not mean that such a package is not useful, it merely means that the package might be directly used by a maintainer instead. This is interesting, because recently one such tool came into existence. It is called Narwhals and it seems to be on track to become critical infrastructure for data science projects. We have the maintainer of Narwhals on the show this week to talk about it.

    To learn more about Narwhals, you can check the repository here: https://github.com/narwhals-dev/narwhals

    This podcast is part of the open efforts over at probabl. To learn more you can check out website or reach out to us on social media.

    Website: https://probabl.ai/

    LinkedIn: https://www.linkedin.com/company/probabl

    Twitter: https://x.com/probabl_ai

    1 hr 1 min
  • Model safety, that's a pickle! with Adrin Jalali - scikit-learn maintainer

    Historically it's always been the case that you would use a pickle file to store a trained scikit-learn model on disk for deployment. Pickles make sense because these are so flexible, but they do carry a security concern. Adrin has been working on a remedy called skops, which is the main topic of this podcast.

    To learn more about skops, make sure to check the documentation: https://skops.readthedocs.io/en/stable/

    1 hr 2 min

About Sample Space

From the publisher's feed

Sample space is a podcast about tools, thoughts and techniques from machine learning practitioners. We talk to toolmakers and practitioners about interesting problems in the real world to find out…