PyTorch Developer Podcast

PyTorch Developer Podcast

By Edward Yang, Team PyTorchTechnology
Download on the App Store

PyTorch Developer Podcast episodes

  • API design via lexical and dynamic scoping

    Lexical and dynamic scoping are useful tools to reason about various API design choices in PyTorch, related to context managers, global flags, dynamic dispatch, and how to deal with BC-breaking changes. I'll walk through three case studies, one from Python itself (changing the meaning of division to true division), and two from PyTorch (device context managers, and torch function for factory functions).

    Further reading.

    • Me unsuccessfully asking around if there was a way to simulate __future__ in libraries https://stackoverflow.com/questions/66927362/way-to-opt-into-bc-breaking-changes-on-methods-within-a-single-module
    • A very old issue asking for a way to change the default GPU device https://github.com/pytorch/pytorch/issues/260 and a global GPU flag https://github.com/pytorch/pytorch/issues/7535
    • A more modern issue based off the lexical module idea https://github.com/pytorch/pytorch/issues/27878
    • Array module NEP https://numpy.org/neps/nep-0037-array-module.html
    22 min
  • Intro to distributed

    Today, Shen Li (mrshenli) joins me to talk about distributed computation in PyTorch. What is distributed? What kinds of things go into making distributed work in PyTorch? What's up with all of the optimizations people want to do here?

    Further reading.

    • PyTorch distributed overview https://pytorch.org/tutorials/beginner/dist_overview.html
    • Distributed data parallel https://pytorch.org/docs/stable/notes/ddp.html
    16 min
  • Double backwards

    Double backwards is PyTorch's way of implementing higher order differentiation. Why might you want it? How does it work? What are some of the weird things that happen when you do this?

    Further reading.

    • Epic PR that added double backwards support for convolution initially https://github.com/pytorch/pytorch/pull/1643
    17 min
  • Functional modules

    Functional modules are a proposed mechanism to take PyTorch's existing NN module API and transform it into a functional form, where all the parameters are explicit argument. Why would you want to do this? What does functorch have to do with it? How come PyTorch's existing APIs don't seem to need this? What are the design problems?

    Further reading.

    • Proposal in GitHub issues https://github.com/pytorch/pytorch/issues/49171
    • Linen design in flax https://flax.readthedocs.io/en/latest/design_notes/linen_design_principles.html
    15 min
  • CUDA graphs

    What are CUDA graphs? How are they implemented? What does it take to actually use them in PyTorch?

    Further reading.

    • NVIDIA has docs on CUDA graphs https://developer.nvidia.com/blog/cuda-graphs/
    • Nuts and bolts implementation PRs from mcarilli: https://github.com/pytorch/pytorch/pull/51436 https://github.com/pytorch/pytorch/pull/46148
    14 min
  • Default arguments

    What do default arguments have to do with PyTorch design? Why are default arguments great for clients (call sites) but not for servers (implementation sites)? In what sense are default arguments a canonicalization to max arity? What problems does this canonicalization cause? Can you canonicalize to minimum arity? What are some lessons to take?

    Further reading. https://github.com/pytorch/pytorch/issues/54613 stop serializing default arguments

    15 min
  • Anatomy of a domain library

    What's a domain library? Why do they exist? What do they do for you? What should you know about developing in PyTorch main library versus in a domain library? How coupled are they with PyTorch as a whole? What's cool about working on domain libraries?

    Further reading.

    • The classic trio of domain libraries is https://pytorch.org/audio/stable/index.html https://pytorch.org/text/stable/index.html and https://pytorch.org/vision/stable/index.html

    Line notes.

    • why do domain libraries exist? lots of domains specific gadgets,
      inappropriate for PyTorch
    • what does a domain library do
      • operator implementations (old days: pure python, not anymore)
        • with autograd support and cuda acceleration
        • esp encoding/decoding, e.g., for domain file formats
          • torchbind for custom objects
          • takes care of getting the dependencies for you
        • esp transformations, e.g., for data augmentation
      • models, esp pretrained weights
      • datasets
      • reference scripts
      • full wheel/conda packaging like pytorch
      • mobile compatibility
    • separate repos: external contributors with direct access
      • manual sync to fbcode; a lot easier to land code! less
        motion so lower risk
    • coupling with pytorch? CI typically runs on nightlies
      • pytorch itself tests against torchvision, canary against
        extensibility mechanisms
      • mostly not using internal tools (e.g., TensorIterator),
        too unstable (this would be good to fix)
    • closer to research side of pytorch; francesco also part of papers
    17 min
  • TensorAccessor

    What's TensorAccessor? Why not just use a raw pointer? What's PackedTensorAccessor? What are some future directions for mixing statically typed and typed erase code inside PyTorch proper?

    Further reading.

    • TensorAccessor source code, short and sweet https://github.com/pytorch/pytorch/blob/master/aten/src/ATen/core/TensorAccessor.h
    • Legacy THCDeviceTensor https://github.com/pytorch/pytorch/blob/master/aten/src/THC/THCDeviceTensor.cuh
    12 min
  • Random number generators

    Why are RNGs important? What is the generator concept? How do PyTorch's CPU and CUDA RNGs differ? What are some of the reasons why Philox is a good RNG for CUDA? Why doesn't the generator class have virtual methods for getting random numbers? What's with the next normal double and what does it have to do with Box Muller transform? What's up with csprng?

    Further reading.

    • CUDAGeneratorImpl has good notes about CUDA graph interaction and pointers to all of the rest of the stuff https://github.com/pytorch/pytorch/blob/1dee99c973fda55e1e9cac3d50b4d4982b6c6c26/aten/src/ATen/CUDAGeneratorImpl.h
    • Transform uniformly distributed random numbers to other distributions with https://github.com/pytorch/pytorch/blob/master/aten/src/ATen/core/TransformationHelper.h
    • torchcsprng https://github.com/pytorch/csprng
    15 min
  • vmap

    What is vmap? How is it implemented? How does our implementation compare to JAX's? What is a good way of understanding what vmap does? What's up with random numbers? Why are there some issues with the vmap that PyTorch currently ships?

    Further reading.

    • Tracking issue for vmap support https://github.com/pytorch/pytorch/issues/42368
    • BatchedTensor source code https://github.com/pytorch/pytorch/blob/master/aten/src/ATen/BatchedTensorImpl.h , logical-physical transformation helper code https://github.com/pytorch/pytorch/blob/master/aten/src/ATen/VmapTransforms.h (well documented, worth a read)
    • functorch, the better, more JAX-y implementation of vmap https://github.com/facebookresearch/functorch
    • Autodidax https://jax.readthedocs.io/en/latest/autodidax.html which contains a super simple vmap implementation that is a good model for the internal implementation that PyTorch has
    18 min

About PyTorch Developer Podcast

From the publisher's feed

The PyTorch Developer Podcast is a place for the PyTorch dev team to do bite sized (10-20 min) topics about all sorts of internal development topics in PyTorch.

More shows like PyTorch Developer Podcast

Talk Python To Me by Michael Kennedy

Talk Python To Me

582 Listeners

Science Weekly by The Guardian

Science Weekly

419 Listeners

Cautionary Tales with Tim Harford by Pushkin Industries

Cautionary Tales with Tim Harford

5,084 Listeners

All-In with Chamath, Jason, Sacks & Friedberg by All-In Podcast, LLC

All-In with Chamath, Jason, Sacks & Friedberg

10,185 Listeners