PyTorch Developer Podcast

PyTorch Developer Podcast

By Edward Yang, Team PyTorchTechnology
Download on the App Store

PyTorch Developer Podcast episodes

  • Code generation

    Why does PyTorch use code generation as part of its build process? Why doesn't it use C++ templates? What things is code generation used for? What are the pros/consof using code generation? What are some other ways to do the same things we currently do with code generation?

    Further reading.

    • Top level file for the new code generation pipeline https://github.com/pytorch/pytorch/blob/master/tools/codegen/gen.py
    • Out of tree external backend code generation from Brian Hirsh: https://github.com/pytorch/xla/issues/2871
    • Documentation for native_functions.yaml https://github.com/pytorch/pytorch/blob/master/aten/src/ATen/native/README.md (have you seen this README before? Yes you've seen this README before. Imma post it again.)

    Outline:

    • High level: reduce the amount of code in PyTorch, easier to develop
    • Strongly typed python
    • Stuff we're using codegen for
      • Meta point: stuff c++ metaprogramming can't do
      • C++ apis (functions, methods on classes)
        • Especially for forwarding (operator dot doko)
        • Prototypes for c++ to implement
      • YAML files used by external frameworks for binding (accidental)
      • Python arg parsing
      • pyi generation
      • Autograd classes for saving saved data
      • Otherwise complicated constexpr computation (e.g., parsing JIT
        schema)
    • Pros
      • Better surface syntax (native_functions.yaml, jit schema,
        derivatives.yaml)
      • Better error messages (template messages famously bad)
      • Easier to organize complicated code; esp nontrivial input
        data structure
      • Easier to debug by looking at generated code
    • Con
      • Not as portable (template can be used by anyone)
      • Less good modeling for C++ type based metaprogramming (we've replicated a crappy version of C++ type system in our codegen)
    • Counterpoints in the design space
      • C++ templates: just as efficient
      • Boxed fallback: simpler, less efficient
    • Open question: can you have best of both worlds, e.g., with partially evaluated interpreters?
    17 min
  • Why is autograd so complicated

    Why is autograd so complicated? What are the constraints and features that go into making it complicated? What's up with it being written in C++? What's with derivatives.yaml and code generation? What's going on with views and mutation? What's up with hooks and anomaly mode? What's reentrant execution? Why is it relevant to checkpointing? What's the distributed autograd engine?

    Further reading.

    • Autograd notes in the docs https://pytorch.org/docs/stable/notes/autograd.html
    • derivatives.yaml https://github.com/pytorch/pytorch/blob/master/tools/autograd/derivatives.yaml
    • Paper on autograd engine in PyTorch https://openreview.net/pdf/25b8eee6c373d48b84e5e9c6e10e7cbbbce4ac73.pdf
    16 min
  • __torch_function__

    What is __torch_function__? Why would I want to use it? What does it have to do with keeping extra metadata on Tensors or torch.fx? How is it implemented? Why is __torch_function__ a really popular way of extending functionality in PyTorch? What makes it different from the dispatcher extensibility mechanism? What are some downsides of it being written this way? What are we doing about it?

    Further reading.

    • __torch_function__ RFC: https://github.com/pytorch/rfcs/blob/master/RFC-0001-torch-function-for-methods.md
    • One of the original GitHub issues tracking the overall design discussion https://github.com/pytorch/pytorch/issues/24015
    • Documentation for using __torch_function__ https://pytorch.org/docs/stable/notes/extending.html#extending-torch
    17 min
  • TensorIterator

    You walk into the whiteboard room to do a technical interview. The interviewer looks you straight in the eye and says, "OK, can you show me how to add the elements of two lists together?" Confused, you write down a simple for loop that iterates through each element and adds them together. Your interviewer rubs his hands together evilly and cackles, "OK, let's make it more complicated."

    What does TensorIterator do? Why the heck is TensorIterator so complicated? What's going on with broadcasting? Type promotion? Overlap checks? Layout? Dimension coalescing? Parallelization? Vectorization?

    Further reading.

    • PyTorch TensorIterator internals https://labs.quansight.org/blog/2020/04/pytorch-tensoriterator-internals/
    • Why is TensorIterator so slow https://dev-discuss.pytorch.org/t/comparing-the-performance-of-0-4-1-and-master/136
    • Broadcasting https://pytorch.org/docs/stable/notes/broadcasting.html and type promotion https://pytorch.org/docs/stable/tensor_attributes.html#type-promotion-doc
    18 min
  • native_functions.yaml

    What does native_functions.yaml have to do with the TorchScript compiler? What multiple use cases is native_functions.yaml trying to serve? What's up with the JIT schema type system? Why isn't it just Python types? What the heck is the (a!) thingy inside the schema? Why is it important that I actually annotate all of my functions accurately with this information? Why is my seemingly BC change to native_functions.yaml actually breaking people's code? Do I have to understand the entire compiler to understand how to work with these systems?

    Further reading.

    • native_functions.yaml README https://github.com/pytorch/pytorch/blob/master/aten/src/ATen/native/README.md
    • Tracking issue for serializing default arguments https://github.com/pytorch/pytorch/issues/54613
    • Test for BC breaking changes in native_functions.yaml https://github.com/pytorch/pytorch/blob/master/test/backward_compatibility/check_backward_compatibility.py
    16 min
  • Serialization

    What is serialization? Why do I care about it? How is serialization done in general in Python? How does pickling work? How does PyTorch implement pickling for its objects? What are some pitfalls of pickling implementation? What does backwards compatibility and forwards compatibility mean in the context of serialization? What's the difference between directly pickling and using torch.save/load? So what the heck is up with JIT/TorchScript serialization? Why did we use zip files? What were some design principles for the serialization format? Why are there two implementations of serialization in PyTorch? Is the fact that PyTorch uses pickling for serialization mean that our serialization format is insecure?

    Further reading.

    • TorchScript serialization design doc https://github.com/pytorch/pytorch/blob/master/torch/csrc/jit/docs/serialization.md
    • Evolution of serialization formats over time https://github.com/pytorch/pytorch/issues/31877
    • Code pointers:
      • Tensor __reduce_ex__ https://github.com/pytorch/pytorch/blob/de845020a0da39e621db984515bc1cce03f526ea/torch/_tensor.py#L97-L178
      • Python side serialization https://github.com/pytorch/pytorch/blob/de845020a0da39e621db984515bc1cce03f526ea/torch/serialization.py#L384-L499
      • C++ side serialization https://github.com/pytorch/pytorch/tree/master/torch/csrc/jit/serialization
    18 min
  • Continuous integration

    How is our CI put together? What is the history of the CI? What constraints are under the CI? Why does the CI use Docker? Why are build and test split into two phases? Why are some parts of the CI so convoluted? How does the HUD work? What kinds of configurations is PyTorch tested under? How did we decide what configurations to test?  What are some of the weird CI configurations? What's up with the XLA CI? What's going on with the Facebook internal builds? 

    Further reading.

    • The CI HUD for viewing the status of master https://ezyang.github.io/pytorch-ci-hud/build/pytorch-master
    • Structure of CI https://github.com/pytorch/pytorch/blob/master/.circleci/README.md
    • How to debug Windows problems on CircleCI https://github.com/pytorch/pytorch/wiki/Debugging-Windows-with-Remote-Desktop-or-CDB-(CLI-windbg)-on-CircleCI
    17 min
  • Stacked diffs and ghstack

    What's a stacked diff? Why might you want to do it? What does the workflow for stacked diffs with ghstack look like? How to use interactive rebase to edit earlier diffs in my stack? How can you actually submit a stacked diff to PyTorch? What are some things to be aware of when using ghstack?

    Further reading.

    • The ghstack repository https://github.com/ezyang/ghstack/
    • A decent explanation of how the stacked diff workflow works on Phabricator, including how to do rebases https://kurtisnusbaum.medium.com/stacked-diffs-keeping-phabricator-diffs-small-d9964f4dcfa6
    13 min
  • Shared memory

    What is shared memory? How is it used in your operating system? How is it used in PyTorch? What's shared memory good for in deep learning? Why use multiple processes rather than one process on a single node? What's the point of PyTorch's shared memory manager? How are allocators for shared memory implemented? How does CUDA shared memory work? What is the difference between CUDA shared memory and CPU shared memory? How did we implement safer CUDA shared memory?

    Further reading.

    • Implementations of vanilla shared memory allocator https://github.com/pytorch/pytorch/blob/master/aten/src/TH/THAllocator.cpp and the fancy managed allocator https://github.com/pytorch/pytorch/blob/master/torch/lib/libshm/libshm.h
    • Multiprocessing best practices describes some things one should be careful about when working with shared memory https://pytorch.org/docs/stable/notes/multiprocessing.html
    • More details on how CUDA shared memory works https://pytorch.org/docs/stable/multiprocessing.html#multiprocessing-cuda-sharing-details
    11 min
  • Automatic mixed precision

    What is automatic mixed precision? How is it implemented? What does it have to do with mode dispatch keys, fallthrough kernels? What are AMP policies? How is its cast caching implemented? How does torchvision also support AMP? What's up with Intel's CPU autocast implementation?

    Further reading.

    • Autocast implementation lives at https://github.com/pytorch/pytorch/blob/master/aten/src/ATen/autocast_mode.cpp
    • How to add autocast implementations to custom operators that are out of tree https://pytorch.org/tutorials/advanced/dispatcher.html#autocast
    • CPU autocasting PR https://github.com/pytorch/pytorch/pull/57386
    15 min

About PyTorch Developer Podcast

From the publisher's feed

The PyTorch Developer Podcast is a place for the PyTorch dev team to do bite sized (10-20 min) topics about all sorts of internal development topics in PyTorch.

More shows like PyTorch Developer Podcast

Talk Python To Me by Michael Kennedy

Talk Python To Me

582 Listeners

Science Weekly by The Guardian

Science Weekly

419 Listeners

Cautionary Tales with Tim Harford by Pushkin Industries

Cautionary Tales with Tim Harford

5,084 Listeners

All-In with Chamath, Jason, Sacks & Friedberg by All-In Podcast, LLC

All-In with Chamath, Jason, Sacks & Friedberg

10,185 Listeners