
Sign up to save your podcasts
Or


Guanhua Wang is a Senior Researcher in DeepSpeed Team at Microsoft. Before Microsoft, Guanhua earned his Computer Science PhD from UC Berkeley. Domino: Communication-Free LLM Training Engine
// MLOps Podcast #278 with Guanhua "Alex" Wang, Senior Researcher at Microsoft.
// Abstract
Given the popularity of generative AI, Large Language Models (LLMs) often consume hundreds or thousands of GPUs to parallelize and accelerate the training process. Communication overhead becomes more pronounced when training LLMs at scale. To eliminate communication overhead in distributed LLM training, we propose Domino, which provides a generic scheme to hide communication behind computation. By breaking the data dependency of a single batch training into smaller independent pieces, Domino pipelines these independent pieces of training and provides a generic strategy of fine-grained communication and computation overlapping. Extensive results show that compared with Megatron-LM, Domino achieves up to 1.3x speedup for LLM training on Nvidia DGX-H100 GPUs.
// Bio
Guanhua Wang is a Senior Researcher in the DeepSpeed team at Microsoft. His research focuses on large-scale LLM training and serving. Previously, he led the ZeRO++ project at Microsoft, which helped reduce over half of model training time inside Microsoft and LinkedIn. He also led and was a major contributor to Microsoft Phi-3 model training. He holds a CS PhD from UC Berkeley, advised by Prof Ion Stoica.
// MLOps Swag/Merchhttps://shop.mlops.community/
// Related Links
Website: https://guanhuawang.github.io/
DeepSpeed hiring: https://www.microsoft.com/en-us/research/project/deepspeed/opportunities/
Large Model Training and Inference with DeepSpeed // Samyam Rajbhandari // LLMs in Prod Conference: https://youtu.be/cntxC3g22oU
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, blogs, newsletters, and more: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Guanhua on LinkedIn: https://www.linkedin.com/in/guanhua-wang/
Timestamps:
[00:00] Guanhua's preferred coffee
[00:17] Takeaways
[01:36] Please like, share, leave a review, and subscribe to our MLOps channels!
[01:47] Phi model explanation
[06:29] Small Language Models optimization challenges
[07:29] DeepSpeed overview and benefits
[10:58] Crazy unimplemented crazy AI ideas
[17:15] Post-training vs QAT
[19:44] Quantization over distillation
[24:15] Using Lauras
[27:04] LLM scaling sweet spot
[28:28] Quantization techniques
[32:38] Domino overview
[38:02] Training performance benchmark
[42:44] Data dependency-breaking strategies
[49:14] Wrap up
Thanks to the High Signal Podcast by Delphina: https://go.mlops.community/HighSignalPodcast
Aditya Naganath is an experienced investor currently working with Kleiner Perkins. He has a passion for connecting with people over coffee and discussing various topics related to tech, products, ideas, and markets.
AI's Next Frontier // MLOps Podcast #277 with Aditya Naganath, Principal at Kleiner Perkins.
// Abstract
LLMs have ushered in an unmistakable supercycle in the world of technology. The low-hanging use cases have largely been picked off. The next frontier will be AI coworkers who sit alongside knowledge workers, doing work side by side. At the infrastructure level, one of the most important primitives invented by man - the data center- is being fundamentally rethought in this new wave.
// Bio
Aditya Naganath joined Kleiner Perkins’ investment team in 2022 with a focus on artificial intelligence, enterprise software applications, infrastructure, and security. Prior to joining Kleiner Perkins, Aditya was a product manager at Google, focusing on growth initiatives for the next billion users team. He previously was a technical lead at Palantir Technologies and formerly held software engineering roles at Twitter and Nextdoor, where he was a Kleiner Perkins fellow. Aditya earned a patent during his time at Twitter for a technical analytics product he co-created.
Originally from Mumbai, India, Aditya graduated magna cum laude from Columbia University with a bachelor’s degree in Computer Science and an MBA from Stanford University. Outside of work, you can find him playing guitar with a hard rock band, competing in chess or on the squash courts, and fostering puppies. He is also an avid poker player.
// MLOps Swag/Merch
https://shop.mlops.community/
// Related Links
Faith's Hymn by Beautiful Chorus: https://open.spotify.com/track/1bDv6grQB5ohVFI8UDGvKK?si=4b00752eaa96413b
Substack: https://adityanaganath.substack.com/?utm_source=substack&utm_medium=web&utm_campaign=substack_profile
With thanks to the High Signal Podcast by Delphina: https://go.mlops.community/HighSignalPodcast
Building the Future of AI in Software Development // Varun Mohan // MLOps Podcast #195 - https://youtu.be/1DJKq8StuTo
Do Re MI for Training Metrics: Start at the Beginning // Todd Underwood // AIQCON - https://youtu.be/DxyOlRdCofo
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, blogs, newsletters, and more: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Aditya on LinkedIn: https://www.linkedin.com/in/aditya-naganath/
Timestamps:
[00:00] Aditya's preferred coffee
[00:07] Takeaways
[01:33] Please like, share, leave a review, and subscribe to our MLOps channels!
[01:44] Investing in the AI frenzy
[02:23] Team dynamics insights
[04:57] Evaluating fad companies
[05:39] AI infrastructure and MLOps
[07:58] Challenges in MLOps Standardization
[08:59] ML vs Data platforms
[14:08] LLMOps vs MLOps
[17:52] Together vs Competitors
[21:19] AI application areas
[27:43] High Signal Podcast by Delphina Ad
[28:36] AI co-pilot in coding
[33:50] LLM providers overview
[37:30] AI paradigms and competition
[46:01] GPU failure management strategies
[52:58] Inference drives cloud revenue
[55:37] Wrap up
Dr Vincent Moens is an Applied Machine Learning Research Scientist at Meta and an author of TorchRL and TensorDict in Pytorch. PyTorch for Control Systems and Decision Making
// MLOps Podcast #276 with Vincent Moens, Research Engineer at Meta.
// Abstract
PyTorch is widely adopted across the machine learning community for its flexibility and ease of use in applications such as computer vision and natural language processing. However, supporting reinforcement learning, decision-making, and control communities is equally crucial, as these fields drive innovation in areas like robotics, autonomous systems, and game-playing. This podcast explores the intersection of PyTorch and these fields, covering practical tips and tricks for working with PyTorch, an in-depth look at TorchRL, and discussions on debugging techniques, optimization strategies, and testing frameworks. By examining these topics, listeners will understand how to effectively use PyTorch for control systems and decision-making applications.
// Bio
Vincent Moens is a research engineer on the PyTorch core team at Meta, based in London. As the maintainer of TorchRL (https://github.com/pytorch/rl) and TensorDict (https://github.com/pytorch/tensordict), Vincent plays a key role in supporting the decision-making community within the PyTorch ecosystem. Alongside his technical role in the PyTorch community, Vincent also actively contributes to AI-related research projects.
Before joining Meta, Vincent worked as an ML researcher at Huawei and AIG. Vincent holds a Medical Degree and a PhD in Computational Neuroscience.
// MLOps Swag/Merch
https://shop.mlops.community/
// Related Links
Musical recommendation: https://open.spotify.com/artist/1Uff91EOsvd99rtAupatMP?si=jVkoFiq8Tmq0fqK_OIEglg
Website: github.com/vmoens
TorchRL: https://github.com/pytorch/rl
TensorDict: https://github.com/pytorch/tensordict
LinkedIn post: https://www.linkedin.com/posts/vincent-moens-9bb91972_join-the-tensordict-discord-server-activity-7189297643322253312-Wo9J?utm_source=share&utm_medium=member_desktop
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, blogs, newsletters, and more: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Vincent on LinkedIn: https://www.linkedin.com/in/mvi/
Timestamps:
[00:00] Vincent preferred coffee
[00:12] Takeaways
[01:03] PyTorch tips and tricks
[05:30] Documentation Ambiguities and Unintended Guidance
[10:34] Modern copies and trade-offs
[17:56] Modular ML frameworks
[21:01] RL abstraction and generalization
[23:08] Developer Experience vs Functionality
[29:22] Streamlining user workflows
[31:10] Developer experience challenges
[36:50] Torch logs and contributions
[40:29] Well-formatted GitHub issue
[43:13] Testing PyTorch models
[48:45] Exciting PyTorch Development
[51:49] Tool discovery and sharing
[55:05] Wrap up
Matt Van Itallie is the founder and CEO of Sema. Prior to this, they were the Vice President of Customer Support and Customer Operations at Social Solutions.
AI-Driven Code: Navigating Due Diligence & Transparency in MLOps // MLOps Podcast #275 with Matt van Itallie, Founder and CEO of Sema.
// Abstract
Matt Van Itallie, founder and CEO of Sema, discusses how comprehensive codebase evaluations play a crucial role in MLOps and technical due diligence. He highlights the impact of Generative AI on code transparency and explains the Generative AI Bill of Materials (GBOM), which helps identify and manage risks in AI-generated code. This talk offers practical insights for technical and non-technical audiences, showing how proper diligence can enhance value and mitigate risks in machine learning operations.
// Bio
Matt Van Itallie is the Founder and CEO of Sema. He and his team have developed Comprehensive Codebase Scans, the most thorough and easily understandable assessment of a codebase and engineering organization. These scans are crucial for private equity and venture capital firms looking to make informed investment decisions. Sema has evaluated code within organizations that have a collective value of over $1 trillion. In 2023, Sema served 7 of the 9 largest global investors, along with market-leading strategic investors, private equity, and venture capital firms, providing them with critical insights.
In addition, Sema is at the forefront of Generative AI Code Transparency, which measures how much code created by GenAI is in a codebase. They are the inventors behind the Generative AI Bill of Materials (GBOM), an essential resource for investors to understand and mitigate risks associated with AI-generated code.
Before founding Sema, Matt was a Private Equity operating executive and a management consultant at McKinsey. He graduated from Harvard Law School and has had some interesting adventures, like hiking a third of the Appalachian Trail and biking from Boston to Seattle.
Full bio: https://alistar.fm/bio/matt-van-itallie
// MLOps Swag/Merch
https://shop.mlops.community/
// Related LinksWebsite: https://en.m.wikipedia.org/wiki/Michael_Gschwind
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, blogs, newsletters, and more: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Matt on LinkedIn: https://www.linkedin.com/in/mvi/
Timestamps:
[00:00] Matt's preferred coffee
[00:07] Takeaways
[01:25] Please like, share, leave a review, and subscribe to our MLOps channels!
[02:23] Code-based scans overview
[09:16] FinOps Automation and Recommendations
[12:03] Code quality evaluation layers
[16:10] Bridging Tech-Biz gap
[21:16] Measurable insights for leadership
[29:09] Startup prioritization and metrics
[32:53] GenAI Code Deletion Insights
[37:07] AI vs Developer Expertise
[41:47] GenAI Copyright Concerns
[46:41] Open Source Risks AI
[53:57] Code Defensibility and Acquisitions
[56:19] Wrap up
Dr. Michael Gschwind is a Director / Principal Engineer for PyTorch at Meta Platforms. At Meta, he led the rollout of GPU Inference for production services.
// MLOps Podcast #274 with Michael Gschwind, Software Engineer, Software Executive at Meta Platforms.
// Abstract
Explore the role in boosting model performance, on-device AI processing, and collaborations with tech giants like ARM and Apple. Michael shares his journey from gaming console accelerators to AI, emphasizing the power of community and innovation in driving advancements.
// Bio
Dr. Michael Gschwind is a Director / Principal Engineer for PyTorch at Meta Platforms. At Meta, he led the rollout of GPU Inference for production services. He led the development of MultiRay and Textray, the first deployment of LLMs at a scale exceeding a trillion queries per day shortly after its rollout. He created the strategy and led the implementation of PyTorch donation optimization with Better Transformers and Accelerated Transformers, bringing Flash Attention, PT2 compilation, and ExecuTorch into the mainstream for LLMs and GenAI models. Most recently, he led the enablement of large language models on-device AI with mobile and edge devices.
// MLOps Swag/Merch
https://mlops-community.myshopify.com/
// Related Links
Website: https://en.m.wikipedia.org/wiki/Michael_Gschwind
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, blogs, newsletters, and more: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Michael on LinkedIn: https://www.linkedin.com/in/michael-gschwind-3704222/?utm_source=share&utm_campaign=share_via&utm_content=profile&utm_medium=ios_app
Timestamps:
[00:00] Michael's preferred coffee
[00:21] Takeaways
[01:59] Please like, share, leave a review, and subscribe to our MLOps channels!
[02:10] Gaming to AI Accelerators
[11:34] Torch Chat goals
[18:53] Pytorch benchmarking and competitiveness
[21:28] Optimizing MLOps models
[24:52] GPU optimization tips
[29:36] Cloud vs On-device AI
[38:22] Abstraction across devices
[42:29] PyTorch developer experience
[45:33] AI and MLOps-related antipatterns
[48:33] When to optimize
[53:26] Efficient edge AI models
[56:57] Wrap up
//Abstract
Luke Marsden, is a passionate technology leader. Experienced in consultant, CEO, CTO, tech lead, product, sales, and engineering roles. Proven ability to conceive and execute a product vision from strategy to implementation, while iterating on product-market fit.
We Can All Be AI Engineers and We Can Do It with Open Source Models // MLOps Podcast #273 with Luke Marsden, CEO of HelixML.
// Abstract
In this podcast episode, Luke Marsden explores practical approaches to building Generative AI applications using open-source models and modern tools. Through real-world examples, Luke breaks down the key components of GenAI development, from model selection to knowledge and API integrations, while highlighting the data privacy advantages of open-source solutions.
// Bio
Hacker & entrepreneur. Founder at helix.ml. Career spanning DevOps, MLOps, and now LLMOps. Working on bringing business value to local, open-source LLMs.
// MLOps Swag/Merch
https://mlops-community.myshopify.com/
// Related LinksWebsite: https://helix.ml
About open source AI: https://blog.helix.ml/p/the-open-source-ai-revolution
Ratatat Cream on Chrome: https://open.spotify.com/track/3s25iX3minD5jORW4KpANZ?si=719b715154f64a5f
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, blogs, newsletters, and more: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Luke on LinkedIn: https://www.linkedin.com/in/luke-marsden-71b3789/
Timestamps:
[00:00] Michael's preferred coffee
[00:21] Takeaways
[01:59] Please like, share, leave a review, and subscribe to our MLOps channels!
[02:10] Gaming to AI Accelerators
[11:34] Torch Chat goals
[18:53] Pytorch benchmarking and competitiveness
[21:28] Optimizing MLOps models
[24:52] GPU optimization tips
[29:36] Cloud vs On-device AI
[38:22] Abstraction across devices
[42:29] PyTorch developer experience
[45:33] AI and MLOps-related antipatterns
[48:33] When to optimize
[53:26] Efficient edge AI models
[56:57] Wrap up
// Abstract
This panel speaks about the diverse landscape of AI agents, focusing on how they integrate voice interfaces, GUIs, and small language models to enhance user experiences. They'll also examine the roles of these agents in various industries, highlighting their impact on productivity, creativity, and user experience, and how these empower developers to build better solutions while addressing challenges like ensuring consistent performance and reliability across different modalities when deploying AI agents in production.
//Bio
Speakers:
Diego Oppenheimer - Co-founder @ Guardrails AI
Jazmia Henry - Founder and CEO @ Iso AI
Rogerio Bonatti - Researcher @ Microsoft
Julia Kroll - Applied Engineer @ Deepgram
Joshua Alphonse - Director of Developer Relations @ PremAIA Prosus | MLOps Community Production
Lauren Kaplan is a sociologist and writer. She earned her PhD in Sociology at Goethe University Frankfurt and worked as a researcher at the University of Oxford and UC Berkeley. The Impact of UX Research in the AI Space
// MLOps Podcast #272 with Lauren Kaplan, Sr UX Researcher.
// Abstract
In this MLOps Community podcast episode, Demetrios and UX researcher Lauren Kaplan explore how UX research can transform AI and ML projects by aligning insights with business goals and enhancing user and developer experiences. Kaplan emphasizes the importance of stakeholder alignment, proactive communication, and interdisciplinary collaboration, especially in adapting company culture post-pandemic. They discuss UX’s growing relevance in AI, challenges like bias, and the use of AI in research, underscoring the strategic value of UX in driving innovation and user satisfaction in tech.
// Bio
Lauren is a sociologist and writer. She earned her PhD in Sociology at Goethe University Frankfurt and worked as a researcher at the University of Oxford and UC Berkeley. Passionate about homelessness and AI, Lauren joined UCSF and later Meta. Lauren recently led UX research at a global AI chip startup and is currently seeking new opportunities to further her work in UX research and AI. At Meta, Lauren led UX research for 1) Privacy-Preserving ML and 2) PyTorch.
Lauren has worked on NLP projects such as Word2Vec analysis of historical HIV/AIDS documents presented at TextXD, UC Berkeley, 2019. Lauren is passionate about understanding technology and advocating for the people who create and consume AI. Lauren has published over 30 peer-reviewed research articles in domains including psychology, medicine, sociology, and more.”
// MLOps Swag/Merchhttps:
mlops-community.myshopify.com/
// Related Links
Podcast on AI UX https://open.substack.com/pub/aistudios/p/how-to-do-user-research-for-ai-products?r=7hrv8&utm_medium=ios
2024 State of AI Infra at Scale Research Report
https://ai-infrastructure.org/wp-content/uploads/2024/03/The-State-of-AI-Infrastructure-at-Scale-2024.pdf
Privacy-Preserving ML UX Public Article
https://www.ttclabs.net/research/how-to-help-people-understand-privacy-enhancing-technologies
Homelessness research and more: https://scholar.google.com/citations?user=24zqlwkAAAAJ&hl=en
Agents in Production: https://home.mlops.community/public/events/aiagentsinprod
Mk.gee Si (Bonus Track): https://open.spotify.com/track/1rukW2Wxnb3GGlY0uDWIWB?si=4d5b0987ad55444a
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, blogs, newsletters, and more: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Lauren on LinkedIn: https://www.linkedin.com/in/laurenmichellekaplan?utm_source=share&utm_campaign=share_via&utm_content=profile&utm_medium=ios_app
Timestamps:
[00:00] Lauren's Introduction
[00:07] Join the AI Agents in Production Conference on November 13!
[01:26] Takeaways
[02:13] Please like, share, leave a review, and subscribe to our MLOps channels!
[02:26] UX research overview
[03:31] UX research methods
[06:22] Effective interview strategies
[10:33] Broader UX Understanding
[14:42] Data Synthesis and Prioritization
[20:28] Measuring Impact in ML
[27:32] Phased Project Rollout
[31:36] UXR in Startups vs Big Companies
[40:32] AI Research Project Scope
[48:03] Increasing UX Maturity
[51:15] Career Paths in UX Research
[55:51] UX Beyond Tech
[59:31] Engineering User Research Tools
[1:06:16] Wrap up
Dr. Petar Tsankov is a researcher and entrepreneur in the field of Computer Science and Artificial Intelligence (AI). EU AI Act - Navigating New Legislation
// MLOps Podcast #271 with Petar Tsankov, Co-Founder and CEO of LatticeFlow AI.
Big thanks to LatticeFlow for sponsoring this episode!
// Abstract
Dive into AI risk and compliance. Petar Tsankov, a leader in AI safety, talks about turning complex regulations into clear technical requirements and the importance of benchmarks in AI compliance, especially with the EU AI Act. We explore his work with big AI players and the EU on safer, compliant models, covering topics from multimodal AI to managing AI risks. He also shares insights on "Comply," an open-source tool for checking AI models against EU standards, making compliance simpler for AI developers. A must-listen for those tackling AI regulation and safety.
// Bio
Co-founder & CEO at LatticeFlow AI, building the world's first product enabling organizations to build performant, safe, and trustworthy AI systems. Before starting LatticeFlow AI, Petar was a senior researcher at ETH Zurich working on the security and reliability of modern systems, including deep learning models, smart contracts, and programmable networks.
Petar has co-created multiple publicly available security and reliability systems that are regularly used:
= ERAN, the world's first scalable verifier for deep neural networks: https://github.com/eth-sri/eran
= VerX, the world's first fully automated verifier for smart contracts: https://verx.ch
= Securify, the first scalable security scanner for Ethereum smart contracts: https://securify.ch
= DeGuard, de-obfuscates Android binaries: http://apk-deguard.com
= SyNET, the first scalable network-wide configuration synthesis tool: https://synet.ethz.ch
Petar also co-founded ChainSecurity, an ETH spin-off that, within 2 years, became a leader in formal smart contract audits and was acquired by PwC Switzerland in 2020.
// MLOps Swag/Merch
https://mlops-community.myshopify.com/
// Related Links
Website: https://latticeflow.ai/
--------------- ✌️Connect With Us ✌️ -------------
Join our Slack community: https://go.mlops.community/slack
Follow us on Twitter: @mlopscommunity
Sign up for the next meetup: https://go.mlops.community/register
Catch all episodes, blogs, newsletters, and more: https://mlops.community/
Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/
Connect with Petar on LinkedIn: https://www.linkedin.com/in/petartsankov/
Timestamps:
[00:00] Petar's preferred coffee
[00:13] Takeaways
[01:12] AI Governance and EU Compliance
[04:16] AI Governance and Model Management
[09:48] AI Governance and Risk Management
[11:38] COMPL-AI
[15:02] EU AI Act Compliance Challenges
[17:16] EU AI Act Actionability
[22:26] Compliance and toxicity issues
[25:48] Model benchmarking and metrics
[28:13] Copyright and model evaluation
[32:28] SOC AI certification
[33:07] EU AI Act Gaps
[37:05] Integrity, safety, and compliance
[41:13] Benchmarking and Overfitting Concerns
[43:15] Tiered compliance approaches
[46:03] Bridging Law and Tech
[48:00] Multimodal AI Feature
[51:45] AI Risk and Mitigation
[55:37] Future Directions for Multimodal Models
[57:33] Wrap up
From the publisher's feed

1,289 Listeners

286 Listeners

1,089 Listeners

622 Listeners

582 Listeners

304 Listeners

337 Listeners

203 Listeners

565 Listeners

512 Listeners

141 Listeners

102 Listeners

222 Listeners

685 Listeners

30 Listeners