Rapid Synthesis: Delivered under 30 mins..ish, or it's on me!

Scaling Multi-Tenant ML Inference on Kubernetes: Workday's Strategy


Listen Later

Workday's engineering team tackled the challenge of scaling machine learning inference for numerous customers by devising a "bin packed shards" strategy on Kubernetes. This approach, detailed in their Medium article from January 2022, involves grouping multiple tenants' ML models into shared units called shards, aiming for efficient resource usage, particularly memory. Kubernetes handles the deployment and scaling of these shards, while Istio's Virtual Services manage the routing of tenant-specific requests. The strategy offers benefits like cost reduction and independent model management but also presents complexities in initial design and ongoing operation, focusing on a balance between efficiency and manageability.

...more
View all episodesView all episodes
Download on the App Store

Rapid Synthesis: Delivered under 30 mins..ish, or it's on me!By Benjamin Alloul πŸ—ͺ πŸ…½πŸ…ΎπŸ†ƒπŸ…΄πŸ…±πŸ…ΎπŸ…ΎπŸ…ΊπŸ…»πŸ…Ό