Training Machine Learning (ML) models on Kubernetes
Update: 2024-05-31
Description
In this episode of the Kubernetes Bytes podcast, Bhavin sits down with Bernie Wu, VP Strategic Partnerships and AI/CXL/Kubernetes Initiatives at Memverge. They discuss about how Kubernetes is the most popular platform to run AI model training and model inferencing jobs. The discussion dives into model training, talking about different phases of a DAG, and then talk about how Memverge can help users with efficient and cost-effective model checkpoints. The discussion goes into topics like saving costs by using spot instances, hot restart of training jobs, reclaiming unused GPU resources, etc.
Check out our website at https://kubernetesbytes.com/
Episode Sponsor: Nethopper
- Learn more about KAOPS: @nethopper.io
- For a supported-demo: info@nethopper.io
- Try the free version of KAOPS now! https://mynethopper.com/auth
Cloud Native News:
- https://www.aquasec.com/blog/linguistic-lumberjack-understanding-cve-2024-4323-in-fluent-bit/
- https://kubernetes.io/blog/2024/05/20/completing-cloud-provider-migration/
- https://thenewstack.io/introducing-aks-automatic-managed-kubernetes-for-developers/
- https://www.harness.io/blog/harness-to-acquire-split
Show Links:
- https://www.linkedin.com/in/berniewu/
- https://criu.org/Main_Page
- https://memverge.com/
- https://youtu.be/tY8YOMRuqWI?si=yB3hHqLUpYPZ-KWN
- https://youtu.be/ND4seSKpJHI?si=shh0iuA9qC-dO6eb
Timestamps:
Comments
Top Podcasts
The Best New Comedy Podcast Right Now – June 2024The Best News Podcast Right Now – June 2024The Best New Business Podcast Right Now – June 2024The Best New Sports Podcast Right Now – June 2024The Best New True Crime Podcast Right Now – June 2024The Best New Joe Rogan Experience Podcast Right Now – June 20The Best New Dan Bongino Show Podcast Right Now – June 20The Best New Mark Levin Podcast – June 2024
In Channel