April 07, 2025

Advanced LLM Optimization techniques

15 minutes

Welcome to another Data Architecture Elevator podcast! Today's discussion is hosted by Paolo Platter supported by our experts Antonino Ingargiola and Irene Donato.

In this episode, we explore effective strategies for optimizing large language models (LLMs) for inference tasks with multimodal data like audio, text, images, and video.

We discuss the shift from online APIs to hosted models, choosing smaller, task-specific models, and leveraging fine-tuning, distillation, quantization, and tensor fusion techniques. We also highlight the role of specialized inference servers such as Triton and Dynamo, and how Kubernetes helps manage horizontal scaling.

Don't forget to follow us on LinkedIn! Enjoy!

...more

View all episodes

By Agile Lab s.r.l.

April 07, 2025

Advanced LLM Optimization techniques

15 minutes

Welcome to another Data Architecture Elevator podcast! Today's discussion is hosted by Paolo Platter supported by our experts Antonino Ingargiola and Irene Donato.

In this episode, we explore effective strategies for optimizing large language models (LLMs) for inference tasks with multimodal data like audio, text, images, and video.

Don't forget to follow us on LinkedIn! Enjoy!

...more

Share Advanced LLM Optimization techniques

Sign up to save your podcasts

Advanced LLM Optimization techniques

Advanced LLM Optimization techniques