
Sign up to save your podcasts
Or


Model sizes are crazy these days with billions and billions of parameters. As Mark Kurtz explains in this episode, this makes inference slow and expensive despite the fact that up to 90%+ of the parameters don’t influence the outputs at all.
Mark helps us understand all of the practicalities and progress that is being made in model optimization and CPU inference, including the increasing opportunities to run LLMs and other Generative AI models on commodity hardware.
Sponsors:
Featuring:
Show Notes:
Upcoming Events:
By Practical AI LLC4.4
189189 ratings
Model sizes are crazy these days with billions and billions of parameters. As Mark Kurtz explains in this episode, this makes inference slow and expensive despite the fact that up to 90%+ of the parameters don’t influence the outputs at all.
Mark helps us understand all of the practicalities and progress that is being made in model optimization and CPU inference, including the increasing opportunities to run LLMs and other Generative AI models on commodity hardware.
Sponsors:
Featuring:
Show Notes:
Upcoming Events:

289 Listeners

1,101 Listeners

169 Listeners

438 Listeners

300 Listeners

347 Listeners

312 Listeners

97 Listeners

138 Listeners

98 Listeners

227 Listeners

649 Listeners

105 Listeners

54 Listeners

34 Listeners