Filtered by speaker:Dwight Churchill×
Searching...
Searching...
2 results for “model distillation” by Dwight Churchill
...and it's still, oh, it's 10,000,000,000 parameters, 20,000,000,000 parameters, whatever. On the inference side, it's similar learnings happening simultaneously. We don't need to do a 100 steps of diffusion for inference, like a 100 denoising steps to
...can distill models and have them work with a few steps of diffusion now. I think we're definitely the most inefficient we'll ever be, and it's only gonna get more and more efficient. It could be a factor of at least an order of magnitude, like, 10 x
Have a podcast?
Get ranked clips, hooks, and ready-to-post copy from your own episodes. Free to try.