Searching...
Searching...
15 results for “growth stage”
growth
Growth stage
So that means that if you're doing something like test time compute and you want to spend a bunch of tokens thinking about what comes next, the longer that that goes, the the the the more tokens you spend on that, that compute grows quadratically in
One thing we found particularly helpful was q k normalization was stabilized, which stabilized training. Another point, that I quickly want to mention is that we reiterated on the story that scaling improves the performance, but it's also important t
So, yeah, so today, we're really excited to talk to you a little bit about that. So first, I'm gonna give a broad overview of kind of the last few years of progress in non post transformer architectures. And then afterwards, Eugene will tell us a lit
And that's kind of a key idea behind diffusion model training, which is that, for any time step t in the process, we can write x t as, our clean data x zero plus, a scaled version of a standard normal, variable. And then a scaling factor sigma of t i
thinking models. And, I think it's reasonably well diffused now, the idea of, that that this is kind of the next scaling paradigm. All analogies are imperfect. What is one way in which thinking fast and slow or system one, system two kinda doesn't tr
...scaling growth. Yeah. And and for listeners, like, prefill will be long input, decode will be long output, for example. Right? Yeah. So, like, decode decode scale I mean, decode is funny because the amount of tokens that you produce scales with the o
reasons to route to different model providers or whatever. But I think that routers are going to eventually go away. And I can understand why it's worth doing it in the short term, because the fact is it is beneficial right now. And if you're buildin
Right? So they have a phase. They have a complex number. And and the value of that phase matters because they you know, zero plus one and zero minus one are completely different states. A p bit, for example, you can have it be 20% of the time sitting
scaling model size and maybe doing a little bit more pretraining. And, you know, especially at the time, it really was about model size. And just sort of doing more uniform scaling of that nature is just going to solve all of your problems. Yeah. And
you can think of the x axis, so kind of what you are scaling as a combination of compute and data, which are kind of similar. And then the y axis is like the held out prediction accuracy over next tokens. We talk about models being autoregressive. It
So, historically, models would be hosted with a single inference engine, and that inference engine would ping pong between two phases. There's pre fill where you're reading the sequence, generating kv cache, which is basically just a set of vectors t
And then at the end, then there's the of the pre training, then there's all the post training. There's many different ways of doing that, different ways of patching it. So there's a whole experiment and phase there, which you can also get a lot of ga
Representing scaling laws where it became more and more formalized that bigger is better across multiple dimensions of what bigger means. So, and but these are all sort of neural networks we're talking about, and we're talking about different archite
Have a podcast?
Get ranked clips, hooks, and ready-to-post copy from your own episodes. Free to try.