Skip to main content

AdamW optimizer

 Nowadays, most LLMs get trained with the AdamW optimizer as opposed to the Adam optimizer. Why? 


There used to be a time when Adam was the king among optimizers, and it didn't make much sense to spend too much time trying to find a better one. This has changed recently, and AdamW has become the default optimizer for the LLM practitioners. 


It all depends on how we apply the regularization terms to the weight parameters. For the typical gradient descent algorithm, if we want to apply an L2​ regularization term, we modify the loss function such that: 


regularized loss = loss + L2 term


Then, we compute the gradient of that new loss to update the model parameters. The goal of the regularization term is to ensure that the weights don't grow too large and it acts as a weight decay mechanism when we update the weights.


In Adam, when we apply the L2 regularization, we regularize the loss function as well but it gets used differently. The loss function is used to compute the first and second moments and when we update the weights, the regularization term is in the numerator and denominator of the gradient update term. Because of it, the effect of the L2 regularization term is minimized and cannot act as a weight decay mechanism.


In AdamW, on the other hand, we DO NOT regularize the loss function and compute the gradient update independent from the regularization term. Only during the weights update do we add the regularization term so that it acts on the weights and not on the loss function. Because of it,  training with AdamW tends to be more stable and leads to models that generalize better! Good to know, right?


Comments

Popular posts from this blog

Python Road map

 

Command on Run

 🔰 23 Important Commands in the RUN (Executer) List 🔹 The command dxdiag: Used to check all the specifications of your device.   🔹 The command cleanmgr: Opens the Disk Cleanup tool.   🔹 The command temp: Accesses temporary files, which we delete as they contribute to slowing down the computer.   🔹 The command regedit: Opens the Registry Editor.   🔹 The command calc: Opens the Calculator.   🔹 The command msconfig: A tool to access programs that run with Windows at startup and disable them to speed up the system.   🔹 The command scandisk: Used for disk checking.   🔹 The command cmd: Opens the Command Prompt for Windows.   🔹 The command defrag: Used to stop and defragment the hard drive.   🔹 The command taskman: Allows you to see what is open in the taskbar and manage it.   🔹 The command pbrush: Opens the Paint program in Windows.   🔹 The command debug: Used to ch...

Is it the end of lora?

 Is this the end of LoRA as a fine-tuning approach? Singular Value fine-tuning is here! We have a new paper in town called the "Transformers Squared". It promises to adapt any LLM to any task without external intervention.  The core idea of the paper is to use Singular Value Decomposition (SVD) to factorize the weight matrices of transformers. During training, we learn different singular values for different tasks. More specifically, we learn to scale the singular values for different tasks. Tasks can be as diverse as math reasoning or coding. The possibilities are endless. During inference, we do a two-pass inference. In the first pass, the LLM decides which "scale" to use for which task. In the second pass, the LLM adapts itself to be a specialist in the task and responds to our prompt. How cool is that? Paper Title: Transformer2: Self-adaptive LLMs Paper: https://sakana.ai/transformer-squared Blog: https://arxiv.org/abs/2501.06252 Video Explanation: https://youtu...