Tokenization

A Tokenizer:
For example:
notion image
The compression ratio is number of bytes per token, one could increase compression ratio by increasing vocabulary size (number of possible token values increases), leading to sparsity.
There are approximately 150K Unicode characters, which means if you use every Unicode characters as a token(this is called character_tokenizer), you will have to face to problems:
Problem 1: this is a very large vocabulary. Problem 2: many characters are quite rare (e.g., 🌍), which is inefficient use of the vocabulary.
Another choice is to use the result of Unicode-8:
this is byte tokenizer, the vocabulary is small however the compression ratio is one.
Another ways is to split strings into words:
What's good: each token is meaningful (since humans invented words). vocabulary_size = "Number of distinct chunks in the training data"
Compression ratio is good, but vocabulary size can be huge.
Moreover:
  • Many words are rare and the model won't learn much about them.
  • This doesn't obviously provide a fixed vocabulary size.
  • New words we haven't seen during training get a special UNK token, which is ugly and can mess up perplexity calculations.
Popular tokenizer: Byte-Pair Encoding (BPE)
Basic idea: train the tokenizer on raw text to construct a vocabulary tailored to the data.
Intuition: common sequences of bytes are represented by a single token, rare sequences are represented by many tokens.
Sketch: start with each byte as a token, and successively merge the most common pair of adjacent tokens.

Resource Accounting

Question: How long would it take to train a 70B parameter model on 15T tokens on 1024 H100s?
According to
,so
Question: What's the largest model that can you can train on 8 H100s using AdamW?
Mixed precision training:
  • Training with fp32 works, but requires lots of memory.
  • Training with fp16 and even bf16 is risky, and you can get instability.
  • Use bf16 for parameters, activations, and gradients
  • Use fp32 for optimizer states(also for exponents)
By default, tensors are stored into memory of CPUs, so if you want to calculate it on GPUs, you need to move the sensor from CPUs to GPUs and move it back after calculation has been done.
notion image
You can also create tensors on GPUs.
 
Then we need to talk about calculation of tensor, we usually use einops
sum:
reduce:
rearrange:
 
Then let’s talk about the computational costs, we usually use two notes:
  • FLOPs: floating-point operations (measure of computation done)
  • FLOP/s: floating-point operations per second (also written as FLOPS), which is used to measure the speed of hardware.
For two matrix , if we multiply them we need D times multiplication and D-1 times addition per output element, so we need:
and for matrix , if we do element-wise multiplication on them, we need FLOPs.
 
You can measure the practical FLOP/s and compare to ideal FLOP/s, then you can calculate Model FLOPs Utilization(MFU):
Usually, MFU of ≥ 0.5 is quite good!
 
So the important thing is that we need to figure out the bottleneck during computation, which means we need to know where the bottleneck coming from.
Let’s begin with an example:
here we need to move x from HBM to computation unit, and then we need to compute Relu and move y back to HBM:
What is the bottleneck?
  • Memory-bound: communication time > computation time
  • Compute-bound: computation time > communication time
So in this case, ReLU is memory-bound.
 
Or we can see accelerator intensity versus arithmetic intensity
What is the bottleneck?
  • Memory-bound: arithmetic intensity < accelerator intensity
  • Compute-bound: arithmetic intensity > accelerator intensity
In general, we'll find ourselves memory bound.
And interesting is that if you change relu to geLu, gelu will do more flops than relu, so the arithmetic_intensity rise to 5, however when the problem is still memory bond so we won’t get faster or slower. So even the arithmetic intensity increase but time didn’t change the FLOPs/s will increase when facing memory bound, and when the bound type shift to compute bound the FLOPs/s will follow the speed of software.
notion image
For vector product the arithmetic intensity is 0.5, for matrix vector product is 1, however for matrix multiplication is n/3 where n is shape of matrix.
So when training for transformer what happens is compute-bound but for inference what mainly doing is matrix vector product, so it is compute-bond.
 
Now Let’s talk about a particular case:
notion image
We have known that for forward propagation, a layer bring 2*B*D*D FLOPs, and for every weight matrix it brings 64 parameters, which means 4*64 bytes for float32.
And for backward propagation, we need to compute gradients and store the tensor graph, usually every weight matrix and output have a same shape matrix.
Let’s focus on one layer:
so we know that num_backward_flops = (2 * B * D * D) + (2 * B * D * D), which means for a full process of train we have to calculate 3 * 2 * B * D * D = 6 * B * (D^2), here B is the number of data and D^2 is number of parameters, this seems just work for easy MLPs but it turns out to be a good approximation for Transformers for short context lengths as well.
Another things is optimizer, for example AdaGrad, it needs to compute and store the square of grad as its optimizer state, so means to compute and store a matrix which has same shape of parameters:
and for memory cost:
here gradient has same shape as parameter for that when we use loss.backward the algorithm will release the non-leaf tensors automatically.
here activation memory is stored for backward propagation, and if you have enough memory, another way is to just store the value before activation, drop the activation and recompute them when doing backward propagation.

Architecture

Normalization in or not in residual

notion image
Almost all modern transformers LLM use pre-norm but why?
notion image
The pre-norm has better train stability for it has more stable gradients in deep networks:
notion image
for post-norm which means:
so that:
will cause numerical instability.
So the key thing is whether put normalization in residual stream, we can also use post norm with normalization out of residual stream(Like Grok, Gemma 2). And even some models put normalization pre and post FFN and attention only if it is out of residual stream.
notion image

LayerNorm or RMSNorm

LayerNorm means:
RMSNorm means:
where:
There is not specific reason to use RMSNorm rather than LayerNorm, however, RMSNorm is usually faster and didn’t bring precision cost.
But it is interesting that in fact Layer-Norm just accounts for a insignificant FLOPs(typically < 0.5%), but it accounts for a much more runtime(5% to 15%) because it is strictly memory-bandwidth bound. Its arithmetic intensity is around 0.5 to 1 FLOP/byte.
notion image
More generally we always drop the bias in FFN.

Activations

notion image
In modern LLM, researchers have notice the importance of gate, so heuristically use it in activation:
notion image
when we use GLU style activation, we have to store 3 matrix instead of 2 matrix so an idea is to use smaller ouput dimension by the factor 2/3 to keep the same number of parameters, but it is just a general rule instead of an iron rule.
notion image

Serial versus Parallel

notion image

Position Embedding

notion image
We have been familiar with the first two embeddings. Relative embedding doesn’t add PE to tokens vector instead it just directly participant into attention score calculation through which is defined as:
where k is a predefined maximum clipping distance.
But modern models usually use RoPE(rotary position embedding), the motivation of RoPE is that we want to find an embedding which makes the inner product of two embedded token vectors just has connection with their original vector and their relative position:
Relative embeddings which directly modify attention matrix can’t fulfill this destination.
Starting from two dimensional case, we have two original semantic vectors , so we can rotate them by rotation matrix:
so we can embed them:
so we get:
so the inner product is just about the original semantic vector and their relative distance.
For high embedding dimension, we can just divide into several pairs and rotate the two dimensional vector of each pair separately, which means
notion image
in practice we write semantic vector as to compute, firstly we compute for every pair:
we share same for same pair of different vector. then we compute rotation matrix element:
so:
where . so:
so the inner product of embedded vector is summarize of of every pair.
Here we use different angles for different pair for two reason:
  • Avoiding Periodic Collisions
Trigonometric function is periodic so it may cause aliasing between i-th token and i plus -th token, different theta brings different frequency so it will never aliasing.
  • Having stronger expression
As we say Different theta brings different frequency, every pair are designed to express information of different fequency.
It is worth noting that absolute PE usually used in input layer while RoPE or other modern PE usually use in every layer of network.

Hyperparameters

feed-forward ratio

Some consensus:
so for GLU, it turns to be:
feed-forward ratio(d_ff/d_model) is an typical hyperparameter
notion image
there are a huge basin when use single-digit numbers of feed-forward ratio.

head dim ratio

we usually follow
which means the ratio between head dim and hidden dim is 1.

aspect ratio

we define aspect ratio as ratio between dimension of model and number of layers.
notion image
however if the model going too deep, we will face the problem of parallelization, so we always want to go wide.
notion image
so aspect ratio is safe near 100.

vocabulary sizes

notion image

Regularization

Nowadays we have tons of corpus on internet, so during pre-training we just even do a single pass on a corpus, which means overfitting is not a problem anymore, then what is the function of regularization(like weight decay or dropout)?
However weight decay is still a popular skill even when dropout seems to be mystifying, but the reason is not the rule of regularization which we measure by validation loss versus training loss:
notion image
The rule of weight decay is that it brings progress when it interact with cosine decay optimization:
notion image
the solid lines mean the primary training and dashed lines mean continued training at different checkpoint. In the two plot we can know that with cosine LR decay, weight decay could be helpful for training which is a little bit counterintuitive.

Stability

You don’t want to get spikes in LLM training for that you usually can’t get fine final models when it happens. Here are some reason why you may get spikes or grad explosion.
  • Softmax, for its exponentials and dividing by zero.
the solution is to use z-loss regularization:
which means give penalty on z so that prevent training instability.
  • QK norm
Adding normalization(RMS) on Q,K matrices is a key trick in keep training stable.
notion image

Faster Decoding

A key problem happens in decoding. When using kv cache, we need to transport kV to calculation unit and compute logits of new token with new Q. The transportation of KV bring a very low arbitrary intensity. MQA and GQA was introduced to solve this problem.
notion image
MQA(multi-query attention) use just one head for Key and Value matrices and broadcasting them to all heads when reasoning, meanwhile leading to lower expressiveness.
So GQA(grouped-query attention) use multiple(but less than Query) heads for Key and Value matrices, and use key-query ratio to control the expressiveness.
To turn a MHA model to a GQA model is easy, first thing is to do Heads pooling:
and then do continuing pre-training (uptraining) on of the original pre-training tokens.
notion image
notion image

Hybrid Attention

Another advantage pre-norm brings is that the output is simple addition of outputs of all layers:
so that we can make some layer sparse(e.g. sliding window attention) or linear and so forth to accelerate computation.
However recent research has shown that too many non-standard attention layers will cause expressiveness degradation, so trade-off between them are important.
notion image
This is what main researches nowadays work on.

Attention Alternatives

Linear Attention

The motivation is to remove the key non-linear part of standard attention: softmax
so that the computation complexity drop from to , the latter one is linear to .
One of the key reasons making linear attention an important work is that if you write it in incremental format it will be similar to recurrent neural network:
here is the memory at t-th step, it has fixed shape which means fixed memory volume.
But simple linear attention didn’t work well for strongly pooling and smoothing the memory .

Mamba

Mamba introduce Gate mechanism into recurrent step of linear attention:
where and is the gate. Then we get:
which offer the way to be trained parallel.

Gated delta net

Gated delta net turn to be:
where . So like LSTM, is like forget gate and is like update gate. And here is meant to remove the similar information in memory to .

Sparse Attention

What Deepseek V3 stype sparse attention do is to filter a subset of tokens to compute. First is to compute weight of all tokens to current token:
where
and then filter tokens by Top-k
and compute attention on the subset:
It it worthing noting that DSA style sparse attention is not used for training, filter is training as a plugin after the model training.

Mixture of Experts

Basics

MoE brings a lot of advantage through its sparsity, we can increase parameters of the model without increasing the FLOPs.
notion image
And MoE usually training faster:
notion image
what we do is to construct multiple FFN for single layer:
here t means t-th token, l means l-th layer, i means i-th FFN. And we need to decide which FFNs we want to use, we train a vector for i-th FFN of l-th layer, and compute:
as the score of i-th FFN, then get probability to choose i-th layer:
and then choose top-k FFNs and compute their output, and sum them with weight :
the second item is for residual connection.

Shared MoEs

In shared MoEs, we maintain some FFNs to be used in every token:
notion image
It’s easy to understand, and the effectiveness is significant.
notion image

Training MoEs

Major challenge in training MoEs is that if we want to sparsity, the gating decision are not differentiable.
The popular way is shown below:
notion image
it is a heuristic algorithm, here P_i means the probability mass on expert i, and f_i means the numbers of token choosing expert i. When backward propagating, we stop the gradient of f_i, so when we penalize loss we actually penalize large P_i, which leading to sparsity. Here f_i is design to prevent cheating when router modulates the probability but still using just single expert.
At perfect balance, , then loss = alpha.
Fine-tune on MoEs especially on experts will occur overfitting.
Another way to train MoEs is to use upcycling:
notion image
notion image
教育心理学Numerical Renormalization Group
Loading...