weight
A weight is a trainable value that determines how strongly a signal influences a neural network’s computations.
A weight is a trainable numerical value that determines how strongly a signal influences a subsequent computation in a neural network. Weights are adjusted during training and commonly stored in matrices or tensors. Unlike hyperparameters such as the learning rate, they are learned rather than set in advance.
The weights of a large language model are stored in model files and loaded into memory for local execution. The phrase “model weights” can also refer collectively to learned parameters such as biases, but it does not include the model architecture or executable code.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
…measured cold start took 67 seconds end to end, including 15 seconds to load model weights into memory, while the GPU execution time remained 297 milliseconds.
It keeps the complete set of specialized model weights in memory but uses only a fraction to calculate each token.
Generating the reply is more sequential: for each new token, the system must repeatedly read model weights from unified memory.