GPU
A GPU is a processor for graphics and parallel computation that is also used to train and run AI models.
A graphics processing unit, or GPU, is a processor designed to carry out many calculations in parallel. Unlike a CPU, which handles a broad range of tasks, it is particularly suited to parallel workloads and is used for both graphics and AI model training and inference.
GPU memory capacity and bandwidth are important hardware considerations when running large language models.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
That leaves plenty of room on a 16GB or 24GB consumer GPU for the context, the text the model keeps in view while it works.
And unlike training a frontier model on thousands of GPUs, building a harness is ordinary software engineering, open to independent developers.
Gemma 4 ships under the permissive Apache 2.0 license in five sizes, from a 2B model small enough for a phone to a 31B model that fits on a single GPU.
The company started by offering older NVIDIA GPUs that were not being used for its own model training.
Open-source options such as Flux and Z-Image can run on capable local hardware or rented GPUs, and their weights can be downloaded and customized.
…can make an image-generation service cheaper to run by deploying Python model code without a custom Dockerfile and shutting GPU workers down when they are idle.
Each could propose code changes, submit training jobs to a shared GPU cluster, read the resulting logs, and decide what to try next.
…AI's Colossus 2 facility in Memphis has deployed 550,000 NVIDIA GB200 and GB300 graphics processing units, or GPUs, the specialized chips used for AI workloads.
…sing hardware, check its fast memory: VRAM, or graphics-card memory, on a dedicated GPU; or unified memory, the shared memory available to an Apple Silicon Mac.
Both machines had 80 GPU cores.