Memory, not a processor.
High-bandwidth memory (HBM) is a stacked DRAM architecture designed to transfer large amounts of data through a wide interface. It can be integrated close to a CPU, GPU, or other processing device. (Micron: high-bandwidth memory)
HBM uses multiple memory dies arranged vertically, with connections between them. This makes “stacked” a description of its physical organization, while “high bandwidth” describes its emphasis on data transfer. (Micron: high-bandwidth memory)
Do not confuse memory with the engine consuming that memory. A GPU performs computation; the HBM connected to it stores and supplies data. (Intel: CPU vs. GPU) (Micron: high-bandwidth memory) A useful first step when reading an accelerator diagram is to identify those two roles separately.
Capacity is not bandwidth.
Memory bandwidth describes a data-transfer rate. HBM achieves a high transfer rate through a wide, highly parallel memory interface. (Micron: high-bandwidth memory) Memory capacity instead describes the amount of information that can be stored.
How much data can be held
How much data can move each second
When comparing devices, write capacity and bandwidth on separate lines. Then ask what the application actually needs to store and move, rather than treating a larger value in one column as an answer to both questions.
How stacked memory connects.
HBM stacks connect memory dies using through-silicon vias and bonding structures. Package-level connections can place the stacks alongside a processor using an interposer. (Micron: high-bandwidth memory)
- DRAM diesMemory layers store data
- Vertical linksConnections pass between layers
- Wide interfaceData moves to and from the processor
There are two different levels of integration here. The memory stack is vertical; the larger package can arrange that stack beside a compute die. Samsung’s 2.5D example combines logic and HBM on a silicon interposer, illustrating how the two descriptions can apply together. (Samsung: heterogeneous integration)
For that reason, “3D memory” and “2.5D package” are not automatically contradictory labels. Read which structure each term is describing, then use the specific manufacturer drawing to understand the implementation. (Samsung: heterogeneous integration)
Why AI discussions mention HBM.
GPUs execute highly parallel AI workloads, and HBM is designed to supply high memory bandwidth near processing devices. (Intel: CPU vs. GPU) (Micron: high-bandwidth memory) The memory subsystem and the compute engine are therefore separate parts of understanding an accelerator.
HBM is not a synonym for all high-performance memory, and DRAM is not a synonym for permanent storage. DRAM needs refresh and loses its contents when power is removed; SRAM uses a different cell structure and does not require periodic refresh while powered. (Samsung: DRAM and SRAM)
- Identify the device: Record the HBM generation and exact manufacturer specification.
- Check the units: Distinguish capacity, bandwidth, and transfer rate per connection.
- Check the scope: Is a number for one stack, one device, or a whole system?
- Read the conditions: Separate interface capabilities from workload measurements.
Continue with chiplets and advanced packaging to see how multiple dies become one component. The companion CPU, GPU, and NPU guide explains the processing roles.
Keep the terms separate.
Is HBM an AI chip?
HBM is memory, not the processor executing a neural network. It can supply data to an accelerator inside a closely integrated package. (Micron: high-bandwidth memory)
Is HBM a type of DRAM?
Yes. HBM uses stacked DRAM dies and a wide memory interface; it is an architecture within the broader DRAM family. (Micron: high-bandwidth memory)
Does more bandwidth mean more capacity?
No. A capacity figure describes an amount of data; a bandwidth figure describes a rate. Keep the two quantities separate when reading the stack specifications. (Micron: high-bandwidth memory)