CUDA cores vs Tensor cores
CUDA cores run general parallel math one operation at a time. Tensor cores complete a small matrix multiply per clock for AI math.
CUDA cores are Nvidia's general-purpose parallel units, each a small arithmetic engine executing one thread's instruction per clock in FP32, FP64, or integer math. Tensor cores are separate hardware, added with the Volta architecture in 2017, and each performs a whole small matrix multiply-accumulate per clock instead of one scalar operation. That is exactly the math a neural network repeats millions of times per pass.
On an A100, the 6,912 CUDA cores deliver about 19.5 TFLOPS of FP32. The 432 tensor cores on the same die reach 312 TFLOPS at FP16, roughly 16 times more, since one tensor instruction replaces dozens of scalar operations. The catch is precision. Tensor hardware only applies to matrix-heavy work at FP16, BF16, TF32, INT8, and now FP4 on Blackwell parts. The general cores still handle rendering, simulation, and anything needing full FP32 or FP64 accuracy.
A spec sheet quoting AI TOPS or tensor TFLOPS is describing tensor throughput, not general compute. For training and inference, tensor core generation and count are the numbers that matter. For rendering, CAD, and simulation, compare CUDA core count and FP32 or FP64 throughput instead.
Sources
Source | Publisher |
|---|---|
NVIDIA | |
NVIDIA | |
NVIDIA |
- Publisher
NVIDIA
- Publisher
NVIDIA
- Publisher
NVIDIA
Last verified August 29, 2026.
- CUDA core count
- Tensor core
- matrix multiply accumulate
- FP16 throughput
- TFLOPS
- mixed precision
- TF32
- AI TOPS
- Nvidia GPU architecture