Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

NVIDIA’s GP100 is the high-end compute GPU in the Pascal generation, built for workloads that benefit from high memory bandwidth and strong double-precision performance. The Tesla P100 is its best-known data-center implementation: NVIDIA listed 16GB of HBM2, 720GB/s of memory bandwidth, and 5.3 TFLOPS of FP64 throughput for that configuration. GP100 also appeared in the Quadro GP100 workstation card.

What GP100 is—and how it relates to Tesla P100

GP100 is a GPU architecture and chip design, not the name of a single retail card. NVIDIA’s full GP100 design has 60 streaming multiprocessors (SMs), while the Tesla P100 configuration uses 56. That distinction matters when comparing the complete chip’s architecture counts with the specifications of a particular product.

The full design groups its SMs into six graphics processing clusters (GPCs) and 30 texture-processing clusters. It has 3,840 FP32 CUDA cores, eight 512-bit memory controllers for a 4,096-bit aggregate interface, and 4MB of L2 cache. The Tesla P100 and Quadro GP100 are products built around GP100; the figures for the full chip should not be mistaken for a guarantee that every product exposes every resource.

Tesla P100 specifications at a glance

NVIDIA’s 2016 launch specifications for the Tesla P100 included the following figures. These are manufacturer-listed peak performance and bandwidth specifications, not a promise that an application will sustain those rates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Nvidia Tesla P100 900-2H400-0000-000 GPU Computing Processor - 16 GB - HBM2 - PCIE 3.0 X16 (Certified Refurbished)
  • GPU Computing Processor
  • 16GB HBM2
  • PCIe 3.0 x16
  • Fanless - Passive Cooling
  • 3584 CUDA Cores
Specification Tesla P100 figure What it describes
GPU configuration 56 SMs The P100 implementation, rather than the full 60-SM GP100 design
Memory 16GB HBM2 High-bandwidth on-package memory
Memory bandwidth 720GB/s Peak bandwidth listed by NVIDIA
FP64 performance 5.3 TFLOPS Peak double-precision throughput listed by NVIDIA
FP32 performance 10.6 TFLOPS Peak single-precision throughput listed by NVIDIA
FP16 performance 21.2 TFLOPS Peak half-precision throughput listed by NVIDIA
NVLink bandwidth 160GB/s bidirectional The listed GPU interconnect bandwidth; useful in multi-GPU system designs

Why GP100 is notable for double precision

Each GP100 SM contains 64 FP32 CUDA cores and 32 FP64 units. That gives the architecture a 2:1 single-to-double-precision throughput ratio: its theoretical FP64 rate is half its FP32 rate. NVIDIA contrasted this with a 3:1 ratio in the earlier Kepler GK110 design, making GP100 unusually capable in FP64 for its generation.

For scientific simulation, numerical modeling, and other workloads that genuinely execute many double-precision operations, this ratio is more informative than the CUDA-core count alone. A large FP64 peak still does not guarantee application speed: the code must use the relevant precision, keep the GPU supplied with work, and avoid bottlenecks elsewhere in the system.

What HBM2 bandwidth does—and does not—mean

The P100’s 720GB/s figure reflects its HBM2 memory subsystem and wide memory interface. High bandwidth can benefit kernels that repeatedly move substantial data and are limited by memory traffic. It cannot, by itself, make every program faster: a compute-bound workload, inefficient access pattern, synchronization overhead, or slow host-side data preparation may limit performance instead.

Memory capacity is a separate constraint from bandwidth. The P100’s 16GB can hold a working set only if the application’s data and intermediate results fit within that capacity, or can be managed effectively through transfers and partitioning. The listed 160GB/s bidirectional NVLink bandwidth concerns GPU-to-GPU communication; it does not replace local memory capacity or guarantee linear scaling across multiple GPUs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pascal features relevant to compute workloads

NVIDIA’s Pascal technical overview describes several GP100 capabilities beyond raw throughput:

  • Native FP16 arithmetic: supports half-precision operations, which can be useful when a workload and its accuracy requirements permit reduced precision.
  • Unified Memory improvements: hardware page faulting and a 49-bit virtual address space support managed-memory behavior across CPU and GPU address spaces. They do not mean that all data movement is free or that a workload’s memory footprint is unlimited.
  • FP64 atomic add: enables supported atomic addition operations on double-precision values.
  • Compute preemption: provides hardware support for preempting compute work, relevant to scheduling and responsiveness.

GP100 is a Pascal-generation device with CUDA compute capability 6.x. The CUDA tuning documentation also describes ECC-protected memory structures for this generation. Exact feature availability and behavior depend on the product, driver, CUDA software stack, and system configuration.

Rank #4
NVIDIA 900-2G610-0000-000 Tesla P40 24GB GDDR5 PCIE 3.0 X16 Passive Cooling
  • Series: Tesla P40, Model: 900-2G610-0000-000
  • GPU Architecture: NVIDIA Pascal, Single-Precision Performance:12 TeraFLOPS
  • Integer Operations (INT8):47 TOPS (Tera-Operations per Second), GPU Memory:24 GB
  • Memorty Bandwidth:346 GB/s, System Interface:PCI Express 3.0 x16
  • Max Power:250W, Enhanced Programmability with Page Migration Engine:Yes, ECC Protection:Yes, Server-Optimized for Data Center Deployment:Yes, Hardware-Accelerated Video Engine:1x Decode Engine, 2x Encode Engine

Which products use GP100?

NVIDIA Tesla P100

The Tesla P100 is the principal data-center accelerator based on GP100. NVIDIA’s 2016 specifications list 16GB HBM2, 720GB/s memory bandwidth, and 160GB/s bidirectional NVLink bandwidth. P100 systems and cards were offered in different physical configurations, so a used-unit listing’s PCIe or SXM form factor must be checked rather than inferred from the GPU name.

NVIDIA Quadro GP100

NVIDIA also announced the Quadro GP100 for professional workstations, listing 16GB HBM2 and support for connecting two cards with NVLink for 32GB across the pair. That aggregate figure should not be read as an assurance that software sees one automatically pooled 32GB memory space; application and system support determine how multiple GPUs and their memory are used.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
COMeap NVIDIA Graphics Card Power Cable 030-0571-000 CPU 8 Pin Male to Dual PCIe 8 Pin Female Adapter for Tesla K80/M40/M60/P40/P100 4.7 inches (2-Pack) (12cm)
  • 『CPU 8P - Dual PCIe 8P』CPU 8 pin male end to plug into the NVIDIA graphics card, dual PCIe 8 pin female ends to plug into the 8 pin(6+2) connector of power supply;
  • 『Compatibility』Compatible with Tesla K80/M40/M60/P40/P100, 170hx nvidia cmp other NVIDIA graphics card with CPU 8 pin port, etc.;
  • 『Note』The 8 pin male end is CPU 8 pin, not pci-e 8 pin, which was only designed for NVIDIA graphics card with CPU 8 pin port. If you connect it with other incompatible devices, it will definitely burn or damage the motherboards, PSUs or graphics cards and we won’t take any responsibility for wrongly using or installing. Please carefully check the compatible types or contact us if you are not sure about it;
  • 『Parameter』Length(including connectors): 4-inch(10cm), Gauge: 1007-16AWG(standard tin-coating copper wire), Maximum power: 600W, Quantity:2pcs, Self-adhesive tape*1pcs;
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to judge whether a P100 fits a workload

Start with the actual bottleneck and software requirements rather than the headline TFLOPS figure. A P100 is a more plausible candidate when its precision support, capacity, and system integration align with the job.

  • Precision: Check whether the application is FP64-bound, FP32-bound, or able to use FP16 without unacceptable loss of accuracy.
  • Memory behavior: Estimate the working-set size and determine whether performance is limited by bandwidth, capacity, or data access patterns.
  • Scaling: For multi-GPU work, verify that the application uses NVLink appropriately and measure communication overhead as well as compute.
  • Software support: Confirm CUDA compute capability 6.x support in the specific application, framework, compiler, and driver combination you intend to run.
  • Host platform: Verify card form factor, chassis fit, power delivery, cooling, and system support. Historical GPU specifications do not establish compatibility with a particular modern server or workstation.

Is Tesla P100 still useful for HPC?

It can remain useful for a compatible HPC or technical-computing workload that benefits from Pascal’s FP64 throughput, HBM2 bandwidth, or NVLink, especially where an existing system and software stack support it. That is a workload- and system-specific judgment, not a general recommendation: the cited launch specifications do not establish current operating-system support, 2026 availability or pricing, seller condition, or compatibility with any particular chassis. For a second-hand accelerator, establish the exact form factor and host requirements before purchase, then confirm that the target software still supports the relevant Pascal device.

Quick Recap

Bestseller No. 1
Nvidia Tesla P100 900-2H400-0000-000 GPU Computing Processor - 16 GB - HBM2 - PCIE 3.0 X16 (Certified Refurbished)
Nvidia Tesla P100 900-2H400-0000-000 GPU Computing Processor - 16 GB - HBM2 - PCIE 3.0 X16 (Certified Refurbished)
GPU Computing Processor; 16GB HBM2; PCIe 3.0 x16; Fanless - Passive Cooling; 3584 CUDA Cores
$169.99
SaleBestseller No. 2
Bestseller No. 4
NVIDIA 900-2G610-0000-000 Tesla P40 24GB GDDR5 PCIE 3.0 X16 Passive Cooling
NVIDIA 900-2G610-0000-000 Tesla P40 24GB GDDR5 PCIE 3.0 X16 Passive Cooling
Series: Tesla P40, Model: 900-2G610-0000-000; GPU Architecture: NVIDIA Pascal, Single-Precision Performance:12 TeraFLOPS
$345.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.