What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google Cloud Run supports AI inference on NVIDIA L4 and NVIDIA RTX PRO 6000 Blackwell GPUs, attached to containerized services as managed capacity. Each instance gets one GPU, and the service can scale to zero. GPU workloads require instance-based billing, however, and GPU time is billed for the full instance lifecycle—not only while a request is being processed. Current configuration details and supported regions are documented in Google Cloud’s Cloud Run GPU documentation.

What Cloud Run GPU provides

Cloud Run adds a supported NVIDIA GPU to a containerized service instance. Google manages the GPU drivers, and its documentation describes GPU capacity as available on demand without reservations. GPU-backed services can run LLM inference and other compute-intensive tasks such as video transcoding and 3D rendering.

The service documentation lists two GPU options. Their VRAM is separate from the instance’s system memory:

GPU VRAM Minimum service configuration Documented regions
NVIDIA L4 24 GB 4 CPU and 16 GiB memory asia-southeast1, asia-south1 (invitation only), europe-west1, europe-west4, us-central1, us-east4
NVIDIA RTX PRO 6000 Blackwell 96 GB 20 CPU and 80 GiB memory asia-southeast1, asia-south2, europe-west4, us-central1

These configuration and region details come from Google’s current service documentation, accessed in 2026; check the live page for updates and project-specific capacity or quota caveats. Both GPU types are documented with NVIDIA driver version 580.x.x (13.0). One GPU is available per instance. In a sidecar setup, only one container can have the GPU attached.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
  • AI Performance: 767 AI TOPS
  • OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode)
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • A 2.5-slot design maximizes compatibility and cooling efficiency for superior performance in small chassis

Google says a supported GPU instance with its preinstalled drivers can start in approximately five seconds, meaning the container processes can use the GPU. That figure is not an end-to-end model-serving cold-start time: model download or loading and inference can add time.

Choose a GPU based on model fit and deployment constraints

NVIDIA L4

The L4 has 24 GB of documented VRAM and a lower minimum CPU and memory configuration than the RTX PRO 6000 Blackwell. Its listed regions include four locations also listed for Blackwell—asia-southeast1, europe-west4, and us-central1—plus europe-west1, us-east4, and invitation-only asia-south1. Whether it suits a model depends on that model’s memory needs and your project’s available quota and capacity.

NVIDIA RTX PRO 6000 Blackwell

Blackwell has 96 GB of documented VRAM and requires at least 20 CPU and 80 GiB of instance memory. Its documented region list is shorter than L4’s, so confirm location support before selecting it. More VRAM may accommodate workloads that do not fit within L4’s documented VRAM, but these specifications alone do not establish a universal performance or cost advantage.

Rank #2
msi Gaming RTX 3050 Ventus 2X 6G OC Graphics Card (NVIDIA RTX 3050, 96-Bit, Boost Clock: 1492 MHz, 6GB GDDR6 14 Gbps, HDMI/DP, Ampere Architecture)
  • Chipset: GeForce RTX 3050
  • Boost Clock / Memory: 1492 MHz / 14 Gbps
  • Video Memory: 6GB GDDR6
  • Memory Interface: 96-bit
  • Output: DisplayPort x 1 (v1.4a) / HDMI 2.1a x 2

Compare the options using your model’s memory requirements, the minimum instance configuration, supported regions, project quota and current pricing. Google’s documentation does not establish a universal price or performance winner.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check region, quota and capacity before deploying

GPU region support and quota are not interchangeable with guaranteed physical capacity. Google’s documentation says initial quota is granted by region on first deployment using the documented non-zonal-redundancy configuration: up to three L4 GPUs, or the equivalent of three RTX PRO 6000 Blackwell GPUs (3,000 milliGPUs). Larger needs require a quota increase. The stated initial quota does not guarantee that capacity will be available under every demand condition.

  • Check the current GPU region list against the region you plan to use; the two GPU types have different lists.
  • Check quota in the project and region where the service will run. Some regions have additional capacity or quota caveats.
  • For a deployment larger than the initial documented allowance, request a quota increase and confirm capacity before relying on it.

Understand billing and zonal redundancy

GPU services must use instance-based billing. Google bills GPU time for the full instance lifecycle, and minimum instances are charged at the full rate while idle. The GPU feature has no per-request GPU fee. Scaling to zero can avoid keeping an instance running between periods of demand, but it does not change the lifecycle billing rule for an instance that is running.

Rank #3
GIGABYTE GeForce RTX 5070 WINDFORCE OC SFF 12G Graphics Card, 12GB 192-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N5070WF3OC-12GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5070
  • Integrated with 12GB GDDR7 192bit memory interface
  • PCIe 5.0
  • NVIDIA SFF ready

Cloud Run’s documented default enables GPU zonal redundancy. The setting changes the balance between GPU availability during zonal disruption and GPU-second cost:

Setting Capacity and failover GPU-second cost
Zonal redundancy enabled Capacity is reserved across multiple zones to improve the chance of handling traffic shifted after a zonal outage. Higher than with redundancy disabled.
Zonal redundancy disabled Failover is best-effort and depends on unused GPU capacity being available. Lower than with redundancy enabled.

The applicable service SLA depends on the redundancy configuration. Google’s current documentation should be consulted for the applicable terms and current prices; no dollar amount is stated here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Deploy an inference service using Google’s example

Google’s Gemma 4 and vLLM Cloud Run codelab, updated May 7, 2026, demonstrates serving Gemma 4 E2B with an RTX PRO 6000 Blackwell GPU. Its example configures 20 CPU, 80 GiB of memory, one GPU, disabled GPU zonal redundancy, a service account, no unauthenticated access, and a startup probe. It enables the Cloud Run, Cloud Build and Artifact Registry APIs.

Rank #4
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4. System Requirements: Minimum 850W PSU with 16-pin 12V-2x6 (12VHPWR) connector required. Verify before purchasing.
  • Military-grade components deliver rock-solid power and longer lifespan for ultimate durability. Compatibility: 348mm (13.7") length, 3.6 slots, 4.3 lbs. Confirm case clearance and slot spacing. GPU bracket included.
  • Protective PCB coating helps protect against short circuits caused by moisture, dust, or debris
  • 3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans
  • Phase-change GPU thermal pad helps ensure optimal thermal performance and longevity, outlasting traditional thermal paste for graphics cards under heavy loads
  1. Confirm prerequisites. Choose a supported region, verify quota and capacity for the project, and check that the selected GPU’s documented minimum resources fit the service.
  2. Enable the required APIs. Follow the codelab’s steps to enable Cloud Run, Cloud Build and Artifact Registry for the project.
  3. Configure the service and container. Use the codelab’s vLLM and Gemma 4 E2B example as a starting point, including its service account, authentication setting and startup probe.
  4. Review the GPU and reliability settings. The example uses one RTX PRO 6000 Blackwell GPU and disables zonal redundancy. Choose that setting only if best-effort GPU failover is acceptable for your service.
  5. Validate the example against your workload. Before deployment, confirm current image tags and flags, region support, model requirements and project quota. A tutorial configuration is not a guarantee of performance or availability for another model or production workload.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What Google’s launch announcements do—and do not—show

Google’s August 21, 2024 announcement described the NVIDIA L4 GPU as a preview feature. Later, Google announced general availability and reported approximately 19 seconds to first token for a specific Gemma 3 4B example, including startup, model loading and inference. That is a vendor-reported result for that workload, not a general latency guarantee or an independent benchmark. Current service documentation now lists both L4 and RTX PRO 6000 Blackwell; use it for present-day configuration and region details.

The 2024 announcement quoted Anne Hecht, then NVIDIA’s Senior Director of Product Marketing, saying: “With the addition of NVIDIA L4 Tensor GPU and NVIDIA NIM support, Cloud Run provides users a real-time, fast-scaling AI inference platform to help customers accelerate their AI projects and get their solutions to market faster — with minimal infrastructure management overhead.” This is a vendor statement in Google Cloud’s announcement, not an independent performance assessment.

Google’s general-availability announcement also quoted Dave Salvator, director of accelerated computing products at NVIDIA: “Serverless GPU acceleration represents a major advancement in making cutting-edge AI computing more accessible. With seamless access to NVIDIA L4 GPUs, developers can now bring AI applications to production faster and more cost-effectively than ever before.” Treat that as an attributed vendor quotation, not evidence that Cloud Run is the least expensive option for every workload.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

SaleBestseller No. 1
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
AI Performance: 767 AI TOPS; OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode); Powered by the NVIDIA Blackwell architecture and DLSS 4
$792.02
Bestseller No. 2
msi Gaming RTX 3050 Ventus 2X 6G OC Graphics Card (NVIDIA RTX 3050, 96-Bit, Boost Clock: 1492 MHz, 6GB GDDR6 14 Gbps, HDMI/DP, Ampere Architecture)
msi Gaming RTX 3050 Ventus 2X 6G OC Graphics Card (NVIDIA RTX 3050, 96-Bit, Boost Clock: 1492 MHz, 6GB GDDR6 14 Gbps, HDMI/DP, Ampere Architecture)
Chipset: GeForce RTX 3050; Boost Clock / Memory: 1492 MHz / 14 Gbps; Video Memory: 6GB GDDR6
$259.97
Bestseller No. 3
GIGABYTE GeForce RTX 5070 WINDFORCE OC SFF 12G Graphics Card, 12GB 192-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N5070WF3OC-12GD Video Card
GIGABYTE GeForce RTX 5070 WINDFORCE OC SFF 12G Graphics Card, 12GB 192-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N5070WF3OC-12GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5070; Integrated with 12GB GDDR7 192bit memory interface
$1,000.53
Bestseller No. 4
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans; Auto-Extreme precision automated manufacturing helps ensure higher reliability
$1,831.31

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.