Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In a September 2026 benchmark, Gemma 4 E2B, E4B, and 12B decoded single requests on an Amazon SageMaker NVIDIA T4 at roughly 0.8 times the rate reported for an L4. With 16 simultaneous requests, the T4’s relative throughput fell further, to 0.53–0.63 times the L4 figures. The tested T4 and L4 outputs matched byte for byte on a 40-question, temperature-zero check, but that small test does not establish general quality equivalence.

What the T4-versus-L4 results show

The benchmark author reported the following single-request decode speeds for three Gemma 4 variants. The L4 numbers come from a related benchmark run the day before, not from a simultaneous, controlled comparison.

Model T4 decode L4 decode T4/L4 ratio
E2B 108.5 tokens/s 141.7 tokens/s 0.77x
E4B 65.6 tokens/s 79.9 tokens/s 0.82x
12B 28.5 tokens/s 35.0 tokens/s 0.81x

At 16 parallel requests, the T4/L4 throughput ratios were 0.63 for E2B, 0.57 for E4B, and 0.53 for 12B. That difference matters if your endpoint serves concurrent traffic: the roughly 0.8x single-request headline does not predict the relative performance under load. These figures are the benchmark author’s measurements, not a guarantee for another deployment or workload. Benchmark details

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How comparable were the benchmark runs?

The T4 tests used a SageMaker ml.g4dn.xlarge instance in us-east-2, vLLM 0.30.0 in an AWS container modified with a Turing patch, FP16, and a host image using driver 580. The benchmark author reports one account and one deployment per model on September 30, 2026.

The related L4 results used ml.g6.xlarge on September 29, 2026, with the stock container and BF16. Because the instance, date, container configuration, and data type differed, this is not a same-software hardware-only comparison. Throughput testing used the AWS CLI from one client machine. The measurements covered single-request decode, throughput up to 16 parallel requests, and a 40-question correctness check at temperature zero. Benchmark details

What the matching answers do—and do not—mean

For each of E2B, E4B, and 12B, the benchmark reports byte-for-byte matching T4 and L4 outputs on its 40 questions. Its correctness table records 36/40 for E2B, 36/40 for E4B, and 40/40 for 12B. Those results describe that sample and setup; they cannot establish that the devices produce equivalent answers across broader prompts, settings, or workloads.

Which Gemma 4 models fit on one T4?

The tested E2B, E4B, and 12B configurations served on a single T4. The 26B-A4B attempt loaded its weights but then failed with an out-of-memory error, making 12B the largest model the author successfully served on one T4 in this configuration. This is a result for the reported setup, not a universal capacity limit for every quantization, context length, or serving stack.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For 12B, the reported T4 KV cache held 20,354 tokens—about 2.48 requests at the configured 8,192-token context. The L4 table reported a larger KV cache for each of the three tested sizes. Cache capacity affects how many tokens and concurrent requests can be kept in memory; do not treat a model’s weight fit alone as proof that a target context and concurrency will fit.

Hourly price versus cost per generated token

The benchmark author reported on-demand SageMaker rates in us-east-2 from an AWS Price List API lookup on September 30, 2026. The per-million-token figures below are the author’s calculations using those hourly rates and measured throughput at 16 parallel requests; they are not universal prices or forecasts.

Model T4 hourly rate L4 hourly rate T4 cost per million output tokens L4 cost per million output tokens
E2B $0.736 $1.1267 $0.260 $0.249
E4B $0.736 $1.1267 $0.419 $0.369
12B $0.736 $1.1267 $0.945 $0.760

The lower T4 hourly rate may suit light traffic where an instance spends substantial time idle. In the benchmark’s 16-request calculations, however, the L4 had lower cost per million output tokens for all three models. Verify current AWS rates, availability, and capacity in your region before making a cost or architecture decision. Benchmark and pricing methodology

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Deployment details to verify before reproducing the T4 setup

The benchmark attributes T4 support to a Turing patch in a derived vLLM container. It also says the CUDA 13 container requires InferenceAmiVersion to be set to the driver-580 host on ml.g4dn. These are version-sensitive implementation details, not a general AWS-supported deployment recipe. Check current SageMaker API, container, driver, and vLLM compatibility before adapting them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AWS announced SageMaker JumpStart availability for Gemma 4 E4B, 26B-A4B, and 31B on April 29, 2026, with deployment through SageMaker Studio or the SageMaker Python SDK. That announcement establishes a JumpStart path for those named variants; it does not establish that the benchmark’s custom T4 configuration is supported through JumpStart. AWS also notes E4B audio input capabilities. AWS announcement

A separate AWS Builder Center article tested Gemma 4 QAT formats on SageMaker L4 instances, using vLLM 0.30.0 in us-east-2. It reported results for E2B, E4B, 12B, 26B-A4B, and 31B; its 4-bit embeddings and lm_head setup decoded 1.12–1.39 times faster than the compared 16-bit embeddings setup and matched answers on that article’s test for the sizes shown. This is separate implementation context, not independent validation of the T4-versus-L4 comparison. AWS Builder Center QAT benchmark

How to choose between T4 and L4 for this workload

  • For mostly one-at-a-time requests: Use the reported 0.77–0.82 T4/L4 decode ratios as a starting point, while accounting for the different run dates and software configurations.
  • For concurrent serving: Compare throughput at the concurrency you expect. In this benchmark, the T4’s relative throughput was 0.53–0.63 of the L4 at 16 parallel requests.
  • For model fit: Validate weights, context length, KV-cache headroom, and concurrency together. The reported 26B-A4B T4 attempt ran out of memory, while 12B served.
  • For cost: Compare both instance-hours and output-token cost at realistic utilization. The hourly-rate advantage and measured 16-way cost-per-token advantage went to different instances.
  • For reproducibility: Match or explicitly account for region, instance type, date, driver, data type, container, vLLM version, and any Turing patch.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.