Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In one reported test serving google/gemma-4-E2B-it with a hand-written pure-JAX implementation, an AWS g6.2xlarge produced 48.4 decode tokens per second versus 12.9 on a g5g.2xlarge—about 3.7 times the measured decode throughput. That is a result for this model, implementation, configuration, and test setup, not a general promise that g6 will serve every LLM 3.7 times faster.

What the benchmark compared

The benchmark author used spot instances and the same byte-identical payload on both: build 51bc52c9e2e9, configuration ple4 + int8_lm_head, and tpu_jax_weight_bytes of 6,155,450,950. The model was google/gemma-4-E2B-it, served using a hand-written pure-JAX port without PyTorch, vLLM, or torch_xla. The figures below are the author’s reported measurements, not independently reproduced results. Read the benchmark and its setup.

Measure or setup g5g.2xlarge g6.2xlarge
GPU NVIDIA T4G, Turing, SM 7.5 NVIDIA L4, Ada, SM 8.9
Host architecture Graviton2, aarch64 x86_64
Reported decode gauge throughput across prompt sweeps 12.9–13.0 tokens/s 48.3–48.5 tokens/s
Headline decode comparison 12.9 tokens/s 48.4 tokens/s; approximately 3.7× g5g

The g6 measurement was run in us-east-1d. The common model payload and configuration make this a useful observed comparison, but they do not make it a controlled test of the GPU alone: the host architectures and base images differed as well.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Decode throughput is not end-to-end throughput

The headline compares decode gauge rates. End-to-end rates include prompt prefill, so they were lower and declined as the input prompt grew. The author attributes that decline to prefill processing. These reported rates show why a decode number by itself may not predict the speed a user experiences on requests with long prompts.

#1 Best Overall
Yahboom RDK X5 4GB Development Board Kit Computing Power Openclaw Support Servo Control for Mechanical Engineers
  • 【Core parameters】★AI performance: 10TOPS★CPU: 8 octa-core Cortex A55 @ 1.5GHZ ★GPU: 32GFLOPS ★Memory: 4GB/8GB ★Power consumption: MAX 25W ★YOLOv5 algorithm frame rate: High performance mode: 28~30fps
  • 【Out-of-the-box Ready, Flexible Configuration】We provide a complete kit for developers from beginner to advanced, including: board, aluminum case, MIPI camera, binocular depth camera, IMU inertial navigation module, LiDAR, power supply, mouse, keyboard, display, AI voice module, and more. No need to purchase additional compatible accessories — get started with your project development right away.
  • 【Strong Compatibility】It comes with a variety of compatible accessories. The aluminum case comes with a cooling fan, which is wear-resistant and effectively dissipates heat and protects the RDK X5. The IMX219 camera/depth camera provides AI visual images and depth images. The radar supports ROS2 mapping, navigation and tracking. The 7-inch IPS HD touch display supports RDK X5/Raspberry Pi 5/Jetson series development boards. A 64GB TF card is provided with Ubuntu-related image files.
  • 【Support LLM】RDK X5 development board supports many leading large models such as DeepSeek-R1, Qwen, Gemma, etc. Users can realize multi-modal recognition of pictures and texts through the RDK large model gateway; support local deployment of DeepSeek-R1 large model to achieve efficient and low-latency AI reasoning. Greatly improve response speed and stability, and give smart devices more powerful autonomous decision-making capabilities.
  • 【Tutorials provided】Provide innovative solutions for the robot era, support multiple complex models and the latest algorithms such as Transfomer, RWKV, Occupancy, Stere0, Perception, etc., and accelerate the rapid implementation of intelligent applications; Yahboom provides data tutorials for development boards and related accessories.
Input prompt length g5g end-to-end g6 end-to-end
41 tokens 12.43 tokens/s 46.23 tokens/s
521 tokens 11.28 tokens/s 42.87 tokens/s
2,057 tokens 8.22 tokens/s 34.57 tokens/s
3,593 tokens 27.55 tokens/s

The benchmark reports a g6 end-to-end rate at 3,593 input tokens but no matching g5g value for that prompt length. The reported decode gauge stayed comparatively stable across the listed sweeps; end-to-end throughput did not.

Why the author says g5g lagged

For this pure-JAX serving implementation, the author’s profiler analysis attributed 87% of g5g decode time to dtype conversion and an fp32 path, compared with 0% on g6. The author also reported g5g reaching 26% of its memory-bandwidth roofline and g6 roughly 100%. These are the author’s interpretations of this run and implementation, not general performance characteristics for every workload on those instance families.

Rank #2
Yahboom ROS2 Python Programming Robot Kit Jetson Powered Educational Robot AI Voice Module, AI Vision Mapping Navigation 360°Move for Mechanical Engineers(Sup Ver-with Orin NX Super 16GB)
  • ROS robotic learning kit for multiple versions: Yahboom provides 4 development board versions of ROSMASRER X3, you can freely choose jetson series development board or Raspberry Pi 5, based on the different performance issues of these development boards, The smoothness of operation is worth considering, Fully compatible with Jetson Orin SUPER Kit.
  • In-depth exploration of AI algorithms and intelligent robots: ROSMASRER X3 is equipped with a depth camera, lidar, and voice interaction module, which can realize ROS operating system, RTAB 3D mapping navigation, PCL 3D point cloud, SLAM mapping navigation, Machine vision applications, Voice interactive control, Python programming, STM32 development, MediaPipe development, YOLO model training, TensorRT acceleration (Note: Different features depend on the version you choose)
  • Rich course materials and professional after-sales support team: We provides 103 dual-language video courses, and online technical assistance (China time). The course content includes: ROSMASTER X3 assembly, Linux operating system, ROS and openCV series courses, depth camera and lidar mapping and navigation explanation, from simple to in-depth learning of mapping and navigation, this is an in-depth learning process, but we recommend that there are Programming basic users to use this robot kit
  • Multi-platform linkage: rosmaster X3 supports a variety of remote control methods such as mobile phone APP, handle, ROS system, computer keyboard, etc. It can control your robot car at any time, import your code, and is an artificial intelligence robot that listens to your instructions. Note: The Map Navigation APP only supports Android phones
  • Application field: rosmaster X3 provides an exploration model for professionals, can learn algorithms, obtain terrain in an unknown field, can deeply learn AI visual recognition, research autonomous driving, explore 3D object recognition, etc.Fully upgraded the ROS2 course.

Tensor Core utilization was reported as zero on both GPUs; the author said the reason was unexplained. That result is a reminder that having accelerator hardware does not guarantee a particular serving stack will use every relevant feature efficiently.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How strong is the 3.7× conclusion?

It is a meaningful result for deciding what to test next, but not conclusive evidence of a universal g6 advantage of that size. The author reports one run per instance, with three repeats per sweep cell and medians. The g5g profile was reproduced on a second instance; the g6 profile was measured once. Alongside the GPU, the machine architecture and base image differed.

Rank #3
MPC 1969 Plymouth Barracuda 1:25 Scale Model KIt
  • 1:25 scale, skill level 2, paint & glue required
  • 129 parts
  • Molded in white, clear, transparent red, with some chrome-plated parts
  • Black vinyl tires
  • Metal Axles

For a deployment decision, test the exact model and serving stack you plan to run. Keep the workload and software configuration consistent, and record both decode and end-to-end throughput. Include prompt length, output length, precision, concurrency, latency, region, and repeated runs in the comparison; otherwise, a faster decode gauge may not translate to faster or cheaper production requests.

What the test says about context length

On g5g, the author used MAX_MODEL_LEN=4096. A prompt of 4,105 tokens served, while a 5,120-token prompt failed on a prefill transient. This is an observation from the described test, not a platform-wide context limit or guarantee for other configurations.

Rank #4
AMT Skill 2 Model Kit 1966 Plymouth Barracuda Funny Car Hemi Under Glass 1/25 Scale Model
  • AMT
  • Accurate Scale Model Kit
  • Detailed Pictorial Instructions Included
  • Will Require Paints, Glue and Modelling Tools to Complete
  • Glue and Paint needed for assembly - NOT INCLUDED

Does g6 cost less per token?

The benchmark did not measure instance price or price per token, and its throughput result cannot answer that question by itself. Cost depends on current pricing and the workload’s useful output, latency, utilization, and deployment conditions. AWS’s separate SageMaker comparison illustrates the additional measurements needed by reporting throughput, latency, and cost per output token for its own tested models and workloads; those results do not establish the cost or value of this EC2 g5g/g6 test. See AWS’s separate SageMaker benchmark.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What to verify before choosing an instance

Use the benchmark as a reason to run your own workload test, then confirm the deployment details that can change over time. AWS’s EC2 accelerated-computing instance specifications are the official reference for instance family details. Check current regional availability, account quotas, and prices for your intended region and purchase option: the benchmark’s spot capacity and region do not establish what will be available or cost-effective elsewhere.

Quick Recap

Bestseller No. 3
MPC 1969 Plymouth Barracuda 1:25 Scale Model KIt
MPC 1969 Plymouth Barracuda 1:25 Scale Model KIt
1:25 scale, skill level 2, paint & glue required; 129 parts; Molded in white, clear, transparent red, with some chrome-plated parts
$29.50
Bestseller No. 4
AMT Skill 2 Model Kit 1966 Plymouth Barracuda Funny Car Hemi Under Glass 1/25 Scale Model
AMT Skill 2 Model Kit 1966 Plymouth Barracuda Funny Car Hemi Under Glass 1/25 Scale Model
AMT; Accurate Scale Model Kit; Detailed Pictorial Instructions Included; Will Require Paints, Glue and Modelling Tools to Complete
$34.00
SaleBestseller No. 5
Bendix Talos Missile Anti-Aircraft Missile with Three Figures 1/40 Scale Model Kit
Bendix Talos Missile Anti-Aircraft Missile with Three Figures 1/40 Scale Model Kit
1/40 scale unpainted plastic model assembly kit; Re-release kit released from the level in 1957
$23.37
Best Value
Sale
Bendix Talos Missile Anti-Aircraft Missile with Three Figures 1/40 Scale Model Kit
  • This plastic model requires assembly and painting. Adhesives, tools, paints, etc., sold separately
  • 1/40 scale unpainted plastic model assembly kit
  • Re-release kit released from the level in 1957
  • Mold Color: White, Olive Drab
  • Atlantis Modelive (USA) Imported Plastic Model

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.