What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Meta’s TLX-based Jagged Flash Attention (JFA) kernel is designed for packed, variable-length sequences on NVIDIA Blackwell. In the PyTorch Blog article published October 1, 2026, its authors report that, on their tested BF16 B200 jagged workloads, JFA averaged about 13% faster forward and about 50% faster backward than the May 2026 FlashAttention-4 (FA4) implementation. Those results describe a particular production-style workload and benchmark setup—not a universal advantage across attention workloads or Blackwell GPUs.

What jagged attention represents

Jagged attention operates on variable-length sequences stored consecutively in packed query, key, and value tensors. Rather than padding each sequence to the length of the longest one, the operation uses offsets to identify where each sequence begins and ends. The public JFA API documents query and key/value offsets as prefix sums, each with one more element than the batch size.

That representation can avoid computation on padding tokens. The PyTorch article says padding can waste up to 50% of compute in the GEM context, citing an earlier training-system explanation; that figure is workload context, not a general estimate for all jagged workloads.

What the B200 comparison shows—and what it does not

The PyTorch Blog authors compared their TLX implementation with the May 2026 FA4 implementation using BF16 on B200 GPUs. They report separate outcomes for a production-style Hierarchical Seed Pooling (HSP) broadcast-query jagged workload and an equal-length, LLM-style dense workload. The averages below belong to those reported benchmark regimes; they should not be read as like-for-like results across all attention shapes.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
NVD RTX PRO 6000 Blackwell Professional Workstation Edition Graphics Card for AI, Design, Simulation, Engineering - 96GB DDR7 ECC Memory - 4th Gen RT/5th Gen Tensor Core GPU - OEM Packaging
  • PLEASE NOTE: Exporting an NVIDIA RTX Pro 6000 GPU outside the US requires strict adherence to the U.S. Export Administration Regulations (EAR) and issuance of an export license from the Bureau of Industry and Security (BIS). Compliance and Know Your Customer (KYC) screening may be required as a condition of order acceptance. [NVIDIA Blackwell Streaming Multiprocessor] The new SM features increased processing throughput, and new neural shaders that integrate neural networks inside of programmable shaders | DLSS 4: Multi Frame Generation ensures ultra-smooth frame pacing for lifelike simulations.
  • [Double-Flow-Through Design] The RTX PRO 6000 Blackwell features a double-flow-through cooling design, optimizing efficiency and airflow to sustain peak performance under 600W power loads. | [5th Gen Tensor Cores] Deliver up to 3X the performance of the previous generation and support for FP4 precision for faster AI model processing times with reduced memory usage, enabling local fine-tuning of LLMs and generative AI | [4th Gen Ray Tracing Cores] Double the ray-triangle intersection rate of the previous generation to create photoreal, physically accurate scenes and immersive 3D designs with RTX Mega Geometry, which enables up to 100X more ray-traced triangles.
  • [PCIe Gen 5] Support for PCIe Gen 5 provides double the bandwidth of PCIe Gen 4, improving data-transfer speeds from CPU memory and unlocking faster performance for data-intensive tasks like AI, data science, and 3D modeling. | [GDDR7 Memory] With 96 GB of GPU memory and 1.8 TB ps bandwidth, it can tackle massive 3D and AI projects, fine-tune AI models locally, explore large-scale VR environments, and drive larger multi-app workflows.
  • [DisplayPort 2.1] Achieve unparalleled visual clarity and performance, driving high resolution displays at up to 8K at 240 Hz and 16K at 60 Hz. Increased bandwidth enables seamless multi-monitor setups while HDR and higher color depth support ensures superior color accuracy for precision work, such as video editing, 3D design, and live broadcasting.
  • [Universal MIG] Divide a single RTX PRO 6000 Blackwell into multiple isolated instances, each with dedicated resources, allowing for concurrent execution of multiple workloads, optimized GPU utilization, and secure isolation of different applications or users. [WARRANTY] 3 YR Manufacturer's Warranty. Bulk OEM Packaging. Retail Packaging is NOT included.
Reported benchmark regime Author-reported result How to interpret it
Production-style broadcast-query jagged shapes, BF16 on B200 About 13% average forward advantage and about 50% average backward advantage over the May 2026 FA4 implementation The averages cover the jagged shapes tested. TLX trails FA4 in forward on the longest sequences at high density.
Equal-length LLM-style dense shapes, BF16 on B200 About 87% of FA4 forward performance and about 17% better backward performance These are the article’s separate dense-workload results, not the jagged production-case comparison.

The authors say they swept sparsity from highly variable sequence lengths toward nearly uniform lengths and used Nsight Compute hardware counters, ptxas spill information, and TritonBench profiler-measure runs for optimization and latency comparisons. These are the authors’ reported methods and measurements; they are not an independent reproduction.

FA4’s own paper provides separate Blackwell context: its authors report up to 1.3× speedup over cuDNN 9.13 and 2.7× over Triton on B200 BF16, reaching 1,613 TFLOPs/s (about 71% utilization) under that paper’s benchmark settings. Those results do not reproduce or validate the JFA-versus-FA4 comparison.

Why the kernel uses TLX

The PyTorch authors say their earlier Triton baseline left pipeline depth, on-chip data movement, and much of scheduling to the compiler. TLX gives the implementation more direct control over hardware-aware memory operations, asynchronous work, barriers, and warp specialization. The aim is to overlap data loading, softmax work, and matrix operations rather than leave those stages waiting on one another.

JFA uses a persistent, warp-specialized structure with explicit shared-memory and tensor-memory allocation, asynchronous Tensor Memory Accelerator (TMA) transfers and matrix multiply-accumulate (MMA) operations, and barriers to coordinate pipeline stages. It also uses Cluster Launch Control (CLC) and scheduling changes to balance work across streaming multiprocessors (SMs). That balance matters because jagged sequences produce uneven tile counts: some SMs can finish early while others still have substantial work.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The optimization work described by the authors includes balancing jagged tiles across SMs, staging dQ work, releasing tensor memory earlier, and peeling loops. The kernel also adopts FA4’s two-CTA collaborative MMA design for a constrained backward path. FA4’s paper discusses why these kinds of choices matter on Blackwell, where tensor-core throughput must be considered alongside other bottlenecks such as shared-memory traffic and exponential operations.

Rank #2
NVIDIA RTX PRO 4000 Blackwell Graphics Card - 24GB GDDR7 ECC Memory, PCIe 5.0 x16, 4X DisplayPort 2.1b, Single Slot Full Height AI Workstation GPU, Retail Packaging
  • Professional GPU with Blackwell Architecture
  • Blackwell Architecture
  • 24GB GDDR7 with PCIe 5.0 & Ray Tracing
  • AI Workstation

Why broadcast queries make backward attention different

In the production HSP case described in the article, one dense query is broadcast across the jagged sequences. The forward pass can use that query for each sequence, but its backward gradient must combine contributions from across the batch. That cross-program reduction is a scheduling and coordination challenge, not just a matter of running the same per-sequence calculation faster.

For the production broadcast-query case with head dimension 128, the authors report that their two-CTA backward path improved throughput by about 12%—equivalent to about 11% lower latency—compared with their single-CTA path. This is a scoped ablation of the two backward paths, not the headline comparison against FA4.

What the public package supports

The facebookresearch/ads_model_kernel_library repository’s tlx_jfa README documents the package capabilities and the conditions for its specialized backward path:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Capability or constraint Package documentation
Attention forms Jagged self-attention and cross-attention; PMA (broadcast-query) attention
Windows and groups Symmetric sliding windows; grouped-query attention in forward
Backward support Autograd backward; the general one-CTA backward handles cases outside the specialized two-CTA path
Two-CTA backward conditions Broadcast-query PMA, head dimension 128, one query group, no sliding window, and load balancing enabled
Hardware and head dimension Blackwell SM100 or newer; head dimension no greater than 128
Documented setup path Python 3.12 is listed as an available setup path; the README specifies fbtriton==3.6.1 and gives a PyTorch CUDA 12.8 wheel example

These are repository documentation details, not a guarantee that a particular local software environment will install or run successfully. The package requirements also mean the B200 benchmark results should not be taken as evidence of compatibility or performance on earlier GPU generations.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

MXFP8 and block-sparse experiments

MXFP8

The article describes an MXFP8 variant that uses E4M3 values, E8M0 block scales, and TLX block-scaled MMA while reusing the kernel skeleton. The authors report forward performance above FA4’s FP8 kernel and backward parity with FA4 at dense. These are their claims for the described experiments; the article does not establish that result for other shapes or settings.

Rank #3
PNY VCNRTXPRO2000B-PB NVIDIA RTX PRO 2000 Blackwell 16GB GDDR7 128B Graphics Cards
  • Form Factor: Plug-in Card
  • Cooler Type: Active Cooler
  • Maximum Power Consumption: 70W
  • Length: 6.6
  • Height: 2.7

Block-sparse attention

The described block-sparse design first runs a scoring kernel that pools query and key blocks and selects top-k key/value blocks. The attention kernel then iterates over the selected blocks. The article says this experimental variant supports broadcast queries, grouped-query attention, and windowing. For the tested sequence lengths at a 0.5 selection ratio, the authors report forward speedups of roughly 1.3–1.5× over dense attention; that result is not a general guarantee for sparse attention.

Implementation size and the meaning of “SOTA”

The PyTorch authors report that their TLX kernel is about 3.2K lines, compared with roughly 10K lines for FA4 CuteDSL kernels. This is their approximate code-size comparison and supports an argument about implementation size and maintainability; line count alone does not measure engineering effort, readability, or ease of extending either kernel.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The article frames the work as a road toward state-of-the-art FA4 performance on Blackwell. Its reported jagged backward average is ahead of the stated comparator on the tested shapes, while jagged forward trails in the longest, high-density cases and the separate dense forward comparison reaches about 87% of FA4 performance. “SOTA” therefore should not be read as a claim that this kernel wins across every shape, dtype, attention variant, or Blackwell system.

How to evaluate the result for your workload

  • Match the sequence layout: packed jagged inputs and broadcast queries differ from equal-length batches and per-sequence queries.
  • Compare forward and backward separately; the reported relative results differ substantially between passes.
  • Keep dtype, GPU model, and comparator version attached to every performance figure.
  • Check supported masks, sliding-window behavior, grouped-query configuration, and the two-CTA backward restrictions against your intended use.
  • Treat the published measurements as evidence for the authors’ tested benchmark mix, not as a substitute for measuring your own shapes and software stack.

For GEM, the authors write: “Attention is the single slowest kernel in GEM.” That statement describes the priority behind this particular optimization effort; it does not establish that attention is the bottleneck in other systems.

Quick Recap

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.