Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To run VMAF through FFmpeg’s CUDA path on Windows, you run a Linux container through Docker Desktop’s WSL 2 backend, pass your NVIDIA GPU into that container, and use an FFmpeg build that includes both Netflix’s libvmaf library and the libvmaf_cuda filter. That filter accepts only CUDA frames, so both the distorted and the reference video must reach it as CUDA frames.

What you need before you start

  • A supported NVIDIA GPU. Docker’s GPU support page for Windows is the checklist to follow for Docker Desktop GPU passthrough.
  • Windows updated to a current build, with an NVIDIA Windows driver that supports GPU paravirtualization for WSL 2.
  • A current WSL 2 Linux kernel, installed or updated with wsl --update.
  • Docker Desktop with the WSL 2 backend enabled.
  • Microsoft documents CUDA support for Windows 11 and for Windows 10 version 21H2. Minimum driver and kernel versions change over time, so confirm them on the current NVIDIA and Microsoft pages before you rely on a specific number.
  • Your reference and distorted videos stored on a local drive.

Step 1: Confirm the GPU is visible inside a container

Do this before touching FFmpeg. Most failed attempts stop at this stage, and debugging FFmpeg will not fix a container that cannot see the GPU.

  1. Open PowerShell and run wsl --update.
  2. Install the current NVIDIA Windows driver with WSL support.
  3. Open Docker Desktop and go to Settings > General. Confirm that Use the WSL 2 based engine is checked.
  4. Go to Settings > Resources > WSL integration and enable integration for the Ubuntu distribution you plan to use.
  5. Run a CUDA base image with nvidia-smi:
docker run --rm --gpus all <a CUDA base image> nvidia-smi

A table listing your GPU and driver version means the container can see the card. NVIDIA’s WSL documentation notes that nvidia-smi has a reduced feature set under WSL 2, so missing metrics do not by themselves indicate a fault. If the command fails, fix the Docker, driver, or WSL configuration first.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 2: Build an FFmpeg image with libvmaf and CUDA

Netflix’s VMAF Docker documentation describes using the NVIDIA Container Toolkit and a Dockerfile.ffmpeg build for FFmpeg with CUDA support and the VMAF filter. Start from that upstream file instead of writing your own build from scratch.

#1 Best Overall
Sale
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
  • AI Performance: 767 AI TOPS
  • OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode)
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Axial-tech fan design features a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • A 2.5-slot design maximizes compatibility and cooling efficiency for superior performance in small chassis

Configure flags

FFmpeg’s filter documentation lists --enable-nonfree --enable-ffnvcodec --enable-libvmaf as configure flags to use after libvmaf is installed. These flags are necessary but not sufficient. The build also depends on a compatible CUDA toolchain and FFmpeg version, so follow the current upstream Dockerfile rather than assembling a build from the flags alone.

Why the VMAF-CUDA build comes from source

NVIDIA’s technical blog states that “VMAF-CUDA must be built from the source.” Plan for a source build rather than expecting a prebuilt package to include the CUDA filter.

Check that the filter exists

Inside the container, run the following. The output should include libvmaf_cuda.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5070 Ti
  • Integrated with 16GB GDDR7 256bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system
ffmpeg -hide_banner -filters | grep vmaf

If the filter is missing, the build did not include libvmaf or the CUDA support, and you should rebuild before continuing.

Step 3: Start the container with your video files

In WSL, the Windows C: drive appears as /mnt/c. The example below mounts a folder named videos on C: into the container at /data. The image tag my-ffmpeg-vmaf is a name chosen for the build from Step 2; use whatever tag you gave yours.

docker run --rm -it --gpus all 
  -e NVIDIA_DRIVER_CAPABILITIES=compute,video 
  -v /mnt/c/videos:/data -w /data 
  my-ffmpeg-vmaf bash

The --gpus all flag exposes the GPU, and NVIDIA_DRIVER_CAPABILITIES=compute,video covers the compute and video decoding paths that the CUDA pipeline uses. Netflix’s example uses these same patterns.

Rank #3
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4. System Requirements: Minimum 850W PSU with 16-pin 12V-2x6 (12VHPWR) connector required. Verify before purchasing.
  • Military-grade components deliver rock-solid power and longer lifespan for ultimate durability. Compatibility: 348mm (13.7") length, 3.6 slots, 4.3 lbs. Confirm case clearance and slot spacing. GPU bracket included.
  • Protective PCB coating helps protect against short circuits caused by moisture, dust, or debris
  • 3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans
  • Phase-change GPU thermal pad helps ensure optimal thermal performance and longevity, outlasting traditional thermal paste for graphics cards under heavy loads

Step 4: Run the CUDA filter graph

The following command adapts FFmpeg’s documented CUDA example. Run it inside the container, from the /data directory. It has not been verified on every Windows, WSL, driver, and GPU combination, so treat it as a working starting point and confirm the output on your own system.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
ffmpeg 
  -hwaccel cuda -hwaccel_output_format cuda -i distorted.mp4 
  -hwaccel cuda -hwaccel_output_format cuda -i reference.mp4 
  -filter_complex "[0:v]scale_cuda=format=yuv420p[dist];[1:v]scale_cuda=format=yuv420p[ref];[dist][ref]libvmaf_cuda=log_fmt=json:log_path=output.json" 
  -f null -

What each part does

  • -hwaccel cuda -hwaccel_output_format cuda is applied to each input so that decoded frames stay on the GPU.
  • scale_cuda=format=yuv420p converts the pixel format on the GPU. It is applied to both inputs so the filter receives matching formats.
  • libvmaf_cuda takes the distorted video as its first input and the reference as its second. Reversing them changes which file is treated as the reference.
  • log_fmt=json and log_path=output.json write the scores to a JSON file in /data, which is the same folder as your videos on the Windows side.
  • -f null - discards the decoded output, so the score log is the result you keep.

Pixel format and alignment

Netflix’s example says that 4:2:0 video decoded as NV12 needs conversion to 4:2:0 with scale_cuda. It says other formats, such as yuv444p or yuv422p, may be passed from the decoder without that conversion. Do not assume either case. Read the actual format of each file with ffprobe -v error -select_streams v:0 -show_entries stream=pix_fmt,width,height,r_frame_rate -of default=noprint_wrappers=1 distorted.mp4 and repeat for the reference, then decide whether the conversion is needed.

No single preprocessing recipe covers every pair of files. Check that both videos have the same dimensions, frame rate, and timing before you read the scores. If they differ, align them first, because a mismatched pair produces scores that do not describe the same frames.

Rank #4
Sale
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5060
  • Integrated with 8GB GDDR7 128bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

Reading the JSON log and comparing results

The log contains the per-frame scores and the pooled VMAF metrics for the run. Use the pooled figures for a summary and the per-frame entries when you need to locate the frames that drop in quality.

A community discussion on Reddit asked whether CPU and GPU VMAF scores are the same. The sources available for this guide do not establish that they are identical across all VMAF versions, models, pixel formats, or inputs. If you need to compare a CUDA score with a CPU score, run both on the same aligned files with the same VMAF model and settings, then compare the results yourself. Do not treat a single pair of runs as proof of equivalence for other content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance expectations

NVIDIA’s technical blog, published in 2024, reports “up to 37x lower per-frame latency at 4K” and “up to 4.4x higher throughput in FFmpeg” for VMAF-CUDA, compared with a dual Intel Xeon 8480 CPU system. These are vendor-reported figures for the configurations NVIDIA tested. They are not independent measurements and not a guaranteed speedup. Your result depends on resolution, pixel format, the decoder path, the GPU, and where your files are stored. Files on the Windows drive accessed through /mnt/c may be slower to read than files stored inside the Linux filesystem, so keep working copies in the Linux filesystem if throughput matters.

Best Value
ASUS TUF Gaming GeForce RTX 5070 12GB GDDR7 OC EditionGaming Graphics Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4 OC mode: 2640MHz/Default mode: 2610MHz (Boost Clock)
  • Military-grade components deliver rock-solid power and longer lifespan for ultimate durability
  • Protective PCB coating helps protect against short circuits caused by moisture, dust, or debris
  • 3.125-slot design with massive fin array optimized for airflow from three Axial-tech fans
  • Phase-change GPU thermal pad helps ensure optimal thermal performance and longevity, outlasting traditional thermal paste for graphics cards under heavy loads
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting

Symptom Likely cause What to check
The container cannot see the GPU, or nvidia-smi fails in Step 1 Docker Desktop is not using the WSL 2 engine, the driver lacks WSL support, or WSL is out of date Repeat the Step 1 settings, run wsl --update, and reinstall the current NVIDIA driver
nvidia-smi shows fewer metrics than on native Linux Reduced feature set under WSL 2, as stated in NVIDIA’s WSL documentation Confirm that the GPU and driver version are listed; this does not indicate a failure
libvmaf_cuda is not listed by the filter check The FFmpeg build lacks libvmaf or the CUDA filter support Run ffmpeg -hide_banner -filters | grep vmaf and rebuild from the upstream Dockerfile
The filter graph rejects the frames An input is not decoded to CUDA frames Confirm that both inputs include -hwaccel cuda -hwaccel_output_format cuda
Pixel format or dimension mismatch The two files differ in format, size, or frame rate Read both files with ffprobe, apply scale_cuda with the same format to both, and align dimensions and timing
No speed gain over the CPU Decoding or file reads are the bottleneck, or the input format forces extra work Move working copies into the Linux filesystem and check the decoder output format

Choosing a Docker setup on Windows

This guide uses the Docker Desktop route because Docker documents GPU passthrough for Linux containers on Windows through its WSL 2 backend. The table compares it with running Docker Engine inside a WSL distribution.

Route GPU passthrough support in the documentation used here Practical note
Docker Desktop with the WSL 2 backend Documented for --gpus Linux-container passthrough on Windows Follow Step 1 exactly, including the WSL integration setting
Docker Engine inside a WSL distribution Not covered by the documentation used for this guide; not stated Verify GPU access separately before building FFmpeg on this route

Neither route is established here as faster or easier for every user.

Quick Recap

SaleBestseller No. 1
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
ASUS Dual GeForce RTX 5060 Ti 16GB GDDR7 OC Edition Gaming Graphics Card
AI Performance: 767 AI TOPS; OC mode: 2632 MHz (OC mode)/ 2602 MHz (Default mode); Powered by the NVIDIA Blackwell architecture and DLSS 4
$786.37
Bestseller No. 2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5070 Ti; Integrated with 16GB GDDR7 256bit memory interface
$1,162.49
Bestseller No. 3
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
ASUS TUF Gaming GeForce RTX™ 5080 16GB GDDR7 OC Edition Graphics Card
3.6-slot design with massive fin array optimized for airflow from three Axial-tech fans; Auto-Extreme precision automated manufacturing helps ensure higher reliability
$1,831.31
SaleBestseller No. 4
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5060; Integrated with 8GB GDDR7 128bit memory interface
$459.99
Bestseller No. 5
ASUS TUF Gaming GeForce RTX 5070 12GB GDDR7 OC EditionGaming Graphics Card
ASUS TUF Gaming GeForce RTX 5070 12GB GDDR7 OC EditionGaming Graphics Card
3.125-slot design with massive fin array optimized for airflow from three Axial-tech fans; Auto-Extreme precision automated manufacturing helps ensure higher reliability
$937.39

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.