Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

NVLM 1.0 is NVIDIA’s 2024 family of multimodal large language models, but the publicly downloadable release is narrower: the NVLM-1.0-D-72B decoder-only checkpoint. It accepts text and images and returns text, uses Qwen2-72B-Instruct with an InternViT-6B vision encoder, and is released for non-commercial use under CC BY-NC 4.0. The documented unquantized path is aimed at multi-GPU NVIDIA systems rather than ordinary laptops.

What NVLM 1.0 is—and what it is not

NVIDIA describes NVLM 1.0 as a family of “frontier-class” multimodal LLMs designed to understand images and text while retaining or improving text-only abilities. The research compares three integration strategies: NVLM-D, NVLM-X and NVLM-H.

A multimodal LLM should not be confused with a general media-generation model. NVLM-D reads image and text inputs and generates text. The released checkpoint does not create images, video or audio. Its model card lists text and image input, text output and a maximum token length of 128K tokens.

The paper was submitted on September 17, 2024 and revised on October 22, 2024. NVIDIA’s release-era results should therefore be read as 2024 comparisons, not as a current 2026 leaderboard.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
  • Powered by Radeon RX 9070 XT
  • WINDFORCE Cooling System
  • Hawk Fan
  • Server-grade Thermal Conductive Gel
  • RGB Lighting

Read the primary research in the NVLM 1.0 paper and the implementation details in NVIDIA’s NVLM-D-72B model card.

What NVIDIA actually released

Research architecture Public status
NVLM-D Public NVLM-1.0-D-72B checkpoint, weights and inference materials.
NVLM-X Discussed and evaluated in the paper; not presented in the public model card as an equivalent downloadable checkpoint.
NVLM-H Proposed and evaluated hybrid design; availability should not be assumed from the NVLM-D release.

In other words, NVLM 1.0 is the research family name, while nvidia/NVLM-D-72B is the public model most readers can download. The model card points to Megatron-Core-related training code, but that is not the same as a turnkey reproduction of NVIDIA’s complete original training pipeline.

How the three architectures differ

NVLM-D: decoder-only multimodal fusion

NVLM-D incorporates visual features into the language model’s main processing path. The design is intended to let one decoder reason jointly over visual and textual information, with particular relevance to OCR, documents and visual reasoning.

NVLM-X: cross-attention

NVLM-X gives language layers access to image representations through cross-attention. NVIDIA presents this approach as potentially more efficient for some high-resolution workloads because visual information need not be inserted throughout the full language sequence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
NVD RTX PRO 6000 Blackwell Professional Workstation Edition Graphics Card for AI, Design, Simulation, Engineering - 96GB DDR7 ECC Memory - 4th Gen RT/5th Gen Tensor Core GPU - OEM Packaging
  • PLEASE NOTE: Exporting an NVIDIA RTX Pro 6000 GPU outside the US requires strict adherence to the U.S. Export Administration Regulations (EAR) and issuance of an export license from the Bureau of Industry and Security (BIS). Compliance and Know Your Customer (KYC) screening may be required as a condition of order acceptance. [NVIDIA Blackwell Streaming Multiprocessor] The new SM features increased processing throughput, and new neural shaders that integrate neural networks inside of programmable shaders | DLSS 4: Multi Frame Generation ensures ultra-smooth frame pacing for lifelike simulations.
  • [Double-Flow-Through Design] The RTX PRO 6000 Blackwell features a double-flow-through cooling design, optimizing efficiency and airflow to sustain peak performance under 600W power loads. | [5th Gen Tensor Cores] Deliver up to 3X the performance of the previous generation and support for FP4 precision for faster AI model processing times with reduced memory usage, enabling local fine-tuning of LLMs and generative AI | [4th Gen Ray Tracing Cores] Double the ray-triangle intersection rate of the previous generation to create photoreal, physically accurate scenes and immersive 3D designs with RTX Mega Geometry, which enables up to 100X more ray-traced triangles.
  • [PCIe Gen 5] Support for PCIe Gen 5 provides double the bandwidth of PCIe Gen 4, improving data-transfer speeds from CPU memory and unlocking faster performance for data-intensive tasks like AI, data science, and 3D modeling. | [GDDR7 Memory] With 96 GB of GPU memory and 1.8 TB ps bandwidth, it can tackle massive 3D and AI projects, fine-tune AI models locally, explore large-scale VR environments, and drive larger multi-app workflows.
  • [DisplayPort 2.1] Achieve unparalleled visual clarity and performance, driving high resolution displays at up to 8K at 240 Hz and 16K at 60 Hz. Increased bandwidth enables seamless multi-monitor setups while HDR and higher color depth support ensures superior color accuracy for precision work, such as video editing, 3D design, and live broadcasting.
  • [Universal MIG] Divide a single RTX PRO 6000 Blackwell into multiple isolated instances, each with dedicated resources, allowing for concurrent execution of multiple workloads, optimized GPU utilization, and secure isolation of different applications or users. [WARRANTY] 3 YR Manufacturer's Warranty. Bulk OEM Packaging. Retail Packaging is NOT included.

NVLM-H: hybrid integration

NVLM-H combines decoder-only and cross-attention mechanisms. Its purpose is to balance the direct multimodal reasoning and OCR benefits associated with NVLM-D against the high-resolution efficiency advantages of cross-attention. It is not universally superior: results depend on image resolution, sequence length, serving software and hardware.

Dynamic image tiling and 1-D tile tags

Large images contain details that disappear when the entire frame is reduced to one small representation. NVLM can divide a high-resolution image into tiles, allowing the vision encoder to inspect local text, chart labels, document regions and other fine details.

  • Benefit: More access to small text and localized visual information.
  • Cost: Additional tiles create more visual tokens, memory traffic and latency.
  • Limit: More tiles do not automatically solve global reasoning, reading order or spatial integration.

NVIDIA’s reported contribution is a 1-D tile-tagging scheme that adds textual or positional structure to the tiled inputs. The paper attributes gains in OCR and multimodal reasoning to this organization. In production, tiling is a quality-versus-efficiency control: the best setting depends on image content and the serving budget.

Backbone, vision encoder and software

Component Public NVLM-D-72B detail
Language backbone Qwen2-72B-Instruct
Vision encoder InternViT-6B
Architecture Decoder-only transformer
Runtime PyTorch
Listed hardware NVIDIA Hopper; NVIDIA reports testing on H100
Operating system Linux
Context limit listed by model card 128K tokens

The Hugging Face adaptation modifies tokenizer and model code for vision-related special tokens and multi-GPU inference. Its documented loading path uses trust_remote_code=True, so operators should review and pin that code in an isolated environment.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
msi Gaming GeForce GT 1030 4GB DDR4 64-bit HDCP Support DirectX 12 DP/HDMI Single Fan OC Graphics Card (GT 1030 4GD4 LP OC)
  • Chipset: NVIDIA GeForce GT 1030
  • Video Memory: 4GB DDR4
  • Boost Clock: 1430 MHz
  • Memory Interface: 64-bit
  • Output: DisplayPort x 1 (v1.4a) / HDMI 2.0b x 1

Training approach

Multimodal pretraining

NVIDIA describes a curated mixture of image captions, image-text pairs, natural images, charts, documents, scene descriptions, OCR-oriented examples and mathematical reasoning data.

Supervised fine-tuning

The SFT mixture covers visual instruction following, document and chart understanding, diagrams, general knowledge, mathematics and text-only tasks.

The reported lesson is that data quality and task diversity can matter more than raw dataset size. High-quality text-only and multimodal mathematics data were deliberately retained to avoid the text-performance degradation seen in some multimodal models. This is NVIDIA’s finding for its recipe, not a universal rule: outcomes vary with the base model, optimization method, data mixture and evaluation design.

What the benchmark numbers show

The following are NVIDIA-reported scores for the Hugging Face adaptation, not independent validation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
ARDIYES GT 740 4GB GDDR5 Low Profile GPU Graphics Card, 4X HDMI Ports for Quad Multi-Monitor Setup, PCI Express 3.0 x16, Silent Cooling, Ideal for Office and Home Theater
  • Robust 4GB Memory & Quad Display Ready: Equipped with 4GB of fast GDDR5 memory to smoothly handle daily graphics tasks. Features four built-in HDMI ports, enabling a seamless quad-monitor setup directly out of the box—perfect for multi-tasking offices, digital signage, or trading desks.
  • Plug-and-Play Installation & Wide Compatibility: Utilizes a standard PCI Express interface for broad compatibility with most desktop PCs. Offers straightforward plug-and-play installation and stable driver support for modern Windows and Linux operating systems, ensuring a hassle-free setup.
  • Quiet, Cool & Compact Design: Engineered with a silent fan and efficient cooling system for near-silent operation, making it ideal for noise-sensitive environments. Its low-profile design fits easily into small form factor cases, with both half-height and full-height brackets included for flexible installation.
  • Enhanced Multimedia & Everyday Performance: Delivers smooth 1080P video playback and supports hardware-accelerated decoding, offering an excellent experience for home theater PCs (HTPC). Provides capable performance for everyday applications, multimedia tasks.
  • Complete Package & Reliable Support: Includes the graphics card, both low-profile and standard brackets, a quick start guide, and screwdriver, which make it simple and quick setup process.
Multimodal benchmark NVLM-D 72B score
MMMU validation / test 58.7 / 54.9
MathVista 65.2
OCRBench 852
AI2D 94.2
ChartQA 86.0
DocVQA 92.6
TextVQA 82.6
RealWorldQA 69.5
VQAv2 85.4
Text-only benchmark Hugging Face adaptation
MMLU 81.7
GSM8K 93.2
MATH 73.1
HumanEval 89.0
Reported average accuracy 84.3

The model card reports a 4.5-point average improvement over its listed Qwen2-72B-Instruct comparison for the Hugging Face implementation; the Megatron implementation is listed with a 4.3-point improvement. NVIDIA also reports Megatron MMMU validation/test of 59.7 / 54.6 and OCRBench of 853. The model card attributes these small differences to the Megatron and Hugging Face codebases.

Prompts, preprocessing, decoding, hardware, competitor versions and evaluation dates can change results. “State of the art” should therefore be attributed to the release-era paper or model card, not treated as a current 2026 claim.

Running NVLM-D-72B

Transformers pipeline

The lowest-friction documented route is:

pip install transformers torch
from transformers import pipeline

pipe = pipeline(
    "image-text-to-text",
    model="nvidia/NVLM-D-72B",
    trust_remote_code=True,
)

messages = [{
    "role": "user",
    "content": [
        {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
        {"type": "text", "text": "What animal is on the candy?"}
    ]
}]
print(pipe(text=messages))

Direct model loading

import torch
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "nvidia/NVLM-D-72B",
    torch_dtype=torch.bfloat16,
    low_cpu_mem_usage=True,
    use_flash_attn=False,
    trust_remote_code=True,
).eval()

A 72B model in bfloat16 has substantial memory requirements. The model card provides a device-map example that reserves part of GPU 0 for the vision encoder and spreads 80 language-model layers across available GPUs. Treat multi-GPU execution as the practical expectation for the documented unquantized path; a typical consumer GPU should not be assumed to run it.

vLLM and SGLang

pip install vllm
vllm serve "nvidia/NVLM-D-72B"

The model card shows OpenAI-compatible requests at http://localhost:8000/v1/chat/completions, including an image_url object.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
ASRock Radeon RX 7600 Challenger Pro 8GB OC, AMD RDNA 3, 8GB GDDR6, PCIe 4.0, Triple Fans, 0dB Silent, 2695MHz Boost, Triple Fan Graphics Card
  • System Compatibility Note: 2.5‑slot card measuring 303 mm (L) x 131 mm (W) x 45 mm (H); requires a single 8‑pin power connector and a recommended 550W power supply. Please verify chassis clearance and power supply capacity before purchase.
  • Dedicated Support: Please contact us directly through Amazon for any product questions or assistance you may require.
  • AMD RDNA 3 Architecture with AI & Ray Tracing Acceleration: Powered by 32 RDNA 3 Compute Units featuring 3rd Gen Ray Tracing Accelerators and 2nd Gen AI Accelerators, delivering lifelike lighting, shadows, and superior machine learning performance for enhanced gaming and content creation.
  • Powerful 1080p & 1440p Gaming Engine: Features a max boost clock of up to 2695 MHz, a game clock of 2280 MHz, and 2048 stream processors, ensuring outstanding frame rates in the latest titles.
  • 8GB High‑Speed GDDR6 Memory: Equipped with 8GB of GDDR6 memory on a 128‑bit interface running at 18 Gbps, delivering up to 288 GB/s bandwidth for high‑resolution textures and demanding game workloads.
pip install sglang
python3 -m sglang.launch_server 
  --model-path "nvidia/NVLM-D-72B" 
  --host 0.0.0.0 
  --port 30000

SGLang exposes the documented endpoint at http://localhost:30000/v1/chat/completions. Support can vary with framework, CUDA and model revisions.

Version and security precautions

  • Use an isolated environment and review remote model code before enabling trust_remote_code=True.
  • Pin model revisions, Python, PyTorch, Transformers, CUDA and serving-engine versions.
  • For close reproduction, use the supplied environment based on nvcr.io/nvidia/pytorch:23.09-py3; NVIDIA warns that version changes can alter results.
  • Do not infer equivalent support on AMD, Apple Silicon, CPU-only systems or every NVIDIA generation from the Hopper/H100 listing.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

License and production reality

The public checkpoint is marked for non-commercial use under CC BY-NC 4.0 in the model card, alongside applicable base-model terms. Public weights do not automatically mean commercially permissive open source.

“Production-grade multimodality” in NVIDIA’s framing means strong image-and-text performance while preserving text capabilities. It does not establish low latency, low cost, safety, privacy compliance, support SLAs, broad hardware compatibility or reliable performance on private company documents.

OCR accuracy also is not document reliability. Models can misread tables, lose reading order, confuse labels and values, invent missing words, fail on low contrast or dense pages, and make arithmetic errors from charts. Consequential workflows need confidence checks, page-level citations, structured extraction validation and human review.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Who should consider NVLM-D?

Need Fit
Academic multimodal research Strong candidate for studying large decoder-only vision-language systems.
OCR, charts and document prototypes Potentially strong, subject to validation on representative data.
NVIDIA H100-based internal service Plausible evaluation target with multi-GPU capacity.
Commercial SaaS License review is mandatory; poor default without permission.
Laptop or single consumer GPU Poor fit without substantial optimization or quantization.
Image generation Not suitable; the released model outputs text only.
Managed API with SLA A hosted, commercially licensed alternative is usually simpler.

Questions to answer before deployment

  1. Is the use internal research, a prototype or commercial?
  2. Can the team satisfy CC BY-NC 4.0 and the base-model terms?
  3. How many GPUs and how much usable memory remain after serving overhead?
  4. Is bfloat16 practical, or is quantization required?
  5. Does the selected serving engine support the model’s custom code and image format?
  6. Will images be stored, logged or sent to third-party infrastructure?
  7. How will uncertain OCR and hallucinated document facts be detected?
  8. Would a smaller or newer multimodal model meet the requirement at lower cost?

Verdict

NVLM 1.0 is an important research release because it compares decoder-only, cross-attention and hybrid multimodal designs and reports that a carefully mixed multimodal recipe need not damage text performance. NVLM-D-72B makes that work inspectable through public weights and serving examples. Its practical limits are equally clear: 72B-class hardware, custom software, non-commercial licensing, text-only output and release-era benchmark evidence. It is best approached as a serious research and internal evaluation target—not as a turnkey, commercially cleared replacement for a managed multimodal service.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.