NVLM 1.0 is NVIDIA’s 2024 family of multimodal large language models, but the publicly downloadable release is narrower: the NVLM-1.0-D-72B decoder-only checkpoint. It accepts text and images and returns text, uses Qwen2-72B-Instruct with an InternViT-6B vision encoder, and is released for non-commercial use under CC BY-NC 4.0. The documented unquantized path is aimed at multi-GPU NVIDIA systems rather than ordinary laptops.
What NVLM 1.0 is—and what it is not
NVIDIA describes NVLM 1.0 as a family of “frontier-class” multimodal LLMs designed to understand images and text while retaining or improving text-only abilities. The research compares three integration strategies: NVLM-D, NVLM-X and NVLM-H.
A multimodal LLM should not be confused with a general media-generation model. NVLM-D reads image and text inputs and generates text. The released checkpoint does not create images, video or audio. Its model card lists text and image input, text output and a maximum token length of 128K tokens.
The paper was submitted on September 17, 2024 and revised on October 22, 2024. NVIDIA’s release-era results should therefore be read as 2024 comparisons, not as a current 2026 leaderboard.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
- Powered by Radeon RX 9070 XT
- WINDFORCE Cooling System
- Hawk Fan
- Server-grade Thermal Conductive Gel
- RGB Lighting
Read the primary research in the NVLM 1.0 paper and the implementation details in NVIDIA’s NVLM-D-72B model card.
What NVIDIA actually released
| Research architecture | Public status |
|---|---|
| NVLM-D | Public NVLM-1.0-D-72B checkpoint, weights and inference materials. |
| NVLM-X | Discussed and evaluated in the paper; not presented in the public model card as an equivalent downloadable checkpoint. |
| NVLM-H | Proposed and evaluated hybrid design; availability should not be assumed from the NVLM-D release. |
In other words, NVLM 1.0 is the research family name, while nvidia/NVLM-D-72B is the public model most readers can download. The model card points to Megatron-Core-related training code, but that is not the same as a turnkey reproduction of NVIDIA’s complete original training pipeline.
How the three architectures differ
NVLM-D: decoder-only multimodal fusion
NVLM-D incorporates visual features into the language model’s main processing path. The design is intended to let one decoder reason jointly over visual and textual information, with particular relevance to OCR, documents and visual reasoning.
NVLM-X: cross-attention
NVLM-X gives language layers access to image representations through cross-attention. NVIDIA presents this approach as potentially more efficient for some high-resolution workloads because visual information need not be inserted throughout the full language sequence.
Rank #2
- PLEASE NOTE: Exporting an NVIDIA RTX Pro 6000 GPU outside the US requires strict adherence to the U.S. Export Administration Regulations (EAR) and issuance of an export license from the Bureau of Industry and Security (BIS). Compliance and Know Your Customer (KYC) screening may be required as a condition of order acceptance. [NVIDIA Blackwell Streaming Multiprocessor] The new SM features increased processing throughput, and new neural shaders that integrate neural networks inside of programmable shaders | DLSS 4: Multi Frame Generation ensures ultra-smooth frame pacing for lifelike simulations.
- [Double-Flow-Through Design] The RTX PRO 6000 Blackwell features a double-flow-through cooling design, optimizing efficiency and airflow to sustain peak performance under 600W power loads. | [5th Gen Tensor Cores] Deliver up to 3X the performance of the previous generation and support for FP4 precision for faster AI model processing times with reduced memory usage, enabling local fine-tuning of LLMs and generative AI | [4th Gen Ray Tracing Cores] Double the ray-triangle intersection rate of the previous generation to create photoreal, physically accurate scenes and immersive 3D designs with RTX Mega Geometry, which enables up to 100X more ray-traced triangles.
- [PCIe Gen 5] Support for PCIe Gen 5 provides double the bandwidth of PCIe Gen 4, improving data-transfer speeds from CPU memory and unlocking faster performance for data-intensive tasks like AI, data science, and 3D modeling. | [GDDR7 Memory] With 96 GB of GPU memory and 1.8 TB ps bandwidth, it can tackle massive 3D and AI projects, fine-tune AI models locally, explore large-scale VR environments, and drive larger multi-app workflows.
- [DisplayPort 2.1] Achieve unparalleled visual clarity and performance, driving high resolution displays at up to 8K at 240 Hz and 16K at 60 Hz. Increased bandwidth enables seamless multi-monitor setups while HDR and higher color depth support ensures superior color accuracy for precision work, such as video editing, 3D design, and live broadcasting.
- [Universal MIG] Divide a single RTX PRO 6000 Blackwell into multiple isolated instances, each with dedicated resources, allowing for concurrent execution of multiple workloads, optimized GPU utilization, and secure isolation of different applications or users. [WARRANTY] 3 YR Manufacturer's Warranty. Bulk OEM Packaging. Retail Packaging is NOT included.
NVLM-H: hybrid integration
NVLM-H combines decoder-only and cross-attention mechanisms. Its purpose is to balance the direct multimodal reasoning and OCR benefits associated with NVLM-D against the high-resolution efficiency advantages of cross-attention. It is not universally superior: results depend on image resolution, sequence length, serving software and hardware.
Dynamic image tiling and 1-D tile tags
Large images contain details that disappear when the entire frame is reduced to one small representation. NVLM can divide a high-resolution image into tiles, allowing the vision encoder to inspect local text, chart labels, document regions and other fine details.
- Benefit: More access to small text and localized visual information.
- Cost: Additional tiles create more visual tokens, memory traffic and latency.
- Limit: More tiles do not automatically solve global reasoning, reading order or spatial integration.
NVIDIA’s reported contribution is a 1-D tile-tagging scheme that adds textual or positional structure to the tiled inputs. The paper attributes gains in OCR and multimodal reasoning to this organization. In production, tiling is a quality-versus-efficiency control: the best setting depends on image content and the serving budget.
Backbone, vision encoder and software
| Component | Public NVLM-D-72B detail |
|---|---|
| Language backbone | Qwen2-72B-Instruct |
| Vision encoder | InternViT-6B |
| Architecture | Decoder-only transformer |
| Runtime | PyTorch |
| Listed hardware | NVIDIA Hopper; NVIDIA reports testing on H100 |
| Operating system | Linux |
| Context limit listed by model card | 128K tokens |
The Hugging Face adaptation modifies tokenizer and model code for vision-related special tokens and multi-GPU inference. Its documented loading path uses trust_remote_code=True, so operators should review and pin that code in an isolated environment.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
- Chipset: NVIDIA GeForce GT 1030
- Video Memory: 4GB DDR4
- Boost Clock: 1430 MHz
- Memory Interface: 64-bit
- Output: DisplayPort x 1 (v1.4a) / HDMI 2.0b x 1
Training approach
Multimodal pretraining
NVIDIA describes a curated mixture of image captions, image-text pairs, natural images, charts, documents, scene descriptions, OCR-oriented examples and mathematical reasoning data.
Supervised fine-tuning
The SFT mixture covers visual instruction following, document and chart understanding, diagrams, general knowledge, mathematics and text-only tasks.
The reported lesson is that data quality and task diversity can matter more than raw dataset size. High-quality text-only and multimodal mathematics data were deliberately retained to avoid the text-performance degradation seen in some multimodal models. This is NVIDIA’s finding for its recipe, not a universal rule: outcomes vary with the base model, optimization method, data mixture and evaluation design.
What the benchmark numbers show
The following are NVIDIA-reported scores for the Hugging Face adaptation, not independent validation.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Rank #4
- Robust 4GB Memory & Quad Display Ready: Equipped with 4GB of fast GDDR5 memory to smoothly handle daily graphics tasks. Features four built-in HDMI ports, enabling a seamless quad-monitor setup directly out of the box—perfect for multi-tasking offices, digital signage, or trading desks.
- Plug-and-Play Installation & Wide Compatibility: Utilizes a standard PCI Express interface for broad compatibility with most desktop PCs. Offers straightforward plug-and-play installation and stable driver support for modern Windows and Linux operating systems, ensuring a hassle-free setup.
- Quiet, Cool & Compact Design: Engineered with a silent fan and efficient cooling system for near-silent operation, making it ideal for noise-sensitive environments. Its low-profile design fits easily into small form factor cases, with both half-height and full-height brackets included for flexible installation.
- Enhanced Multimedia & Everyday Performance: Delivers smooth 1080P video playback and supports hardware-accelerated decoding, offering an excellent experience for home theater PCs (HTPC). Provides capable performance for everyday applications, multimedia tasks.
- Complete Package & Reliable Support: Includes the graphics card, both low-profile and standard brackets, a quick start guide, and screwdriver, which make it simple and quick setup process.
| Multimodal benchmark | NVLM-D 72B score |
|---|---|
| MMMU validation / test | 58.7 / 54.9 |
| MathVista | 65.2 |
| OCRBench | 852 |
| AI2D | 94.2 |
| ChartQA | 86.0 |
| DocVQA | 92.6 |
| TextVQA | 82.6 |
| RealWorldQA | 69.5 |
| VQAv2 | 85.4 |
| Text-only benchmark | Hugging Face adaptation |
|---|---|
| MMLU | 81.7 |
| GSM8K | 93.2 |
| MATH | 73.1 |
| HumanEval | 89.0 |
| Reported average accuracy | 84.3 |
The model card reports a 4.5-point average improvement over its listed Qwen2-72B-Instruct comparison for the Hugging Face implementation; the Megatron implementation is listed with a 4.3-point improvement. NVIDIA also reports Megatron MMMU validation/test of 59.7 / 54.6 and OCRBench of 853. The model card attributes these small differences to the Megatron and Hugging Face codebases.
Prompts, preprocessing, decoding, hardware, competitor versions and evaluation dates can change results. “State of the art” should therefore be attributed to the release-era paper or model card, not treated as a current 2026 claim.
Running NVLM-D-72B
Transformers pipeline
The lowest-friction documented route is:
pip install transformers torch
from transformers import pipeline
pipe = pipeline(
"image-text-to-text",
model="nvidia/NVLM-D-72B",
trust_remote_code=True,
)
messages = [{
"role": "user",
"content": [
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
{"type": "text", "text": "What animal is on the candy?"}
]
}]
print(pipe(text=messages))
Direct model loading
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained(
"nvidia/NVLM-D-72B",
torch_dtype=torch.bfloat16,
low_cpu_mem_usage=True,
use_flash_attn=False,
trust_remote_code=True,
).eval()
A 72B model in bfloat16 has substantial memory requirements. The model card provides a device-map example that reserves part of GPU 0 for the vision encoder and spreads 80 language-model layers across available GPUs. Treat multi-GPU execution as the practical expectation for the documented unquantized path; a typical consumer GPU should not be assumed to run it.
vLLM and SGLang
pip install vllm
vllm serve "nvidia/NVLM-D-72B"
The model card shows OpenAI-compatible requests at http://localhost:8000/v1/chat/completions, including an image_url object.
Best Value
- System Compatibility Note: 2.5‑slot card measuring 303 mm (L) x 131 mm (W) x 45 mm (H); requires a single 8‑pin power connector and a recommended 550W power supply. Please verify chassis clearance and power supply capacity before purchase.
- Dedicated Support: Please contact us directly through Amazon for any product questions or assistance you may require.
- AMD RDNA 3 Architecture with AI & Ray Tracing Acceleration: Powered by 32 RDNA 3 Compute Units featuring 3rd Gen Ray Tracing Accelerators and 2nd Gen AI Accelerators, delivering lifelike lighting, shadows, and superior machine learning performance for enhanced gaming and content creation.
- Powerful 1080p & 1440p Gaming Engine: Features a max boost clock of up to 2695 MHz, a game clock of 2280 MHz, and 2048 stream processors, ensuring outstanding frame rates in the latest titles.
- 8GB High‑Speed GDDR6 Memory: Equipped with 8GB of GDDR6 memory on a 128‑bit interface running at 18 Gbps, delivering up to 288 GB/s bandwidth for high‑resolution textures and demanding game workloads.
pip install sglang
python3 -m sglang.launch_server
--model-path "nvidia/NVLM-D-72B"
--host 0.0.0.0
--port 30000
SGLang exposes the documented endpoint at http://localhost:30000/v1/chat/completions. Support can vary with framework, CUDA and model revisions.
Version and security precautions
- Use an isolated environment and review remote model code before enabling
trust_remote_code=True. - Pin model revisions, Python, PyTorch, Transformers, CUDA and serving-engine versions.
- For close reproduction, use the supplied environment based on
nvcr.io/nvidia/pytorch:23.09-py3; NVIDIA warns that version changes can alter results. - Do not infer equivalent support on AMD, Apple Silicon, CPU-only systems or every NVIDIA generation from the Hopper/H100 listing.
License and production reality
The public checkpoint is marked for non-commercial use under CC BY-NC 4.0 in the model card, alongside applicable base-model terms. Public weights do not automatically mean commercially permissive open source.
“Production-grade multimodality” in NVIDIA’s framing means strong image-and-text performance while preserving text capabilities. It does not establish low latency, low cost, safety, privacy compliance, support SLAs, broad hardware compatibility or reliable performance on private company documents.
OCR accuracy also is not document reliability. Models can misread tables, lose reading order, confuse labels and values, invent missing words, fail on low contrast or dense pages, and make arithmetic errors from charts. Consequential workflows need confidence checks, page-level citations, structured extraction validation and human review.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Who should consider NVLM-D?
| Need | Fit |
|---|---|
| Academic multimodal research | Strong candidate for studying large decoder-only vision-language systems. |
| OCR, charts and document prototypes | Potentially strong, subject to validation on representative data. |
| NVIDIA H100-based internal service | Plausible evaluation target with multi-GPU capacity. |
| Commercial SaaS | License review is mandatory; poor default without permission. |
| Laptop or single consumer GPU | Poor fit without substantial optimization or quantization. |
| Image generation | Not suitable; the released model outputs text only. |
| Managed API with SLA | A hosted, commercially licensed alternative is usually simpler. |
Questions to answer before deployment
- Is the use internal research, a prototype or commercial?
- Can the team satisfy CC BY-NC 4.0 and the base-model terms?
- How many GPUs and how much usable memory remain after serving overhead?
- Is bfloat16 practical, or is quantization required?
- Does the selected serving engine support the model’s custom code and image format?
- Will images be stored, logged or sent to third-party infrastructure?
- How will uncertain OCR and hallucinated document facts be detected?
- Would a smaller or newer multimodal model meet the requirement at lower cost?
Verdict
NVLM 1.0 is an important research release because it compares decoder-only, cross-attention and hybrid multimodal designs and reports that a carefully mixed multimodal recipe need not damage text performance. NVLM-D-72B makes that work inspectable through public weights and serving examples. Its practical limits are equally clear: 72B-class hardware, custom software, non-commercial licensing, text-only output and release-era benchmark evidence. It is best approached as a serious research and internal evaluation target—not as a turnkey, commercially cleared replacement for a managed multimodal service.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

