Dense Prediction Transformers (DPTs) can turn an image into a pixel-level semantic map: each pixel is assigned a class such as road, sky, person, or building. For a practical starting point, use Hugging Face Transformers with DPTForSemanticSegmentation and the Intel/dpt-large-ade checkpoint. That checkpoint predicts a fixed set of ADE20K scene classes; it does not identify arbitrary objects or separate two objects that share the same class.
DPT is a broader architecture for dense image predictions, including both semantic segmentation and monocular depth estimation. This guide focuses on segmentation, explains the model’s output, and shows how to run and inspect an inference result.
What image segmentation predicts
Image classification assigns a label to an entire image, while object detection identifies objects with boxes. Segmentation instead assigns predictions to image locations, producing a spatial map aligned with the scene.
- Semantic segmentation assigns a class to each pixel. Pixels belonging to different cars may all be labeled “car,” without identifying which car is which.
- Instance segmentation assigns both a class and an individual object identity, so two cars receive separate masks.
- Panoptic segmentation combines semantic labels for scene regions with instance masks for countable objects.
The DPT checkpoint used below is a semantic-segmentation model. It is not an instance-segmentation system or a general-purpose object cutout tool.
Recommended Free Tools
#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
What makes DPT a dense-prediction model
A dense-prediction model produces an output for many or all spatial positions in an image. Semantic segmentation produces a discrete class per pixel; monocular depth estimation produces a continuous depth-like value per pixel. Surface normals, optical flow, and saliency are other examples of dense tasks. DPT names the broader architecture, not segmentation alone; the Hugging Face DPT documentation describes both segmentation and depth estimation.
Convolutional neural networks (CNNs) build representations from local neighborhoods and expand their receptive fields through successive layers. Vision transformers represent image content as patches or tokens and use self-attention to mix information across distant parts of the image. This global feature interaction can help when a region’s meaning depends on the surrounding scene, but it does not make transformers universally better than CNNs. DPT can require substantial memory and computation, and results depend on pretraining, resolution, decoder design, and the similarity between the model’s training data and the target images.
How DPT turns an image into a segmentation map
- Preprocess the image. The checkpoint’s image processor resizes, normalizes, and converts the image into tensors in the form expected by the model.
- Embed image patches. The image is represented as visual tokens corresponding to spatial patches or transformed visual features.
- Encode global and staged features. Transformer self-attention mixes information between tokens. Intermediate encoder stages retain representations at different levels.
- Reassemble spatial features. DPT converts token sequences back into image-like feature maps and recovers multiple feature resolutions.
- Fuse and decode. A convolutional decoder progressively fuses and upsamples those representations; the segmentation head produces class scores, or logits, for each output location.
- Resize and select classes. Logits are resized to the desired image dimensions, then the highest-scoring class at each pixel is selected.
The original Vision Transformers for Dense Prediction paper describes combining multi-stage transformer features with a convolutional decoder to produce dense predictions. It reported 49.02% mIoU on ADE20K for semantic segmentation under the paper’s experimental setup; that is a historical result, not a current benchmark guarantee for every DPT checkpoint. Read the paper.
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
Choose segmentation or depth estimation deliberately
DPT has task-specific heads and checkpoints. The Hugging Face API distinguishes DPTForSemanticSegmentation from DPTForDepthEstimation; one model’s output should not be interpreted as the other task’s result.
| Task | Typical output | How to interpret it |
|---|---|---|
| Semantic segmentation | Class scores, commonly shaped like (batch, classes, height, width) |
Take argmax across classes at each pixel to get integer class IDs. The available labels depend on the checkpoint. |
| Monocular depth estimation | One continuous depth-like value per pixel | Represents estimated scene geometry according to that checkpoint; it does not identify semantic classes. |
A depth visualization is not a segmentation mask, and a segmentation class map is not a depth estimate. DPT does not necessarily produce both at once: task heads and checkpoints are separate.
Run pretrained semantic segmentation in Python
Use a supported Python and PyTorch environment with the current Hugging Face Transformers package installed. Pin package versions and, where reproducibility matters, the checkpoint revision in your own environment. The exact dependency versions can affect preprocessing and outputs; the example below uses the documented AutoImageProcessor and DPTForSemanticSegmentation interfaces.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
import numpy as np
import torch
import torch.nn.functional as F
from PIL import Image
from transformers import AutoImageProcessor, DPTForSemanticSegmentation
checkpoint = "Intel/dpt-large-ade"
image = Image.open("input.jpg").convert("RGB")
processor = AutoImageProcessor.from_pretrained(checkpoint)
model = DPTForSemanticSegmentation.from_pretrained(checkpoint)
model.eval()
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model.to(device)
inputs = {
key: value.to(device)
for key, value in processor(images=image, return_tensors="pt").items()
}
with torch.no_grad():
outputs = model(**inputs)
logits = outputs.logits
logits = F.interpolate(
logits,
size=(image.height, image.width),
mode="bilinear",
align_corners=False,
)
segmentation = logits.argmax(dim=1)[0].cpu().numpy()
print(segmentation.shape) # (image height, image width)
print(np.unique(segmentation)) # class IDs present in this image
The output of argmax is a two-dimensional integer array, not an RGB image. The logits may have different spatial dimensions from the model input, so resize the continuous logits before selecting classes. Bilinear interpolation is appropriate for logits; for an already discrete class-ID mask, use nearest-neighbor interpolation if resizing is unavoidable.
Create a quick visual check
A generated palette is useful for checking whether the model separates regions, but its colors are arbitrary and do not name the classes.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallnum_classes = logits.shape[1]
rng = np.random.default_rng(42)
palette = rng.integers(
0, 256, size=(num_classes, 3), dtype=np.uint8
)
mask_rgb = palette[segmentation]
mask_image = Image.fromarray(mask_rgb)
mask_image.save("segmentation-mask.png")
overlay = Image.blend(
image.convert("RGBA"),
mask_image.convert("RGBA"),
alpha=0.5,
)
overlay.save("segmentation-overlay.png")
For a meaningful ADE20K visualization, use the checkpoint’s label mapping and palette. Do not infer class names from color alone: an arbitrary palette is only a visual aid. Consult the DPT model documentation and Hugging Face semantic-segmentation guide for documented model and task APIs.
Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Manage device and memory limits
The example selects CUDA when available and otherwise uses the CPU; it makes no runtime promise because speed depends on hardware, image dimensions, batch size, precision, and software versions. For a first run on constrained hardware, process one image at a time and consider a smaller or hybrid model. Reducing resolution can reduce memory use but can also remove small structures. Tiling very large images may help fit memory, but tile boundaries can introduce seams and tiles lose some global context.
Interpret and evaluate the results
The Intel/dpt-large-ade checkpoint is oriented to ADE20K semantic classes. It can predict only labels represented by that checkpoint; it is not open-vocabulary and cannot be expected to segment user-defined categories without suitable training or a different model. A model trained on general scene imagery may also transfer poorly to medical scans, satellite imagery, microscopy, industrial inspection, or other specialized domains.
For labeled evaluation, mean Intersection over Union (mIoU) summarizes class overlap:
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
IoU_c = TP_c / (TP_c + FP_c + FN_c)
mIoU = (1 / C) × Σ IoU_c, where C is the number of evaluated classes. IoU measures overlap between the predicted region and ground truth for a class. Because mIoU averages class-level scores, it can expose weak results on rare classes that overall pixel accuracy might obscure. Comparisons are meaningful only when dataset split, label mapping, preprocessing, resolution, and evaluation protocol match.
Also inspect per-class IoU, pixel accuracy, frequency-weighted IoU, and boundary metrics such as boundary F-score or boundary IoU. For deployment, measure latency, peak memory, and throughput on the intended hardware. A single aggregate score or attractive overlay cannot show every failure that matters in an application.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Common failure modes and what to check
- Confused classes: visually similar categories such as road and sidewalk, wall and building, or floor and carpet may be mixed. Inspect per-class results and the label mapping, not only the overlay.
- Small or thin structures disappear: wires, poles, signs, distant pedestrians, and fine boundaries can be lost in patch representations or decoder upsampling. Higher input resolution may help at the cost of memory and latency.
- Jagged edges, holes, or isolated regions: resizing and model predictions can create boundary artifacts. If you apply connected-component filtering, morphological operations, or a conditional random field, validate the change against ground truth; post-processing can remove legitimate detail.
- Unexpected colors or labels: check whether the output is class IDs, whether IDs use the checkpoint’s mapping, and whether the palette matches that mapping. RGB colors themselves carry no inherent class meaning.
- Poor performance in a new domain: night, fog, infrared, fisheye, aerial, medical, or factory images may differ substantially from ordinary scene imagery. Evaluate on representative labeled data and consider fine-tuning rather than assuming reliable transfer.
- Logits do not match image size: resize logits to the target dimensions before
argmax. Enlarging a low-resolution class-ID map instead can produce blocky or distorted boundaries. - Download, compatibility, or reproducibility problems: verify the checkpoint name and network access, then record Python, PyTorch, Transformers, processor configuration, checkpoint revision, device, precision, and resizing method. These choices can change results.
When to choose DPT—and when not to
DPT is worth considering when the goal is dense semantic scene understanding, global context is useful, the available checkpoint’s labels and domain are close to the task, and the deployment budget can accommodate its compute needs. It is a poor fit if the required labels are absent, if the goal is instance identities or text-prompted arbitrary masks, or if the deployment requires real-time inference on low-power hardware without a suitable optimized configuration.
| Alternative | Consider it when | Important distinction |
|---|---|---|
| CNN-based systems such as U-Net- or DeepLab-style models | You need mature tooling, a domain-specific model, or a configuration suited to constrained deployment. | Performance and compute depend on the encoder, decoder, data, and training setup; global context may require architectural choices. |
| SegFormer | You want a transformer-based semantic-segmentation family with an efficiency-oriented lightweight decoder. | It is a different architecture, not simply another name for DPT. |
| Mask2Former | Mask-level prediction is central, especially for semantic, instance, or panoptic segmentation. | Its task capabilities differ from a fixed-label semantic DPT checkpoint. |
| Segment Anything-family models | You need promptable or interactive masks. | Promptable masks solve a different problem from assigning a fixed semantic class to every pixel. |
| Open-vocabulary segmentation models | Categories need to be specified through text or vary beyond a fixed training label set. | Prompt sensitivity and domain transfer introduce different evaluation and consistency concerns. |
Original DPT code and current implementation choice
The original Intel repository is useful for studying the research implementation and legacy reproduction. It includes separate scripts, run_monodepth.py and run_segmentation.py; segmentation outputs go to output_semseg, and the segmentation script supports -t dpt_hybrid and -t dpt_large. The repository is archived and states that Intel no longer maintains it, including bug fixes, releases, or updates. Its documented Python 3.7, PyTorch 1.8.0, OpenCV 4.5.1, and timm 0.4.5 environment describes historical code, not recommended current installation requirements. See the original DPT repository.
For a new Python inference workflow, the Hugging Face Transformers implementation is the more practical starting point: it exposes a semantic-segmentation class, pretrained loading, image preprocessing, and output handling. The original paper remains useful for understanding the architecture, while the deployed checkpoint and your validation data determine whether it is suitable for a particular application.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

