What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can run OFA-YOLO object detection on a Zynq UltraScale+ MPSoC by preparing an image for the model, quantizing it for the DPU, running inference, and decoding the returned tensors into image-space boxes. Aleksei Rostov’s January 2025 Hackster project describes this workflow with Vitis AI 3.0’s OFA-YOLO model and a DPUCZDX8G. Treat its performance figures as results from that project—not as general guarantees or a confirmed compatibility recipe for every Zynq board.

What the OFA-YOLO deployment does

AMD/Xilinx included OFA-YOLO for object detection among the models described in the Vitis AI 3.0 release notes. Rostov’s project implements inference on a Zynq UltraScale+ system using the DPUCZDX8G, with Python code that names the vart and xir libraries. The project assumes a Linux environment in which the Vitis AI 3.0 libraries are already configured; it is not a from-scratch guide to installing the toolchain or building a DPU design.

The described model input is a 640×640×3 tensor. It produces three output tensors with grids of 80×80, 40×40 and 20×20, each with 255 channels. The example model configuration specifies 80 classes. These dimensions describe the project’s configuration, not every possible OFA-YOLO model or deployment.

How to run inference, step by step

The essential path is image preprocessing, DPU execution, output decoding and postprocessing. Use the tensor metadata and configuration belonging to the model you actually deploy; do not substitute thresholds or quantization values from a different artifact.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
  1. Prepare the image. Resize the source image to the model’s 640×640 input and apply the scaling and normalization expected by that model.
  2. Quantize the input. Read the input tensor’s fixed-point metadata and use it to convert the prepared data to the INT8 representation expected by the DPU. The project describes quantization as part of the preprocessing path.
  3. Run the DPU. Submit the input tensor to the DPU runner and retrieve the model’s output tensors. The example uses Vitis AI’s vart and xir libraries.
  4. Dequantize the outputs. Apply each output tensor’s own fixed-point metadata before treating its values as model scores or box predictions.
  5. Decode detections at all three scales. Interpret the output grids using the model’s YOLO-style grid and anchor configuration to recover box coordinates, objectness and class scores. Use the matching model configuration rather than assuming anchor values or decoding details.
  6. Filter and suppress candidates. Apply confidence filtering, then non-maximum suppression (NMS) to remove overlapping detections. The project’s model configuration and its non-optimized sample code use different threshold values, so there is no single threshold pair to copy as universal. Select the values specified for your model and application.
  7. Restore source-image coordinates. Map the surviving boxes from model-input coordinates back to the original frame, accounting for the resize and any other preprocessing transform. Render labels only if the application needs a visual display.

For an offline-image application, the camera is unnecessary: image files can supply the model input. A live-camera application also needs a capture path. The project names a Logitech C270 for its camera demonstration.

Hardware and software prerequisites to verify

The author’s page contains an important bill-of-materials mismatch. Its “Things used” list names a Trenz Electronic TE0821-02-2AE91PA module and TE0703 carrier board. The narrative describing tests instead names a TE0820-03-2AI21FA module and a TE0703-06 carrier board. Confirm the exact tested module and carrier combination with the project author before purchasing hardware; the discrepancy prevents treating either pair as a definitive reproduction list.

Rank #2
RCTCBRZVTW FPGA Development Board Zynq UltraScale+ MPSoC XCZU2CG AI(AXU2CGA Video Package)
  • Stability: Long-term stable use
  • Maintenance: Easy to maintain
  • Easy to install: Simple operation
  • Application: Wide range of applications
  • Correct use: correct use can extend the product life

Also verify the complete deployment combination: board and carrier revision, DPU design and configuration, Vitis AI version, model artifact, and matching compiler and runtime components. The cited Vitis AI 3.0 material establishes historical inclusion of OFA-YOLO in that release context, not present-day support for a particular Trenz board or compatibility with a newer toolchain. AMD/Xilinx’s Vitis AI repository describes the stack as supporting inference on Xilinx hardware, but the sources cited for this project do not establish that this exact board-and-artifact combination works with the latest releases.

The project page describes a free, non-optimized Python sample and says optimized implementation resources are available by contacting the author or making a donation. That paid optimized code is the author’s offering, not an official AMD/Xilinx distribution; the page’s description is not an independent verification of its performance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
AMD ZU15EG Development Board Zynq UltraScale+ ARM FPGA Platform with 4GB DDR4 PS 2GB DDR4 PL FMC HPC SFP HDMI SATA MIPI AI Video Processing Educational Kit (MIPI Package)
  • ARM plus FPGA Hybrid Architecture:Powered by AMD Xilinx Zynq UltraScale Plus XCZU15EG with ARM Cortex-A53 and FPGA logic, delivering powerful heterogeneous computing performance for embedded development.
  • Large-Capacity DDR4 Memory:Equipped with 4GB DDR4 for ARM (PS) and 2GB DDR4 for FPGA (PL), ideal for high-speed data processing, real-time signal processing, and AI acceleration workloads.
  • Rich High-Speed Interfaces:Includes FMC HPC, SFP, SATA, MIPI CSI, Mini DisplayPort, and 4K HDMI input and output. Perfect for image processing, video capture, and ultra-high bandwidth applications.
  • Ideal for AI and Video Applications:Widely used in artificial intelligence, 4K video systems, edge computing, and deep learning inference. Supports DisplayPort interface for high-resolution display integration.
  • Full Development Resources Included:Comes with schematics, Verilog HDL demos, and hands-on experiment guidelines. Supports fast prototyping for research, education, and product development.

What the reported accuracy and speed comparisons show

Rostov reports evaluating a full model and models at 30% and 50% sparsity, with COCO metrics calculated using pycocotools. In the reported comparison, the full model achieved higher average precision (AP) and average recall (AR) across object sizes, while pruning improved throughput at an accuracy cost that was especially apparent for small and medium objects. The retrieved project text does not provide the underlying AP/AR values or a full results table, so readers cannot use it to quantify the trade-off or independently compare the models.

The author also reports that a multithreaded C++ implementation took about 20 milliseconds less per model than the multithreaded Python implementation. The measured interval is specifically the time to upload data to the DPU runner and retrieve it—not a complete application’s end-to-end frame time. The author separately characterizes the non-optimized Python example as about ten times slower than the multithreaded implementation, but the retrieved text does not include the detailed timing table or enough benchmark setup to reproduce that ratio. Neither comparison should be treated as a general performance guarantee for other boards, software builds or workloads.

Rank #4
AMD ZU15EG Development Board Zynq UltraScale+ ARM FPGA Platform with 4GB DDR4 PS 2GB DDR4 PL FMC HPC SFP HDMI SATA MIPI AI Video Processing Educational Kit (ADDA Package)
  • ARM plus FPGA Hybrid Architecture:Powered by AMD Xilinx Zynq UltraScale Plus XCZU15EG with ARM Cortex-A53 and FPGA logic, delivering powerful heterogeneous computing performance for embedded development.
  • Large-Capacity DDR4 Memory:Equipped with 4GB DDR4 for ARM (PS) and 2GB DDR4 for FPGA (PL), ideal for high-speed data processing, real-time signal processing, and AI acceleration workloads.
  • Rich High-Speed Interfaces:Includes FMC HPC, SFP, SATA, MIPI CSI, Mini DisplayPort, and 4K HDMI input and output. Perfect for image processing, video capture, and ultra-high bandwidth applications.
  • Ideal for AI and Video Applications:Widely used in artificial intelligence, 4K video systems, edge computing, and deep learning inference. Supports DisplayPort interface for high-resolution display integration.
  • Full Development Resources Included:Comes with schematics, Verilog HDL demos, and hands-on experiment guidelines. Supports fast prototyping for research, education, and product development.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to choose a configuration for your application

  • Prioritize detection quality: Start with the full model as the project’s reported comparison found it more accurate than the tested pruned versions. Validate with your own images and class distribution, especially if small or medium objects matter.
  • Consider pruning when throughput is limiting: The project reports a throughput benefit for its pruned models, alongside lower AP/AR. Measure the accuracy change on the target task before selecting a sparsity level; the published text does not supply the numerical results needed to predict that change.
  • Separate DPU time from application speed: Measure image capture, preprocessing, runner submission and retrieval, decoding, NMS and display or output separately. The project’s approximately 20 ms C++/Python difference applies only to its stated upload-and-retrieval interval.
  • Benchmark the implementation you will ship: Threading, language choice, model artifact, tensor handling and postprocessing all affect the full pipeline. The project’s comparisons do not establish that C++ or a particular pruning level will be fastest or most accurate in a different deployment.

Sources and evidence limits

The implementation details and author-reported results above come from Aleksei Rostov’s Hackster project, published January 24, 2025. Historical Vitis AI 3.0 model-zoo context comes from AMD/Xilinx release notes, and the broader description of the inference stack comes from the official AMD/Xilinx Vitis AI repository. The project text available for these claims does not expose complete AP/AR results, a full timing table, or a consistent tested bill of materials. It therefore supports an explanation of the pipeline and a qualified account of the author’s comparisons, but not an independent benchmark or a definitive current hardware compatibility claim.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.