Tutorial - How to run AddaxAI on AMD GPUs

Hi everyone,

I’m posting this to share an update on the PyTorch modification used by AddaxAI to run on older AMD GPUs, like my 5600XT.

I’ve put the full step-by-step guide in this repository:

so I wouldn’t clutter the forum with a massive block of text.

A few points:

I did all of this with the help of Opus 5.5 High, guiding it through certain processes using my basic knowledge of Linux.

I don’t have any Linux expertise; I simply looked into “The Rock’s” repository to see if it was possible to run the program with an AMD GPU.

The tutorial in the repo was written with AI assistance, since, as I mentioned, I’m no Linux expert.

The program runs perfectly on my CachyOS setup; it’s even installed as a native application, just like opening a program on Windows.

This was tested only on my machine (specs below). Anyone with newer GPUs that support ROCm should use the official AMD ROCm stack. Newer GPUs should actually be much faster than mine.

Experiment

Unofficial and experimental. TheRock wheels are nightly builds, and this was tested on one machine only.

GPU AMD Radeon RX 5600 XT (Navi 10, gfx1010, 6 GB)
CPU, RAM AMD Ryzen 7 2700 (8 cores, 16 threads), 32 GB
OS CachyOS (Arch-based), kernel 7.2, Mesa 26.2
AddaxAI 7.10.0, Linux .deb
PyTorch torch 2.14.0 + torchvision 0.29.0a0, TheRock build 10.2.0a20260930

Results on my machine

Detection (MD5A), classification (SpeciesNet) and embeddings (DINOv2) all ran on the GPU.

Run Whole job Details
16 images 34 s MegaDetector ~5.7 images/s once warmed up (9 s to load the model, 4 s for the first image)
363 videos (1080p, 15 fps, ~10 s each, H.264) 16 min detection 10.5 min (1.7 s per video), SpeciesNet 5 min, embeddings 20 s; 2,661 detections

GPU load peaked at 99% and VRAM use at +1.9 GB. The kernel log showed no amdgpu errors. During the video run the GPU was busy only about half the time (see Limitations).

Limitations

  • TensorFlow models stay on the CPU: NEO-MNCN-v1-0 (Neotropics), PAM-SDZWA-v1, SOCAL-IRC-v3-6, TAS-BB-v1 and TERRAI-NEP-v1. There is no TensorFlow ROCm build for these GPUs.
  • CAM-AI4G-v1 (Colombian Amazon) uses a separate environment with torch 2.4 (env-pywildlife), which this guide does not touch, so it also runs on the CPU.
  • During video analysis the GPU is idle about half the time, because the steps run one after another. MegaDetector first decodes every frame of a clip on the CPU (about 0.4 s for a 10-second 1080p clip here). Then each sampled frame (one per second of video by default) is prepared on the CPU and run on the GPU, about 0.13 s per frame. Most of the time goes to the sampled frames, so a lower video sampling rate in AddaxAI should be faster (not measured).
  • A warning about libMIOpenCKGroupedConv_gfx1010.so missing is harmless.
  • The tested build (10.2.0a20260930) leaves AMD’s index in early November 2026; the index keeps about six weeks. After that, download needs a newer build that still includes torch 2.14.0: pick a date at https://nightly.repo.amd.com/rocm/whl-next/torch/, put it in the BUILD= line of addaxai-rocm, then run download, apply and test. Newer builds are untested.

One thing worth mentioning is that, based on my tests on Wednesday, each video took about 5.9 seconds to analyze. Opus spotted an error in my implementation and reduced that time to 1.7 seconds.

I hope this helps people with AMD GPUs.