Hi everyone,
I’m posting this to share an update on the PyTorch modification used by AddaxAI to run on older AMD GPUs, like my 5600XT.
I’ve put the full step-by-step guide in this repository:
so I wouldn’t clutter the forum with a massive block of text.
A few points:
I did all of this with the help of Opus 5.5 High, guiding it through certain processes using my basic knowledge of Linux.
I don’t have any Linux expertise; I simply looked into “The Rock’s” repository to see if it was possible to run the program with an AMD GPU.
The tutorial in the repo was written with AI assistance, since, as I mentioned, I’m no Linux expert.
The program runs perfectly on my CachyOS setup; it’s even installed as a native application, just like opening a program on Windows.
This was tested only on my machine (specs below). Anyone with newer GPUs that support ROCm should use the official AMD ROCm stack. Newer GPUs should actually be much faster than mine.
Experiment
Unofficial and experimental. TheRock wheels are nightly builds, and this was tested on one machine only.
| GPU | AMD Radeon RX 5600 XT (Navi 10, gfx1010, 6 GB) |
| CPU, RAM | AMD Ryzen 7 2700 (8 cores, 16 threads), 32 GB |
| OS | CachyOS (Arch-based), kernel 7.2, Mesa 26.2 |
| AddaxAI | 7.10.0, Linux .deb |
| PyTorch | torch 2.14.0 + torchvision 0.29.0a0, TheRock build 10.2.0a20260930 |
Results on my machine
Detection (MD5A), classification (SpeciesNet) and embeddings (DINOv2) all ran on the GPU.
| Run | Whole job | Details |
|---|---|---|
| 16 images | 34 s | MegaDetector ~5.7 images/s once warmed up (9 s to load the model, 4 s for the first image) |
| 363 videos (1080p, 15 fps, ~10 s each, H.264) | 16 min | detection 10.5 min (1.7 s per video), SpeciesNet 5 min, embeddings 20 s; 2,661 detections |
GPU load peaked at 99% and VRAM use at +1.9 GB. The kernel log showed no amdgpu errors. During the video run the GPU was busy only about half the time (see Limitations).
Limitations
- TensorFlow models stay on the CPU: NEO-MNCN-v1-0 (Neotropics), PAM-SDZWA-v1, SOCAL-IRC-v3-6, TAS-BB-v1 and TERRAI-NEP-v1. There is no TensorFlow ROCm build for these GPUs.
- CAM-AI4G-v1 (Colombian Amazon) uses a separate environment with torch 2.4 (
env-pywildlife), which this guide does not touch, so it also runs on the CPU. - During video analysis the GPU is idle about half the time, because the steps run one after another. MegaDetector first decodes every frame of a clip on the CPU (about 0.4 s for a 10-second 1080p clip here). Then each sampled frame (one per second of video by default) is prepared on the CPU and run on the GPU, about 0.13 s per frame. Most of the time goes to the sampled frames, so a lower video sampling rate in AddaxAI should be faster (not measured).
- A warning about
libMIOpenCKGroupedConv_gfx1010.somissing is harmless. - The tested build (
10.2.0a20260930) leaves AMD’s index in early November 2026; the index keeps about six weeks. After that,downloadneeds a newer build that still includes torch 2.14.0: pick a date at https://nightly.repo.amd.com/rocm/whl-next/torch/, put it in theBUILD=line ofaddaxai-rocm, then rundownload,applyandtest. Newer builds are untested.
One thing worth mentioning is that, based on my tests on Wednesday, each video took about 5.9 seconds to analyze. Opus spotted an error in my implementation and reduced that time to 1.7 seconds.
I hope this helps people with AMD GPUs.