Back to projects
IRGPU 4 weeks 4 pers.

Real-Time Video Motion Detection — GPU Porting

CPU → GPU porting of a real-time motion detection algorithm, accelerated by ×24 using CUDA and 6 Nsight-guided optimizations.

C++ CUDA GStreamer Nsight Systems Nsight Compute
Aperçu du projet IRGPU

Context & Problem Statement

Real-time motion detection is a fundamental building block of video surveillance systems, automated video stream analysis, and robotics. The objective of this project was to design a background subtraction filter capable of operating at 30 FPS or higher in High Definition, constrained to rely strictly on NVIDIA CUDA and the GStreamer multimedia framework.

Technical Challenges


Pipeline Overview

Processing is executed frame by frame across 5 sequential stages, where each pixel is processed independently — an architecture perfectly suited for GPU parallelization:

01 Source Image Raw RGB
02 Estimated Background K=3 Reservoirs Model
03 Motion Mask L₁ Norm
04 Morpho Opening Erosion + Dilation
05 Hysteresis Threshold 4-connected Propagation
06 Final Result Motion Coloring

The Motion Detection Algorithm

Step 1 — Background Estimation

The most complex stage of the pipeline. For each pixel, a set of K = 3 color reservoirs (color + weight) is maintained to model possible background colors (noise, lighting shifts, etc.).

For each new frame and each pixel:

The estimated background for each pixel is simply the color of the reservoir with the highest weight.

Step 2 — Motion Mask Calculation

The difference between the current pixel and the estimated background is measured via the L₁ norm (mean absolute difference across R, G, B). A high score indicates a pixel in motion.

Step 3 — Morphological Opening

The raw mask contains noise (leaves, video compression artifacts). A morphological opening on a disk of radius R=3 eliminates false positives:

Step 4 — Hysteresis Thresholding

Ensures spatial consistency: pixels with a high score (> 45) are strong seeds. Low score pixels (> 20) are retained only if they are adjacent to a strong pixel (propagation via 4-connectivity), until global convergence.

Step 5 — Motion Coloring

Validated pixels are colored in transparent red on top of the original image, visualizing motion while preserving the underlying video content.


From Sequential C++ to GPU: A Profiling-Driven Approach

Optimization was not performed blindly — every architectural decision was justified by precise profiling metrics using NVIDIA Nsight Systems (global temporal analysis) and NVIDIA Nsight Compute (fine-grained GPU kernel analysis).

Baseline: C++ Implementation (CPU)

The reference sequential implementation processes pixels one by one in two nested loops. It serves as the ground truth to validate the accuracy of each GPU version (SSIM = 1.0000).

Implementation Time (s) Throughput (FPS) Speedup
C++ Reference 616.78 s 5.29 FPS 1.00×

5.29 FPS — live video processing is impossible.

Naive CUDA Port (×9.24)

The initial CUDA version consists of a direct transposition: each pixel is mapped to a GPU thread (2D grid, 16×16 blocks). Without any additional optimizations, moving to GPU immediately crosses the 30 FPS milestone.

Implementation Time (s) Throughput (FPS) Speedup SSIM
C++ Reference 616.78 s 5.29 FPS 1.00× 1.0000
Naive CUDA 52.72 s 48.92 FPS 9.24× 0.9951

Major memory bottlenecks remained, revealed by profiling.


The 6 Major Optimizations

Opti 1 & 2 — Lazy Memory + Switch to float (×18)

Opti 1: The initial version reallocated GPU buffers for every frame. With Lazy Initialization, memory is allocated once at startup, leaving only 2 PCIe transfers per frame (incoming frame → GPU, result → CPU).

Opti 2: Nsight Compute emitted an explicit warning regarding double precision overhead on consumer GPUs. Switching to float with lroundf() significantly accelerated floating-point operations.

Nsight Compute: FP64 () overhead impact warning on consumer GPU
Nsight Compute: FP64 (double) overhead impact warning on consumer GPU
Implementation Throughput (FPS) Speedup
Naive CUDA 48.92 FPS 9.24×
CUDA Lazy Mem + Float 95.43 FPS 18.03×

Opti 3 — Replacing cuRAND with a Lightweight LCG (×18.7)

Nsight Systems reveals that cuRAND allocates a 48-byte structure per pixel in VRAM for its internal state: on HD 1080p video, this consumes ~95 MB just for random generation!

Nsight Systems: VRAM overhead caused by
Nsight Systems: VRAM overhead caused by curandState
Resolution curandState Allocation Size
320×240 ~3.5 MB
1920×1080 ~94.9 MB
cuRAND Allocation (320×240) cuRAND Allocation (1080p)
cuRAND 320x240 cuRAND 1080p

Solution: A Linear Congruential Generator (LCG) computed on the fly from the pixel index and frame number — zero additional VRAM bytes required.

Throughput cuRAND Throughput fast_rand
Throughput cuRAND Throughput fast_rand

Opti 4 — Hysteresis Thresholding in Shared Memory (×23.5)

Hysteresis propagation requires multiple passes until convergence. Without optimization, each iteration triggers CPU/GPU synchronizations and VRAM bandwidth saturation.

VRAM memory access saturation without Shared Memory
VRAM memory access saturation without Shared Memory

Solution: 16×16 tiling in Shared Memory with a +1 pixel halo. Propagation from strong pixels to neighboring weak pixels occurs locally, without global VRAM accesses.

VRAM Analysis Before VRAM Analysis After
Before Shared Memory After Shared Memory

Result: VRAM requests were reduced 8-fold (from 139 K to 17.7 K per iteration), adding +26% pure execution speedup on this kernel alone.

Opti 5 — Tiled Morphological Opening (×24.1)

Erosion and dilation read 29 neighboring pixels per thread directly from global VRAM. Tiling into Shared Memory (halo 2×R) with disk offsets stored in Constant Memory cut VRAM traffic in half (10.69 → 5.44 GB/s).

Opti 6 — 32×8 Block Geometry (×24.5)

Nsight Compute revealed that a 32×8 block geometry (256 threads) aligns perfectly with the size of a CUDA warp (32 threads) in the horizontal direction, maximizing memory coalescence during image row accesses.

Nsight Compute: memory coalescence and warp alignment (32×8 blocks)
Nsight Compute: memory coalescence and warp alignment (32×8 blocks)

CUDA Parallel Design Patterns Analysis

Within the GPU architecture of the project, we evaluated key parallel design patterns:


Results & Benchmarks

Summary Table

Implementation Time (s) Throughput (FPS) Speedup SSIM Accuracy
C++ Reference 616.78 s 5.29 FPS 1.00× 1.0000
Naive CUDA 52.72 s 48.92 FPS 9.24× 0.9951
CUDA Lazy Mem + Float 18.18 s 95.43 FPS 18.03× 0.9950
CUDA Fast Random 17.01 s 98.96 FPS 18.70× 0.9950
CUDA Shared Mem Hysteresis 10.80 s 124.34 FPS 23.49× 0.9949
CUDA Tiling Opening 10.65 s 127.57 FPS 24.10× 0.9949
CUDA Final (32×8 Blocks) 10.56 s 129.51 FPS 24.47× 0.9949

Global Performance Comparison

FPS throughput comparison and overall speedup across dataset
FPS throughput comparison and overall speedup across dataset

From 5.29 FPS to 129.51 FPS: a ×24.47 speedup with near-perfect visual accuracy (SSIM = 0.9949), proving that every optimization stage — from memory layout to random number generation — contributed directly to the final result.


Real-Time Live Video Motion Detection Demo