GPU-Accelerated Satellite Constellation Simulation

GPU-Accelerated Satellite Constellation Simulation

← Back to projects

Context: Independent Performance Engineering Project · 2026
Focus: CUDA · Parallel Computing · Benchmarking · Correctness Validation

I built a C++17/CUDA simulation for evaluating satellite motion and ground-point coverage on both CPU and GPU.

The project focuses on more than kernel acceleration. It compares serial CPU and CUDA execution under controlled workloads, separates kernel-level speedup from true end-to-end performance, validates CPU/GPU consistency, and measures how keeping simulation state resident on the GPU changes the cost of repeated timesteps.

View source on GitHub →

Project snapshot

MetricResult
ImplementationC++17 · CUDA
Largest single-step workload1,000,000 satellites
CPU time, 1M satellites × 1 step93.0149 ms
GPU kernel time, 1M satellites × 1 step1.9140 ms
GPU end-to-end time, 1M satellites × 1 step11.0439 ms
Single-step kernel speedup48.55×
Single-step end-to-end speedup8.73×
Largest resident workload1,000,000 satellites × 100 steps
Resident kernel speedup49.25×
Resident end-to-end speedup46.84×
Automated tests4
Final benchmark validationAll workloads passed

The challenge

Satellite updates and coverage checks are naturally parallel across independent satellites, but raw kernel speed is only part of GPU performance.

A practical GPU implementation also has to account for:

The project therefore focused on four questions:

Simulation model

The simulator intentionally uses a simplified orbital model so that the project can isolate the computational structure of GPU acceleration.

Each satellite is represented by:

The satellite follows a circular orbit in the Cartesian x-y plane:

x = r cos(theta)
y = r sin(theta)
z = 0

The orbital angle evolves as:

theta(t + dt) = theta(t) + omega * dt

and is wrapped into the interval [0, 2π).

The goal is not high-fidelity astrodynamics. The simplified model provides a deterministic workload for comparing serial CPU execution with parallel CUDA execution.

Ground-point coverage

Coverage is evaluated from the satellite-to-ground line-of-sight vector and the ground point’s radial vector.

For satellite position s and ground position g:

LOS = s - g

The elevation angle is computed from the line-of-sight direction relative to the local upward direction at the ground point.

A satellite is considered visible when its elevation exceeds the configured minimum elevation angle. The benchmark uses a 10° minimum elevation threshold.

This same coverage logic is shared between the CPU and GPU paths so that performance comparisons can be paired with explicit correctness checks.

CPU and GPU execution

The serial CPU implementation processes satellites one after another:

Satellite 0 -> update -> coverage
Satellite 1 -> update -> coverage
Satellite 2 -> update -> coverage
...

The CUDA implementation maps independent satellites to GPU threads:

Thread 0 -> Satellite 0
Thread 1 -> Satellite 1
Thread 2 -> Satellite 2
...

Two CUDA kernels perform the main workload:

update_satellites_kernel
        |
        v
compute_coverage_kernel

This mapping is well suited to CUDA because position updates and coverage evaluations are independent across satellites.

Benchmark design

I benchmarked two execution regimes:

  1. Single-step execution
  2. GPU-resident multi-step execution

The distinction is important because it separates a transfer-heavy use case from a compute-heavy workload where device memory residency can amortize transfer cost.

Single-step benchmark

One single-step GPU execution includes:

Host -> Device satellite transfer
            |
            v
      Position update
            |
            v
     Coverage evaluation
            |
            v
Device -> Host coverage transfer

Each workload size is executed 20 times per benchmark execution, and the median is reported.

The end-to-end measurement includes:

GPU memory allocation is excluded from timing, and correctness validation is performed outside the timed region.

Single-step results

Final values are the median across three independent Release-mode benchmark executions.

SatellitesCPU computeGPU kernelsGPU end-to-endKernel speedupEnd-to-end speedup
1,0000.0826 ms0.0102 ms0.0417 ms8.06×1.98×
10,0000.8548 ms0.0288 ms0.1632 ms30.66×5.24×
100,0009.5530 ms0.2020 ms1.2694 ms47.41×7.55×
1,000,00093.0149 ms1.9140 ms11.0439 ms48.55×8.73×

At the largest single-step workload, the CUDA kernels are almost 49× faster than the serial CPU compute path. The end-to-end gain is lower because transfers and synchronization remain a meaningful part of total runtime.

GPU-resident multi-step execution

A simulation normally advances the same satellite state across many timesteps.

The resident benchmark therefore performs only one initial host-to-device transfer, keeps the constellation on the GPU for 100 timesteps, and transfers the final result back only once:

Host -> Device
      |
      v
+----------------------+
| Position update      |
| Coverage evaluation  | x 100 timesteps
+----------------------+
      |
      v
Device -> Host

Each resident workload is executed 5 times per benchmark execution.

Resident multi-step results

Final values are again the median across three independent benchmark executions.

WorkloadCPU totalGPU kernel loopGPU end-to-endKernel speedupEnd-to-end speedup
100,000 × 100 steps878.12 ms19.72 ms20.78 ms44.52×42.20×
1,000,000 × 100 steps8740.33 ms177.36 ms186.61 ms49.25×46.84×

For the largest resident workload, the simulation performs 100,000,000 satellite-step evaluations.

Keeping the state on the GPU reduces the relative cost of data movement so that end-to-end acceleration approaches the raw kernel acceleration.

This is the main systems result of the project: data residency and transfer strategy matter almost as much as kernel parallelization when optimizing GPU workloads.

Correctness validation

Performance measurements are only useful if the CPU and GPU implementations agree.

The project includes four automated tests:

TestPurpose
satelliteInitial position, updates and angle wrapping
ground_pointGround-point construction and coordinate access
coverageCovered and non-covered elevation-angle cases
cpu_gpu_consistencyCPU/GPU position and coverage agreement

The CPU/GPU consistency test runs both implementations from the same initial constellation and simulation parameters.

Every benchmark workload also validates final CPU and GPU coverage results. The resident benchmark additionally checks sampled final satellite positions after repeated updates.

All workloads passed validation in all three final benchmark executions.

GitHub Actions verifies the CUDA build and CPU-side tests on the hosted runner. Device-dependent CPU/GPU consistency validation is performed locally when CUDA-capable hardware is available.

Benchmark reproducibility

The benchmark uses deterministic satellite initialization based on a golden-angle phase distribution rather than random initialization.

The published measurements use:

3 independent benchmark executions

Single-step:
    20 repetitions per workload
    median reported

Resident multi-step:
    5 repetitions per workload
    100 timesteps
    median reported

The final published values are the median of the three independent execution medians.

This setup reduces sensitivity to individual timing fluctuations while keeping the comparison between CPU and GPU workloads reproducible.

Benchmark environment

The published results were measured on:

ComponentConfiguration
CPUIntel Core i5-8250U @ 1.60 GHz
CPU baselineSingle-threaded
GPUNVIDIA GeForce MX150
Compute capability6.1
GPU memory~1994 MiB
CUDA Toolkit12.4
NVCC12.4.131
C++ compilerG++ 13.4.0
Build typeRelease
CUDA architecturesm_61

The measured speedups are hardware- and system-dependent. They should be interpreted as results for this tested environment rather than universal CUDA performance claims.

Engineering the implementation

The codebase separates reusable simulation logic from executables, tests and benchmarks.

Key components include:

The project uses C++17, CUDA, CMake and CTest, with dedicated benchmark and validation paths rather than mixing measurement logic into the main simulation executable.

Scope and limitations

This project is a CUDA parallel-computing and performance-engineering exercise, not a production astrodynamics package.

The current model intentionally excludes:

The CPU reference is single-threaded, so the reported speedups measure CUDA acceleration relative to that serial baseline.

The benchmark therefore demonstrates acceleration of the implemented computational workload, not the performance of a complete operational satellite-constellation simulator.

What I contributed

What this project demonstrates

This project demonstrates how I approach performance engineering as a combination of parallelization, measurement, correctness and systems-level reasoning.

The main result is not simply that CUDA kernels execute faster than serial CPU code. The project shows how host-device data movement can limit application-level acceleration, and how retaining repeatedly used state on the GPU can convert a transfer-dominated workload into a compute-dominated one.

It also demonstrates experience with low-level performance measurement, reproducible benchmarking, CPU/GPU consistency testing and maintainable C++/CUDA project structure.

Core technologies: C++17 · CUDA · GPU Computing · Parallel Computing · Performance Optimization · Benchmarking · CMake · CTest · GitHub Actions · Linux

View source on GitHub →

← Back to projects