GPU vs CPU backtesting: when graphics cards help, and when they don't
Testing one strategy is a sequential job. Testing millions of different strategies is not, and that difference is why GPUs can change how much of a search space you are able to explore.
AlphaPrime team ·
Two different kinds of work
Inside a single backtest, each bar depends on the bar before it: the position, the open orders and the equity all carry forward. That part is sequential. You cannot compute bar 5,000 before bar 4,999.
A strategy search is different. When you generate and test many candidate strategies, each candidate's backtest is independent of the others. Candidate 1 does not need anything from candidate 2. Work like this, many independent tasks of the same shape, is what parallel hardware is built for.
What a CPU does well
A desktop CPU has a small number of fast, flexible cores, typically 8 to 32. Each core handles branching logic, irregular memory access and complex control flow well. Running a strategy search on a CPU usually means giving each core its own candidates. Doubling the cores roughly doubles throughput, until memory bandwidth becomes the limit.
What a GPU does differently
A modern NVIDIA GPU has thousands of simpler cores grouped into streaming multiprocessors (SMs). It is fastest when many threads run the same instructions on different data at the same time. For strategy search that means:
- Many candidates are evaluated together in a batch, each on its own thread or group of threads.
- Price data is loaded into GPU memory once and read by every candidate.
- Results such as scores and statistics are reduced on the GPU, so only the summary has to travel back to the CPU.
Done well, this keeps every SM busy. Measured on a single RTX 3090, AlphaPrime averaged 0.052 ms per strategy evaluation and completed 1.268 billion strategy evaluations in 18 hours 23 minutes.
Where the limits are
GPUs are not automatically faster. A GPU implementation has to deal with:
- Branch divergence. Threads that take different paths through the code wait for each other. Strategies with very different logic in the same batch can slow a batch down.
- Memory. Long histories, many symbols and per-trade records compete for GPU memory. Batch sizes have to fit.
- Transfer costs. Moving data between CPU and GPU memory is slow compared with computing on it. A design that copies data back and forth for every candidate loses most of the advantage.
- Small jobs. If you only need to test a few hundred strategies, setup costs dominate and a CPU may be just as quick.
That is why the engine has to be designed for the GPU from the start. Moving an existing CPU backtester onto a graphics card rarely gives the same benefit.
How to read a speed claim
A number like "millions of strategies per hour" means little on its own. Before comparing, check:
- Hardware: which GPU or CPU, and how many.
- Data: how many bars per candidate, and what bar size.
- What counts as a strategy: fully backtested and scored, or only partly evaluated and filtered early.
- Setup: the strategy complexity and the search settings.
- Average versus peak: total strategies divided by total time, or the best moment.
One measured example
For reference, AlphaPrime's published benchmark on an NVIDIA RTX 3090 completed 1,268,200,000 strategy evaluations in 18 hours 23 minutes of continuous running, an average of 0.052 ms per strategy evaluation. It was measured on 2026-10-05 with a real research setup: 37-minute data from 2015-01-01 to 2026-09-25, 114,371 bars per candidate, TensorMOEA/D multi-objective search, counting the evaluations completed across generation, full backtesting, scoring and selection. Actual speed depends on the graphics card, data length and research setup.

Why speed matters for research quality
Faster search is not only about waiting less. It changes what you can afford to check: more walk-forward windows, more robustness tests per candidate, and more careful handling of multiple testing. Speed is most useful when it is spent on validation, not only on generating more candidates.
What AlphaPrime does here
AlphaPrime is built around GPU batches for strategy generation, backtesting and scoring. GPU mode supports NVIDIA GeForce GTX 10 series and newer on Windows 10/11 64-bit; the CPU and GPU editions have the same features, and the GPU edition computes on an NVIDIA graphics card for more speed. Walk-forward analysis runs on NVIDIA graphics cards.

