Skip to main content
← All articles

PC Optimization · Inside Hone

How to compare two FPS benchmark results

Check whether your benchmark runs match, read average FPS alongside frametimes, and decide when a before-and-after comparison needs repeating.

Illustrated checklist for matching the scene, settings, and measurement before comparing two benchmark runs.
In this article

Before calling a higher FPS number an improvement, check that both runs measured the same task. Match the scene, settings and capture method, with only your intended change between them. If those conditions differ, rerun the test. If they match but the result keeps reversing, leave the comparison inconclusive.

Check whether the runs belong together

Open both reports and compare their setup notes before comparing their scores. Use this checklist to decide whether to keep the pair or collect another run.

CheckWhat should matchIf it does not match
WorkloadGame build, scene, route and cameraRepeat the same scene. Keep different scenes as separate comparisons.
PC and settingsHardware, driver, resolution, graphics, upscaling, FPS cap and sync state, except the change being testedRestore the baseline setup or describe the comparison as testing several changes together.
Capture windowStart and end criteria, plus durationCapture the same section again. Avoid trimming only the worse run.
Tool and metricTool version, timing stream and low-FPS calculationReprocess compatible captures consistently, or capture both again.
Frames includedGame process, dropped-frame policy and generated-frame treatmentMatch the capture options and record what the report includes.
PreparationWarm-up and cache preparation procedureRepeat using the same procedure for both configurations.

This is a practical comparison checklist. It does not prescribe a universal warm-up time or cooldown. Microsoft’s testing methodology (opens in a new tab) illustrates why setup, recording and rerunning belong in a defined procedure; its device-lab timings are not requirements for every gaming PC.

Keep the scene relevant

A built-in benchmark can give you a repeatable workload, but it may miss the part of the game you want to improve. CapFrameX’s scene-selection guide (opens in a new tab) discusses that limitation. Use the same repeatable scene on both sides, then keep your conclusion specific to it. A result from that scene cannot establish performance throughout the game.

Check what the capture counted

Frame generation needs an explicit note. Record whether it was enabled and whether the reported stream includes generated frames. Also record whether frames that were not displayed were excluded. Neither capture policy should silently change between runs.

PresentMon’s documentation (opens in a new tab) distinguishes MsBetweenPresents, which measures intervals between Present calls, from DisplayedTime, which describes how long a frame was displayed. Those are different measurements. A Present-call rate alone does not establish displayed cadence or input latency. PresentMon’s generated-frame classification also requires application or driver instrumentation; do not assume every capture can identify those frames.

Read beyond average FPS

Once the setup matches, read average FPS beside the complete frametime trace: the sequence of measured frame intervals. The average summarizes the run but loses the position of individual slow frames. A long interval marks a pause in that measured stream; it does not, by itself, identify the cause or prove a visible display freeze.

Check the definition behind a low-FPS number before comparing it. A percentile cutoff, an average over the slowest fraction and a time-weighted low summarize different things. CapFrameX explains these metric differences (opens in a new tab). Matching labels such as “1% low” are insufficient without matching calculations, capture boundaries and exclusions.

If the average rises while lows fall, retain both observations. Inspect the traces and repeat before deciding what changed.

Repeat when the conclusion is uncertain

Keep each run’s result instead of saving only the best-looking pair. Compare the spread of the repeated results for each configuration. Some tail measurements can move with a few isolated slow frames, as the CapFrameX metrics guide describes.

If a small difference changes direction across comparable runs, call the result inconclusive. This guide sets no universal run count or percentage threshold. Repeating a mismatched test also leaves the mismatch intact; fix the checklist failures first.

Save enough detail to repeat the comparison

Keep the two configurations, your checklist notes and all individual results together. State what changed and which scene you tested. If you changed a whole profile, describe a profile comparison: it cannot tell you how much each setting contributed.

Bring the same comparison notes with you so the next test answers the same question.

Frequently asked questions

Are two benchmark runs enough?

One before-and-after pair gives you a comparison, but does not show how much repeated runs vary. Repeat both configurations and keep each result. If a small difference keeps changing direction, leave the outcome inconclusive.

Can I compare 1% lows from different tools?

Only after checking that the calculation, timing stream, capture window and exclusions agree. A percentile cutoff and an average of the slowest fraction are different summaries, even when their labels look similar.

What if average FPS improves but the lows get worse?

Keep both observations. Check that the low calculation matches, inspect the complete frametime traces and repeat the comparison. A higher average alone does not establish smoother play, and one timing spike does not identify its cause.

Does a built-in benchmark represent the whole game?

No. It can provide a repeatable scene, but that scene may differ from the demanding gameplay you care about. Keep conclusions specific to the tested scene, settings and PC.

Jonathan Houle

Written by

Jonathan Houle

Digital Marketing Manager