SLP Tuner Launcher v1.0.2 (Major SPEED BOOST) with Benchmark & Autotune built-in!

I started this new thread because, with the new Autotune feature and “Take Screenshot” button built directly into the launcher, it is now extremely easy for users to share their actual results.

The previous thread had a lot of comments along the lines of “I feel like it’s faster” or “I don’t notice any difference,” often without including hardware configuration, settings, benchmarks, or other information that would help me diagnose the results or help the community understand what actually works and what doesn’t.

I’d like people to use this thread to share their findings along with screenshots, so we can build a much more useful collection of real-world results across different hardware and output resolutions.

The launcher’s Conservative Starting Point has been tweaked many times, but it is still essentially a ballpark estimate extrapolated from benchmark data collected on my own machine. I do not recommend simply sticking with that starting point, because it has not actually been tested or optimized for your particular hardware and target output resolution.

Instead, I recommend running Autotune and letting the algorithm determine the optimal Conservative and Aggressive settings for your hardware at the output resolution you actually intend to use.

This is important because the same parameters can affect performance very differently at different output resolutions. A setting that works well at 1920x1080, for example, may behave very differently at 1440x1080, or 960x720.

Faster SLP Preflight

This version also includes a feature intended to fix the extremely long pre-flight loading time before SLP even starts processing.

It has worked very well on my machine, but I’d really like to hear whether it fixes the problem for other users as well.

To test it, make sure you check:

“Enable faster threaded SLP preflight”

If you try it, please post your results here — ideally with screenshots (use Take screenshot button) showing your hardware, resolution, settings, and Autotune results. The more real benchmark data we collect, the better we can understand what settings work best across different systems.

Topaz SLP Tuning Launcher v1.0.2

v1.0.2 is a major reliability, safety, monitoring, and automatic-tuning release for the portable Windows launcher.

Why I made it

I have tried nearly every AI video enhancer I could find, from precision/traditional models such as Proteus, Iris, Rhea, UniFab, Aiarty, VikPea, and Nero to generative/diffusion restorers such as SEEDVR2, FlashVSR, SLM, SLP, and LTX-2.5 LoRA. For my material and standards, SLP is the undisputed king of quality—but it is slow. I built this launcher to make SLP faster, safer to tune, easier to understand, and much more enjoyable to use. SLP once ran slowly enough on my workstation that I sometimes used SEEDVR2 for less-important videos; with the gain I now get from this launcher, I personally no longer have a reason to do that. Your results will depend on your own hardware, source, driver, and output resolution.

Important RAM + VRAM correction

The v1.0.0 and v1.0.1 starting presets used installed VRAM but did not account for system RAM. v1.0.1 also made several preset recipes more aggressive than v1.0.0, which could leave VRAM-rich but RAM-constrained systems vulnerable to paging or OOM. A separate migration defect could silently refresh an older serialized built-in preset by name and substitute the newer, more memory-intensive recipe.

v1.0.2 replaces that approach with a conservative starting-point matrix selected from both physical RAM and VRAM. It also preserves the exact values of a differing older recipe as a custom or unsaved preset instead of silently replacing it. The matrix is only a rough 1080p starting point, not a guarantee. Run at least a Standard AutoTune at every output resolution you use: VAE encode/decode tiling and other settings can gain speed at one geometry but become neutral or slower at another.

Benchmark / System AutoTune

  • Builds guarded Quick, Standard, or Thorough plans from a deterministic clip or your own source. Standard discards the first warm-up chunk and measures the next two; Thorough repeats that three-chunk measurement three times in randomized order with drift controls.
  • Compares an explicitly verified baseline cuDNN runtime with the upgraded runtime without repeatedly modifying the installed Topaz DLLs. Prepare for offline run downloads, hashes, extracts, and freezes both child-local runtimes before the campaign.
  • Runs bounded one-factor tests, then derives and separately validates conservative and aggressive combined presets. Successfully measured, warning-free custom-suite rows are included in that decision evidence.
  • Tracks whole-system RAM, Windows commit, dedicated VRAM, workload deltas, warning evidence, exact settings revisions, output integrity, and matched repetition baselines.
  • Exports an evidence PDF report with speed, memory, efficiency, and frontier views.
  • Adds compact, responsive Benchmark controls; explicit mode hover help; editable custom suites; resolution-aware temporal and tile candidates; and sortable live Results/efficiency tables.

For the most trustworthy results, close unnecessary applications and leave the machine as idle as possible during testing. After any required online preparation is complete, consider disconnecting from the internet to reduce the chance of background downloads, updates, cloud syncing, or other network-related activity interfering with the benchmark. Even light typing, web browsing, or other seemingly minor activity can reduce measured FPS.

It is also important to keep the system thermally stable throughout the test. Make sure the computer has adequate ventilation, and use an open window or air conditioning if necessary. During a long Autotune campaign, the CPU, GPU, case, and even the room itself can gradually heat-soak, causing performance to drift over time.

During one of my approximately 19-hour test sessions, I experienced a drastic 16–17% drop in FPS across the board. Because of this, I strongly recommend minimizing anything that could introduce variability into the Autotune results. The goal is to keep the testing environment as consistent as possible so the settings selected by Autotune reflect the actual performance of your hardware rather than temporary background activity or thermal conditions.

Speed and quality

The launcher does not increase speed at the expense of output quality. It uses the same SLP model and exposes the same runner’s scheduling, chunking, overlap, tiling, memory, cuDNN, and preflight behavior. My extensive SEEDVR2 testing found that pushing these controls can slightly improve quality while using more memory: larger or untiled spatial regions create fewer artificial tile boundaries, and lower overlap reduces the amount of independently processed imagery that must be blended. Larger temporal chunks can give the model longer continuous context and reduce chunk-boundary resets, potentially improving motion continuity and jitter.

Those are tradeoffs, not universal rules. Zero spatial or temporal overlap can expose seams; this project observed localized seams at zero overlap and therefore keeps guarded floors. Larger or untiled VAE geometry can consume far more RAM/VRAM and can even run slower at some resolutions. If you have substantial spare headroom - especially 48 GB VRAM or more -you may still test larger or untiled spatial regions for quality even when they are not the fastest setting. Topaz already starts at a relatively large chunk 121, so the visible improvement over stock SLP is likely smaller than the differences I saw in SEEDVR2.

Judge representative output yourself with SKV89 Video Compare. Static frames can reveal spatial boundary/blending differences, but they cannot demonstrate smoother motion or reduced jitter. Use modes 2, 3, or 4 to play the videos side by side.

Live monitoring, warnings, and history

  • Adds process-attributed SLP used RAM, SLP used VRAM, and SLP shared GPU memory current/peak readings with the phase at which each peak occurred, alongside whole-system capacity and peak readings.
  • Rising SLP dedicated VRAM near capacity together with sustained shared-GPU growth and rising RAM/commit is a useful visual signal of spill or thrashing. Shared GPU memory alone is not proof and never triggers an alert by itself.
  • Adds sustained RAM, commit, VRAM, likely-spill, and no-progress warnings. The launcher never terminates anything automatically. If the exact worker PID can be verified as neuroserver.exe, the user may explicitly terminate only that process tree; Topaz Video itself is never killed.
  • Reconciles log state with the live process list so canceled or missing runners no longer leave a stale VAE-decode phase or warning behind.
  • Makes long-chunk tracking resilient to repeated intermediate RUNNING records and long log gaps, so the active counter and FPS history do not reset without a genuine later runner start.
  • Adds persistent Completed files, including SLP output format, and separate Error history / resume evidence. Recovery is an honest frame-0 retry under captured settings; unsafe pixel-changing mid-file continuation is not exposed.
  • Adds a redacted, rotating diagnostic Live log and bounded searchable/exportable on-screen history.

Settings and preflight

  • Adds Detect & restore Topaz defaults. Unknown native values are never guessed; a successful default restore disables faster threaded preflight.
  • Makes Apply Settings visibly active for every pending launch change, including the threaded-preflight checkbox. Ordinary SLP values apply to later SLP runners in the existing launched Topaz session. Threaded preflight is chosen when Topaz starts, so apply it, fully close Topaz, and use Launch Topaz here.
  • Keeps the opt-in threaded SLP 2.6 frame-count hook process-local and reversible. It adds decoder threading to the same exact ffprobe -count_frames operation, falls back to native behavior, and does not alter Topaz, FFmpeg, Windows, or source media.
  • Adds Take screenshot, which maximizes the launcher, waits for stable window geometry, then copies the complete launcher to the clipboard. Open the relevant tab and paste it into a support or results reply with Ctrl+V; without the settings/monitor/results evidence, a report cannot be meaningfully diagnosed and offers little help to the community.

cuDNN 9.24 and lower-memory NVIDIA cards

cuDNN 9.24.0.43 for CUDA 12 remains the fastest package tested by this project with Topaz’s CUDA 12.8 runtime. Compatible 12 GB and 16 GB NVIDIA cards should theoretically be able to benefit from that runtime even when SLP tuning values remain stock, because the cuDNN comparison does not require selecting larger chunks or tiles. This is not a promise for hardware I do not own. Please do not ask me whether your hardware configuration will benefit from this app or ask me to predict its exact gain. Run AutoTune at your desired SLP output resolution and share the complete results so others with the same hardware can learn from them.

Settings and cuDNN management

Topaz SLP Tuning Launcher settings and cuDNN management

Monitor / FPS and completed-file history

Topaz SLP Tuning Launcher Monitor and completed files

Benchmark / System AutoTune plan

Topaz SLP Tuning Launcher Benchmark and System AutoTune plan

Benchmark results and efficiency

Topaz SLP Tuning Launcher Benchmark results and efficiency

RAM + VRAM starting guide

Topaz SLP Tuning Launcher RAM and VRAM starting guide

I’m not sure why this thread does not show up under the General section where I post it. I just want to emphasize that this app should not be viewed as being in conflict with Topaz’s interests. If anything, its purpose is to make Topaz SLP faster, easier, and more enjoyable to use.

By helping users get better performance out of SLP on their existing hardware, the launcher strengthens what is already, in my opinion, the undisputed leader in generative video restoration. Faster processing makes the overall experience substantially better, which should translate into happier customers who are more likely to keep using SLP rather than looking elsewhere for faster or better alternatives.

The goal is not to replace or compete with anything Topaz provides - it is simply to help users get the most out of SLP. This launcher app is simply a COMPANION to Topaz SLP.

Thanks for the new update! regarding the autotune resolution, should that be the upscaled output resolution or source resolution?
I typically upscale 480x360 3x to 1440x1080p
or 640x480 2x to 1280x960p.

Should I set autotune resolution to source or the upscale output resolution?

Tried to run the autotune, and after it went through everything, all of them failed. Not sure what is wrong???

Despite the autotune feature not working, I went ahead and used my test footage that I have been testing with on 1.0.0 and used my settings that worked in that version, and this is what I got…

Depite checking the option for faster preload, for my apple prores, I went from arounf 10 sec to 25 sec, and the entire thing took nearly 28 mins to encode start to finish, up from around 13-14 mins in 1.0.0, and the fps went from 2.2 fps in 1.0.0 down to 1.2 fps. As much as I hate to report it, the new version is performing worse by around 50%

I seemingly have that speed problem with 1.0.2, too: With identical settings it’s much slower for me than 1.0.0.

I used identical settings and didn’t change any background activity for both versions of the launcher. 1.0.0 completed 2 chunks in 6:41 whereas 1.0.2 took 7:06 for only 1 chunk.

That is the output resolution. It is the output resolution that determines the speed and peak vram and peak sysram.

The slowdown came from vram oom and it is spilling over to shared memory. But if the settings were the same from v1.0.0 that worked without oom, this is very strange. I fed the code from v1.0.0 and v1.0.2 into AI and found no issues. I also ran tests yesterday with v1.0.0, v1.0.1, and v1.0.2 with the same settings and the ram and vram usage was the same so this is all very strange. Can you confirm it is the exact same settings and output resolution that you tried between v1.0.0 and v1.0.1.

Yes I confirmed there are issues with the autotune. Let me look into that.

That’s very interesting. Thanks for sharing. Maybe these can fix the dirty fix problem. But right now, I just want to get the core function of the app, which is autotune / speed boost working universally. It is easy to get it to work for one machine but getting it to work for limitless amount of possible hardware configurations and potentially different Topaz versions, proves to be far more difficult than I thought. I’m starting to understand why Topaz just chose 16gb vram as the optimal level to tune for. I’m going to take break for awhile and get back to real life and real work after I solve the autotune issue.

When updating cuDNN on the new SLP Tuner.

Thanks for sharing those screenshots. The difference between 6:41 and 7:06 is pretty minor and can just be random fluctuations between runs due to many possible factors. I experienced this often in my own tests when running on the same launcher at the same output resolution. But your vram is at the limit. Let me do some more work on this.

You’re welcome. As you previously wrote, it’s similar to SLP SeedVR2.6, and these improvements stem from using other filters in parallel. I could actually be a beta tester for your application; you could send me beta versions, and I’ll test them before release, because it works on your hardware but gives errors on other hardware. In fact, I used version 1.0.2 today, and it gave errors, so I closed it.

I think that might be the pathname shown in the error message being too long. I just tested it again on my system and it works. But the next version will be more tolerant.

I’m thinking of making v1.0.2 a beta and will release another v1.0.2 for public beta testing. v1.0.3 will be the stable version.

The experimental preload helps tremendously. For this current job, preload on 1.0.0 took roughly 25mins, on 1.0.2 it was less than 5mins.

However, I ran into OOM error on 1.0.2 even with stock/Topaz default settings. I have switched back to 1.0.0 for the current job:

Current job config on 1.0.0:

Current job’s stats via 1.0.0:

Current job’s stats (launched via 1.0.0) stats viewed through 1.0.2:

Sorry, it’s 6:41 for two (2) chunks of 121 with 1.0.0 and 7:06 for one (1) chunk of 121 with 1.02., so less than half the speed.

yea that’s due to vram oom. Once you see vram spilling over to shared gpu memory, it slows down drastically

I hope these tests are useful for your developement: I found what appears to be a reproducible v1.0.2 VRAM/commit regression on an RTX 4090 tested today. I performed a clean reinstall of Topaz Video 1.7.0 to its default path before retesting both launcher versions.

System and workload:

  • RTX 4090 24 GB

  • Ryzen 7 9850X3D

  • 48 GB installed system RAM (47.3 GiB reported)

  • Topaz Video 1.7.0, SLP 2.6

  • 30-second 640×480 source, enhanced 2× to 1280×960

  • Topaz-native cuDNN 9.7.1.26 in every test

  • PyTorch 2.7.0+cu128 / CUDA 12.8

  • Experimental faster threaded preflight disabled

  • Warm-up run performed before the measured 30-second export

Direct Topaz 1.7.0, without tuner launcher:

  • Render time: 11:43

  • System RAM peak: 23.7 GiB

  • Dedicated VRAM peak: 16.3 GiB

  • Completed successfully

Launcher v1.0.2, “Stock Topaz” preset with no SLP tuning override:

  • Render time: 11:56

  • System RAM peak: 24.0 GiB

  • Dedicated VRAM peak: 22.9 GiB

  • SLP-used VRAM peak: 21.6 GiB

  • Shared GPU memory peak: 11.6 GiB

  • Completed successfully

The large VRAM difference between direct Topaz and the v1.0.2 Stock Topaz launch is already present before applying custom tuning.

I then compared these exact effective settings:

  • Temporal chunk: 361

  • Temporal overlap: 8

  • VAE operator cap: 4 GiB

  • VAE temporal micro-batch: 4

  • Encode tiled: enabled, 640 / overlap 80

  • Decode tiled: enabled, 480 / overlap 32

  • Full model pool: off

  • DiT block and MLP chunks: Topaz defaults

  • Attention-window group: 10

With launcher v1.0.0:

  • Render time: 10:19

  • System RAM peak: 27.0 GiB

  • Dedicated VRAM peak: 16.0 GiB

  • Completed successfully

With launcher v1.0.2 and the same displayed/effective SLP values:

  • System RAM reached 38.3 GiB

  • Dedicated VRAM reached 23.5 GiB

  • Shared GPU memory reached 29.2 GiB

  • SLP shared GPU memory reached 29.1 GiB

  • Windows commit reached 64.6/65.0 GiB

  • The run entered severe spill/thrashing and was not allowed to complete

I also tested the more aggressive 481 temporal chunk plus 640 decode tile:

Launcher v1.0.0:

  • Render time: 10:03

  • System RAM peak: 29.7 GiB

  • Dedicated VRAM peak: 19.3 GiB

  • Completed successfully

Launcher v1.0.2:

  • Topaz reported “Out of memory” after 2:10

  • System RAM peak: 42.4 GiB

  • Dedicated VRAM peak: 23.5 GiB

  • SLP-used VRAM peak: 22.7 GiB

  • Shared GPU memory peak: 29.0 GiB

I verified the effective child-settings lines. The visible SLP parameters match between v1.0.0 and v1.0.2. The one additional visible variable in v1.0.2 is:

TOPAZ_SLP_COLOR_CORRECTION=wavelet

Since wavelet is also the native Topaz default, I would not expect that alone to explain the difference.

Do any of these improvements help with the ghosting in SLP that was also a huge issue in SEEDVR?

I tried the settings in your description using v1.0.0 and v1.0.2 but couldn’t see a difference. But regardless, just keep a copy of v1.0.0 as a fall back. Use v1.0.2.1 for autotune to determine the best settings for your hardware configuration at the target output resolution. If v1.0.2.1 gives you oom issues, you can just copy in those optimum settings into v1.0.0. You will lose the pre-flight/loading fix though. Launch Topaz with v1.0.0. But you can still launch v1.0.2.1 for the improved live readings.