VIDEN

Video Enhancer

A Python CLI that extracts video frames, enhances them with Real-ESRGAN through NCNN/Vulkan, and rebuilds the video with its original audio.

Upscaling comparison

Input / bicubicReal-ESRGAN
Examples from the original Real-ESRGAN project, rather than outputs from my video pipeline.

Technologies

Python 3 / FFmpeg / ffprobe / NCNN / Vulkan

How it works

Frame extraction

Read dimensions and frame rate with ffprobe, then decode the video into numbered RGB frames. Optional working dimensions preserve aspect ratio.

What I built

GPU inference
A subprocess runs the NCNN backend with local .param and .bin model files. Model selection, scale, and tile size are passed explicitly.
Memory control
Tiled processing limits GPU memory use. The CLI accepts 2× and 4× scaling; the selected model must support the requested scale.
Video handling
ffprobe reads dimensions and frame rate. Working dimensions preserve aspect ratio and are rounded to even values for video encoding.
Progress & cleanup
Progress counts completed PNG frames. Backend failures retain a log during processing; temporary working files are cleaned up when the run finishes.
Audio
The pipeline extracts the original audio when present and muxes it back into the reconstructed video.

Model & quality tradeoffs

The CLI loads different NCNN model files, with RealESRGAN_General_WDN_x4_v3 as the current default. Scale compatibility depends on the selected model. Smaller tiles reduce memory pressure; denoising can smooth useful texture.

Frames are enhanced independently. Temporal consistency is not explicitly modeled, so video evaluation also needs to consider flicker and changing textures between frames.

My work is the orchestration and video workflow. The restoration model and NCNN/Vulkan runtime come from the upstream projects. Generated detail is an estimate of texture, not a guarantee of the original content.