VIDEN
Video Enhancer
A Python CLI that extracts video frames, enhances them with Real-ESRGAN through NCNN/Vulkan, and rebuilds the video with its original audio.
Upscaling comparison
Technologies
Python 3 / FFmpeg / ffprobe / NCNN / Vulkan
How it works
Frame extraction
Read dimensions and frame rate with ffprobe, then decode the video into numbered RGB frames. Optional working dimensions preserve aspect ratio.
What I built
- GPU inference
- A subprocess runs the NCNN backend with local
.paramand.binmodel files. Model selection, scale, and tile size are passed explicitly. - Memory control
- Tiled processing limits GPU memory use. The CLI accepts 2× and 4× scaling; the selected model must support the requested scale.
- Video handling
- ffprobe reads dimensions and frame rate. Working dimensions preserve aspect ratio and are rounded to even values for video encoding.
- Progress & cleanup
- Progress counts completed PNG frames. Backend failures retain a log during processing; temporary working files are cleaned up when the run finishes.
- Audio
- The pipeline extracts the original audio when present and muxes it back into the reconstructed video.
Model & quality tradeoffs
The CLI loads different NCNN model files, with RealESRGAN_General_WDN_x4_v3 as the current default. Scale compatibility depends on the selected model. Smaller tiles reduce memory pressure; denoising can smooth useful texture.
Frames are enhanced independently. Temporal consistency is not explicitly modeled, so video evaluation also needs to consider flicker and changing textures between frames.
My work is the orchestration and video workflow. The restoration model and NCNN/Vulkan runtime come from the upstream projects. Generated detail is an estimate of texture, not a guarantee of the original content.