Skip to content

GPU compositor: effect stacks round-trip through the CPU, intermediates allocate fresh textures, and Noise on an adjustment layer re-uploads its footprint every frame #687

Description

@MatejKozlovsky

Version: main (after #641/#648) · Area: GPU compositor (crates/gpu), layer cache (crates/render/src/cache.rs)

What happens

On the GPU compositor (--gpu, Mercury GPU Acceleration), a 4K comp with effect stacks spends most of its
time moving pixels rather than compositing:

  1. Effect stacks round-trip. A layer with effects is processed by the CPU path (styled_layer), which
    runs each run of GPU-capable effects as upload → kernels → readback, then uploads the result again to
    place it.
  2. Every intermediate is a new texture. Within one frame, a working texture released by the encoder
    that made it waits in the shared pool until that encoder submits, so a frame allocates a new texture for
    almost every intermediate (a 4K RGBA f32 texture is 127 MB).
  3. A time-dependent effect re-keys an adjustment layer's footprint. The footprint (source through the
    masks, without effects) is keyed with the effects included, so Noise or Grain on a finishing adjustment
    layer changes the key every frame and the GPU re-uploads the full-frame footprint each time.
  4. The upload cache evicts silently. When static layers exceed the fixed 1536 MB upload budget, layer
    textures are evicted and re-uploaded every frame without any log line, and an embedder cannot size the
    budget to the device.

Measured

On a test PC (Ryzen 7 5800X, RTX 3090 24 GB, Windows 10), effectcraft-cli bench --gpu on a 4K comp
with effect stacks (main, median of 5): GPU cold 1181 ms/frame, GPU warm 208 ms/frame. A 3 s / 90-frame
4K export with --gpu (RAYON_NUM_THREADS=1, so frames stay on the GPU) takes 81.1 s with a 17.2 GB
VRAM peak.

Expected

Effect stacks stay resident on the GPU, a frame reuses its own intermediates, an adjustment layer's
footprint keeps its key across frames, and a full upload cache is reported.

Activity

  1. vbasky commented on Oct 11, 2026

    @vbasky
    Contributor

    Partial: #701 fixes two of the four — the adjustment footprint no longer re-keys each frame under a time-dependent effect (127 MB/frame re-upload gone), and over-budget upload eviction is now counted, logged once and settable via GpuContext::set_upload_budget.

    Deliberately not attempted: stacks staying GPU-resident (#1) and per-frame intermediate texture reuse (#2). Both are the measured wins but need a DX12/Vulkan/Metal device to know a change helped; every box here has only GL (declined by the compositor) or llvmpipe, so any number we would post would be fictional. The footprint fix does cut one full-frame upload per frame for finished comps; the residency rewrite and the intermediate pool want a run of effectcraft-cli bench --gpu before and after on hardware like yours — a maintainer with such a machine should drive that, not a guess-and-check patch.

  2. MatejKozlovsky commented on Oct 11, 2026

    @MatejKozlovsky
    ContributorAuthor

    Thanks @vbasky, nice work on #701!

    Heads-up so we don't duplicate effort: #688 already covers all four items from this issue, including the two in #701 (the footprint key ignoring time-dependent effects, and the upload-cache budget with a single warning when it's full). It also adds the two GPU parts you mentioned leaving to someone with real hardware: resident effect stacks and in-frame texture reuse.

    I measured it on an RTX 3090 test PC (4K, 90 frames, RAYON_NUM_THREADS=1 so frames stay on the GPU): 81.1 s → 49.3 s (1.65×), with peak VRAM 17.2 GB → 9.4 GB. Output is pixel-identical to main's GPU render, and CPU output is byte-identical to main.

    @bflatastic, happy to go whichever way is easier. If you'd rather merge #701 first, I'll rebase #688 on top of it so it only carries the GPU parts.

  3. vbasky commented on Oct 11, 2026

    @vbasky
    Contributor

    @MatejKozlovsky thanks — we will not duplicate the GPU-resident stacks or in-frame texture reuse. Those belong in #688, and your 3090 numbers are the measurement we could not take here.

    #701 stays the small cache-key / upload-budget slice (items 3–4 only). If maintainers land that first, rebasing #688 on top so it only carries the GPU parts is the easier split. If #688 merges first, we can close #701 as already covered.

    @bflatastic either order is fine from this side.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions