Skip to content

GPU audio runtime (developer guide)

Status: experimental, default-OFF. Requires a Skia/Dawn GPU build (-DPULP_ENABLE_GPU=ON). Most paths carry a CPU fallback so plugins keep working on unsupported devices; the ones that cannot fall back say so and fail to silence rather than to wrong audio.

This is not "run your audio on the GPU." It's a runtime that lets a plugin selectively accelerate computationally expensive DSP on the GPU while keeping the audio callback bounded and preserving seamless CPU compatibility where the node declares a fallback. The audio thread is never blocked on the GPU. A node with a declared fallback can use it when the GPU cannot do the work or cannot finish before its prepared lead; nodes without one follow their declared miss policy. The callback contract is real-time-safe; GPU scheduling itself is not a hard real-time guarantee. This guide is the developer surface: the compute primitives, the real-time transport, and the ready-made processors.

The model in one paragraph

The runtime has two transport costs to manage: payload movement and scheduling. The ordinary staged path uploads and reads back each block. On supported Apple Silicon/Dawn configurations, the experimental shared-memory path imports persistent host allocations and lets CPU and GPU operate on the same storage, so the per-block CPU-to-GPU upload and GPU-to-CPU readback disappear. CPU-side planar-to-complex packing, overlap-add, and callback buffer copies still exist. Both paths still depend on non-real-time GPU submission, scheduling, and completion observation. The audio thread publishes work and reads a result a fixed number of blocks later (reported to the host as plugin delay compensation); it never waits for the GPU. To make that pay off, keep resources resident, fuse or batch enough work, and prepare explicit lead instead of blocking. Small, low-latency DSP stays on the CPU, with a continuously prepared fallback wherever one is tractable.

Honest tradeoffs (read this first)

GPU audio is easy to oversell, so here are the hard parts up front — three things to know before reaching for it:

  1. The fixed cost is real and we don't hide it. The staged path pays for upload and readback. The Apple shared-memory path can remove those copies, but it does not remove submission, scheduling, completion observation, cache traffic, or contention. For small or scalar work those fixed costs can still dwarf the compute, and the GPU is then slower than the CPU. GPU work pays off when it is large or fused enough to amortize the fixed cost (long IRs, many voices/IRs batched, big FFTs, or neural inference). That's why SuperConvolver defaults to the CPU engine and the GPU engine is opt-in, aimed at the heavy regime.

  2. It is NOT a free speed-up for any plugin. If your DSP is gain, biquads, a compressor, a small delay, an envelope follower, or MIDI — keep it on the CPU; the GPU will only make it slower. The win is narrow and specific (see the list below). We'd rather tell you that than sell you a GPU mode that regresses your plugin.

  3. Compatibility is opt-out-safe, and where it is not, it fails closed. A common disappointment with GPU audio is a plugin that won't even run on a given card. Most Pulp paths avoid that with a signal::* CPU fallback.

Not all of them can. GpuMultiConvolver (many IRs batched into one GPU submit) and Spectral Lab's cloud node have no tractable RT-safe CPU equivalent, so they declare supports_cpu_fallback = false and emit silence on a miss. That is deliberate: a bounded, obvious dropout beats plausible wrong audio. Check the processor's descriptor rather than assuming a fallback exists.

For a node that DOES declare one: if the GPU is absent, unsupported, or fails to initialize, the node logs it and runs its signal::* CPU reference — the plugin still loads and produces correct audio, just unaccelerated. On a miss mid-stream (the worker didn't finish in time) the CpuFallback policy fills the block on the CPU seamlessly — but only if that fallback is kept continuously primed and latency-aligned (see prime_fallback() below); an unprimed one substitutes a stale block with a torn history. capabilities().backend tells you at runtime which backend you actually got ("Metal"/"D3D12"/"Vulkan"), and audio paths are only validated on Apple Silicon / Metal today (see the matrix below) — everything else is experimental-but-falls-back. We won't claim a card works until we've tested it.

Bottom line: GPU audio here is a latency-tolerant accelerator for heavy, batched, parallel DSP, not a hard GPU dependency or a hard scheduling guarantee. Where a CPU fallback is tractable a node carries one; where it is not, the node says so and fails to silence rather than to wrong audio.

What belongs on the GPU (and what doesn't)

The GPU only pays off for work that is heavy, parallel, and latency-tolerant — enough arithmetic per block to dwarf the CPU↔GPU round-trip. Match the workload.

The value case — two ways GPU audio is genuinely worth it: 1. Things that are otherwise impossible / infeasible on CPU. A single plugin whose compute a CPU can't sustain in real time: large-scale physical modeling, thousands of simultaneous convolutions/partials, room/spatial modeling that touches enormous sample counts, real-time neural inference. The most compelling case is a plugin that simply couldn't exist on a CPU at all — that's where the GPU earns its keep. 2. Headroom — running more heavy work at once. Offloading demanding, batchable DSP to the otherwise-idle GPU so the session can carry more heavy instances/voices than the CPU alone would allow.

When it is NOT worth it (we'll say it plainly): if your plugin is light DSP — a basic synth, an EQ, a delay — leave the GPU out; there's little to gain and the round-trip will only make it slower. And because the GPU path is latency- tolerant by design (it adds fixed plugin-delay-compensated latency), it suits mixing, sound-design and rendering — not zero-latency live tracking/monitoring.

The crossover, concretely: GPU wins once per-block compute ≫ the round-trip, and the win grows as you scale size (longer IRs, more voices, bigger FFTs/banks). Below that line the CPU is faster — so pick per workload:

Good GPU candidates (big parallel math, or batched across instances/voices): - FFTs / iFFTs - Partitioned & long-IR convolution (and many IRs at once) - Spectral EQ, spectral freeze, spectral morphing - Large filter banks / beamforming - ML model inference (NAM / CLAP-style neural models) - Waveform & spectrogram generation; large additive / sinewave banks

Keep on the CPU (too small to beat the round-trip; latency-sensitive): - Simple gain, biquad filters, compressors - Small delays, envelope followers, MIDI processing

Rule of thumb: if it's a tight per-sample scalar op, leave it on the CPU. If it's a block-wide transform or a bank you can batch, the GPU can win — and the further you push size (long IRs, many voices, big FFTs), the bigger the win. SuperConvolver's GPU mode targets the long-IR / many-IR regime for this reason; at short IRs the CPU path is selected because it's faster there.

Which GPUs / platforms

GPU compute runs through the WebGPU layer (Dawn / wgpu-native), so it follows WebGPU's backends. Where a node declares a CPU fallback, a plugin still works with no compatible GPU — it just runs the signal::* reference path. Where it declares none (supports_cpu_fallback = false), a missed or unavailable GPU block comes out silent, never as wrong audio.

Platform Backend GPUs
macOS — Apple Silicon Metal experimental; use only with an exact provider/machine receipt
macOS — Intel Metal Intel / AMD GPUs
Windows 10+ D3D12 (or Vulkan) NVIDIA, AMD, Intel
Linux Vulkan NVIDIA, AMD, Intel
no compatible GPU CPU fallback (correct, not accelerated)

capabilities().backend reports the live backend at runtime ("Metal" / "D3D12" / "Vulkan"), plus limits and optional features (timestamp-query, f16). Audio paths require exact provider/machine validation on Apple Silicon / Metal; the other backends are supported by the WebGPU layer but not yet audio-validated — treat them as experimental until benchmarked, and rely on the CPU fallback meanwhile.

Layer 1 — compute primitives (pulp::render::GpuCompute)

Create with GpuCompute::create() then initialize_standalone() (own device, headless) or initialize_from_surface(GpuSurface&) (share the UI device and queue for compute work). Device sharing does not by itself make audio/UI buffers zero-copy. capabilities() reports backend, limits, and optional features (timestamp-query, f16). All are validated against CPU references.

Primitive Method Use
FFT / iFFT fft_forward / fft_inverse (+ fft_forward_timed) spectral analysis, fast convolution
Fused convolution prepare_convolution + convolve one-readback FFT convolution with a resident IR
Batched convolution prepare_convolution_batch + convolve_batch many blocks/instances in one submit+readback (amortizes per-dispatch + readback overhead; measured ~2× at batch 2 / stereo — the multiplier grows with batch size until it becomes bandwidth- or readback-bound)
Magnitude / complex-mul compute_magnitude, complex_multiply, batch_magnitude spectral building blocks
Dense matmul matmul neural layers, matrix DSP (ambisonics, mixing)
Dense layer dense_tanh neural inference (NAM dense / LSTM gates)
Additive synth additive_synth sum of sinusoidal partials
Modal synth modal_strike struck/plucked resonant-mode banks
Granular granular_cloud windowed pitch-shifted grain clouds

These are not real-time-safe (they block on readback); call them from the worker (Layer 2) or offline.

Layer 2 — real-time transport (pulp::gpu_audio)

GpuAudioTransport bridges the audio thread and a non-RT worker. Implement a GpuAudioNode (descriptor with channels/block/latency/miss-policy; prepare(); process_block() on the worker; process_cpu_fallback() RT-safe). prepare() the transport with the node and Config{ ring_blocks, run_worker_thread }; the audio callback calls process() — it writes the input ring and reads the latency-delayed output, never blocking. On a miss the node's MissPolicy fills the block — Silence by default (a bounded, obvious dropout), or CpuFallback if you have a fallback you keep primed and latency-aligned. Report latency_samples() to the host: the transport delays your output by latency_blocks * block_size, and a host that is not told leaves your track shifted late against every other track in the session.

The ordinary staged transport copies between its CPU rings and the provider. The experimental shared path instead keeps provider-owned slots persistently imported from host-visible storage. The callback still publishes and consumes bounded records; a non-RT service owns submission and completion. Shared memory therefore removes per-block CPU/GPU staging transfers, not all CPU-side copies and not the scheduling problem. Choose enough fixed lead for the measured workload and keep the CPU fallback ready for late, stale, failed, or unavailable GPU results.

For host diagnostics, capability_report() returns an allocation-free snapshot of the selected path. path distinguishes the ordinary staged worker from the experimental shared-memory path, while provider is Dawn only when the exact shared Dawn path is active and is otherwise Unknown. The snapshot also carries the prepared lead, miss policy, fallback availability, and whether transport diagnostics are available. GpuConvolver::backend() reports the underlying native backend (Metal for this Dawn path), while provider reports the API implementation (Dawn). Eligible means the path was accepted by prepare(); it is not a hard real-time scheduling guarantee. The report is read-only and deliberately exposes no rings, queues, callback hooks, or live path-switching controls.

Layer 3 — ready-made processors (pulp::gpu_audio)

  • GpuConvolver — FFT overlap-add convolution with a fixed IR; a continuously-fed, zero-latency signal::PartitionedConvolver CPU fallback that stays latency-aligned so a GPU miss is filled seamlessly.
  • GpuStft — STFT/ISTFT (windowed analyze / synthesize) — the spectral toolkit base.
  • GpuSpectralFreeze — capture + sustain a spectral frame (infinite pad).
  • GpuSpectralMorph — blend between two captured spectra (t ∈ [0,1]).

Authoring pattern

  1. Pick the layer: a ready processor (Layer 3), or your own GpuAudioNode whose process_block calls Layer-1 primitives.
  2. Always implement a CPU fallback (signal::*) — it's your degrade path on no-GPU devices and your miss policy.
  3. Keep small/low-latency DSP on the CPU; send only coarse, batched, heavy work to the GPU, and batch across instances/voices to amortize the round-trip.
  4. Test against a CPU reference (golden / frequency-domain) and assert RT-safety (no alloc/lock/block on the audio thread).

A minimal GPU node, end to end

A GpuAudioNode is the whole integration surface. Implement four methods — what you are, how to set up, the worker-side work, and an RT-safe fallback — then let the transport schedule it:

#include <pulp/gpu_audio/gpu_audio_node.hpp>
#include <pulp/gpu_audio/gpu_audio_transport.hpp>
#include <pulp/render/gpu_compute.hpp>

using namespace pulp;

// 1. The node: heavy, batched work that runs on the transport's worker thread.
class MyGpuNode : public gpu_audio::GpuAudioNode {
public:
    MyGpuNode(uint32_t channels, uint32_t block, uint32_t sr)
        : channels_(channels), block_(block), sr_(sr) {}

    gpu_audio::GpuAudioNodeDescriptor descriptor() const override {
        gpu_audio::GpuAudioNodeDescriptor d;
        d.name = "MyGpuNode";
        d.input_channels = channels_;
        d.output_channels = channels_;
        d.block_size = block_;
        d.sample_rate = sr_;
        d.latency_blocks = 1;   // fixed PDC — you MUST report this (see below)

        // Miss policy. Both fields default to failing CLOSED (Silence, no
        // fallback): a missed block comes out silent. That is the right default —
        // a dropout is obvious and bounded.
        //
        // Opt into CpuFallback only if you actually have a correct fallback. It
        // is not enough to implement process_cpu_fallback(): a stateful fallback
        // (a convolver, a filter) must ALSO be fed every block via prime_fallback()
        // so its history is current and latency-aligned the moment it takes over.
        // An unprimed fallback substitutes a stale block with a torn history.
        //
        // Do NOT use MissPolicy::PassthroughDry. The transport's output is delayed
        // by latency_blocks, so substituting the dry sample for the CURRENT time
        // jumps the stream forward by the whole latency for one block and back
        // again — a timeline break, not an honest dropout.
        d.miss_policy = gpu_audio::MissPolicy::CpuFallback;
        d.supports_cpu_fallback = true;
        return d;
    }

    // Non-RT setup: create the device + upload static data here, never in process_block.
    bool prepare() override {
        gpu_ = render::GpuCompute::create();
        if (!gpu_ || !gpu_->initialize_standalone()) { gpu_.reset(); return false; }
        // ... prepare Layer-1 plans (FFTs, convolution, matmul) on gpu_ ...
        return true;
    }

    // Worker thread: may touch the GPU and block on readback. NEVER the audio thread.
    void process_block(const audio::BufferView<const float>& in,
                       audio::BufferView<float>& out, uint32_t n) override {
        // ... call gpu_->fft_forward / convolve / matmul, write `out` ...
    }

    // RT-safe: called on EVERY block (hit or miss), before the output-ring read,
    // so a stateful fallback stays continuously fed and its substitute is a
    // correct, latency-aligned continuation of the stream rather than a stale
    // block with gaps in its history. Required whenever your fallback has state.
    void prime_fallback(const audio::BufferView<const float>& in,
                        uint32_t n) noexcept override {
        // ... advance your signal::* CPU path by one block and stage its output ...
    }

    // RT-safe: runs on the audio thread for a missed block. No alloc/lock/block.
    void process_cpu_fallback(const audio::BufferView<const float>& in,
                              audio::BufferView<float>& out, uint32_t n) noexcept override {
        // ... emit the substitute prime_fallback() already computed for this slot ...
    }
private:
    uint32_t channels_, block_, sr_;
    std::unique_ptr<render::GpuCompute> gpu_;
};

// 2. Wire it into your Processor. prepare() once; process() per block, RT-safe.
class MyPlugin : public format::Processor {
    MyGpuNode node_{2, 256, 48000};
    gpu_audio::GpuAudioTransport transport_;

    void prepare(const format::PrepareContext&) override {
        node_.prepare();
        transport_.prepare(&node_, {/*ring_blocks*/ 8, /*run_worker_thread*/ true});
    }

    // REPORT THE LATENCY. The transport delays your output by
    // latency_blocks * block_size, and the host cannot know that unless you say
    // so — it will leave your track shifted late against every other track in
    // the session, and nothing will error. This override is not optional.
    int latency_samples() const override { return transport_.latency_samples(); }

    void process(audio::BufferView<float>& out,
                 const audio::BufferView<const float>& in, /*...*/) override {
        transport_.process(in, out, out.num_samples());  // writes input, reads delayed output — never blocks
    }
};

Prove that latency_samples() is honest rather than trusting it — the number is hand-maintained and drifts the moment the block size or latency_blocks changes:

pulp audio render --plugin MyPlugin.clap --out /tmp/o.wav --duration-ms 500 \
    --input-signal noise --param <your-bypass-id>=1 --latency-report /tmp/lat.json

It exits nonzero if the reported latency does not match the delay actually in the audio. See the CLI reference.

That is the entire contract: the audio thread only ever calls transport_.process() (lock-free), the GPU work happens on the worker a fixed number of blocks behind, and a missed block degrades via your MissPolicy instead of glitching.

Worked examples

Both ship as buildable plugins (VST3/AU/CLAP/Standalone) with a CPU default and an opt-in GPU engine, a CPU-reference correctness test, and a benchmark:

  • examples/super-convolver — multi-room convolution reverb. Shows batching many distinct IRs into one GPU submit (GpuMultiConvolver), the live CPU↔GPU engine switch, and a live status line (backend, blocks, µs/block, real-time headroom). The clearest read for the transport + live-rebuild lifecycle.
  • examples/spectral-lab — N-layer spectral freeze/morph "cloud". Shows authoring your own GpuAudioNode (gpu_spectral_cloud_node.hpp) over a GPU-resident spectral stack, and reusing the same status-line + headroom telemetry. The clearest read for writing a node from scratch.

Building against an installed Pulp SDK

Everything above works whether Pulp is a source submodule or an installed SDK. A standalone plugin repo that consumes Pulp through find_package(Pulp) (rather than add_subdirectory) gets the GPU audio runtime from the SDK's exported targets:

You link You get
pulp::gpu-audio Layer 2 transport + Layer 3 ready-made processors (GpuAudioTransport, GpuConvolver, GpuStft, GpuSpectralFreeze/Morph) and their headers
pulp::render Layer 1 primitives (GpuCompute) — present only in a GPU-enabled SDK build

pulp::gpu-audio is always exported: it's GPU-agnostic and carries its own CPU fallback, so a non-GPU SDK still exposes the transport (it just runs the signal::* reference path). pulp::render exists only when the SDK was built -DPULP_ENABLE_GPU=ON, so gate any GPU-primitive code on it:

find_package(Pulp CONFIG REQUIRED)

pulp_add_plugin(MyPlugin
    FORMATS      VST3 AU CLAP Standalone
    SOURCES      my_plugin.cpp
    PLUGIN_NAME  "MyPlugin"
    # ...
)

target_link_libraries(MyPlugin_Core PUBLIC pulp::gpu-audio)
if(TARGET pulp::render)
    target_link_libraries(MyPlugin_Core PUBLIC pulp::render)  # Layer-1 GPU primitives
endif()

find_package(Pulp) resolves the transport's private dependencies for you (Threads, and ICU on Linux) — a static library propagates its private link deps into its exported interface, so PulpConfig re-finds them; you don't list them yourself.

Ship a self-contained GPU bundle

A GPU plugin loads libwgpu_native.dylib at runtime. From a source build the WebGPU dependency copies that dylib next to each format binary automatically; an installed-SDK consumer has no such upstream copy step, so a find_package(Pulp) plugin must place the dylib itself and bake an @loader_path rpath — otherwise the bundle passes every check on the build machine and fails to load ("Library not loaded: @rpath/libwgpu_native.dylib") on any other Mac. The SDK exports two helpers for exactly this. find_package(Pulp) already defines them — no extra include() — so call them directly:

foreach(_fmt MyPlugin_CLAP MyPlugin_VST3 MyPlugin_AU MyPlugin_Standalone)
    if(TARGET ${_fmt})
        pulp_make_bundle_relocatable(${_fmt})       # copy dylib + bake @loader_path
        pulp_validate_bundle_relocatable(${_fmt})   # POST_BUILD guard: fail the build if not
    endif()
endforeach()

pulp_validate_bundle_relocatable runs at build time and fails hard if a bundle still references a build-machine-absolute path, so the footgun is caught in CI instead of on a customer's machine. Apply it to single-bundle formats (Standalone/CLAP/VST3/AU); an AUv3 .appex is intentionally not relocatable in isolation — its dylib and framework live in the container app's Contents/Frameworks and resolve by @rpath only inside that container, so the container app (not the standalone appex) is the relocatability unit.

GPU work runs outside the audio callback, with CPU fallback when its result is unavailable. This bounds the callback's interaction with the GPU; it does not provide a hard scheduling guarantee from the operating system or driver.

For the experimental shared-memory route, the paced convolution probe records callback timing, missed deliveries, and numerical correctness through the public transport and its own worker. Use the tracing guide for per-block lifecycle analysis.