GPU audio runtime (developer guide)¶
Status: experimental, default-OFF. Requires a Skia/Dawn GPU build (
-DPULP_ENABLE_GPU=ON). Most paths carry a CPU fallback so plugins keep working on unsupported devices; the ones that cannot fall back say so and fail to silence rather than to wrong audio.
This is not "run your audio on the GPU." It's a runtime that lets a plugin selectively accelerate computationally expensive DSP on the GPU while keeping the audio callback bounded and preserving seamless CPU compatibility where the node declares a fallback. The audio thread is never blocked on the GPU. A node with a declared fallback can use it when the GPU cannot do the work or cannot finish before its prepared lead; nodes without one follow their declared miss policy. The callback contract is real-time-safe; GPU scheduling itself is not a hard real-time guarantee. This guide is the developer surface: the compute primitives, the real-time transport, and the ready-made processors.
The model in one paragraph¶
The runtime has two transport costs to manage: payload movement and scheduling. The ordinary staged path uploads and reads back each block. On supported Apple Silicon/Dawn configurations, the experimental shared-memory path imports persistent host allocations and lets CPU and GPU operate on the same storage, so the per-block CPU-to-GPU upload and GPU-to-CPU readback disappear. CPU-side planar-to-complex packing, overlap-add, and callback buffer copies still exist. Both paths still depend on non-real-time GPU submission, scheduling, and completion observation. The audio thread publishes work and reads a result a fixed number of blocks later (reported to the host as plugin delay compensation); it never waits for the GPU. To make that pay off, keep resources resident, fuse or batch enough work, and prepare explicit lead instead of blocking. Small, low-latency DSP stays on the CPU, with a continuously prepared fallback wherever one is tractable.
Honest tradeoffs (read this first)¶
GPU audio is easy to oversell, so here are the hard parts up front — three things to know before reaching for it:
-
The fixed cost is real and we don't hide it. The staged path pays for upload and readback. The Apple shared-memory path can remove those copies, but it does not remove submission, scheduling, completion observation, cache traffic, or contention. For small or scalar work those fixed costs can still dwarf the compute, and the GPU is then slower than the CPU. GPU work pays off when it is large or fused enough to amortize the fixed cost (long IRs, many voices/IRs batched, big FFTs, or neural inference). That's why SuperConvolver defaults to the CPU engine and the GPU engine is opt-in, aimed at the heavy regime.
-
It is NOT a free speed-up for any plugin. If your DSP is gain, biquads, a compressor, a small delay, an envelope follower, or MIDI — keep it on the CPU; the GPU will only make it slower. The win is narrow and specific (see the list below). We'd rather tell you that than sell you a GPU mode that regresses your plugin.
-
Compatibility is opt-out-safe, and where it is not, it fails closed. A common disappointment with GPU audio is a plugin that won't even run on a given card. Most Pulp paths avoid that with a
signal::*CPU fallback.
Not all of them can. GpuMultiConvolver (many IRs batched into one GPU
submit) and Spectral Lab's cloud node have no tractable RT-safe CPU
equivalent, so they declare supports_cpu_fallback = false and emit silence
on a miss. That is deliberate: a bounded, obvious dropout beats plausible
wrong audio. Check the processor's descriptor rather than assuming a
fallback exists.
For a node that DOES declare one: if the GPU is absent, unsupported, or fails
to initialize, the node logs it and runs its signal::* CPU reference — the
plugin still loads and produces correct audio, just unaccelerated. On a miss
mid-stream (the worker didn't finish in time) the CpuFallback policy fills
the block on the CPU seamlessly — but only if that fallback is kept
continuously primed and latency-aligned (see prime_fallback() below); an
unprimed one substitutes a stale block with a torn history.
capabilities().backend tells you at runtime
which backend you actually got ("Metal"/"D3D12"/"Vulkan"), and audio paths are
only validated on Apple Silicon / Metal today (see the matrix below) —
everything else is experimental-but-falls-back. We won't claim a card works
until we've tested it.
Bottom line: GPU audio here is a latency-tolerant accelerator for heavy, batched, parallel DSP, not a hard GPU dependency or a hard scheduling guarantee. Where a CPU fallback is tractable a node carries one; where it is not, the node says so and fails to silence rather than to wrong audio.
What belongs on the GPU (and what doesn't)¶
The GPU only pays off for work that is heavy, parallel, and latency-tolerant — enough arithmetic per block to dwarf the CPU↔GPU round-trip. Match the workload.
The value case — two ways GPU audio is genuinely worth it: 1. Things that are otherwise impossible / infeasible on CPU. A single plugin whose compute a CPU can't sustain in real time: large-scale physical modeling, thousands of simultaneous convolutions/partials, room/spatial modeling that touches enormous sample counts, real-time neural inference. The most compelling case is a plugin that simply couldn't exist on a CPU at all — that's where the GPU earns its keep. 2. Headroom — running more heavy work at once. Offloading demanding, batchable DSP to the otherwise-idle GPU so the session can carry more heavy instances/voices than the CPU alone would allow.
When it is NOT worth it (we'll say it plainly): if your plugin is light DSP — a basic synth, an EQ, a delay — leave the GPU out; there's little to gain and the round-trip will only make it slower. And because the GPU path is latency- tolerant by design (it adds fixed plugin-delay-compensated latency), it suits mixing, sound-design and rendering — not zero-latency live tracking/monitoring.
The crossover, concretely: GPU wins once per-block compute ≫ the round-trip, and the win grows as you scale size (longer IRs, more voices, bigger FFTs/banks). Below that line the CPU is faster — so pick per workload:
Good GPU candidates (big parallel math, or batched across instances/voices): - FFTs / iFFTs - Partitioned & long-IR convolution (and many IRs at once) - Spectral EQ, spectral freeze, spectral morphing - Large filter banks / beamforming - ML model inference (NAM / CLAP-style neural models) - Waveform & spectrogram generation; large additive / sinewave banks
Keep on the CPU (too small to beat the round-trip; latency-sensitive): - Simple gain, biquad filters, compressors - Small delays, envelope followers, MIDI processing
Rule of thumb: if it's a tight per-sample scalar op, leave it on the CPU. If it's a block-wide transform or a bank you can batch, the GPU can win — and the further you push size (long IRs, many voices, big FFTs), the bigger the win. SuperConvolver's GPU mode targets the long-IR / many-IR regime for this reason; at short IRs the CPU path is selected because it's faster there.
Which GPUs / platforms¶
GPU compute runs through the WebGPU layer (Dawn / wgpu-native), so it follows
WebGPU's backends. Where a node declares a CPU fallback, a plugin still works
with no compatible GPU — it just runs the signal::* reference path. Where it
declares none (supports_cpu_fallback = false), a missed or unavailable GPU
block comes out silent, never as wrong audio.
| Platform | Backend | GPUs |
|---|---|---|
| macOS — Apple Silicon | Metal | experimental; use only with an exact provider/machine receipt |
| macOS — Intel | Metal | Intel / AMD GPUs |
| Windows 10+ | D3D12 (or Vulkan) | NVIDIA, AMD, Intel |
| Linux | Vulkan | NVIDIA, AMD, Intel |
| no compatible GPU | — | CPU fallback (correct, not accelerated) |
capabilities().backend reports the live backend at runtime ("Metal" / "D3D12" /
"Vulkan"), plus limits and optional features (timestamp-query, f16). Audio paths
require exact provider/machine validation on Apple Silicon / Metal; the other
backends are supported by the WebGPU layer but not yet audio-validated — treat
them as experimental until benchmarked, and rely on the CPU fallback meanwhile.
Layer 1 — compute primitives (pulp::render::GpuCompute)¶
Create with GpuCompute::create() then initialize_standalone() (own device,
headless) or initialize_from_surface(GpuSurface&) (share the UI device and
queue for compute work). Device sharing does not by itself make audio/UI buffers
zero-copy. capabilities() reports backend, limits, and optional features
(timestamp-query, f16). All are validated against CPU references.
| Primitive | Method | Use |
|---|---|---|
| FFT / iFFT | fft_forward / fft_inverse (+ fft_forward_timed) |
spectral analysis, fast convolution |
| Fused convolution | prepare_convolution + convolve |
one-readback FFT convolution with a resident IR |
| Batched convolution | prepare_convolution_batch + convolve_batch |
many blocks/instances in one submit+readback (amortizes per-dispatch + readback overhead; measured ~2× at batch 2 / stereo — the multiplier grows with batch size until it becomes bandwidth- or readback-bound) |
| Magnitude / complex-mul | compute_magnitude, complex_multiply, batch_magnitude |
spectral building blocks |
| Dense matmul | matmul |
neural layers, matrix DSP (ambisonics, mixing) |
| Dense layer | dense_tanh |
neural inference (NAM dense / LSTM gates) |
| Additive synth | additive_synth |
sum of sinusoidal partials |
| Modal synth | modal_strike |
struck/plucked resonant-mode banks |
| Granular | granular_cloud |
windowed pitch-shifted grain clouds |
These are not real-time-safe (they block on readback); call them from the worker (Layer 2) or offline.
Layer 2 — real-time transport (pulp::gpu_audio)¶
GpuAudioTransport bridges the audio thread and a non-RT worker. Implement a
GpuAudioNode (descriptor with channels/block/latency/miss-policy; prepare();
process_block() on the worker; process_cpu_fallback() RT-safe). prepare()
the transport with the node and Config{ ring_blocks, run_worker_thread }; the
audio callback calls process() — it writes the input ring and reads the
latency-delayed output, never blocking. On a miss the node's MissPolicy
fills the block — Silence by default (a bounded, obvious dropout), or
CpuFallback if you have a fallback you keep primed and latency-aligned. Report
latency_samples() to the host: the transport delays your output by
latency_blocks * block_size, and a host that is not told leaves your track
shifted late against every other track in the session.
The ordinary staged transport copies between its CPU rings and the provider. The experimental shared path instead keeps provider-owned slots persistently imported from host-visible storage. The callback still publishes and consumes bounded records; a non-RT service owns submission and completion. Shared memory therefore removes per-block CPU/GPU staging transfers, not all CPU-side copies and not the scheduling problem. Choose enough fixed lead for the measured workload and keep the CPU fallback ready for late, stale, failed, or unavailable GPU results.
For host diagnostics, capability_report() returns an allocation-free snapshot
of the selected path. path distinguishes the ordinary staged worker from the
experimental shared-memory path, while provider is Dawn only when the exact
shared Dawn path is active and is otherwise Unknown. The snapshot
also carries the prepared lead, miss policy, fallback availability, and whether
transport diagnostics are available. GpuConvolver::backend() reports the
underlying native backend (Metal for this Dawn path), while provider reports
the API implementation (Dawn). Eligible means the path was accepted by
prepare(); it is not a hard real-time scheduling guarantee. The report is
read-only and deliberately exposes no rings, queues, callback hooks, or live
path-switching controls.
Layer 3 — ready-made processors (pulp::gpu_audio)¶
GpuConvolver— FFT overlap-add convolution with a fixed IR; a continuously-fed, zero-latencysignal::PartitionedConvolverCPU fallback that stays latency-aligned so a GPU miss is filled seamlessly.GpuStft— STFT/ISTFT (windowed analyze / synthesize) — the spectral toolkit base.GpuSpectralFreeze— capture + sustain a spectral frame (infinite pad).GpuSpectralMorph— blend between two captured spectra (t ∈ [0,1]).
Authoring pattern¶
- Pick the layer: a ready processor (Layer 3), or your own
GpuAudioNodewhoseprocess_blockcalls Layer-1 primitives. - Always implement a CPU fallback (
signal::*) — it's your degrade path on no-GPU devices and your miss policy. - Keep small/low-latency DSP on the CPU; send only coarse, batched, heavy work to the GPU, and batch across instances/voices to amortize the round-trip.
- Test against a CPU reference (golden / frequency-domain) and assert RT-safety (no alloc/lock/block on the audio thread).
A minimal GPU node, end to end¶
A GpuAudioNode is the whole integration surface. Implement four methods — what
you are, how to set up, the worker-side work, and an RT-safe fallback — then let
the transport schedule it:
#include <pulp/gpu_audio/gpu_audio_node.hpp>
#include <pulp/gpu_audio/gpu_audio_transport.hpp>
#include <pulp/render/gpu_compute.hpp>
using namespace pulp;
// 1. The node: heavy, batched work that runs on the transport's worker thread.
class MyGpuNode : public gpu_audio::GpuAudioNode {
public:
MyGpuNode(uint32_t channels, uint32_t block, uint32_t sr)
: channels_(channels), block_(block), sr_(sr) {}
gpu_audio::GpuAudioNodeDescriptor descriptor() const override {
gpu_audio::GpuAudioNodeDescriptor d;
d.name = "MyGpuNode";
d.input_channels = channels_;
d.output_channels = channels_;
d.block_size = block_;
d.sample_rate = sr_;
d.latency_blocks = 1; // fixed PDC — you MUST report this (see below)
// Miss policy. Both fields default to failing CLOSED (Silence, no
// fallback): a missed block comes out silent. That is the right default —
// a dropout is obvious and bounded.
//
// Opt into CpuFallback only if you actually have a correct fallback. It
// is not enough to implement process_cpu_fallback(): a stateful fallback
// (a convolver, a filter) must ALSO be fed every block via prime_fallback()
// so its history is current and latency-aligned the moment it takes over.
// An unprimed fallback substitutes a stale block with a torn history.
//
// Do NOT use MissPolicy::PassthroughDry. The transport's output is delayed
// by latency_blocks, so substituting the dry sample for the CURRENT time
// jumps the stream forward by the whole latency for one block and back
// again — a timeline break, not an honest dropout.
d.miss_policy = gpu_audio::MissPolicy::CpuFallback;
d.supports_cpu_fallback = true;
return d;
}
// Non-RT setup: create the device + upload static data here, never in process_block.
bool prepare() override {
gpu_ = render::GpuCompute::create();
if (!gpu_ || !gpu_->initialize_standalone()) { gpu_.reset(); return false; }
// ... prepare Layer-1 plans (FFTs, convolution, matmul) on gpu_ ...
return true;
}
// Worker thread: may touch the GPU and block on readback. NEVER the audio thread.
void process_block(const audio::BufferView<const float>& in,
audio::BufferView<float>& out, uint32_t n) override {
// ... call gpu_->fft_forward / convolve / matmul, write `out` ...
}
// RT-safe: called on EVERY block (hit or miss), before the output-ring read,
// so a stateful fallback stays continuously fed and its substitute is a
// correct, latency-aligned continuation of the stream rather than a stale
// block with gaps in its history. Required whenever your fallback has state.
void prime_fallback(const audio::BufferView<const float>& in,
uint32_t n) noexcept override {
// ... advance your signal::* CPU path by one block and stage its output ...
}
// RT-safe: runs on the audio thread for a missed block. No alloc/lock/block.
void process_cpu_fallback(const audio::BufferView<const float>& in,
audio::BufferView<float>& out, uint32_t n) noexcept override {
// ... emit the substitute prime_fallback() already computed for this slot ...
}
private:
uint32_t channels_, block_, sr_;
std::unique_ptr<render::GpuCompute> gpu_;
};
// 2. Wire it into your Processor. prepare() once; process() per block, RT-safe.
class MyPlugin : public format::Processor {
MyGpuNode node_{2, 256, 48000};
gpu_audio::GpuAudioTransport transport_;
void prepare(const format::PrepareContext&) override {
node_.prepare();
transport_.prepare(&node_, {/*ring_blocks*/ 8, /*run_worker_thread*/ true});
}
// REPORT THE LATENCY. The transport delays your output by
// latency_blocks * block_size, and the host cannot know that unless you say
// so — it will leave your track shifted late against every other track in
// the session, and nothing will error. This override is not optional.
int latency_samples() const override { return transport_.latency_samples(); }
void process(audio::BufferView<float>& out,
const audio::BufferView<const float>& in, /*...*/) override {
transport_.process(in, out, out.num_samples()); // writes input, reads delayed output — never blocks
}
};
Prove that latency_samples() is honest rather than trusting it — the number is
hand-maintained and drifts the moment the block size or latency_blocks changes:
pulp audio render --plugin MyPlugin.clap --out /tmp/o.wav --duration-ms 500 \
--input-signal noise --param <your-bypass-id>=1 --latency-report /tmp/lat.json
It exits nonzero if the reported latency does not match the delay actually in the audio. See the CLI reference.
That is the entire contract: the audio thread only ever calls transport_.process()
(lock-free), the GPU work happens on the worker a fixed number of blocks behind,
and a missed block degrades via your MissPolicy instead of glitching.
Worked examples¶
Both ship as buildable plugins (VST3/AU/CLAP/Standalone) with a CPU default and an opt-in GPU engine, a CPU-reference correctness test, and a benchmark:
examples/super-convolver— multi-room convolution reverb. Shows batching many distinct IRs into one GPU submit (GpuMultiConvolver), the live CPU↔GPU engine switch, and a live status line (backend, blocks, µs/block, real-time headroom). The clearest read for the transport + live-rebuild lifecycle.examples/spectral-lab— N-layer spectral freeze/morph "cloud". Shows authoring your ownGpuAudioNode(gpu_spectral_cloud_node.hpp) over a GPU-resident spectral stack, and reusing the same status-line + headroom telemetry. The clearest read for writing a node from scratch.
Building against an installed Pulp SDK¶
Everything above works whether Pulp is a source submodule or an installed SDK.
A standalone plugin repo that consumes Pulp through find_package(Pulp) (rather
than add_subdirectory) gets the GPU audio runtime from the SDK's exported
targets:
| You link | You get |
|---|---|
pulp::gpu-audio |
Layer 2 transport + Layer 3 ready-made processors (GpuAudioTransport, GpuConvolver, GpuStft, GpuSpectralFreeze/Morph) and their headers |
pulp::render |
Layer 1 primitives (GpuCompute) — present only in a GPU-enabled SDK build |
pulp::gpu-audio is always exported: it's GPU-agnostic and carries its own CPU
fallback, so a non-GPU SDK still exposes the transport (it just runs the
signal::* reference path). pulp::render exists only when the SDK was built
-DPULP_ENABLE_GPU=ON, so gate any GPU-primitive code on it:
find_package(Pulp CONFIG REQUIRED)
pulp_add_plugin(MyPlugin
FORMATS VST3 AU CLAP Standalone
SOURCES my_plugin.cpp
PLUGIN_NAME "MyPlugin"
# ...
)
target_link_libraries(MyPlugin_Core PUBLIC pulp::gpu-audio)
if(TARGET pulp::render)
target_link_libraries(MyPlugin_Core PUBLIC pulp::render) # Layer-1 GPU primitives
endif()
find_package(Pulp) resolves the transport's private dependencies for you
(Threads, and ICU on Linux) — a static library propagates its private link
deps into its exported interface, so PulpConfig re-finds them; you don't list
them yourself.
Ship a self-contained GPU bundle¶
A GPU plugin loads libwgpu_native.dylib at runtime. From a source build the
WebGPU dependency copies that dylib next to each format binary automatically; an
installed-SDK consumer has no such upstream copy step, so a find_package(Pulp)
plugin must place the dylib itself and bake an @loader_path rpath — otherwise
the bundle passes every check on the build machine and fails to load
("Library not loaded: @rpath/libwgpu_native.dylib") on any other Mac. The SDK
exports two helpers for exactly this. find_package(Pulp) already defines them —
no extra include() — so call them directly:
foreach(_fmt MyPlugin_CLAP MyPlugin_VST3 MyPlugin_AU MyPlugin_Standalone)
if(TARGET ${_fmt})
pulp_make_bundle_relocatable(${_fmt}) # copy dylib + bake @loader_path
pulp_validate_bundle_relocatable(${_fmt}) # POST_BUILD guard: fail the build if not
endif()
endforeach()
pulp_validate_bundle_relocatable runs at build time and fails hard if a bundle
still references a build-machine-absolute path, so the footgun is caught in CI
instead of on a customer's machine. Apply it to single-bundle formats
(Standalone/CLAP/VST3/AU); an AUv3 .appex is intentionally not relocatable in
isolation — its dylib and framework live in the container app's
Contents/Frameworks and resolve by @rpath only inside that container, so the
container app (not the standalone appex) is the relocatability unit.
GPU work runs outside the audio callback, with CPU fallback when its result is unavailable. This bounds the callback's interaction with the GPU; it does not provide a hard scheduling guarantee from the operating system or driver.
For the experimental shared-memory route, the paced convolution probe records callback timing, missed deliveries, and numerical correctness through the public transport and its own worker. Use the tracing guide for per-block lifecycle analysis.