Labs / QUBICENGINE HANDBOOK

CPU/GPU scheduling

Change producer and consumer durations, inspect frame-slot waits, and see why extra buffering has a cost.

Inspect the relationship#

Change producer and consumer durations, inspect frame-slot waits, and see why extra buffering has a cost.

INTERACTIVE ILLUSTRATIONSIMULATION / BROWSER

CPU and GPU, in parallel This illustration uses canvas. The explanation below describes the same process.

Change a control to inspect the result. Values describe the simulation, not native engine benchmarks. Open full lab ↗

Try these experiments#

  1. Set CPU to five milliseconds and GPU to ten. Compare one and two frames in flight.
  2. Make the CPU slower than the GPU. Observe where the GPU waits for a new submission.
  3. Increase frames in flight and inspect the orange CPU-slot waits. Buffering changes overlap without making each GPU task faster.

What the illustration calculates#

The simulation has one CPU producer and one graphics queue. A CPU frame can start only after the CPU finishes its preceding frame and the chosen frame slot’s prior GPU use completes. A GPU frame starts after both its CPU submission and the preceding queue work complete. Six fixed-duration frames are shown. The lines and counters are modeled durations, not browser timing or native engine benchmarks. Presentation scheduling, variable task costs, copy/compute queues, and display latency are intentionally outside this simplified view.

Connect it to frame contexts#

Each FrameContext owns an allocator and transient storage. Its completion value protects reuse. More frames in flight need more retained storage and can put additional old input before presentation. A production renderer chooses a latency/memory/throughput tradeoff based on measured behavior.

The ownership rule
text · REFERENCE EXCERPT
slot can reset when completedFence >= slot.lastSubmission
resource can retire when every last-use completion point passes

DX12 expresses queue completion with ID3D12Fence values. Vulkan can use submission fences or timeline values for host completion, while WSI binary semaphores have additional image/presentation lifecycle rules. One abstraction does not erase those differences.

What to inspect when it is wrong#

A long allocator wait is a symptom of work that has not completed. A long CPU task can leave the GPU idle. Adding more buffers may hide the wait but increase latency or memory. An arbitrary sleep may conceal a race without proving safety. Read native DX12 synchronization, Vulkan resource synchronization, and frame lifecycle.

Search titles and full article text.