Frame time is the useful unit#
At 60 frames per second, a frame budget is about 16.67 milliseconds. At 144, it is about 6.94 milliseconds. Saving two milliseconds has different FPS effects at different starting points, so compare frame times instead of averaging FPS values and guessing where work went. CPU and GPU stages overlap. Total latency depends on dependencies, buffering, and presentation. A fast CPU update cannot compensate for a GPU pass that exceeds the budget. An idle GPU may mean the CPU submits too late, rather than a graphics feature being too expensive.
Start with a question#
Capture a repeatable scene, a warm-up period, and relevant camera motion. Record CPU update, extraction, recording, queue waits, and GPU pass timestamps. Inspect distributions and spikes, not only one average. Separate cold asset loading and pipeline compilation from steady-state performance. QubicEngine’s diagnostics tag work by system and pass. GPU timing queries are read back after completion; waiting immediately for the result changes the workload being measured. Native captures can be investigated with tools such as PIX on Windows. Browser laboratory durations are explicitly simulations.
CPU and GPU, in parallel This illustration uses canvas. The explanation below describes the same process.
Reduce the work you proved expensive#
| Bottleneck | Useful technique | What it costs |
|---|---|---|
| Draw submission | Batching, instancing, indirect commands | Sorting, buffer generation, more complex culling |
| Vertex work | Frustum culling, LOD, mesh reduction | Bounds tests, asset variants, transition artifacts |
| Pixel work | Early depth, reduced overdraw, simpler shaders | Extra passes or lower visual fidelity |
| Bandwidth | Compression, smaller formats, locality | Decode support, precision loss, artifacts |
| CPU iteration | Dense arrays, job decomposition | Data movement and scheduling overhead |
| Loading stalls | Async decode, staging rings, caching | Memory budget and lifetime tracking |
| Synchronization | Frame contexts and planned dependencies | More retained resources and possibly more latency |
A technique can make another bottleneck worse. A depth prepass adds geometry and submission work while reducing expensive shading. Instancing reduces command count but can complicate material diversity and visibility. Occlusion queries can stall if read synchronously.
Visibility and level of detail#
Spend work on what you can see This illustration uses canvas. The explanation below describes the same process.
Frustum culling rejects bounds outside the camera volume. The lab uses point centers for clarity; a real mesh needs conservative bounds so partially visible objects are not incorrectly rejected. Occlusion culling asks whether another surface hides those bounds and typically uses a depth hierarchy with temporal conservatism. LOD changes geometry or shading complexity based on projected size and error. Distance is a useful teaching proxy, but a large far object can still need detail. Use hysteresis to prevent rapid switching and a transition strategy that matches asset content. Culling has a CPU or GPU cost of its own.
Implementation: memory belongs to a budget
Frame-owned upload rings reserve aligned spans. Persistent resources use heaps or allocations chosen by memory intent. Temporary render targets can share memory only when their GPU lifetimes do not overlap and placement requirements are compatible. Pooling removes repeated allocation costs but can retain memory long after a scene stops needing it. Count resident GPU bytes, retained staging bytes, active descriptors, and pending retirement separately. A logical asset release does not instantly reduce physical usage if earlier frames still reference it. If a memory graph grows without bound, inspect retirement completion and cache eviction policies before adding a larger pool.
Change one thing, verify one cause#
Disable a costly pass or change its resolution to test a GPU hypothesis. Reduce object count to test submission or geometry hypotheses. Turn off asynchronous work only in a controlled capture to reveal dependency behavior. Keep the scene and settings comparable. Threading a tiny task can cost more than executing it directly. Queueing more work can hide waits while increasing latency and memory. Lowering texture resolution may save bandwidth but will not fix a main-thread collision spike. Record the expected effect, measure it, and retain the change only when the result supports the explanation.
Related reading#
Study render-graph lifetimes, advanced GPU-driven rendering, and native synchronization. QubicEngine’s documentation does not invent native benchmark results or hardware performance claims.