The CPU does not need to name every visible draw#
In a conventional renderer the CPU selects objects and records their draw calls. In a GPU-driven path, compute work can test instance bounds, choose detail, compact visible records, and produce indirect arguments. The graphics stage consumes that generated work with fewer CPU submissions. QubicEngine’s reference design stores instance, mesh, and material tables in persistent GPU buffers. Per-frame extraction updates changed records. A visibility pass writes compacted instance indices and indirect argument ranges; rendering consumes them after the required dependency.
Spend work on what you can see This illustration uses canvas. The explanation below describes the same process.
The laboratory uses CPU point tests and distance-based LOD as a readable model. The native reference path uses conservative bounds and projected-error choices. Its triangle counts are illustrative costs, not measurements of your GPU.
Culling and indirect commands#
Frustum culling tests bounds against camera planes. A depth hierarchy can reject occluded bounds with temporal conservatism. An indirect command contains exactly the fields expected by the backend’s command signature or draw-indirect contract. Buffer capacity and counter reset are correctness conditions. DX12 ExecuteIndirect uses a command signature and argument/count buffers. Vulkan indirect commands use the appropriate draw-indirect operations and feature limits, with separate descriptor/instance-data contracts. The backend maps the engine’s draw intent rather than assuming binary-identical command records.
Reference visibility output
struct VisibleInstance {
uint32_t instanceIndex;
uint32_t selectedLod;
};
struct DrawIndexedArguments {
uint32_t indexCount;
uint32_t instanceCount;
uint32_t firstIndex;
int32_t baseVertex;
uint32_t firstInstance;
};These are engine records. Each backend validates its required native argument layout and alignment. Compute writes require a barrier to indirect-argument consumption and shader reads. An overflow policy must bound output writes and report the event; silently overrunning a visible-list buffer is not a performance shortcut.
Mesh processing and clusters#
Offline mesh processing improves index/vertex locality, builds LODs, and partitions geometry into bounded clusters. Clusters store conservative bounds and error metadata. Small clusters improve culling granularity but increase metadata and scheduling overhead. Material boundaries and deforming geometry constrain what can be grouped. A virtualized geometry system can select a hierarchy of clusters and stream needed pages. It needs residency tracking, conservative parent fallback, page tables, and a feedback budget. Missing fine pages must still produce a valid coarse representation, not holes.
Virtualization is an asset and scheduling system#
The reference design cooks a cluster hierarchy and page pack, retains coarser resident levels, and requests finer pages according to projected error and budget. AssetStore owns source identity and cooked versions; the renderer owns residency descriptors and frame-safe publication. Streaming priority includes visibility and request age to avoid starvation. Nanite is an Epic Games technology. General hierarchical cluster selection and virtualized geometry can be explained independently, but this documentation does not claim Epic’s proprietary implementation or equivalent behavior.
Tradeoffs and failure modes#
GPU-driven work reduces CPU submission overhead when that overhead matters. It adds compute passes, data structures, barriers, and debugging complexity. Tiny scenes may become slower. Indirect arguments can make a capture less intuitive unless diagnostics preserve instance/material identities. Objects popping incorrectly may have poor LOD error or hysteresis. Occlusion flicker can come from nonconservative bounds or unstable history. Draw corruption can come from argument layouts, counters, or missing write-to-indirect barriers. Start with a readable CPU reference result and compare visibility sets before optimizing the path. Continue with large-world streaming and render-graph dependencies.