Vulkan / SECONDARY BACKEND

Resources, layouts & synchronization

Allocate memory deliberately, transfer pixels, express stage/access dependencies, and retire work with the right host and queue completion facts.

Images and buffers need bound memory#

Create a VkBuffer or VkImage with explicit usages, query its memory requirements, choose a compatible memory type, allocate or suballocate storage, and bind it. Memory-type bits restrict which types are legal; desired properties guide the choice among them. Staging buffers need host-visible memory. Device-local image storage is normally useful for repeatedly sampled textures. Host-coherent memory avoids explicit cache flushes for ordinary writes, while noncoherent memory needs aligned flush ranges. Host visibility is not permission to overwrite data still read by the GPU.

Transfer an image#

Decode into validated CPU pixels, populate staging memory, and flush writes when required. Create the image with transfer-destination and sampled usages. Transition from its initial layout to TRANSFER_DST_OPTIMAL, record vkCmdCopyBufferToImage, then establish shader-read access and SHADER_READ_ONLY_OPTIMAL. VkBufferImageCopy describes offsets, subresource layers, and extent. A zero bufferRowLength/bufferImageHeight means tightly packed according to the copy rules. These fields use texel units, not DX12 byte RowPitch. Compressed formats have block constraints. Query and honor actual memory/copy requirements instead of porting a DX12 alignment formula unchanged.

INTERACTIVE ILLUSTRATIONSIMULATION / BROWSER

Follow a texture to the GPU This illustration uses canvas. The explanation below describes the same process.

Change a control to inspect the result. Values describe the simulation, not native engine benchmarks. Open full lab ↗

The lab displays DX12 row pitch by design. The shared lifecycle also applies to Vulkan, but its staging description and layout transition code are backend-specific.

Synchronization2 expresses the dependency#

cpp · REFERENCE EXCERPT
VkImageMemoryBarrier2 barrier{VK_STRUCTURE_TYPE_IMAGE_MEMORY_BARRIER_2};
barrier.srcStageMask = VK_PIPELINE_STAGE_2_TRANSFER_BIT;
barrier.srcAccessMask = VK_ACCESS_2_TRANSFER_WRITE_BIT;
barrier.dstStageMask = VK_PIPELINE_STAGE_2_FRAGMENT_SHADER_BIT;
barrier.dstAccessMask = VK_ACCESS_2_SHADER_SAMPLED_READ_BIT;
barrier.oldLayout = VK_IMAGE_LAYOUT_TRANSFER_DST_OPTIMAL;
barrier.newLayout = VK_IMAGE_LAYOUT_SHADER_READ_ONLY_OPTIMAL;
barrier.srcQueueFamilyIndex = VK_QUEUE_FAMILY_IGNORED;
barrier.dstQueueFamilyIndex = VK_QUEUE_FAMILY_IGNORED;
barrier.image = image;
barrier.subresourceRange = {VK_IMAGE_ASPECT_COLOR_BIT, 0, 1, 0, 1};
VkDependencyInfo dependency{VK_STRUCTURE_TYPE_DEPENDENCY_INFO};
dependency.imageMemoryBarrierCount = 1;
dependency.pImageMemoryBarriers = &barrier;
vkCmdPipelineBarrier2(commandBuffer, &dependency);

This excerpt covers one mip/layer copied and sampled on the same queue family, using a device with synchronization2 enabled. Compute sampling requires a matching destination stage, and a mip chain requires the relevant range. Queue-family transfers require paired release/acquire semantics rather than IGNORED indices.

Fences, binary semaphores, and timelines#

A fence communicates submission completion to the host. Binary semaphores synchronize queue/WSI operations and must obey their signal/wait lifecycle. Timeline semaphores use increasing values, can coordinate queues, and can be waited on by the host. The reference backend uses timelines for upload and retirement bookkeeping; swap-chain acquisition/presentation uses supported binary WSI semantics. Submitting commands in order is not a general substitute for memory visibility dependencies. Stage/access masks establish which producers and consumers participate. Overly broad masks can serialize unrelated work; missing masks can create data races that only appear on another device.

Cross-queue ownership#

If upload uses a separate family and an exclusive resource, a release transfers ownership from the upload family and an acquire receives it on the consumer family. A semaphore carries the execution dependency between submissions. If families are the same, ownership transfer is unnecessary, but access ordering remains necessary. Keep staging storage until the upload completion value passes. Keep the image and its descriptors until all readers finish. A material uses a fallback while upload is pending, then resolves the ready version through the same AssetStore publication boundary as DX12.

Memory pools and troubleshooting#

Suballocation reduces many small native allocations but must satisfy alignment, type compatibility, and lifetime. Aliasing and transient image reuse follow physical requirements as well as graph lifetimes. Descriptor sets and image views also outlive their submitted references. Corruption after a host write can mean missing noncoherent flushing or early staging reuse. Layout validation errors mean the tracked usage/layout plan diverged from recording. Hangs can come from waiting on a value never signaled after an error. Read Khronos synchronization and Synchronization2.

Search titles and full article text.