This is part 8 of SoftGPU. “Wave,” “warp,” and “subgroup” mean different things on different vendors. SoftGPU must define a software semantic machine without pretending it is gfx1201 silicon.

SoftGPU vocabulary

TermSoftGPU meaning
Grid / workgroup / local idSame indexing formulas as the functional IR in part 7
WaveContiguous block of wave_size flat local ids inside a workgroup
LaneIndex of a workitem inside its SoftGPU wave (flat_local % wave_size)
Group memoryPer-workgroup CPU arena; not R9700 LDS capacity evidence
BarrierSoftGPU generation sync between barrier-separated program segments
DivergenceStructured if / while with per-lane masks; reconverge after the construct

Partial waves (workgroup size not divisible by wave_size) are first-class: the last wave simply has fewer live lanes. SoftGPU’s wave_size is a software parameter, 32 or 64, not R9700 wavefront evidence.

The barrier model

SoftGPU splits a program body on top-level barrier ops. For each segment, every wave runs the segment to completion (lockstep within the wave for divergent if). Then the next segment begins. That is a deterministic wave_barrier schedule.

Divergent barriers (a barrier inside if or while) are unsupported and fail validation. They are not silently linearized.

Atomics, honestly

atomic_add on global or group returns the previous i32 value. SoftGPU applies a functional sequential atomic on the host arena. Named scope and order fields are recorded in the IR for honesty. They do not claim AMDGPU memory-model semantics.

What this is not

  • Not gfx1201 wavefront scheduling
  • Not a proof of HIP kernel success via AQL
  • Not a sanitizer completeness claim — that is the next part, and even there the subset is declared

Next: building GPU sanitizers on SoftGPU.

Canonical: thanos.github.io.