<<<CUDA C++ Coursegrid · block · warp · lane
Back to the roadmap
Global memory

Memory coalescing simulator

Adjust stride, start offset, and access width. Watch the 32 addresses, the sector count, and the bandwidth efficiency.

Access pattern

Preset
1
0

Non-zero offset can misalign sectors

Bytes per thread
Sector transactions
4
Ideal 4
Bandwidth efficiency
100.0%
Bytes moved 128 B
Bytes used 128 B
Wasted 0 B
in[i]: ideal case. 32 floats fill exactly 4 sectors.

Device-memory sector view (32 B per cell)

Fully usedPartially usedUntouched

Addresses of the 32 threads in the warp

t0
+0
s0
t1
+4
s0
t2
+8
s0
t3
+12
s0
t4
+16
s0
t5
+20
s0
t6
+24
s0
t7
+28
s0
t8
+32
s1
t9
+36
s1
t10
+40
s1
t11
+44
s1
t12
+48
s1
t13
+52
s1
t14
+56
s1
t15
+60
s1
t16
+64
s2
t17
+68
s2
t18
+72
s2
t19
+76
s2
t20
+80
s2
t21
+84
s2
t22
+88
s2
t23
+92
s2
t24
+96
s3
t25
+100
s3
t26
+104
s3
t27
+108
s3
t28
+112
s3
t29
+116
s3
t30
+120
s3
t31
+124
s3
Access being simulated
int i = blockIdx.x * blockDim.x + threadIdx.x;float v = in[i]; // one warp (32 threads) produces 4 32B sector transactions// moved 128 B, used 128 B, efficiency 100.0%