Back to the roadmapGlobal memory Access being simulated
Memory coalescing simulator
Adjust stride, start offset, and access width. Watch the 32 addresses, the sector count, and the bandwidth efficiency.
Access pattern
Preset
1
0
Non-zero offset can misalign sectors
Bytes per thread
Sector transactions
4
Ideal 4
Bandwidth efficiency
100.0%
Bytes moved 128 B
Bytes used 128 B
Wasted 0 B
in[i]: ideal case. 32 floats fill exactly 4 sectors.
Device-memory sector view (32 B per cell)
Fully usedPartially usedUntouched
Addresses of the 32 threads in the warp
t0
+0
s0
t1
+4
s0
t2
+8
s0
t3
+12
s0
t4
+16
s0
t5
+20
s0
t6
+24
s0
t7
+28
s0
t8
+32
s1
t9
+36
s1
t10
+40
s1
t11
+44
s1
t12
+48
s1
t13
+52
s1
t14
+56
s1
t15
+60
s1
t16
+64
s2
t17
+68
s2
t18
+72
s2
t19
+76
s2
t20
+80
s2
t21
+84
s2
t22
+88
s2
t23
+92
s2
t24
+96
s3
t25
+100
s3
t26
+104
s3
t27
+108
s3
t28
+112
s3
t29
+116
s3
t30
+120
s3
t31
+124
s3
CUDA C++
int i = blockIdx.x * blockDim.x + threadIdx.x;float v = in[i]; // one warp (32 threads) produces 4 32B sector transactions// moved 128 B, used 128 B, efficiency 100.0%Related lesson: Memory coalescing: the same kernel, a 5x gap