<<<CUDA C++ Coursegrid · block · warp · lane
Back to the roadmap
Shared memory

Bank conflict simulator

Pick an access pattern and padding, then see how many threads land on each bank, the worst conflict degree, and how many cycles it takes.

Shared-memory access pattern

Pattern
32
0

The classic +1 makes the row width odd, coprime with 32

Max conflict degree
32×
Access serializes over 32 cycles
Banks used
1 / 32
32-way conflict. This shared-memory access takes 32 cycles instead of 1, so effective bandwidth drops to 3%. Try setting padding to 1 so the row width becomes odd.

Where the 32 banks land

Each column is a bank. Each square on the bar is a thread that hits it. Taller bar means more serialized cycles.

0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31

Per-thread mapping

t0
w0
bank 0
t1
w32
bank 0
t2
w64
bank 0
t3
w96
bank 0
t4
w128
bank 0
t5
w160
bank 0
t6
w192
bank 0
t7
w224
bank 0
t8
w256
bank 0
t9
w288
bank 0
t10
w320
bank 0
t11
w352
bank 0
t12
w384
bank 0
t13
w416
bank 0
t14
w448
bank 0
t15
w480
bank 0
t16
w512
bank 0
t17
w544
bank 0
t18
w576
bank 0
t19
w608
bank 0
t20
w640
bank 0
t21
w672
bank 0
t22
w704
bank 0
t23
w736
bank 0
t24
w768
bank 0
t25
w800
bank 0
t26
w832
bank 0
t27
w864
bank 0
t28
w896
bank 0
t29
w928
bank 0
t30
w960
bank 0
t31
w992
bank 0
Access being simulated
1__shared__ float tile[32][32];2 3float v = tile[threadIdx.x][0];        // column access4// word = tid * 32   →   bank = (tid * 32) % 325// Row width 32 shares a factor with 32: still conflicts
This model uses 4-byte banks and 32 banks, matching float shared-memory accesses. 64-bit double accesses are split by hardware on modern GPUs and the rules differ slightly. In practice you still fix them with padding or swizzle.