Pick an access pattern and padding, then see how many threads land on each bank, the worst conflict degree, and how many cycles it takes.
Shared-memory access pattern
Pattern
32
0
The classic +1 makes the row width odd, coprime with 32
Max conflict degree
32×
Access serializes over 32 cycles
Banks used
1 / 32
32-way conflict. This shared-memory access takes 32 cycles instead of 1, so effective bandwidth drops to 3%. Try setting padding to 1 so the row width becomes odd.
Where the 32 banks land
Each column is a bank. Each square on the bar is a thread that hits it. Taller bar means more serialized cycles.
0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
Per-thread mapping
t0
w0
bank 0
t1
w32
bank 0
t2
w64
bank 0
t3
w96
bank 0
t4
w128
bank 0
t5
w160
bank 0
t6
w192
bank 0
t7
w224
bank 0
t8
w256
bank 0
t9
w288
bank 0
t10
w320
bank 0
t11
w352
bank 0
t12
w384
bank 0
t13
w416
bank 0
t14
w448
bank 0
t15
w480
bank 0
t16
w512
bank 0
t17
w544
bank 0
t18
w576
bank 0
t19
w608
bank 0
t20
w640
bank 0
t21
w672
bank 0
t22
w704
bank 0
t23
w736
bank 0
t24
w768
bank 0
t25
w800
bank 0
t26
w832
bank 0
t27
w864
bank 0
t28
w896
bank 0
t29
w928
bank 0
t30
w960
bank 0
t31
w992
bank 0
Access being simulated
CUDA C++
1__shared__floattile[32][32];23floatv=tile[threadIdx.x][0];// column access4// word = tid * 32 → bank = (tid * 32) % 325// Row width 32 shares a factor with 32: still conflicts
This model uses 4-byte banks and 32 banks, matching float shared-memory accesses. 64-bit double accesses are split by hardware on modern GPUs and the rules differ slightly. In practice you still fix them with padding or swizzle.