Instruction latency analysis usually focuses on performance optimizationβmaking code run as fast as possible. The Assembly Hall of Shame takes the opposite approach: searching for the absolute floor of single-instruction performance.
π Current Champions πStrategy: Use fxrstor64 to load 512-byte FPU/MMX/XMM state from a high-latency MMIO region in the PCIe fabric, then starve the fabric while the load is in flight β a fleet of hammer cores pounds a different high-latency MMIO register with tight 4-byte reads, saturating the PCIe root complex and endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue behind all that contending traffic.
Contender: AMD Ryzen 7 5800H
; CPU 0 β timed instruction movl $0xfcc68830, %rsi fxrstor64 %rsi ; CPUs 1..N β hammer loop against a different high-latency location movl 0xfcc68858, %eaxπ Score: 198,002,498,236 cycles
π Time: 62 seconds
A spec-violating unaligned ymm0 load that forced non-posted dword transactions from stalled GPU registers was used to break the fundamental design of System Management Mode in smiiiiiiiiiiiiiiii.
vmovdqu 0xfcc003b1, %ymm0 Instructions may use whatever setup is necessary, but only a single instruction is eligible to be scored. Trapped/emulated/virtualized instructions may only time the trap, not the handler. Instructions must not be interruptible. rep movs, pause, etc. are disqualified. Times are normalized based on the CPU base clock frequency. All platforms must be in their factory stock configurations - no hardware modifications.Strategy: nop does nothing. It opens the leaderboard accordingly.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
Score: 1 cycles
Time: 0 nanoseconds
Strategy: Regular nop was too short, but how do we make nothing take longer? Try a lonnnnnng nop.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
data16 data16 data16 data16 data16 data16 data16 nopl 0x00000000(%%eax,%%eax,1)Score: 20 cycles
Time: 7 nanoseconds
Strategy: Just a reference instruction to get our bearings.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
Score: 49 cycles
Time: 18 nanoseconds
Strategy: Use 128-bit dividend (rdx:rax=2:0) with small divisor to push the quotient above the ceiling imposed by sign-extension, driving the longest path through the divider microcode.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
xorq %rax, %rax ; rax = 0 (low 64 bits of dividend) movq $2, %rdx ; rdx = 2 (high 64 bits: full dividend = 2^65) movq $5, %rbx ; divisor β quotient = 2^65/5 β 7.4Γ10^18 idivq %rbxScore: 77 cycles
Time: 28 nanoseconds
Strategy: Use maximum nesting depth (31) to force 30 display-pointer loads and pushes through the microcode display-walk path.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
enter $0, $31 ; 0 bytes allocated, nesting depth 31 (maximum)Score: 112 cycles
Time: 41 nanoseconds
Strategy: Try a small denormal to trigger an FP microcode assist.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movabsq $0x0000000000000001, %rax movq %rax, -8(%rsp) fldl -8(%rsp)Score: 133 cycles
Time: 49 nanoseconds
Strategy: Just ensure the cache line is dirty.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
clflush (%rax) ; rax -> dirty cache line resident in L3Score: 165 cycles
Time: 60 nanoseconds
Strategy: Use exponent 0x7ff to reach 'special value' processing in microcode; positive/negative, NaN/inf doesn't seem to make a difference, go with QNaN.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movabsq $0x7fffffffffffffff, %rax movq %rax, -8(%rsp) fldl -8(%rsp) fsinScore: 257 cycles
Time: 94 nanoseconds
Strategy: Saturate all write-combining line-fill buffers with movnti stores to distinct cache lines, forcing mfence to drain the full LFB write path to the uncore before retiring.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movnti %r9, 0*64(%rdi) ; Γ16 distinct cache lines β saturate the write-combining LFBs ; β¦ movnti %r9, 15*64(%rdi) mfence ; must drain all pending LFB writes before retiringScore: 326 cycles
Time: 120 nanoseconds
Strategy: Nothing for now, just check how long it takes to invalidate the TLB.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
Score: 352 cycles
Time: 110 nanoseconds
Strategy: Hit x87 FP microcode assist path by using denormal source operand.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
fldl subnorm ; 1e-310: value < DBL_MIN, biased exponent = 0 faddl subnorm ; source is subnormal β FP microcode assistScore: 677 cycles
Time: 249 nanoseconds
Strategy: Align lock-prefixed operand to straddle cache-line boundary, forcing CPU to assert the external bus lock rather than using the fast MESI cache-coherence path.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
; split_ptr % 64 == 63 β dword spans bytes 63 (line N) and 64β66 (line N+1) lock xaddl %r9d, (%rdi)Score: 865 cycles
Time: 319 nanoseconds
Strategy: Use subnormal divisor, hardware hands control to microcode assist, assist normalizes operand, performs the division, then restores architectural state.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
movabsq $0x3ff0000000000000, %rax ; 1.0 (normal dividend) movq %rax, -8(%rsp) fldl -8(%rsp) ; ST(0) = 1.0 movabsq $0x0000002000000000, %rax ; 6.79e-313 (subnormal divisor) movq %rax, -8(%rsp) fdivl -8(%rsp) ; ST(0) = 1.0 / subnormal β FP assistScore: 883 cycles
Time: 325 nanoseconds
Strategy: Use rakefield to find the highest latency CPUID leaves.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
Score: 1248 cycles
Time: 460 nanoseconds
Strategy: Execute in a tight loop to deplete the hardware entropy pool faster than it can be refilled, forcing subsequent calls to stall while the entropy source recovers.
Contender: Intel(R) Core(TM) i7-8559U CPU @ 2.70GHz
Score: 5,579 cycles
Time: 2.057 microseconds
Strategy: Use project:nightshyft to identify high latency MSRs. MCG_CTL on Zen look like a winner: may be a microcode quiesce and synchronize on MCA error banks across hardware units, some potentially off-die, requiring fabric-level communication rather than a simple local register write.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
movl $0x17b, %ecx ; MCG_CTL wrmsrScore: 34,304 cycles
Time: 10.742 microseconds
Strategy: Target an I/O port that straddles a NIC device register boundary, triggering the device to quiesce its TX DMA engine on each write.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
mov $0xf019, %dx outl %eax, %dxScore: 49,857 cycles
Time: 15.580 microseconds
Strategy: Use project:nightshyft to identify high latency model-specific-registers: VIA uses an undocumented register at 0x133 that gives wildly high response time. No idea what it does.
Contender: VIA Eden Processor 800MHz
movl $0x133, %ecx ; undocumented MSR rdmsrScore: 161,602 cycles
Time: 202.004 microseconds
Strategy: Fully load L1/L2/L3 caches with dirty lines to force DRAM writeback of entire hierarchy.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
Score: 1,616,480 cycles
Time: 506.165 microseconds
Strategy: Target I/O port mapped to an ACPI PM block where an unaligned 4-byte read decodes into multiple non-posted loads from wherever this port goes.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
mov $0x0413, %dx inl %dx, %eaxScore: 12,524,415 cycles
Time: 3.921769 milliseconds
Strategy: Use mmiotic to identify high-latency deadspace in PCIe fabric, hit unkown GPU register.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
Score: 443,937,696 cycles
Time: 139.010268 milliseconds
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 8-byte MMIO read to get two dword register accesses, which isn't technically allowed but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
Score: 887,716,864 cycles
Time: 277.971228 milliseconds
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 16-byte MMIO read to get four dword register accesses, which isn't technically allowed but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
vmovdqu 0xfcc003b0, %xmm0Score: 1,774,555,776 cycles
Time: 555.664133 milliseconds
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 32-byte MMIO read to get eight dword register accesses, which still isn't technically allowed but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
vmovdqu 0xfcc003b0, %ymm0Score: 3,549,079,296 cycles
Time: 1.111345034 s
Strategy: Search MMIO space for slowest registers in PCIe fabric, hit unknown GPU register, use 32-byte unaligned MMIO read to get nine dword register accesses, which is even less allowed than the aligned version, but works anyway.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
vmovdqu 0xfcc003b1, %ymm0Score: 4,453,212,256 cycles
Time: 1.394428818 seconds
Strategy: Use mmiotic to identify high-latency deadspace in PCIe fabric, isolate region near 0's and offset state to avoid MXCSR corruption (possibly VGA buffer?), use fxrstor64 to load 512-byte FPU/MMX/XMM state from MMIO, forcing CPU to process 512 bytes of I/O transactions through slowest available memory aperture.
Contender: AMD Ryzen 7 5800H
movl $0xfcc68830, %rsi fxrstor64 %rsiScore: 74,584,168,512 cycles
Time: 23.354502677 seconds
Strategy: Extend fxrstor64 (baseline) by starving the fabric while the load is in flight β a fleet of hammer cores pounds a different high-latency MMIO register with tight 4-byte reads, saturating the PCIe root complex and endpoint with non-posted transactions, so CPU 0's 512-byte fxrstor64 must queue behind all that contending traffic.
Contender: AMD Ryzen 7 5800H with Radeon Graphics (Trigkey S5)
; CPU 0 β timed instruction movl $0xfcc68830, %rsi fxrstor64 %rsi ; CPUs 1..N β hammer loop against a different high-latency location movl 0xfcc68858, %eaxπ Score: 198,002,498,236 cycles
π Time: 62 seconds
??. xrstor64 (AMX, MMIO)Strategy: Leverage extended AVX state in Sapphire Rapids with MMIO approach from fxrstor64: xsave state area is 8KB vs 512 bytes, 16x size -> 1,000,000,000,000 cycles
Contender: TODO
; XCR0 must enable AMX components (bits 17-18); state area ~8KB xrstor64 (%rsi) ; rsi -> MMIO region, same technique as fxrstor64The assembly hall-of-shame is a research effort from Christopher Domas (@xoreaxeaxeax).