MAPLE's hardware-software co-design allows programs to perform long-latency memory accesses asynchronously from the core, avoiding pipeline stalls, and enabling greater memory parallelism (MLP).
-
Updated
Feb 22, 2024 - C
MAPLE's hardware-software co-design allows programs to perform long-latency memory accesses asynchronously from the core, avoiding pipeline stalls, and enabling greater memory parallelism (MLP).
Low-level command-line tool for measuring CPU and Metal GPU memory bandwidth, synthetic LLM decode and prefill memory traffic, cache and main-memory latency, access-pattern performance, TLB behavior, and two-thread cache-line handoff protocol latency on Apple Silicon Macs.
Modern Memory Bandwidth and Latency Benchmarks
Experimental study and analysis on the effect of using a wide range of different supply voltage values on the reliability, latency, and retention characteristics of DDR3L DRAM SO-DIMMs
This repository contains the code to benchmark CPU cache miss latency and branch misprediction penalty
General-purpose compile-time Expression Templates library for C++
A CPU benchmark suite that shows its work. Six workloads, calibrated measurement windows, barrier-synchronised threads, robust statistics, full machine provenance, and Ed25519-signed results.
Analysis of cache friendly implementations as part of my research
This repository contains the code to test GPU memory
To associate your repository with the memory-latency topic, visit your repo's landing page and select "manage topics."