IJAM: Volume 38, No. 4(2025)

PREFETCHING TECHNIQUES IN MODERN MULTICORE PROCESSORS: PRINCIPLES, TAXONOMY, DESIGN TRADE-OFFS, AND EMERGING DIRECTIONS

 

Purnendu Das1, Bishwa Ranjan Roy2,*

 

1,2Department of Computer Science, Assam University, Silchar

 

Abstract. Prefetching is one of the most effective microarchitectural techniques for tolerating the widening processor–memory latency gap. Its role becomes more complex in multicore processors because each core may generate speculative memory traffic that competes for shared cache capacity, on-chip-network bandwidth, memory-controller queues, and DRAM service. This paper reviews the evolution of prefetching from software hints, next-line and stride mechanisms to history-, delta-, spatial-, temporal-, perceptron-, and reinforcement-learning-based designs. It first explains the organization of modern multicore cache hierarchies and the latency-hiding objective of prefetching. It then develops a unified design taxonomy based on trigger source, context, prediction representation, lookahead, target cache, filtering, and feedback control. Representative techniques—including stream buffers, GHB, SMS, AMPM, VLDP, BOP, SPP, Domino, Bingo, PPF, MLOP, DSPatch, IPCP, Pythia, and Berti—are compared qualitatively. Particular attention is given to multicore interference, bandwidth-aware throttling, prefetch-aware cache and memory management, security leakage, and evaluation methodology. The review concludes that future prefetchers should be multilevel, resource-aware, workload-adaptive, and explicitly secure rather than optimized only for single-core address-prediction accuracy.

 

Download paper from here

 

How to cite this paper?
Source: International Journal of Applied Mathematics
ISSN printed version: 1311-1728
ISSN on-line version: 1314-8060
Year: 2025
Volume: 38
Issue: 4

References

[1] S. Mittal, “A survey of recent prefetching techniques for processor caches,” ACM Computing Surveys, vol. 49, no. 2, Art. 35, 2016, doi: 10.1145/2907071.

[2] W. A. Wulf and S. A. McKee, “Hitting the memory wall: Implications of the obvious,” ACM SIGARCH Computer Architecture News, vol. 23, no. 1, pp. 20-24, 1995.

[3] N. P. Jouppi, “Improving direct-mapped cache performance by the addition of a small fully-associative cache and prefetch buffers,” in Proc. 17th Int. Symp. Computer Architecture (ISCA), 1990, pp. 364-373.

[4] T. C. Mowry, M. S. Lam, and A. Gupta, “Design and evaluation of a compiler algorithm for prefetching,” in Proc. 5th Int. Conf. Architectural Support for Programming Languages and Operating Systems (ASPLOS), 1992, pp. 62-73.

[5] J. W. C. Fu, J. H. Patel, and B. L. Janssens, “Stride directed prefetching in scalar processors,” in Proc. 25th Annu. Int. Symp. Microarchitecture (MICRO), 1992, pp. 102-110.

[6] D. Joseph and D. Grunwald, “Prefetching using Markov predictors,” IEEE Trans. Computers, vol. 48, no. 2, pp. 121-133, 1999.

[7] K. J. Nesbit and J. E. Smith, “Data cache prefetching using a global history buffer,” in Proc. 10th Int. Symp. High-Performance Computer Architecture (HPCA), 2004, pp. 96-105.

[8] S. Somogyi, T. F. Wenisch, A. Ailamaki, B. Falsafi, and A. Moshovos, “Spatial memory streaming,” in Proc. 33rd Int. Symp. Computer Architecture (ISCA), 2006, pp. 252-263.

[9] I. Hur and C. Lin, “Memory prefetching using adaptive stream detection,” in Proc. 39th Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2006, pp. 397-408.

[10] S. Srinath, O. Mutlu, H. Kim, and Y. N. Patt, “Feedback directed prefetching: Improving the performance and bandwidth-efficiency of hardware prefetchers,” in Proc. 13th Int. Symp. High-Performance Computer Architecture (HPCA), 2007, pp. 63-74.

[11] M. Ferdman, T. F. Wenisch, A. Ailamaki, B. Falsafi, and A. Moshovos, “Temporal instruction fetch streaming,” in Proc. 41st Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2008, pp. 1-10.

[12] Y. Ishii, M. Inaba, and K. Hiraki, “Access map pattern matching for data cache prefetch,” in Proc. Int. Conf. Supercomputing (ICS), 2009, pp. 499-500.

[13] S. Somogyi, T. F. Wenisch, A. Ailamaki, and B. Falsafi, “Spatio-temporal memory streaming,” in Proc. 36th Int. Symp. Computer Architecture (ISCA), 2009, pp. 69-80.

[14] E. Ebrahimi, C. J. Lee, O. Mutlu, and Y. N. Patt, “Coordinated control of multiple prefetchers in multi-core systems,” in Proc. 42nd Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2009, pp. 316-326.

[15] A. Jain and C. Lin, “Linearizing irregular memory accesses for improved correlated prefetching,” in Proc. 46th Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2013, pp. 247-259.

[16] S. H. Pugsley et al., “Sandbox prefetching: Safe run-time evaluation of aggressive prefetchers,” in Proc. 20th Int. Symp. High-Performance Computer Architecture (HPCA), 2014, pp. 626-637.

[17] M. Shevgoor, S. Koladiya, R. Balasubramonian, C. Wilkerson, S. H. Pugsley, and Z. Chishti, “Efficiently prefetching complex address patterns,” in Proc. 48th Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2015.

[18] P. Michaud, “Best-offset hardware prefetching,” in Proc. IEEE Int. Symp. High Performance Computer Architecture (HPCA), 2016, pp. 469-480.

[19] J. Kim, S. H. Pugsley, P. V. Gratz, A. L. N. Reddy, C. Wilkerson, and Z. Chishti, “Path confidence based lookahead prefetching,” in Proc. 49th Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2016, pp. 1-12.

[20] M. Bakhshalipour, P. Lotfi-Kamran, and H. Sarbazi-Azad, “Domino temporal data prefetcher,” in Proc. IEEE Int. Symp. High Performance Computer Architecture (HPCA), 2018, pp. 131-142.

[21] S. Kondguli and M. Huang, “Division of labor: A more effective approach to prefetching,” in Proc. 45th Annu. Int. Symp. Computer Architecture (ISCA), 2018, pp. 83-95.

[22] Y. Shin, H. Kim, D. Kim, J. Jeong, and J. Hur, “Unveiling hardware-based data prefetcher, a hidden source of information leakage,” in Proc. ACM SIGSAC Conf. Computer and Communications Security (CCS), 2018.

[23] M. Bakhshalipour, P. Lotfi-Kamran, and H. Sarbazi-Azad, “Bingo spatial data prefetcher,” in Proc. IEEE Int. Symp. High Performance Computer Architecture (HPCA), 2019, pp. 399-411.

[24] E. Bhatia, G. Chacon, S. H. Pugsley, E. Teran, P. V. Gratz, and D. A. Jiménez, “Perceptron-based prefetch filtering,” in Proc. 46th Annu. Int. Symp. Computer Architecture (ISCA), 2019, doi: 10.1145/3307650.3322207.

[25] M. Shakerinava, M. Bakhshalipour, P. Lotfi-Kamran, and H. Sarbazi-Azad, “Multi-lookahead offset prefetching,” in Proc. 3rd Data Prefetching Championship (DPC3), 2019.

[26] R. Bera, A. V. Nori, O. Mutlu, and S. Subramoney, “DSPatch: Dual spatial pattern prefetcher,” in Proc. 52nd Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2019, pp. 531-544.

[27] S. Pakalapati and B. Panda, “Bouquet of instruction pointers: Instruction pointer classifier-based spatial hardware prefetching,” in Proc. 47th Annu. Int. Symp. Computer Architecture (ISCA), 2020, doi: 10.1109/ISCA45697.2020.00021.

[28] R. Bera, K. Kanellopoulos, A. V. Nori, T. Shahroodi, S. Subramoney, and O. Mutlu, “Pythia: A customizable hardware prefetching framework using online reinforcement learning,” in Proc. 54th Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2021.

[29] S. Jiang et al., “Matryoshka: A coalesced delta sequence prefetcher,” in Proc. 30th Int. Conf. Parallel Architectures and Compilation Techniques (PACT), 2021, doi: 10.1145/3472456.3473510.

[30] A. Navarro-Torres et al., “Berti: An accurate local-delta data prefetcher,” in Proc. 55th Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2022, pp. 975-991.

[31] B. Panda et al., “CLIP: Load criticality based data prefetching for bandwidth-constrained systems,” in Proc. 56th Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2023, doi: 10.1145/3613424.3614245.

[32] Y. Chen et al., “AfterImage: Leaking control flow data and tracking load operations via hardware prefetcher,” in Proc. 28th ACM Int. Conf. Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2023, doi: 10.1145/3575693.3575719.

[33] S. Nath et al., “Secure prefetching for secure cache systems,” in Proc. 57th Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2024, IEEE Xplore document 10764627.

[34] Y. Cui et al., “A highly effective page and PC based delta prefetcher,” ACM Trans. Architecture and Code Optimization, 2024, doi: 10.1145/3675398.

[35] A. Navarro-Torres et al., “A complexity-effective local delta prefetcher,” IEEE Trans. Computers, 2025, doi: 10.1109/TC.2025.3533086.

[36] M. Li et al., “Profile-guided temporal prefetching,” in Proc. 52nd Annu. Int. Symp. Computer Architecture (ISCA), 2025, doi: 10.1145/3695053.3731070.

[37] R. Bera et al., “Multi-agent reinforcement learning for multicore prefetching,” in Proc. 58th Annu. IEEE/ACM Int. Symp. Microarchitecture (MICRO), 2025, doi: 10.1145/3725843.3756096.

[38] F. Jiang et al., “SpectrePrefetch: Undermining cache-centric secure speculation defenses or secure cache designs via hardware prefetchers,” 2025, IEEE Xplore document 11240638.

 

IJAM

o                 Home

o                 Contents

o                 Editorial Board