De Lima JPC, Santos PC, Alves MA, Beck AC, Carro L (2018)
Publication Type: Conference contribution
Publication year: 2018
Publisher: Association for Computing Machinery, Inc
Pages Range: 113-120
Conference Proceedings Title: 2018 ACM International Conference on Computing Frontiers, CF 2018 - Proceedings
Event location: Ischia, ITA
ISBN: 9781450357616
Scaling existing architectures to large-scale data-intensive applications is limited by energy and performance losses caused by off-chip memory communication and data movements in the cache hierarchy. Processing-in-Memory (PIM) has been recently revisited to address the issues of memory and power wall, mainly due to the maturity of 3D-stacking manufacturing technology and the increasing demand for bandwidth and parallel access in emerging data-centric applications. Recent studies have shown a wide variety of processing mechanisms to be placed in the logic layer of 3D-stacked memories, not to mention the already available 3Dstacked DRAMs, such as Micron's Hybrid Memory Cube (HMC). Nevertheless, a few studies compare PIM accelerators to each other and have made efforts to indicate the trade-offs between power, area, and performance. In this paper, we review different state-ofthe- art 3D-stacked in-memory accelerators, and we analyze them considering important constraints regarding area and power due to critical embedded nature of PIM. Aiming to point in the direction of massive parallel PIM designs, we take the simplest design found in this survey, and we explore the architectural design space to meet the constraints imposed by HMC. Our results show that the most straightforward approach can provide the highest performance while consuming the lowest amount of area and power, which makes it the most suitable design found in this survey for an energy-efficient in-memory accelerator, whether it goes in High- Performance Computing or Embedded Systems. For instance, the outstanding point in the design space indicates that a performance density of 320 GBps/mm2 and a performance efficiency of 0.6 GBps/ mW can be achieved in the best scenario, that is, when a massive parallel application reaches the peak bandwidth.
APA:
De Lima, J.P.C., Santos, P.C., Alves, M.A., Beck, A.C., & Carro, L. (2018). Design space exploration for PIM architectures in 3D-stacked memories. In 2018 ACM International Conference on Computing Frontiers, CF 2018 - Proceedings (pp. 113-120). Ischia, ITA: Association for Computing Machinery, Inc.
MLA:
De Lima, João Paulo C., et al. "Design space exploration for PIM architectures in 3D-stacked memories." Proceedings of the 15th ACM International Conference on Computing Frontiers, CF 2018, Ischia, ITA Association for Computing Machinery, Inc, 2018. 113-120.
BibTeX: Download