000 05666nam a22005057i 4500
999 _c96058
_d96058
001 13462108
005 20260710113446.0
006 m||||||||d||||||||
007 cr||n|||||||||
008 231018s2024 nju o 000 0 eng d
020 _a9781394171880
020 _a9781394171910
_q(electronic bk. : oBook)
020 _a1394171919
_q(electronic bk. : oBook)
020 _a9781394171897
_qelectronic book
020 _a1394171897
_qelectronic book
020 _a9781394171903
_q(electronic bk.)
020 _a1394171900
_q(electronic bk.)
020 _z9781394171880
_qhardcover
020 _z1394171889
_qhardcover
024 7 _a10.1002/9781394171910
_2doi
035 9 _a(GOBI)99996250713
035 _a(OCoLC)1404053066
037 _a10296182
_bIEEE
040 _aYDX
_beng
_erda
_cYDX
_dYDX
_dIEEEE
_dDG1
041 _aeng
050 4 _aQA76.87
_b.M86 2024
082 0 4 _a006.3/2
_223/eng/20231025
100 1 _aMunir, Arslan,
_eauthor.
245 1 0 _aAccelerators for convolutional neural networks /
_cArslan Munir, Joonho Kong, Mahmood Azhar Qureshi.
264 1 _aHoboken, New Jersey :
_bJohn Wiley & Sons, Inc.,
_c[2024]
300 _a1 online resource.
336 _atext
_btxt
_2rdacontent
337 _acomputer
_bc
_2rdamedia
338 _aonline resource
_bcr
_2rdacarrier
505 0 _aAbout the Authors xiii -- Preface xv -- Part I Overview 1 -- 1 Introduction 3 -- 1.1 History and Applications 5 -- 1.2 Pitfalls of High-Accuracy DNNs/CNNs 6 -- 1.2.1 Compute and Energy Bottleneck 6 -- 1.2.2 Sparsity Considerations 9 -- 1.3 Chapter Summary 11 -- 2 Overview of Convolutional Neural Networks 13 -- 2.1 Deep Neural Network Architecture 13 -- 2.2 Convolutional Neural Network Architecture 15 -- 2.3 Popular CNN Models 26 -- 2.4 Popular CNN Datasets 30 -- 2.5 CNN Processing Hardware 31 -- 2.6 Chapter Summary 37 -- Part II Compressive Coding for CNNs 39 -- 3 Contemporary Advances in Compressive Coding for CNNs 41 -- 3.1 Background of Compressive Coding 41 -- 3.2 Compressive Coding for CNNs 43 -- 3.3 Lossy Compression for CNNs 43 -- 3.4 Lossless Compression for CNNs 44 -- 3.5 Recent Advancements in Compressive Coding for CNNs 48 -- 3.6 Chapter Summary 50 -- 4 Lossless Input Feature Map Compression 51 -- 4.1 Two-Step Input Feature Map Compression Technique 52 -- 4.2 Evaluation 55 -- 4.3 Chapter Summary 57 -- 5 Arithmetic Coding and Decoding for 5-Bit CNN Weights 59 -- 5.1 Architecture and Design Overview 60 -- 5.2 Algorithm Overview 63 -- 5.3 Weight Decoding Algorithm 67 -- 5.4 Encoding and Decoding Examples 69 -- 5.5 Evaluation Methodology 74 -- 5.6 Evaluation Results 75 -- 5.7 Chapter Summary 84 -- Part III Dense CNN Accelerators 85 -- 6 Contemporary Dense CNN Accelerators 87 -- 6.1 Background on Dense CNN Accelerators 87 -- 6.2 Representation of the CNNWeights and Feature Maps in Dense Format 87 -- 6.3 Popular Architectures for Dense CNN Accelerators 89 -- 6.4 Recent Advancements in Dense CNN Accelerators 92 -- 6.5 Chapter Summary 93 -- 7 iMAC: Image-to-Column and General Matrix Multiplication-Based Dense CNN Accelerator 95 -- 7.1 Background and Motivation 95 -- 7.2 Architecture 97 -- 7.3 Implementation 99 -- 7.4 Chapter Summary 100 -- 8 NeuroMAX: A Dense CNN Accelerator 101 -- 8.1 RelatedWork 102 -- 8.2 Log Mapping 103 -- 8.3 Hardware Architecture 105 -- 8.4 Data Flow and Processing 108 -- 8.5 Implementation and Results 118 -- 8.6 Chapter Summary 124 -- Part IV Sparse CNN Accelerators 125 -- 9 Contemporary Sparse CNN Accelerators 127 -- 9.1 Background of Sparsity in CNN Models 127 -- 9.2 Background of Sparse CNN Accelerators 128 -- 9.3 Recent Advancements in Sparse CNN Accelerators 131 -- 9.4 Chapter Summary 133 -- 10 CNN Accelerator for In Situ Decompression and Convolution of Sparse Input Feature Maps 135 -- 10.1 Overview 135 -- 10.2 Hardware Design Overview 135 -- 10.3 Design Optimization Techniques Utilized in the Hardware Accelerator 140 -- 10.4 FPGA Implementation 141 -- 10.5 Evaluation Results 143 -- 10.6 Chapter Summary 149 -- 11 Sparse-PE: A Sparse CNN Accelerator 151 -- 11.1 RelatedWork 155 -- 11.2 Sparse-PE 156 -- 11.3 Implementation and Results 174 -- 11.4 Chapter Summary 184 -- 12 Phantom: A High-Performance Computational Core for Sparse CNNs 185 -- 12.1 RelatedWork 189 -- 12.2 Phantom 190 -- 12.3 Phantom-2D 201 -- 12.4 Experiments and Results 209 -- 12.5 Chapter Summary 218 -- Part V HW/SW Co-Design and Co-Scheduling for CNN Acceleration 221 -- 13 State-of-the-Art in HW/SW Co-Design and Co-Scheduling for CNN Acceleration 223 -- 13.1 HW/SW Co-Design 223 -- 13.2 HW/SW Co-Scheduling 228 -- 13.3 Chapter Summary 230 -- 14 Hardware/Software Co-Design for CNN Acceleration 231 -- 14.1 Background of iMAC Accelerator 231 -- 14.2 Software Partition for iMAC Accelerator 232 -- 14.3 Experimental Evaluations 235 -- 14.4 Chapter Summary 237 -- 15 CPU-Accelerator Co-Scheduling for CNN Acceleration 239 -- 15.1 Background and Preliminaries 240 -- 15.2 CNN Acceleration with CPU-Accelerator Co-Scheduling 242 -- 15.3 Experimental Results 251 -- 15.4 Chapter Summary 257 -- 16 Conclusions 259 -- References 265 -- Index 285.
588 _aDescription based on online resource; title from digital title page (viewed on October 25, 2023).
650 0 _aNeural networks (Computer science)
655 4 _aElectronic books.
700 1 _aKong, Joonho,
_eauthor.
700 1 _aQureshi, Mahmood Azhar,
_eauthor.
776 0 8 _iPrint version:
_z1394171889
_z9781394171880
_w(OCoLC)1368339666
856 4 0 _uhttps://onlinelibrary.wiley.com/book/10.1002/9781394171910
_yFull text is available at Wiley Online Library. Click here to view.
901 _aYBPebook
942 _2ddc
_cER