HyCA: A Hybrid Computing Architecture for Fault Tolerant Deep Learning

Liu, Cheng; Chu, Cheng; Xu, Dawen; Wang, Ying; Wang, Qianlong; Li, Huawei; Li, Xiaowei; Cheng, Kwang-Ting

Computer Science > Hardware Architecture

arXiv:2106.04772 (cs)

[Submitted on 9 Jun 2021 (v1), last revised 27 Oct 2021 (this version, v2)]

Title:HyCA: A Hybrid Computing Architecture for Fault Tolerant Deep Learning

Authors:Cheng Liu, Cheng Chu, Dawen Xu, Ying Wang, Qianlong Wang, Huawei Li, Xiaowei Li, Kwang-Ting Cheng

View PDF

Abstract:Hardware faults on the regular 2-D computing array of a typical deep learning accelerator (DLA) can lead to dramatic prediction accuracy loss. Prior redundancy design approaches typically have each homogeneous redundant processing element (PE) to mitigate faulty PEs for a limited region of the 2-D computing array rather than the entire computing array to avoid the excessive hardware overhead. However, they fail to recover the computing array when the number of faulty PEs in any region exceeds the number of redundant PEs in the same region. The mismatch problem deteriorates when the fault injection rate rises and the faults are unevenly distributed. To address the problem, we propose a hybrid computing architecture (HyCA) for fault-tolerant DLAs. It has a set of dot-production processing units (DPPUs) to recompute all the operations that are mapped to the faulty PEs despite the faulty PE locations. According to our experiments, HyCA shows significantly higher reliability, scalability, and performance with less chip area penalty when compared to the conventional redundancy approaches. Moreover, by taking advantage of the flexible recomputing, HyCA can also be utilized to scan the entire 2-D computing array and detect the faulty PEs effectively at runtime.

Subjects:	Hardware Architecture (cs.AR)
Cite as:	arXiv:2106.04772 [cs.AR]
	(or arXiv:2106.04772v2 [cs.AR] for this version)
	https://doi.org/10.48550/arXiv.2106.04772

Submission history

From: Cheng Chu [view email]
[v1] Wed, 9 Jun 2021 02:26:02 UTC (9,480 KB)
[v2] Wed, 27 Oct 2021 11:40:05 UTC (5,951 KB)

Computer Science > Hardware Architecture

Title:HyCA: A Hybrid Computing Architecture for Fault Tolerant Deep Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Hardware Architecture

Title:HyCA: A Hybrid Computing Architecture for Fault Tolerant Deep Learning

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators