Rethinking Transformer-Based Blind-Spot Network for Self-Supervised Image Denoising

Abstract

Blind-spot networks (BSN) have been prevalent neural architectures in self-supervised image denoising (SSID). However, most existing BSNs are conducted with convolution layers. Although transformers have shown the potential to overcome the limitations of convolutions in many image restoration tasks, the attention mechanisms may violate the blind-spot requirement, thereby restricting their applicability in BSN. To this end, we propose to analyze and redesign the channel and spatial attentions to meet the blind-spot requirement. Specifically, channel self-attention may leak the blind-spot information in multi-scale architectures, since the downsampling shuffles the spatial feature into channel dimensions. To alleviate this problem, we divide the channel into several groups and perform channel attention separately. For spatial self-attention, we apply an elaborate mask to the attention matrix to restrict and mimic the receptive field of dilated convolution. Based on the redesigned channel and window attentions, we build a Transformer-based Blind-Spot Network (TBSN), which shows strong local fitting and global perspective abilities. Furthermore, we introduce a knowledge distillation strategy that distills TBSN into smaller denoisers to improve computational efficiency while maintaining performance. Extensive experiments on real-world image denoising datasets show that TBSN largely extends the receptive field and exhibits favorable performance against state-of-the-art SSID methods.

Cite

Text

Li et al. "Rethinking Transformer-Based Blind-Spot Network for Self-Supervised Image Denoising." AAAI Conference on Artificial Intelligence, 2025. doi:10.1609/AAAI.V39I5.32506

Markdown

[Li et al. "Rethinking Transformer-Based Blind-Spot Network for Self-Supervised Image Denoising." AAAI Conference on Artificial Intelligence, 2025.](https://mlanthology.org/aaai/2025/li2025aaai-rethinking/) doi:10.1609/AAAI.V39I5.32506

BibTeX

@inproceedings{li2025aaai-rethinking,
  title     = {{Rethinking Transformer-Based Blind-Spot Network for Self-Supervised Image Denoising}},
  author    = {Li, Junyi and Zhang, Zhilu and Zuo, Wangmeng},
  booktitle = {AAAI Conference on Artificial Intelligence},
  year      = {2025},
  pages     = {4788-4796},
  doi       = {10.1609/AAAI.V39I5.32506},
  url       = {https://mlanthology.org/aaai/2025/li2025aaai-rethinking/}
}