Top

International Journal of Computer Assisted Radiology and Surgery

Published in:

Open Access 08-01-2024 | Cataract | Original Article

DeepPyramid+: medical image segmentation using Pyramid View Fusion and Deformable Pyramid Reception

Authors: Negin Ghamsarian, Sebastian Wolf, Martin Zinkernagel, Klaus Schoeffmann, Raphael Sznitman

Published in: International Journal of Computer Assisted Radiology and Surgery | Issue 5/2024

Abstract

Purpose

Semantic segmentation plays a pivotal role in many applications related to medical image and video analysis. However, designing a neural network architecture for medical image and surgical video segmentation is challenging due to the diverse features of relevant classes, including heterogeneity, deformability, transparency, blunt boundaries, and various distortions. We propose a network architecture, DeepPyramid+, which addresses diverse challenges encountered in medical image and surgical video segmentation.

Methods

The proposed DeepPyramid+ incorporates two major modules, namely “Pyramid View Fusion” (PVF) and “Deformable Pyramid Reception” (DPR), to address the outlined challenges. PVF replicates a deduction process within the neural network, aligning with the human visual system, thereby enhancing the representation of relative information at each pixel position. Complementarily, DPR introduces shape- and scale-adaptive feature extraction techniques using dilated deformable convolutions, enhancing accuracy and robustness in handling heterogeneous classes and deformable shapes.

Results

Extensive experiments conducted on diverse datasets, including endometriosis videos, MRI images, OCT scans, and cataract and laparoscopy videos, demonstrate the effectiveness of DeepPyramid+ in handling various challenges such as shape and scale variation, reflection, and blur degradation. DeepPyramid+ demonstrates significant improvements in segmentation performance, achieving up to a 3.65% increase in Dice coefficient for intra-domain segmentation and up to a 17% increase in Dice coefficient for cross-domain segmentation.

Conclusions

DeepPyramid+ consistently outperforms state-of-the-art networks across diverse modalities considering different backbone networks, showcasing its versatility. Accordingly, DeepPyramid+ emerges as a robust and effective solution, successfully overcoming the intricate challenges associated with relevant content segmentation in medical images and surgical videos. Its consistent performance and adaptability indicate its potential to enhance precision in computerized medical image and surgical video analysis applications.

This paper is an extended version of DeepPyramid [6], featuring minor enhancements in the DPR module.

The PyTorch implementation of DeepPyramid \(+\) is publicly available at https://github.com/Negin-Ghamsarian/DeepPyramid_Plus.

This paper aims to design a dedicated network tailored to address medical image and video segmentation challenges, emphasizing various modalities but not within a multi-modal training framework. We substantiate the efficacy of our model through distinctive validations across diverse medical image and video datasets.

Ghamsarian N, Taschwer M, Putzgruber-Adamitsch D, Sarny S, Schoeffmann K (2021) Relevance detection in cataract surgery videos by spatio-temporal action localization. In: 2020 25th International conference on pattern recognition (ICPR), pp 10720–10727

Ghamsarian N (2020) Enabling relevance-based exploration of cataract videos. In: Proceedings of the 2020 international conference on multimedia retrieval, pp 378–382

Ghamsarian N, Amirpourazarian H, Timmerer C, Taschwer M, Schöffmann K (2020) Relevance-based compression of cataract surgery videos using convolutional neural networks. In: Proceedings of the 28th ACM international conference on multimedia, pp 3577–3585

Ghamsarian N, Taschwer M, Putzgruber-Adamitsch D, Sarny S, El-Shabrawi Y, Schoeffmann K (2021) LensID: a CNN-RNN-based framework towards lens irregularity detection in cataract surgery videos. In: Medical image computing and computer assisted intervention—MICCAI 2021: 24th international conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part VIII 24. Springer, pp 76–86

Huang X, Wang H, She C, Feng J, Liu X, Hu X, Chen L, Tao Y (2022) Artificial intelligence promotes the diagnosis and screening of diabetic retinopathy. Front Endocrinol 13:946915CrossRef

Ghamsarian N, Taschwer M, Sznitman R, Schoeffmann K (2022) Deeppyramid: Enabling pyramid view and deformable pyramid reception for semantic segmentation in cataract surgery videos. In: Medical image computing and computer assisted intervention—MICCAI 2022: 25th international conference, Singapore, September 18–22, 2022, Proceedings, Part V. Springer, pp 276–286

Ronneberger O, Fischer P, Brox T (2015) U-net: Convolutional networks for biomedical image segmentation. In: Navab N, Hornegger J, Wells WM, Frangi AF (eds) Medical Image Computing and Computer-Assisted Intervention—MICCAI 2015. Springer, Cham, pp 234–241

Chen X, Zhang R, Yan P (2019) Feature fusion encoder decoder network for automatic liver lesion segmentation. In: 2019 IEEE 16th international symposium on biomedical imaging (ISBI 2019), pp 430–433

Ni Z-L, Bian G-B, Zhou X-H, Hou Z-G, Xie X-L, Wang C, Zhou Y-J, Li R-Q, Li Z (2019) Raunet: Residual attention u-net for semantic segmentation of cataract surgical instruments. In: Gedeon T, Wong KW, Lee M (eds) Neural Information Processing. Springer, Cham, pp 139–149CrossRef

10.

Gu Z, Cheng J, Fu H, Zhou K, Hao H, Zhao Y, Zhang T, Gao S, Liu J (2019) Ce-net: Context encoder network for 2d medical image segmentation. IEEE Trans Med Imaging 38(10):2281–2292CrossRefPubMed

11.

Ni Z-L, Bian G-B, Wang G-A, Zhou X-H, Hou Z-G, Chen H-B, Xie X-L (2020) Pyramid attention aggregation network for semantic segmentation of surgical instruments. Proc AAAI Conf Artif Intell 34(07):11782–11790

12.

Ni Z-L, Bian G-B, Wang G-A, Zhou X-H, Hou Z-G, Xie X-L, Li Z, Wang Y-H (2021) Barnet: bilinear attention network with adaptive receptive fields for surgical instrument segmentation. In: Proceedings of the twenty-ninth international conference on international joint conferences on artificial intelligence, pp 832–838

13.

Feng S, Zhao H, Shi F, Cheng X, Wang M, Ma Y, Xiang D, Zhu W, Chen X (2020) CPFNet: Context pyramid fusion network for medical image segmentation. IEEE Trans Med Imaging 39(10):3008–3018CrossRefPubMed

14.

Zhou Z, Siddiquee MMR, Tajbakhsh N, Liang J (2020) Unet++: Redesigning skip connections to exploit multiscale features in image segmentation. IEEE Trans Med Imaging 39(6):1856–1867CrossRefPubMed

15.

Roy AG, Navab N, Wachinger C (2019) Recalibrating fully convolutional networks with spatial and channel “squeeze and excitation" blocks. IEEE Trans Med Imaging 38(2):540–549CrossRefPubMed

16.

Ghamsarian N, Taschwer M, Putzgruber-Adamitsch D, Sarny S, El-Shabrawi Y, Schöffmann K (2021) Recal-net: Joint region-channel-wise calibrated network for semantic segmentation in cataract surgery videos. In: Neural information processing: 28th international conference, ICONIP 2021, Sanur, Bali, Indonesia, December 8–12, 2021, Proceedings, Part III 28. Springer, pp 391–402

17.

Zhao H, Shi J, Qi X, Wang X, Jia J (2017) Pyramid scene parsing network. In: Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR)

18.

Chen L-C, Papandreou G, Kokkinos I, Murphy K, Yuille AL (2018) Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans Pattern Anal Mach Intell 40(4):834–848CrossRefPubMed

19.

Chen L-C, Zhu Y, Papandreou G, Schroff F, Adam H (2018) Encoder-decoder with atrous separable convolution for semantic image segmentation. In: Proceedings of the European conference on computer vision (ECCV)

20.

Ghamsarian N, El-Shabrawi Y, Nasirihaghighi S, Putzgruber-Adamitsch D, Zinkernagel M, Wolf S, Schoeffmann K, Sznitman R (2023) Cataract-1K: cataract surgery dataset for scene segmentation, phase recognition, and irregularity detection. arXiv preprint https://arxiv.org/abs/2312.06295

21.

Bodenstedt S, Speidel S, Allan M, Stoyanov D, Maier-Hein L, Kenngott H, Wagner M (2015) Multi-instrument EndoVis challenge dataset. https://endovissub-instrument.grand-challenge.org/

22.

Leibetseder A, Schoeffmann K, Keckstein J, Keckstein S (2022) Endometriosis detection and localization in laparoscopic gynecology. Multimed Tools Appl 81(5):6191–6215CrossRef

23.

Liu Q, Dou Q, Yu L, Heng PA (2020) MS-Net: multi-site network for improving prostate segmentation with heterogeneous MRI data. IEEE Trans Med Imaging

24.

Bogunovic H, Venhuizen F, Klimscha S, Apostolopoulos S, Bab-Hadiashar A, Bagci U, Beg MF, Bekalo L, Chen Q, Ciller C, Gopinath K, Gostar AK, Jeon K, Ji Z, Kang SH, Koozekanani DD, Lu D, Morley D, Parhi KK, Park HS, Rashno A, Sarunic M, Shaikh S, Sivaswamy J, Tennakoon R, Yadav S, De Zanet S, Waldstein SM, Gerendas BS, Klaver C, Sánchez CI, Schmidt-Erfurth U (2019) Retouch: the retinal oct fluid detection and segmentation benchmark and challenge. IEEE Trans Med Imaging 38(8):1858–1874CrossRefPubMed

25.

Grammatikopoulou M, Flouty E, Kadkhodamohammadi A, Quellec G, Chow A, Nehme J, Luengo I, Stoyanov D (2021) CaDIS: Cataract dataset for surgical RGB-image segmentation. Med Image Anal 71:102053CrossRefPubMed

26.

27.

Xiao T, Liu Y, Zhou B, Jiang Y, Sun J (2018) Unified perceptual parsing for scene understanding. In: Proceedings of the European conference on computer vision (ECCV), pp 418–434

28.

Ghamsarian N, Gamazo Tejero J, Márquez-Neila P, Wolf S, Zinkernagel M, Schoeffmann K, Sznitman R (2023) Domain adaptation for medical image segmentation using transformation-invariant self-training. In: International conference on medical image computing and computer-assisted intervention. Springer, pp 331–341

Title: DeepPyramid+: medical image segmentation using Pyramid View Fusion and Deformable Pyramid Reception
Authors: Negin Ghamsarian
Sebastian Wolf
Martin Zinkernagel
Klaus Schoeffmann
Raphael Sznitman
Publication date: 08-01-2024
Publisher: Springer International Publishing
Keyword: Cataract
Published in: International Journal of Computer Assisted Radiology and Surgery / Issue 5/2024
Print ISSN: 1861-6410
Electronic ISSN: 1861-6429
DOI: https://doi.org/10.1007/s11548-023-03046-2

Keynote webinar | Spotlight on medication adherence

Springer Medicine

DeepPyramid+: medical image segmentation using Pyramid View Fusion and Deformable Pyramid Reception

Abstract

Purpose

Methods

Results

Conclusions

Keynote webinar | Spotlight on medication adherence

Springer Medicine

Abstract

Purpose

Methods

Results

Conclusions

Please log in to get access to this content

Other articles of this Issue 5/2024

Assistance by adaptative damping on a complex bimanual task in laparoscopic surgery

Efficient EndoNeRF reconstruction and its application for data-driven surgical simulation

SF-TMN: SlowFast temporal modeling network for surgical phase recognition

Pose-based tremor type and level analysis for Parkinson’s disease from video

PELE scores: pelvic X-ray landmark detection with pelvis extraction and enhancement

OCT-based intra-cochlear imaging and 3D reconstruction: ex vivo validation of a robotic platform