Skip to main navigation menu Skip to main content Skip to site footer

Probabilistic Medical Image Segmentation with Calibrated Uncertainty Estimation across Heterogeneous Clinical Images, Variable Acquisition Conditions, and Anatomies

Artificial Intelligence and Intelligent Systems Research, Volume 1, Issue 4, 2026 cover

Abstract

Medical image segmentation is a foundational task in computational pathology and radiology, facilitating precise anatomical localization, volumetric analysis, and downstream treatment planning. Despite the paradigm-shifting success of deep learning architectures in these domains, standard deterministic models remain notoriously overconfident and fail to provide reliable uncertainty estimates, particularly when deployed in heterogeneous clinical environments. This paper presents a comprehensive investigation into probabilistic medical image segmentation, emphasizing the critical requirement of calibrated uncertainty estimation across a wide spectrum of imaging modalities, variable acquisition conditions, and diverse anatomical structures. We systematically explore the theoretical underpinnings of Bayesian neural networks, stochastic latent variable models, and ensemble methodologies tailored for dense prediction tasks. Furthermore, we introduce robust calibration mechanisms designed to align the confidence of predictive models with their empirical accuracy, mitigating the severe risks associated with miscalibrated overconfidence in life-critical diagnostic systems. Through extensive theoretical analysis and empirical evaluation across multiple clinical datasets, encompassing neuro-oncology, hepatology, and cardiology, this research demonstrates that appropriately calibrated probabilistic frameworks not only yield competitive segmentation fidelity but also produce highly interpretable uncertainty maps. These uncertainty maps effectively capture both aleatoric and epistemic uncertainties, serving as vital signals for human-in-the-loop clinical workflows. The integration of advanced calibration techniques ensures that the segmentation models maintain diagnostic reliability even under significant domain shifts caused by disparate scanner protocols and inherent physiological variability.

Keywords

Medical Image Segmentation, Probabilistic Modeling, Uncertainty Calibration, Clinical Heterogeneity

PDF

References

  1. 1. Li, X., Yang, F., Chen, L., & Cai, H. (2016, July). Saliency transfer: An example-based method for salient object detection. In IJCAI (pp. 3411-3417).
  2. 2. Zhang B, Wang X, Kao H, et al. Consensus Matrix: A Role Specialized Multi Agent Framework for Structured Collaborative Decision Making in Agentic Visual Media Workflows[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2026: 4613 4622.
  3. 3. Wang, H., Wei, C., Ren, W., Liu, J., Lin, F., & Chen, W. (2026). Rationalrewards: Reasoning rewards scale visual generation both training and test time.arXiv preprint arXiv:2604.11626.
  4. 4. Li, Q., Luo, T., Jiang, M., Jiang, Z., Hou, C., & Li, F. (2025, April). Semi-supervised multi-view multi-label learning with view-specific transformer and enhanced pseudo-label. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 39, No. 17, pp. 18430-18438).
  5. 5. Li, Q., Luo, T., Jiang, M., Liao, J., & Jiang, Z. (2024, October). Deep incomplete multi-view network semi-supervised multi-label learning with unbiased loss. In Proceedings of the 32nd ACM International Conference on Multimedia (pp. 9048-9056).
  6. 6. Mi, L., Wang, W., Tu, W., He, Q., Kong, R., Fang, X., ... & Liu, Y. (2025, March). Empower vision applications with lora lmm. In Proceedings of the Twentieth European Conference on Computer Systems (pp. 261-277).
  7. 7. Zhao, H., Wang, Q., Zhan, G., Min, W., Zou, Y., & Cui, S. (2022). Need only one more point (NOOMP): Perspective adaptation crowd counting in complex scenes. IEEE Transactions on Multimedia, 25, 1414-1426.
  8. 8. Peng, Qucheng, et al. "3d vision-language gaussian splatting." International Conference on Learning Representations. Vol. 2025. 2025.
  9. 9. Simonyan, K., & Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. In Proceedings of ICLR.
  10. 10. Yang, Y., Liao, Y. H., Bortins, I., Baldwin, D. P., & Zhang, S. (2024). Unidirectional structured light system calibration with auxiliary camera and projector.Optics and Lasers in Engineering,175, 107984.
  11. 11. Xia, Y., Ye, T., Huang, J., Mei, X., & Ma, J. (2026, March). Probabilistic deformation consistency for unsupervised shape matching. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 40, No. 13, pp. 10951-10959).
  12. 12. Qi Z, Wang S, Su C, et al. Self-regulated learning for egocentric video activity anticipation[J]. IEEE transactions on pattern analysis and machine intelligence, 2021, 45(6): 6715-6730.
  13. 13. Yao, S., Guo, J., Li, J., Ou, J., Feng, Y., Hu, J., & Liu, D. (2025). Adversarial hard negative samples for continual relation extraction. Applied Soft Computing, 181, 113365.
  14. 14. Wang, J., Malawade, A. V., Zhou, J., Yu, S. Y., & Al Faruque, M. A. (2024, January). Rs2g: Data-driven scene-graph extraction and embedding for robust autonomous perception and scenario understanding. In 2024 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) (pp. 7478-7487). IEEE.
  15. 15. Xia, Y., & Ma, J. (2026). Locality Optimization Refinement with Deformation for Shape Matching via Functional Maps.International Journal of Computer Vision,134(2), 76.
  16. 16. Liu, C., Ma, L., Zhang, X. T., Zhang, Y., Zhang, H., Yang, X., & Tian, F. (2026). Boosting omni-modal language models: Staged post-training with visually debiased evaluation. arXiv preprint arXiv:2605.12034.
  17. 17. Yang, Y., Ma, X., Li, C., Zheng, Z., Zhang, Q., Huang, G., ... & Zhao, Q. (2021). Believe what you see: Implicit constraint approach for offline multi-agent reinforcement learning. Advances in Neural Information Processing Systems, 34, 10299-10312.
  18. 18. Cong, P., Xu, Y., Ren, Y., Zhang, J., Xu, L., Wang, J., ... & Ma, Y. (2023, June). Weakly supervised 3d multi-person pose estimation for large-scale scenes based on monocular camera and single lidar. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 37, No. 1, pp. 461-469).
  19. 19. Huang, H., Zhang, J., Zhang, J., Xu, J., & Wu, Q. (2020). Low-rank pairwise alignment bilinear network for few-shot fine-grained image classification.IEEE Transactions on Multimedia,23, 1666-1680.
  20. 20. Xia, Y., & Ma, J. (2022). Locality-guided global-preserving optimization for robust feature matching.IEEE Transactions on Image Processing,31, 5093-5108.
  21. 21. Qi, Z., Yuan, Y., Ruan, X., Wang, S., Zhang, W., & Huang, Q. (2024, March). Bias-conflict sample synthesis and adversarial removal debias strategy for temporal sentence grounding in video. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, No. 5, pp. 4533-4541).
  22. 22. Chen, L., Jiang, X., Liu, X., & Zhou, Z. (2021). Logarithmic norm regularized low-rank factorization for matrix and tensor completion. IEEE Transactions on Image Processing, 30, 3434-3449.
  23. 23. Zhao, Y., Li, Z., Guo, X., & Lu, Y. (2022). Alignment-guided temporal attention for video action recognition. Advances in Neural Information Processing Systems, 35, 13627-13639.
  24. 24. Yang, Y., & Zhang, S. (2025). Pixel-wise calibration for a multi-focus microscopic 3D imaging system.Optics and Lasers in Engineering,194, 109127.
  25. 25. Guo, Zixin, Kai Zhao, and Luyan Zhang. "InstanceRSR: Real-World Super-Resolution via Instance-Aware Representation Alignment." ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2026.
  26. 26. Qu, W., Wang, J., Gong, Y., Huang, X., & Xiao, L. (2025, June). An end-to-end robust point cloud semantic segmentation network with single-step conditional diffusion models. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 27325-27335). IEEE.
  27. 27. Simonyan, K., & Zisserman, A. (2014). Two-stream convolutional networks for action recognition in videos. In Advances in Neural Information Processing Systems.
  28. 28. Zhang, W., Huang, M., Zhou, Y., Zhang, J., Yu, J., Wang, J., & Xu, L. (2024). Both2hands: Inferring 3d hands from both text prompts and body dynamics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2393-2404).
  29. 29. Lee, Y., Gao, P., Chen, Z., Fan, W., Jing, G., & Hu, Y. (2025, June). Boosting audio-visual segmentation via triple-modalities alignment. In 2025 IEEE International Conference on Multimedia and Expo (ICME) (pp. 1-6). IEEE.
  30. 30. Qi, Z., Yuan, Y., Ruan, X., Wang, S., Zhang, W., & Huang, Q. (2024). Collaborative debias strategy for temporal sentence grounding in video. IEEE Transactions on Circuits and Systems for Video Technology, 34(11), 10972-10986.
  31. 31. Xia, Y., Lu, Y., Gao, Y., & Ma, J. (2024, March). Locality preserving refinement for shape matching with functional maps. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, No. 6, pp. 6207-6215).
  32. 32. Gao, P., Lee, Y., Zhang, H., Chen, Z., & Liu, X. (2024, December). Dynamic identity-guided attention network for visible-infrared person re-identification. In International Conference on Neural Information Processing (pp. 364-379). Singapore: Springer Nature Singapore.
  33. 33. Ding, Y., Li, Z., Huang, D., Zhang, K., Li, Z., & Feng, W. (2022). Adaptive range guided multi-view depth estimation with normal ranking loss. In Proceedings of the Asian Conference on Computer Vision (pp. 1892-1908).
  34. 34. Wang, H., Feng, W., Yu, J., Liu, C., Nie, P., Lin, F., ... & Wei, C. (2026). Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation.arXiv preprint arXiv:2607.05382.
  35. 35. Qu, D., & Ma, Y. (2026). Fourier–Transformer Mixer Network for Efficient Video Scene Graph Prediction.Engineering Proceedings,120(1), 16.