Abstract A crucial factor indicating the fluid transport capability of rock porous media (RPM) like rocks is permeability. This study introduces a novel deep learning approach called ConViT to enhance the precision and effectiveness of permeability estimation using rock imagery. The ConViT model merges the strengths of Vision Transformer (ViT) and Convolutional Neural Networks (CNN), featuring an innovative Gated Positional Self-Attention (GPSA) module that seamlessly combines the benefits of attention mechanisms with convolutional architectures for RPM analysis. The proposed model uses binarized images of RPM as input, extracts features through GPSA, and ultimately enables end-to-end permeability prediction. Additionally, EigenCAM is used to identify the areas of concern of the RPM to enhance the interpretability of the result. An evaluation of multiple deep learning architectures was performed using the porous materials dataset. Findings reveal that the ConViT model achieves precise permeability estimation for RPM. The coefficient of determination (R2), mean square error (MSE), and mean absolute percentage error (MAPE) of the test set are 0.9807, 0.0013, and 0.2663 respectively. Leveraging the model's transfer learning capability, the research findings were successfully applied to permeability prediction of tight sandstone cast thin sections from the Shaximiao Formation, achieving a prediction accuracy of 82.04%. The research results provide a new method for end-to-end permeability evaluation of RPM.
Zhai et al. (Thu,) studied this question.