Volume 28 , Issue 1 , June 2026 , Pages 60-74
1 Department of Computer Science, College of Science, University of Sulaimani, Iraq
Software cost estimation is one of the unresolved challenges in software engineering, which drives project success rates and organizational budgets. Conventional machine learning solutions fail to learn the relationships among features with different nonlinearities. This work introduces a Transformer-based novel architecture for software cost estimation based on Multi-Head Self-Attention. The novelties of the architecture lie in the use of deep Transformer blocks, which help learn hierarchical patterns from representational features created by a specific embedding layer and pooling methods. Extensive experimentation on 897 samples of nine publicly accessible datasets (albrecht, china, desharnais, finnish, isbsg10, kemerer, kitchenham, maxwell, miyazaki94) supports its efficiency. For instance, Transformer-based approach provides MAE of 2196.30 and R² of 0.3291–31% improvements over the best baseline (ElasticNet) and outperform Ridge, Random Forest, XGBoost, LightGBM and Gradient Boosting. Furthermore, Feature1, Feature2 and Feature3 are the most relevant for accurate predictions while ten-fold cross-validation achieved mean MAE of 2308.43 confirms generalizability and stability across different conditions.