Transformer-Based Attention Mechanism for Software Cost Estimation with Feature Importance Analysis

Volume 28 , Issue 1 , June 2026 , Pages 60-74

Authors

Hawar Othman 1

1 Department of Computer Science, College of Science, University of Sulaimani, Iraq

DOI logo 10.17656/sujpas.1142

Keywords

Abstract


Software cost estimation is one of the unresolved challenges in software engineering, which drives project success rates and organizational budgets. Conventional machine learning solutions fail to learn the relationships among features with different nonlinearities. This work introduces a Transformer-based novel architecture for software cost estimation based on Multi-Head Self-Attention. The novelties of the architecture lie in the use of deep Transformer blocks, which help learn hierarchical patterns from representational features created by a specific embedding layer and pooling methods. Extensive experimentation on 897 samples of nine publicly accessible datasets (albrecht, china, desharnais, finnish, isbsg10, kemerer, kitchenham, maxwell, miyazaki94) supports its efficiency. For instance, Transformer-based approach provides MAE of 2196.30 and R² of 0.3291–31% improvements over the best baseline (ElasticNet) and outperform Ridge, Random Forest, XGBoost, LightGBM and Gradient Boosting. Furthermore, Feature1, Feature2 and Feature3 are the most relevant for accurate predictions while ten-fold cross-validation achieved mean MAE of 2308.43 confirms generalizability and stability across different conditions.

Statistics
  • Article view63
  • Downloads5
  • First online25 June 2026
  • Published at25 June 2026

  • RIS
  • BibTeX
  • EndNote
  • Mendeley
  • APA (7th edition)
  • MLA (9th edition)
  • Chicago
  • Harvard
  • IEEE
  • Vancouver