Skip to main content
Back to timeline
Natural Sciences and Applied TechnologySource publication:

E-commerce CLV comparison: a neural network reaches RMSE 297.06 on exported scoring outputs, beating three regression and ensemble models

Synopsis

Using 17,049 e-commerce transaction records from 5,000 customers observed between January 2023 and February 2024, the study built a customer-level customer lifetime value (CLV) prediction pipeline with 33 modelling attributes and compared ridge-based Linear Regression, Random Forest, Gradient Boosted Trees and a feed-forward Neural Network under a common 10-fold cross-validation design in Altair AI Studio (RapidMiner), then converted predicted value into three actionable segments of 2,763 low-value, 1,252 medium-value and 985 high-value customers.

AI-generated editorial illustration: Machine Learning-Based Customer Lifetime Value Prediction in E-Commerce: A Comparative Study of Regression, Ensemble and Neural Models

Interpretation

The study assembles an end-to-end pipeline from raw transaction records to customer-level CLV prediction, aggregating and transforming transactions into 33 modelling attributes that include recency, tenure, purchase frequency, average order value and discount-related measures, together with encoded demographic and behavioural variables. Rather than treating CLV as a purely statistical exercise, the work links feature engineering, model comparison and customer segmentation into one deployment-oriented flow. The data basis is stated explicitly: 17,049 transaction records, 5,000 customers, an observation window from January 2023 to February 2024, and 33 modelling attributes.

Among the four regression approaches, the feed-forward Neural Network attains the lowest error on the exported prediction files with RMSE 297.06, versus 2,289.58 for ridge-based Linear Regression, 2,053.98 for Random Forest and 1,363.53 for Gradient Boosted Trees; the Neural Network also reaches MAE=187.86, R²=0.9968 and correlation=0.9984. The comparison rests primarily on a common 10-fold cross-validation stream, while predictions exported for all 5,000 customer records are reported separately as scoring diagnostics because the two evaluation paths are not equivalent for every learner. The model ranking comes from the common cross-validation stream, whereas RMSE, MAE, R² and correlation come from the exported prediction files; the text distinguishes these two evidence sources.

The study converts predicted customer value into three actionable segments: 2,763 low-value, 1,252 medium-value and 985 high-value customers. This step connects predictive accuracy directly to customer portfolio decisions, making CLV output usable for tiered operations rather than stopping at error metrics. The segment sizes come directly from predictions over all 5,000 customer records, and the three counts sum to the total customer base.

The study argues that evaluation provenance, temporal separation of predictors and target, and non-negative prediction constraints are essential when highly accurate scoring outputs are interpreted for deployment. The work brings evaluation provenance and deployment constraints explicitly into the discussion instead of reporting only a model ranking. This judgement is supported by the text's statement that the cross-validation stream and the exported scoring path are not equivalent, and by its emphasis on temporal separation and non-negative constraints.

Perspective

The pipeline targets e-commerce settings with customer-level transaction history, and suits teams that turn aggregated attributes such as recency, tenure, purchase frequency, average order value and discount measures into customer value estimates for tiered operations. The findings concern 5,000 customers within the January 2023 to February 2024 observation window; model comparison rests primarily on a common 10-fold cross-validation stream in Altair AI Studio (RapidMiner), with exported predictions serving as scoring diagnostics. For readers reusing the pipeline, its value lies in how it organises feature construction, the distinction between evaluation paths, and the step from prediction to segmentation.

The current read is incomplete in scope: figures, full result tables and the per-model comparison across the cross-validation and exported scoring paths are not included, so the specific differences for each learner across the two evaluation paths cannot be checked here. Readers may still watch how temporal separation and non-negative constraints are implemented in deployment, and how stable highly accurate scoring outputs remain for customer portfolio decisions. These are open questions for further validation and application.

Sources