Cross-Dataset Parcel Workload Priors for Last-Mile Capacity Forecasting: Integrating Package Segmentation with Delivery Operation Signals
DOI:
https://doi.org/10.51903/jtie.v5i1.564Keywords:
Computer Vision, Package Segmentation, Last-Mile Delivery, Parcel Capacity Planning, Cross-Dataset Distributional ProxyAbstract
This paper evaluates a cross-dataset framework for parcel segmentation, next-day last-mile operational-intensity forecasting, and capacity-oriented error analysis. The operational study uses 4,514,661 delivery tasks and 6,136,147 pickup tasks from LaDe-D and LaDe-P across five cities (May–October 2022), while the visual study uses 2,197 Package Segmentation images with 7,643 annotated package instances. A segmentation model estimates package count, foreground area, and instance-area dispersion to construct an additive workload descriptor. Because the images are not paired with operational records, city-day visual variables are represented as cross-dataset distributional proxies derived from empirical-rank mapping. Segmentation baselines are evaluated independently of forecasting. The random-forest pixel classifier achieves the best segmentation performance (IoU 0.5323, Dice 0.6450), outperforming YOLO11n-seg (IoU 0.3439, Dice 0.4591). Operational intensity is represented by the first principal component of delivery orders, active couriers, areas of interest, regions, and area types, explaining 89.19% of training variance. On a 230 city-day chronological holdout, ElasticNet Operational achieves the best forecasting accuracy (MAE 2.0521). Within XGBoost models, Vision-Prior records MAE 2.4988, slightly outperforming Delivery+Pickup (MAE 2.5154) and a shuffled-prior control (MAE 2.5973). However, the 0.0166 MAE improvement is not statistically significant under a paired moving-block bootstrap (95% CI: −0.0558 to 0.0319). A separate analysis of 6,112 Amazon routes and 1,457,175 packages similarly shows only a 0.24% MAE reduction from parcel-mix features. Overall, the proposed proxy provides only limited gains, while the operational ElasticNet remains the strongest baseline. Time- and location-paired visual observations are needed to establish practical operational benefits from computer vision.
References
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5-32. https://doi.org/10.1023/A:1010933404324
Chen, S., He, S., & Sun, E. (2024). Risk-bounded GPU resource oversubscription via conformal demand envelopes in production AI clusters. JACS, 4(5), 119-134. https://doi.org/10.69987/JACS.2024.40509
Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785-794). https://doi.org/10.1145/2939672.2939785
Crainic, T. G., & Laporte, G. (1997). Planning models for freight transportation. European Journal of Operational Research, 97(3), 409-438. https://doi.org/10.1016/S0377-2217(96)00298-6
Dai, N., Hu, X., Xu, K., Yuan, Y., & Lu, Z. (2025). Research on dimension measurement algorithm for parcel boxes in high-speed sorting system. Scientific Reports. https://doi.org/10.1038/s41598-025-07730-y
Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189-1232. https://doi.org/10.1214/aos/1013203451
Gevaers, R., Van de Voorde, E., & Vanelslander, T. (2014). Cost modelling and simulation of last-mile characteristics in an innovative B2C supply chain environment with implications on urban areas and cities. Procedia - Social and Behavioral Sciences, 125, 398-411. https://doi.org/10.1016/j.sbspro.2014.01.1483
Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.
Han, S., Liu, X., Han, X., Wang, G., & Wu, S. (2020). Visual sorting of express parcels based on multi-task deep learning. Sensors, 20(23), 6785. https://doi.org/10.3390/s20236785
He, K., Gkioxari, G., Dollár, P., & Girshick, R. (2017). Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (pp. 2961-2969). https://doi.org/10.1109/ICCV.2017.322
He, S., Chang, X., & Sun, E. (2024). Cross-cloud transfer learning for AI training capacity forecasting under workload and topology distribution shift. JACS, 4(1), 100-120. https://doi.org/10.69987/JACS.2024.40108
He, S., Li, C., & Rao, H. (2025). Few-shot cold-start workload forecasting for new AI inference tenants with time-series foundation models. Journal of Technology Informatics and Engineering, 4(1), 306-324. https://doi.org/10.51903/jtie.v4i1.546
He, S., Nie, J., & Li, C. (2026). Power-aware inventory planning for AI infrastructure using job-level forecasting and LLM workload explanations. Journal of Technology Informatics and Engineering, 5(1), 341-359. https://doi.org/10.51903/jtie.v5i1.548
He, S., Tu, H., & Liu, I. (2023). Safe PD capacity forecasting with time-series foundation models and calibrated uncertainty for heterogeneous GPU clusters. JACS, 3(4), 48-66. https://doi.org/10.69987/JACS.2023.30404
Hyndman, R. J., & Athanasopoulos, G. (2021). Forecasting: Principles and practice (3rd ed.). OTexts.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W. Y., Dollár, P., & Girshick, R. (2023). Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 4015-4026).
Lin, T. Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., & Zitnick, C. L. (2014). Microsoft COCO: Common objects in context. In D. Fleet, T. Pajdla, B. Schiele, & T. Tuytelaars (Eds.), Computer Vision - ECCV 2014 (pp. 740-755). Springer. https://doi.org/10.1007/978-3-319-10602-1_48
Ling, Z., Xin, Q., Lin, Y., Su, G., & Shui, Z. (2024). Optimization of autonomous driving image detection based on RFAConv and triplet attention. In Proceedings of the 2nd International Conference on Software Engineering and Machine Learning (SEML 2024).
Lu, S., & Zhou, D. (2023). LLM-augmented customer representation learning for next-purchase prediction in online retail. JACS, 3(3), 50-64. https://doi.org/10.69987/JACS.2023.30305
Lu, Z., Dai, N., Hu, X., Xu, K., & Yuan, Y. (2024). Research on high-speed classification and location algorithm for logistics parcels based on a monocular camera. Scientific Reports, 14, 15901. https://doi.org/10.1038/s41598-024-66941-x
McKinney, W. (2010). Data structures for statistical computing in Python. In Proceedings of the 9th Python in Science Conference (pp. 56-61). https://doi.org/10.25080/Majora-92bf1922-00a
Merchán, D., Arora, J., Pachon, J., Konduri, K., Winkenbach, M., Parks, S., & Noszek, J. (2022). 2021 Amazon Last Mile Routing Research Challenge: Data Set. Transportation Science, 58(1), 8-11. https://doi.org/10.1287/trsc.2022.1173
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., VanderPlas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825-2830.
Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 779-788). https://doi.org/10.1109/CVPR.2016.91
Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In N. Navab, J. Hornegger, W. M. Wells, & A. Frangi (Eds.), Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015 (pp. 234-241). Springer. https://doi.org/10.1007/978-3-319-24574-4_28
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., & Fei-Fei, L. (2015). ImageNet large scale visual recognition challenge. International Journal of Computer Vision, 115, 211-252. https://doi.org/10.1007/s11263-015-0816-y
Savelsbergh, M., & Van Woensel, T. (2016). 50th anniversary invited article: City logistics: Challenges and opportunities. Transportation Science, 50(2), 579-590. https://doi.org/10.1287/trsc.2016.0675
Simchi-Levi, D., Kaminsky, P., & Simchi-Levi, E. (2007). Designing and managing the supply chain: Concepts, strategies and case studies (3rd ed.). McGraw-Hill.
Ultralytics. (2024). Package Segmentation Dataset. https://docs.ultralytics.com/datasets/segment/package-seg/
Wu, L., Wen, H., Hu, H., Mao, X., Xia, Y., Shan, E., Zheng, J., Lou, J., Liang, Y., Yang, L., Zimmermann, R., Lin, Y., & Wan, H. (2024). LaDe: The first comprehensive last-mile express dataset from industry. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. https://doi.org/10.1145/3637528.3671548
Xin, Q. (2025). Uncertainty-aware late fusion for 3D perception (confidence calibration + fusion rule learning). Journal of Technology Informatics and Engineering, 4(1), 215-238. https://doi.org/10.51903/jtie.v4i1.485
Xin, Q. (2026a). LiDAR-camera object-level fusion for multi-target tracking using JPDA and EKF: A reproducible empirical study on a PandaSet-parameterised five-sequence dataset. Journal of Technology Informatics and Engineering, 5(1), 54-76. https://doi.org/10.51903/jtie.v5i1.486
Xin, Q. (2026b). Probabilistic bike-sharing demand forecasting under changing weather and seasonal regimes with transformer-based models. Transport Findings. https://doi.org/10.32866/001c.157499
Xin, Q. (2026c). Self-supervised customer representation learning for segmentation and next-purchase prediction on UCI online retail. Journal of Information and Technology, 14(1). https://doi.org/10.32664/j-intech.v14i01.2229
Ye, T., Ma, R., & Luo, S. (2024). Vision-language traffic sign recognition for self-driving: Robust classification, uncertainty calibration, and LLM safety explanations on a GTSRB-compatible benchmark. JACS, 4(10), 84-102. https://doi.org/10.69987/JACS.2024.41007
Zhang, L., Ma, R., & Greg, P. (2025). Digital-twin dispatching for urban mobility via spatio-temporal transformers and offline reinforcement learning. Journal of Technology Informatics and Engineering, 4(2), 337-363. https://doi.org/10.51903/jtie.v4i2.501
Zhang, L., & Zhang, E. (2024). LLM-explained graph traffic forecasting for DOT corridor operations: Full empirical evaluation on METR-LA and PEMS-BAY with crash evidence. JACS, 4(5), 135-145. https://doi.org/10.69987/JACS.2024.40510
Zhao, S., Bai, J., & Roberson, D. (2025). Multi-horizon GPU demand forecasting with workload semantics and operational risk curves: An empirical study on Alibaba Clusterdata GPU trace. Journal of Technology Informatics and Engineering, 4(3), 544-571. https://doi.org/10.51903/jtie.v4i3.498
Zhong, Z. S., Ma, R., & Zhao, H. (2023a). Human-uncertainty distillation for calibrated vision models on CIFAR-10H. JACS, 3(2), 77-89. https://doi.org/10.69987/JACS.2023.30206
Zhong, Z. S., Zhang, L., & Ma, D. (2023b). Spatial RAG for urban crash hotspot discovery and safety countermeasure recommendation. JACS, 3(12), 45-61. https://doi.org/10.69987/JACS.2023.31206
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Ruiyan Ma, Long Zhang, Annie Bai

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

