[1]
2026. Token-Burst-Aware Capacity Planning for LLM Inference Services: Request Arrival, Token Demand, and Failure Risk Modeling from BurstGPT Traces. Journal of Technology Informatics and Engineering. 5, 2 (Aug. 2026), 142–164. DOI:https://doi.org/10.51903/jtie.v5i2.565.