Implementation and performance evaluation of a private AI cloud based on an OpenAI-compatible API in a higher education environment
Main Article Content
Abstract
The use of Large Language Models (LLMs) through public cloud services still faces challenges such as token-based cost schemes, dependence on internet connectivity, data privacy risks, and limited control over infrastructure. These conditions drive the need for a private AI cloud implementation capable of providing inference services independently within an institutional environment. This study designs and implements a private AI cloud in a higher education environment using a client–server architecture based on two virtual machines (VMs). It uses a Tesla T4 GPU to run the quantized language model Qwen3.6-28B-REAP20-A3B (Q4_K_M) via llama.cpp. At the same time, it provides an OpenAI-Compatible API service, and the Service VM runs Open WebUI as the user interface. Evaluation was conducted using a synthetic burst workload to represent worst-case conditions. The evaluation follows a partial factorial design covering 13 of 30 possible combinations. A complete concurrency sweep (1, 3, 5, 10, and 20 users) was performed for the chat load class under both parallelism configurations (--parallel 1 and --parallel 4). In comparison, the long-prompt load classes were tested at a single concurrency point (5 users). However, validation using real agent applications is outside the scope of this study. Results show that at a single processing slot configuration (--parallel 1), system throughput reached a stable plateau of around 38.5 tokens/second even though GPU utilization was only around 60%, while latency increased linearly as the number of users grew. This finding indicates that the pattern is consistent with a bottleneck arising from queue serialization rather than from GPU computational capacity limits. Activating continuous batching with four processing slots (--parallel 4) increased throughput to around 65 tokens/second (~1.7×) and reduced latency under high load by around 35%, while GPU utilization remained below maximum capacity (~58%). In contrast, long-prompt loads increased GPU utilization to 97–100% peak (60–63% mean), indicating brief GPU saturation during prefill as the primary limiting factor. All test scenarios achieved a 100% protocol success rate (600 s timeout); practically viable capacity is up to 3 concurrent users when both SLO thresholds (TTFT ≤ 2 s, RT ≤ 30 s) are applied simultaneously, or up to 10 if only the response-time criterion is considered. This study contributes architectural documentation, a validated evaluation protocol, and service capacity characterization that educational institutions with limited computing resources can replicate.
Downloads
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors retain copyright of their work and grant Jurnal Ilmiah Teknologi Informasi Asia the right of first publication. The work is simultaneously licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0) that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
While the editorial board endeavors to ensure accuracy, they accept no responsibility for the content of articles. Liability rests solely with the respective authors.
This is an open-access journal. All articles are immediately available to read and reuse upon publication under the CC BY 4.0 license. Users may share and adapt the material for any purpose, including commercial use, provided they comply with the license terms.
References
Chen, K., Zhou, X., Lin, Y., Feng, S., Shen, L., & Wu, P. (2025). A survey on privacy risks and protection in large language models. In Journal of King Saud University - Computer and Information Sciences (Vol. 37, Number 7). Springer International Publishing. https://doi.org/10.1007/s44443-025-00177-1
Chen, Q., Chen, X., & Huang, K. (2025). SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference. http://arxiv.org/abs/2507.06567
Crompton, H., Burke, D., Nickel, C., Bozkurt, A., Miao, F., Sharples, M., et al. (2026). Governing generative AI in higher education: A global Delphi study on policy and practice. International Journal of Educational Technology in Higher Education, 23(1), 21. https://doi.org/10.1186/s41239-026-00602-z
Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120), 1–39.
Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2022). GPTQ: Accurate post-training quantization for generative pre-trained transformers. arXiv. https://doi.org/10.48550/arXiv.2210.17323
Kakolyris, A. K., Masouros, D., Xydis, S., & Soudris, D. (2024). SLO-Aware GPU DVFS for Energy-Efficient LLM Inference Serving. IEEE Computer Architecture Letters, 23(2), 150–153. https://doi.org/10.1109/LCA.2024.3406038
Khalil, A., Heilles, G., Parraga, M., & Heilles, S. (2025). Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware. https://doi.org/10.48550/arXiv.2512.23029
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP '23) (pp. 611–626). ACM. https://doi.org/10.1145/3600006.3613165
Lasby, M., Lazarevich, I., Sinnadurai, N., Lie, S., Ioannou, Y., & Thangarasa, V. (2025). REAP the Experts: Why Pruning Prevails for One-Shot MoE compression. http://arxiv.org/abs/2510.13999
Malakhov, K. S. (2025). Deploying LLMs on CPU-only Environments with llama.cpp Library Set: MedLocalGPT Project Case. CEUR Workshop Proceedings (repository: https://github.com/knowledge-ukraine/medlocalgpt)
Morell-Mengual, V., Fernández-García, O., Berenguer, C., Ortega-Barón, J., Gil-Llario, M. D., & Estruch-García, V. (2025). Characteristics, motivations and attitudes of students using ChatGPT and other language model-based chatbots in higher education. Education and Information Technologies, 30(15), 22257–22274. https://doi.org/10.1007/s10639-025-13650-1
Nyamsuren, E. (2025). Evaluating quantized Large Language Models for code generation on low-resource language benchmarks. Journal of Computer Languages, 84, 101351. https://doi.org/10.1016/J.COLA.2025.101351
Rakka, M., Fouda, M. E., Khargonekar, P., & Kurdahi, F. (2024). A Review of State-of-the-art Mixed-Precision Neural Network Frameworks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12), 7793–7812. https://doi.org/10.1109/TPAMI.2024.3394390
Stuhlmann, L., Argerich, M. F., & Fürst, J. (2025). Bench360: Benchmarking Local LLM Inference from 360 Degrees. http://arxiv.org/abs/2511.16682
Wang, J., Du, H., Niyato, D., Kang, J., Xiong, Z., Kim, D. I., & Letaief, K. B. (2025). Toward Scalable Generative Ai via Mixture of Experts in Mobile Edge Networks. IEEE Wireless Communications, 32(1), 142–149. https://doi.org/10.1109/MWC.003.2400046
Wang, Z., Li, S., Zhou, Y., Li, X., Gu, R., Cam-Tu, N., Tian, C., & Zhong, S. (2024). Revisiting SLO and goodput metrics in LLM serving. arXiv. https://doi.org/10.48550/arXiv.2410.14257
Wu, B., Zhong, Y., Zhang, Z., Liu, S., Liu, F., Sun, Y., Huang, G., Liu, X., & Jin, X. (2026). FastServe: Iteration-level preemptive scheduling for large language model inference. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26) (pp. 57–74). USENIX Association.
Xiang, Y., Li, X., Qian, K., Zhang, Y., Yu, W., Zhai, E., Jin, X., & Zhou, J. (2026). ServeGen: Workload characterization and generation of large language model serving in production. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26) (pp. 1845–1859). USENIX Association.
Yan, B., Li, K., Xu, M., Dong, Y., Zhang, Y., Ren, Z., & Cheng, X. (2025). On protecting the data privacy of Large Language Models (LLMs) and LLM agents: A literature review. In High-Confidence Computing (Vol. 5, Number 2). Shandong University. https://doi.org/10.1016/j.hcc.2025.100300
Yigci, D., Eryilmaz, M., Yetisen, A. K., Tasoglu, S., & Ozcan, A. (2025). Large Language Model-Based Chatbots in Higher Education. In Advanced Intelligent Systems (Vol. 7, Number 3). John Wiley and Sons Inc. https://doi.org/10.1002/aisy.202400429
Yu, G. I., Jeong, J. S., Kim, G. W., Kim, S., & Chun, B. G. (2022). Orca: A distributed serving system for Transformer-based generative models. In Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) (pp. 521–538). USENIX Association.
Zhang, Z. J., Shi, J., & Tang, S. (2025). Cloud or On-Premises? A Strategic View of Large Language Model Deployment. https://doi.org/10.2139/ssrn.5296479
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., & Zhang, H. (2024). DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) (pp. 193–210). USENIX Association.
Zhou, Z., Ning, X., Hong, K., Fu, T., Xu, J., Li, S., Lou, Y., Wang, L., Yuan, Z., Li, X., Yan, S., Dai, G., Zhang, X.-P., Dong, Y., & Wang, Y. (2024). A survey on efficient inference for large language models. arXiv. https://doi.org/10.48550/arXiv.2404.14294
Zhu, X., Li, J., Liu, Y., Ma, C., & Wang, W. (2024). A Survey on Model Compression for Large Language Models. Transactions of the Association for Computational Linguistics, 12, 1556–1577. https://doi.org/10.1162/tacl_a_00704