Implementation and performance evaluation of a private AI cloud based on an OpenAI-compatible API in a higher education environment

Main Article Content

Mabrur Roh Bintang Jaya
Lia Farokhah

Abstract

The use of Large Language Models (LLMs) through public cloud services still faces challenges such as token-based cost schemes, dependence on internet connectivity, data privacy risks, and limited control over infrastructure. These conditions drive the need for a private AI cloud implementation capable of providing inference services independently within an institutional environment. This study designs and implements a private AI cloud in a higher education environment using a client–server architecture based on two virtual machines (VMs). It uses a Tesla T4 GPU to run the quantized language model Qwen3.6-28B-REAP20-A3B (Q4_K_M) via llama.cpp. At the same time, it provides an OpenAI-Compatible API service, and the Service VM runs Open WebUI as the user interface. Evaluation was conducted using a synthetic burst workload to represent worst-case conditions. The evaluation follows a partial factorial design covering 13 of 30 possible combinations. A complete concurrency sweep (1, 3, 5, 10, and 20 users) was performed for the chat load class under both parallelism configurations (--parallel 1 and --parallel 4). In comparison, the long-prompt load classes were tested at a single concurrency point (5 users). However, validation using real agent applications is outside the scope of this study. Results show that at a single processing slot configuration (--parallel 1), system throughput reached a stable plateau of around 38.5 tokens/second even though GPU utilization was only around 60%, while latency increased linearly as the number of users grew. This finding indicates that the pattern is consistent with a bottleneck arising from queue serialization rather than from GPU computational capacity limits. Activating continuous batching with four processing slots (--parallel 4) increased throughput to around 65 tokens/second (~1.7×) and reduced latency under high load by around 35%, while GPU utilization remained below maximum capacity (~58%). In contrast, long-prompt loads increased GPU utilization to 97–100% peak (60–63% mean), indicating brief GPU saturation during prefill as the primary limiting factor. All test scenarios achieved a 100% protocol success rate (600 s timeout); practically viable capacity is up to 3 concurrent users when both SLO thresholds (TTFT ≤ 2 s, RT ≤ 30 s) are applied simultaneously, or up to 10 if only the response-time criterion is considered. This study contributes architectural documentation, a validated evaluation protocol, and service capacity characterization that educational institutions with limited computing resources can replicate.

Downloads

Download data is not yet available.

Article Details

How to Cite
Jaya, M. R. B., & Farokhah, L. (2026). Implementation and performance evaluation of a private AI cloud based on an OpenAI-compatible API in a higher education environment. Jurnal Ilmiah Teknologi Informasi Asia, 20(2), 120–130. https://doi.org/10.32815/jitika.1288
Section
Article

References

Chen, K., Zhou, X., Lin, Y., Feng, S., Shen, L., & Wu, P. (2025). A survey on privacy risks and protection in large language models. In Journal of King Saud University - Computer and Information Sciences (Vol. 37, Number 7). Springer International Publishing. https://doi.org/10.1007/s44443-025-00177-1

Chen, Q., Chen, X., & Huang, K. (2025). SlimCaching: Edge Caching of Mixture-of-Experts for Distributed Inference. http://arxiv.org/abs/2507.06567

Crompton, H., Burke, D., Nickel, C., Bozkurt, A., Miao, F., Sharples, M., et al. (2026). Governing generative AI in higher education: A global Delphi study on policy and practice. International Journal of Educational Technology in Higher Education, 23(1), 21. https://doi.org/10.1186/s41239-026-00602-z

Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120), 1–39.

Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2022). GPTQ: Accurate post-training quantization for generative pre-trained transformers. arXiv. https://doi.org/10.48550/arXiv.2210.17323

Kakolyris, A. K., Masouros, D., Xydis, S., & Soudris, D. (2024). SLO-Aware GPU DVFS for Energy-Efficient LLM Inference Serving. IEEE Computer Architecture Letters, 23(2), 150–153. https://doi.org/10.1109/LCA.2024.3406038

Khalil, A., Heilles, G., Parraga, M., & Heilles, S. (2025). Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware. https://doi.org/10.48550/arXiv.2512.23029

Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP '23) (pp. 611–626). ACM. https://doi.org/10.1145/3600006.3613165

Lasby, M., Lazarevich, I., Sinnadurai, N., Lie, S., Ioannou, Y., & Thangarasa, V. (2025). REAP the Experts: Why Pruning Prevails for One-Shot MoE compression. http://arxiv.org/abs/2510.13999

Malakhov, K. S. (2025). Deploying LLMs on CPU-only Environments with llama.cpp Library Set: MedLocalGPT Project Case. CEUR Workshop Proceedings (repository: https://github.com/knowledge-ukraine/medlocalgpt)

Morell-Mengual, V., Fernández-García, O., Berenguer, C., Ortega-Barón, J., Gil-Llario, M. D., & Estruch-García, V. (2025). Characteristics, motivations and attitudes of students using ChatGPT and other language model-based chatbots in higher education. Education and Information Technologies, 30(15), 22257–22274. https://doi.org/10.1007/s10639-025-13650-1

Nyamsuren, E. (2025). Evaluating quantized Large Language Models for code generation on low-resource language benchmarks. Journal of Computer Languages, 84, 101351. https://doi.org/10.1016/J.COLA.2025.101351

Rakka, M., Fouda, M. E., Khargonekar, P., & Kurdahi, F. (2024). A Review of State-of-the-art Mixed-Precision Neural Network Frameworks. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12), 7793–7812. https://doi.org/10.1109/TPAMI.2024.3394390

Stuhlmann, L., Argerich, M. F., & Fürst, J. (2025). Bench360: Benchmarking Local LLM Inference from 360 Degrees. http://arxiv.org/abs/2511.16682

Wang, J., Du, H., Niyato, D., Kang, J., Xiong, Z., Kim, D. I., & Letaief, K. B. (2025). Toward Scalable Generative Ai via Mixture of Experts in Mobile Edge Networks. IEEE Wireless Communications, 32(1), 142–149. https://doi.org/10.1109/MWC.003.2400046

Wang, Z., Li, S., Zhou, Y., Li, X., Gu, R., Cam-Tu, N., Tian, C., & Zhong, S. (2024). Revisiting SLO and goodput metrics in LLM serving. arXiv. https://doi.org/10.48550/arXiv.2410.14257

Wu, B., Zhong, Y., Zhang, Z., Liu, S., Liu, F., Sun, Y., Huang, G., Liu, X., & Jin, X. (2026). FastServe: Iteration-level preemptive scheduling for large language model inference. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26) (pp. 57–74). USENIX Association.

Xiang, Y., Li, X., Qian, K., Zhang, Y., Yu, W., Zhai, E., Jin, X., & Zhou, J. (2026). ServeGen: Workload characterization and generation of large language model serving in production. In 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26) (pp. 1845–1859). USENIX Association.

Yan, B., Li, K., Xu, M., Dong, Y., Zhang, Y., Ren, Z., & Cheng, X. (2025). On protecting the data privacy of Large Language Models (LLMs) and LLM agents: A literature review. In High-Confidence Computing (Vol. 5, Number 2). Shandong University. https://doi.org/10.1016/j.hcc.2025.100300

Yigci, D., Eryilmaz, M., Yetisen, A. K., Tasoglu, S., & Ozcan, A. (2025). Large Language Model-Based Chatbots in Higher Education. In Advanced Intelligent Systems (Vol. 7, Number 3). John Wiley and Sons Inc. https://doi.org/10.1002/aisy.202400429

Yu, G. I., Jeong, J. S., Kim, G. W., Kim, S., & Chun, B. G. (2022). Orca: A distributed serving system for Transformer-based generative models. In Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) (pp. 521–538). USENIX Association.

Zhang, Z. J., Shi, J., & Tang, S. (2025). Cloud or On-Premises? A Strategic View of Large Language Model Deployment. https://doi.org/10.2139/ssrn.5296479

Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., & Zhang, H. (2024). DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) (pp. 193–210). USENIX Association.

Zhou, Z., Ning, X., Hong, K., Fu, T., Xu, J., Li, S., Lou, Y., Wang, L., Yuan, Z., Li, X., Yan, S., Dai, G., Zhang, X.-P., Dong, Y., & Wang, Y. (2024). A survey on efficient inference for large language models. arXiv. https://doi.org/10.48550/arXiv.2404.14294

Zhu, X., Li, J., Liu, Y., Ma, C., & Wang, W. (2024). A Survey on Model Compression for Large Language Models. Transactions of the Association for Computational Linguistics, 12, 1556–1577. https://doi.org/10.1162/tacl_a_00704

Similar Articles

1 2 3 4 5 6 7 8 > >> 

You may also start an advanced similarity search for this article.