Small Models as Control Planes for Enterprise GenAI

Authors

  • Carlos Mendoza Author
    Competing Interests

    AI,ML

DOI:

https://doi.org/10.5281/zenodo.22173002

Keywords:

Generative AI Orchestration, Enterprise GenAI Control Planes, Cost-Efficient Model Serving, Small Language Models (SLMs), Large Language Model Optimization, Model Sparsity and Quantization, GenAI Workload Characterization, Intelligent Request Routing, Model Caching Strategies, Personalized GenAI Services, Enterprise AI Cost Governance, Control-Plane Decision Logic, Distributed GenAI Architectures, Model Selection and Policy Enforcement, AI-Driven Service Digitalization, Adaptive Model Configuration, Execution Cost Optimization, Latent Enterprise GenAI Use Cases, Scalable GenAI Infrastructure, Next-Generation AI Orchestration Frameworks.

Abstract

Generative AI applications such as chatbots and text-to-image systems create demand for GenAI models that address the common information needs of diverse stakeholders in a responsive and personalized manner. Yet the effort to train, host, and serve these models can incur substantial cost and complexity. Many large language models can answer a wide range of questions, but risk being underutilized for specific enterprise workloads. At the same time, enterprise-integrated GenAI services are supporting the digitalization of business processes at an unprecedented scale, revealing latent use cases for specialized models or adjusted configurations of the same model that reflect the cost profiles of these systems. These factors suggest that deploying smaller models to manage the GenAI orchestration layer across an enterprise might yield significant cost savings. A control plane design based on the concept of GenAI orchestration is proposed, along with a set of cost-efficiency principles for implementing this functionality.

Control planes are responsible for decision-making and policy enforcement across a distributed system. Making cost-effectiveness an explicit design goal when architecting an orchestration layer introduces additional considerations beyond those that typically inform the design of control planes. Model size, sparsity, quantization, caching, and workload characterization shape the trade-offs governing the overall cost of model execution, create opportunities for realizing cost savings, and identify workload patterns that can further inform cost-saving measures.

References

1. Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120), 1–39.

2. Lester, B., Al-Rfou, R., & Constant, N. (2021). The power of scale for parameter-efficient prompt tuning. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 3045–3059.

3. Ganesh, S. K., Subbareddy, K., Gopi, A., Davuluri, P. N., Mannar, B. R., & Karnawat, A. T. (2026, June). Fraudulent Credit Card Transaction Detection Using a Hybrid Ensemble-Anomaly Detection Framework. In 2026 6th International Conference on Intelligent Technologies (CONIT) (pp. 1-6). IEEE.

4. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. International Conference on Learning Representations.

5. Li, Y., Wei, F., Zhang, J., & Zhang, X. (2023). ELLA-V: Stable neural network adaptation for efficient language model serving. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 1–15.

6. Yandamuri, U. S., Loganathan, R., Davuluri, P. S. L. N., Rani, P. R. S., Kolla, S. H., & Nagubandi, A. R. (2026). Adaptive Intelligence Networks for Humancentered Enterprise Automation and Governance. In 2026 Third International Conference on Innovations in Cybersecurity and Data Science (ICICDS) (pp. 678–683). IEEE. 2026 Third International Conference on Innovations in Cybersecurity and Data Science (ICICDS). https://doi.org/10.1109/icicds70526.2026.11604640

7. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. Advances in Neural Information Processing Systems, 36, 10088–10115.

8. Chen, L., Zaharia, M., & Zou, J. (2023). FrugalGPT: How to use large language models while reducing cost and improving performance. arXiv preprint arXiv:2305.05176.

9. Banoth, S., Santoshi Kumari, M., Lalitha, P., Deepthi, P., Kumar Peddi, R., & Kavitha, P. (2026). An Efficient Hybrid K-Means and Random Forest-Based Approach for Cloud Malware Detection and Privacy Protection. In 2026 6th International Conference on Intelligent Technologies (CONIT) (pp. 1–6). IEEE. 2026 6th International Conference on Intelligent Technologies (CONIT). https://doi.org/10.1109/conit69683.2026.11621822

10. Jiang, D., Ren, X., & Lin, B. Y. (2023). LLM-Blender: Ensembling large language models with pairwise ranking and generative fusion. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, 14165–14178.

11. Mangalampalli, B. M., Peddi, R. K., Kolla, S. K., Reddy, V. A. R., Mangala, N., & Seenu, A. (2026, June). Explainable Clinical Graph Intelligence Framework for Longitudinal Risk Modeling and Care Pathway Optimization. In 2026 Third International Conference on Innovations in Cybersecurity and Data Science (ICICDS) (pp. 1-6). IEEE.

12. Leviathan, Y., Kalman, M., & Matias, Y. (2023). Fast inference from transformers via speculative decoding. Proceedings of the 40th International Conference on Machine Learning, 19274–19286.

13. Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient memory management for large language model serving with PagedAttention. Proceedings of the 29th Symposium on Operating Systems Principles, 611–626.

14. Mahadevan, S. (2026). Governed Serverless Automation Ecosystem for AI-Driven Enterprise Integration and Sustainable Cloud Operations. International Journal of Computer Information Systems and Industrial Management Applications, 18(16s), 1561-1570.

15. Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36, 68539–68551.

16. Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., et al. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.

17. Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., et al. (2023). Mistral 7B. arXiv preprint arXiv:2310.06825.

18. Vakkalagadda, T., & Rajamanickam, V. (2026). Agentic AI Architectures for Next-Generation Investment Advisory and Portfolio Intelligence. International Journal of Computer Information Systems and Industrial Management Applications, 18(9s), 1206-1219.

19. Ding, D., Mallick, A., Wang, C., Sim, R., Mukherjee, S., Ruehle, V., Lakshmanan, L. V. S., & Awadallah, A. (2024). Hybrid LLM: Cost-efficient and quality-aware query routing. International Conference on Learning Representations.

20. Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., et al. (2024). Mixtral of experts. arXiv preprint arXiv:2401.04088.

21. Abdin, M., Jacobs, S. A., Awan, A. A., Aneja, J., Awadalla, H., Awadallah, A., et al. (2024). Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219.

22. Bhavani, B. D., SR, S., Loganathan, R., & Nagaraj, S. (2026, April). Evolutionary Gravitational Neocognitron Neural Network, Snow Leopard Optimization and Deep Graph Reinforcement Learning for Routing Protocol in WSN. In 2026 2nd International Conference on Intelligent Systems and Computational Networks (ICISCN) (pp. 1-8). IEEE.

23. Mei, K., Zhu, X., Xu, W., Hua, W., Jin, M., Li, Z., Xu, S., Ye, R., Ge, Y., & Zhang, Y. (2024). AIOS: LLM agent operating system. arXiv preprint arXiv:2403.16971.

24. Zhuang, N., Tao, M., Zhang, C., Jin, Y., Xu, K., Chen, L., Huang, S., & Feng, Y. (2024). Harder task needs more experts: Dynamic routing in MoE models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 12883–12895.

25. Piyush Kumar Pareek. (2026). Human-Centric Machine Learning Frameworks for Scalable Software Quality Prediction. Journal of Intelligent Decision Making and Information Science, 3(5s), 1525–1543. https://doi.org/10.59543/jidmis.v3.1284

26. Araujo, V., Moens, M.-F., & Tuytelaars, T. (2024). Learning to route for dynamic adapter composition in continual learning with language models. Findings of the Association for Computational Linguistics: EMNLP 2024, 687–696.

27. Ge, Y., Ren, Y., Hua, W., Xu, S., Tan, J., & Zhang, Y. (2023). LLM as OS, agents as apps: Envisioning AIOS, agents and the AIOS-agent ecosystem. arXiv preprint arXiv:2312.03815.

28. Loganathan, R. (2026). SaaS Entitlement Governance: A Reproducible Model for Detecting and Reclaiming Over-Provisioned Access at Enterprise Scale. Journal of Intelligent Decision Making and Information Science, 3(7s), 2519-2530.

29. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. International Conference on Learning Representations.

30. Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Hambro, E., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36, 68539–68551.

31. Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, J., et al. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18, 186345.

32. Davuluri, P. S. L. (2023). AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems. AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems (December 15, 2023).

33. Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., et al. (2023). The rise and potential of large language model based agents: A survey. arXiv preprint arXiv:2309.07864.

34. Wang, Y., Zeng, Z., Li, Y., Huang, Z., Chen, Z., & others. (2024). A survey on large language model based autonomous agents. ACM Computing Surveys, 57(6), 1–45.

35. Behera, A. P., Champati, J. P., Morabito, R., Tarkoma, S., & Gross, J. (2025). Towards efficient multi-LLM inference: Characterization and analysis of LLM routing and hierarchical techniques. arXiv preprint arXiv:2506.06579.

36. She, J., Zheng, W., Liu, Z., Wang, H., Xing, E. P., Yao, H., & Ho, Q. (2025). Token level routing inference system for edge devices. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 1–10.

37. Ahuja, K., Kumar, S., Goyal, N., & others. (2025). Adaptive LLM routing under budget constraints. Findings of the Association for Computational Linguistics: EMNLP 2025, 11320–11340.

38. Zhang, D., Song, J., Bi, Z., Song, X., Yuan, Y., Wang, T., Yeong, J., & Hao, J. (2025). Mixture of experts in large language models. arXiv preprint arXiv:2507.11181.

39. Johnson, W., & Lee, C. (2026). Evaluating small language models for front-door routing: A harmonized benchmark and synthetic-traffic experiment. arXiv preprint arXiv:2604.02367.

40. Piyush Kumar Pareek. (2026). Self-Evolving Analytics Pipelines for Reliable AI-Augmented Software Systems. Journal of Intelligent Decision Making and Information Science, 3(5s), 1558–1578. https://doi.org/10.59543/jidmis.v3.1286

41. Li, J. (2026). The evolution of mixture-of-experts architectures in large language models: Routing, topology, load balancing, and expert parallelism. arXiv preprint arXiv:2608.08650.

Additional Files

Published

2026-03-07

Data Availability Statement

None

How to Cite

Small Models as Control Planes for Enterprise GenAI. (2026). The American Journal of Analytics and Artificial Intelligence (AJAAI), 4(01). https://doi.org/10.5281/zenodo.22173002

Similar Articles

11-20 of 39

You may also start an advanced similarity search for this article.