Edit-Signal Prompt Optimization: A Production Method for Continuous Quality Improvement in Industrial LLM Systems

Authors

  • Agzamkhodjaev Saydolimkhon Nodirovich

Keywords:

Large language models, LLM-as-a-Judge, software quality assurance, MLOps, industrial AI systems, deterministic validation, human-centered AI, evaluation metrics, automated correction, simulation

Abstract

Deploying large language models (LLMs) in production business systems exposes a critical reliability gap: generated outputs frequently contain structural errors, logical inconsistencies, and hallucinations that erode user trust and block downstream operations. Existing remediation approaches, RLHF, manual prompt engineering, and offline evaluation, are either prohibitively expensive, fail to scale, or do not operate on live production traffic. This paper introduces Edit-Signal Prompt Optimization (ESPO), a closed-loop prompt refinement method in which (1) production user edits are systematically aggregated into prompt-level corrections without model fine-tuning, (2) LLM-judge rationales are converted into targeted rewrite directives rather than used solely as pass/fail signals, and (3) these two feedback streams are unified into a continuous improvement cycle operating on live production traffic. ESPO is embedded within a multi-level evaluation architecture combining deterministic validation, LLM-as-a-Judge semantic assessment, and human-in-the-loop feedback. The method was designed, deployed, and validated at Treater, Inc., an AI-powered retail execution platform processing communications across 24,000+ retail store locations. Over an eight-week evaluation window encompassing approximately 15,000 generated reports, ESPO reduced critical errors by 40% and decreased user edit rates by 34%, demonstrating continuous improvement without model retraining. To the author’s knowledge, this is the first integration and productionization of edit-derived prompt optimization as a systematic method for industrial LLM systems.

Author Biography

  • Agzamkhodjaev Saydolimkhon Nodirovich

    Founding Engineer, Treater, Inc, New York, USA

References

[1] Gartner. (2023, October 11). Gartner says more than 80% of enterprises will have used generative AI APIs or deployed generative AI-enabled applications by 2026. Retrieved from: https://www.gartner.com/en/newsroom/press-releases/2023-10-11-gartner-says-more-than-80-percent-of-enterprises-will-have-used-generative-ai-apis-or-deployed-generative-ai-enabled-applications-by-2026 (date accessed: November 14, 2025).

[2] Zhang, T., Ladhak, F., Durmus, E., Liang, P., McKeown, K., & Hashimoto, T. B. (2024). Benchmarking large language models for news summarization. Transactions of the Association for Computational Linguistics, 12, 39–57. https://doi.org/10.1162/tacl_a_00632.

[3] Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023). G-Eval: NLG evaluation using GPT-4 with better human alignment. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 2511–2522). https://doi.org/10.18653/v1/2023.emnlp-main.153.

[4] OpenAI. (2023). GPT-4 technical report. arXiv. https://doi.org/10.48550/arXiv.2303.08774.

[5] Rudd, E. M., Andrews, C., & Tully, P. (2025). A practical guide for evaluating LLMs and LLM-reliant systems. arXiv. https://doi.org/10.48550/arXiv.2506.13023.

[6] Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., & Hashimoto, T. B. (2023). AlpacaFarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Systems, 36, 30039–30069. https://doi.org/10.48550/arXiv.2305.14387.

[7] Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837. https://doi.org/10.48550/arXiv.2201.11903.

[8] Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. arXiv. https://doi.org/10.48550/arXiv.2306.05685.

[9] Chen, L., Zaharia, M., & Zou, J. (2023). FrugalGPT: How to use large language models while reducing cost and improving performance. arXiv. https://doi.org/10.48550/arXiv.2305.05176.

[10] Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744. https://doi.org/10.48550/arXiv.2203.02155.

[11] Khattab, O., Singhvi, A., Maheshwari, P., Zhang, Z., Santhanam, K., Vardhamanan, S., Haq, S., Sharma, A., Joshi, T. T., Moazam, H., Miller, H., Zaharia, M., & Potts, C. (2023). DSPy: Compiling declarative language model calls into self-improving pipelines. arXiv. https://doi.org/10.48550/arXiv.2310.03714.

[12] Es, S., James, J., Espinosa-Anke, L., & Schockaert, S. (2024). RAGAs: Automated evaluation of retrieval augmented generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations (pp. 150–158). https://doi.org/10.18653/v1/2024.eacl-demo.16.

[13] Confident AI. (2024). DeepEval: The LLM evaluation framework. Retrieved from: https://github.com/confident-ai/deepeval (date accessed: January 9, 2026).

[14] TruEra. (2024). TruLens: Evaluation and tracking for LLM experiments and AI agents. Retrieved from: https://github.com/truera/trulens (date accessed: February 18, 2026).

Downloads

Published

2026-07-31

Issue

Section

Articles

How to Cite

Agzamkhodjaev Saydolimkhon Nodirovich. (2026). Edit-Signal Prompt Optimization: A Production Method for Continuous Quality Improvement in Industrial LLM Systems. American Scientific Research Journal for Engineering, Technology, and Sciences, 104(1), 264-273. https://asrjetsjournal.org/American_Scientific_Journal/article/view/12292