|
DOI: 10.14489/vkit.2026.06.pp.027-035
Юрлов П. Я., Сгибнев И. В., Намнанов А. Б., Кузьминов В. А., Штраух А. С., Кутырев Я. И., Визильтер Ю. В. МЕТОД ОЦЕНКИ ТЕХНИЧЕСКИХ ТРЕБОВАНИЙ ПО ЗАДАННЫМ КРИТЕРИЯМ С ПОМОЩЬЮ БОЛЬШИХ ЯЗЫКОВЫХ МОДЕЛЕЙ (c. 27-35)
Аннотация. Посвящена проблеме оценки качества технических требований в авиационной отрасли, где критически важны корректность и однозначность формулировок. Предложен метод использования больших языковых моделей для автоматизированной оценки требований по критериям однозначности, корректности и единичности. Разработанный метод промптинга представляет собой модификацию подхода Plan-and-Solve, дополненную механизмом самокритики. Алгоритм включает четыре этапа, выполняемых в рамках одного вызова модели: составление плана, пошаговое выполнение, критика собственного решения и его улучшение. Экспериментальная проверка проводилась на специализированном наборе данных авиационных требований. Результаты исследования показали, что внедрение этапа самокритики обеспечивает прирост качества классификации (метрики F1) по всем рассматриваемым критериям по сравнению с базовым методом. Представлены ограничения текущего подхода. Указаны пути дальнейшего совершенствования метода, включая разделение этапов на независимые вызовы модели и интеграцию предметных знаний.
Ключевые слова: анализ требований; обработка естественного языка; нейронные сети; большие языковые модели; NLP; LLM.
Yurlov P. Ya., Sgibnev I. V., Namnanov A. B., Kuzminov V. A., Shtraukh A. S., Kutyrev Ya. I., Vizilter Yu. V. A CRITERIA-BASED EVALUATION METHOD OF TECHNICAL REQUIREMENTS WITH LARGE LANGUAGE MODELS (pp. 27-35)
Abstract. Technical requirements are a key element in the development of safety-critical systems where their quality directly affects system certification and reliability. However, the automation of requirement verification is hindered by the use of natural language formulations and a lack of large annotated datasets in specialised domains. Recent advances in LLMs offer a promising solution by enabling complex language understanding without task-specific fine-tuning. This paper proposes a criteria-based method for evaluating technical requirements using LLMs that combines planning and self-critique. The approach extends the Plan-and-Solve prompting paradigm by adding a reflection stage in which the model reviews and refines its own solutions. The method is applied to requirement classification according to three criteria: clarity, correctness and atomicity. Experiments are conducted on a dataset of aviation requirements. Using the GPT-4.1-nano model, the proposed method is compared with a baseline Plan-and-Solve approach and evaluated using the F1 metric. The results show that incorporating self-critique yields an improvement in classification performance across all criteria. The findings indicate that self-reflective prompting can enhance the performance of LLMs in requirement evaluation tasks without additional training data. At the same time, the performance remains modest which highlights the inherent complexity of the problem. Future work needs to focus on advanced prompting architectures and deeper integration of domain knowledge.
Keywords: Requirements analysis; Natural language processing; Neural networks; Large language models; NLP; LLM.
П. Я. Юрлов, И. В. Сгибнев, А. Б. Намнанов, В. А. Кузьминов, А. С. Штраух, Я. И. Кутырев, Ю. В. Визильтер (Государственный научно-исследовательский институт авиационных систем, Москва, Россия) E-mail:
Этот e-mail адрес защищен от спам-ботов, для его просмотра у Вас должен быть включен Javascript
P. Ya. Yurlov, I. V. Sgibnev, A. B. Namnanov, V. A. Kuzminov, A. S. Shtraukh, Ya. I. Kutyrev, Yu. V. Vizilter (State Research Institute of Aviation Systems, Moscow, Russia) E-mail:
Этот e-mail адрес защищен от спам-ботов, для его просмотра у Вас должен быть включен Javascript
1. Управление требованиями в аэрокосмической отрасли // Visure. URL: https://visuresolutions.com/ru/авиакосмическая-промышленность-и-оборона/управление-требованиями (дата обращения: 01.12.2025). 2. Cimatti A., Roveri M., Susi A., Tonetta S. Formalization and validation of safety-critical requirements // arXiv preprint arXiv:1003.1741. 2010. 3. Writing Effective Requirements Specifications // NASA SATC. URL: https://www.csc.kth.se/utbildning/kth/kurser/DD1363/NASARequirements.html (дата обращения: 01.12.2025). 4. Rangineni S. An analysis of data quality requirements for machine learning development pipelines frameworks // International Journal of Computer Trends and Technology. 2023. Т. 71. №. 9. С. 16–27. 5. Finetuned language models are zero-shot learners / J. Wei, M. Bosma, V. Zhao et al. // arXiv preprint arXiv:2109.01652. 2021. 6. Boehm B. W. Software engineering economics // IEEE transactions on Software Engineering. 1984. No. 1. P. 4–21. 7. RTCA DO-178C. Software Considerations in Airborne Systems and Equipment Certification. Radio Technical Commission for Aeronautics (RTCA), Inc., 2011. 8. SAE ARP4754A. Guidelines for Development of Civil Aircraft and Systems. SAE International, 2010. 9. ISO/IEC/IEEE 29148:2018. Systems and software engineering – Life cycle processes – Requirements engineering. Geneva, Swizeland: ISO /IEC/. IEEE, 2018. 10. International Councilon Systems Engineering (INCOSE). Guide for Writing Requirements. INCOSE-TP-2010-006-04. San Diego, CA, USA: International Council on Systems Engineering, 2019. 11. Fagan M. E. Design and code inspections to reduce errors in program development //IBM Systems Journal. 1999. V. 38. No. 2.3. P. 258–287. 12. Cheng B. H. C., Atlee J. M. Research directions in requirements engineering // Future of software engineering (FOSE'07). 23–25 May 2007. Minneapolis, Minnesota, USA. P. 285–303. 13. Mavin A., Wilkinson P., Harwood A., Novak M. Easy Approach to Requirements Syntax (EARS) // Proceedings of the 17th IEEE International Requirements Engineering Conference, 31 Aug – 4 September 2009. Atlanta, GA, USA. IEEE, 2009. P. 317–322. 14. Detecting terminological ambiguity in user stories: Tool and experimentation / F. Dalpiaz, I. van der Schalk, S. Brinkkemper et al. //Information and Software Technology. 2019. V. 110. P. 3–16. 15. Machine learning for requirements engineering (ML4RE): A systematic literature review complemented by practitioners’ voices from Stack Overflow / T. Li, X. Zhang, Y. Wang, Q. Zhou et al. // Information and Software Technology. 2024. V. 172. Art. 107477. 16. Natural language processing for requirements engineering: A systematic mapping study / L. Zhao, W. Alhoshan, A. Ferrari et al. // ACM Computing Surveys (CSUR). 2021. V. 54. No. 3. P. 1–41. 17. Attention is all you need / A. Vaswani, N. Shazeer, N. Parmar et al. // Advances in neural information processing systems (NeurIPS 2017). 4–9 December 2017. LongBearch, CA, USA. 18. Devlin J., Chang M.-W., Lee K., Toutenova K. Bert: Pre-training of deep bidirectional transformers for language understanding // Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 2–7 June 2019. Minneapolis, Minnesota, USA. P. 4171–4186. 19. Анализ текстов на естественном языке в задаче оценки требований к проектированию конструкции авиационного бортового оборудования / Б. В. Вишняков, И. В. Сгибнев, Ю. А. Солоделов и др. // Вестник компьютерных и информационных технологий. 2024. Т. 21, № 4. C. 17–29. DOI: 10.14489/vkit.2024.04.pp.017-029 20. A survey of large language models / W. X. Zhao, K. Zhou, J. Li et al. // arXiv preprint arXiv:2303.18223. 2023. V. 1. No. 2. 21. Language models are few-shot learners / T. Brown, B. Mann, N. Ryder et al. // Advances in neural information processing systems (NeurIPS 2020). V. 33. P. 1877–1901. 22. Emergent abilities of large language models / J. Wei, Y. Tay, R. Bommasani et al. // arXiv preprint arXiv:2206.07682. –2022. 23. Domain specialization as the key to make large language models disruptive: A comprehensive survey / C. Ling, X. Zhao, J. Lu et al. //ACM Computing Surveys. 2025. V. 58. No. 3. P. 1–39. 24. Chain-of-thought prompting elicits reasoning in large language models / J. Wei, X. Wang, D. Sohuurmans et al. // Advances in neural information processing systems (NeurIPS 2022). 28 November – 9 December 2022. New Orleans, Louisiana, USA. P. 24824–24837. 25. Least-to-most prompting enables complex reasoning in large language models / D. Zhou, N. Scharli, L. Hon et al. // arXiv preprint arXiv:2205.10625. 2022. 26. Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models / L. Wang, W. Xu, Y. Lan et al. // arXiv preprint arXiv:2305.04091. 2023. 27. Self-refine: Iterative refinement with self-feed¬back / A. Madaan, N. Tandon, P. Gupta et al. // Advances in Neural Information Processing Systems (NeurIPS 2023). 10–16 December 2023. New Orleans, Louisiane, USA. P. 46534–46594. 28. Reflexion: Language agents with verbal reinforcement learning / N. Shinn, F. Cassano, A. Gopinath et al. // Advances in Neural Information Processing Systems (NeurIPS 2023). 10–16 December 2023. New Orleans, Louisiana, USA. P. 8634-8652. 29. GPT-4.1-nano // OpenAI Platform. URL: https://platform.openai.com/docs/models/gpt-4.1-nano (дата обращения: 01.12.2025).
1. Visure. (n.d.). Upravlenie trebovaniyami v aerokosmicheskoi otrasli [Requirements management in the aerospace industry]. Retrieved December 1, 2025, from https://visuresolutions.com/ru/авиакосмическая-промышленность-и-оборона/управление-требованиями. [in Russian language]. 2. Cimatti, A., Roveri, M., Susi, A., & Tonetta, S. (2010). Formalization and validation of safety critical requirements. arXiv preprint. arXiv:1003.1741 3. NASA SATC. (n.d.). Writing effective requirements specifications. Retrieved December 1, 2025, from https://www.csc.kth.se/utbildning/kth/kurser/DD1363/NASARequirements.html 4. Rangineni, S. (2023). An analysis of data quality requirements for machine learning development pipelines frameworks. International Journal of Computer Trends and Technology, 71(9), 16–27. 5. Wei, J., Bosma, M., Zhao, V., et al. (2021). Finetuned language models are zero shot learners. arXiv preprint. arXiv:2109.01652 6. Boehm, B. W. (1984). Software engineering economics. IEEE Transactions on Software Engineering, (1), 4–21. 7. RTCA. (2011). Software considerations in airborne systems and equipment certification (RTCA DO 178C). Radio Technical Commission for Aeronautics. 8. SAE International. (2010). Guidelines for development of civil aircraft and systems (SAE ARP4754A). 9. ISO/IEC/IEEE. (2018). Systems and software engineering – Life cycle processes – Requirements engineering (ISO/IEC/IEEE 29148:2018). International Organization for Standardization. 10. International Council on Systems Engineering (INCOSE). (2019). Guide for writing requirements (INCOSE TP 2010 006 04). INCOSE. 11. Fagan, M. E. (1999). Design and code inspections to reduce errors in program development. IBM Systems Journal, 38(2.3), 258–287. 12. Cheng, B. H. C., & Atlee, J. M. (2007, May 23–25). Research directions in requirements engineering. In Future of Software Engineering (FOSE’07) (pp. 285–303). Minneapolis, MN, United States. 13. Mavin, A., Wilkinson, P., Harwood, A., & Novak, M. (2009, August 31 – September 4). Easy approach to requirements syntax (EARS). In Proceedings of the 17th IEEE International Requirements Engineering Conference (pp. 317–322). Atlanta, GA, United States. IEEE. 14. Dalpiaz, F., van der Schalk, I., Brinkkemper, S., et al. (2019). Detecting terminological ambiguity in user stories: Tool and experimentation. Information and Software Technology, 110, 3–16. 15. Li, T., Zhang, X., Wang, Y., & Zhou, Q. (2024). Machine learning for requirements engineering (ML4RE): A systematic literature review complemented by practitioners’ voices from Stack Overflow. Information and Software Technology, 172, Article 107477. 16. Zhao, L., Alhoshan, W., Ferrari, A., et al. (2021). Natural language processing for requirements engineering: A systematic mapping study. ACM Computing Surveys, 54(3), 1–41. 17. Vaswani, A., Shazeer, N., Parmar, N., et al. (2017, December 4–9). Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS 2017). Long Beach, CA, United States. 18. Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019, June 2–7). BERT: Pre training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171–4186). Minneapolis, MN, United States. 19. Vishnyakov, B. V., Sgibnev, I. V., Solodelov, Yu. A., et al. (2024). Analysis of natural language texts in the problem of evaluating design requirements for aircraft onboard equipment. Vestnik komp'yuternykh i informatsionnykh tekhnologii, 21(4), 17–29. [in Russian language]. https://doi.org/10.14489/vkit.2024.04.pp.017-029 20. Zhao, W. X., Zhou, K., Li, J., et al. (2023). A survey of large language models. arXiv preprint. arXiv:2303.18223 21. Brown, T., Mann, B., Ryder, N., et al. (2020). Language models are few shot learners. In Advances in Neural Information Processing Systems (NeurIPS 2020) (Vol. 33, pp. 1877–1901). 22. Wei, J., Tay, Y., Bommasani, R., et al. (2022). Emergent abilities of large language models. arXiv preprint. arXiv:2206.07682 23. Ling, C., Zhao, X., Lu, J., et al. (2025). Domain specialization as the key to make large language models disruptive: A comprehensive survey. ACM Computing Surveys, 58(3), 1–39. 24. Wei, J., Wang, X., Schuurmans, D., et al. (2022, November 28 – December 9). Chain of thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (NeurIPS 2022) (pp. 24824–24837). New Orleans, LA, United States. 25. Zhou, D., Schärli, N., Hou, L., et al. (2022). Least to most prompting enables complex reasoning in large language models. arXiv preprint. arXiv:2205.10625 26. Wang, L., Xu, W., Lan, Y., et al. (2023). Plan and solve prompting: Improving zero shot chain of thought reasoning by large language models. arXiv preprint. arXiv:2305.04091 27. Madaan, A., Tandon, N., Gupta, P., et al. (2023, December 10–16). Self refine: Iterative refinement with self feedback. In Advances in Neural Information Processing Systems (NeurIPS 2023) (pp. 46534–46594). New Orleans, LA, United States. 28. Shinn, N., Cassano, F., Gopinath, A., et al. (2023, December 10–16). Reflexion: Language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS 2023) (pp. 8634–8652). New Orleans, LA, United States. 29. OpenAI. (n.d.). GPT 4.1 nano. OpenAI Platform. Retrieved December 1, 2025, from https://platform.openai.com/docs/models/gpt-4.1-nano
Статью можно приобрести в электронном виде (PDF формат).
Стоимость статьи 700 руб. (в том числе НДС 20%). После оформления заказа, в течение нескольких дней, на указанный вами e-mail придут счет и квитанция для оплаты в банке.
После поступления денег на счет издательства, вам будет выслан электронный вариант статьи.
Для заказа скопируйте doi статьи:
10.14489/vkit.2026.06.pp.027-035
и заполните форму
Отправляя форму вы даете согласие на обработку персональных данных.
.
This article is available in electronic format (PDF).
The cost of a single article is 700 rubles. (including VAT 20%). After you place an order within a few days, you will receive following documents to your specified e-mail: account on payment and receipt to pay in the bank.
After depositing your payment on our bank account we send you file of the article by e-mail.
To order articles please copy the article doi:
10.14489/vkit.2026.06.pp.027-035
and fill out the form
.
|