<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE root>
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:ali="http://www.niso.org/schemas/ali/1.0/" article-type="research-article" dtd-version="1.2" xml:lang="ru"><front><journal-meta><journal-id journal-id-type="publisher-id">News of the Kabardino-Balkarian Scientific Center of the Russian Academy of Sciences</journal-id><journal-title-group><journal-title xml:lang="en">News of the Kabardino-Balkarian Scientific Center of the Russian Academy of Sciences</journal-title><trans-title-group xml:lang="ru"><trans-title>Известия Кабардино-Балкарского научного центра РАН</trans-title></trans-title-group></journal-title-group><issn publication-format="print">1991-6639</issn><issn publication-format="electronic">2949-1940</issn></journal-meta><article-meta><article-id pub-id-type="publisher-id">265564</article-id><article-id pub-id-type="doi">10.35330/1991-6639-2024-26-4-54-61</article-id><article-id pub-id-type="edn">KZKDOT</article-id><article-categories><subj-group subj-group-type="toc-heading" xml:lang="ru"><subject>Информатика и информационные процессы</subject></subj-group><subj-group subj-group-type="toc-heading" xml:lang="en"><subject>Informatics and information processes</subject></subj-group><subj-group subj-group-type="article-type"><subject>Research Article</subject></subj-group></article-categories><title-group><article-title xml:lang="en">A method for assessing the degree of confidence in the self-explanations of GPT models</article-title><trans-title-group xml:lang="ru"><trans-title>Метод оценки степени доверия к самообъяснениям GPT-моделей</trans-title></trans-title-group></title-group><contrib-group><contrib contrib-type="author"><name-alternatives><name xml:lang="ru"><surname>Лукьянов</surname><given-names>Андрей Николаевич</given-names></name><name xml:lang="en"><surname>Lukyanov</surname><given-names>Andrey N.</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="en"><p>Student, Research Assistant, Center for Advanced Studies in Artificial Intelligence</p></bio><bio xml:lang="ru"><p>студент,<bold> </bold>лаборант-исследователь,<bold> </bold>Центр перспективных исследований в искусственном интеллекте</p></bio><email>andreylukianovai@gmail.com</email><xref ref-type="aff" rid="aff1"/></contrib><contrib contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-4089-6580</contrib-id><contrib-id contrib-id-type="spin">8583-3592</contrib-id><name-alternatives><name xml:lang="en"><surname>Tramova</surname><given-names>Aziza M.</given-names></name><name xml:lang="ru"><surname>Трамова</surname><given-names>Азиза Мухамадияевна</given-names></name></name-alternatives><address><country country="RU">Russian Federation</country></address><bio xml:lang="ru"><p>д-р экон. наук, профессор кафедры информатики</p></bio><bio xml:lang="en"><p>Doctor of Economic Sciences, Professor of the Department of Informatics</p></bio><email>Tramova.AM@rea.ru</email><xref ref-type="aff" rid="aff1"/></contrib></contrib-group><aff-alternatives id="aff1"><aff><institution xml:lang="ru">Российский экономический университет им. Г. В. Плеханова</institution></aff><aff><institution xml:lang="en">Plekhanov Russian University of Economics</institution></aff></aff-alternatives><content-language>ru</content-language><pub-date date-type="pub" iso-8601-date="2024-10-07" publication-format="electronic"><day>07</day><month>10</month><year>2024</year></pub-date><pub-date date-type="collection"><year>2024</year></pub-date><volume>26</volume><issue>4</issue><issue-title xml:lang="en"/><issue-title xml:lang="ru"/><fpage>54</fpage><lpage>61</lpage><history><date date-type="received" iso-8601-date="2024-10-06"><day>06</day><month>10</month><year>2024</year></date><date date-type="accepted" iso-8601-date="2024-10-06"><day>06</day><month>10</month><year>2024</year></date></history><permissions><copyright-statement xml:lang="en">Copyright ©; 2024, Лукьянов А.N., Трамова А.M.</copyright-statement><copyright-statement xml:lang="ru">Copyright ©; 2024, Лукьянов А.Н., Трамова А.М.</copyright-statement><copyright-year>2024</copyright-year><copyright-holder xml:lang="en">Лукьянов А.N., Трамова А.M.</copyright-holder><copyright-holder xml:lang="ru">Лукьянов А.Н., Трамова А.М.</copyright-holder><ali:free_to_read xmlns:ali="http://www.niso.org/schemas/ali/1.0/"/><license><ali:license_ref xmlns:ali="http://www.niso.org/schemas/ali/1.0/">https://creativecommons.org/licenses/by/4.0</ali:license_ref></license></permissions><self-uri xlink:href="https://journals.rcsi.science/1991-6639/article/view/265564">https://journals.rcsi.science/1991-6639/article/view/265564</self-uri><abstract xml:lang="en"><p>With the rapid growth in the use of generative neural network models for practical tasks, the problem of explaining their decisions is becoming increasingly acute. As neural network-based solutions are being introduced into medical practice, government administration, and defense, the demands for interpretability of such systems will undoubtedly increase. In this study, we aim to propose a method for verifying the reliability of self-explanations provided by models post factum by comparing the attention distribution of the model during the generation of the response and its explanation.<bold> </bold>The authors propose and develop methods for numerical evaluation of answers reliability provided by generative pre-trained transformers. It is proposed to use the Kullback-Leibler divergence over the attention distributions of the model during the issuance of the response and the subsequent explanation. Additionally, it is proposed to compute the ratio of the model's attention between the original query and the generated explanation to understand how much the self-explanation was influenced by its own response. An algorithm for recursively computing the model's attention across the generation steps is proposed to obtain these values.<bold> </bold>The study demonstrated the effectiveness of the proposed methods, identifying metric values corresponding to correct and incorrect explanations and responses.<bold> </bold>We analyzed the currently existing methods for determining the reliability of generative model responses, noting that the overwhelming majority of them are challenging for an ordinary user to interpret. In this regard, we proposed our own methods, testing them on the most widely used generative models available at the time of writing. As a result, we obtained typical values for the proposed metrics, an algorithm for their computation, and visualization.</p></abstract><trans-abstract xml:lang="ru"><p>Со стремительным ростом использования генеративных нейросетевых моделей для решения практических задач все более остро встает проблема объяснения их решений. По мере ввода решений на основе нейросетей в медицинскую практику, государственное управление и сферу обороны требования к таким системам в плане их интерпретируемости однозначно будут расти. В данной работе предложен метод проверки достоверности само-объяснений, которые модели дают постфактум, посредством сравнения распределения внимания модели во время генерации ответа и его объяснения. Авторами предложены и разработаны методы для численной оценки степени достоверности ответов генеративных предобученных трансформеров. Предлагается использовать расхождение Кульбака – Лейблера над распределениями внимания модели во время выдачи ответа и следующего за этим объяснения. Также предлагается вычислять отношение внимания модели между изначальным запросом и сгенерированным объяснением с целью понять, насколько само-объяснение было обусловлено собственным ответом. Для получения данных величин предлагается алгоритм для рекурсивного вычисления внимания модели по шагам генерации. В результате исследования была продемонстрирована работа предложенных методов, найдены значения метрик, соответствующие корректным и некорректным объяснениям и ответам. Был проведен анализ существующих в настоящий момент методов определения достоверности ответов генеративных моделей, причем подавляющее большинство из них сложно интерпретируемые обычным пользователем. В связи с этим мы выдвинули собственные методы, проверив их на наиболее широко используемых на момент написания генеративных моделях, находящихся в открытом доступе. В результате мы получили типичные значения для предложенных метрик, алгоритм их вычисления и визуализации.</p></trans-abstract><kwd-group xml:lang="ru"><kwd>нейронные сети</kwd><kwd>метрики</kwd><kwd>языковые модели</kwd><kwd>интерпретируемость</kwd><kwd>LLM</kwd><kwd>GPT</kwd><kwd>XAI</kwd></kwd-group><kwd-group xml:lang="en"><kwd>neural networks</kwd><kwd>metrics</kwd><kwd>language models</kwd><kwd>interpretability</kwd><kwd>GPT</kwd><kwd>LLM</kwd><kwd>XAI</kwd></kwd-group><funding-group/></article-meta></front><body></body><back><ref-list><ref id="B1"><label>1.</label><mixed-citation>Vaswani A., Shazeer N., Parmar N. et al. Attention is all you need. Advances in neural information processing systems. 2017. No. 3. URL: https://arxiv.org/abs/1706.03762</mixed-citation></ref><ref id="B2"><label>2.</label><mixed-citation>Dosovitskiy A., Beyer L., Kolesnikov A. et al. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations, 2020. URL: https://arxiv.org/abs/2010.11929</mixed-citation></ref><ref id="B3"><label>3.</label><mixed-citation>Selvaraju R.R., Cogswell M., Das A. et al. Grad-CAM: Visual explanations from deep networks via gradient-based localization. URL: https://arxiv.org/abs/1610.02391</mixed-citation></ref><ref id="B4"><label>4.</label><mixed-citation>Ribeiro M.T., Singh S., Guestrin C. "Why should I trust you?": Explaining the Predictions of Any Classifier. URL: https://arxiv.org/abs/1602.04938</mixed-citation></ref><ref id="B5"><label>5.</label><mixed-citation>Lundberg S., Lee S. A unified approach to interpreting model predictions. URL: https://arxiv.org/abs/1705.07874</mixed-citation></ref><ref id="B6"><label>6.</label><mixed-citation>Jesse Vig. Visualizing attention in transformer-based language representation models. URL: https://arxiv.org/abs/1904.02679</mixed-citation></ref><ref id="B7"><label>7.</label><mixed-citation>Bereska L., Gavves E. Mechanistic interpretability for AI Safety – A review. URL: https://arxiv.org/abs/2404.14082</mixed-citation></ref><ref id="B8"><label>8.</label><mixed-citation>Lewis P., Perez E., Piktus A. et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. URL: https://arxiv.org/abs/2005.11401</mixed-citation></ref><ref id="B9"><label>9.</label><mixed-citation>Wei J., Wang X., Schuurmans D. et al. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. URL: https://arxiv.org/abs/2201.11903</mixed-citation></ref><ref id="B10"><label>10.</label><mixed-citation>Pfau J., Merrill W., Bowman S.R. Let's think dot by dot: Hidden computation in transformer language models. URL: https://arxiv.org/abs/2404.15758</mixed-citation></ref><ref id="B11"><label>11.</label><mixed-citation>Abnar S., Zuidema W. Quantifying attention flow in transformers. URL: https://arxiv.org/abs/2005.00928</mixed-citation></ref><ref id="B12"><label>12.</label><mixed-citation>Touvron H., Lavril T., Izacard G. et al. LLaMA: Open and efficient foundation language models. URL: https://arxiv.org/abs/2302.13971</mixed-citation></ref><ref id="B13"><label>13.</label><mixed-citation>Jiang A.Q., Sablayrolles A., Mensch A. et al. Mistral 7B. URL: https://arxiv.org/abs/2310.06825</mixed-citation></ref><ref id="B14"><label>14.</label><mixed-citation>Tunstall L., Beeching E., Lambert N. et al. Zephyr: Direct distillation of LM alignment. URL: https://arxiv.org/abs/2310.16944</mixed-citation></ref><ref id="B15"><label>15.</label><mixed-citation>Gu A., Dao T. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. URL: https://arxiv.org/abs/2312.00752</mixed-citation></ref><ref id="B16"><label>16.</label><mixed-citation>Ali A., Zimerman I., Wolf L. The Hidden Attention of Mamba Models. URL: https://arxiv.org/abs/2403.01590</mixed-citation></ref></ref-list></back></article>
