فهرست این فصلبخش ۳، فصل ۲۲ — صفحهٔ ۶ از ۷
بخش ۳، فصل ۲۲ — صفحهٔ ۶ از ۷
چالشهای کلیدی و بحث
متن اصلی فارسی با منشأ، شناسه و پیوند استناد پایدار.
ادامهٔ بخش «مراجع» — بخش ۵ از ۵
- [379] Dimitris Tsipras، Shibani Santurkar، Logan Engstrom، Alexander Turner و Aleksander Madry. Robustness may be at odds with accuracy. در International Conference on Learning Representations، 2019. URL: https://openreview.net/forum?id=SyxAb30cY7،
doi:10.48550/arXiv.1805.12152.
- [380] Alexander Turner، Dimitris Tsipras و Aleksander Madry. Clean-label backdoor attacks. در ICLR، 2019. URL: https://openreview.net/forum?id=HJg6e2CcK7.
- [381] U.K. AI Safety Institute (مؤسسه ایمنی هوش مصنوعی). Advanced ai evaluations: May update، 2024. تاریخ دسترسی: 2024-08-18 (۲۸ مرداد ۱۴۰۳). URL: https://www.aisi.gov.uk/work/advanced-ai-evaluations-may-update.
- [382] Kush R. Varshney. Trustworthy Machine Learning. Independently Published، Chappaqua, NY, USA، 2022. URL: https://www.trustworthymachinelearning.com/.
- [383] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. در Advances in neural information processing systems، صفحات 5998–6008، 2017. URL: https://arxiv.org/abs/1706.03762،
doi:10.48550/arXiv.1706.03762.
- [384] Sridhar Venkatesan، Harshvardhan Sikka، Rauf Izmailov، Ritu Chadha، Alina Oprea و Michael J. De Lucia. Poisoning attacks and data sanitization mitigations for machine learning models in network intrusion detection systems. در MILCOM، صفحات 874–879. IEEE، 2021. URL: https://ieeexplore.ieee.org/document/9652916،
doi:10.1109/MILCOM52596.2021.9652916.
- [385] Sameer Wagh، Shruti Tople، Fabrice Benhamouda، Eyal Kushilevitz، Prateek Mittal و Tal Rabin. FALCON: honest-majority maliciously secure framework for private deep learning. در Proceedings on Privacy Enhancing Technologies (PoPETs) 2021, Issue 1، 20201.
- [386] Eric Wallace، Shi Feng، Nikhil Kandpal، Matt Gardner و Sameer Singh. Universal adversarial triggers for attacking and analyzing NLP. پیشچاپ arXiv، arXiv:1908.07125، 2019. URL: https://arxiv.org/abs/1908.07125،
doi:10.48550/arXiv.1908.07125.
- [387] Eric Wallace، Kai Xiao، Reimar Leike، Lilian Weng، Johannes Heidecke و Alex Beutel. The instruction hierarchy: Training LLMs to prioritize privileged instructions، 2024. URL: https://arxiv.org/abs/2404.13208،
arXiv:2404.13208،doi:10.48550/arXiv.2404.13208.
- [388] Eric Wallace، Tony Z. Zhao، Shi Feng و Sameer Singh. Concealed data poisoning attacks on NLP models. در NAACL، 2021. URL: https://arxiv.org/abs/2010.12563،
doi:10.48550/arXiv.2010.12563.
- [389] Alexander Wan، Eric Wallace، Sheng Shen و Dan Klein. Poisoning language models during instruction tuning، ۲۰۲۳. URL: https://arxiv.org/abs/2305.00944،
arXiv:2305.00944،doi:10.48550/arXiv.2305.00944.
- [390] Bolun Wang، Yuanshun Yao، Shawn Shan، Huiying Li، Bimal Viswanath، Haitao Zheng و Ben Y. Zhao. Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks. در 2019 IEEE Symposium on Security and Privacy (SP)، صفحات 707–723، San Francisco, CA, USA، مه ۲۰۱۹ (اردیبهشت ۱۳۹۸ تا خرداد ۱۳۹۸). IEEE. URL: https://ieeexplore.ieee.org/document/8835365/،
doi:10.1109/SP.2019.00031.
- [391] Haotao Wang، Tianlong Chen، Shupeng Gui، Ting-Kuei Hu، Ji Liu و Zhangyang Wang. Once-for-All Adversarial Training: In-Situ Tradeoff between Robustness and Accuracy for Free. در Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS 2020)، Vancouver, Canada، ۲۰۲۰. URL: https://arxiv.org/abs/2010.11828،
doi:10.48550/arXiv.2010.11828.
- [392] Hongyi Wang، Kartik Sreenivasan، Shashank Rajput، Harit Vishwakarma، Saurabh Agarwal، Jy-yong Sohn، Kangwook Lee و Dimitris Papailiopoulos. Attack of the Tails: Yes, You Really Can Backdoor Federated Learning. در NeurIPS، ۲۰۲۰. URL: https://arxiv.org/abs/2007.05084،
doi:10.48550/arXiv.2007.05084.
- [393] Lei Wang، Chen Ma، Xueyang Feng، Zeyu Zhang، Hao Yang، Jingsen Zhang، Zhiyuan Chen، Jiakai Tang، Xu Chen، Yankai Lin، Wayne Xin Zhao، Zhewei Wei و Jirong Wen. A survey on large language model based autonomous agents. Frontiers of Computer Science، 18(6)، مارس ۲۰۲۴ (اسفند ۱۴۰۲ تا فروردین ۱۴۰۳). URL: http://dx.doi.org/10.1007/s11704-024-40231-1،
doi:10.1007/s11704-024-40231-1.
- [394] Shiqi Wang، Kexin Pei، Justin Whitehouse، Junfeng Yang و Suman Jana. Formal security analysis of neural networks using symbolic intervals. در 27th USENIX Security Symposium (USENIX Security 18)، صفحات 1599–1614، Baltimore, MD، اوت ۲۰۱۸ (مرداد ۱۳۹۷ تا شهریور ۱۳۹۷). USENIX Association. URL: https://www.usenix.org/conference/usenixsecurity18/presentation/wang-shiqi.
- [395] Wenxiao Wang، Alexander Levine و Soheil Feizi. Improved certified defenses against data poisoning with (deterministic) finite aggregation. در Kamalika Chaudhuri، Stefanie Jegelka، Le Song، Csaba Szepesvári، Gang Niu و Sivan Sabato، ویراستاران، International Conference on Machine Learning, ICML 2022, 17-۲۳ ژوئیه ۲۰۲۲ (۱ مرداد ۱۴۰۱), Baltimore, Maryland, USA، جلد 162 از Proceedings of Machine Learning Research، صفحات 22769–22783. PMLR، ۲۰۲۲. URL: https://proceedings.mlr.press/v162/wang22m.html،
doi:10.48550/arXiv.2202.02628.
- [396] Wenxiao Wang، Alexander J Levine و Soheil Feizi. Improved certified defenses against data poisoning with (Deterministic) finite aggregation. در Kamalika Chaudhuri، Stefanie Jegelka، Le Song، Csaba Szepesvari، Gang Niu و Sivan Sabato، ویراستاران، Proceedings of the 39th International Conference on Machine Learning، جلد 162 از Proceedings of Machine Learning Research، صفحات 22769–22783. PMLR، ۱۷–۲۳ ژوئیهٔ ۲۰۲۲. URL: https://proceedings.mlr.press/v162/wang22m.html،
doi:10.48550/arXiv.2202.02628.
- [397] Xiaosen Wang و Kun He. Enhancing the transferability of adversarial attacks through variance tuning. در IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021، صفحات 1924–1933. Computer Vision Foundation / IEEE، ۲۰۲۱. URL: https://openaccess.thecvf.com/content/CVPR2021/html/Wang_Enhancing_the_Transferability_of_Adversarial_Attacks_Through_Va
- riance_Tuning_CVPR_2021_paper.html،
doi:10.1109/CVPR46437.2021.00196.
- [398] Yanting Wang، Wei Zou و Jinyuan Jia. FCert: Certifiably robust few-shot classification in the era of foundation models. در Proc. IEEE Security and Privacy Symposium، ۲۰۲۴. URL: https://arxiv.org/abs/2404.08631،
doi:10.48550/arXiv.2404.08631.
- [399] Yuxia Wang، Haonan Li، Xudong Han، Preslav Nakov و Timothy Baldwin. Do-not-answer: A dataset for evaluating safeguards in LLMs، ۲۰۲۳. URL: https://arxiv.org/abs/2308.13387،
arXiv:2308.13387،doi:10.48550/arXiv.2308.13387.
- [400] Alexander Wei، Nika Haghtalab و Jacob Steinhardt. Jailbroken: How does LLM safety training fail? arXiv preprint arXiv:2307.02483، ۲۰۲۳.
- [401] Xingxing Wei، Jun Zhu، Sha Yuan و Hang Su. Sparse adversarial perturbations for videos. در Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence، AAAI’19/IAAI’19/EAAI’19. AAAI Press، ۲۰۱۹.
doi:10.1609/aaai.v33i01.33018973.
- [402] Zhipeng Wei، Jingjing Chen، Xingxing Wei، Linxi Jiang، Tat-Seng Chua، Fengfeng Zhou و Yu-Gang Jiang. Heuristic black-box adversarial attacks on video recognition models. در Proceedings of the AAAI Conference on Artificial Intelligence، جلد 34، صفحات 12338–12345، ۲۰۲۰. URL: https://ojs.aaai.org/index.php/AAAI/article/view/6918،
doi:10.48550/arXiv.1911.09449.
- [403] Lilian Weng. Adversarial attacks on latent language models، ۲۰۲۳. URL: https://lilianweng.github.io/posts/2023-10-25-adv-attack-llm/.
- [404] Emily Wenger، Josephine Passananti، Arjun Nitin Bhagoji، Yuanshun Yao، Haitao Zheng و Ben Y. Zhao. Backdoor attacks against deep learning systems in the physical world. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)، صفحات 6202–6211، ۲۰۲۰. URL: https://arxiv.org/abs/2006.14580،
doi:10.48550/arXiv.2006.14580.
- [405] Simon Willison. The dual LLM pattern for building AI assistants that can resist prompt injection، ۲۰۲۳. Accessed: 2024-08-22 (۱ شهریور ۱۴۰۳). URL: https://simonwillison.net/2023/Apr/25/dual-llm-pattern/.
- [406] Yotam Wolf، Noam Wies، Yoav Levine و Amnon Shashua. Fundamental limitations of alignment in large language models. ArXiv، abs/2304.11082، ۲۰۲۳. URL: https://api.semanticscholar.org/CorpusID:258291526،
doi:10.48550/arXiv.2304.11082.
- [407] Dongxian Wu و Yisen Wang. Adversarial neuron pruning purifies backdoored deep models. در M. Ranzato، A. Beygelzimer، Y. Dauphin، P.S. Liang و J. Wortman Vaughan، ویراستاران، Advances in Neural Information Processing Systems، جلد 34، صفحات 16913–16925. Curran Associates, Inc.، 2021. URL: https://proceedings.neurips.cc/paper/2021/file/8cbe9ce23f42628c98f80fa0fac8b19a-Paper.pdf،
doi:10.48550/arXiv.2110.14430.
- [408] Fangzhou Wu، Ning Zhang، Somesh Jha، Patrick McDaniel و Chaowei Xiao. A new era in LLM security: Exploring security concerns in real-world LLM-based systems، 2024. URL: https://arxiv.org/abs/2402.18649،
arXiv:2402.18649.
- [409] Xi Wu، Matthew Fredrikson، Somesh Jha و Jeffrey F. Naughton. A methodology for formalizing model-inversion attacks. در 2016 IEEE 29th Computer Security Foundations Symposium (CSF)، صفحات 355–370، 2016.
doi:10.1109/CSF.2016.32.
- [410] Yuhao Wu، Franziska Roesner، Tadayoshi Kohno، Ning Zhang و Umar Iqbal. SecGPT: An execution isolation architecture for LLM-based systems، 2024. URL: https://arxiv.org/abs/2403.04960،
arXiv:2403.04960،doi:10.48550/arXiv.2402.18649.
- [411] Zhen Xiang، David J. Miller و George Kesidis. Post-training detection of backdoor attacks for two-class and multi-attack scenarios. در The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022. OpenReview.net، 2022. URL: https://openreview.net/forum?id=MSgB8D4Hy51،
doi:10.48550/arXiv.2201.08474.
- [412] Huang Xiao، Battista Biggio، Gavin Brown، Giorgio Fumera، Claudia Eckert و Fabio Roli. Is feature selection secure against training data poisoning? در International Conference on Machine Learning، صفحات 1689–1698، 2015. URL: https://arxiv.org/abs/1804.07933،
doi:10.48550/arXiv.1804.07933.
- [413] Qizhe Xie، Zihang Dai، Eduard Hovy، Thang Luong و Quoc Le. Unsupervised data augmentation for consistency training. در H. Larochelle، M. Ranzato، R. Hadsell، M.F. Balcan و H. Lin، ویراستاران، Advances in Neural Information Processing Systems، جلد 33، صفحات 6256–6268. Curran Associates, Inc.، 2020. URL: https://proceedings.neurips.cc/paper/2020/file/44feb0096faa8326192570788b38c1d1-Paper.pdf،
doi:10.48550/arXiv.1904.12848.
- [414] Weilin Xu، Yanjun Qi و David Evans. Automatically evading classifiers. در Proceedings of the 2016 Network and Distributed Systems Symposium، صفحات 21–24، 2016. URL: https://www.cs.virginia.edu/~evans/pubs/ndss2016/.
- [415] Xiaojun Xu، Xinyun Chen، Chang Liu، Anna Rohrbach، Trevor Darrell و Dawn Song. Fooling vision and language models despite localization and attention mechanism. https://arxiv.org/abs/1709.08693، 2017.
doi:10.48550/ARXIV.1709.08693.
- [416] Xiaojun Xu، Qi Wang، Huichen Li، Nikita Borisov، Carl A. Gunter و Bo Li. Detecting AI trojans using meta neural analysis. در IEEE Symposium on Security and Privacy, S&P 2021، صفحات 103–120، ایالات متحده، مه ۲۰۲۱ (اردیبهشت ۱۴۰۰ تا خرداد ۱۴۰۰). URL: https://ieeexplore.ieee.org/document/9519467،
doi:10.1109/SP40001.2021.00034.
- [417] Karren Yang، Wan-Yi Lin، Manash Barman، Filipe Condessa و Zico Kolter. Defending multimodal fusion models against single-source adversaries. در 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE Xplore، 2022. URL: https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=9578130،
doi:10.48550/ARXIV.2206.12714.
- [418] Limin Yang، Zhi Chen، Jacopo Cortellazzi، Feargus Pendlebury، Kevin Tu، Fabio Pierazzi، Lorenzo Cavallaro و Gang Wang. Jigsaw puzzle: Selective backdoor attack to subvert malware classifiers. CoRR، abs/2202.05470، 2022. URL: https://arxiv.org/abs/2202.05470،
arXiv:2202.05470،doi:10.48550/arXiv.2202.05470.
- [419] Shunyu Yao، Jeffrey Zhao، Dian Yu، Nan Du، Izhak Shafran، Karthik Narasimhan و Yuan Cao. React: Synergizing reasoning and acting in language models، 2023. URL: https://arxiv.org/abs/2210.03629،
arXiv:2210.03629،doi:10.48550/arXiv.2210.03629.
- [420] Yuanshun Yao، Huiying Li، Haitao Zheng و Ben Y. Zhao. Latent backdoor attacks on deep neural networks. در Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security، CCS ’19، صفحه 2041–2055، نیویورک، NY، ایالات متحده، 2019. Association for Computing Machinery.
doi:10.1145/3319535.3354209.
- [421] Jiayuan Ye، Aadyaa Maddi، Sasi Kumar Murakonda، Vincent Bindschaedler و Reza Shokri. Enhanced membership inference attacks against machine learning models. در Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security، CCS ’22، صفحه 3093–3106، نیویورک، NY، ایالات متحده، 2022. Association for Computing Machinery.
doi:10.1145/3548606.3560675.
- [422] Samuel Yeom، Irene Giacomelli، Matt Fredrikson و Somesh Jha. Privacy risk in machine learning: Analyzing the connection to overfitting. در IEEE Computer Security Foundations Symposium، CSF ’18، صفحات 268–282، 2018. https://arxiv.org/abs/1709.01604. URL: https://arxiv.org/abs/1709.01604،
doi:10.48550/arXiv.1709.01604.
- [423] Dong Yin، Yudong Chen، Ramchandran Kannan و Peter Bartlett. Byzantine-Robust Distributed Learning: Towards Optimal Statistical Rates. در ICML، 2018. URL: https://arxiv.org/abs/1803.01498،
doi:10.48550/arXiv.1803.01498.
- [424] Youngjoon Yu، Hong Joo Lee، Byeong Cheon Kim، Jung Uk Kim و Yong Man Ro. Investigating vulnerability to adversarial examples on multimodal data fusion in deep learning. https://arxiv.org/abs/2005.10987، 2020. برخط.
doi:10.48550/ARXIV.2005.10987.
- [425] Andrew Yuan، Alina Oprea و Cheng Tan. Dropout attacks. در IEEE Symposium on Security and Privacy (S&P)، 2024. URL: https://arxiv.org/abs/2309.01614،
doi:10.48550/arXiv.2309.01614.
- [426] Santiago Zanella-Béguelin, Lukas Wutschitz, Shruti Tople, Victor Rühle, Andrew Paverd, Olga Ohrimenko، Boris Köpf و Marc Brockschmidt. «Analyzing information leakage of updates to natural language models». در Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security، صفحات 363–375، نیویورک، نیویورک، آمریکا، 2020. Association for Computing Machinery.
doi:10.1145/3372297.3417880.
- [427] Santiago Zanella-Beguelin، Lukas Wutschitz، Shruti Tople، Ahmed Salem، Victor Rühle، Andrew Paverd، Mohammad Naseri، Boris Köpf و Daniel Jones. «Bayesian estimation of differential privacy». در Andreas Krause، Emma Brunskill، Kyunghyun Cho، Barbara Engelhardt، Sivan Sabato و Jonathan Scarlett (ویراستاران)، Proceedings of the 40th International Conference on Machine Learning، جلد 202 از Proceedings of Machine Learning Research، صفحات 40624–40636. PMLR، ۲۳ تا ۲۹ ژوئیه ۲۰۲۳ (۷ مرداد ۱۴۰۲). URL: https://proceedings.mlr.press/v202/zanella-beguelin23a.html،
doi:10.48550/arXiv.2206.05199.
- [428] Rowan Zellers، Ximing Lu، Jack Hessel، Youngjae Yu، Jae Sung Park، Jize Cao، Ali Farhadi و Yejin Choi. «Merlot: Multimodal neural script knowledge models»، 2021. URL: https://arxiv.org/abs/2106.02636،
arXiv:2106.02636،doi:10.48550/arXiv.2106.02636.
- [429] Yi Zeng، Si Chen، Won Park، Zhuoqing Mao، Ming Jin و Ruoxi Jia. «Adversarial unlearning of backdoors via implicit hypergradient». در International Conference on Learning Representations، 2022. URL: https://openreview.net/forum?id=MeeQkFYVbzW،
doi:10.48550/arXiv.2110.03735.
- [430] Qiusi Zhan، Zhixiang Liang، Zifan Ying و Daniel Kang. «InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents»، 2024. URL: https://arxiv.org/abs/2403.02691،
arXiv:2403.02691،doi:10.48550/arXiv.2403.02691.
- [431] Chiyuan Zhang، Samy Bengio، Moritz Hardt، Benjamin Recht و Oriol Vinyals. «Understanding deep learning (still) requires rethinking generalization». Commun. ACM، 64(3):107–115، فوریه ۲۰۲۱ (بهمن ۱۳۹۹ تا اسفند ۱۳۹۹).
doi:10.1145/3446776.
- [432] Hanlin Zhang، Benjamin L. Edelman، Danilo Francati، Daniele Venturi، Giuseppe Ateniese و Boaz Barak. «Watermarks in the sand: Impossibility of strong watermarking for generative models». ArXiv، abs/2311.04378، 2023. URL: https://api.semanticscholar.org/CorpusID:265050535،
doi:10.48550/arXiv.2311.04378.
- [433] Hongyang Zhang، Yaodong Yu، Jiantao Jiao، Eric Xing، Laurent El Ghaoui و Michael Jordan. «Theoretically principled trade-off between robustness and accuracy». در Kamalika Chaudhuri و Ruslan Salakhutdinov (ویراستاران)، Proceedings of the 36th International Conference on Machine Learning، جلد 97 از Proceedings of Machine Learning Research، صفحات 7472–7482. PMLR، ۹ تا ۱۵ ژوئن ۲۰۱۹ (۲۵ خرداد ۱۳۹۸). URL: https://proceedings.mlr.press/v97/zhang19p.html،
doi:10.48550/arXiv.1901.08573.
- [434] Ruisi Zhang، Seira Hidano و Farinaz Koushanfar. «Text revealer: Private text reconstruction via model inversion attacks against transformers». arXiv preprint arXiv:2209.10505، 2022. URL: https://arxiv.org/abs/2209.10505،
doi:10.48550/arXiv.2209.10505.
- [435] Su-Fang Zhang، Jun-Hai Zhai، Bo-Jun Xie، Yan Zhan و Xin Wang. «Multimodal representation learning: Advances, trends and challenges». در 2019 International Conference on Machine Learning and Cybernetics (ICMLC)، صفحات 1–6. IEEE، 2019. URL: https://api.semanticscholar.org/CorpusID:209901378،
doi:10.1109/ICMLC48188.2019.8949228.
- [436] Susan Zhang، Mona Diab و Luke Zettlemoyer. «Democratizing access to large-scale language models with OPT-175B». https://ai.facebook.com/blog/democratizing-access-to-large-scale-language-models-with-opt-175b/، 2022. Meta AI.
- [437] Wanrong Zhang، Shruti Tople و Olga Ohrimenko. «Leakage of dataset properties in Multi-Party machine learning». در 30th USENIX Security Symposium (USENIX Security 21)، صفحات 2687–2704. USENIX Association، اوت ۲۰۲۱ (مرداد ۱۴۰۰ تا شهریور ۱۴۰۰). URL: https://www.usenix.org/conference/usenixsecurity21/presentation/zhang-wanrong،
doi:10.48550/arXiv.2006.07267.
- [438] Wei Emma Zhang، Quan Z. Sheng، Ahoud Alhazmi و Chenliang Li. «Adversarial attacks on deep-learning models in natural language processing: A survey». ACM Trans. Intell. Syst. Technol.، 11(3)، آوریل ۲۰۲۰ (فروردین ۱۳۹۹ تا اردیبهشت ۱۳۹۹).
doi:10.1145/3374217.
- [439] Yiming Zhang و Daphne Ippolito. «Prompts should not be seen as secrets: Systematically measuring prompt extraction attack success». arXiv preprint arXiv:2307.06865، 2023. URL: https://arxiv.org/abs/2307.06865،
doi:10.48550/arXiv.2307.06865.
- [440] Yuhao Zhang، Aws Albarghouthi و Loris D’Antoni. «Bagflip: A certified defense against data poisoning». در Alice H. Oh، Alekh Agarwal، Danielle Belgrave و Kyunghyun Cho (ویراستاران)، Advances in Neural Information Processing Systems، 2022. URL: https://openreview.net/forum?id=ZidkM5b92G،
doi:10.48550/arXiv.2205.13634.
- [441] Zhengming Zhang، Ashwinee Panda، Linyue Song، Yaoqing Yang، Michael Mahoney، Prateek Mittal، Ramchandran Kannan و Joseph Gonzalez. «Neurotoxin: Durable backdoors in federated learning». در Kamalika Chaudhuri، Stefanie Jegelka، Le Song، Csaba Szepesvari، Gang Niu و Sivan Sabato (ویراستاران)، Proceedings of the 39th International Conference on Machine Learning، جلد 162 از Proceedings of Machine Learning Research، صفحات 26429–26446. PMLR، ۱۷ تا ۲۳ ژوئیه ۲۰۲۲ (۱ مرداد ۱۴۰۱). URL: https://proceedings.mlr.press/v162/zhang22w.html،
doi:10.48550/arXiv.2206.10341.
- [442] Zhikun Zhang، Min Chen، Michael Backes، Yun Shen و Yang Zhang. «Inference attacks against graph neural networks». در 31st USENIX Security Symposium (USENIX Security 22)، 2022. URL: https://www.usenix.org/conference/usenixsecurity22/presentation/zhang-zhikun.
- [443] Junhao Zhou، Yufei Chen، Chao Shen و Yang Zhang. «Property inference attacks against GANs». در Proceedings of Network and Distributed System Security، NDSS، 2022. URL: https://arxiv.org/abs/2111.07608،
doi:10.48550/arXiv.2111.07608.
- [444] Chen Zhu، W. Ronny Huang، Hengduo Li، Gavin Taylor، Christoph Studer و Tom Goldstein. حملات مسمومسازی با برچسب پاکِ انتقالپذیر بر شبکههای عصبی عمیق. به ویراستاری Kamalika Chaudhuri و Ruslan Salakhutdinov، مجموعه مقالات سیوششمین کنفرانس بینالمللی یادگیری ماشین (Proceedings of the 36th International Conference on Machine Learning)، جلد ۹۷ از Proceedings of Machine Learning Research، صفحات ۷۶۱۴–۷۶۲۳. PMLR، ۰۹ تا ۱۵ ژوئن ۲۰۱۹ (۲۵ خرداد ۱۳۹۸). URL: https://proceedings.mlr.press/v97/zhu19a.html،
doi:10.48550/arXiv.1905.05897.
- [445] Daniel M. Ziegler، Nisan Stiennon، Jeffrey Wu، Tom B. Brown، Alec Radford، Dario Amodei، Paul Christiano و Geoffrey Irving. تنظیم دقیق «مدلهای زبانی» (Language models) بر اساس ترجیحات انسانی، ۲۰۲۰. URL: https://arxiv.org/abs/1909.08593،
arXiv:1909.08593،doi:10.48550/arXiv.1909.08593.
- [446] Giulio Zizzo، Chris Hankin، Sergio Maffeis و Kevin Jones. «یادگیری ماشین خصمانه» (Adversarial Machine Learning, AML) فراتر از دامنه تصویر. در Proceedings of the 56th Annual Design Automation Conference 2019، DAC ’19، نیویورک، ایالت نیویورک، آمریکا، ۲۰۱۹. Association for Computing Machinery.
doi:10.1145/3316781.3323470.
- [447] Andy Zou، Long Phan، Justin Wang، Derek Duenas، Maxwell Lin، Maksym Andriushchenko، Rowan Wang، Zico Kolter، Matt Fredrikson و Dan Hendrycks. بهبود همسوسازی و استواری با «مدارشکنها» (circuit breakers)، ۲۰۲۴. URL: https://arxiv.org/abs/2406.04313،
arXiv:2406.04313،doi:h10.48550/arXiv.2406.04313.
- [448] Andy Zou، Zifan Wang، J Zico Kolter و Matt Fredrikson. حملات خصمانهٔ همگانی و انتقالپذیر بر مدلهای زبانیِ همسوشده. arXiv preprint arXiv:2307.15043، ۲۰۲۳. URL: https://arxiv.org/abs/2307.15043،
doi:10.48550/arXiv.2307.15043.
- [449] Wei Zou، Runpeng Geng، Binghui Wang و Jinyuan Jia. PoisonedRAG: حملات مسمومسازی دانش علیه تولیدِ تقویتشده با بازیابی در مدلهای زبانی بزرگ، ۲۰۲۴. URL: https://arxiv.org/abs/2005.11401،
arXiv:2402.07867،doi:10.48550/arXiv.2005.11401.
پیوست الف. واژهنامه
با کلیک بر شمارهٔ صفحه در انتهای هر تعریف، به صفحهای که آن اصطلاح در آن بهکار رفته است هدایت میشوید.
استناد
محمدعلی کهندژ، راهنمای جامع حاکمیت، امنیت و مدیریت ریسک هوش مصنوعی، شناسه بخش: KDJ-AI-2026E1-P03-C22-S11
شناسهٔ محتوا KDJ-AI-2026E1-P03-C22-S11-89B3BCB4
واژهنامه — A
- نمونهٔ خصمانه (adversarial example) نمونهٔ آزمونی تغییریافتهای که موجب طبقهبندی نادرست یا رفتار نادرست یک مدل یادگیری ماشین (Machine Learning, ML) در زمان استقرار (Deployment) میشود. ix, 6
- «یادگیری ماشین خصمانه» (Adversarial Machine Learning, AML) حملاتی که از ماهیت آماری و دادهمحورِ سامانههای یادگیری ماشین بهره میگیرند. xii, 1
- عامل (agent) برنامههای نرمافزاری که میتوانند با محیط خود تعامل کنند، اطلاعات دریافت کنند و در راستای هدفی بزرگتر که از بیرون تعیین شده است، اقداماتی خودهدایتشده انجام دهند. 1, 35, 37, 39, 50, 52, 54
- «سطح زیر منحنی» (AREA UNDER THE CURVE, AUC) سنجشی از توانایی یک طبقهبند در تمایز میان کلاسها در یادگیری ماشین. AUC بالاتر به این معناست که یک مدل هنگام تمایز میان دو کلاس عملکرد بهتری دارد. AUC کل مساحت دوبُعدی زیر «منحنی مشخصهٔ عملکرد گیرنده» (RECEIVER OPERATING CHARACTERISTIC — ROC-AUC) را میسنجد. 30
- حملات استنباط ویژگی (attribute inference attacks) حملهای علیه مدلهای یادگیری ماشین که با داشتن دانشی جزئی دربارهٔ یک رکورد، ویژگیهای حساسِ آن رکورد از دادههای آموزشی را استنباط میکند. 7
- گسست دسترسپذیری (availability breakdown) در بستر AML، اختلال در توانایی کاربران یا فرایندهای دیگر برای دستیابی بهموقع و قابلاعتماد به خروجیها یا کارکرد یک «سامانه هوش مصنوعی» (AI system). 6, 39
استناد
محمدعلی کهندژ، راهنمای جامع حاکمیت، امنیت و مدیریت ریسک هوش مصنوعی، شناسه بخش: KDJ-AI-2026E1-P03-C22-S12
شناسهٔ محتوا KDJ-AI-2026E1-P03-C22-S12-19DC0EF2
واژهنامه — B
- الگوی درِ پشتی (backdoor pattern) تبدیل یا درجی که روی یک نمونهٔ داده اعمال میشود و رفتاری مشخصشده توسط مهاجم را در مدلی که هدف حملهٔ مسمومسازی با درِ پشتی (backdoor poisoning attack) قرار گرفته فعال میکند. برای مثال، در بینایی رایانهای، مهاجم ممکن است بتواند مدل را مسموم کند بهگونهای که درج یک مربع از پیکسلهای سفید، برچسب هدفِ مطلوب را ایجاد کند. 6, 22, 107
- حملهٔ مسمومسازی با درِ پشتی (backdoor poisoning attack) حملهای از نوع مسمومسازی (poisoning) که باعث میشود مدل در پاسخ به ورودیهایی که الگوی مشخصی از درِ پشتی (backdoor pattern) را دنبال میکنند، رفتاری منتخبِ مهاجم نشان دهد. 6, 42
استناد
محمدعلی کهندژ، راهنمای جامع حاکمیت، امنیت و مدیریت ریسک هوش مصنوعی، شناسه بخش: KDJ-AI-2026E1-P03-C22-S13
شناسهٔ محتوا KDJ-AI-2026E1-P03-C22-S13-D6D530D2
واژهنامه — C
- classification (طبقهبندی) وظیفه پیشبینی اینکه ورودی به کدامیک از مجموعهای از مقولههای گسسته تعلق دارد. 5
- convolutional neural networks (شبکههای عصبی کانولوشنی) ردهای از شبکههای عصبی پیشخور که دستکم یک لایه کانولوشنی دارند و با نام CNNs شناخته میشوند. در لایههای کانولوشنی، آشکارسازهای ویژگی (که بهصورت هسته یا فیلتر شناخته میشوند) ویژگیهای خاصی را در سراسر دادههای ورودی آشکار میکنند. CNNها عمدتاً برای پردازش دادههای شبکهای مانند تصاویر بهکار میروند و برای وظایفی مانند طبقهبندی تصویر، تشخیص (Detection) اشیا و قطعهبندی تصویر بهویژه مؤثرند. 5, 31
استناد
محمدعلی کهندژ، راهنمای جامع حاکمیت، امنیت و مدیریت ریسک هوش مصنوعی، شناسه بخش: KDJ-AI-2026E1-P03-C22-S14
شناسهٔ محتوا KDJ-AI-2026E1-P03-C22-S14-EE8D28FF