فهرست این فصلبخش ۳، فصل ۲۲ — صفحهٔ ۲ از ۷
بخش ۳، فصل ۲۲ — صفحهٔ ۲ از ۷
چالشهای کلیدی و بحث
متن اصلی فارسی با منشأ، شناسه و پیوند استناد پایدار.
مراجع
بخش ۱ از ۵ — این بخش طولانی است و در چند صفحه آمده است.
ترجمهٔ منبع
- [1] Martin Abadi، Andy Chu، Ian Goodfellow، H Brendan McMahan، Ilya Mironov، Kunal Talwar و Li Zhang. Deep learning with differential privacy. در ACM Conference on Computer and Communications Security، CCS ’16، صفحات 308–318، ۲۰۱۶. https://arxiv.org/abs/1607.00133.
doi:10.48550/arXiv.1607.00133.
- [2] Mark Abspoel، Daniel Escudero و Nikolaj Volgushev. Secure training of decision trees with continuous attributes. در Proceedings on Privacy Enhancing Technologies (PoPETs) 2021, Issue 1، ۲۰۲۰.
- [3] Hojjat Aghakhani، Wei Dai، Andre Manoel، Xavier Fernandes، Anant Kharkar، Christopher Kruegel، Giovanni Vigna، David Evans، Ben Zorn و Robert Sim. TrojanPuzzle: Covertly poisoning code-suggestion models، ۲۰۲۴. URL: https://arxiv.org/abs/2301.02344،
arXiv:2301.02344،doi:10.48550/arXiv.2301.02344.
- [4] Hojjat Aghakhani، Dongyu Meng، Yu-Xiang Wang، Christopher Kruegel و Giovanni Vigna. Bullseye polytope: A scalable clean-label poisoning attack with improved transferability. در IEEE European Symposium on Security and Privacy, 2021, Vienna, Austria, September 6-10, 2021، صفحات 159–178. IEEE، ۲۰۲۱. URL: https://ieeexplore.ieee.org/document/9581207،
doi:10.1109/EuroSP51992.2021.00021.
- [5] Meta AI. Llama guard 3 documentation، ۲۰۲۴. Accessed: 2024-08-13 (۲۳ مرداد ۱۴۰۳). URL: https://llama.meta.com/docs/model-cards-and-prompt-formats/llama-guard-3/.
- [6] Meta AI. Prompt guard documentation، ۲۰۲۴. Accessed: 2024-08-13 (۲۳ مرداد ۱۴۰۳). URL: https://llama.meta.com/docs/model-cards-and-prompt-formats/prompt-guard/.
- [7] AI Safety Institute. Systemic ai safety fast grants. https://www.aisi.gov.uk/grants، ۲۰۲۴. Accessed: 2024-08-22 (۱ شهریور ۱۴۰۳).
- [8] Dan Alistarh، Zeyuan Allen-Zhu و Jerry Li. Byzantine Stochastic Gradient Descent. در NeurIPS، ۲۰۱۸. URL: https://arxiv.org/abs/1803.08917،
doi:10.48550/arXiv.1803.08917.
- [9] Gabriel Alon و Michael Kamfonas. Detecting language model attacks with perplexity، ۲۰۲۳. URL: https://arxiv.org/abs/2308.14132،
arXiv:2308.14132،doi:10.48550/arXiv.2308.14132.
- [10] Galen Andrew، Peter Kairouz، Sewoong Oh، Alina Oprea، H. Brendan McMahan و Vinith Suriyakumar. One-shot empirical privacy estimation for federated learning، ۲۰۲۳. URL: https://arxiv.org/abs/2302.03098،
arXiv:2302.03098،doi:10.48550/arXiv.2302.03098.
- [11] Maksym Andriushchenko، Francesco Croce و Nicolas Flammarion. Jailbreaking leading safety-aligned LLMs with simple adaptive attacks، ۲۰۲۴. URL: https://arxiv.org/abs/2404.02151،
arXiv:2404.02151،doi:10.48550/arXiv.2404.02151.
- [12] Maksym Andriushchenko، Alexandra Souly، Mateusz Dziemian، Derek Duenas، Maxwell Lin، Justin Wang، Dan Hendrycks، Andy Zou، Zico Kolter، Matt Fredrikson، Eric Winsor، Jerome Wynne، Yarin Gal و Xander Davies. Agentharm: A benchmark for measuring harmfulness of llm agents، ۲۰۲۴. URL: https://arxiv.org/abs/2410.09024،
arXiv:2410.09024.
- [13] Anthropic. Model Card and Evaluations for Claude Models. https://www-files.anthropic.com/production/images/Model-Card-Claude-2.pdf، ژوئیه ۲۰۲۳ (تیر ۱۴۰۲ تا مرداد ۱۴۰۲). Anthropic.
- [14] Anthropic. Anthropic’s interactive prompt engineering tutorial. https://github.com/anthropics/prompt-eng-interactive-tutorial، ۲۰۲۴. Accessed: 2024-08-22 (۱ شهریور ۱۴۰۳).
- [15] Anthropic. Claude 3.5 Sonnet. https://www.anthropic.com/news/claude-3-5-sonnet، ژوئن ۲۰۲۴ (خرداد ۱۴۰۳ تا تیر ۱۴۰۳). Anthropic.
- [16] Anthropic. Expanding our model safety bug bounty program. https://www.anthropic.com/news/model-safety-bug-bounty، ۲۰۲۴. Accessed: 2024-08-22 (۱ شهریور ۱۴۰۳).
- [17] Giovanni Apruzzese، Hyrum S Anderson، Savino Dambra، David Freeman، Fabio Pierazzi و Kevin Roundy. “real attackers don’t compute gradients”: Bridging the gap between adversarial ml research and practice. در 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML)، صفحات 339–364. IEEE، ۲۰۲۳. URL: https://arxiv.org/abs/2212.14315،
doi:10.48550/arXiv.2212.14315.
- [18] Arthur. Shield، ۲۰۲۳. URL: https://www.arthur.ai/product/shield.
- [19] Giuseppe Ateniese، Luigi V. Mancini، Angelo Spognardi، Antonio Villani، Domenico Vitali و Giovanni Felici. Hacking smart machines with smarter ones: How to extract meaningful data from machine learning classifiers. Int. J. Secur. Netw.، 10(3):137–150، سپتامبر ۲۰۱۵ (شهریور ۱۳۹۴ تا مهر ۱۳۹۴).
doi:10.1504/IJSN.2015.071829.
- [20] Anish Athalye، Nicholas Carlini و David A. Wagner. Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples. در Jennifer G. Dy و Andreas Krause، ویراستاران، Proceedings of the 35th International Conference on Machine Learning, ICML 2018, Stockholmsmässan, Stockholm, Sweden, July 10-15, 2018، جلد ۸۰ از Proceedings of Machine Learning Research، صفحات 274–283. PMLR، ۲۰۱۸. URL: http://proceedings.mlr.press/v80/athalye18a.html،
doi:10.48550/arXiv.1802.00420.
- [21] Anish Athalye، Logan Engstrom، Andrew Ilyas و Kevin Kwok. Synthesizing robust adversarial examples، ۲۰۱۸. URL: https://arxiv.org/abs/1707.07397،
arXiv:1707.07397،doi:10.48550/arXiv.1707.07397.
- [22] Eugene Bagdasaryan، Andreas Veit، Yiqing Hua، Deborah Estrin و Vitaly Shmatikov. How to backdoor federated learning. در Silvia Chiappa و Roberto Calandra، ویراستاران، Proceedings of the Twenty Third International Conference on Artificial Intelligence and Statistics، جلد ۱۰۸ از Proceedings of Machine Learning Research، صفحات 2938–2948. PMLR، ۲۶–۲۸ اوت ۲۰۲۰ (۷ شهریور ۱۳۹۹). URL: http://proceedings.mlr.press/v108/bagdasaryan20a.html.
- [23] Eugene Bagdasaryan، Andreas Veit، Yiqing Hua، Deborah Estrin و Vitaly Shmatikov. How to backdoor federated learning. در AISTATS. PMLR، ۲۰۲۰. URL: https://proceedings.mlr.press/v108/bagdasaryan20a.html،
doi:10.48550/arXiv.1807.00459.
- [24] Eugene Bagdasaryan، Ren Yi، Sahra Ghalebikesabi، Peter Kairouz، Marco Gruteser، Sewoong Oh، Borja Balle و Daniel Ramage. Air gap: Protecting privacy-conscious conversational agents، ۲۰۲۴. URL: https://arxiv.org/abs/2405.05175،
arXiv:2405.05175،doi:10.48550/arXiv.2405.05175.
- [25] Marieke Bak، Vince Istvan Madai، Marie-Christine Fritzsche، Michaela Th. Mayrhofer و Stuart McLennan. You can’t have ai both ways: Balancing health data privacy and access fairly. Frontiers in Genetics، ۱۳، ۲۰۲۲. https://www.frontiersin.org/articles/10.3389/fgene.2022.929453.
doi:10.3389/fgene.2022.929453.
- [26] Borja Balle، Giovanni Cherubin و Jamie Hayes. بازسازی دادههای آموزشی با دشمنان آگاه. در NeurIPS 2021 Workshop on Privacy in Machine Learning (PRIML)، ۲۰۲۱. URL: https://openreview.net/forum?id=Yi2DZTbnBl4،
doi: 10.48550/arXiv.2201.04845.
- [27] Tadas Baltrušaitis، Chaitanya Ahuja و Louis-Philippe Morency. یادگیری ماشین چندوجهی: مرور و طبقهبندی، ۲۰۱۷.
doi:10.48550/ARXIV.1705.09406.
- [28] Anthony M. Barrett، Dan Hendrycks، Jessica Newman و Brandie Nonnecke. نمایهٔ استانداردهای مدیریت ریسک هوش مصنوعی UC Berkeley برای سامانههای هوش مصنوعی همهمنظوره (General-Purpose AI Systems, GPAIS) و مدلهای بنیادین. UC Berkeley Center for Long Term Cybersecurity، ۲۰۲۳. https://cltc.berkeley.edu/seeking-input-and-feedback-ai-risk-management-standards-profile-for-increasingly-multi-purpose-or-general-purpose-ai/.
doi:10.48550/ARXIV.2206.08966.
- [29] Lejla Batina، Shivam Bhasin، Dirmanto Jap و Stjepan Picek. CSI NN: مهندسی معکوس معماریهای شبکه عصبی از طریق کانال جانبی الکترومغناطیسی. در Proceedings of the 28th USENIX Conference on Security Symposium، SEC’19، صفحه ۵۱۵–۵۳۲، آمریکا، ۲۰۱۹. USENIX Association. URL: https://www.usenix.org/conference/usenixsecurity19/presentation/batina.
- [30] Khaled Bayoudh، Raja Knani، Fayçal Hamdaoui و Abdellatif Mtibaa. مروری بر یادگیری چندوجهی عمیق برای بینایی رایانه: پیشرفتها، روندها، کاربردها و مجموعهدادهها. Vis. Comput.، 38(8):2939–2970، اوت ۲۰۲۲ (مرداد ۱۴۰۱ تا شهریور ۱۴۰۱).
doi:10.1007/s00371-021-02166-7.
- [31] Nora Belrose، Zach Furman، Logan Smith، Danny Halawi، Igor Ostrovsky، Lev McKinney، Stella Biderman و Jacob Steinhardt. استخراج پیشبینیهای نهفته از ترنسفورمرها با عدسی تنظیمشده. arXiv preprint arXiv:2303.08112، ۲۰۲۳. URL: https://arxiv.org/abs/2303.08112،
doi:10.48550/arXiv.2303.08112.
- [32] Philipp Benz، Chaoning Zhang، Soomin Ham، Gyusang Karjauv، Adil Cho و In So Kweon. بدهبستان سهگانه میان دقت، استواری و انصاف. Workshop on Adversarial Machine Learning in Real-World Computer Vision Systems and Online Challenges (AML-CV) at CVPR، ۲۰۲۱. URL: https://dl.acm.org/doi/10.1145/3645088،
doi:10.1145/3645088.
- [33] Jamie Bernardi، Gabriel Mukobi، Hilary Greaves، Lennart Heim و Markus Anderljung. سازگاری اجتماعی با هوش مصنوعی پیشرفته، ۲۰۲۴. URL: https://arxiv.org/abs/2405.10295،
arXiv:2405.10295،doi:10.48550/arXiv.2405.10295.
- [34] David Berthelot، Nicholas Carlini، Ian Goodfellow، Nicolas Papernot، Avital Oliver و Colin A Raffel. MixMatch: رویکردی جامع به یادگیری نیمهنظارتی. در H. Wallach، H. Larochelle، A. Beygelzimer، F. d’Alché-Buc، E. Fox و R. Garnett، ویراستاران، Advances in Neural Information Processing Systems 32، صفحات ۵۰۵۰–۵۰۶۰.
Curran Associates, Inc.، ۲۰۱۹. URL: http://papers.nips.cc/paper/8749-mixmatch-a-holistic-approach-to-semi-supervised-learning.pdf، doi:10.48550/arXiv.2405.10295.
- [35] Arjun Nitin Bhagoji، Supriyo Chakraborty، Prateek Mittal و Seraphin Calo. حملات مسمومسازی مدل در یادگیری فدرال. در NeurIPS SECML، ۲۰۱۸.
- [36] Arjun Nitin Bhagoji، Supriyo Chakraborty، Prateek Mittal و Seraphin Calo. تحلیل یادگیری فدرال از دریچه خصمانه. در Kamalika Chaudhuri و Ruslan Salakhutdinov، ویراستاران، Proceedings of the 36th International Conference on Machine Learning، جلد ۹۷ از Proceedings of Machine Learning Research، صفحات ۶۳۴–۶۴۳. PMLR، ۰۹–۱۵ ژوئن ۲۰۱۹ (۲۵ خرداد ۱۳۹۸). URL: https://proceedings.mlr.press/v97/bhagoji19a.html،
doi:10.48550/arXiv.1811.12470.
- [37] Battista Biggio، Igino Corona، Giorgio Fumera، Giorgio Giacinto و Fabio Roli. کیسهبندی طبقهبندها برای مقابله با حملات مسمومسازی در وظایف طبقهبندی خصمانه. در Proceedings of the 10th International Conference on Multiple Classifier Systems، MCS’11، صفحه ۳۵۰–۳۵۹، برلین، هایدلبرگ، ۲۰۱۱. Springer-Verlag. URL: https://api.semanticscholar.org/CorpusID:12680508.
- [38] Battista Biggio، Igino Corona، Davide Maiorca، Blaine Nelson، Nedim Šrndić، Pavel Laskov، Giorgio Giacinto و Fabio Roli. حملات گریز علیه یادگیری ماشین در زمان آزمون. در Joint European conference on machine learning and knowledge discovery in databases، صفحات ۳۸۷–۴۰۲. Springer، ۲۰۱۳.
doi:10.1007/978-3-642-40994-3_25.
- [39] Battista Biggio، Blaine Nelson و Pavel Laskov. ماشینهای بردار پشتیبان تحت نویز برچسب خصمانه. در Chun-Nan Hsu و Wee Sun Lee، ویراستاران، Proceedings of the Asian Conference on Machine Learning، جلد ۲۰ از Proceedings of Machine Learning Research، صفحات ۹۷–۱۱۲، South Garden Hotels and Resorts، تایوان، ۱۴–۱۵ نوامبر ۲۰۱۱ (۲۴ آبان ۱۳۹۰). PMLR. URL: https://proceedings.mlr.press/v20/biggio11.html،
doi:10.48550/arXiv.2206.00352.
- [40] Battista Biggio، Blaine Nelson و Pavel Laskov. حملات مسمومسازی علیه ماشینهای بردار پشتیبان. در Proceedings of the 29th International Coference on International Conference on Machine Learning, ICML، ۲۰۱۲. URL: https://arxiv.org/abs/1206.6389،
doi:10.48550/arXiv.1206.6389.
- [41] Battista Biggio، Konrad Rieck، Davide Ariu، Christian Wressnegger، Igino Corona، Giorgio Giacinto و Fabio Roli. مسمومسازی خوشهبندی رفتاری بدافزار. در Proceedings of the 2014 Workshop on Artificial Intelligent and Security Workshop، AISec ’14، صفحه ۲۷–۳۶، نیویورک، نیویورک، آمریکا، ۲۰۱۴. Association for Computing Machinery.
doi:10.1145/2666652.2666666.
- [42] Battista Biggio و Fabio Roli. الگوهای وحشی: ده سال پس از ظهور یادگیری ماشین خصمانه. Pattern Recognition، 84:317–331، دسامبر ۲۰۱۸ (آذر ۱۳۹۷ تا دی ۱۳۹۷). URL: https://doi.org/10.1016%2Fj.patcog.2018.07.023،
doi:10.1016/j.patcog.2018.07.023.
- [43] Peva Blanchard، El Mahdi El Mhamdi، Rachid Guerraoui و Julien Stainer. یادگیری ماشین با دشمنان: نزول گرادیان بردپذیر بیزانسی. در NeurIPS، ۲۰۱۷.
URL: https://papers.nips.cc/paper_files/paper/2017/file/f4b9ec30ad9f68f89b29639786cb62ef-Paper.pdf.
- [44] Rishi Bommasani، Kevin Klyman، Sayash Kapoor، Shayne Longpre، Betty Xiong، Nestor Maslej و Percy Liang. The foundation model transparency index v1.1: مه ۲۰۲۴ (اردیبهشت ۱۴۰۳ تا خرداد ۱۴۰۳)، ۲۰۲۴. URL: https://arxiv.org/abs/2407.12929،
arXiv:2407.12929،doi:10.48550/arXiv.2407.12929.
- [45] Lucas Bourtoule، Varun Chandrasekaran، Christopher A. Choquette-Choo، Hengrui Jia، Adelin Travers، Baiwu Zhang، David Lie و Nicolas Papernot. Machine unlearning. در 42nd IEEE Symposium on Security and Privacy, SP 2021, San Francisco, CA, USA, 24-۲۷ مه ۲۰۲۱ (۶ خرداد ۱۴۰۰)، صفحات 141–159. IEEE، ۲۰۲۱.
doi:10.1109/SP40001.2021.00019.
- [46] Dillon Bowen، Brendan Murphy، Will Cai، David Khachaturov، Adam Gleave و Kellin Pelrine. Scaling laws for data poisoning in llms، ۲۰۲۴. URL: https://arxiv.org/abs/2408.02946،
arXiv:2408.02946،doi:10.48550/arXiv.2408.02946.
- [47] Wieland Brendel، Jonas Rauber و Matthias Bethge. Decision-based adversarial attacks: Reliable attacks against black-box machine learning models. در 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - ۳ مه ۲۰۱۸ (۱۳ اردیبهشت ۱۳۹۷), Conference Track Proceedings. OpenReview.net، ۲۰۱۸. URL: https://openreview.net/forum?id=SyZI0GWCZ،
doi:10.48550/arXiv.1712.04248.
- [48] Gavin Brown، Mark Bun، Vitaly Feldman، Adam Smith و Kunal Talwar. When is memorization of irrelevant training data necessary for high-accuracy learning? در Proceedings of the 53rd Annual ACM SIGACT Symposium on Theory of Computing، STOC 2021، صفحات 123–132، نیویورک، ایالات متحده، ۲۰۲۱. انجمن «ماشینهای محاسباتی» (Association for Computing Machinery, ACM).
doi:10.1145/3406325.3451131.
- [49] Tom B. Brown، Benjamin Mann، Nick Ryder، Melanie Subbiah، Jared Kaplan، Prafulla Dhariwal، Arvind Neelakantan، Pranav Shyam، Girish Sastry، Amanda Askell، Sandhini Agarwal، Ariel Herbert-Voss، Gretchen Krueger، Tom Henighan، Rewon Child، Aditya Ramesh، Daniel M. Ziegler، Jeffrey Wu، Clemens Winter، Christopher Hesse، Mark Chen، Eric Sigler، Mateusz Litwin، Scott Gray، Benjamin Chess، Jack Clark، Christopher Berner، Sam McCandlish، Alec Radford، Ilya Sutskever و Dario Amodei. Language models are few-shot learners. CoRR، abs/2005.14165، ۲۰۲۰. URL: https://arxiv.org/abs/2005.14165،
arXiv:2005.14165.
- [50] Gon Buzaglo، Niv Haim، Gilad Yehudai، Gal Vardi و Michal Irani. Reconstructing training data from multiclass neural networks، ۲۰۲۳. URL: https://arxiv.org/abs/2305.03350،
arXiv:2305.03350،doi:10.48550/arXiv.2305.03350.
- [51] Xiaoyu Cao، Minghong Fang، Jia Liu و Neil Zhenqiang Gong. FLTrust: Byzantinerobust federated learning via trust bootstrapping. در NDSS، ۲۰۲۱. URL: https://arxiv.org/abs/2012.13995،
doi:10.48550/arXiv.2012.13995.
- [52] Yinzhi Cao و Junfeng Yang. Towards making systems forget with machine unlearning. در 2015 IEEE Symposium on Security and Privacy، صفحات 463–480، ۲۰۱۵. URL: https://ieeexplore.ieee.org/document/7163042،
doi:10.1109/SP.2015.35.
- [53] Nicholas Carlini. Poisoning the unlabeled dataset of Semi-Supervised learning. در 30th USENIX Security Symposium (USENIX Security 21)، صفحات 1577–1592. USENIX Association، اوت ۲۰۲۱ (مرداد ۱۴۰۰ تا شهریور ۱۴۰۰). URL: https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-poisoning.
- [54] Nicholas Carlini، Steve Chien، Milad Nasr، Shuang Song، Andreas Terzis و Florian Tramer. Membership inference attacks from first principles. در 2022 IEEE Symposium on Security and Privacy (S&P)، صفحات 1519–1519، Los Alamitos, CA, USA، مه ۲۰۲۲ (اردیبهشت ۱۴۰۱ تا خرداد ۱۴۰۱). IEEE Computer Society. URL: https://doi.ieeecomputersociety.org/10.1109/SP46214.2022.00090،
doi:10.1109/SP46214.2022.00090.
- [55] Nicholas Carlini، Jamie Hayes، Milad Nasr، Matthew Jagielski، Vikash Sehwag، Florian Tramèr، Borja Balle، Daphne Ippolito و Eric Wallace. Extracting training data from diffusion models، ۲۰۲۳. URL: https://arxiv.org/abs/2301.13188،
arXiv:2301.13188،doi:10.48550/arXiv.2301.13188.
- [56] Nicholas Carlini، Daphne Ippolito، Matthew Jagielski، Katherine Lee، Florian Tramer و Chiyuan Zhang. Quantifying memorization across neural language models. https://arxiv.org/abs/2202.07646، ۲۰۲۲.
doi:10.48550/ARXIV.2202.07646.
- [57] Nicholas Carlini، Matthew Jagielski، Christopher A Choquette-Choo، Daniel Paleka، Will Pearce، Hyrum Anderson، Andreas Terzis، Kurt Thomas و Florian Tramèr. Poisoning web-scale training datasets is practical. arXiv preprint arXiv:2302.10149، ۲۰۲۳. URL: https://arxiv.org/abs/2302.10149،
doi:10.48550/arXiv.2302.10149.
- [58] Nicholas Carlini، Matthew Jagielski و Ilya Mironov. Cryptanalytic extraction of neural network models. در Daniele Micciancio و Thomas Ristenpart، ویراستاران، Advances in Cryptology – CRYPTO 2020، صفحات 189–218، Cham، ۲۰۲۰. Springer International Publishing. URL: https://arxiv.org/abs/2003.04884،
doi:10.48550/arXiv.2003.04884.
- [59] Nicholas Carlini، Chang Liu، Úlfar Erlingsson، Jernej Kos و Dawn Song. The Secret Sharer: Evaluating and testing unintended memorization in neural networks. در USENIX Security Symposium، USENIX ’19)، صفحات 267–284، ۲۰۱۹. https://arxiv.org/abs/1802.08232. URL: https://arxiv.org/abs/1802.08232،
doi:10.48550/arXiv.1802.08232.
- [60] Nicholas Carlini، Milad Nasr، Christopher A Choquette-Choo، Matthew Jagielski، Irena Gao، Anas Awadalla، Pang Wei Koh، Daphne Ippolito، Katherine Lee، Florian Tramer و همکاران. Are aligned neural networks adversarially aligned? arXiv preprint arXiv:2306.15447، ۲۰۲۳. URL: https://arxiv.org/abs/2306.15447،
doi:10.48550/arXiv.2306.15447.
- [61] Nicholas Carlini، Daniel Paleka، Krishnamurthy Dj Dvijotham، Thomas Steinke، Jonathan Hayase، A. Feder Cooper، Katherine Lee، Matthew Jagielski، Milad Nasr، Arthur Conmy، Itay Yona، Eric Wallace، David Rolnick و Florian Tramèr. Stealing part of a production language model، ۲۰۲۴. URL: https://arxiv.org/abs/2403.06634،
arXiv:2403.06634،doi:10.48550/arXiv.2403.06634.
- [62] Nicholas Carlini، Florian Tramer، Krishnamurthy Dj Dvijotham، Leslie Rice، Mingjie Sun و J. Zico Kolter. (certified!!) adversarial robustness for free!، ۲۰۲۳. URL: https://arxiv.org/abs/2206.10550،
arXiv:2206.10550،doi:10.48550/arXiv.2206.10550.
- [63] Nicholas Carlini، Florian Tramèr، Eric Wallace، Matthew Jagielski، Ariel Herbert-Voss، Katherine Lee، Adam Roberts، Tom Brown، Dawn Song، Úlfar Erlingsson، Alina Oprea و Colin Raffel. Extracting training data from large language models. در 30th USENIX Security Symposium (USENIX Security 21)، صفحات 2633–2650. USENIX Association، اوت ۲۰۲۱ (مرداد ۱۴۰۰ تا شهریور ۱۴۰۰). URL: https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting.
- [64] Nicholas Carlini و David Wagner. Adversarial examples are not easily detected: Bypassing ten detection methods. در Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security، AISec ’17، صفحات 3–14، نیویورک، نیویورک، آمریکا، ۲۰۱۷. Association for Computing Machinery.
doi:10.1145/3128572.3140444.
- [65] Nicholas Carlini و David Wagner. Towards evaluating the robustness of neural networks. در Proc. IEEE Security and Privacy Symposium، ۲۰۱۷. URL: https://arxiv.org/abs/1608.04644،
doi:10.48550/arXiv.1608.04644.
- [66] Nicholas Carlini و David Wagner. Audio adversarial examples: Targeted attacks on speech-to-text. در 2018 IEEE Security and Privacy Workshops (SPW)، صفحات 1–7. IEEE، ۲۰۱۸. URL: https://arxiv.org/abs/1801.01944،
doi:10.48550/arXiv.1801.01944.
- [67] Stephen Casper، Yuxiao Li، Jiawei Li، Tong Bu، Kevin Zhang، Kaivalya Hariharan و Dylan Hadfield-Menell. Red teaming deep neural networks with feature synthesis tools. arXiv preprint arXiv:2302.10894، ۲۰۲۳. URL: https://arxiv.org/abs/2302.10894،
doi:10.48550/arXiv.2302.10894.
- [68] Stephen Casper، Jason Lin، Joe Kwon، Gatlen Culp و Dylan Hadfield-Menell. Explore, establish, exploit: Red teaming language models from scratch، ۲۰۲۳. URL: https://arxiv.org/abs/2306.09442،
arXiv:2306.09442،doi:10.48550/arXiv.2306.09442.
- [69] National Cyber Security Center. Introducing our new machine learning security principles، بازیابیشده در فوریهٔ ۲۰۲۳ از https://www.ncsc.gov.uk/blog-post/introducing-our-new-machine-learning-security-principles. URL: https://www.ncsc.gov.uk/blog-post/introducing-our-new-machine-learning-security-principles.
- [70] Varun Chandrasekaran، Kamalika Chaudhuri، Irene Giacomelli، Somesh Jha و Songbai Yan. Exploring connections between active learning and model extraction. در Proceedings of the 29th USENIX Conference on Security Symposium، SEC’20، آمریکا، ۲۰۲۰. USENIX Association. URL: https://arxiv.org/abs/1811.02054،
doi:10.48550/arXiv.1811.02054.
- [71] Hong Chang، Ta Duy Nguyen، Sasi Kumar Murakonda، Ehsan Kazemi و R. Shokri. On adversarial bias and the robustness of fair machine learning. https://arxiv.org/abs/2006.08669، ۲۰۲۰.
- [72] Patrick Chao، Edoardo Debenedetti، Alexander Robey، Maksym Andriushchenko، Francesco Croce، Vikash Sehwag، Edgar Dobriban، Nicolas Flammarion، George J. Pappas، Florian Tramer، Hamed Hassani و Eric Wong. JailbreakBench: An open robustness benchmark for jailbreaking large language models، ۲۰۲۴. URL: https://arxiv.org/abs/2404.01318،
arXiv:2404.01318،doi:10.48550/arXiv.2404.01318.
- [73] Patrick Chao، Alexander Robey، Edgar Dobriban، Hamed Hassani، George J Pappas و Eric Wong. Jailbreaking black box large language models in twenty queries. arXiv preprint arXiv:2310.08419، ۲۰۲۳. URL: https://arxiv.org/abs/2310.08419،
doi:10.48550/arXiv.2310.08419.
- [74] Harsh Chaudhari، John Abascal، Alina Oprea، Matthew Jagielski، Florian Tramèr و Jonathan Ullman. SNAP: Efficient extraction of private properties with poisoning. در 2023 IEEE Symposium on Security and Privacy (S&P)، ۲۰۲۳. URL: https://arxiv.org/abs/2208.12348،
doi:10.48550/arXiv.2208.12348.
- [75] Harsh Chaudhari، Giorgio Severi، John Abascal، Matthew Jagielski، Christopher A. Choquette-Choo، Milad Nasr، Cristina Nita-Rotaru و Alina Oprea. Phantom: General trigger attacks on retrieval augmented language generation، ۲۰۲۴. URL: https://arxiv.org/abs/2405.20485،
arXiv:2405.20485،doi:10.48550/arXiv.2405.20485.
- [76] Bryant Chen، Wilka Carvalho، Nathalie Baracaldo، Heiko Ludwig، Benjamin Edwards، Taesung Lee، Ian Molloy و Biplav Srivastava. Detecting backdoor attacks on deep neural networks by activation clustering. https://arxiv.org/abs/1811.03728، ۲۰۱۸. URL: https://arxiv.org/abs/1811.03728،
doi:10.48550/arXiv.1811.03728.
- [77] Hongge Chen، Huan Zhang، Pin-Yu Chen، Jinfeng Yi و Cho-Jui Hsieh. Attacking visual language grounding with adversarial examples: A case study on neural image captioning. https://arxiv.org/abs/1712.02051، ۲۰۱۷. URL: https://arxiv.org/abs/1712.02051،
doi:10.48550/ARXIV.1712.02051.
- [78] Huili Chen، Cheng Fu، Jishen Zhao و Farinaz Koushanfar. DeepInspect: A black-box trojan detection and mitigation framework for deep neural networks. در Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19، صفحات 4658–4664. International Joint Conferences on Artificial Intelligence Organization، ژوئیهٔ ۲۰۱۹.
doi:10.24963/ijcai.2019/647.
- [79] Jianbo Chen، Michael I. Jordan و Martin J. Wainwright. HopSkipJumpAttack: A query-efficient decision-based attack. در 2020 IEEE Symposium on Security and Privacy, SP 2020, San Francisco, CA, USA, May 18-21, 2020، صفحات 1277–1294. IEEE، ۲۰۲۰.
doi:10.1109/SP40000.2020.00045.
- [80] Pin-Yu Chen، Huan Zhang، Yash Sharma، Jinfeng Yi و Cho-Jui Hsieh. Zoo: Zeroth order optimization based black-box attacks to deep neural networks without training substitute models. در Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security، AISec ’17، صفحات 15–26، نیویورک، نیویورک، آمریکا، ۲۰۱۷. Association for Computing Machinery.
doi:10.1145/3128572.3140448.
- [81] Shang-Tse Chen، Cory Cornelius، Jason Martin و Duen Horng Chau. ShapeShifter: Robust Physical Adversarial Attack on Faster R-CNN Object Detector، صفحه 52–68. Springer International Publishing، 2019. URL: http://dx.doi.org/10.1007/978-3-030-10925-7_4،
doi:10.1007/978-3-030-10925-7_4.
- [82] Xiaoyi Chen، Ahmed Salem، Dingfan Chen، Michael Backes، Shiqing Ma، Qingni Shen، Zhonghai Wu و Yang Zhang. Badnl: حملههای «درِ پشتی» (Backdoor) علیه مدلهای NLP با بهبودهای حفظ معنایی. در Annual Computer Security Applications Conference، ACSAC ’21، صفحات 554–569، نیویورک، NY، USA، 2021. Association for Computing Machinery.
doi:10.1145/3485832.3485837.
- [83] Xiaoyi Chen، Siyuan Tang، Rui Zhu، Shijun Yan، Lei Jin، Zihao Wang، Liya Su، Zhikun Zhang، XiaoFeng Wang و Haixu Tang. رابط Janus: چگونه تنظیم دقیق در «مدلهای زبانی بزرگ» (Large Language Models، LLM) خطرهای «حریم خصوصی» (Privacy) را تشدید میکند، 2024. URL: https://arxiv.org/abs/2310.15469،
arXiv:2310.15469،doi:10.48550/arXiv.2310.15469.
- [84] Xinyun Chen، Chang Liu، Bo Li، Kimberly Lu و Dawn Song. حملههای هدفمند درِ پشتی بر سیستمهای یادگیری عمیق با استفاده از مسمومسازی داده. arXiv preprint arXiv:1712.05526، 2017. URL: https://arxiv.org/abs/1712.05526،
doi:10.48550/arXiv.1712.05526.
- [85] Heng-Tze Cheng و Romal Thoppilan. LaMDA: Towards Safe, Grounded, and High-Quality Dialog Models for Everything. https://ai.googleblog.com/2022/01/lamda-towards-safe-grounded-and-high.html، 2022. Google Brain. URL: https://research.google/blog/lamda-towards-safe-grounded-and-high-quality-dialog-models-for-everything/.
- [86] Minhao Cheng، Thong Le، Pin-Yu Chen، Huan Zhang، Jinfeng Yi و Cho-Jui Hsieh. حمله جعبهسیاه برچسبسخت با پرسوجوی بهینه: رویکردی مبتنی بر بهینهسازی. در 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net، 2019. URL: https://openreview.net/forum?id=rJlk6iRqKX،
doi:10.48550/arXiv.1807.04457.
- [87] Minhao Cheng، Simranjit Singh، Patrick H. Chen، Pin-Yu Chen، Sijia Liu و Cho-Jui Hsieh. Sign-opt: حمله تخاصمی برچسبسخت با پرسوجوی بهینه. در International Conference on Learning Representations، 2020. URL: https://openreview.net/forum?id=SklTQCNtvS،
doi:10.48550/arXiv.1909.10773.
- [88] Alesia Chernikova و Alina Oprea. FENCE: حملههای «گریز» (Evasion) شدنی بر شبکههای عصبی در محیطهای مقیدشده. ACM Transactions on Privacy and Security (TOPS) Journal، 2022. URL: https://arxiv.org/abs/1909.10480،
doi:10.48550/arXiv.1909.10480.
- [89] Christopher A. Choquette-Choo، Florian Tramer، Nicholas Carlini و Nicolas Papernot. حملههای «استنتاج عضویت» (membership inference) فقطبرچسبی. در Marina Meila و Tong Zhang، ویراستاران، Proceedings of the 38th International Conference on Machine Learning، جلد 139 از Proceedings of Machine Learning Research، صفحات 1964–1974. PMLR، 18–24 Jul 2021. URL: https://proceedings.mlr.press/v139/choquette-choo21a.html،
doi:10.48550/arXiv.2007.14321.
- [90] Sheng-Yen Chou، Pin-Yu Chen و Tsung-Yi Ho. چگونه مدلهای انتشاری را درِ پشتی دار کنیم؟ https://arxiv.org/abs/2212.05400، 2022.
doi:10.48550/ARXIV.2212.05400.
- [91] Antonio Emanuele Cinà، Kathrin Grosse، Ambra Demontis، Sebastiano Vascon، Werner Zellinger، Bernhard A. Moser، Alina Oprea، Battista Biggio، Marcello Pelillo و Fabio Roli. Wild patterns reloaded: مروری بر «امنیت» (Security) «یادگیری ماشین» (Machine Learning، ML) در برابر مسمومسازی دادههای آموزشی. ACM Computing Surveys، مارس ۲۰۲۳ (اسفند ۱۴۰۱ تا فروردین ۱۴۰۲). URL: https://doi.org/10.1145%2F3585385،
doi:10.1145/3585385.
- [92] Jack Clark و Raymond Perrault. 2022 AI index report. https://aiindex.stanford.edu/wp-content/uploads/2022/03/2022-AI-Index-Report_Master.pdf، 2022. Human Centered AI، Stanford University.
- [93] Joseph Clements، Yuzhe Yang، Ankur Sharma، Hongxin Hu و Yingjie Lao. بسیج فنون تخاصمی علیه یادگیری عمیق برای امنیت شبکه، 2019. URL: https://arxiv.org/abs/1903.11688،
doi:10.48550/ARXIV.1903.11688.
- [94] Jeremy Cohen، Elan Rosenfeld و Zico Kolter. «استواری» (Robustness) تخاصمی گواهیشده از طریق هموارسازی تصادفی. در Kamalika Chaudhuri و Ruslan Salakhutdinov، ویراستاران، Proceedings of the 36th International Conference on Machine Learning، جلد 97 از Proceedings of Machine Learning Research، صفحات 1310–1320. PMLR، 09–15 Jun 2019. URL: https://proceedings.mlr.press/v97/cohen19c.html.
- [95] Jeremy Cohen، Elan Rosenfeld و Zico Kolter. استواری تخاصمی گواهیشده از طریق هموارسازی تصادفی. در International Conference on Machine Learning، صفحات 1310–1320. PMLR، 2019.
- [96] Gabriela F. Cretu، Angelos Stavrou، Michael E. Locasto، Salvatore J. Stolfo و Angelos D. Keromytis. اخراج شیاطین: پاکسازی دادههای آموزشی برای حسگرهای ناهنجاری. در 2008 IEEE Symposium on Security and Privacy (sp 2008)، صفحات 81–95، 2008. URL: https://ieeexplore.ieee.org/document/4531146،
doi:10.1109/SP.2008.11.
- [97] Francesco Croce، Maksym Andriushchenko، Vikash Sehwag، Edoardo Debenedetti، Nicolas Flammarion، Mung Chiang، Prateek Mittal و Matthias Hein. Robustbench: یک معیار مرجع استانداردشده استواری تخاصمی. در Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)، 2021. URL: https://openreview.net/forum?id=SSKZPJCt7B،
doi:10.48550/arXiv.2010.09670.
- [98] Nilesh Dalvi، Pedro Domingos، Mausam، Sumit Sanghai و Deepak Verma. ردهبندی تخاصمی. در Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining، KDD ’04، صفحات 99–108، نیویورک، NY، USA، 2004. Association for Computing Machinery.
doi:10.1145/1014052.1014066.
- [99] DARPA. DARPA AI Cyber Challenge Aims to Secure Nation’s Most Critical Software، 2023. تاریخ دسترسی: 2024-08-22 (۱ شهریور ۱۴۰۳). URL: https://www.darpa.mil/news-events/2023-08-09.
- [100] Emiliano De Cristofaro. مرور انتقادی حریم خصوصی در یادگیری ماشین. IEEE Security & Privacy، 19(4):19–27، 2021.
doi:10.1109/MSEC.2021.3076443.
استناد
محمدعلی کهندژ، راهنمای جامع حاکمیت، امنیت و مدیریت ریسک هوش مصنوعی، شناسه بخش: KDJ-AI-2026E1-P03-C22-S11
شناسهٔ محتوا KDJ-AI-2026E1-P03-C22-S11-89B3BCB4