فهرست این فصلبخش ۳، فصل ۲۲ — صفحهٔ ۵ از ۷
بخش ۳، فصل ۲۲ — صفحهٔ ۵ از ۷
چالشهای کلیدی و بحث
متن اصلی فارسی با منشأ، شناسه و پیوند استناد پایدار.
ادامهٔ بخش «مراجع» — بخش ۴ از ۵
- [287] Dario Pasquini، Martin Strohmeier، و Carmela Troncoso. «Neural Exec: یادگیری (و یادگیری از) محرکهای اجرا برای حملههای تزریق دستور (Prompt Injection)»، ۲۰۲۴. URL: https://arxiv.org/abs/2403.03792،
arXiv:2403.03792،doi:10.48550/arXiv.2403.03792.
- [288] Arpita Patra، Thomas Schneider، Ajith Suresh و Hossein Yalame. ABY2.0: Improved Mixed-Protocol secure Two-Party computation. در سیامین USENIX Security Symposium (USENIX Security 21)، صفحات ۲۱۶۵–۲۱۸۲. USENIX Association، اوت ۲۰۲۱ (مرداد ۱۴۰۰ تا شهریور ۱۴۰۰). URL: https://www.usenix.org/conference/usenixsecurity21/presentation/patra.
- [289] Andrea Paudice، Luis Muñoz-González و Emil C. Lupu. Label sanitization against label flipping poisoning attacks. در Carlos Alzate، Anna Monreale، Haytham Assem، Albert Bifet، Teodora Sandra Buda، Bora Caglayan، Brett Drury، Eva García-Martín، Ricard Gavaldà، Stefan Kramer، Niklas Lavesson، Michael Madden، Ian Molloy، Maria-Irina Nicolae و Mathieu Sinn، ویراستاران، Nemesis/UrbReas/SoGood/IWAISe/GDM@PKDD/ECML، جلد ۱۱۳۲۹ از Lecture Notes in Computer Science، صفحات ۵–۱۵. Springer، ۲۰۱۸. URL: http://dblp.uni-trier.de/db/conf/pkdd/nemesis2018.html#PaudiceML18.
- [290] Hammond Pearce، Baleegh Ahmad، Benjamin Tan، Brendan Dolan-Gavitt و Ramesh Karri. Asleep at the keyboard? Assessing the security of GitHub Copilot’s code contributions، ۲۰۲۱. URL: https://arxiv.org/abs/2108.09293،
arXiv:2108.09293،doi:10.48550/arXiv.2108.09293.
- [291] R. Perdisci، D. Dagon، Wenke Lee، P. Fogla و M. Sharif. Misleading worm signature generators using deliberate noise injection. در 2006 IEEE Symposium on Security and Privacy (S&P’06)، Berkeley/Oakland, CA، ۲۰۰۶. IEEE. URL: http://ieeexplore.ieee.org/document/1623998/،
doi:10.1109/SP.2006.26.
- [292] Ethan Perez، Saffron Huang، Francis Song، Trevor Cai، Roman Ring، John Aslanides، Amelia Glaese، Nat McAleese و Geoffrey Irving. تیم قرمز (Red Teaming) کردن مدلهای زبانی با مدلهای زبانی. پیشچاپ arXiv با شناسه arXiv:2202.03286، ۲۰۲۲. URL: https://arxiv.org/abs/2202.03286،
doi:10.48550/arXiv.2202.03286.
- [293] Neehar Peri، Neal Gupta، W. Ronny Huang، Liam Fowl، Chen Zhu، Soheil Feizi، Tom Goldstein و John P. Dickerson. Deep k-NN defense against clean-label data poisoning attacks. در Adrien Bartoli و Andrea Fusiello، ویراستاران، Computer Vision – ECCV 2020 Workshops، صفحات ۵۵–۷۰، Cham، ۲۰۲۰. Springer International Publishing. URL: https://arxiv.org/abs/1909.13374،
doi:10.48550/arXiv.1909.13374.
- [294] Neil Perry، Megha Srivastava، Deepak Kumar و Dan Boneh. Do users write more insecure code with AI assistants? در Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security، CCS ’23. ACM، نوامبر ۲۰۲۳ (آبان ۱۴۰۲ تا آذر ۱۴۰۲). URL: http://dx.doi.org/10.1145/3576915.3623157،
doi:10.1145/3576915.3623157.
- [295] Fabio Pierazzi، Feargus Pendlebury، Jacopo Cortellazzi و Lorenzo Cavallaro. Intriguing properties of adversarial ML attacks in the problem space. در 2020 IEEE Symposium on Security and Privacy (S&P)، صفحات ۱۳۰۸–۱۳۲۵. IEEE Computer Society، ۲۰۲۰. URL: https://doi.ieeecomputersociety.org/10.1109/SP40000.2020.00073،
doi:10.1109/SP40000.2020.00073.
- [296] Julien Piet، Maha Alrashed، Chawin Sitawarin، Sizhe Chen، Zeming Wei، Elizabeth Sun، Basel Alomair و David Wagner. Jatmo: تزریق دستور (Prompt Injection) defense by task-specific finetuning، ۲۰۲۴. URL: https://arxiv.org/abs/2312.17673،
arXiv:2312.17673،doi:10.48550/arXiv.2312.17673.
- [297] Krishna Pillutla، Galen Andrew، Peter Kairouz، H. Brendan McMahan، Alina Oprea و Sewoong Oh. Unleashing the power of randomization in auditing differentially private ML. در Advances in Neural Information Processing Systems، ۲۰۲۳. URL: https://arxiv.org/abs/2305.18447،
doi:10.48550/arXiv.2305.18447.
- [298] PromptArmor و Kai Greshake. Data exfiltration from writer.com with indirect prompt injection، ۲۰۲۳. URL: https://promptarmor.substack.com/p/data-exfiltration-from-writercom.
- [299] Jonathan Protzenko، Bryan Parno، Aymeric Fromherz، Chris Hawblitzel، Marina Polubelova، Karthikeyan Bhargavan، Benjamin Beurdouche، Joonwon Choi، Antoine Delignat-Lavaud، Cédric Fournet، Natalia Kulatova، Tahina Ramananandro، Aseem Rastogi، Nikhil Swamy، Christoph Wintersteiger و Santiago Zanella-Beguelin. EverCrypt: A fast, verified, cross-platform cryptographic provider. در Proceedings of the IEEE Symposium on Security and Privacy (Oakland)، مه ۲۰۲۰ (اردیبهشت ۱۳۹۹ تا خرداد ۱۳۹۹). URL: https://eprint.iacr.org/2019/757،
doi:10.1109/SP40000.2020.00114.
- [300] Xiangyu Qi، Yi Zeng، Tinghao Xie، Pin-Yu Chen، Ruoxi Jia، Prateek Mittal و Peter Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to!، ۲۰۲۳. URL: https://arxiv.org/abs/2310.03693،
arXiv:2310.03693.
- [301] Gauthama Raman M. R.، Chuadhry Mujeeb Ahmed و Aditya Mathur. یادگیری ماشین (ML) for intrusion detection in industrial control systems: Challenges and lessons from experimental evaluation. Cybersecurity، ۴(۲۷)، ۲۰۲۱. URL: https://arxiv.org/abs/2202.11917،
doi:10.48550/arXiv.2202.11917.
- [302] Aida Rahmattalabi، Shahin Jabbari، Himabindu Lakkaraju، Phebe Vayanos، Max Izenberg، Ryan Brown، Eric Rice و Milind Tambe. Fair influence maximization: A welfare optimization approach. در Proceedings of the AAAI Conference on Artificial Intelligence 35th، ۲۰۲۱. URL: https://arxiv.org/abs/2006.07906،
doi:10.48550/arXiv.2006.07906.
- [303] Adnan Siraj Rakin، Md Hafizul Islam Chowdhuryy، Fan Yao و Deliang Fan. DeepSteal: Advanced model extractions leveraging efficient weight stealing in memories. در 2022 IEEE Symposium on Security and Privacy (S&P)، صفحات ۱۱۵۷–۱۱۷۴، ۲۰۲۲. URL: https://ieeexplore.ieee.org/document/9833743،
doi:10.1109/SP46214.2022.9833743.
- [304] Dhanesh Ramachandram و Graham W. Taylor. Deep multimodal learning: A survey on recent advances and trends. IEEE Signal Processing Magazine، ۳۴(۶):۹۶–۱۰۸، ۲۰۱۷. URL: https://ieeexplore.ieee.org/document/8103116،
doi:10.1109/MSP.2017.2738401.
- [305] Javier Rando و Florian Tramèr. دور زدن محدودیتهای مدل همگانی (Universal Jailbreak) backdoors from poisoned human feedback، ۲۰۲۴. URL: https://arxiv.org/abs/2311.14455،
arXiv:2311.14455،doi:10.48550/arXiv.2311.14455.
- [306] Traian Rebedea، Razvan Dinu، Makesh Sreedhar، Christopher Parisien و Jonathan Cohen. «NeMo Guardrails: A toolkit for controllable and safe LLM applications with programmable rails» (جعبهابزاری برای برنامههای کاربردی LLM قابلکنترل و ایمن با ریلهای برنامهپذیر)، ۲۰۲۳. URL: https://arxiv.org/abs/2310.10501،
arXiv:2310.10501،doi:10.48550/arXiv.2310.10501.
- [307] Johann Rehberger. «Data exfiltration via markdown injection - exploiting chatgpt's webpilot plugin» (برونکشی داده از طریق تزریق markdown — بهرهکشی از افزونه webpilot در ChatGPT)، ۱۶ مه ۲۰۲۳ (۲۶ اردیبهشت ۱۴۰۲). Embrace The Red، تاریخ دسترسی: 2024-08-18 (۲۸ مرداد ۱۴۰۳). URL: https://embracethered.com/blog/posts/2023/chatgpt-webpilot-data-exfil-via-markdown-injection/.
- [308] SNYK Report. «AI Code Security and Trust: Organizations must change their approach» (امنیت و اعتماد در کد AI: سازمانها باید رویکرد خود را تغییر دهند). https://go.snyk.io/2023-ai-code-security-report-dwn-typ.html?aliId=eyJpIjoiUDFvdzRSdHI0dm5rVktvSSIsInQiOiJxOElRU2dQdkdqQm03ZjNLSDFFVkxBPT0ifQ%253D%253D، ۲۰۲۳. Human Centered AI، دانشگاه استنفورد.
- [309] Maria Rigaki و Sebastian Garcia. «A survey of privacy attacks in machine learning» (مروری بر حملات حریم خصوصی در یادگیری ماشین). ACM Comput. Surv.، 56(4)، نوامبر ۲۰۲۳ (آبان ۱۴۰۲ تا آذر ۱۴۰۲).
doi:10.1145/3624010.
- [310] Maria Rigaki و Sebastian Garcia. «A survey of privacy attacks in machine learning» (مروری بر حملات حریم خصوصی در یادگیری ماشین). ACM Comput. Surv.، 56(4)، نوامبر ۲۰۲۳ (آبان ۱۴۰۲ تا آذر ۱۴۰۲).
doi:10.1145/3624010.
- [311] Rishi Bommasani و همکاران. «On the opportunities and risks of foundation models» (درباره فرصتها و خطرهای مدلهای پایه)، ۲۰۲۴. URL: https://arxiv.org/abs/2108.07258،
arXiv:2108.07258.
- [312] Alexander Robey، Eric Wong، Hamed Hassani و George J Pappas. «SmoothLLM: Defending large language models against jailbreaking attacks» (SmoothLLM: دفاع از مدلهای زبانی بزرگ در برابر حملات دور زدن محدودیتهای مدل). ArXiv، abs/2310.03684، ۲۰۲۳. URL: https://api.semanticscholar.org/CorpusID:263671542،
doi:10.48550/arXiv.2310.03684.
- [313] Robust Intelligence. AI Firewall، ۲۰۲۳. URL: https://www.robustintelligence.com/platform/ai-firewall.
- [314] Elan Rosenfeld، Ezra Winston، Pradeep Ravikumar و Zico Kolter. «Certified robustness to label-flipping attacks via randomized smoothing» (استواری گواهیشده در برابر حملات وارونهسازی برچسب از طریق هموارسازی تصادفی). در International Conference on Machine Learning، صفحات 8230–8241. PMLR، ۲۰۲۰. URL: https://arxiv.org/abs/1902.02918،
doi:10.48550/arXiv.1902.02918.
- [315] Benjamin IP Rubinstein، Blaine Nelson، Ling Huang، Anthony D Joseph، Shing-hon Lau، Satish Rao، Nina Taft و J Doug Tygar. «Antidote: understanding and defending against poisoning of anomaly detectors» (Antidote: فهم و دفاع در برابر مسمومسازی آشکارسازهای ناهنجاری). در Proceedings of the 9th ACM SIGCOMM conference on Internet measurement، صفحات 1–14، ۲۰۰۹.
doi:10.1145/1644893.1644895.
- [316] Mark Russinovich، Ahmed Salem و Ronen Eldan. «Great, now write an article about that: The crescendo multi-turn LLM jailbreak attack» (عالی، اکنون درباره آن مقالهای بنویس: حمله چندنوبتی Crescendo برای دور زدن محدودیتهای مدل LLM). ArXiv، abs/2404.01833، ۲۰۲۴. URL: https://api.semanticscholar.org/CorpusID:268856920،
doi:10.48550/arXiv.2404.01833.
- [317] Alexandre Sablayrolles، Matthijs Douze، Cordelia Schmid، Yann Ollivier و Hervé Jégou. «White-box vs black-box: Bayes optimal strategies for membership inference» (جعبهسفید در برابر جعبهسیاه: راهبردهای بهینه بیزی برای استنتاج عضویت). در ICML، جلد 97 از Proceedings of Machine Learning Research، صفحات 5558–5567. PMLR، ۲۰۱۹. URL: https://arxiv.org/abs/1908.11229،
doi:10.48550/arXiv.1908.11229.
- [318] Carl Sabottke، Octavian Suciu و Tudor Dumitras. «Vulnerability disclosure in the age of social media: Exploiting Twitter for predicting real-world exploits» (افشای آسیبپذیری در عصر رسانههای اجتماعی: بهرهگیری از توییتر برای پیشبینی اکسپلویتهای دنیای واقعی). در 24th USENIX Security Symposium (USENIX Security 15)، صفحات 1041–1056، واشینگتن دیسی، اوت ۲۰۱۵ (مرداد ۱۳۹۴ تا شهریور ۱۳۹۴). USENIX Association. URL: https://www.usenix.org/conference/usenixsecurity15/technical-sessions/presentation/sabottke.
- [319] Vinu Sankar Sadasivan، Aounon Kumar، Sriram Balasubramanian، Wenxiao Wang و Soheil Feizi. «Can ai-generated text be reliably detected?» (آیا متن تولیدشده با هوش مصنوعی را میتوان بهطور قابلاعتماد شناسایی کرد؟)، ۲۰۲۴. URL: https://arxiv.org/abs/2303.11156،
arXiv:2303.11156،doi:10.48550/arXiv.2303.11156.
- [320] Vinu Sankar Sadasivan، Shoumik Saha، Gaurang Sriramanan، Priyatham Kattakinda، Atoosa Chegini و Soheil Feizi. «Fast adversarial attacks on language models in one gpu minute» (حملات خصمانه سریع روی مدلهای زبانی در یک دقیقه GPU)، ۲۰۲۴. URL: https://arxiv.org/abs/2402.15570،
arXiv:2402.15570،doi:10.48550/arXiv.2402.15570.
- [321] Ahmed Salem، Giovanni Cherubin، David Evans، Boris Köpf، Andrew Paverd، Anshuman Suri، Shruti Tople و Santiago Zanella-Béguelin. «SoK: Let the privacy games begin! A unified treatment of data inference privacy in machine learning» (SoK: بگذارید بازیهای حریم خصوصی آغاز شود! برخورد یکپارچه با حریم خصوصی استنتاج داده در یادگیری ماشین). https://arxiv.org/abs/2212.10986، ۲۰۲۲.
doi:10.48550/ARXIV.2212.10986.
- [322] Ahmed Salem، Rui Wen، Michael Backes، Shiqing Ma و Yang Zhang. «Dynamic backdoor attacks against machine learning models» (حملات درِ پشتی پویا علیه مدلهای یادگیری ماشین). https://arxiv.org/abs/2003.03675، ۲۰۲۰.
doi:10.48550/ARXIV.2003.03675.
- [323] Roman Samoilenko. «New prompt injection attack on ChatGPT web version. markdown images can steal your chat data» (حمله جدید تزریق دستور روی نسخه وب ChatGPT؛ تصاویر markdown میتوانند دادههای گفتگوی شما را بربایند)، ۲۰۲۳. URL: https://systemweakness.com/new-prompt-injection-attack-on-chatgpt-web-version-ef717492c5c2.
- [324] Scale AI. «Adversarial robustness leaderboard» (جدول ردهبندی استواری خصمانه). https://scale.com/leaderboard/adversarial_robustness، ۲۰۲۴. تاریخ دسترسی: 2024-08-22 (۱ شهریور ۱۴۰۳).
- [325] Oscar Schwartz. «In 2016, Microsoft's racist chatbot revealed the dangers of online conversation: The bot learned language from people on Twitter—but it also learned values» (در سال ۲۰۱۶، چتبات نژادپرست مایکروسافت خطرهای گفتگوی برخط را آشکار کرد: این ربات زبان را از کاربران توییتر آموخت — اما ارزشها را نیز آموخت). https://spectrum.ieee.org/in-2016-microsofts-racist-chatbot-revealed-the-dangers-of-online-conversation، ۲۰۱۹. IEEE Spectrum.
- [326] R. Schwartz، A. Vassilev، K. Greene، L. Perine، A. Burt و P. Hall. «Towards a Standard for Identifying and Managing Bias in Artificial Intelligence» (بهسوی استانداردی برای شناسایی و مدیریت سوگیری در هوش مصنوعی). https://doi.org/10.6028/NIST.SP.1270، ۲۰۲۲. Special Publication (NIST SP) 800-1270، مؤسسه ملی استاندارد و فناوری آمریکا، گترزبورگ، مریلند. URL: https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.1270.pdf،
doi:10.6028/NIST.SP.1270.
- [327] Avi Schwarzschild، Micah Goldblum، Arjun Gupta، John P Dickerson و Tom Goldstein. «Just how toxic is data poisoning? A unified benchmark for backdoor and data poisoning attacks» (مسمومسازی داده چقدر سمّی است؟ یک معیار سنجش یکپارچه برای حملات درِ پشتی و مسمومسازی داده). https://arxiv.org/abs/2006.12557، ۲۰۲۰. arXiv.
doi:10.48550/ARXIV.2006.12557.
- [328] Leo Schwinn، David Dobre، Sophie Xhonneux، Gauthier Gidel و Stephan Gunnemann. «Soft prompt threats: Attacking safety alignment and unlearning in opensource LLMs through the embedding space» (تهدیدهای دستور نرم: حمله به همراستاسازی ایمنی و فرایادگیری در مدلهای زبانی بزرگ متنباز از طریق فضای embedding)، ۲۰۲۴. URL: https://arxiv.org/abs/2402.09063،
arXiv:2402.09063،doi:10.48550/arXiv.2402.09063.
- [329] Giorgio Severi، Jim Meyer، Scott Coull و Alina Oprea. «Explanation-guided backdoor poisoning attacks against malware classifiers» (حملات مسمومسازی درِ پشتی هدایتشده با توضیحپذیری علیه طبقهبندهای بدافزار). در 30th USENIX Security Symposium (USENIX Security 2021)، ۲۰۲۱. URL: https://www.usenix.org/conference/usenixsecurity21/presentation/severi.
- [330] Ali Shafahi، W Ronny Huang، Mahyar Najibi، Octavian Suciu، Christoph Studer، Tudor Dumitras، و Tom Goldstein. «قورباغههای مسموم! حملات مسمومسازی داده با برچسب تمیز و هدفمند بر شبکههای عصبی». در Advances in Neural Information Processing Systems، صفحات 6103–6113، ۲۰۱۸. URL: https://arxiv.org/abs/1804.00792،
doi:10.48550/arXiv.1804.00792.
- [331] Shawn Shan، Arjun Nitin Bhagoji، Haitao Zheng، و Ben Y. Zhao. «جرمیابی مسمومسازی: ردیابی حملات مسمومسازی داده در شبکههای عصبی». در 31st USENIX Security Symposium (USENIX Security 22)، صفحات 3575–3592، Boston, MA، اوت ۲۰۲۲ (مرداد ۱۴۰۱ تا شهریور ۱۴۰۱). USENIX Association. URL: https://www.usenix.org/conference/usenixsecurity22/presentation/shan.
- [332] Mahmood Sharif، Sruti Bhagavatula، Lujo Bauer، و Michael K. Reiter. «لوازم یک جنایت: حملات واقعی و پنهانکارانه بر بازشناسی چهرهٔ مدرن». در Proceedings of the 23rd ACM SIGSAC Conference on Computer and Communications Security، اکتبر ۲۰۱۶ (مهر ۱۳۹۵ تا آبان ۱۳۹۵). URL: https://www.ece.cmu.edu/~lbauer/papers/2016/ccs2016-face-recognition.pdf،
doi:10.1145/2976749.2978392.
- [333] Vasu Sharma، Ankita Kalra، Vaibhav، Simral Chaudhary، Labhesh Patel، و LP Morency. «توجه کن و حمله کن: حملات تخاصمیِ هدایتشده با مکانیزم توجه بر مدلهای پاسخ به پرسش تصویری». https://nips2018vigil.github.io/static/papers/accepted/33.pdf، ۲۰۱۸.
- [334] Ryan Sheatsley، Blaine Hoak، Eric Pauley، Yohan Beugin، Michael J. Weisman، و Patrick McDaniel. «دربارهٔ استواری قیود دامنه». در Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security، CCS ’21، صفحه 495–515، New York, NY, USA، ۲۰۲۱. Association for Computing Machinery.
doi:10.1145/3460120.3484570.
- [335] Virat Shejwalkar و Amir Houmansadr. «دستکاری بیزانسی: بهینهسازی حملات و دفاعهای مسمومسازی مدل برای یادگیری فدرال». در NDSS، ۲۰۲۱. URL: https://www.ndss-symposium.org/wp-content/uploads/ndss2021_6C-3_24498_paper.pdf.
- [336] Virat Shejwalkar، Amir Houmansadr، Peter Kairouz، و Daniel Ramage. «بازگشت به نقطهٔ آغاز: ارزیابی انتقادی حملات مسمومسازی بر یادگیری فدرال در محیط عملیاتی». در 43rd IEEE Symposium on Security and Privacy, SP 2022, San Francisco, CA, USA, May 22-26, 2022، صفحات 1354–1371. IEEE، ۲۰۲۲.
doi:10.1109/SP46214.2022.9833647.
- [337] Xinyue Shen، Zeyuan Chen، Michael Backes، Yun Shen، و Yang Zhang. «‹هر کاری بکن›: توصیف و ارزیابی پرامپتهای دور زدن محدودیتهای مدل در محیط واقعی بر مدلهای زبانی بزرگ». arXiv preprint arXiv:2308.03825، ۲۰۲۳. URL: https://arxiv.org/abs/2308.03825،
doi:10.48550/arXiv.2308.03825.
- [338] Xinyue Shen، Zeyuan Chen، Michael Backes، Yun Shen، و Yang Zhang. «‹هر کاری بکن›: توصیف و ارزیابی پرامپتهای دور زدن محدودیتهای مدل در محیط واقعی بر مدلهای زبانی بزرگ». CoRR، abs/2308.03825، ۲۰۲۳. URL: https://doi.org/10.48550/arXiv.2308.03825،
arXiv:2308.03825،doi:10.48550/ARXIV.2308.03825.
- [339] Xinyue Shen، Yiting Qu، Michael Backes، و Yang Zhang. «حملات سرقت پرامپت علیه مدلهای تولید متن به تصویر». arXiv preprint arXiv:2302.09923، ۲۰۲۳. URL: https://arxiv.org/abs/2302.09923،
doi:10.48550/arXiv.2302.09923.
- [340] Abhay Sheshadri، Aidan Ewart، Phillip Guo، Aengus Lynch، Cindy Wu، Vivek Hebbar، Henry Sleight، Asa Cooper Stickland، Ethan Perez، Dylan Hadfield-Menell، و Stephen Casper. «آموزش تخاصمی نهانِ هدفمند، استواری در برابر رفتارهای مضرّ پایدار را در LLMها بهبود میبخشد»، ۲۰۲۴. URL: https://arxiv.org/abs/2407.15549،
arXiv:2407.15549،doi:10.48550/arXiv.2407.15549.
- [341] Cong Shi، Tianfang Zhang، Zhuohang Li، Huy Phan، Tianming Zhao، Yan Wang، Jian Liu، Bo Yuan، و Yingying Chen. «حملهٔ درِ پشتی مستقل از موقعیت در دامنهٔ صوت از طریق محرکهای نامحسوس». در Proceedings of the 28th Annual International Conference on Mobile Computing And Networking، MobiCom ’22، صفحه 583–595، New York, NY, USA، ۲۰۲۲. Association for Computing Machinery.
doi:10.1145/3495243.3560531.
- [342] Reza Shokri، Marco Stronati، Congzheng Song، و Vitaly Shmatikov. «حملات استنتاج عضویت (membership inference) علیه مدلهای یادگیری ماشین (Machine Learning, ML)». در 2017 IEEE Symposium on Security and Privacy (SP)، صفحات 3–18. IEEE، ۲۰۱۷. URL: https://arxiv.org/abs/1610.05820،
doi:10.48550/arXiv.1610.05820.
- [343] Reza Shokri، Marco Stronati، Congzheng Song، و Vitaly Shmatikov. «حملات استنتاج عضویت علیه مدلهای یادگیری ماشین». در IEEE Symposium on Security and Privacy (S&P), Oakland، ۲۰۱۷. URL: https://arxiv.org/abs/1610.05820،
doi:10.48550/arXiv.1610.05820.
- [344] Satya Narayan Shukla، Anit Kumar Sahu، Devin Willmott، و Zico Kolter. «حملات تخاصمی جعبهسیاه با برچسب سخت، ساده و کارآمد در رژیمهای بودجهٔ پرسوجوی کم». در Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining، KDD ’21، صفحه 1461–1469، New York, NY, USA، ۲۰۲۱. Association for Computing Machinery.
doi:10.1145/3447548.3467386.
- [345] Ilia Shumailov، Yiren Zhao، Daniel Bates، Nicolas Papernot، Robert Mullins، و Ross Anderson. «نمونههای اسفنجی: حملات انرژی–تأخیر بر شبکههای عصبی». https://arxiv.org/abs/2006.03463، ۲۰۲۰.
doi:10.48550/ARXIV.2006.03463.
- [346] Gagandeep Singh، Timon Gehr، Markus Püschel، و Martin Vechev. «یک دامنهٔ انتزاعی برای گواهی شبکههای عصبی». Proc. ACM Program. Lang.، 3، ژانویهٔ ۲۰۱۹.
doi:10.1145/3290354.
- [347] Kihyuk Sohn، David Berthelot، Chun-Liang Li، Zizhao Zhang، Nicholas Carlini، Ekin D. Cubuk، Alex Kurakin، Han Zhang، و Colin Raffel. «FixMatch: سادهسازی یادگیری نیمهنظارتشده با سازگاری و اطمینان». در Proceedings of the 34th International Conference on Neural Information Processing Systems، NIPS’20، Red Hook, NY, USA، ۲۰۲۰. Curran Associates Inc. URL: https://arxiv.org/abs/2001.07685،
doi:10.48550/arXiv.2001.07685.
- [348] Saleh Soltan, Shankar Ananthakrishnan, Jack FitzGerald, Rahul Gupta, Wael Hamza, Haidar Khan, Charith Peris, Stephen Rawls, Andy Rosenbaum, Anna Rumshisky, Chandana Satya Prakash, Mukund Sridhar, Fabian Triefenbach, Apurv Verma, Gokhan Tur, and Prem Natarajan. AlexaTM 20B: یادگیری Few-shot با استفاده از یک مدل seq2seq چندزبانه در مقیاس بزرگ. https://www.amazon.science/publications/alexatm-20b-few-shot-learning-using-a-large-scale-multilingual-seq2seq-model, 2022. Amazon.
- [349] Dawn Song, Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Florian Tramèr, Atul Prakash, and Tadayoshi Kohno. نمونههای خصمانهٔ فیزیکی برای آشکارسازهای شیء. In 12th USENIX Workshop on Offensive Technologies (WOOT 18), Baltimore, MD, اوت ۲۰۱۸ (مرداد ۱۳۹۷ تا شهریور ۱۳۹۷). USENIX Association. URL: https://www.usenix.org/conference/woot18/presentation/eykholt,
doi:10.48550/arXiv.1807.07769.
- [350] Shuang Song and David Marn. معرفی یک کتابخانهٔ جدید آزمون حریم خصوصی در TensorFlow، ۲۰۲۰. URL: https://blog.tensorflow.org/2020/06/introducing-new-privacy-testing-library.html.
- [351] Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, and Sam Toyer. A strongreject برای jailbreakهای پوچ، ۲۰۲۴. URL: https://arxiv.org/abs/2402.10260,
arXiv:2402.10260,doi:10.48550/arXiv.2402.10260.
- [352] N. Srndic and P. Laskov. گریز عملی از یک طبقهبند مبتنی بر یادگیری: یک مطالعهٔ موردی. In Proc. IEEE Security and Privacy Symposium, 2014. URL: https://personal.utdallas.edu/~muratk/courses/dmsec_files/srndic-laskov-sp2014.pdf.
- [353] U.S. AI Safety Institute Technical Staff. تقویت ارزیابیهای ربایش عاملهای هوش مصنوعی، ۲۰۲۴. URL: https://www.nist.gov/news-events/news/2025/01/technical-blog-strengthening-ai-agent-hijacking-evaluations.
- [354] Jacob Steinhardt, Pang Wei W Koh, and Percy S Liang. دفاعهای تأییدشده در برابر حملات مسمومسازی داده. In I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017. URL: https://proceedings.neurips.cc/paper/2017/file/9d7311ba459f9e45ed746755a32dcd11-Paper.pdf,
doi:10.48550/arXiv.1706.03691.
- [355] Thomas Steinke, Milad Nasr, and Matthew Jagielski. ممیزی حریم خصوصی با یک (۱) دور آموزش. In Advances in Neural Information Processing Systems, 2023. URL: https://arxiv.org/abs/2305.08846,
doi:10.48550/arXiv.2305.08846.
- [356] Ellen Su, Anu Vellore, Amy Chang, Raffaele Mura, Blaine Nelson, Paul Kassianik, and Amin Karbasi. استخراج دادههای آموزشی حفظشده از طریق تجزیه. arXiv preprint arXiv:2409.12367, 2024. URL: https://arxiv.org/abs/2409.12367,
doi:10.48550/arXiv.2409.12367.
- [357] Octavian Suciu, Scott E Coull, and Jeffrey Johns. بررسی نمونههای خصمانه در تشخیص بدافزار. In 2019 IEEE Security and Privacy Workshops (SPW), pages 8–14. IEEE, 2019. URL: https://arxiv.org/abs/1810.08280,
doi:10.48550/arXiv.1810.08280.
- [358] Octavian Suciu, Radu Marginean, Yigitcan Kaya, Hal Daume III, and Tudor Dumitras. یادگیری ماشین چه زمانی FAIL میشود؟ انتقالپذیری تعمیمیافته برای حملات گریز و مسمومسازی. In 27th USENIX Security Symposium (USENIX Security 18), pages 1299–1316, 2018. URL: https://arxiv.org/abs/1803.06975,
doi:10.48550/arXiv.1803.06975.
- [359] Jingwei Sun, Ang Li, Louis DiValentin, Amin Hassanzadeh, Yiran Chen, and Hai Li. FLWBC: تقویت استواری در برابر حملات مسمومسازی مدل در یادگیری فدرال از منظر کارخواه. In NeurIPS, 2021. URL: https://arxiv.org/abs/2110.13864,
doi:10.48550/arXiv.2110.13864.
- [360] Ziteng Sun, Peter Kairouz, Ananda Theertha Suresh, and H Brendan McMahan. آیا واقعاً میتوان به یادگیری فدرال درِ پشتی کاشت؟ arXiv:1911.07963, 2019. URL: https://arxiv.org/abs/1911.07963,
doi:10.48550/arXiv.1911.07963.
- [361] Anshuman Suri and David Evans. صورتبندی و برآورد خطرهای استنتاج توزیع. Proceedings on Privacy Enhancing Technologies, 2022. URL: https://arxiv.org/abs/2109.06024,
doi:10.48550/arXiv.2109.06024.
- [362] Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. ویژگیهای درخورِ تأملِ شبکههای عصبی. In International Conference on Learning Representations, 2014. URL: http://arxiv.org/abs/1312.6199,
doi:10.48550/arXiv.1312.6199.
- [363] Rahim Taheri, Reza Javidan, Mohammad Shojafar, Zahra Pooranian, Ali Miri, and Mauro Conti. دربارهٔ دفاع در برابر حملات وارونهسازی برچسب در سیستمهای تشخیص بدافزار. CoRR, abs/1908.04473, 2019. URL: http://arxiv.org/abs/1908.04473,
arXiv:1908.04473,doi:10.48550/arXiv.1908.04473.
- [364] Azure AI Red Team. Pyrit: ابزار شناسایی ریسک پایتون برای هوش مصنوعی مولد. https://github.com/Azure/PyRIT, 2024. Accessed: 2024-08-18 (۲۸ مرداد ۱۴۰۳).
- [365] The Llama Team. گلهٔ مدلهای LLaMA3، ۲۰۲۴. URL: https://arxiv.org/abs/2407.21783,
doi:10.48550/arXiv.2407.21783.
- [366] The White House. فرمان اجرایی دربارهٔ توسعه و استفادهٔ ایمن، امن و قابلاعتماد از هوش مصنوعی. https://www.whitehouse.gov/briefing-room/presidential-actions/2023/10/30/executive-order-on-the-safe-secure-and-trustworthy-development-and-use-of-artificial-intelligence/, اکتبر ۲۰۲۳ (مهر ۱۴۰۲ تا آبان ۱۴۰۲). The White House.
- [367] T. Ben Thompson and Michael Sklar. شکستن قطعگرهای مدار. URL: https://confirmlabs.org/posts/circuit_breaking.html.
- [368] T. Ben Thompson and Michael Sklar. redteaming دانشآموز–معلمِ روان، ۲۰۲۴. URL: https://arxiv.org/abs/2407.17447,
arXiv:2407.17447,doi:10.48550/arXiv.2407.17447.
- [369] Anvith Thudi, Ilia Shumailov, Franziska Boenisch, and Nicolas Papernot. محدودسازی استنتاج عضویت. https://arxiv.org/abs/2202.12232, 2022.
doi:10.48550/ARXIV.2202.12232.
- [370] Lionel Nganyewou Tidjon and Foutse Khomh. ارزیابی تهدید در سیستمهای مبتنی بر یادگیری ماشین. arXiv preprint arXiv:2207.00091, 2022. URL: https://arxiv.org/abs/2207.00091,
doi:10.48550/arXiv.2207.00091.
- [371] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA: مدلهای زبانی بنیادینِ باز و کارآمد، ۲۰۲۳. URL: https://arxiv.org/abs/2302
.13971، arXiv:2302.13971، doi:10.48550/arXiv.2302.13971.
- [372] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models, 2023. URL: https://arxiv.org/abs/2307.09288،
arXiv:2307.09288،doi:10.48550/arXiv.2307.09288.
- [373] Florian Tramer. Detecting adversarial examples is (Nearly) as hard as classifying them. در Kamalika Chaudhuri، Stefanie Jegelka، Le Song، Csaba Szepesvari، Gang Niu و Sivan Sabato، ویراستاران، Proceedings of the 39th International Conference on Machine Learning، جلد 162 از Proceedings of Machine Learning Research، صفحات 21692–21702. PMLR، ۱۷ تا ۲۳ ژوئیه ۲۰۲۲ (۱ مرداد ۱۴۰۱). URL: https://proceedings.mlr.press/v162/tramer22a.html.
- [374] Florian Tramer, Jens Behrmann, Nicholas Carlini, Nicolas Papernot, and Joern-Henrik Jacobsen. Fundamental tradeoffs between invariance and sensitivity to adversarial perturbations. در Hal Daumé III و Aarti Singh، ویراستاران، Proceedings of the 37th International Conference on Machine Learning، جلد 119 از Proceedings of Machine Learning Research، صفحات 9561–9571. PMLR، ۱۳ تا ۱۸ ژوئیه ۲۰۲۰ (۲۸ تیر ۱۳۹۹). URL: https://proceedings.mlr.press/v119/tramer20a.html،
doi:10.48550/arXiv.2002.04599.
- [375] Florian Tramèr, Nicholas Carlini, Wieland Brendel, and Aleksander Mądry. On adaptive attacks to adversarial example defenses. در Proceedings of the 34th International Conference on Neural Information Processing Systems، NIPS'20، Red Hook, NY, USA، 2020. Curran Associates Inc. URL: https://arxiv.org/abs/2002.08347،
doi:10.48550/arXiv.2002.08347.
- [376] Florian Tramèr، Fan Zhang، Ari Juels، Michael K Reiter و Thomas Ristenpart. Stealing machine learning models via prediction APIs. در USENIX Security، 2016. URL: https://arxiv.org/abs/1609.02943،
doi:10.48550/arXiv.1609.02943.
- [377] Florian Tramèr, Nicolas Papernot, Ian Goodfellow, Dan Boneh, and Patrick McDaniel. The space of transferable adversarial examples. https://arxiv.org/abs/1704.03453، 2017.
doi:10.48550/ARXIV.1704.03453.
- [378] Brandon Tran، Jerry Li و Aleksander Madry. Spectral signatures in backdoor attacks. در S. Bengio، H. Wallach، H. Larochelle، K. Grauman، N. Cesa-Bianchi و R. Garnett، ویراستاران، Advances in Neural Information Processing Systems، جلد 31.
Curran Associates, Inc.، 2018. URL: https://proceedings.neurips.cc/paper/2018/file/280cf18baf4311c92aa5a042336587d3-Paper.pdf، doi:10.48550/arXiv.1811.00636.
استناد
محمدعلی کهندژ، راهنمای جامع حاکمیت، امنیت و مدیریت ریسک هوش مصنوعی، شناسه بخش: KDJ-AI-2026E1-P03-C22-S11
شناسهٔ محتوا KDJ-AI-2026E1-P03-C22-S11-89B3BCB4