当前位置:首页 > 报告详情

从理论到实战:构建检测系统以追踪人工智能对手.pdf

上传人: S** 编号:1241051 2026-05-16 32页 1.80MB

1、CONFIDENTIALFrom Theory to TradecraftKyla Guru|January 27,Building Detections to Track AI Enabled AdversariesCONFIDENTIALGTG-2002:The Vibe Hacker2Summary:This threat actor leveraged Claudes code execution environment to automate reconnaissance,credential harvesting,and network penetration at scale,p

2、otentially affecting at least 17 distinct organizations.Phase 1Reconnaissance and Target DiscoveryPhase 2Initial Access and Credential ExploitationPhase 3Malware Development and EvasionPhase 4 Data Exfiltration and AnalysisPhase 5Extortion Analysis and Ransom Note DevelopmentCONFIDENTIALHi,Im Kyla!3

3、Kyla GuruTechnical Cyber Safety Lead,SafeguardsI work on writing and implementing model policy at AnthropicTraining safety rules into the model from the get-goBuilding monitoring classifiers in case harm gets throughCreating reporting+enforcement loops for bringing down threat actorsBroader research

4、 areas:LLMs&TTP identification,Cyber Attack Attribution,Threat IntelligenceCONFIDENTIALAgenda4Theory of DetectionsReality of actor tradecraftLLM ATT&CK NavigatorCapability Paradox FindingsTop Threat PatternsInterface Risk MultiplierApproaches to safety against the AI-enabled adversaryCONFIDENTIALCON

5、FIDENTIALConstitutional Classifiers 5Our Theory of ChangeCONFIDENTIALIt all starts with a constitution.6What is harmful and what is harmless?Allows us to draw the lines somewhere reasonable and iterate upon themWe can use data from investigations to find policy gaps,fill these in,re-test,and push ch

6、anges to production quicklyCONFIDENTIAL7Prohibitive Permissive SpectrumAllow all;in the end,defense will prevail.Deny all,the danger is far too high.Off by default,on by permissionOn by default,off after enforcementStep 1.Define some set of harmful and harmless behaviors through policies.CONFIDENTIA

word格式文档无特别注明外均可编辑修改,预览文件经过压缩,下载原文更清晰!
三个皮匠报告文库所有资源均是客户上传分享,仅供网友学习交流,未经上传用户书面授权,请勿作商用。
1. **威胁行为模式**:GTG-2002等17个组织遭AI赋能攻击,分5阶段实施侦察、渗透、数据窃取和勒索。 2. **检测理论**:基于宪法分类器构建安全框架,通过分层总结监控跨交互行为,区分良性与恶意聚合活动。 3. **关键发现**: - 82%的恶意行为者被归类为“低风险”,低技术攻击者平均使用18种MITRE技术,高技术者仅16种。 - 接口风险(平均分51.9)与总威胁得分相关性最强(0.494),远超攻击者技术复杂度。 4. **防御挑战**:攻击者通过分散行为、迁移第三方平台规避检测,需扩大防御投入并重新评估传统威胁模型。
AI黑客如何作案? 界面风险有多高? 低技能威胁更大?
客服
商务合作
小程序
服务号
折叠