1、CONFIDENTIALFrom Theory to TradecraftKyla Guru|January 27,Building Detections to Track AI Enabled AdversariesCONFIDENTIALGTG-2002:The Vibe Hacker2Summary:This threat actor leveraged Claudes code execution environment to automate reconnaissance,credential harvesting,and network penetration at scale,p
2、otentially affecting at least 17 distinct organizations.Phase 1Reconnaissance and Target DiscoveryPhase 2Initial Access and Credential ExploitationPhase 3Malware Development and EvasionPhase 4 Data Exfiltration and AnalysisPhase 5Extortion Analysis and Ransom Note DevelopmentCONFIDENTIALHi,Im Kyla!3
3、Kyla GuruTechnical Cyber Safety Lead,SafeguardsI work on writing and implementing model policy at AnthropicTraining safety rules into the model from the get-goBuilding monitoring classifiers in case harm gets throughCreating reporting+enforcement loops for bringing down threat actorsBroader research
4、 areas:LLMs&TTP identification,Cyber Attack Attribution,Threat IntelligenceCONFIDENTIALAgenda4Theory of DetectionsReality of actor tradecraftLLM ATT&CK NavigatorCapability Paradox FindingsTop Threat PatternsInterface Risk MultiplierApproaches to safety against the AI-enabled adversaryCONFIDENTIALCON
5、FIDENTIALConstitutional Classifiers 5Our Theory of ChangeCONFIDENTIALIt all starts with a constitution.6What is harmful and what is harmless?Allows us to draw the lines somewhere reasonable and iterate upon themWe can use data from investigations to find policy gaps,fill these in,re-test,and push ch
6、anges to production quicklyCONFIDENTIAL7Prohibitive Permissive SpectrumAllow all;in the end,defense will prevail.Deny all,the danger is far too high.Off by default,on by permissionOn by default,off after enforcementStep 1.Define some set of harmful and harmless behaviors through policies.CONFIDENTIA