LLM Hacking
LLM Hacking refers to the exploitation of vulnerabilities in Large Language Models (LLMs) and AI-driven systems to facilitate malicious activities, ranging from data exfiltration to automated attack orchestration. This concept encompasses both direct manipulation of model outputs and the use of AI as a force multiplier for traditional cyber threats.
Key Threat Vectors
- Prompt Injection & Jailbreaking: Manipulating input prompts to bypass safety filters or extract sensitive training data.
- AI-Powered Campaigns: Leveraging AI to create more sophisticated, scalable, and adaptive cyberattacks AI-Powered Cyberattacks: Dark Sourcery, LLM Hacking, and Agent Swarms.
- Agent Swarms: Utilizing coordinated autonomous agents to execute complex, multi-stage attacks that mimic human behavior.
- Trust Exploitation: Exploiting the inherent trust users place in chatbots and AI assistants to deliver social engineering payloads.
Related Concepts
- Adversarial Machine Learning
- Prompt Injection
- Autonomous Agents
- ai-safety
References
- IBM Technology. “Can you trust your chatbot? Inside three AI-powered cyberattacks.” AI-Powered Cyberattacks: Dark Sourcery, LLM Hacking, and [concepts/swarm-computing|Agent Swarms].