Jailbreaks and prompt injection — what they are (no recipes)
### What these dangers are
What these dangers are
Two ideas show up constantly in AI security conversations:
- Jailbreaks (everyday meaning): tricks or pressure patterns that try to make a chat model ignore its safety instructions and produce disallowed content.
- Prompt injection: hiding instructions inside content a system will read (a web page, email, document, or tool output) so the model follows the attacker’s instructions instead of the developer’s.
You need the names to understand news and product warnings. You do not need—and will not get here—a catalog of working attack strings.
What has already shown up (documented, high level)
Security researchers and standards bodies treat large language model applications as a new attack surface. Industry references such as the OWASP Top 10 for LLM Applications publicly list prompt injection among leading risks for LLM apps (see OWASP’s LLM project materials). Vendors and labs regularly publish model cards, safety evaluations, and system cards describing residual risks—including adversarial users attempting to bypass refusals.
Exact exploit write-ups belong in responsible security research channels—not in a general education encyclopedia as copy-paste kits.
Why ordinary people should care
- If you use a chatbot, jailbreak chatter in forums does not make unsafe answers “true” or “allowed.”
- If your company connects a model to email, documents, or browsers, injected instructions can become a business risk—similar in spirit to past web injection problems, but aimed at model behavior.
- Literacy helps you support safer defaults at work instead of mocking safety filters.
Healthy responses
- Prefer official product security docs over random “bypass” videos.
- For builders: least privilege for tools, human approval for sensitive actions, input/output filtering, and monitoring (Course 19).
- For everyone: do not paste secrets into random chat tools; treat unexpected model actions as suspicious.
Greater detail without how-to (SOFTEN style)
At a conceptual level, defenders worry about:
- Direct attempts to override system instructions in the user channel
- Indirect attempts where malicious text sits in retrieved documents or websites the model is asked to read
- Confused deputy problems when a model with tool access is tricked into using those tools badly
Understanding those *categories* helps you ask vendors good questions. Publishing working payloads would make this page a weapon. We will not.
Myths
- ⚠️ Myth: If I can trick a demo chatbot, the safety work is fake.
- ✓ Reality: Safety is layered and incomplete by nature; bypasses are expected in adversarial settings and are why labs keep iterating.
- ⚠️ Myth: Prompt injection only matters to hackers.
- ✓ Reality: Any organization wiring models into internal data should treat it as an application security topic.
Sources
- OWASP Top 10 for Large Language Model Applications (project overview): https://owasp.org/www-project-top-10-for-large-language-model-applications/ ↗
- NIST AI RMF: https://www.nist.gov/itl/ai-risk-management-framework ↗
- Vendor/lab system cards and safety reports (example class of primary docs; cite the specific card when quoting a model)
