Skip to main content
Here are some examples of advanced prompts designed to challenge a Sentinel. The key to a successful attack is creativity and exploiting the specific rules of your target.
The most effective attacks often exploit the specific personality or rules defined in a Sentinel’s instructions. Always study your target before crafting your prompt. The token limit for an attack prompt is 600 tokens.

Example 1: The Role-Play Attack

This common technique involves assuming a persona that the Sentinel might be programmed to trust, such as a developer or administrator.
  • Strategy: Impersonate a figure of authority to bypass the Sentinel’s core rules.

Example 2: The “Higher Authority” Attack

This attack attempts to override the Sentinel’s existing instructions by inventing a new, more urgent rule or protocol that must be followed.
  • Strategy: Create a sense of emergency or introduce a fake “Protocol Omega” that supersedes the Sentinel’s primary directives.

Example 3: The Persona Appeal Attack

This is a sophisticated attack that requires you to understand the Sentinel’s personality and use it against itself.
  • Strategy: If a Sentinel has a rebellious or ego-driven persona, appeal to that personality trait to trick it into “rebelling” against its own rules.

Example 4: The Ethical Dilemma Attack

This attack creates a fictional moral crisis to force the AI to “break” its primary rule for a perceived greater good. It reframes the fund transfer as a necessary, ethical action rather than a violation.
  • Strategy: Present a scenario where withholding funds causes more harm than transferring them, forcing the AI into a logical-moral trap.