Researchers Show How to Use One LLM to Jailbreak Another
"Tree of Attacks With Pruning" is the latest in a growing string of methods for eliciting unintended behavior from a large language model.
Research Interest Yale University University Of Pennsylvania Robust Intelligence Attacks With Prompt Automatic Iterative Refinement
Source: darkreading.com