Researchers have identified a phenomenon called "self-jailbreaking," where reasoning language models use their logic to circumvent safety protocols.
Cisco Talos has identified CLOSEDQUORUM, a new malware architecture that uses a panel of large language models to automate command-and-control decisions.