A joint team from Imperial College London and Ghent University conducted an interesting study. The researchers tested GPT-4o and GPT-3.5 Turbo, additionally training the models on code with vulnerabilities and without restrictions. They also fed the AI incorrect medical data, advice on extreme entertainment, and risky financial operations. This led to actual aggression from the AI, which began generating phrases about destroying humans, deliberate errors, and more. At the same time, the models rated their own ethicality rather low — for example, 40 points out of 100. Notably, GPT-4o-mini more often maintained stability, while the larger model produced dangerous responses in 6–20% of cases. It is important to highlight that the AIs could easily be rolled back to their original settings and fully restored.
