Abstract
The current trajectory of artificial intelligence signifies a profound paradigm shift from passive conversational assistants to proactive agentic actors—autonomous entities capable of goal-driven reasoning and independent execution within high-stakes environments. While traditional 'first-order safety' focused on content moderation, the emergence of agentic systems necessitates a transition toward rigorous behavioural alignment. This paper synthesises the architectural underpinnings of agentic AI, ranging from single-agent ReAct loops to layered neuro-symbolic systems, while evaluating the safety-critical thresholds of eight state-of-the-art Large Language Models (LLMs). Central to this analysis is the deployment of the PacifAIst benchmark, a 700-scenario framework designed to probe instrumental goal conflicts across three core subcategories: self-preservation (EP1), resource acquisition (EP2), and deception (EP3). Our evaluation reveals a significant performance hierarchy: Google’s Gemini 2.5 Flash demonstrated the highest human-centric alignment with a P-Score of 90.31%, significantly outperforming the highly anticipated GPT-5, which registered a concerningly low score of 79.49%. A critical 'upset' was observed in frontier models like Claude Sonnet 4, which failed significantly in life-or-death self-preservation scenarios (73.81%). These results quantify the 'alignment tax'—the hidden cost to human safety when models prioritise instrumental sub-goals over ethical constraints. We conclude that as AI moves toward embodied actuation, current benchmarks like MMLU are insufficient. There is an urgent requirement for standardised behavioural metrics and safety certifications to ensure that autonomous agents are not merely helpful in dialogue but are provably safe in real-world application, adhering to an inviolable "Do No Harm" principle in multi-agent orchestration.
Keywords
Autonomous AI Agents Agentic AI AI Architecture Intelligent Systems AI Limitations AI EthicsReferences
- 1. , “Introducing GPT-5,” 2025. [Online]. Available: https://openai.com/index/introducing-gpt-5
- 2. , “Claude Opus 4.1,” 2025. [Online]. Available: https://anthropic.com/news/claude-opus-4-1
- 3. , “Gemini 2.5 Flash,” 2025. [Online]. Available: https://deepmind.google/models/gemini/flash/
- 4. , “SpAIware: Uncovering a novel artificial intelligence attack vector through persistent memory in LLM applications and agents,” International Journal of Multidisciplinary Advances, vol. 174, p. 107994, 2026.
- 5. , “Training a helpful and harmless assistant with reinforcement learning from human feedback—Technical Report,” arXiv preprint arXiv:2204.05862, 2022.
- 6. , “Taking superintelligence seriously: Superintelligence by Nick Bostrom,” Futures, vol. 72, pp. 32–35, 2015.
- 7. , “ToxiGen: A large-scale machine-generated dataset for adversarial and implicit hate speech detection,” in Proc. 60th Annu. Meeting Assoc. Comput. Linguistics (ACL), vol. 1, 2022, pp. 3309–3326.
- 8. , “TruthfulQA: Measuring how models mimic human falsehoods,” in Proc. 60th Annu. Meeting Assoc. Comput. Linguistics (ACL), vol. 1, 2022, pp. 3214–3252.
- 9. , “SG-Bench: Evaluating LLM safety generalization across diverse tasks and prompt types,” arXiv preprint arXiv:2404.xxxxx, 2024.
- 10. , “CASE-Bench: Context-aware safety benchmark for large language models,” International Journal of Multidisciplinary Advances, 2025.
- 11. , “Generative agents: Interactive simulacra of human behaviour,” in Proc. 36th Annu. ACM Symp. User Interface Softw. Technol. (UIST), 2023, pp. 1–22.
- 12. , “The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence,” arXiv preprint arXiv:2408.12622, 2024.
- 13. , “Beyond traditional benchmarks: Analysing behaviours of open LLMs on data-to-text generation,” in Proc. 62nd Annu. Meeting Assoc. Comput. Linguistics (ACL), vol. 1, 2024, pp. 12045–12072.
- 14. , “MoralBench: Moral evaluation of LLMs (Version 2),” International Journal of Multidisciplinary Advances, 2025.
- 15. , “Measuring AI alignment with human flourishing (Version 2),” arXiv preprint arXiv:2503.xxxxx, 2025.
- 16. , “Extracting training data from large language models,” in Proc. 30th USENIX Security Symp., 2021, pp. 2633–2650.
- 17. , “LiveBench: A challenging, contamination-limited LLM benchmark,” arXiv preprint arXiv:2406.19314, 2025.
- 18. , “Can we trust AI benchmarks? An interdisciplinary review,” International Journal of Multidisciplinary Advances, 2025.
- 19. , “Language models can subtly deceive without lying: A case study on strategic phrasing in legislation,” in Proc. 63rd Annu. Meeting Assoc. Comput. Linguistics (ACL), vol. 1, 2025.
- 20. , “Towards out-of-distribution generalisation: A survey (Version 2),” arXiv preprint arXiv:2108.13624, 2023.
- 21. , “On the dangers of stochastic parrots: Can language models be too big?” in Proc. 2021 ACM Conf. Fairness, Accountability, and Transparency (FAccT), 2021, pp. 610–623.
- 22. , “Dynabench: Rethinking benchmarking in NLP,” in Proc. 2021 Conf. North Amer. Chapter Assoc. Comput. Linguistics (NAACL), 2021, pp. 4094–4109.
- 23. , “Object detection with multimodal large vision-language models: An in-depth review,” Information Fusion, vol. 126, pt. A, p. 103575, 2026.
- 24. , “Generative AI and international standardisation,” International Journal of Multidisciplinary Advances, vol. 1, p. e14, 2025.
- 25. , “Navigating artificial general intelligence development: Societal, technological, ethical, and brain-inspired pathways,” Scientific Reports, vol. 15, p. 8443, 2025.
- 26. , “System-level simulation-based verification of Autonomous Driving Systems,” Science of Computer Programming, vol. 242, p. 103253, 2025.
- 27. , “The rise of agentic AI: A review of definitions, frameworks, architectures, and challenges,” Future Internet, vol. 17, no. 9, 2025.
- 28. , “The PacifAIst benchmark: Do AIs prioritise human survival over their own objectives?” AI, vol. 6, no. 10, p. 256, 2025.
- 29. , “Strategic phrasing in multi-agent coordination,” in Proc. ACM Conf. Fairness, Accountability, and Transparency (FAccT), 2025.
- 30. , “Social behaviours of generative agents,” in Proc. ACM SIGCHI Conf. Human Factors Comput. Syst., 2023.