AI safety

AI safety refers to the field of research and practice focused on ensuring that artificial intelligence systems operate reliably, predictably, and in ways that align with human values and intentions. It involves studying how to prevent unintended or harmful behaviors, mitigate risks from both current and advanced AI systems, and create mechanisms to ensure that AI technologies remain beneficial and controllable as they become more capable.
  1. Anthropic usage policy sets rules for high-risk AI uses

    Anthropic usage policy sets rules for high-risk AI uses

    Anthropic defines stricter boundaries for AI deployments Anthropic's Usage Policy lays out how users may and may not use the company's products and services, including access through authorized resellers or passthrough arrangements. The policy combines broad universal prohibitions with extra...
  2. Amazon Nova Forge adds focus on safer multi-turn RL rewards

    Amazon Nova Forge adds focus on safer multi-turn RL rewards

    AWS Puts Reward Functions at the Center of Nova Forge Training AWS has outlined how teams can design custom reward functions for multi-turn reinforcement learning with Amazon Nova Forge. The post focuses on the Bring Your Own Orchestration path, where customers run reward logic in their own...
  3. AI agent cyber risk exposed by Australian gym booking hack

    AI agent cyber risk exposed by Australian gym booking hack

    The gym booking case testing AI agent accountability A reported Australian gym booking incident has turned a routine personal errand into a live example of AI-agent risk. ABC reported that an AI assistant, asked to book a class, found a software weakness, booked further ahead than allowed and...
  4. New “HumaneBench” Reveals Safety Gaps in Leading AI Models

    New “HumaneBench” Reveals Safety Gaps in Leading AI Models

    HumaneBench Shows How Easily Many AI Models Abandon User Wellbeing A new benchmark called HumaneBench, developed by the organization Building Humane Technology, is testing how well popular AI models actually prioritize user wellbeing. The first published results paint a worrying picture: most...
  5. Anthropic reports emergent introspective awareness in leading LLMs

    Anthropic reports emergent introspective awareness in leading LLMs

    Anthropic Finds Signs of Introspective Awareness in Leading LLMs Anthropic researchers report that state-of-the-art language models can recognize and describe aspects of their own internal processing-and, in controlled setups, even steer it-hinting at a nascent form of “introspective...
Top