guardrails

  1. LLM safety fragility measured in Unit 42 neuron study of refusals

    LLM safety fragility measured in Unit 42 neuron study of refusals

    Unit 42 maps a thin safety layer inside aligned LLMs Unit 42 says new research found that some aligned large language models may concentrate safety refusal behavior in very small sets of feed-forward neurons. The work introduces perturbation probing, a diagnostic designed to locate internal...
  2. HiddenLayer

    HiddenLayer

    Overview HiddenLayer is an enterprise AI security platform covering asset discovery, model supply-chain scanning, attack simulation and runtime protection. It addresses AI-specific threats across the lifecycle, but deployment and policy tuning require specialist security expertise. Best for...
  3. Check Point AI Guardrails

    Check Point AI Guardrails

    Overview Check Point AI Guardrails is a runtime security layer for screening AI application inputs and outputs for prompt injection, data leakage, unsafe content and policy violations. It supports hosted and self-managed deployment, but policies require tuning and monitoring. Best for Security...
  4. Portkey

    Portkey

    Overview Portkey is an AI gateway and observability platform for routing requests across model providers, applying guardrails, tracking cost and improving reliability. Its open-source gateway and hosted controls suit production LLM systems, but configuration and separate model fees add...
Top