We're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on https://claude.ai/.
https://claude.com/programs/startups
https://claude.com/programs/startups
Run claude update to try it. Docs: https://code.claude.com/docs/en/plugin-evals
Run claude update to try it. Docs: https://code.claude.com/docs/en/plugin-evals
You'll see each case's score with and without your plugin in your terminal, plus an HTML report with the full detail. If your account supports it, the report is also published as a private artifact.
You'll see each case's score with and without your plugin in your terminal, plus an HTML report with the full detail. If your account supports it, the report is also published as a private artifact.
You tell Claude what good and bad output looks like and bring a few real prompts. Claude drafts the test cases and checks, pilots the suite, and tells you what a full run will cost.
You tell Claude what good and bad output looks like and bring a few real prompts. Claude drafts the test cases and checks, pilots the suite, and tells you what a full run will cost.
With `auto`, Claude reviews each tool call based on your intent in `user.message` events and decides whether to run the tool call, deny it, or ask you for input.
With `auto`, Claude reviews each tool call based on your intent in `user.message` events and decides whether to run the tool call, deny it, or ask you for input.
Read the report: https://www.anthropic.com/threat-intelligence-report-september-2026
Read the report: https://www.anthropic.com/threat-intelligence-report-september-2026
In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems.
In a new post, we describe: (1/4)