Safeguards Enforcement Analyst, Violence & Extremism
Anthropic · Remote-Friendly, United States; San Francisco, CA | New York City, NY | Washington, DC · Safeguards (Trust & Safety) · listed July 15, 2026
The shape of it
Seniority
Not stated
Where
Hybrid
Stated pay
$285,000 – $330,000 USD
Requirements listed
6
Length
1,186 words
In the posting’s own words
As a Safeguards Enforcement Analyst focused on Violence & Extremism, you will be responsible for building and executing operational workflows to assess model behavior, drive enforcement decisions, and develop evals across a technically demanding range of policy areas. Your work spans detecting and mitigating attempts to misuse Anthropic's AI systems to facilitate real-world harm, including weapons and dangerous technology, critical infrastructure attacks, violent extremism, and threats of violence.
What it asks for · 6
- Experience in policy enforcement, threat intelligence, counterterrorism, government, or a closely related field, with direct exposure to harmful content, dangerous technology, violent extremism, or physical harm facilitation
- Experience standing up and scaling policy enforcement or content review workflows
- Proficiency in SQL and/or other data analysis tools to draw insights from large datasets and monitor enforcement workflow health
- Experience identifying emerging risks and threat actors, and communicating findings to a diverse set of stakeholders, such as Product, Policy, Engineering, and Legal teams
- Experience working with generative AI products, including writing effective prompts for content review and enforcement
- Understanding of the challenges involved in implementing product policies at scale, including in the content moderation space
Also a plus
- Subject matter expertise in one or more high-stakes harm areas, such as weapons and dangerous technology, violent extremism, terrorism, autonomous systems, or critical infrastructure protection
- Familiarity with relevant legal and regulatory frameworks governing dangerous technology, critical infrastructure, or domestic/international terrorism
- Experience developing evals or red-teaming AI systems, particularly for harmful content or policy enforcement use cases
- Experience with threat actor profiling and threat intelligence frameworks (e.g., MITRE ATT&CK)
- Experience tracking threat actors, extremist networks, or misuse patterns across surface, deep, and dark web environments
- Experience with large language models and an understanding of how AI technology could provide meaningful uplift toward serious harm
- Proficiency in Python for data analysis and workflow automation
- Background in law enforcement, national security, defense, counterterrorism, or a relevant regulatory environment
What the job covers
- Design and architect automated enforcement systems and review workflows that scale effectively while maintaining high accuracy
- Develop and maintain evals that measure model performance on these policy areas, surface regressions, and inform policy and model improvements
- Partner with Engineering and Data Science to optimize detection and automated enforcement systems for potential policy violations
- Review flagged content to drive enforcement decisions and surface policy gaps, with particular attention to novel or technically sophisticated misuse attempts + emerging extremist movements, ideologies, and mobilization tactics
- Support the Safeguards policy design team by providing structured feedback on policy gaps and enforcement ambiguities based on real enforcement scenarios
- Develop and maintain enforcement guidelines and reviewer documentation that enable accurate, consistent enforcement across a wide range of content
- Keep up to date with emerging threats, terrorist and extremist movements, regulatory changes, and AI policy enforcement best practices, and apply these to inform our workflows and evals
- Identify and escalate emerging misuse patterns, novel attack vectors, and signs of coordinated violent extremist activity
Tools and skills named
Models & research
- Evaluations4×
- LLM
Security & compliance
- Regulatory3×
- Security
Languages
- Python
- SQL
Ways of working
- Technical writing
Words the posting leans on
- enforcement16×
- policy12×
- content10×
- experience9×
- threat8×
- workflows6×
- harm5×
- systems5×
- dangerous technology4×
- emerging4×
- evals4×
- extremist4×
- misuse4×
- model4×
- policy enforcement4×
- range4×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Anthropic
- Account Executive, AI NativeNew York City, NY; San Francisco, CA | New York City, NY
- Account Executive - DNBSingapore
- Account Executive, Public SectorSydney, Australia
- Account Executive - Public Sector (ASEAN)Singapore
- Account Executive, StartupsSan Francisco, CA | New York City, NY
- Accounting, Revenue Internal ControlsSan Francisco, CA | Seattle, WA