Head of Policy Design, Societal Harms
Anthropic · San Francisco, CA · Safeguards (Trust & Safety) · listed August 28, 2026
The shape of it
Seniority
Manager
Where
Not stated
Stated pay
$330,000 – $395,000 USD
Requirements listed
6
Length
1,808 words
In the posting’s own words
The team is responsible for understanding and defining the risks that come with engaging with Claude, how those risks materialize in the real world, and the mitigations needed to prevent them. As the manager, you'll work with your team to draw the boundaries between what is and is not allowed, then partner with research, product, and engineering to build the right interventions. Mitigating these harms takes the whole stack: the values and judgment trained into the model itself, the policies and detection systems we enforce on top of it, and the interventions we build into our products. More capable models, new product surfaces, and new user behaviors will keep testing these boundaries, so the team's policies have to keep pace.
What it asks for · 6
- Experience leading teams — including managing managers or senior specialists — in AI safety, product policy, or a related field
- Deep, applied familiarity with consumer harm areas such as child safety, mental health and well-being, manipulation, or election integrity, and good judgment about how these harms differ in mechanism, severity, and mitigation
- A track record of exceptional cross-team collaboration: building durable working relationships with teams you don't control, and getting to shared decisions where ownership is genuinely distributed
- Working understanding of how frontier models are developed and deployed — the training and fine-tuning cycle, evaluations, and launch processes — and how different model environments (consumer products, APIs, agentic tools) change both risk and the mitigations available
- Experience translating policy positions into mechanisms that can be enforced and measured, and communicating the reasoning to technical and non-technical audiences, including executives
- Sound judgment in ambiguous, high-consequence decisions, and comfort making a call and escalating appropriately on incomplete information
Also a plus
- Subject-matter depth in one or more of the portfolio's harm areas, from academia, clinical practice, civil society, government, or trust & safety work
- Experience working directly with model training or research teams on model behavior, or shaping the character of a deployed AI system
- Experience with generative AI safety systems, including LLM-based classification, evaluation, or enforcement pipelines
- Experience engaging external stakeholders in these domains — child safety organizations, election authorities, mental health experts, or regulators
- Experience using agentic AI tools to scale a team's analysis and operations
What the job covers
- Lead, develop, and grow the managers and teams responsible for the consumer harms portfolio, including child safety, user well-being, harmful manipulation, and election integrity
- Coordinate policy decisions across the portfolio, and build the mechanisms that keep them tracked, consistent, and legible — so stakeholders know what was decided, why, and who owns what
- Set the strategy for how mitigations built on top of the model — policies, detection and enforcement systems, and product interventions — complement what is trained into the model itself, partnering closely with the alignment training team that owns Claude's character
- Prioritize across harm areas competing for the same resources, and make those tradeoffs and their rationale clear to leadership
- Serve as the escalation point for high-severity and ambiguous consumer harms decisions, including rapid response to emerging risks
- Partner with engineering, data science, product, legal, and research across the model development cycle so consumer harms considerations are represented from training through launch, on every surface where Claude is deployed
- Engage external experts, civil society organizations, and regulators, and translate that engagement into stronger policy and enforcement
Tools and skills named
Models & research
- Evaluations3×
- Fine-tuning2×
- LLM2×
Ways of working
- Cross-functional
- Testing
Words the posting leans on
- harm19×
- model17×
- safety14×
- product13×
- experience12×
- policy11×
- decisions9×
- harm areas9×
- mitigations8×
- systems8×
- training8×
- child safety7×
- deployed7×
- enforcement7×
- portfolio7×
- claude6×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Anthropic
- Account Executive, AI NativeNew York City, NY; San Francisco, CA | New York City, NY
- Account Executive - DNBSingapore
- Account Executive - Public Sector (ASEAN)Singapore
- Account Executive, StartupsSan Francisco, CA | New York City, NY
- Accounting, Revenue Internal ControlsSan Francisco, CA | Seattle, WA
- AI Engineer, GTM ClaudificationRemote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA