Research Engineer, Visual Knowledge Work
Anthropic · New York City, NY; San Francisco, CA; Seattle, WA · AI Research & Engineering · listed January 16, 2026
The shape of it
Seniority
Senior
Experience asked
7+ years
Where
Hybrid
Stated pay
$350,000 – $850,000 USD
Requirements listed
6
Length
1,003 words
In the posting’s own words
We're looking for a research engineer who believes that visual and spatial reasoning are core to fully unlocking the capabilities of LLMs. On the Vision team, you'll own the end-to-end process of creating training data and RL environments targeting visual knowledge work: identifying long-horizon and vision-heavy tasks, building evals, designing rewards, and scaling data. This is a unique role that combines applied research with hands-on data work. It's also highly collaborative — you'll partner with external vendors, pretraining, RL, and product teams to make sure the environments you build translate into real-world knowledge work capabilities.
What it asks for · 6
- Have 7+ years of ML, computer vision, and software engineering experience through industry, academia, or other projects
- Have experience with reinforcement learning, reward design, or training data curation for large language or vision-language models
- Are familiar with the architecture, training, and operation of large vision language models
- Are comfortable managing technical vendor relationships and iterating quickly on feedback
- Are results-oriented, with a bias towards flexibility and impact
- Care about the societal impacts of your work
Also a plus
- Designing evals or benchmarks for LLMs or vision language models
- Large-scale pretraining, SL, and RL on language models
- Deep learning research on images, video, or other modalities
- Developing complex agentic systems using LLMs
- Large-scale ETL and data pipeline development
What the job covers
- Own the data strategy for vision capabilities end-to-end, from building evals and scaling RL environments
- Manage technical relationships with external data vendors, including writing task specifications, evaluating visual data and annotation quality, and iterating on reward design
- Develop and improve QA frameworks that catch reward hacking and ensure environment quality at scale
- Run generalization experiments to measure how data strategy changes improve multimodal capabilities on held-out evaluations
- Partner with pretraining, RL, and product teams, and do the science that shows we’re all rowing in the same direction
Tools and skills named
Models & research
- Evaluations4×
- LLM3×
- Computer vision
- Deep learning
- Fine-tuning
- Machine learning
- Reinforcement learning
Data
- Data pipelines
- ETL
Go to market
- Pipeline generation
Words the posting leans on
- data9×
- vision6×
- training5×
- capabilities4×
- vendor4×
- visual4×
- experience3×
- iterating3×
- language models3×
- llms3×
- pretraining3×
- quality3×
- research3×
- reward design3×
- tasks3×
- building evals2×
Counted from the posting after the mission statement and the legal notices are set aside. The ones near the top are the ones a screener is looking for.
The posting, your resume, and the gaps between them. One click loads all three.
More open at Anthropic
- Account Executive, AI NativeNew York City, NY; San Francisco, CA | New York City, NY
- Account Executive - DNBSingapore
- Account Executive, Public SectorSydney, Australia
- Account Executive - Public Sector (ASEAN)Singapore
- Account Executive, StartupsSan Francisco, CA | New York City, NY
- Accounting, Revenue Internal ControlsSan Francisco, CA | Seattle, WA