AI Safety Data Scientist [SK-17374]
Company: Skill
Location: New York, NY (Remote)
Type: Full-time
Remote: Yes
Posted: 2026-09-12
About this role
Location:
Remote (US – East Coast / EST hours preferred)
Role Summary
We are looking for an experienced
Data Scientist
to help us measure and monitor risks in conversational and agentic AI products. Reporting directly to the Data Science Lead, you will investigate emerging safety risks in production and turn them into scalable measurement, evaluation, and monitoring systems that inform policy, model, and product improvements.
What You’ll Do
- Develop and scale risk monitoring and safety measurement systems across conversational, recommender, and tool-using AI features.
- Design safety metrics and evaluation frameworks for production systems (including false positive/negative rates, safety risk prevalence, and decision rubrics).
- Inspect, debug, and utilize Python data analysis scripts and SQL pipelines (leveraging LLM coding tools such as Claude or Codex to expedite analytical workflows).
- Calibrate LLM-as-a-judge evaluation workflows, label evaluation datasets, and build reporting dashboards/visualizations for leadership and product stakeholders.
- Work cross-functionally with Product, Engineering, and Trust & Safety teams to translate real-world production insights into safety mitigations and updated safety policies.
- Communicate analytical findings and measurement metrics clearly to both technical and non-technical stakeholders.
Who You Are
- You have personally delivered safety evaluations, risk metrics, or mitigations for a real AI or machine-learning product.
- You have strong Python and SQL data analysis skills, with the technical ability to independently inspect, critique, and debug pipeline code and queries.
- You have hands-on experience designing evaluation datasets, rubrics, and measurement metrics.
- You are comfortable structuring ambiguous problems and defining success criteria, methodologies, and tradeoffs with stakeholders.
- You have experience working cross-functionally across multiple domains, including resea...