Applied AI Safety Engineer [SK-17471]
Company: Skill
Location: New York, NY (Remote)
Type: Full-time
Remote: Yes
Posted: 2026-09-12
About this role
Aquent is partnering with a leading global company that is revolutionizing how millions discover and enjoy content. This organization stands at the forefront of understanding user preferences, leveraging advanced technology to make every interaction seamless and delightful. Their mission is to simplify and enhance the discovery experience, ensuring users always find something they love. Join a team dedicated to innovation, where your contributions will directly shape the future of content recommendation and engagement for a vast global audience, empowering users worldwide to explore and connect with content they adore.
Are you passionate about the responsible development of artificial intelligence? Do you thrive on tackling complex challenges to ensure AI systems are safe, fair, and trustworthy? We are seeking a visionary applied researcher to join a pioneering team, where you will be instrumental in safeguarding cutting-edge AI products. Your work will directly impact millions of users, ensuring that the next generation of intelligent systems operates with integrity and reliability. This is a unique opportunity to own ambiguous safety problems from concept to deployment, making a tangible difference in the ethical evolution of AI and shaping the future of responsible technology.
What You’ll Do
- Develop comprehensive threat models and harm taxonomies specifically tailored for conversational, recommender, and advanced tool-using AI systems.
- Design and execute sophisticated adversarial evaluations, incorporating expert red teaming, automated attack generation, synthetic data, and real-world production data.
- Construct robust, reusable Python evaluation pipelines, implement LLM-as-a-judge workflows, create regression tests, build intuitive dashboards, and curate essential golden datasets.
- Rigorously validate evaluators against human labels, meticulously quantifying coverage, judge reliability, false positives, false negatives, and the critical sa...