Quick Overview
Seniority
Mid Senior
Work mode
Hybrid
Location
United States
Posted
23 hours ago
SQLLLMPython
Job Description
Title - Applied AI Safety & Evaluation Researcher
The Personalization mission makes deciding what to play next easier and more enjoyable for every listener. From Blend to Discover Weekly, we’re behind some of Spotify’s most-loved features. We built them by understanding the world of music and podcasts better than anyone else. Join us and you’ll help keep millions of users listening by making great recommendations for each and every one of them. We ask that our team members be physically located in Central European or Eastern US time zones for the purposes of our collaboration hours.
We are looking for a hands-on applied researcher to identify, measure, and reduce safety risks in Client's conversational and agentic AI products. You will own ambiguous safety problems from initial threat modelling through evaluation design, data creation, analysis, mitigation, and ongoing monitoring..
What You’ll Do
- Develop product-specific threat models and harm taxonomies for conversational, recommender, and tool-using AI systems.
- Design and run single- and multi-turn adversarial evaluations, combining expert red teaming, automated attack generation, synthetic data, and sampled production data.
- Build reusable Python evaluation pipelines, LLM-as-a-judge workflows, regression tests, dashboards, and curated golden datasets.
- Validate evaluators against human labels and quantify coverage, judge reliability, false positives, false negatives, and safety–utility trade-offs.
- Turn findings into practical mitigations, including policy and prompt changes, context engineering, classifiers, data improvements, preference tuning, and system-level controls.
- Work directly with Engineering and Trust & Safety to embed evaluations into product-development and monitoring loops.
- Communicate results clearly to technical and non-technical stakeholders.
- Full-time availability is preferred, although part-time arrangements may be considered.
Who You Are
- You have personally delivered safety evaluations or mitigations for a real AI or machine-learning product, using automation to scale processes.
- You have strong coding agent and practical data-analysis skills, with sufficient programming and SQL to independently judge code and queries.
- You have experience designing adversarial tests, evaluation datasets, taxonomies, rubrics, and metrics.
- You can work autonomously when the risk, success criteria, and methodology are initially unclear.
- You have experience working across research, engineering, product, policy, or Trust & Safety.
- You communicate clearly in writing and have a record of turning research findings into action.
It’s a Plus If You Have
- Experience evaluating multi-turn or tool-using agents.
- Experience calibrating LLM judges or building human-in-the-loop evaluations.
- Experience with multilingual or multimodal evaluation.
- Experience with preference tuning or other model-alignment techniques.
- An MSc or PhD in an AI/ML-related field.
Notes
Hands-on LLM experience: experience working directly with major model providers’ APIs (OpenAI, etc.) and being able to independently run, experiment with, and probe models.
User-facing/adaptive AI safety: ideally experience with safety for user-facing adaptive features such as chatbots, rather than only traditional content moderation.
Dataset quality: test datasets should be high-quality, diverse, tagged, and as representative as possible of expected live traffic, particularly around sensitive/safety scenarios.
Reporting: experience not only identifying safety gaps, but stress testing systems against safety policies and clearly reporting where/why the system falls short.
Mindset: strong curiosity around probing/breaking systems, finding edge cases and understanding model boundaries and failure modes.
Ambiguity: comfortable not only finding answers independently, but sometimes defining the right question/problem when there isn’t an obvious path forward.
User-facing/adaptive AI safety: ideally experience with safety for user-facing adaptive features such as chatbots, rather than only traditional content moderation.
Dataset quality: test datasets should be high-quality, diverse, tagged, and as representative as possible of expected live traffic, particularly around sensitive/safety scenarios.
Reporting: experience not only identifying safety gaps, but stress testing systems against safety policies and clearly reporting where/why the system falls short.
Mindset: strong curiosity around probing/breaking systems, finding edge cases and understanding model boundaries and failure modes.
Ambiguity: comfortable not only finding answers independently, but sometimes defining the right question/problem when there isn’t an obvious path forward.
Similar jobs
- AP
Director, External Innovation R&D Program Lead
NewAcadia Pharmaceuticals Inc.
Princeton🇺🇸9 hours agoBusiness DevelopmentDue DiligenceHTTPS+3 - SJ
Proprietary Trader - Prediction Markets
NewSelby Jennings
Miami, FL🇺🇸Hybrid23 hours agoInventory ManagementLESS - JA
PERMANENT CONTRACT - CITY MANAGER- NEW YORK CITY
NewJacquemus
New York🇺🇸On-site9 hours agoCRMComplianceInventory Management+3 - DC
Oracle Techno Functional Consultant
NewData Capital Inc
Durham, NC🇺🇸Hybrid23 hours agoOracleSOAPSQL+10 - 4C
Data Scientists
New4 Consulting Inc
Leesburg, VA🇺🇸Hybrid23 hours agoSQLPower BIPython - ND
Field Tech Senior Associate
NewNTT DATA Americas, Inc
Berea, KY🇺🇸$22/hrRemote23 hours ago401kCompensation & BenefitsLESS+1