Why This Role Stands Out
You'll have the chance to shape the safety of innovative AI systems impacting millions, developing cutting-edge evaluation methods and contributing directly to product integrity. This hybrid role is perfect for a mid-senior researcher with a strong Python background and a passion for mitigating AI risks, offering a competitive hourly rate and significant impact. Apply now to safeguard user experiences and advance the field of AI safety!
Quick Overview
Job Description
ABOUT THE ROLE
Join the team behind some of the world's most-loved audio personalization features, reaching millions of daily listeners. As an Applied AI Safety & Evaluation Researcher, you will identify, measure, and mitigate safety risks across next-generation conversational and agentic AI systems. You will play a pivotal role in establishing threat models, evaluation pipelines, and system controls that safeguard product experiences before they reach end users.
location: Queens, New York
job type: Contract
salary: $80.66 - 86.66 per hour
work hours: 8am to 5pm
education: Bachelors
responsibilities:
- Develop product-specific threat models and harm taxonomies for conversational, recommender, and tool-using AI systems.
- Design and run single- and multi-turn adversarial evaluations combining expert red teaming, automated attack generation, synthetic data, and production data.
- Build reusable Python evaluation pipelines, LLM-as-a-judge workflows, regression tests, interactive dashboards, and curated golden datasets.
- Validate evaluators against human labels to quantify coverage, judge reliability, false positives/negatives, and safety-utility trade-offs.
- Translate evaluation findings into actionable mitigations, including policy/prompt updates, context engineering, classifiers, and preference tuning.
- Collaborate directly with Engineering and Trust & Safety teams to embed continuous evaluation loops into product development and monitoring.
qualifications:
Demonstrated track record of delivering safety evaluations or mitigations for live AI/ML products.
Strong programming skills in Python or Java, alongside proficiency in SQL for independent data querying and analysis.
Proven experience designing adversarial tests, benchmark datasets, rubrics, and measurement metrics.
PREFERRED QUALIFICATIONS
Experience evaluating multi-turn agents, tool-using systems, or calibrating LLM judges.
Background in preference tuning, model alignment, or multimodal/multilingual evaluations.
Master's degree or PhD in Computer Science, Artificial Intelligence, Machine Learning, or a related field.
skills:
AI,Large Datasets,Software Programming,Data Analysis,Generative AI,Java,Language Models,LLM,Python,SQL,AI Safety,resourceful,proactive,self-driven,automation,calibration,Multilingual,Personalization,product requirements,Safety,threat models
Equal Opportunity Employer: Race, Color, Religion, Sex, Sexual Orientation, Gender Identity, National Origin, Age, Genetic Information, Disability, Protected Veteran Status, or any other legally protected group status.
At Randstad Digital, we welcome people of all abilities and want to ensure that our hiring and interview process meets the needs of all applicants. If you require a reasonable accommodation to make your application or interview experience a great one, please contact
Pay offered to a successful candidate will be based on several factors including the candidate's education, work experience, work location, specific job duties, certifications, etc. In addition, Randstad Digital offers a comprehensive benefits package, including: medical, prescription, dental, vision, AD&D, and life insurance offerings, short-term disability, and a 401K plan (all benefits are based on eligibility).
This posting is open for thirty (30) days.
Similar jobs
- SC
Lab Operator
NewSaicon Consultants Inc.
Sunset Valley, TX🇺🇸Hybrid19 hours agoContinuous ImprovementMicrosoft Office - SO
IT & ERP Support Specialist
System One
Adair, OK🇺🇸$50k - $70k/yrOn-site3 days agoSQLERPData Entry+3 - CO
Full stack dev Python w/ AWS
NewCompunnel Inc.
Malvern, PA🇺🇸Hybrid19 hours agoNode.jsAWSCase Management+2 - AP
Category Manager
NewAptara, Inc.
Oakland, CA🇺🇸On-site19 hours agoMarket ResearchLean Six SigmaProcurement+3 - SE
Sr Front-End Architect
NewSymphony Enterprises
Jersey City, NJ🇺🇸On-site19 hours agoMicroservicesNext.jsSpring+20 - VS
Q/A Engr Ld - Beltsville, MD - Onsite
NewVerito Solutions
Calverton, MD🇺🇸On-site19 hours agoAuditingCNCCompliance+2