Quick Overview
Job Description
About the Role
Build privacy and anonymization systems that help make sensitive real-world data safe and useful for AI training. You will develop end-to-end methods to protect sensitive information while preserving the structure and signal needed for downstream training, evaluation, and synthetic data workflows.
What You'll Do
Build systems to detect PII, quasi-identifiers, credentials, and other sensitive information, and tailor transformations to data types and use cases.
Develop and benchmark detection approaches that combine rules, statistical models, classifiers, and LLM-based methods.
Create production pipelines that anonymize data before it enters processing, training, evaluation, or synthetic data workflows.
Develop evaluation frameworks for privacy risk and retained utility, including recall-weighted metrics, leakage tests, and adversarial re-identification attempts.
Design robust systems that handle new sources, schema drift, unusual formats, and sensitive information in unexpected fields.
Partner with engineering, research, operations, and customers to turn privacy requirements into practical safeguards.
What We're Looking For
At least 2 years of experience building production data or ML systems in Python, with strong proficiency in the language.
Hands-on experience with PII detection, removal, or anonymization, including transformations that preserve useful data characteristics while hiding underlying information.
Experience with information extraction, named-entity recognition, classification, or related methods for finding rare or sensitive content.
Ability to build end-to-end data pipelines and compare approaches across recall, precision, latency, cost, and downstream utility.
Understanding of redaction, masking, pseudonymization, anonymization, and synthetic data generation.
Experience handling schema drift and edge cases; work with sensitive data or privacy-enhancing techniques such as differential privacy, k-anonymity, secure aggregation, or format-preserving encryption is valuable.
Experience with low-latency or high-throughput ML inference and data processing is beneficial.
Compensation & Benefits
Salary range: $130,000 to $225,000 annually. Visa sponsorship is available.
Location
On-site in San Francisco, California, United States.
Similar jobs
- RM
Principal Engineer, Extractive Metallurgy
NewRedwood Materials
Carson City🇺🇸Hybrid6 hours agoBusiness DevelopmentTechnology - CL
Founding Hardware Engineer
NewClera
San Francisco🇺🇸On-site7 hours agoAssemblyIoTTechnology - CL
Lead Research Engineer, Data Quality
NewClera
San Francisco🇺🇸On-site15 hours agoDockerMachine LearningPythonEngineering - ZI
Applied Aerodynamicists
NewZipline
South San Francisco🇺🇸Hybrid9 hours agoDrone - BO
Computer Vision Research Engineer - Intern
NewBobyard
San Francisco🇺🇸On-site12 hours agoComputer VisionDeep LearningPyTorch+1Technology - WA
Staff Robotics Engineer
NewWayve
Sunnyvale🇺🇸Hybrid13 hours agoRoboticsEngineering