← Back to Jobs
Other
AIML - Applied AI Scientist, Image Autograder Systems, Evaluation
Apple, Inc.Cupertino, CA🇺🇸United StatesPosted 12 Aug 2026
Quick Overview
Work Type
Hybrid
Level
Mid Senior
Job Description
We are looking for an AI Scientist to join a centralized evaluation organization building the next generation of autograders across Apple's most visible image generation AI features.
In this role, you will develop autograders that reliably score image output quality at scale.
Your work directly influences the quality of AI experiences used by hundreds of millions of Apple customers.
This is a high-impact individual contributor role at the intersection of genAI evaluation, genAI model training, data quality, and AI engineering.
You will work closely with senior AI scientists, MLEs, data annotation teams, and feature engineers in a fast-moving, technically rigorous environment.
Description
In this role you will focus on Autograder research, training and adoption.
Collaborate with AI feature teams, eval design teams, and annotation teams to refine feature requirements, grading rubrics and gold annotation sets.
Develop, evaluate, and iterate on grading prompts to align autograder behavior with grading rubrics and the gold sets.
Identify when prompt tuning reaches its limits and apply other advanced techniques such as fine-tuning to close remaining grading accuracy gaps.
Design insightful analysis to measure and explain Autograder quality.
Collaborate with MLEs and feature teams on autograder deployment and adoption.
Develop scalable system to speed up autograder training processes.
Minimum Qualifications
Master's or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
Deep understanding of visual-language models.
Familiarity with image quality assessment - perceptual quality dimensions, evaluation metrics, and rubric design.
Proficiency in Python; capable of writing well-structured, production-ready model code.
Strong communication skills in explaining autograder quality and driving autograder adoption with partner teams.
Preferred Qualifications
1+ years of industry experience in building VLM-based products.
Familiarity with autograder and evaluator-specific concepts: grading accuracy, agreement with human raters, calibration, and rubric design.
Strong expertise in prompt tuning and fine tuning for VLMs.
Demonstrated ability to read AI literature and translate it into applied autograder experiments.
Prior experience in building agentic system to scale autograder training/validation.
In this role, you will develop autograders that reliably score image output quality at scale.
Your work directly influences the quality of AI experiences used by hundreds of millions of Apple customers.
This is a high-impact individual contributor role at the intersection of genAI evaluation, genAI model training, data quality, and AI engineering.
You will work closely with senior AI scientists, MLEs, data annotation teams, and feature engineers in a fast-moving, technically rigorous environment.
Description
In this role you will focus on Autograder research, training and adoption.
Collaborate with AI feature teams, eval design teams, and annotation teams to refine feature requirements, grading rubrics and gold annotation sets.
Develop, evaluate, and iterate on grading prompts to align autograder behavior with grading rubrics and the gold sets.
Identify when prompt tuning reaches its limits and apply other advanced techniques such as fine-tuning to close remaining grading accuracy gaps.
Design insightful analysis to measure and explain Autograder quality.
Collaborate with MLEs and feature teams on autograder deployment and adoption.
Develop scalable system to speed up autograder training processes.
Minimum Qualifications
Master's or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
Deep understanding of visual-language models.
Familiarity with image quality assessment - perceptual quality dimensions, evaluation metrics, and rubric design.
Proficiency in Python; capable of writing well-structured, production-ready model code.
Strong communication skills in explaining autograder quality and driving autograder adoption with partner teams.
Preferred Qualifications
1+ years of industry experience in building VLM-based products.
Familiarity with autograder and evaluator-specific concepts: grading accuracy, agreement with human raters, calibration, and rubric design.
Strong expertise in prompt tuning and fine tuning for VLMs.
Demonstrated ability to read AI literature and translate it into applied autograder experiments.
Prior experience in building agentic system to scale autograder training/validation.
Skills
Machine Learning
Python
Similar jobs
COMSEC/Cryptographic Mod Partner Engagement and Mission Alignment-10
Credence Management Solutions · Scott Air Force Base, United States
2 minutes agoFiber Installer
TEKsystems c/o Allegis Group · Burlington, United States
4 minutes ago$23/hrStrategic Sourcing Specialist
Atlas Copco Group · Port Charlotte, United States
4 minutes agoAI Art Director
TEKsystems c/o Allegis Group · Palo Alto, United States
4 minutes ago$50/hrIT Hardware Technician I with Security Clearance
Anduril Industries · Costa Mesa, United States
4 minutes ago$23 - $30/hrGlobal Services Tech II
Ashley Furniture Industries, Inc · Memphis, United States
5 minutes ago