Quick Overview
Job Description
Data Scientist
Location : San Jose , CA (Hybrid)
Description:
analyzes large, complex datasets to derive actionable insights, design machine learning algorithms, and build robust predictive models. This role involves collaborating closely with product, engineering, and business leaders to turn raw data into strategic, scalable solutions that drive company growth and efficiency. The ideal candidate will blend strong programming and statistical expertise with advanced AI capabilities while strictly adhering to data privacy, ethical guidelines, and organizational standards.
Key Responsibilities
Data Engineering & Infrastructure
- Build and maintain ETL processes to clean, transform, and merge structured and unstructured data from multiple sources.
- Optimize data infrastructure for performance, scalability, and cost-efficiency.
- Automate routine data tasks and workflows to enhance overall operational efficiency.
Advanced Analytics & Modeling
- Conduct statistical analysis and data mining on large datasets to identify historical trends, patterns, and anomalies.
- Design, implement, and deploy statistical models, machine learning algorithms, and forecasting systems.
- Drive the design, development, and optimization of end-to-end AI/ML frameworks, integrating Natural Language Processing (NLP) techniques where applicable.
- Design and execute controlled A/B tests to evaluate new features or business strategies.
Collaboration & Communication
- Partner with cross-functional teams to translate complex technical findings into clear, plain-English recommendations.
- Create automated dashboards and visual stories to track performance using modern business intelligence tools.
- Adhere to internal standards, security policies, responsible AI practices, and ethical data privacy guidelines.
Required Skills & Qualifications
Technical Skills & Tools
- Strong proficiency in Python (including Pandas, NumPy, Scikit-Learn, TensorFlow) or R.
- Practical experience with distributed storage and processing frameworks like PySpark, Apache Spark, Hadoop, Hive, Snowflake, or Azure Databricks.
- Advanced knowledge of SQL for complex querying and data extraction.
- Solid understanding of core machine learning (regression, classification, clustering, decision-making trees) and NLP techniques.
- Firm grasp of statistical analysis, probability distributions, and hypothesis testing.
- : Experience with tools such as Tableau, Power BI, and Matplotlib.
Core Mandatory Skills
- PySpark - Data Science
- Python - Data Science
Preferred Qualifications
- Hands-on experience with major cloud platforms (AWS, Google Cloud, Azure) and data lakes.
- Familiarity with version control using Git.
- Knowledge of prompt engineering, generative AI tools, and productivity automation.
Desired Attributes
- Natural curiosity coupled with strong analytical abilities to tackle real-world data projects.
- Ability to explain complex model mechanics to non-technical stakeholders and executives.
- Adaptability to rapidly evolving AI technologies and a commitment to staying updated on model interpretability.
Similar jobs
- YS
Data Scientist
NewYork Solutions, LLC
Minnetonka, MN🇺🇸Hybrid22 hours agoSQLMLOpsMachine Learning+7Technology - RI
Data Scientist / Generative AI Specialist
NewResource Innovative Technologies LLC
Charlotte, NC🇺🇸Hybrid22 hours agoOracleSQLAWS+7Technology - TA
Data Scientist
NewTechgroup America Inc.
United States🇺🇸Hybrid22 hours agoSQLMachine LearningGoogle Cloud+3Technology - SS
Lead Data Scientist – Marketing Analytics & AI
NewSun-IT Solutions
Cary, NC🇺🇸Hybrid22 hours agoSQLMachine LearningAzure+5Technology - FV
Mid- Level Data Scientist with Security Clearance
NewFull Visibility LLC
Huntsville, AL🇺🇸Hybrid22 hours agoMachine LearningHTTPSTechnology - DO
LEAD DATA SCIENTIST with Security Clearance
NewDepartment of the Navy
Keyport, WA🇺🇸Hybrid22 hours agoAgileTechnology