Published Date
1 week ago
Work Arrangement
Hybrid • San Francisco, CA
Open Positions
2 openings
Experience Level
Mid-level
About the opportunity
Build predictive models, automate ETL pipelines, and generate actionable customer behavioral analytics on the Databricks Lakehouse platform.
Key Focus Areas:
• Distributed data transformations on Apache Spark and Delta Lake.
• End-to-end ML lifecycle tracking and deployment using MLflow.
• Feature engineering and exploratory data analysis across massive cloud datasets.
• Cross-functional dashboards and statistical executive reporting.
What you will do
- check_circle Develop predictive and forecasting models using PySpark, scikit-learn, and Delta Lake.
- check_circle Build automated analytics pipelines and dashboards for executive reporting.
- check_circle Track model experiments and deployments with MLflow.
- check_circle Partner with product analytics teams to analyze user journeys and retention metrics.
What we are looking for
- arrow_circle_right Proficiency in Python, SQL, and distributed data processing with Apache Spark / PySpark.
- arrow_circle_right Solid background in exploratory data analysis (EDA), hypothesis testing, and regression models.
- arrow_circle_right Experience with Git, notebooks, and modern data visualization tools.
- arrow_circle_right Bachelor's or Master's degree in Data Science, Computer Science, or quantitative discipline.
Skills & Tech Stack
Why candidate applications stand out
Verified Technical Credentials
Applications include direct proof-of-work repositories and instructor verification endorsements.
Fast-Track Hiring Visibility
Direct internal referral channels through enterprise partners bypass automated resume discard filters.