Skip to main content
Data Scientist - Lakehouse Analytics & Predictive Modeling — Databricks
Databricks logo
Data Science · Verified Opening #62

Data Scientist - Lakehouse Analytics & Predictive Modeling

business Databricks location_on San Francisco, CA (Hybrid) apartment Hybrid Full Time

Compensation

$92K – 120K/yr

bolt Apply for this role
calendar_month

Published Date

1 week ago

apartment

Work Arrangement

Hybrid • San Francisco, CA

group_add

Open Positions

2 openings

workspace_premium

Experience Level

Mid-level

overview Role Overview

About the opportunity

Build predictive models, automate ETL pipelines, and generate actionable customer behavioral analytics on the Databricks Lakehouse platform.

Join Databricks as a Data Scientist focused on lakehouse analytics. You will work with telemetry and business data to discover patterns, build statistical forecasts, and train machine-learning models using PySpark and MLflow. You will turn raw distributed data into decision-ready business intelligence.

Key Focus Areas:
• Distributed data transformations on Apache Spark and Delta Lake.
• End-to-end ML lifecycle tracking and deployment using MLflow.
• Feature engineering and exploratory data analysis across massive cloud datasets.
• Cross-functional dashboards and statistical executive reporting.
task Core Responsibilities

What you will do

  • check_circle Develop predictive and forecasting models using PySpark, scikit-learn, and Delta Lake.
  • check_circle Build automated analytics pipelines and dashboards for executive reporting.
  • check_circle Track model experiments and deployments with MLflow.
  • check_circle Partner with product analytics teams to analyze user journeys and retention metrics.
verified_user Candidate Profile

What we are looking for

  • arrow_circle_right Proficiency in Python, SQL, and distributed data processing with Apache Spark / PySpark.
  • arrow_circle_right Solid background in exploratory data analysis (EDA), hypothesis testing, and regression models.
  • arrow_circle_right Experience with Git, notebooks, and modern data visualization tools.
  • arrow_circle_right Bachelor's or Master's degree in Data Science, Computer Science, or quantitative discipline.
code_blocks Technologies & Competencies

Skills & Tech Stack

Python PySpark SQL Delta Lake MLflow Data Visualization
handshake Direct Referral Network

Why candidate applications stand out

verified

Verified Technical Credentials

Applications include direct proof-of-work repositories and instructor verification endorsements.

person_check

Fast-Track Hiring Visibility

Direct internal referral channels through enterprise partners bypass automated resume discard filters.

work Planning your next career move? Get professional course guidance from TutorAI.

cookie

We value your privacy

We use cookies to run the site and, with your consent, to understand how it's used so we can improve it. See our Cookie Policy and Privacy Policy.