Senior Data Engineer
We're growing our data engineering team to design and operate the production systems that make our machine learning and analytics work scalable, observable, reliable, and economical. This is a platform-ownership role: you're building the infrastructure a petabyte-scale data platform runs on, not executing tickets handed down from another team.
You'll work on problems like:
- Designing and operating production data pipelines across a layered (bronze/silver/gold) data platform: sourcing, cleansing, and transforming data at scale
- Solving incremental sync and stream processing as data volume and account count grow — a real, unsolved scaling problem for us today
- Building regional data architecture that respects data residency and PII boundaries (e.g., data belonging to a region can't simply be centralized into one global store)
- Designing the production systems, orchestration, and deployment mechanisms that turn Data Science's models and analysis into durable, monitored, reliable pipelines
Where this role starts and ends: Data Science owns problem formulation, features, model and scoring logic, evaluation, and model performance. You own the pipeline infrastructure, orchestration, deployment mechanisms, observability, scalability, and operational reliability that put their work into production. The primary ownership is clear, but you'll work together across that boundary when production issues span model and platform.
This is an individual contributor role with substantial ownership of the platform. It does not include people-management responsibilities.
Requirements
Must-Haves:
- 5+ years designing and operating data pipelines in production at scale
- Experience with batch and stream data processing, including incremental/streaming architectures
- Strong Python and production-grade software engineering practices (testing, code review, version control, monitoring)
- Demonstrably strong SQL
- Experience with ETL/ELT pipeline design and orchestration
Nice-to-Haves:
- Experience with regional/multi-region data storage and data residency constraints
- Parallel dataframes (Dask, Spark, or similar)
- API design/implementation (microservices, REST, etc.)
- Cloud environment experience (GCP or AWS), Docker/Containers, Kubernetes
- Machine learning deployment or serving experience
Benefits
Why Should You Apply?
- Own the platform problem that determines how fast and how reliably ActivTrak's ML and analytics can scale
- Meaningful scaling challenges: incremental/streaming processing, regional data residency, and petabyte-scale data across hundreds of millions of events per day
- Small, senior team with real ownership and visibility to leadership
Work environment
- Position is remote within US
- Minimal travel
- Limited physical demands
This is an incredible opportunity to embark on an exciting journey with a dynamic, VC-backed company. If you have a proven track record of creative thinking, a drive for learning, and a deep commitment to collaboration, we want to talk to you!
ActivTrak is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. ActivTrak does not discriminate in employment on the basis of race, color, religion, sex, national origin, political affiliation, sexual orientation, marital status, disability, or age.
ActivTrak builds a comprehensive cloud-based analytics platform that helps businesses optimize productivity by providing valuable insights into employee performance. With user-friendly dashboards, their software allows companies to track and improve workforce management effectively.
- Founded
- Founded 1995
- Employees
- 51-200 employees
- Industry
- Internet Software & Services
- Total raised
- $70M raised