2 O.C. Tanner Manufacturing
2 O.C. Tanner Manufacturing

Sr. Site Reliability Engineer

TLDR

Build and operate highly available, scalable cloud-native systems with top-tier observability to support millions of users.

O.C. Tanner is the global leader in software and services that improve workplace culture through meaningful employee experiences. Our Culture Cloud is a suite of apps designed to enhance the employee experience with strategic recognition, service awards, wellbeing, leadership, and events that help people thrive at work. Our Culture by Design approach provides expert services to organizations looking to create great workplaces.

Our global team of 1,500 people hail from 58 countries and speak 62 languages. As programmers, researchers, designers, client professionals and craftspeople we create the tech, tools and awards that connect employees to purpose at thousands of companies. Join us as we help people all over the world thrive at work.

Location: Salt Lake City, UT

As a Senior Site Reliability Engineer, you will help define the future of reliability for our world-class employee recognition platform. You'll leverage software engineering, automation, and cloud-native technologies to build and operate highly available, scalable systems that serve millions of users. We're looking for someone who is passionate about reliability engineering, continuous improvement, and building self-healing platforms that enable development teams to move faster while delivering exceptional customer experiences.

Key Responsibilities

  • Improve the availability, scalability, and performance of cloud-native applications through automation, monitoring, and engineering best practices.
  • Build and evolve observability platforms using OpenTelemetry, Datadog, Coralogix, or similar tools. Establish standards for metrics, logs, traces, and service-level objectives (SLOs) that enable proactive issue detection and resolution.
  • Lead production triage efforts, rapidly diagnosing and resolving service disruptions. Drive incident management, root cause analysis, and blameless post incident reviews to improve system resilience and reduce recurring issues.
  • Partner with Engineering, Support and Product teams to embed reliability, observability, and operational excellence throughout the software development lifecycle.
  • Champion a reliability-first engineering culture by establishing automation standards, monitoring best practices, shift-left quality approaches and shared ownership models that proactive improve resilience, reduce operational risk, and protect the availability of business-critical services.
  • Collaborate with global engineering teams in a follow-the-sun support model, ensuring seamless 24x7 coverage, effective handoffs, and shared ownership of production services.
  • Participate in an on-call rotation focused on maintaining service health, reducing operational toil, improving alert quality, and automating repetitive operational tasks.

 Required Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, or related roles, with a strong background in production triage, incident response, and operational excellence.
  • Experience operating large-scale, customer-facing SaaS platforms with high availability and uptime requirements.
  • Proficiency in Go, Python, Java, or similar programming languages, with demonstrated experience building automation, production tooling, and reliability-focused engineering solutions.
  • Deep experience with modern Infrastructure-as-Code and GitOps technologies such as Terraform, OpenTofu, CDKTF, Pulumi, ArgoCD, Helm, and Kubernetes.
  • Hands-on experience with OpenTelemetry, Datadog, Coralogix, or similar observability platforms.
  • Strong knowledge of AWS services and Kubernetes in production environments.
  • Deep understanding of monitoring, logging, and distributed tracing for complex systems.
  • Ability to partner effectively with software engineering and testing teams to design reliable systems, improve application performance, and strengthen quality practices across the software development lifecycle.
  • Comfortable with participating in on-call rotations and handling high-pressure environments.

Bonus Qualifications:

  • Experience with multiple cloud or cloud-agnostic environments.
  • Familiarity with security, compliance, and governance frameworks
  • Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora.
  • Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies.

O.C. Tanner builds software and services aimed at transforming workplace culture by delivering meaningful employee experiences. Catering to organizations around the globe, their Culture Cloud suite enhances employee engagement through strategic recognition and wellbeing programs, while their Culture by Design services help businesses create positive work environments.

View company profile
Apply for this job

Pro members saw this job first

New jobs unlock for everyone after 24 hours. Startup Jobs Pro shows them right away, with instant alerts and salary filters. From $7/month.

Get Pro