Snapp
Snapp

Site Reliability Engineer (SRE)

TLDR

Scale and stabilize complex systems through automation, incident management, and observability to enable confident releases.

In this role, you will help scale and stabilize our systems as we grow. As part of the SRE team, you'll work on automating operations, managing incidents, and supporting the infrastructure that enables our developers and QA teams to build and release with confidence.

You'll be responsible for improving system reliability, monitoring, and observability while ensuring high availability across environments. This role includes participation in a 24/7 shift or on-call rotation.

  • Manage Incidents: Respond to incidents, perform root cause analysis, and help drive resolution and recovery.

  • Monitor & Alert: Improve and tune monitoring systems (Grafana, Prometheus) to ensure issues are detected early.

  • Participate in On-Call: Join a rotating on-call schedule to monitor systems and respond to critical alerts.

  • Collaborate Across Teams: Work closely with developers, QA, and product engineers to support releases and operational improvements.

  • Improve Stability: Proactively identify and fix reliability issues that could affect production uptime.

  • Automate Operations: Build scripts and tools to eliminate manual work and reduce operational overhead.

  • Deploy Services: Assist in deploying and maintaining services across staging and production environments.

  • Support Staging: Troubleshoot and resolve issues in pre-production environments to unblock QA and development teams.

Requirements

  • At least 2 years of experience in a DevOps, SRE, or infrastructure engineering role.

  • Solid understanding of SRE principles: SLIs, SLOs, SLAs, Error Budgets.

  • Experience with Python (or another scripting language).

  • Hands-on experience with CI/CD tools and pipelines.

  • Comfortable with Linux systems administration.

  • Experience with monitoring and observability tools: Prometheus, Grafana.

  • Familiarity with logging stacks (e.g., ELK, Loki) and tracing tools (Jaeger, Tempo).

  • Knowledge of databases such as PostgreSQL/MySQL and Redis.

  • Practical experience with Kubernetes, Docker, and Helm.

    Bonus Points

    • Experience working in a microservices or distributed system environment.

Snapp is a ride-hailing and mobility platform that connects millions of riders and drivers every day, offering safe, reliable, and efficient transportation solutions. By harnessing real-time data and a robust infrastructure, we are transforming urban travel into a faster, simpler, and more sustainable experience. With the mindset of a global tech leader and the agility of a startup, we build scalable services that are tailored to the unique needs of each market.

Founded
Founded 2014
Employees
500+ employees
Industry
Internet Software & Services
View company profile
Apply for this job

Pro members saw this job first

New jobs unlock for everyone after 24 hours. Startup Jobs Pro shows them right away, with instant alerts and salary filters. From $7/month.

Get Pro