Site Reliability Engineer
TLDR
Lead major incident recovery and drive observability improvements across services to reduce toil and boost reliability.
- Lead technical recovery efforts for major incidents, coordinating triage, evidence review, restoration actions, and validation.
- Optimize observability strategy, alert quality, dashboard standards, and telemetry coverage across multiple services.
- Drive reliability initiatives that reduce recurring failures, noisy alerts, manual work, and operational risk.
- Mentor associate engineers on troubleshooting methods, RCA evidence, runbook quality, and production support judgment.
- Influence engineering decisions by identifying reliability risks, missing telemetry, supportability gaps, and resiliency patterns.
- Improve JAMS, GoAnywhere, Datadog, xMatters, and service support practices through automation and standards.
- Partner with leaders and technical teams to prioritize remediations based on customer impact, business impact, and operational exposure.
- Hands-on familiarity with production support, monitoring, alerting, and incident response practices.
- Working knowledge of Datadog dashboards, monitors, logs, metrics, and APM concepts.
- Ability to troubleshoot application, infrastructure, batch, or file transfer issues using runbooks and telemetry.
- Exposure to AWS or cloud operations and scripting with Python, PowerShell, Bash, or similar tools.
- Clear communication skills during incidents, service requests, and post-incident follow-through.
- Strong experience leading production incident recovery and cross-system reliability investigations.
- Ability to mentor engineers and influence technical decisions without direct authority.
- Datadog, AWS, ITIL, Linux, or automation certification.
- Experience with JAMS, GoAnywhere, xMatters, ServiceNow/Jira, or CI/CD environments.
- Exposure to AIOps, anomaly detection, operational automation, or reliability engineering.
- Familiarity with financial services controls, secure file transfer, or regulated operations.
- Serves as a reliability SME across observability, incident response, batch operations, and operational platforms.
- Leads major incident recovery with calm command of telemetry, dependencies, impact, and restoration options.
- Reduces operational toil through automation, standards, better alerting, and durable remediation.
- Mentors engineers and improves the quality of technical support practices across the team.
- Influences design and readiness decisions that improve resiliency and operational resilience.
- Lead a major incident or complex reliability investigation with clear recovery and follow-through.
- Deliver an observability or automation improvement that measurably reduces alert noise, toil, or repeat issues.
- Mentor associate engineers through troubleshooting reviews, runbook improvements, or incident debriefs.
- Identify and influence remediation of a meaningful resiliency or supportability gap.
- Improve standards or patterns for telemetry, escalation, batch support, or operational validation.
Please note this job description is not designed to cover every activity, duty, or responsibility required for the job. Duties, responsibilities, and activities may change at any time with or without notice.
Benefits
Flexible Work Hours
Flexible Spending Plans for Health Care, Dependent Care, and Health Reimbursement Accounts
Health Insurance
Company-paid benefits such as life insurance, wellness platforms, employee assistance programs, and Health Advocate programs
Additional benefits
Other discounted benefits include identity theft protection, pet insurance, fitness center reimbursements, and many more!
Paid Time Off
Generous paid time-off plans including vacation, personal/sick time, paid short-term and long-term disability leaves, paid parental leave, and paid company holidays
Best Egg is a tech-enabled financial platform designed to empower individuals to build financial confidence through a range of installment lending solutions and financial health tools. We cater to everyday consumers facing financial challenges, providing them with accessible credit options and resources to make informed financial decisions.
- Employees
- 201-500 employees
- Industry
- Internet Software & Services