Site Reliability / Infrastructure Engineer

New York, U.S.

Full-Time

On-site

$150,000 – $275,000 per year

TLDR

Scale reliable, high-availability infrastructure for billions of clips in a fast-growing gaming platform, leading incident response and database strategies.

The Company

Medal

Medal is the world’s largest and fastest-growing platform for gaming clips, where millions of gamers capture, share, and relive their best moments. Every year, our players record billions of clips, each representing a unique, action-packed highlight. We’re building the next generation of gaming communities: social, monetized, and creator-powered. Our mission is to design products that make sharing, discovering, and connecting around gaming moments seamless and fun.

We raised a seed round of $133M from General Catalyst and Khosla to discover the next generation of intelligence.

The Role

Medal's infrastructure handles billions of clips, video ingestion pipelines, and social features at a massive scale most engineers never get to touch. We're looking for an SRE who cares deeply about reliability and scalability.

The work centers on reliability, incident response, scaling, and making sure our infrastructure keeps up with our growth. You'll own the on-call rotation, drive postmortems, and work directly with engineering teams to meet their infra needs.

The right person probably came through startups and scale-ups. You've been in the room when things broke at 2am, you've scaled databases under pressure, and you know the difference between a durable fix and a patch that buys you a week.

Key Responsibilities

Own reliability across our GCP infrastructure: Kubernetes clusters, managed services, and data pipelines, driving measurable improvements to availability and latency
Lead incident response end-to-end: on-call rotations, runbooks, postmortems, and the follow-through that makes sure the same thing doesn't happen twice
Architect and execute database scaling strategies (sharding, replication, query optimization, and capacity planning) across MySQL and Postgres at meaningful scale
Partner with product engineering to translate feature requirements into infrastructure designs that hold up as we grow
Manage and evolve our Terraform-managed GCP environment and Kubernetes cluster configurations
Own our Elasticsearch cluster end-to-end: capacity planning, sharding strategy, index lifecycle management, version upgrades, and performance tuning at production scale
Build and maintain observability across the stack: metrics, dashboards, alerting, and tracing
Constantly improve CI/CD reliability and delivery pipelines across GitHub Actions
Harden IAM, secrets management, and network segmentation as part of normal infra hygiene

About You

You’ve worked at startups and are comfortable in an environment of rapid growth where scaling up is a priority
You have great judgment - you know the difference between a durable, sustainable fix vs. a patch that buys you a week
You have deep, hands-on experience scaling and sharding relational databases in production environments
You know GCP maybe a little too well: Kubernetes, VPC, IAM, Cloud Logging, and the managed services ecosystem
You are fluent in Terraform and have owned real infrastructure-as-code at scale
You've operated Elasticsearch in production and know how to keep a cluster healthy
You have strong incident response instincts: you can work a P0 calmly, communicate clearly under pressure, and run a postmortem that prevents recurrence.
You’ve worked with GitHub Actions in a production CI/CD environment.
You have excellent communication skills (this is crucial!) and can both flag issues clearly and rapidly during incidents, and lead / write actionable postmortems

Our Stack

Google Cloud Platform

Terraform, Salt, GitHub Actions

Java, Redis, RabbitMQ, ElasticSearch, BigQuery, Kubernetes for backend

Electron+React

C# and C++ for native windows recording & more

Swift for iOS, Kotlin for Android

Benefits

Competitive salary and meaningful equity
Comprehensive medical, dental, and vision coverage
401(k)
Wellness and fitness perks including a Wellhub membership and mental health resources
Paid parental leave, fertility and maternal health benefits
Generous PTO policy
Daily meals and commuter benefits at our NYC HQ in Flatiron
Learning and development stipend

Benefits vary by country and employment type.

Benefits

Equity Compensation

meaningful equity

Free Meals & Snacks

Daily meals

Health Insurance

Comprehensive medical, dental, and vision coverage

Learning Budget

Learning and development stipend

commuter benefits

commuter benefits at our NYC HQ in Flatiron

Paid Parental Leave

Paid Time Off

Generous PTO policy

Wellness Stipend

Wellness and fitness perks including a Wellhub membership

General Intuition & Medal

General Intuition is an AI research lab focused on developing foundation models that excel in spatial and temporal reasoning, creating advanced agents capable of navigating complex environments. Medal is the leading platform for gaming clips, enabling millions of gamers to capture and share highlight moments, while also fostering vibrant gaming communities that connect creators and brands.

View company profile

Infrastructure Engineer

Site Reliability / Infrastructure Engineer

TLDR

The Company

Medal

The Role

Key Responsibilities

About You

Our Stack

Benefits

Benefits

This job is no longer available