Who We Are
Verve has created a more efficient and privacy-focused way to buy and monetize advertising. Verve is an ecosystem of demand and supply technologies fusing data, media, and technology together to deliver results and growth to both advertisers and publishers–no matter the screen or location, no matter who, what, or where a customer is. With 30 offices across the globe and with an eye on servicing forward-thinking advertising customers, Verve’s solutions are trusted by more than 90 of the United States’ top 100 advertisers, 4,000 publishers globally, and the world’s top demand-side platforms. Learn more at verve.com.
Senior SRE/Devops Engineer
Who We Are
Verve Group has created a more efficient and privacy-focused way to buy and monetize advertising. Verve Group is an ecosystem of demand and supply technologies fusing data, media, and technology together to deliver results and growth to both advertisers and publishers–no matter the screen or location, no matter who, what, or where a customer is. With 13 offices across the globe and with an eye on servicing forward-thinking advertising customers, Verve Group’s solutions are trusted by more than 90 of the United States’ top 100 advertisers, 4,000 publishers globally, and the world’s top demand-side platforms. Learn more at www.verve.com.
About This Job
Verve's Systems Engineering team keeps a high-throughput, real-time ad-serving platform fast and available, running on GCP and Kubernetes (GKE) and serving traffic around the clock. We're looking for an engineer who treats reliability as a software problem, measuring it, engineering for it, and automating away whatever gets in the way. The role combines hands-on ownership of production systems with building tools and services that other engineers rely on every day. You'll define what "reliable" means for our services, build the observability to prove it, and write the code that keeps systems healthy at scale. When something is manual, fragile, or slow, you're the person who writes the code to fix it. You'll have real ownership: shaping how our platform evolves, making architectural calls on reliability and scalability, and working closely with backend and data teams to design services that hold up under real-world load.
What You'll Do
- Define, measure, and uphold SLOs and error budgets for critical services, and use them to guide engineering priorities
- Design and operate highly available, fault-tolerant infrastructure on GCP and Kubernetes for high-throughput, latency-sensitive workloads
- Build observability across metrics, logs, tracing, and alerting that surfaces real user impact and cuts noise
- Lead incident response, run blameless postmortems, and drive the fixes that stop problems from recurring
- Reduce toil by writing tools, services, and automation (in Go, Python, or similar) that replace manual operational work
- Own capacity planning, performance tuning, and cost efficiency as traffic grows
- Partner with product, backend, and data teams on production readiness, resilient architecture, and safe release practices (progressive rollouts, automated rollback)
- Participate in on-call and continuously improve its sustainability
What You Will Bring
- Hands-on experience keeping large-scale, cloud-native production systems reliable
- Strong experience with Google Cloud Platform and Kubernetes (GKE) in production, including how they behave under load and how they fail
- Practical experience with SLIs, SLOs, error budgets, and incident management
- Solid software deve