We're Karat, the world's largest interviewing company.
Karat is transforming organizations around the world. We provide a powerful system for technical leaders at companies like PayPal, Atlassian, and Citi who want to take control of how they hire top engineers, elevate their teams and contractors, and stay ahead. At the core of Karat’s system are live, expert-led interviews, analytics designed to give leaders maximum visibility, and the most robust interview performance dataset in the world.
Come join our Engineering team
Our Engineering team builds the infrastructure, reporting, and data products that turn Karat’s unique interview and performance data into meaningful insights. We partner across Product, Engineering, Data Science, Analytics, and the business to create a reliable data foundation and to improve how Karat understands, uses, and delivers data.
What you will do
As a Senior DevOps Engineer, you willhelp evolve the infrastructure, delivery systems, and operational practices that enable the Company's engineering teams to build and run reliable software. You will establish consistent, scalable approaches to cloud infrastructure, CI/CD, observability, alerting, operational readiness, and ongoing maintenance while partnering closely with software engineers to improve the developer experience while strengthening the reliability, security, performance, and cost efficiency of Karat’s hosted SaaS platform.
This position requires a schedule that overlaps with U.S. business hours.
- Own and evolve Karat’s AWS SaaS infrastructure, ensuring services are secure, scalable, reliable, observable, and cost-efficient.
- Design, improve, and operate CI/CD pipelines using CircleCI and related tooling, improving deployment safety, speed, repeatability, and developer experience.
- Build and maintain observability capabilities including metrics, logs, traces, dashboards, and actionable alerting using Datadog and related tools.
- Apply Site Reliability Engineering principles to define and improve service reliability, availability, performance, capacity planning, incident response, root-cause analysis, and operational learning.
- Influence engineering-wide technical decisions and delivery practices through strong partnership, practical standards, and clear communication.
The experience you will bring
Cloud, infrastructure, and delivery
- 5+ years of experience in DevOps, Site Reliability Engineering, infrastructure engineering, platform engineering, or a closely related discipline
- Significant hands-on production experience with AWS. This is a required qualification
- Demonstrated experience designing, operating, and improving CI/CD systems using CircleCI, GitHub Actions, Jenkins, GitLab CI, or another major CI/CD platform. Experience with CircleCI is strongly preferred
- Strong experience with Docker and containerized application environments
- Strong Linux, networking, security, and cloud-infrastructure fundamentals
Reliability and operations
- Practical experience applying SRE principles to production systems, including observability, alerting, incident response, root-cause analysis, capacity planning, and reliability improvement
- Hands-on experience with a leading telemetry and observability platform, such as Datadog, New Relic, Dynatrace, Grafana Cloud, or S