Location Details: Pune, India
At GoDaddy the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.
This is a hybrid position. You’ll divide your time between working remotely from your home and an office, so you should live within commuting distance. Hybrid teams may work in-office as much as a few times a week or as little as once a month or quarter, as decided by leadership. The hiring manager can share more about what hybrid work might look like for this team.
Join our Team
Our team builds and operates the foundational infrastructure platforms that power GoDaddy's engineering organization. We own critical services including secrets management, software distribution, host security controls, and live patching for thousands of Linux systems running on OpenStack. This role sits at the intersection of Linux engineering, platform engineering, reliability engineering, and security. You will help define how core infrastructure services are designed, operated, automated, and scaled across the enterprise!
What you'll get to do...
- Design, build, and operate highly available, scalable, and secure infrastructure platforms supporting large-scale Linux environments, with a focus on reliability, resiliency, and operational efficiency
- Lead the architecture, implementation, and operation of infrastructure services, including OpenStack, enterprise secrets management, package management, software promotion pipelines, and platform lifecycle management
- Develop and maintain automation solutions using infrastructure-as-code, Ansible, Python, Go, and self-service capabilities to improve efficiency and reduce operational overhead
- Build and improve observability and reliability practices through monitoring, logging, alerting, dashboards, managing incidents, analyzing underlying causes, disaster recovery, and service health reporting
- Collaborate with security, platform, and engineering teams to strengthen infrastructure maturity, troubleshoot complex distributed systems and cloud environments, provide technical leadership and mentorship, and participate in a 24x7 on-call rotation supporting mission-critical platforms
Your experience should include...
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, Linux Systems Engineering, or a related subject area, with experience operating large-scale production environments
- Strong expertise in Linux systems administration, troubleshooting, performance analysis, storage, filesystems, networking, and operating system internals
- Hands-on experience with cloud infrastructure platforms such as OpenStack or comparable technologies, along with a solid understanding of distributed systems, reliability engineering, and operational excellence practices
- Proficiency in automation, configuration management, and software delivery using tools and technologies such as Ansible, Python, Go, Bash, Git-based development workflows, and CI/CD platforms
- Experience with AI-assisted development or troubleshooting, using closed-loop workflows (generate, review, test, refine) and gu