Sr Engineer, Systems Reliability - SRE
Job Description
ABOUT THE ROLE:The Sr Engineer, Site Reliability – PMT Operations is the senior technical anchor of the India SRE team. This role carries the full operational scope of the Engineer, Systems Reliability (JD12) - cross-platform incident response, automation development, and platform support across the PMT portfolio - with additional responsibility for technical mentorship, shift leadership, and driving the team's reliability engineering practice.The Sr Engineer sets the standard. You own the hardest incidents, review the automation your peers build, and close the loop on post-incident learning. You are the technical bridge between the India SRE team and the Domain Architect, and the escalation point when incidents exceed the standard SRE tier's resolution scope.WHAT YOU'LL DO:Everything in JD12 (Engineer, Systems Reliability), plus:Serve as technical lead and escalation point for complex, multi-system, or high-severity incidents across the portfolio.Mentor and develop the India SRE team through code review, pairing, and post-incident coaching.Drive the team's runbook and automation quality: set standards, identify gaps, and ensure completeness.Own blameless post-incident reviews for P1 and P2 incidents; identify and track systemic fixes through to resolution.Design and build the team's AI agent roadmap: identify workflows where agents can reduce toil, prototype agents, and establish patterns the team follows.Partner with the Domain Architect on architecture decisions that affect SRE operations; represent the operational team's perspective in design reviews.Contribute to hiring: participate in technical interviews for SRE candidates.WHAT YOU'LL BRING:6–10 years of experience in SRE, software engineering, or technical platform operations - with a strong software development foundation.Deep cross-platform experience: you have operated and automated across multiple enterprise SaaS and cloud-native systems, not just one.Demonstrated technical leadership: you have led incident response, driven blameless postmortems, and influenced engineering practices across a team.Strong software engineering foundation; able to build production-quality automation and tooling, review others' code credibly, and set a quality bar.Track record of measurable toil reduction through automation, AI agents, or tooling investment.Excellent written communication; comfortable authoring architecture notes, RCA documents, and technical standards.Hands-on experience using AI tools (Claude, GitHub Copilot, or equivalent) to build, operate, and improve engineering workflows - and ability to lead others in doing the same.REQUIRED TECHNICAL SKILLS:AI & agent tooling: Proficient in AI coding/automation assistants (Claude, Copilot, or equivalent). Able to design, build, and maintain AI agents for operational workflows; set team standards for agent design and quality; and lead adoption across the SRE team. Expected to be the team's AI tooling anchor.Scripting proficiency in Python, PowerShell, or Bash for production automation and tooling development.Experience with REST API testing and integration debugging.SQL proficiency for data validation, troubleshooting, and reporting.Working knowledge of multiple monitoring/observability platforms (Splunk, Datadog, AppDynamics, or similar); able to design monitoring strategy, not just configure dashboards.CI/CD pipeline experience (GitHub Actions, Jenkins, or similar) - sufficient to contribute fixes and evaluate DevOps work.Version control proficiency (Git); comfortable with branching strategies and PR review.Experience with cloud environments (AWS, GCP, or Azure) - operational level with architectural awareness.PREFERRED SKILLS:Direct hands-on experience with Workday (HCM, Payroll, or Integrations) and/or UKG Pro WFM at the integration / operations level.Experience with Dell Boomi or similar iPaaS integration platforms.Infrastructure-as-code familiarity (Terraform or CloudFormation).Experience designing or participating in SRE team buildouts, vendor transitions, or greenfield operations programs.Familiarity with identity and access management workflows (SSO, SAML, SCIM).. Requirements: name: TMUS Global Solutions. location: Hyderabad, IN. experience: 6 - 10 years. employmentType: Full-Time. Primary Skills: SQL, Python or Powershell or Bash, Kubernetes, SRE or Site Reailibility, Splunk or Open Telemetry or OTEL or Dynatrace or Cloudwatch or App Dynamics or Grafana or Prometheus or New Relic or Datadog or Nagios, Rest or API, Git or CICD or CI/CD, AI or Artificial Intelligence, AWS or Azure or GCP
