Sponsored
Verified Job IT / Software / Data Analyst

Lead Systems Operations Engineer

Hyderabad, Andhra Pradesh
IT / Software / Data Analyst
#690927
Remote / WFH

Job Description

About this role:

Wells Fargo is seeking a Lead Systems Operations Engineer
Lead Site Reliability Engineer (App SRE) is responsible for driving reliability, automation, observability, and performance for mission‑critical applications and platforms.
This role blends software engineering excellence with operational expertise to deliver stable, scalable, and resilient services, while reducing toil and shifting operations left across the application lifecycle.
The Lead SRE acts as a technical authority and mentor, partnering with application, platform, and DevOps teams to embed reliability into design, delivery, and operations.

In this role, you will:

Lead complex, broad impact initiatives including provision of high-level systems consultation for the technology teams
Work as key participant in large scale planning of computer systems and network infrastructure for Systems Operations functional area
Review and analyze complex technical challenges, as well as escalated support issues related to core business solutions that require in depth evaluation of multiple factors, such as alternatives, enhancements, periodic systems reviews, or improvements to existing systems
Make decisions on technical changes and enhancements
Consult with engineering team on change design requiring solid understanding of technical process controls or standards that influence and drive new initiatives
Collaborate and consult with technical peers, colleagues, and mid to more experienced level managers to resolve systems support issues and achieve goals

Required Qualifications:

5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
Job Expectations:

Partner with application, platform, and business stakeholders to define, implement, and govern SLIs, SLOs, and error budgets, balancing reliability with delivery velocity.
Lead the design and continuous improvement of observability, telemetry, monitoring, and alerting, ensuring actionable insights and reduced alert fatigue.
Identify, prioritize, and implement automation and self‑healing solutions to eliminate operational toil and improve service resilience.
Own and lead production readiness and go‑live activities, including NFR validation, Permit to Operate (PTO), and operational risk assessments.
Provide engineering‑led application production support, acting as an escalation point for complex application and platform issues.
Lead and troubleshoot major incident response (P1/P2/P3), drive in‑depth root cause analysis (RCA), and ensure preventative actions are implemented to achieve long‑term stability.
Influence and guide teams to shift reliability left by embedding SRE practices into design, CI/CD pipelines, and release processes.
Mentor junior engineers and contribute to SRE standards, best practices, and operating models.
Collaborate and consult with technical peers, colleagues, and mid to more experienced level managers to resolve systems support issues and achieve goals

Additional Required Qualifications:

5+ years of hands‑on experience in production application support engineering, with a strong focus on reliability, availability, and operational excellence.
3+ years of experience leading and operating production systems in a Site Reliability Engineering, DevOps, or Reliability Engineering role.
3+ years of experience working with enterprise schedulers and databases, such as Autosys, Oracle, and MS SQL Server.
3+ years of experience supporting applications on Kubernetes / OpenShift platforms.
Strong understanding of observability concepts (metrics, logs, traces, APM) using tools such as AppDynamics, ThousandEyes, Prometheus, Grafana, Splunk, and Aternity.
Solid experience with web‑based applications and application servers.
Proven experience providing technical leadership and hands‑on execution in complex enterprise environments.
Excellent communication and documentation skills, with the ability to influence both technical and non‑technical stakeholders.

Desired Qualifications:

Strong scripting and automation skills using Unix/Linux shell, Python, or Ansible.
Experience with cloud and container platforms, including Red Hat OpenShift (OCP) and modern cloud architecture concepts.
Experience working with COTS platforms in regulated or large‑scale enterprise environments.
View more Lead Jobs in Hyderabad →
Sponsored

Similar Openings in IT / Software / Data Analyst

More jobs you might like

Senior Data Analyst gaurav2 Verified
Jaipur, Rajasthan IT / Software / Data Analyst

Job Description: Senior Data Analyst Position Overview: We are seeking a skilled and experienced Senior Data Analyst to join our dynamic tea...

Posted 22m ago View Details
Full Stack Engineer HR Manager Verified
Mumbai, Maharashtra IT / Software / Data Analyst

Axis My India, is Indias leading Consumer Data Intelligence Company and a Harvard Business School case study, offering research & marketing ...

Posted 29m ago View Details
LaravelCodeIgniterPHP Developer HR Manager Verified
Remote / WFH IT / Software / Data Analyst

Looking for full time Laravel developer with 3+ years of experience with PHP and at least 1 year experience with CodeIgniter or Laravel. Wor...

Posted 29m ago View Details
Technical Officer HR Manager Verified
Remote / WFH IT / Software / Data Analyst

Looking for Officer/Sr. Officer in Technical Department. Responsibilities: • Handling Shift Technical activity related to uniformity & balan...

Posted 29m ago View Details
Remote / WFH IT / Software / Data Analyst

Hiring for PHP Codeigniter Developer at Ahmedabad, Sola. Candidate should be at list 1 to 2 year experience with MVC custom web application ...

Posted 29m ago View Details