Sponsored
Verified Job Other Jobs

NOC / SRE Lead – NOC Operations

Thiruvananthapuram, Kerala
Other Jobs
#793672
Remote / WFH
Phykon

Job Description

ob description
Location: Thiruvananthapuram, Kerala, India (Onsite)

Experience: 8+ Years

Employment Type: Full-Time

Department: Network Operations Center (NOC)

Reports To: [Head of Operations / Engineering – to be confirmed]

About The Role

We are seeking an experienced NOC / SRE Lead to own the health, observability, and incident response for a connected fleet of field-deployed systems and the data center infrastructure that supports it.

This is the keystone L3 role in our Network Operations Center (NOC). You will own deep diagnosis and root-cause analysis on the running system, command the hardest escalations, and build the standards, runbooks, observability architecture, and operator-certification program that the rest of the NOC inherits.

As the L3 Lead, you will act as the escalation point above L2 support teams, taking ownership of the most complex incidents while keeping escalations to L4 product engineering limited to genuine code and firmware issues — protecting both fleet uptime and engineering velocity.

This is a greenfield opportunity. In the near term, the role combines a hands-on L3 technical function with NOC leadership. As the NOC scales, the role is designed to grow into a Principal SRE or NOC Manager track.

Key Responsibilities
• Fleet Health & Observability
• Architect the monitoring stack the NOC runs on — scrape architecture, alert rules, dashboards, and SLOs — rather than only consuming it.
• Own end-to-end visibility of the connected fleet across data center, edge Kubernetes, Linux, and network layers.
• Continuously tune alerting to reduce noise and catch degradation early.
• Data Center & Infrastructure Operations
• Oversee the health and availability of data center and colocation infrastructure — compute, storage, network, power, and cooling dependencies — supporting the fleet.
• Coordinate with data center operators, colocation providers, and remote-hands teams for physical interventions, maintenance windows, and capacity changes.
• Manage hardware fault handling, RMA workflows, and site-level incident response across distributed data center and edge sites.
• Coordinate with telecommunications providers, colocation partners, cloud vendors, and infrastructure service providers to resolve service-impacting issues and maintain operational continuity.
• Incident Management & Response
• Serve as incident commander on the hardest escalations, including Sev1/P1 incidents, and drive them to resolution within SLA.
• Perform deep diagnosis and root-cause analysis on the running system using logs, metrics, and telemetry.
• Diagnose complex issues across:
• Network connectivity, routing, NAT/CGNAT, and tunnels
• Degraded or intermittent links to field-deployed hardware
• Production Kubernetes and Linux platform dependencies at the edge
• Assume end-to-end ownership of major incidents, service disruptions, and customer escalations until resolution and formal closure.
• Act as the highest operational escalation point within the NOC for complex technical and service- impacting incidents.
• Provide timely updates to internal stakeholders, leadership teams, and customers during major incidents and service outages.
• Standards, Runbooks & Operator Certification
• Codify diagnostic procedures, escalation paths, and operational standards that make the NOC repeatable and scalable.
• Own the operator-certification layer that qualifies operators to run and support the fleet.
• Develop and continuously improve runbooks and operational procedures inherited by L1/L2 teams.
• Escalation & Coordination
• Maintain escalation hygiene: reserve L4 product engineering for genuinely complex issues requiring code or firmware changes.
• Clearly define problem scope, business impact, and affected systems when escalating, with supporting logs and evidence.
• Facilitate technical communication across engineering and operations teams during major incidents, providing timely stakeholder updates.
• Continuous Service Improvement
• Lead post-incident reviews and drive corrective actions to closure.
• Identify recurring incidents and contribute preventive improvements and monitoring optimization.
• Client & Stakeholder Engagement
• Act as a primary technical point of contact for customers, partners, vendors, and internal stakeholders on operational and service-related matters.
• Participate in customer meetings, operational reviews, service reviews, and technical discussions as a subject matter expert.
• Provide clear, professional, and executive-level communication during major incidents, service disruptions, and operational escalations.
• Build and maintain strong working relationships with customer technical teams while driving accountability and service excellence.
• Leadership & Team Development
• Provide technical leadership, mentoring, and guidance to L1 and L2 engineers.
• Drive knowledge sharing, operational maturity, and continuous improvement initiatives across the NOC function.
• Support the development and maintenance of operator certification programs and competency frameworks.
• Participate in recruitment, onboarding, training, and capability development activities as the NOC organization scales.

Required Skills & Qualifications
• 8+ years in SRE, production engineering, or senior NOC roles with strong platform depth (not generalist IT-NOC).
• Production Kubernetes and Linux experience, ideally in edge or distributed environments rather than pure cloud.
• Ownership-level experience with Prometheus and Grafana — designing scrape architectures and alert rules, not just reading dashboards.
• Hands-on experience with data center operations — running production infrastructure across data center, colocation, or distributed edge sites.
• Strong practical networking: TCP/IP, NAT/CGNAT, tunnels, and degraded-link diagnosis.
• Demonstrated on-call leadership and Sev1 incident command in production environments.
• Proven ability to codify standards, runbooks, and certification programs for operations teams.
• Strong analytical, troubleshooting, and documentation skills, with excellent communication across engineering and business stakeholders.
• Strong ownership mindset with the ability to independently drive incidents, projects, and operational improvements to completion.
• Ability to remain calm, structured, and decisive during high-severity incidents and customer escalations.
• Excellent stakeholder management and communication skills, including the ability to communicate complex technical concepts to both technical and non-technical audiences.
• Experience interacting directly with enterprise customers, vendors, and senior stakeholders in operational environments.

Tools & Technologies

Observability & Telemetry: Prometheus, Grafana, and related monitoring/alerting platforms

Platform: Kubernetes, Linux (edge/distributed deployments)

Automation: Infrastructure-as-Code tooling (advantage)

Incident & Collaboration: Incident management, ticketing, and on-call platforms

Preferred Qualifications
• Greenfield NOC build experience — standing up a NOC, runbooks, and on-call from scratch (strongest positive signal).
• Prior NOC/SRE experience supporting a hardware or field-deployed product.
• Exposure to OT/BMS environments and comfort bridging IT and industrial systems.
• Relevant certifications (Kubernetes, Linux, networking, or cloud) are an advantage.
• Highly desired: Proven experience building or scaling a Network Operations Center (NOC) from the ground up, including monitoring strategy, runbooks, escalation models, operational processes, and team development.

Work Environment
• Operates within a 24×7 Network Operations Center supporting a live fleet of field-deployed systems and distributed data center infrastructure.
• May require participation in rotational shifts, weekend coverage, planned maintenance activities, and on-call support arrangements.
• Serves as a key escalation point for critical incidents and major customer-impacting events.
• Acts as a primary technical point of contact for customers, partners, and internal stakeholders during operational reviews and service incidents.
• Expected to perform effectively in high-pressure operational environments while maintaining clear communication, leadership, and decision-making capabilities.
• Provides technical leadership, mentoring, and operational guidance to L1 and L2 support teams.
Report this listing
View more Lead Jobs in Thiruvananthapuram →
Sponsored

Similar Openings in Other Jobs

More jobs you might like

Legal Analyst yashraj Verified
Bengaluru, Karnataka Other Jobs

Job description We're looking for Legal Analysts! Responsibilities • Manage and process agreements in partnership with sales, procurement, a...

Posted Just now View Details
Client Acquisition Manager yashraj Verified
Manjēshvar, Kerala Other Jobs

Hire22.ai is India’s 1st Agentic Job Portal for mid and senior hiring. Hire22 enables employers to create a JobCoNCT and initiate the hiring...

Posted 1m ago View Details
Bengaluru, Karnataka Other Jobs

Job description Senior Rust Engineer - Remote Kake is a remote-first company and a people-first global community of senior engineers. Kake e...

Posted 3m ago View Details
Enterprise Sales Director yashraj Verified
Hyderabad, Telangana Other Jobs

Hiring | VP – Business Development | Digital Manufacturing & Data Analytics We’re looking for a senior enterprise sales leader with 16+ ...

Posted 4m ago View Details
Pune, Maharashtra Other Jobs

Key focus areas: ● P&L Owner of a cohort of Business Managers (10-15) in a region/city ● Ownership to Client Satisfaction of all project...

Posted 5m ago View Details
HSEQ Manager yashraj Verified
Bālāpur, Maharashtra Other Jobs

Company Description BASIL HOSPITALITY Pvt Ltd is a hospitality-focused company based in Pune, Maharashtra, India, with operations supporting...

Posted 6m ago View Details
Hyderabad, Telangana Other Jobs

Job description Designation - Key Account Manager – (Placements and Outreach Manager) Edunet Foundation is looking for experienced and resul...

Posted 7m ago View Details