Senior Network Reliability Engineer, Incident Management
Description
The world still has coverage blind spots. You could help eliminate them at Skylo.
Skylo has pioneered a standards-based approach to satellite connectivity. We connect smartphones and IoT devices directly to satellites. No special hardware, no entirely new networks. Just billions of existing devices, suddenly reachable anywhere on Earth. We're not building toward this future. We're already in it.
Our direct-to-device service is live on millions of activated devices across five continents, covering more than 72 million square kilometers, in partnership with leading satellite operators, mobile network operators, Tier-1 chipset makers, and OEMs worldwide. And we're just getting started.
At the heart of it all is Skylo's commercial NTN vRAN: a 3GPP standards-based, cloud-native platform that seamlessly bridges terrestrial and satellite networks. It's the infrastructure that makes true anywhere, anytime connectivity possible.
When you join Skylo, you'll work at the intersection of three markets reshaping how the world stays connected: mass-market consumer devices, automotive, and industrial IoT. Enabling people outdoors and critical workflows in the world's most remote places.
This is a rare chance to work on technology that matters, at a company that's already proving it works
HOW YOU WILL IMPACT SKYLO
As a Network Reliability Engineer (NRE) for Incident Management in the Global Product Support & Customer Success organization, you are the operational nerve center of Skylo's 24x7 incident response function. You own the bridge, from first alert to full-service restoration. On a network where every subscriber is roaming over satellite, structured incident coordination directly determines whether millions of devices stay connected. You operate across all three Skylo hubs (Mountain View, Espoo, Bengaluru) and across every network domain: RAN, 5G Core, Cloud Infrastructure, and OSS, ensuring incidents are triaged, escalated, and resolved with speed and discipline.
At Senior NRE level you independently command bridge calls for Sev 1-4 incidents, support subject matter experts as incident coordinator and function as the communication lead during Sev 1 events, drive the execution of runbooks without supervision, manage the full incident lifecycle end-to-end in the ticketing system, and continuously improve the procedures you operate against. You are not a ticket router, you are the structured coordinator who ensures every incident has a clear owner, a live timeline, and a documented resolution path. Assuring Skylo's network availability commitments to MNO partners and meeting SLA obligations is your primary measure of success.
KEY RESPONSIBILITIES
INCIDENT COMMAND & BRIDGE COORDINATION
- Serve as the central command point during all network degradations, service disruptions, and subscriber-impacting events, opening the bridge, initiating the incident management process, and maintaining command from first alert through full-service restoration.
- Own end-to-end incident lifecycle management for Sev 1-4: initiate bridge calls, identify the cause and impacted domain, page the correct on-call NRE, maintain bridge discipline with clear ownership and timelines, and drive to restoration.
- Prioritize incidents according to urgency and business impact, classifying severity accurately using alarm signatures, subscriber impact data, and domain KPI telemetry from available OSS systems.
- Escalate to subject matter experts in Operations and Engineering teams when critical or time-sensitive resolution is required, providing full technical context, a structured problem statement, and a documented timeline.
- Engage and interface with vendor support teams (RAN vendor, Core vendor, cloud infrastructure) when incident resolution requires external escalation; track vendor SLA response and escalate vendor delays to the domain NRE.
- Support hypercare operations during major network launches, high-risk change windows, and special events, maintaining readiness and acting as first responder for any degradation during the hypercare window.
INCIDENT DOCUMENTATION & LIFECYCLE MANAGEMENT
- Ensure all trouble tickets are created promptly in the incident management system (Jira/ServiceNow) with complete technical details, troubleshooting steps, MOPs followed, and outcome documentation, never closing an incident with an incomplete ticket.
- Produce a structured incident timeline artifact within two hours of closure: sequence of events, alarms triggered, actions taken, bridge participants, and all open action items with assigned owners and due dates.
- Manage the open incident backlog at optimum levels: track ageing tickets, escalate stalled items, and ensure no incident closes without a documented resolution path or a justified deferral.
- Coordinate post-incident review (PIR) scheduling: compile the incident record, gather logs from in-house observability tools and other relevant sources, and deliver a structured problem statement to the domain NREs owning the root cause analysis.
- Handle internal, external, and MNO partner incident escalations and follow-ups; interface with Market Operations, OEM contacts, and partner NOCs for joint incident resolution, ensuring external-facing communications are approved before transmission.
NETWORK AVAILABILITY & SLA ASSURANCE
- Assure that Skylo's operated network meets agreed availability KPIs and MNO partner SLA commitments, proactively tracking availability metrics and flagging degradation trends before they breach SLA thresholds.
- Track top recurring issues and feed continuous improvement inputs: document repeat-incident patterns, identify the operational gap (missing runbook, stale threshold, absent alert), and route findings to the appropriate domain NRE for action.
- Contribute to the weekly and monthly Network Performance Report: incident count by severity and domain, MTTR trends, top issues, SLA compliance summary, and KPI deviation analysis.
- Drive proactive measures for network issue detection and isolation, actively participating in the Service Assurance and Automation domain and providing operational input for closed-loop automation requirements.
ON-CALL & SHIFT OPERATIONS
- Participate in the global follow-the-sun on-call rotation as the incident coordination tier: Espoo shift bridges the US overnight window, while Mountain View and Bengaluru shifts cover their respective regions, maintaining 24x7 continuous coverage.
- Maintain shift handoff hygiene: produce a written end-of-shift summary covering open incidents, degraded components, active change windows, vendor escalations in flight, and priority items for the incoming shift.
- Manage the on-call paging workflow: acknowledge alerts within SLA target windows, escalate to L2 within defined thresholds, and ensure no alert goes unacknowledged across the shift boundary.
- Support planned maintenance and change windows: validate pre-change observability coverage, confirm rollback readiness with the NI team, and execute rollback runbooks if a deployment causes service degradation.
CROSS-FUNCTIONAL COLLABORATION
- Work in close collaboration across multiple Skylo functions during incidents: RAN NRE, Core NRE, Cloud Infrastructure NRE, BOSS (BSS & OSS), Network Implementation, Product Engineering and Market teams, coordinating without creating confusion by maintaining a single source of truth on the bridge.
- Engage appropriate stakeholders based on incident signature, subscriber impact, and domain ownership, avoiding over-escalation and under-escalation with disciplined severity classification.
- Interface with MNO partner NOC teams during shared-impact events: relay technical status updates, manage the partner communication cadence, and escalate partner requests through the correct internal channel.
- Surface repetitive manual incident steps to the service assurance automation team, documenting the step, frequency, and toil cost as structured input to t