wagey.ggwagey.gg
30,244  jobs30,244  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(30,244)/Site Reliability Engineer Role(165)/Skylo Technologies (7) - Senior Network Reliability Engineer, Incident Management
Pro members applied to this job 36 hours before you saw itGet Pro ›
Skylo Technologies

Skylo Technologies - Senior Network Reliability Engineer, Incident Management

Remote - Mountain View, California, United States$125k - $135k+ Equity4d ago
RemoteSeniorNADiagnosticsCloud ComputingSite Reliability EngineerPrincipalBashPythonStaff DevelopmentProduct MarketingReporting

Requirements

• 5-10+ years of experience in telecom/wireless operations, network operations, or NRE in a production 24x7 environment. • Demonstrated ability to independently manage incident bridge calls: open the war room, maintain bridge discipline, drive to resolution, and produce a structured incident record. • Strong understanding of telecom network environments and 5G functional components, with sufficient knowledge of RAN (CUSM, eCPRI, PTP/SyncE), 5G Core (AMF, SMF, UPF), and Cloud/OSS to triage intelligently and escalate with context. • Incident and outage management expertise: ability to prioritize by urgency and impact, manage multiple simultaneous events, and operate effectively under high-pressure 24x7 conditions. • Hands-on experience with at least one observability platform: Grafana dashboards, Prometheus alerting, Loki log queries, or equivalent, with the ability to independently navigate to relevant signals during an active incident. • Kubernetes operational literacy: able to run kubectl get pods, describe a failing pod, read container logs, and identify health issues at the level needed to triage and escalate a platform-layer incident. • Structured written communication: capable of producing clear incident timelines, executive stakeholder updates, and post-incident summaries under time pressure. • Ticketing system proficiency (Jira, ServiceNow, or equivalent): incident lifecycle management, escalation workflows, and backlog hygiene. • On-call tooling experience (PagerDuty or equivalent): alert acknowledgement, escalation policy management, and on-call scheduling. • Ability to span departments and build strong working relationships with RAN, Core, Cloud, Engineering, and external partner teams to drive joint incident resolution. • Preferred • Prior experience in a satellite, NTN, or space-to-ground connectivity operational environment. • Experience with OSS/BSS platforms: FCAPS alarm management, event correlation, SNMP trap handling, or EMS integration. • Familiarity with 3GPP NAS/NG-AP signaling flows, RRC state machine, or IMSI lifecycle procedures. • ITIL Foundation certification or demonstrated practical application of ITIL incident, problem, and change management processes. • Scripting ability in Python or Bash, sufficient to automate repetitive operational tasks, parse log output, or build quick diagnostic utilities. • Experience supporting hypercare operations, major network launches, or high-risk change windows. • Exposure to closed-loop automation frameworks or service assurance platforms used in network operations.

Responsibilities

• At Senior NRE level you independently command bridge calls for Sev 1-4 incidents, support subject matter experts as incident coordinator and function as the communication lead during Sev 1 events, drive the execution of runbooks without supervision, manage the full incident lifecycle end-to-end in the ticketing system, and continuously improve the procedures you operate against. You are not a ticket router, you are the structured coordinator who ensures every incident has a clear owner, a live timeline, and a documented resolution path. Assuring Skylo's network availability commitments to MNO partners and meeting SLA obligations is your primary measure of success. • Incident Command & Bridge Coordination • Own end-to-end incident lifecycle management for Sev 1-4: initiate bridge calls, identify the cause and impacted domain, page the correct on-call NRE, maintain bridge discipline with clear ownership and timelines, and drive to restoration. • Escalate to subject matter experts in Operations and Engineering teams when critical or time-sensitive resolution is required, providing full technical context, a structured problem statement, and a documented timeline. • Engage and interface with vendor support teams (RAN vendor, Core vendor, cloud infrastructure) when incident resolution requires external escalation; track vendor SLA response and escalate vendor delays to the domain NRE. • Support hypercare operations during major network launches, high-risk change windows, and special events, maintaining readiness and acting as first responder for any degradation during the hypercare window. • Incident Documentation & Lifecycle Management • Ensure all trouble tickets are created promptly in the incident management system (Jira/ServiceNow) with complete technical details, troubleshooting steps, MOPs followed, and outcome documentation, never closing an incident with an incomplete ticket. • Produce a structured incident timeline artifact within two hours of closure: sequence of events, alarms triggered, actions taken, bridge participants, and all open action items with assigned owners and due dates. • Manage the open incident backlog at optimum levels: track ageing tickets, escalate stalled items, and ensure no incident closes without a documented resolution path or a justified deferral. • Coordinate post-incident review (PIR) scheduling: compile the incident record, gather logs from in-house observability tools and other relevant sources, and deliver a structured problem statement to the domain NREs owning the root cause analysis. • Handle internal, external, and MNO partner incident escalations and follow-ups; interface with Market Operations, OEM contacts, and partner NOCs for joint incident resolution, ensuring external-facing communications are approved before transmission. • Network Availability & SLA Assurance • Assure that Skylo's operated network meets agreed availability KPIs and MNO partner SLA commitments, proactively tracking availability metrics and flagging degradation trends before they breach SLA thresholds. • Track top recurring issues and feed continuous improvement inputs: document repeat-incident patterns, identify the operational gap (missing runbook, stale threshold, absent alert), and route findings to the appropriate domain NRE for action. • Contribute to the weekly and monthly Network Performance Report: incident count by severity and domain, MTTR trends, top issues, SLA compliance summary, and KPI deviation analysis. • Drive proactive measures for network issue detection and isolation, actively participating in the Service Assurance and Automation domain and providing operational input for closed-loop automation requirements. • On-Call & Shift Operations • Participate in the global follow-the-sun on-call rotation as the incident coordination tier: Espoo shift bridges the US overnight window, while Mountain View and Bengaluru shifts cover their respective regions, maintaining 24x7 continuous coverage. • Maintain shift handoff hygiene: produce a written end-of-shift summary covering open incidents, degraded components, active change windows, vendor escalations in flight, and priority items for the incoming shift. • Manage the on-call paging workflow: acknowledge alerts within SLA target windows, escalate to L2 within defined thresholds, and ensure no alert goes unacknowledged across the shift boundary. • Support planned maintenance and change windows: validate pre-change observability coverage, confirm rollback readiness with the NI team, and execute rollback runbooks if a deployment causes service degradation. • Cross-Functional Collaboration • Work in close collaboration across multiple Skylo functions during incidents: RAN NRE, Core NRE, Cloud Infrastructure NRE, BOSS (BSS & OSS), Network Implementation, Product Engineering and Market teams, coordinating without creating confusion by maintaining a single source of truth on the bridge. • Interface with MNO partner NOC teams during shared-impact events: relay technical status updates, manage the partner communication cadence, and escalate partner requests through the correct internal channel. • Surface repetitive manual incident steps to the service assurance automation team, documenting the step, frequency, and toil cost as structured input to the automation backlog. • Technical Domain Knowledge • Maintain working knowledge of RAN architecture relevant to incident triage: CUSM (Control plane, User plane, Synchronization plane, Management plane), eCPRI/CPRI interfaces, PTP and SyncE timing, Netconf, and NTN-specific RAN alarm patterns. • Operate virtualized NTN infrastructure observability tools: navigate Grafana dashboards, learn and use custom in-house tooling capabilities, run kubectl commands to assess pod health, and correlate OSS alarms with underlying infrastructure events. • Build domain depth progressively across the first 12 months; the Senior NRE role is a structured development track toward Staff NRE and Principal NRE, where independent incident command and domain ownership are the primary accountabilities.

Benefits

• Remote, U.S: $125,000 – $135,000 base salary + equity • These ranges reflects the low and high end of the range Skylo reasonably and generally expects to pay the hired candidate in this role. • Skylo is an equal-opportunity employer and we celebrate diversity. We do not discriminate based on race, religion, color, ancestry, national origin, caste, sex, sexual orientation, parent or caregiver status, political affiliation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service consistent with applicable federal, state, and local laws. • We are also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. Please let us know if you need assistance or accommodation due to a disability.

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

AccelaAccela - Principal Site Reliability Engineer1mo ago
·Remote - Based - US·$160k - $185k/year + Equity
RemoteNAPrincipalInsuranceCloud ComputingSite Reliability EngineerPrincipalBashPythonKubernetesAzureChange Management
Backblaze External WebsiteBackblaze External Website - Sr. Site Reliability Engineer4mo ago
·Remote - USA·$150k - $200k/year + Equity
RemoteNASeniorCloud ComputingSite Reliability EngineerPrincipalGoPythonDocumentationLinuxPerformance Management
Cambridge Mobile TelematicsCambridge Mobile Telematics - Principal Site Reliability Engineer, Machine Learning2d ago
·Cambridge, MA, US·$142k - $178k/year + Equity
In OfficeNAPrincipalCloud ComputingArtificial IntelligenceSite Reliability EngineerPrincipalPythonReportingAWSDatadogTerraformLinuxUbuntuDockerKubernetesRayDatabricks
Bedrock Ocean ExplorationBedrock Ocean Exploration - Senior Site Reliability Engineer, Robotics & Cloud Infrastructure4w ago
·Brooklyn, New York, USA·$164k - $220k/year + Equity
In OfficeNASeniorCloud ComputingRoboticsSite Reliability EngineerTeam ManagementCustomer OnboardingBashPythonGo
Stack AVStack AV - Senior Site Reliability Engineer1mo ago
·Remote - Pittsburgh, PA or Remote
RemoteNASeniorCloud ComputingGovernmentSite Reliability EngineerBashPythonLinuxGCPAWS
AmwellAmwell - Senior Site Reliability Engineer1mo ago
·US - Remote - Hybrid·$129k - $140k/year + Equity
In OfficeNASeniorMental HealthCloud ComputingSite Reliability EngineerBashPythonPuppetTerraformAnsible
replitreplit - Staff Site Reliability Engineer2mo ago
·Remote - Europe
RemoteEMEAStaffCloud ComputingSite Reliability EngineerPrincipalGoPythonMentoringReportingKubernetes
runpodrunpod - Site Reliability Engineer1w ago
·Remote - USA·$150k - $200k/year + Equity
RemoteNASeniorSite Reliability EngineerLinuxPerformance ReviewsReportingSlackPrometheusGrafanaPythonBashGo
Andromeda ClusterAndromeda Cluster - Senior Site Reliability Engineer - AI Infrastructure3mo ago
·Remote - San Francisco, California , United States
RemoteNASeniorSite Reliability EngineerBashGoHelmPythonTerraform

Browse more by category

Show 165 moreSite Reliability EngineerShow 761 morePrincipalShow 321 moreBashShow 4,673 morePythonShow 195 moreStaff DevelopmentShow 3,098 moreProduct MarketingShow 6,355 moreReporting
Privacy·Terms··Contact·FAQ·Wagey on X