Skylo Technologies - Staff Network Reliability Engineer, Service Assurance AIOps & Automation
Requirements
• 14+ years in platform automation, observability/SRE tooling, service assurance development, or network automation engineering — including 4+ years operating at a Principal / Staff-plus / Architect level owning org-wide automation or platform architecture in a 24×7 production environment. • Demonstrated track record architecting a centralized automation or platform framework consumed by multiple operations and/or engineering teams: reference architecture, API and schema governance, and an extensibility model that prevented fragmentation. • Expert software development in Python (primary) with production-grade discipline — modular architecture, testing, CI/CD integration, and code review; Go a strong plus. • Deep DAG-based workflow orchestration mastery: Apache Airflow, Prefect, Dagster, or equivalent — multi-step pipelines with branching, error handling, retry, and human-in-the-loop gates. • Data engineering at architecture level: real-time streaming (Pub/Sub, Kafka), data lake design (BigQuery or equivalent), schema design, DataOps, and data-quality governance. • Production observability stack ownership: Prometheus (PromQL, recording and alerting rules), Grafana, VictoriaMetrics, and log aggregation (Loki or ELK). • AI-assisted development as a daily accelerator — Claude Code, Claude Cowork, GitHub Copilot, or equivalent — with the judgment to set team standards for safe, auditable AI-in-the-loop automation. • Kubernetes and cloud-native architecture (GKE or EKS), Helm authorship, and GitOps via ArgoCD; CI/CD pipeline design (GitLab CI or Jenkins) with security scanning and rollback. • Alert correlation and event processing at scale, plus governance, audit, and security design for production-changing automation. • Executive-grade communication: translate operational and engineering toil into architecture and an operating model, and present automation impact and roadmap to engineering leadership. • Preferred • Automation for telecom or NTN network operations and engineering: 5G Core, RAN, and cloud-infrastructure alarm patterns plus network-engineering change and validation workflows. • Agentic AI development: LLM-powered diagnostic agents, automated RCA assistants, or AI-driven operational copilots. • ServiceNow or Jira automation (API-driven ticket lifecycle, workflow automation, bot integration) and advanced Grafana development (custom plugins, dashboard governance). • Chaos engineering and automated resilience testing; DataOps/MLOps practices (feature stores, model serving, experiment tracking, pipeline versioning) for operational ML. • OSS/BSS integration awareness for provisioning, assurance, inventory, and reconciliation impacts. • Certifications: CKA, AWS/GCP Professional, Confluent Kafka, or Apache Airflow certification
Responsibilities
• Automation Architecture & Technical Strategy • Own the end-to-end support and development of framework and multi-year roadmap for Skylo’s centralized network automation framework — the single platform that Network Operations (Core, RAN, Cloud, IM NRE teams) and Network Engineering consume for automated diagnostics, remediation, change, and validation. • Own the end-to-end support and development of framework and multi-year roadmap • Serve as the design authority for the automation architecture review: set engineering standards, API contracts, data schemas, and governance guardrails that prevent automation fragmentation across domains and between Operations and Engineering. • Serve as the design authority • Define the platform extensibility model so domain NREs and network engineers compose domain-specific workflows from governed, reusable building blocks — without forking framework code or re-implementing primitives. • Define the platform extensibility model • Own build-vs-buy and tooling decisions (orchestration engine, data lake, observability stack, AI tooling), translating them into capital and run-cost implications for engineering leadership with a clear, defensible rationale. • Own build-vs-buy and tooling decisions • Establish the operating model around the framework: ownership boundaries, readiness gates, promotion criteria, and the cross-functional cadence that keeps Operations and Engineering aligned on one automation backlog rather than competing toolchains. • Establish the operating model around the framework • Centralized Framework & Platform Engineering • Architect the centralized framework: reusable health-check modules, diagnostic libraries, remediation actions, and notification/escalation handlers — composable components with versioned, governed interfaces that domain teams chain into workflows. • Architect the centralized framework • Define the automation API layer: standardized interfaces for alert ingestion (Pub/Sub, Prometheus Alertmanager), action execution (kubectl, GCP API, NF CLIs, network configuration interfaces), and result reporting (ITSM update, Slack, dashboard refresh). • Define the automation API layer • Set codebase and delivery standards: testing frameworks, CI/CD pipelines, linting, unit/integration tests, security scanning, and staged rollout via GitLab CI — and hold the organization to them through reference implementations and architecture review. • Set codebase and delivery standards • DAG-Based Orchestration & Closed-Loop Automation • Define the DAG-based orchestration standard (Apache Airflow or equivalent): every workflow a composable, auditable pipeline of diagnostic steps, decision gates, remediation actions, and validation checks. • Define the DAG-based orchestration standard • Architect closed-loop automation for P3/P4 fault classes across both operations and engineering change: trigger → diagnostic DAG → root-cause classification → remediation → validation → auto-close with a full evidence trail and zero human touch for covered classes. • Architect closed-loop automation for P3/P4 fault classes • Own the progressive automation tier model: Tier 1 (auto-diagnose and recommend), Tier 2 (auto-remediate with human approval), Tier 3 (fully autonomous closed-loop) — with confidence scoring and explicit promotion/demotion criteria based on measured incident outcomes. • Own the progressive automation tier model • AI-Assisted Development & Intelligent Operations • Set the team standard for AI-assisted automation development — Claude Code, Claude Cowork, and GitHub Copilot as workflow-integrated force multipliers — and personally prototype reference implementations to prove the bar: from problem statement to staged prototype in hours, not weeks. • Set the team standard for AI-assisted automation development • Architect AI-powered diagnostic agents and alert-correlation engines that correlate Core NF, RAN, Cloud, and OSS signals into actionable root-cause hypotheses and suppress alarm storms — grounded in trusted data, measurable accuracy, and auditability, not novelty. • Architect AI-powered diagnostic agents and alert-correlation engines • Make AI a structural part of the operating model — how the organization detects, isolates, resolves, reports, and learns — with human-in-the-loop gates, model/agent governance, and explicit thresholds for operational trust. • Make AI a structural part of the operating model • Data Engineering & Operational Intelligence • Architect the real-time data pipelines and the operational data lake (BigQuery or equivalent): structure telemetry, incident records, KPI time-series, and automation execution logs into queryable, governed datasets that enable trend analysis, capacity forecasting, and automation-effectiveness measurement. • Architect the real-time data pipelines and the operational data lake • Set DataOps standards: version-controlled pipeline definitions, automated data-quality checks, schema-evolution management, and pipeline health monitoring — the data layer must be as reliable as the network it observes. • Set DataOps standards • Define the dashboarding and analytics standard (Grafana): high-fidelity operational views for NRE teams, network engineers, Incident Managers, and executive leadership, built from a single governed data layer. • Define the dashboarding and analytics standard • Governance, Security & Operational Readiness • Enforce automation governance: every automated action has a defined rollback path, human-in-the-loop validation for P1/P2 actions, audit logging, and safe-handoff to humans when confidence falls below threshold. • Enforce automation governance • Build security into the operating model: least-privilege identity for automation actors, secrets management, full audit trails on production-changing actions, controlled partner/vendor access, and security exceptions that are explicit, time-bound, owned, and visible. • Build security into the operating model • Account for downstream blast radius: ensure automated actions that touch provisioning, inventory, entitlement, assurance, or reporting are designed without creating OSS/BSS or revenue-assurance operational debt. • Account for downstream blast radius • Organizational Leadership & Cross-Functional Partnership • Act as the automation intake and design authority for all NRE domain teams and Network Engineering: convert surfaced toil and engineering pain into prioritized, measurable automation with defined MTTA/MTTR targets and owners. • Act as the automation intake and design authority • Partner with Ops Platform Engineering on shared event schemas, API contracts, and policy engines; partner with Security, OSS/BSS, and Customer Operations to validate operational impact before rollout. • Partner with Ops Platform Engineering • Mentor and raise the bar for Staff and Senior automation developers; establish the technical-leadership and review pipeline for network automation across all hubs. • Mentor and raise the bar • Drive continuous improvement: track automation coverage, MTTA/MTTR impact per workflow, false-positive rate, human-override frequency, and adoption — every automation is measured, not assumed. • Drive continuous improvement
Benefits
• Monthly allowances for wellness and education reimbursement • A generous time off policy, holidays, and the opportunity to temporarily work abroad • Once-in-a-lifetime opportunity to be part of developing and running the world’s first commercial, live direct-to-device satellite network and service • Access to a world-class team and talent across tech domains: software, hardware, chipsets, telecom, satellite and network virtualization • Open, transparent, inclusive culture that blends Silicon Valley, Nordic and South Asia characteristics • EEO Statement • Skylo is an equal-opportunity employer and we celebrate diversity. We do not discriminate on the basis of race, religion, color, ancestry, national origin, caste, sex, sexual orientation, parent or caregiver status, political affiliation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service consistent with applicable federal, state, and local laws. • We are also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. Please let us know if you need assistance or accommodation due to a disability.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT