Panoptyc - Cloud Operations Manager
Requirements
• Cloud Expertise: 5+ years working with AWS (EC2, RDS, S3, ECS, Fargate, IAM, IaC, networking, security groups) with hands-on architecture and troubleshooting experience • Cloud Expertise: • DevOps/CI-CD: Strong background with CI/CD tools (GitHub Actions, Jenkins, GitLab CI, CircleCI) including pipeline design, testing strategies, and deployment automation • DevOps/CI-CD: • People Management: Experience managing and developing technical team members, with patience and skill in coaching less experienced engineers through complex technical concepts • People Management: • Process Orientation: Track record of implementing operational processes that stick - documentation standards, change management, on-call rotations, incident response • Process Orientation: • Security & Compliance: Understanding of cloud security best practices, IAM policies, SOC 2 considerations, and infrastructure-as-code for audit trails • Security & Compliance: • Communication: Ability to translate technical complexity into clear explanations for both engineering teams and non-technical stakeholders • Communication: • Optimizing costs for ML/GPU workloads • Terraform or CloudFormation infrastructure-as-code experience • Kubernetes/container orchestration knowledge • Experience with monitoring/observability tools (DataDog, CloudWatch, Grafana, PagerDuty) • Background in SRE (Site Reliability Engineering) practices • AWS certifications (Solutions Architect, SysOps Administrator) • Experience with SOC 2 compliance and security audits • Scripting skills in Python, Bash, or (bonus points) Ruby for automation • What Success Looks Like • IT operations mistakes decrease significantly through better processes and mentorship • AWS infrastructure follows architecture best practices with documented patterns • CI/CD pipelines are reliable, fast, and trusted by engineering teams • Team member(s) develop stronger cloud engineering skills and confidence • Incident response times improve with clear runbooks and escalation procedures • Infrastructure costs are optimized without sacrificing reliability or performance
Responsibilities
• Team Leadership: Manage and mentor IT operations team member(s), providing technical guidance, establishing clear processes, and developing their cloud engineering capabilities through hands-on coaching • Team Leadership: • Cloud Architecture: Own and grow our AWS environment, implementing proper architecture patterns, cost optimization, security hardening, and disaster recovery procedures across web, AI/ML, and edge computing workloads • Cloud Architecture: • CI/CD Operations: Design and maintain automated CI/CD pipelines for infrastructure and application services, including support for containerized, hardware-dependent, and ML-related deployments • CI/CD Operations: • Process & Standards: Build operational maturity through documentation, runbooks, change management processes, incident response procedures, and knowledge transfer protocols across cloud and edge systems • Process & Standards: • Engineering Support: Partner with engineering teams to provide infrastructure that enables fast, safe deployments while maintaining system reliability and security • Engineering Support: • Incident Management: Lead incident response, conduct blameless post-mortems, and implement preventive measures to reduce recurring issues • Incident Management:
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT