hamsa - DevOps Engineer - Blockchain and Data Platforms
Requirements
• 5+ years of experience in DevOps, SRE, or platform engineering roles • Strong hands-on experience with: Kubernetes (EKS or equivalent), ArgoCD / GitOps workflows • Experience building and operating production-grade distributed systems • Strong experience with observability tooling (metrics, logs, tracing) • Experience with cloud platforms (AWS preferred) • Experience with CI/CD pipelines and release automation • Strong understanding of: containerization (Docker), networking concepts, system reliability and fault tolerance • Experience integrating with enterprise security and DevOps frameworks • Strong troubleshooting and incident response skills • Bachelor’s degree in Computer Science, Engineering, or a related field • MBA or advanced technical degree is a plus
Responsibilities
• deploy and manage services using GitOps (ArgoCD) • operate Kubernetes (EKS) clusters in production • ensure high availability and resilience of critical services • implement end-to-end observability (metrics, logs, traces) • integrate with enterprise tooling for security, identity, and compliance • support application teams with reliable, scalable infrastructure • Kubernetes & Cloud Platform Engineering • Design, deploy, and operate Kubernetes clusters (EKS) for production workloads • Define and manage cluster architecture, networking and ingress, workload isolation and scaling, resource optimization • Ensure platform reliability, performance, and cost efficiency • GitOps & Continuous Deliver • Implement and manage GitOps workflows using ArgoCD • Define deployment standards for micro-services, data pipelines, blockchain-integrated services • Maintain environment consistency across dev, staging, and production • Support safe, auditable, and repeatable releases • Observability & Reliability Engineering • Build and maintain a comprehensive observability stack, including: metrics (e.g., Prometheus), logs (e.g., ELK / OpenSearch or similar), distributed tracing (e.g., OpenTelemetry) • Define SLIs, SLOs, and alerting strategies • Enable proactive detection and resolution of system issues • Partner with engineering teams to improve system resilience and performance • Production Operations & Incident Management • Establish and maintain production readiness standards • Lead or participate in: incident response, root cause analysis (RCA), post-mortems • Continuously improve system reliability and operational processes • Security, Compliance & Enterprise Integration • Integrate platform with institution-grade DevOps and security tooling, including: identity and access management, secrets management, audit logging, vulnerability scanning • Ensure compliance with internal policies and regulatory expectations • Work closely with cybersecurity teams to embed security into the platform • Provide tooling and support for application teams to: deploy services reliably, monitor and debug systems, scale workloads effectively • Improve developer experience through automation and platform abstractions • Infrastructure as Code & Automation • Build and maintain infrastructure using Infrastructure as Code (IaC) (e.g., Terraform or similar) • Automate provisioning, scaling, and operational tasks • Ensure reproducibility and consistency across environments
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT