Greenhouse - Open Role — Staff Software Engineer, Infrastructure
Requirements
• 10+ years of experience in software engineering, platform engineering, infrastructure engineering, or systems operations for highly available production systems. • Experience managing blockchain nodes, including high availability, failover, and operational resilience. • Experience operating infrastructure in high-traffic, customer-critical, or security-sensitive environments. • Strong production experience with PostgreSQL. • Proven experience provisioning and managing infrastructure with Terraform or similar infrastructure-as-code tools. • Deep experience with containerized infrastructure and Kubernetes in highly available environments. • Proficiency with programming and scripting languages such as .NET, Go, Bash, or TypeScript. • Experience with messaging and queueing technologies such as RabbitMQ or AMQP. • Experience with observability tools such as OpenTelemetry, Grafana, Loki, Prometheus, or similar platforms. • Familiarity with GitOps practices using platforms like Argo CD or Flux. • Experience working with confidential computing solutions like AWS Nitro Enclaves, IBM Hyper Protect Virtual Servers, or GCP Confidential Computing. • Excellent ability in solving problems and the ability to diagnose complex distributed system issues. • Strong written and verbal communication skills, with a track record of working effectively across teams. • A focus on clear procedures, strong documentation habits, and a dedication to operational excellence.
Responsibilities
• Design, build, and operate scalable, resilient infrastructure across cloud providers including Azure, AWS, GCP, and IBM Cloud. • Lead infrastructure architecture and reliability improvements for critical custody services. • Implement and improve monitoring, alerting, logging, and observability across distributed systems. • Own and evolve blockchain node infrastructure, including high availability, failover, and provider management. • Build automation for infrastructure provisioning, deployments, testing, failover, incident response, and operational maintenance. • Drive deployment operations and release reliability for platform and product services. • Proactively identify performance bottlenecks, reliability risks, and operational gaps before they impact customers. • Participate in on-call rotations, support production incidents, and help improve incident response and post-incident learning. • Contribute to internal platform tools, services, and developer workflows. • Create clear documentation, runbooks, and operational procedures for critical systems. • Mentor engineers and provide technical guidance on infrastructure, reliability, and platform engineering decisions.
Benefits
• Competitive benefits that cover physical and mental healthcare, retirement, family forming, and family support • Employee giving match • Mobile phone stipend • Take Care of Yourself • R&R days so you can rest and recharge • Generous wellness reimbursement and weekly onsite & virtual programming • Generous vacation policy - work with your manager to take time off when you need it • Industry-leading parental leave policies. Family planning benefits. • Catered lunches, fully-stocked kitchens with premium snacks/beverages, and plenty of fun events • Benefits listed above are for full-time employees.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT