istaridigital.ai - Sr. Cloud Infrastructure Engineer
Requirements
• Minimum of 5 years of experience in DevOps, Infrastructure Engineering, Platform Engineering, or Site Reliability Engineering • Strong hands-on experience with AWS in production environments • Proven experience operating Kubernetes in production • Kubernetes • Strong experience with Terraform and infrastructure-as-code practices • Terraform • Proven experience building or improving CI/CD pipelines and deployment automation • Solid understanding of cloud networking, IAM, secrets management, and operational controls • Experience with monitoring, logging, and observability tooling • Scripting proficiency in Python, Go, Bash, or similar • Python, Go, Bash, or similar • Excellent troubleshooting and problem-solving skills in complex production environments • Strong communication skills with the ability to explain technical concepts to both technical and non-technical stakeholders • Must live/work in the U.S. • Experience operating stateful workloads on Kubernetes, such as databases or message queues, including persistent storage and backup/recovery • Experience with GitOps-based deployment workflows • Experience packaging software for customer-managed or self-hosted deployment (e.g., Helm charts) • Familiarity with compliance or security frameworks such as FedRAMP, NIST, SOC 2, or similar • FedRAMP, NIST, SOC 2, or similar • Experience with PostgreSQL, cloud storage platforms, and production networking patterns • Experience with configuration management tools such as Ansible • Ansible • Experience with additional cloud platforms such as Azure or GCP • Azure or GCP • Experience with service mesh or advanced Kubernetes networking • Experience supporting customer-facing or mission-critical production infrastructure • Top Secret Security Clearance • Istari is a digital engineering software company building open, scalable infrastructure for engineering data. Our platform lets customers simply and securely integrate engineering models across disciplines, organizations, and security levels — whether they're designing prototypes, performing virtual testing, or training AI and autonomy for complex systems. • Focus is rewarded. Finish is remembered. • Facts are friendly. Even when they are not fun. • Fellowship is fundamental. Make others successful.
Responsibilities
• Design, implement, and maintain scalable, reliable infrastructure in AWS • Operate and improve Kubernetes-based environments, including production workloads • Build and maintain infrastructure as code using Terraform • Improve CI/CD pipelines, deployment workflows, and release automation in partnership with engineering teams • Build and maintain the packaging and reference architectures customers use to install our software in their own environments • Strengthen observability across the platform, including monitoring, logging, alerting, and actionable dashboards • Improve developer experience through tooling, environment automation, and self-service infrastructure • Operate within and preserve the established security and compliance posture of our environments • Monitor, troubleshoot, and resolve complex infrastructure issues with clear and timely communication • Participate in incident response and post-incident analysis • Develop and maintain documentation, runbooks, and technical standards • Identify opportunities to improve cost efficiency, performance, and resilience across environments
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT