wagey.ggwagey.gg
30,534  jobs30,534  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(30,534)/Site Reliability Engineer Role(169)/tenex (15) - Staff Site Reliability Engineer
tenex

tenex - Staff Site Reliability Engineer

Remote - USA1w ago
RemoteStaffNACybersecurityCloud ComputingSite Reliability EngineerCloud ArchitectGCPAWSAzureKubernetesTerraformPrometheusELKPulumiGrafanaDatadogGoogle GKETemporalCross-functional Collaboration

Requirements

• SRE & INFRASTRUCTURE EXPERTISE • Core Engineering: 10+ years of experience in SRE, DevOps, or Software/Systems Engineering, particularly in managing production systems at scale. • Cloud Infrastructure: Deep expertise in public cloud environments (AWS, GCP, or Azure) and managing services such as Kubernetes (EKS/GKE), networking, and storage. • Infrastructure as Code: Extensive experience with tools like Terraform, Pulumi, or similar technologies to manage complex infrastructure deployments. • Observability: Hands-on experience with monitoring, logging, and tracing stacks (e.g., Prometheus, Grafana, ELK, Datadog) to drive data-informed reliability decisions. • Distributed Systems: Solid understanding of microservices architecture, distributed databases, and event-driven systems. • Communication: Clear, concise communication skills and a bias for collaborative problem-solving. • Leadership Alignment: Proven track record of guiding multi-stakeholder initiatives and influencing engineering practices across teams. • Analytical Rigor: Strong problem-solving, debugging, and analytical skills, especially in high-pressure environments. • Domain Background: Prior work in cybersecurity, specifically regarding SIEM, EDR, or SOAR infrastructure. • AI/ML Infrastructure: Experience supporting infrastructure for large-scale AI/ML workloads (e.g., GPU scheduling, LLM serving optimization). • Startup Mentality: Background driving high-impact engineering initiatives in high-growth startups or enterprise SaaS. • Strong familiarity with Agentic Workflows such as Agno, Temporal, etc.. • EDUCATION & CERTIFICATIONS • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field. • Relevant certifications (CKA/CKAD, AWS/GCP Professional Cloud Architect, etc.) are a plus.

Responsibilities

• System Resilience: Design, build, and maintain highly available, scalable, and secure infrastructure to support our AI-native cybersecurity platform. • Automation & Tooling: Develop internal tooling and automation to streamline deployment processes, incident response, and capacity planning. • Performance Engineering: Monitor system performance and proactively identify bottlenecks, optimizing infrastructure for low-latency, high-throughput AI workloads. • Incident Management: Lead incident response efforts, conduct post-mortems, and implement long-term solutions to prevent recurring reliability issues. • Infrastructure as Code (IaC): Manage infrastructure via code, driving consistency, auditability, and scalability across our cloud environments (e.g., AWS, GCP). • Cross-Functional Collaboration: Partner with sibling Engineering teams, Product, and Security teams to ensure reliability is baked into our development lifecycle from concept to production.

Benefits

• Opportunity to work with cutting-edge AI-driven cybersecurity technologies and Google SecOps solutions. • Collaborate with a talented and innovative team focused on continuously improving security operations and system reliability. • Competitive salary and benefits package. • A culture of growth and development, with opportunities to expand your knowledge in AI, cybersecurity, and emerging technologies. • If you're passionate about building resilient infrastructure, scaling AI systems, and working at the intersection of reliability and security, we encourage you to apply!

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

GitLabGitLab - Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms1w ago
·Remote - Canada·$126k - $314k/year + Equity
RemoteNAStaffCloud ComputingSite Reliability EngineerTerraformKubernetesRubyGoAWSGCP
Veeam SoftwareVeeam Software - Senior Site Reliability Engineer- FedRamp1w ago
·Remote - USA·$173k - $173k/year
RemoteNASeniorCloud ComputingArtificial IntelligenceSite Reliability EngineerGoJavaC#TypeScriptAWSAzurePrometheusGrafanaELKTerraformKubernetesDocumentationGitB2BPulumiCloseCosmosElastic Stack
talkiatrytalkiatry - Senior Site Reliability Engineer1w ago
·Remote - Americas
RemoteNASeniorCloud ComputingSite Reliability EngineerTerraformPythonTypeScriptTeam ManagementPrometheusDatadogAWSGrafanaNode.jsReactKubernetesDocumentation
Clover HealthClover Health - Senior Site Reliability Engineer1w ago
·Remote - USA·$160k - $160k/year + Equity
RemoteNASeniorCloud ComputingSite Reliability EngineerGCPAzureAWS
PointClickCarePointClickCare - Senior Site Reliability Engineer, AI Infrastructure1mo ago
·Mississauga, Ontario - Hybrid·$139k - $155k/year
In OfficeNASeniorLife SciencesSoftwareCybersecuritySite Reliability EngineerDocumentationTerraformAzureKubernetesDatabricks
PinterestPinterest - Site Reliability Engineer II, tvScientific1mo ago
·San Francisco, California, United States·$114k - $114k/year + Equity
RemoteNAMidCloud ComputingSite Reliability EngineerBashPythonAWSKubernetesTerraform
IroncladIronclad - Senior Staff Site Reliability Engineer2mo ago
·San Francisco, California, United States - Hybrid·$245k - $270k/year + Equity
In OfficeNAStaffSite Reliability EngineerKubernetesTerraformClaudeZedPulumi
SpotifySpotify - Site Reliability Engineer5mo ago
·Remote - New York, NY·$133k - $190k/year
RemoteNACloud ComputingArtificial IntelligenceSite Reliability EngineerAWSGCPTerraformReactPython
SecurityScorecardSecurityScorecard - Senior Site Reliability Engineer1w ago
·Remote - USA·$152k - $195k/year + Equity
RemoteNASeniorArtificial IntelligenceSite Reliability EngineerBashGoPythonKubernetesJenkinsTerraformHelmPrometheusGoogle GKEGovernancePulumiGrafanaDatadogKafkaMLOps

Browse more by category

Show 169 moreSite Reliability EngineerShow 50 moreCloud ArchitectShow 1,145 moreGCPShow 2,770 moreAWSShow 1,200 moreAzureShow 1,517 moreKubernetesShow 828 moreTerraformShow 193 morePrometheusShow 33 moreELKShow 59 morePulumi
Privacy·Terms··Contact·FAQ·Wagey on X