wagey.ggwagey.gg
31,339  jobs31,339  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(31,339)/Site Reliability Engineer Role(176)/GitLab (213) - Site Reliability Engineer
GitLab

GitLab - Site Reliability Engineer

Remote - Bangalore, India+ Equity1mo ago
RemoteAPACCloud ComputingSite Reliability EngineerGoAWSTerraformPrometheusGrafanaDockerAnsibleTalent AcquisitionDocumentationQuality Assurance

Requirements

• Professional experience operating production infrastructure on AWS at scale, including EC2, Auto Scaling Groups, IAM, and VPC networking. • Strong infrastructure-as-code experience with Terraform, including writing and refactoring modules used by other teams. • Proficiency in Go for building and debugging infrastructure tooling, or strong experience in another systems language and willingness to work in Go daily. • Practical knowledge of CI/CD systems and job execution: how pipelines schedule work, how ephemeral build environments get provisioned and torn down, and what makes CI workloads reliable. • Experience with observability practices such as metrics, dashboards, alerting, logging, and service-level-objective-based monitoring, using tools such as Prometheus, Grafana, and OpenSearch. • Experience with on-call rotations and incident management for customer-facing systems. • Strong problem-solving skills, excellent written communication, and comfort working asynchronously across Americas, Europe, Middle East, Africa, and Asia-Pacific time zones. • Direct GitLab Runner experience, familiarity with configuration management such as Ansible, and container tooling such as Docker are a plus. • How GitLab Supports Full-Time Employees • Benefits to support your health, finances, and well-being • Flexible Paid Time Off • Team Member Resource Groups • Equity Compensation & Employee Stock Purchase Plan • Growth and Development Fund • Please note that we welcome interest from candidates with varying levels of experience; many successful candidates do not meet every single requirement. Additionally, studies have shown that people from underrepresented groups are less likely to apply to a job unless they meet every single qualification. If you're excited about this role, please apply and allow our recruiters to assess your application. • Country Hiring Guidelines: GitLab hires new team members in countries around the world. All of our roles are remote, however some roles may carry specific location-based eligibility requirements. Our Talent Acquisition team can help answer any questions about location after starting the recruiting process. • Country Hiring Guidelines:

Responsibilities

• Design, build, and operate AWS infrastructure for Hosted Runners across many single-tenant environments, including Elastic Compute Cloud (EC2), Auto Scaling Groups, Virtual Private Cloud (VPC) networking, subnets, Network Address Translation (NAT), network access control lists, PrivateLink, Identity and Access Management (IAM), and Elastic Container Registry (ECR). • Develop and maintain infrastructure as code using Terraform, contributing to common modules and the deployment tooling that provisions and upgrades runner stacks. • Write Go code for our runner tooling and autoscaling components, including the fleeting instance-lifecycle plugins, zero-downtime deployment command-line interface, and reusable infrastructure toolkits. • Build and improve the GitLab CI/CD pipelines that orchestrate blue/green zero-downtime deployments, automated upgrades, quality assurance validation, and performance testing of runner stacks. • Define and monitor service level objectives for CI job execution, including queue times, job success rates, and fleet saturation, and build the Grafana dashboards, alerts, and runbooks behind them. • Participate in an on-call rotation, handle incidents affecting customer CI/CD workloads, and automate away recurring toil. • Run performance and scale testing that reflects real customer workloads, and tune autoscaling parameters for cost and reliability. • Write documentation and runbooks so the broader team can operate runner stacks consistently.

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

pod-networkpod-network - Site Reliability Engineer (APAC)2mo ago
·APAC (UTC+7 to UTC+10)·$100k - $100k/year + Equity
In OfficeAPACCryptocurrencyCloud ComputingSite Reliability EngineerLinuxDockerPrometheusGrafanaRust
Okta, Inc.Okta, Inc. - Staff Site Reliability Engineer3mo ago
·Bengaluru, India
In OfficeAPACStaffCloud ComputingSoftwareSite Reliability EngineerKubernetesTimeline ManagementAWSGCPHelmTerraformPythonGoGoogle GKELinuxAnsible
Wikimedia FoundationWikimedia Foundation - Senior Site Reliability Engineer, Wikimedia Enterprise3mo ago
·Remote - Anywhere·$31k - $31k/year
RemoteWWSeniorCloud ComputingNonprofitSite Reliability EngineerTerraformPythonAnsibleGoAWS
plaudplaud - Site Reliability Engineer - Singapore1mo ago
·Singapore·Equity
In OfficeAPACSeniorCloud ComputingSite Reliability EngineerGoJavaPythonAWSGCPAzureKubernetes
DevelocityDevelocity - Senior Site Reliability Engineer1mo ago
·Remote - Europe (GMT)·Equity
RemoteEMEASeniorCloud ComputingSoftwareSite Reliability EngineerBashPythonAWSKubernetesPrometheusDocumentationTerraformGrafanaJavaKotlin
DevelocityDevelocity - Staff Site Reliability Engineer1mo ago
·Remote - Europe (GMT)·Equity
RemoteEMEAStaffCloud ComputingSoftwareSite Reliability EngineerBashPythonCoachingNew Hire OnboardingAWSKubernetesDocumentationPrometheusTerraformGrafanaJavaKotlin
talkiatrytalkiatry - Senior Site Reliability Engineer1mo ago
·Remote - Americas
RemoteNASeniorCloud ComputingSite Reliability EngineerTerraformPythonTypeScriptTeam ManagementPrometheusDatadogAWSGrafanaNode.jsReactKubernetesDocumentation
Veeam SoftwareVeeam Software - Senior Site Reliability Engineer- FedRamp1mo ago
·Remote - USA·$173k - $173k/year
RemoteNASeniorCloud ComputingArtificial IntelligenceSite Reliability EngineerGoJavaC#TypeScriptAWSAzurePrometheusGrafanaELKTerraformKubernetesDocumentationGitB2BPulumiCloseCosmosElastic Stack
OKXOKX - DevOps / Site Reliability Engineer3mo ago
·Singapore
In OfficeAPACMidCloud ComputingArtificial IntelligenceSite Reliability EngineerGoJavaPythonReactAlibaba Cloud

Browse more by category

Show 176 moreSite Reliability EngineerShow 1,803 moreGoShow 3,029 moreAWSShow 920 moreTerraformShow 232 morePrometheusShow 289 moreGrafanaShow 892 moreDockerShow 165 moreAnsibleShow 553 moreTalent AcquisitionShow 5,249 moreDocumentation
Privacy·Terms··Contact·FAQ·Wagey on X