wagey.ggwagey.gg
30,165  jobs30,165  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(30,165)/Site Reliability Engineer Role(165)/Cambridge Mobile Telematics (1) - Principal Site Reliability Engineer, Machine Learning
Pro members applied to this job 36 hours before you saw itGet Pro ›
Cambridge Mobile Telematics

Cambridge Mobile Telematics - Principal Site Reliability Engineer, Machine Learning

Cambridge, MA, US$142k - $178k+ Equity2d ago
In OfficePrincipalNACloud ComputingArtificial IntelligenceSite Reliability EngineerPrincipalPythonReportingAWSDatadogTerraformLinuxUbuntuDockerKubernetesRayDatabricks

Requirements

• Bachelor’s degree or equivalent years of experience and/or certification in a related field • 7+ years working in Site Reliability Engineering or Information Technology • Design and document systems, including writing and reviewing code, to automate away problems within your team’s domain • Intermediate to expert experience deploying and maintaining AWS services such as EC2, ECS, EKS, SQS, Lambda, Dynamo, RDS/Aurora, S3, and IAM • Intermediate to expert experience monitoring services and applications using tools such as CloudWatch Metrics, CloudWatch Logs, and Datadog, including defining and configuring alerts and SLO reports • Intermediate to expert experience maintaining the uptime and scalability of AWS compute services used for Machine Learning and Data Science workloads, specifically EC2 and EKS • Intermediate to expert coding skills in at least one programming language; we work primarily in Python • Intermediate to expert experience using Infrastructure as Code platforms and CI/CD pipelines, specifically Terraform, to manage AWS infrastructure and services • In-depth knowledge & experience with Linux operating systems (Amazon Linux, Ubuntu) on EC2 and Docker / Kubernetes • Experience with leading projects in system design, architecture changes, and technology selection

Responsibilities

• Use independent judgment and discretion to own SLOs, error budgets, and the operational health of Ray clusters running on AWS EKS and Databricks workloads on AWS EC2 across multiple accounts and regions • Maintain the observability of uptime, availability, and scalability of EKS Ray and Databricks workloads using CloudWatch and Datadog, including defining alerting that maps to SLOs • Operate and tune EKS Ray workloads at scale including autoscaling, GPU scheduling, and automated failure recovery • Manage Databricks on AWS including workspace administration, cluster policies, Unity Catalog, job orchestration, and IAM Roles and Policies • Maintain ongoing cost visibility, cost optimization, and capacity planning across EC2 and EKS workloads, including through the use of On Demand Capacity Reservations and Spot lifecycle • Perform ongoing maintenance of the underlying EC2 and EKS infrastructure, including regular security updates and operating system upgrades • Codify everything as infrastructure-as-code using Terraform and CI/CD pipelines, enabling updates through Pull Requests with approval workflows, while also automating maintenance tasks to reduce toil • Lead incident response for Data Science and Machine Learning platform outages, run blameless postmortems, and drive systemic remediation, including participating in an on-call rotation • Complete any additional tasks as they arise

Benefits

• Fair and competitive salary based on skills and experience, and annual performance bonus • Equity may be awarded in the form of Restricted Stock Units (RSUs) • Medical, Dental, Vision and Life Insurance, matching 401k, short-term & long-term disability and parental leave • Unlimited Paid Time Off including vacation, sick days & public holidays • Flexible scheduling and work from home policy depending on role and responsibilities • The base salary range for this position is: $142,000 to $177,600. This range is specifically for Cambridge, MA • Work on a mission with real impact: crashes prevented, injuries avoided, lives protected around the world • Recognized innovator in mobility AI, earning top honors including the TIME Industry Leader in AI, a Gold Edison Award, and the Artificial Intelligence Excellence Award for AI for Social Good. CMT is also Great Place to Work Certified • Be part of the team inventing the future of mobility and road safety • Move fast, own outcomes, do work that matters • High ownership, small teams, and direct access to leadership — no layers between your work and its impact • Unlimited PTO, flexible scheduling, competitive salary, annual performance bonus, RSUs, and full benefits including medical, dental, vision, and 401k match • Summer Fridays provide team members with half days to recharge • Comprehensive wellness, education, and employee assistance programs • Commitment to Diversity and Inclusion:

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

AccelaAccela - Principal Site Reliability Engineer1mo ago
·Remote - Based - US·$160k - $185k/year + Equity
RemoteNAPrincipalInsuranceCloud ComputingSite Reliability EngineerPrincipalBashPythonKubernetesAzureChange Management
Skylo TechnologiesSkylo Technologies - Senior Network Reliability Engineer, Incident Management4d ago
·Remote - Mountain View, California, United States·$125k - $135k/year + Equity
RemoteNASeniorDiagnosticsCloud ComputingSite Reliability EngineerPrincipalBashPythonStaff DevelopmentProduct MarketingReporting
KongKong - Staff Site Reliability Engineer1mo ago
·United States·$150k - $210k/year
In OfficeNAStaffArtificial IntelligenceSoftwareSite Reliability EngineerPrincipalKubernetesPlaneTerraformHelmPostgreSQL
Backblaze External WebsiteBackblaze External Website - Sr. Site Reliability Engineer4mo ago
·Remote - USA·$150k - $200k/year + Equity
RemoteNASeniorCloud ComputingSite Reliability EngineerPrincipalGoPythonDocumentationLinuxPerformance Management
PlaysonPlayson - Principal Site Reliability Engineer (Platform Tribe)5mo ago
·European Union
In OfficeEMEAPrincipalInsuranceCloud ComputingSite Reliability EngineerPrincipalGoGitPythonKubernetesAWS
Extreme NetworksExtreme Networks - AI Principal Machine Learning Engineer (10189)1mo ago
·Seattle, Washington, United States·$170k - $240k/year + Equity
In OfficeNAPrincipalCloud ComputingArtificial IntelligenceMachine Learning EngineerPrincipalDockerKubernetesPythonFastAPIAWS
Menlo SecurityMenlo Security - Principal Platform Infrastructure Engineer (Containers)2mo ago
·Remote - Canada·$103k - $103k/year + Equity
RemoteNAPrincipalCloud ComputingArtificial IntelligencePrincipalInfrastructure EngineerPythonKubernetesTerraformGeminiAWS
SpotifySpotify - Site Reliability Engineer5mo ago
·Remote - New York, NY·$133k - $190k/year
RemoteNACloud ComputingArtificial IntelligenceSite Reliability EngineerAWSGCPTerraformReactPython
PinterestPinterest - Site Reliability Engineer II, tvScientific1mo ago
·San Francisco, California, United States·$114k - $114k/year + Equity
RemoteNAMidCloud ComputingSite Reliability EngineerBashPythonAWSKubernetesTerraform

Browse more by category

Show 165 moreSite Reliability EngineerShow 761 morePrincipalShow 4,672 morePythonShow 6,353 moreReportingShow 2,724 moreAWSShow 166 moreDatadogShow 808 moreTerraformShow 733 moreLinuxShow 68 moreUbuntuShow 783 moreDocker
Privacy·Terms··Contact·FAQ·Wagey on X