wagey.ggwagey.gg
31,365  jobs31,365  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(31,365)/Engineering Manager Role(574)/Spotify (76) - Machine Learning Engineering Manager - LLM Serving & Infrastructure
Spotify

Spotify - Machine Learning Engineering Manager - LLM Serving & Infrastructure

New York, NY / Boston, MA / United States of America (Home Mix)$176k - $252k5mo ago
RemoteStaffNAArtificial IntelligenceData AnalyticsEngineering ManagerMachine Learning EngineerMLOpsMLflow

Requirements

• Lead a high-performing engineering team to develop, build, and deploy a high-scale, low-latency LLM Serving Infrastructure. • Drive the implementation of a unified serving layer to support multiple LLM models and inference types (batch, offline eval flows and real-time/streaming). • Lead all aspects of the development of the Model Registry for deploying, versioning, and running LLMs across production environments. • Ensure successful integration with the core Personalization and Recommendation systems to deliver LLM-powered features. • Define and champion standardized technical interfaces and protocols for efficient model deployment and scaling. • Establish and monitor the serving infrastructure's performance, cost, and reliability, including load balancing, autoscaling, and failure recovery. • Collaborate closely with data science, machine learning research, and feature teams (Autoplay, Home, Search, etc.) to drive the active adoption of the serving infrastructure. • Scale up the serving architecture to handle hundreds of millions of users and high-volume inference requests for internal domain-specific LLMs. • Drive Latency and Cost Optimization: partner with SRE and ML teams to implement techniques like quantization, pruning, and efficient batching to minimize serving latency and cloud compute costs. • Develop Observability and Monitoring: build dashboards and alerting for service health, tracing, A/B test traffic, and latency trends to ensure consistency to defined SLAs. • Contribute to Core LPM Serving: focus on the technical strategy for deploying and maintaining the core Large Personalization Model (LPM). • 5+ years of experience in software or machine learning engineering, with 2+ years of experience managing an engineering team. • Hands-on with ML Engineering: you have deep expertise in building, scaling, and governing high-quality ML systems and datasets, including defining data schemas, handling data lineage, and implementing data validation pipelines (e.g., HuggingFace datasets library or similar internal systems). • Deep technical background in building and operating large-scale, high-velocity Machine Learning/MLOps infrastructure, ideally for personalization, recommendation, or Large Language Models (LLMs). • Proven track record to drive complex projects involving multiple partners and federated contribution models ("one source of truth, many contributors"). • Expertise in designing robust, loosely coupled systems with clean APIs and clear separation of concerns (e.g., distinguishing between fast dev-time tools and rigorous production-like systems). • Experience integrating evaluation and testing into continuous integration/continuous deployment (CI/CD) pipelines to enable rapid 'fork-evaluate-merge' developer workflows. • Solid understanding of experiment tracking and results visualization platforms (e.g., MLFlow, custom UIs). • A pragmatic leader who can balance the need for speed with progressive rigor and production fidelity. • This role is based in New York or Boston. • We offer you the flexibility to work where you work best! There will be some in person meetings, but still allows for flexibility to work from home.

Responsibilities

• Lead a high-performing engineering team to develop, build, and deploy a high-scale, low-latency LLM Serving Infrastructure. • Drive the implementation of a unified serving layer to support multiple LLM models and inference types (batch, offline eval flows and real-time/streaming). • Lead all aspects of the development of the Model Registry for deploying, versioning, and running LLMs across production environments. • Ensure successful integration with the core Personalization and Recommendation systems to deliver LLM-powered features. • Define and champion standardized technical interfaces and protocols for efficient model deployment and scaling. • Establish and monitor the serving infrastructure's performance, cost, and reliability, including load balancing, autoscaling, and failure recovery. • Collaborate closely with data science, machine learning research, and feature teams (Autoplay, Home, Search, etc.) to drive the active adoption of the serving infrastructure. • Scale up the serving architecture to handle hundreds of millions of users and high-volume inference requests for internal domain-specific LLMs. • Drive Latency and Cost Optimization: partner with SRE and ML teams to implement techniques like quantization, pruning, and efficient batching to minimize serving latency and cloud compute costs. • Develop Observability and Monitoring: build dashboards and alerting for service health, tracing, A/B test traffic, and latency trends to ensure consistency to defined SLAs. • Contribute to Core LPM Serving: focus on the technical strategy for deploying and maintaining the core Large Personalization Model (LPM).

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

SandboxAQSandboxAQ - Staff Machine Learning Engineer, AI Generation Engine2mo ago
·Canada·$106k - $187k/year
In OfficeNAStaffCloud ComputingArtificial IntelligenceMachine Learning EngineerMLOpsWeights & BiasesMLflowPythonPandas
RedditReddit - Senior Staff Machine Learning Engineer, ML Understanding3mo ago
·Remote - USA·$266k - $372k/year + Equity
RemoteNAStaffArtificial IntelligenceNonprofitMachine Learning EngineerMLOpsMentoring
GreenhouseGreenhouse - Sr Machine Learning Engineer, AI Research2mo ago
·Remote - USA·$185k - $185k/year
RemoteNASeniorArtificial IntelligenceLife InsuranceInsuranceMachine Learning EngineerPythonMLOpsWeights & BiasesMLflowKubeflow
SpotifySpotify - Machine Learning Engineering Manager, Music Mission5mo ago
·Remote - New York, NY·$184k - $263k/year
RemoteNAStaffArtificial IntelligenceData AnalyticsEngineering ManagerMachine Learning EngineerTeam ManagementTeam LeadershipPerformance ManagementCoachingMentoring
PhiloPhilo - Sr. Machine Learning Engineer (Recommendation Systems)4mo ago
·Remote - San Francisco, CA or remote within the U.S.·$175k - $235k/year + Equity
RemoteNASeniorArtificial IntelligenceData AnalyticsMachine Learning EngineerPythonDocumentationMLOpsStrategic Planning
AIFTAIFT - Machine Learning Engineer Lead, Vulcan5mo ago
·Taipei, Hong Kong, Singapore, Japan, Abu Dhabi
In OfficeAPACStaffArtificial IntelligenceMachine Learning EngineerDVCMLOpsMLflowKubeflowAirflow
AffirmAffirm - Engineering Manager, Machine Learning Platform1w ago
·Remote - Canada·$181k - $241k/year + Equity
RemoteNAStaffArtificial IntelligenceEducationEngineering ManagerMachine Learning EngineerCoachingData QualityCross-functional CollaborationBase
RedditReddit - Machine Learning Engineering Manager1mo ago
·Remote - USA·$230k - $322k/year + Equity
RemoteNAStaffArtificial IntelligenceSoftwareEngineering ManagerMachine Learning Engineer
RedditReddit - Ads Conversion Modeling, Machine Learning Engineering Manager3mo ago
·Remote - USA·$230k - $230k/year + Equity
RemoteNAStaffArtificial IntelligenceSoftwareEngineering ManagerMachine Learning Engineer

Browse more by category

Show 574 moreEngineering ManagerShow 375 moreMachine Learning EngineerShow 251 moreMLOpsShow 89 moreMLflow
Privacy·Terms··Contact·FAQ·Wagey on X