wagey.ggwagey.gg
30,434  jobs30,434  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(30,434)/Machine Learning Engineer Role(358)/egen (12) - Lead Machine Learning Engineer, Inference & Performance
egen

egen - Lead Machine Learning Engineer, Inference & Performance

Remote - Europe *4w ago
RemoteStaffEMEAArtificial IntelligenceNonprofitMachine Learning EngineerClient ConsultingPerformance ManagementSQLGoogle GKEPython

Requirements

• Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field • 5+ years of experience in ML/AI engineering, with a meaningful portion focused on performance, infrastructure, or systems • Proven track record of deploying and optimizing models in a production environment • Demonstrated experience profiling and improving GPU utilization for training and/or inference • Experience with Classic Machine Learning (neural nets, training, tuning) is a strong plus • Knowledge of Data Engineering and SQL • ## Personal Attributes: • Ownership: You take pride in your work and see optimizations through from profile to production • Curiosity: Hardware and serving frameworks change fast; you are a lifelong learner who stays ahead of the curve • Rigor: You measure before you optimize and let data, not intuition, guide where you spend effort • Consultative Spirit: You enjoy interacting with clients and can translate technical complexity into business value • Ethics: You prioritize responsible AI development and data privacy

Responsibilities

• Optimize Inference: Build and tune production LLM serving with vLLM and SGLang—maximizing throughput and minimizing latency through batching, paged attention, quantization, and KV-cache strategies • Profile & Accelerate Training: Instrument and profile training runs to find bottlenecks, then resolve them with the right attention implementations (e.g. FlashAttention) tuned to the underlying hardware (H200, GB200) • Engineer for the Hardware: Apply a working understanding of GPU architecture and attention internals to choose the right approach per accelerator, rather than relying on defaults • Serve at Scale: Deploy and operate multiple models within shared GPU clusters on GKE, with autoscaling, efficient bin-packing, and graceful handling of mixed workloads • Drive Efficiency: Own GPU utilization as a first-class metric—measure it, improve throughput-per-dollar, and continuously raise the ceiling on what our fleet can deliver • Collaborate & Consult: Work directly with clients to understand performance, latency, and cost requirements, and translate them into pragmatic serving and training architectures • ## Your Technical Toolkit: • Core Languages: Mastery of Python and shell scripting; comfort reading and reasoning about lower-level (CUDA-adjacent) performance code is a strong plus • Inference Frameworks: Hands-on experience with vLLM, SGLash, or comparable high-performance serving stacks • GPU & Model Internals: Solid grasp of GPU architecture, the fundamentals of LLM inference, and the attention mechanism—including where the bottlenecks live and how FlashAttention and similar techniques address them across hardware generations (H200, GB200) • Profiling: Fluency with profiling tools to diagnose training and inference bottlenecks (compute-bound vs. memory-bound, kernel-level analysis) • Infrastructure: Strong Kubernetes (GKE) experience—deploying and autoscaling multiple models on shared GPU clusters on Google Cloud • Mindset: A strong software engineering foundation—you write clean, maintainable code, measure before optimizing, and understand the full SDLC

Benefits

• This role is eligible for our competitive salary and comprehensive benefits package to support your well-being: • Comprehensive Health Insurance • Paid Leave (Vacation/PTO) • Paid Holidays • Parental Leave • Bereavement Leave • 401 (k) Employer Match • Employee Referral Bonuses • Check out our complete list of benefits here - >https://egen.ai/people/#benefits • Important: All roles are subject to standard hiring verification practices, which may include background checks, employment verification, and other relevant checks. • EEO and Accommodations:

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

happyhotelhappyhotel - Staff Machine Learning Engineer - Pricing & Revenue (m/f/d)2mo ago
·Offenburg, Baden-Württemberg, Germany
In OfficeEMEAStaffGenomicsArtificial IntelligenceMachine Learning EngineerdbtSnowflakeMetabaseSQLPython
PearlPearl - Senior/Lead Machine Learning Engineer4mo ago
·Remote - Ukraine
RemoteEMEAStaffCloud ComputingArtificial IntelligenceMachine Learning EngineerSQLPythonData QualityCursor.NET
Planday from XeroPlanday from Xero - Machine Learning Engineer (Agentic AI)2d ago
·Remote - Denmark, United Kingdom
RemoteEMEAMidArtificial IntelligenceNonprofitMachine Learning EngineerPythonPandasSnowflakedbtDatabricksSQLObservable
constructorconstructor - Senior Machine Learning Engineer: ML Recall1mo ago
·Remote - United Kingdom·$80k - $120k/year + Equity
RemoteEMEASeniorCloud ComputingArtificial IntelligenceMachine Learning EngineerPythonSQLMySQLPySparkAWS
preplypreply - Staff Machine Learning Ops Engineer1mo ago
·London, Greater London, United Kingdom·Equity
In OfficeEMEAStaffCloud ComputingArtificial IntelligenceMachine Learning EngineerVectorPerformance ManagementGCPKubernetesAWS
RedditReddit - Staff Machine Learning Engineer, ML Efficiency1mo ago
·Remote - UK
RemoteEMEAStaffCloud ComputingArtificial IntelligenceMachine Learning EngineerPythonTraining DevelopmentRustC++Java
facultyfaculty - Copy of Lead Machine Learning Engineer1mo ago
·London, United Kingdom, Hybrid
In OfficeEMEAStaffCloud ComputingArtificial IntelligenceMachine Learning EngineerPythonscikit-learnTeam ManagementCoachingAWS
Fireworks AIFireworks AI - Applied Machine Learning Engineer, EMEA1w ago
·London, UK
In OfficeEMEASeniorArtificial IntelligenceMachine Learning EngineerPythonCustomer Success
deliveroodeliveroo - Senior Machine Learning Engineer - Personalisation3mo ago
·London, City of, United Kingdom, Hybrid
In OfficeEMEASeniorArtificial IntelligenceGroceryMachine Learning EngineerPythonStakeholder Management

Browse more by category

Show 358 moreMachine Learning EngineerShow 195 moreClient ConsultingShow 1,075 morePerformance ManagementShow 2,582 moreSQLShow 25 moreGoogle GKEShow 4,717 morePython
Privacy·Terms··Contact·FAQ·Wagey on X