wagey.ggwagey.gg
38,923  jobs38,923  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(38,923)/Infrastructure Engineer Role(236)/spaitial (6) - Machine Learning Systems & Infrastructure Engineer
spaitial

spaitial - Machine Learning Systems & Infrastructure Engineer

London, United Kingdom1mo ago
In OfficeMidEMEACloud ComputingArtificial IntelligenceInfrastructure EngineerMachine Learning EngineerDockerKubernetesTerraformPythonCUDA

Requirements

• 3+ years writing production-quality Python in a large, multi-author codebase, with strong SWE fundamentals (ML systems experience strongly preferred). • Hands-on with modern ML training stacks (PyTorch; DDP/FSDP or comparable); have personally debugged distributed jobs across many GPUs and nodes. • Have shipped non-trivial end-to-end data pipelines at scale — ingestion, transformation, validation, versioning, republish — ideally including real-world sources with rate limits, auth, or undocumented APIs. • Hands-on GPU compute and performance debugging (CUDA/NCCL, GPU utilization, networking bottlenecks, profiling). • Working knowledge of cloud environments (AWS, GCP, or Azure), including object storage, IAM, and cost awareness. • Proficient with containers (Docker, Kubernetes) and comfortable reading and writing IaC (Terraform) for the surfaces you ship. • Strong working knowledge of how to store and query large datasets at scale: SQL fundamentals; relational (e.g., Postgres), analytical (e.g., BigQuery, Snowflake), and embedded (e.g., SQLite) stores; and object storage with caching layers. Familiarity with ML workflow orchestration and experiment tracking (e.g., Kubeflow Pipelines, MLflow). • Experience with monitoring and observability tooling (e.g., Prometheus/Grafana, OpenTelemetry) and CI/CD for infra and ML workflows (e.g., GitHub Actions).

Responsibilities

• Own and evolve the ML systems that enable training, evaluation, and serving of large foundation models — trainer, dataset loaders, checkpointing, and experiment orchestration code. • Distributed training enablement: Improve high-throughput training stacks (e.g., PyTorch DDP/FSDP, NCCL) for performance, stability, and reproducibility, including preemption-safe and sharded checkpointing. • Data systems and pipelines: Build end-to-end Python pipelines that turn third-party capture sources into clean, versioned training datasets — including scraping (e.g., Playwright) and preprocessing — and optimize the underlying storage at petabyte scale (object storage, fuse mounts, caching layers, shared filesystems, and relational / analytical / embedded metadata stores). • ML workflow orchestration and serving: Operate the systems researchers use to launch experiments, data jobs, and production endpoints — workflow engines (e.g., Kubeflow Pipelines, Airflow), GPU schedulers (e.g., Volcano, Slurm), experiment trackers (e.g., MLflow, Weights & Biases), and managed-inference platforms (e.g., Modal, Triton) — and maintain a launcher SDK for one-command runs. • Containerization and packaging: Ship workloads with Docker and Kubernetes; maintain IaC (Terraform) for the surfaces you own and CI/CD pipelines, including self-hosted GPU runners. • Observability and reliability: Monitoring, logging, and alerting for job performance, data-pipeline health, and cost (e.g., Prometheus/Grafana, OpenTelemetry); define SLOs and incident response for the systems you own. • Security and access: Manage secrets, IAM, and network boundaries (e.g., Tailscale, cloud VPC) for the systems you own. • Collaboration: Partner with ML researchers, engineers, and the platform team to unblock training and data work and improve developer experience.

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

GraphcoreGraphcore - Software Infrastructure Kubernetes Engineer4mo ago
·Bristol, UK; Cambridge, UK; Gdańsk, Pomeranian Voivodeship, Poland
In OfficeEMEAArtificial IntelligenceCloud ComputingInfrastructure EngineerC++PythonKubernetesTerraformGo
encordencord - Machine Learning Engineer1w ago
·London·Equity
In OfficeEMEAMidCloud ComputingArtificial IntelligenceMachine Learning EngineerPythonKerasRESTGCPAWSCUDAKubernetesGraphQLFull Stack
coram-aicoram-ai - Software Engineer - Infrastructure4mo ago
·London, United Kingdom - Hybrid
In OfficeEMEAMidCloud ComputingInternet of ThingsSoftware EngineerInfrastructure EngineerAWSTerraformDockerKubernetesPython
trainlinetrainline - Machine Learning Engineer1mo ago
·London, Greater London, United Kingdom - Hybrid·£60k - £75k/year
In OfficeEMEAInsuranceArtificial IntelligenceMachine Learning EngineerDockerTerraform
FractileFractile - Silicon Infrastructure Engineer2w ago
·London - Hybrid·Equity
In OfficeEMEAArtificial IntelligenceDeveloper ToolsInfrastructure EngineerDockerKubernetesTerraformAnsiblePrometheus
TelnyxTelnyx - Infrastructure Engineer (Core)1mo ago
·Remote - Dublin, Ireland; Ho Chi Minh City, Vietnam; Bangalore, India; Warsaw, Poland; Amsterdam, Netherlands
RemoteEMEAMidSoftwareInfrastructure EngineerPythonAnsiblePuppetTerraformKubernetes
GraphcoreGraphcore - 2026 Graduate IT Infrastructure Engineer2d ago
·Bristol, UK
In OfficeEMEAJuniorCloud ComputingArtificial IntelligenceInfrastructure EngineerLinuxWindows ServerDockerKubernetesAWSGCPBashAzurePython
Insider OneInsider One - Senior Machine Learning Engineer (Agentic AI)1mo ago
·Remote - Istanbul, Turkiye·Equity
RemoteEMEASeniorCloud ComputingArtificial IntelligenceSoftwareMachine Learning EngineerPythonSQLAWSKubernetes
facultyfaculty - Copy of Lead Machine Learning Engineer3w ago
·London
In OfficeEMEAStaffCloud ComputingArtificial IntelligenceMachine Learning EngineerTeam ManagementCoachingPythonscikit-learnAWSAzureGCPDockerKubernetesFull StackMentoring

Browse more by category

Show 236 moreInfrastructure EngineerShow 466 moreMachine Learning EngineerShow 1,082 moreDockerShow 1,913 moreKubernetesShow 1,182 moreTerraformShow 6,296 morePythonShow 58 moreCUDA
Privacy·Terms··Contact·FAQ·Wagey on X