• Bachelor's degree or equivalent experience in computer science, engineering, or similar.
• Strong experience with Kubernetes and container orchestration at scale.
• Experience designing and implementing custom Kubernetes operators.
• Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc).
• Experience managing GPU clusters and debugging hardware issues.
• Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.
• Experience with ML-specific orchestration tools (Ray, Slurm).
• Knowledge of GPU scheduling, multi-tenancy, and resource optimization.
• Familiarity with vLLM deployment patterns and configuration.
• Track record of improving operational reliability for ML systems.
• Experience deploying inference systems on large-scale GPU (1,000+) clusters.
• Location: This role is based in Singapore.
• Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.
• Visa sponsorship: We sponsor visas on a case-by-case basis.
• Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.