wagey.ggwagey.gg
31,365  jobs31,365  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(31,365)/Staff Engineer Role(824)/Inferact (7) - Member of Technical Staff, TPU Performance Engineering
Inferact

Inferact - Member of Technical Staff, TPU Performance Engineering

Singapore$155k - $310k+ Equity1w ago
In OfficeStaffAPACArtificial IntelligenceStaff EngineerJAX

Requirements

• Bachelor's degree or equivalent experience in computer science, engineering, systems, machine learning, or similar. • Hands-on experience building or optimizing TPU workloads using JAX, XLA, Pallas, or related compiler and runtime tooling. • Deep understanding of TPU execution, memory behavior, compilation, and performance constraints for ML workloads. • Experience optimizing ML kernels or inference paths such as attention, GEMM, sampling, KV cache, fused kernels, or backend runtime paths. • Strong performance profiling and benchmarking skills, with the ability to use measurements, compiler artifacts, correctness tests, and reproducible benchmarks to guide optimization work. • Experience with vLLM, SGLang, TensorRT-LLM, XLA-based serving, or other LLM inference systems. • Familiarity with batching, KV cache, decoding, serving tradeoffs, and backend performance constraints in production inference systems. • Experience with compiler technologies such as XLA, MLIR, LLVM, Pallas, or other kernel DSLs, including lowering, fusion, and backend code generation. • Knowledge of quantization methods such as INT8, FP8, mixed precision, or TPU-specific numeric formats, including accuracy and performance tradeoffs. • Contributed to vLLM, JAX/XLA, Pallas, PyTorch/XLA, compiler projects, or other open-source ML infrastructure. • Built TPU benchmarking infrastructure or automated performance regression detection for accelerator workloads. • Worked directly with Google TPU ecosystem stakeholders, accelerator platform teams, or early-access programs to ship backend, compiler, or inference performance improvements. • Location: This role is based in Singapore. • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity. • Visa sponsorship: We sponsor visas on a case-by-case basis. • Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.

Benefits

• SGD 200K – SGD 400K • Offers Equity • Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build. • We're looking for a TPU performance engineer to make vLLM a first-class inference engine on Google TPUs. You'll build and optimize TPU backends, compiler integrations, runtime paths, and benchmarking infrastructure using JAX, XLA, Pallas, and related tooling so vLLM can deliver frontier inference performance on TPU hardware. • You'll work at the boundary of inference systems, kernels, compilers, and hardware architecture, improving production-relevant model serving on TPU with clear correctness, latency, and throughput benchmarks. Your work will help make TPU support in vLLM usable, fast, benchmarked, and maintainable. • Visa sponsorship: We sponsor visas on a case-by-case basis. • Visa sponsorship: • Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

inherentinherent - Member of Technical Staff (Post Training)1mo ago
·London, England, United Kingdom
In OfficeEMEAStaffArtificial IntelligenceStaff EngineerPythonJAX
RunwayRunway - Member of Technical Staff, Research Engineer3mo ago
·United States·$270k - $370k/year
In OfficeNAStaffArtificial IntelligenceResearch EngineerStaff EngineerJAX
WeaveWeave - Staff / Senior ML - GenAI Engineer, Voice & Speech13h ago
·Remote - India
RemoteAPACStaffCloud ComputingArtificial IntelligenceStaff EngineerCoachingLearning & DevelopmentB2BJupyterMLflowTritonPythonDVCDagsterAWSGCPObservableData AnalysisKubernetesGoogle GKE
StacklokStacklok - Staff Forward Deployed Engineer - Kubernetes1w ago
·Remote - Singapore
RemoteAPACStaffArtificial IntelligenceOil & GasStaff EngineerKubernetesHelmPrometheusDatadogGrafanaGoMentoring
GitLabGitLab - Staff Fullstack Engineer1w ago
·Remote - Bangalore, India·Equity
RemoteAPACStaffArtificial IntelligenceSoftwareStaff EngineerReportingRubySnowflakeGoDatabricksTalent AcquisitionJiraZendeskThe Graph
primeintellectprimeintellect - Member of Technical Staff - Inference3w ago
·Hybrid - Asia-Pacific *·$312k+/year
In OfficeAPACStaffCloud ComputingArtificial IntelligenceStaff EngineerCUDAPythonAWSGCPKubernetesTritonRustC++RedisKafkaPrometheusTerraformGrafanaAnsible
multiplier-holdingsmultiplier-holdings - Staff Engineer - Knowledge Intelligence Pod1mo ago
·Singapore·Equity
In OfficeAPACStaffArtificial IntelligenceStaff EngineerClose
OKXOKX - Staff/Senior Staff Engineer, Kubernetes1mo ago
·Singapore
In OfficeAPACStaffCloud ComputingArtificial IntelligenceStaff EngineerKubernetesReportingAWSAlibaba CloudShell
bjakcareerbjakcareer - Technical Lead, Machine Learning3mo ago
·Singapore, Orchard Road
In OfficeAPACStaffArtificial IntelligenceTech LeadMachine Learning EngineerJAX

Browse more by category

Show 824 moreStaff EngineerShow 93 moreJAX
Privacy·Terms··Contact·FAQ·Wagey on X