• Bachelor's degree or equivalent experience in computer science, engineering, or similar.
• Deep experience writing CUDA kernels or equivalent (CuTeDSL, Triton, TileLang, Pallas).
• Strong understanding of GPU architecture: memory hierarchy, warp scheduling, tiling, tensor cores.
• Proficiency in C++ and Python with demonstrated ability to write high-performance code.
• Experience with profiling tools (Nsight, rocprof) and performance optimization methodologies.
• Obsession with benchmarks and squeezing every percentage point of speedup.
• Experience with ML-specific kernel optimization (FlashAttention, fused kernels).
• Knowledge of quantization techniques (INT8, FP8, mixed-precision).
• Familiarity with multiple accelerator platforms (NVIDIA, AMD, TPU, Intel).
• Experience with compiler technologies (LLVM, MLIR, XLA).
• Kernel-related contributions to vLLM or other inference engine projects.
• Contributions to open-source GPU, ML systems, or compiler optimization projects
• Written deep technical blogs on GPU optimization.
• Location: This role is based in Singapore.
• Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.
• Visa sponsorship: We sponsor visas on a case-by-case basis.
• Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.