• Bachelor's degree or equivalent experience in computer science, engineering, or similar.
• Deep understanding of transformer architectures and their variants.
• Strong programming skills in Python with experience in PyTorch internals.
• Experience with LLM inference systems (vLLM, TensorRT-LLM, SGLang, TGI).
• Ability to read and implement model architectures and inference techniques from research papers.
• Demonstrate the ability to contribute performant and maintainable code and debug in complex ML codebases.
• Deep understanding of KV-cache memory management, prefix caching, and hybrid model serving.
• Familiarity with RL frameworks and algorithms for LLMs.
• Experience with multimodal inference (audio/image/video/text).
• Contributions to open-source ML or system infrastructure projects.
• Implemented core features in vLLM or other inference engine projects.
• Contributed to vLLM integrations (verl, OpenRLHF, Unsloth, LlamaFactory, etc).
• Written widely-shared technical blogs or side projects on vLLM or LLM inference.
• Location: This role is based in Singapore.
• Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is S$200,000 to S$400,000 annually + equity.
• Visa sponsorship: We sponsor visas on a case-by-case basis.
• Benefits: Inferact offers a generous benefits package, including medical, dental, and vision coverage.