Nebius - Find your role: Open positions
Requirements
• 5+ years of experience in ML/AI systems, with at least 2 years focused on LLMs and generative AI • at least 2 years focused on LLMs and generative AI • Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches • LLM ecosystem • model architectures • fine-tuning approaches • Running LLMs in production: deploying and operating inference workloads • Running LLMs in production • LLM fine-tuning, including supervised fine-tuning (SFT/LoRA) and data preparation/curation; experience with RL-based fine-tuning is a strong plus • LLM fine-tuning • supervised fine-tuning • RL-based fine-tuning • strong plus • LLM evaluation: building task-specific benchmarks and offline/online eval pipelines, including LLM-as-a-judge setups • LLM evaluation • benchmarks • offline/online eval pipelines • LLM-as-a-judge • Inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM, Transformers) • Inference frameworks • Deploying LLM-powered applications using APIs from OpenAI, Anthropic, or open-source models • Strong Python programming skills • Python programming • Excellent communication skills, with the ability to clearly explain technical concepts to diverse audiences • explain technical concepts • Must be fluent in Mandarin Chinese • Mandarin Chinese • Experience with inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM) • Work with multimodal AI models (e.g., vision-language, speech) • Proficiency with DevOps tools (Docker, Kubernetes) • Contributions to open-source ML/AI projects • Preferred technical stack: • Programming Languages: Python • Programming Languages: • ML Frameworks and Libraries: vLLM, TensorRT-LLM, SGLang, Transformers, OpenAI/Anthropic SDKs • ML Frameworks and Libraries • MLOps and DevOps tools: Kubernetes (K8s), Docker, Git • MLOps and DevOps tools • Cloud Platforms: AWS (SageMaker, Bedrock), GCP (Vertex AI), Azure (Azure ML) • Cloud Platforms
Responsibilities
• Optimize LLM inference across various modalities to drive business value and support customer goals • Provide support in supervised and reinforcement learning fine-tuning to maximize model quality for the customers • Design and implement LLM-based solutions using Nebius Token Factory’s inference services • Build production-ready applications leveraging our serverless LLM APIs, including multimodal models (text, vision, audio) and domain-specific models • Provide technical expertise in prompt engineering, RAG architectures and model selection • Collaborate with product and engineering teams to surface customer feedback and shape the platform roadmap • Guide customers in scaling from POC to production with a focus on performance, reliability, and cost efficiency
Benefits
• Competitive compensation • Career growth and learning opportunities • Flexibility and ownership • Collaborative and innovative culture • Opportunity to work on impactful AI projects • International environment and talented teams • What's it like to work at Nebius: • Fast moving - Bold thinking - Constant growth - Meaningful impact - Trust and real ownership - Opportunity to shape the future of AI • Equal Opportunity Statement:
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT