wagey.ggwagey.gg
31,365  jobs31,365  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(31,365)/AI Engineer Role(768)/featherlessai (8) - AI Researcher — Inference Optimization
featherlessai

featherlessai - AI Researcher — Inference Optimization

Remote - (world) - USA *6mo ago
RemoteNAArtificial IntelligenceHospitalsDiagnosticsAI EngineerPythonTritonONNXCUDAClose

Requirements

• Strong background in machine learning, deep learning, or AI systems. • machine learning, deep learning, or AI systems • Hands-on experience optimizing inference for large-scale models. • large-scale models • Proficiency in Python and modern ML frameworks (e.g., PyTorch). • Python • Experience with inference tooling (e.g., Triton, TensorRT, vLLM, ONNX Runtime). • Ability to design experiments and communicate results clearly. • Experience deploying production inference systems at scale. • production inference systems at scale • Familiarity with distributed and multi-GPU inference. • distributed and multi-GPU inference • Experience contributing to open-source ML or inference frameworks. • open-source ML or inference frameworks • Authorship or co-authorship of peer-reviewed research papers in machine learning, systems, or related fields. • Authorship or co-authorship of peer-reviewed research papers • Experience working close to hardware (CUDA, ROCm, profiling tools). • What Success Looks Like • Measurable gains in latency, throughput, and cost efficiency. • latency, throughput, and cost efficiency • Optimized inference systems running reliably in production. • Research ideas successfully translated into deployable systems. • Clear benchmarks and documentation that inform product decisions. • Relevant Research Areas (Bonus) • Long-context inference optimization • Speculative decoding • KV-cache compression and paging • Efficient decoding strategies • Hardware-aware inference design

Responsibilities

• Research and develop techniques to optimize inference performance for large neural networks. • optimize inference performance • Improve latency, throughput, memory efficiency, and cost per inference. • latency, throughput, memory efficiency, and cost per inference • Design and evaluate model-level optimizations (quantization, pruning, KV-cache optimization, architecture-aware simplifications). • model-level optimizations • Implement systems-level optimizations (dynamic batching, kernel fusion, multi-GPU inference, prefill vs decode optimization). • systems-level optimizations • Benchmark inference workloads across hardware accelerators. • Collaborate with engineering teams to deploy optimized inference pipelines. • deploy optimized inference pipelines • Translate research insights into production-ready improvements. • production-ready improvements

Benefits

• Equity compensation is mentioned as a benefit. Response: EQUITY COMPENSATION AVAILABLE • Paid time off (PTO) options are included among benefits. Response: PTO OFFERED • Remote work options are explicitly mentioned as a benefit. Response: REMOTE WORK OPTIONS AVAILABLE

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

PowerlinePowerline - Lead AI Engineer2w ago
·Palo Alto, California, United States
In OfficeNAStaffArtificial IntelligenceAI EngineerPython
Bedrock TalentBedrock Talent - Lead AI Engineer, Stealth Longevity AI Startup2w ago
·Remote - USA·$180k - $200k/year + Equity
RemoteNAStaffCloud ComputingArtificial IntelligenceAI EngineerPython
curaicurai - Senior Applied AI Engineer2w ago
·Remote - USA *·$175k - $200k/year + Equity
RemoteNASeniorLife SciencesArtificial IntelligenceAI EngineerPythonTeam LeadershipMentoringClose
FathomFathom - AI Engineer3mo ago
·Remote - San Francisco, California, United States
RemoteNAArtificial IntelligenceAI EngineerPythonCUDASwiftRayLoom
ZapierZapier - Sr. Applied AI Engineer3mo ago
·Remote - San Francisco, California, United States·$232k - $348k/year
RemoteNASeniorCloud ComputingArtificial IntelligenceAI EngineerTypeScriptPythonDocumentationCloseLater
Farsight AIFarsight AI - AI Engineer6d ago
·Remote - New York City, New York, United States·$200k - $300k/year + Equity
RemoteNAMidArtificial IntelligenceAI EngineerPythonDocumentationRails
EnCharge AIEnCharge AI - AI Compiler Engineer1mo ago
·Remote - USA·$190k - $255k/year
RemoteNAMidArtificial IntelligenceAI EngineerC++PythonMentoring
1mind1mind - AI System Engineer2mo ago
·Remote - USA·$175k - $250k/year
RemoteNAArtificial IntelligenceAI EngineerPythonKubernetesTypeScript
TRM LabsTRM Labs - AI Agent Engineer4mo ago
·San Fracisco , California , United States - Hybrid·$200k - $275k/year + Equity
In OfficeNAArtificial IntelligenceAI EngineerPythonVectorObservable

Browse more by category

Show 768 moreAI EngineerShow 5,064 morePythonShow 54 moreTritonShow 21 moreONNXShow 70 moreCUDAShow 2,649 moreClose
Privacy·Terms··Contact·FAQ·Wagey on X