• Design, develop, and operate large-scale, high-performance infrastructure that powers Confluent Cloud.
• Build foundational software to improve reliability, scalability, and efficiency across cloud environments.
• Work on distributed systems challenges such as consensus algorithms, failover strategies, and resource allocation.
• Collaborate with teams across Confluent to optimize and enhance infrastructure for real-time data streaming use cases.
• Troubleshoot and improve system reliability, observability, and performance across multiple cloud providers (AWS, Azure, GCP).
• 2-5 years of industry experience designing, building, and supporting backend systems in production.
• Strong fundamentals in distributed systems, cloud infrastructure, and networking.
• Experience in building and operating large-scale, high-availability systems.
• Good understanding of cloud platforms (AWS, Azure, or GCP) and their services.
• Proficiency in Java, Scala, C++, Go, or other statically typed languages.
• A self-starter with strong problem-solving skills and the ability to work in a fast-paced environment.
• BS, MS, or PhD in computer science or a related field, or equivalent work experience.
• WHAT GIVES YOU AN EDGE:
• Exposure to model serving, LLM/agent infrastructure, or streaming data systems.
• Note - You don't need a background in ML research or model training — this role is about building and operating the platform that serves AI reliably at scale, not inventing the models.
• READY TO BUILD WHAT'S NEXT? LET’S GET IN MOTION.
• COME AS YOU ARE
• Belonging isn’t a perk here. It’s the baseline. We work across time zones and backgrounds, knowing the best ideas come from different perspectives. And we make space for everyone to lead, grow, and challenge what’s possible.