• Ensuring storage is reliable, predictable, and not a bottleneck for any critical workloads across the company
• Owning performance and stability of storage systems, and continuously improving them as data volumes and workloads grow
• Designing and evolving data placement, resiliency, and lifecycle strategies to balance performance, cost, and reliability
• Ensuring the platform behaves predictably during failures, maintenance, and scaling events
• Improving how storage integrates with compute environments (GPU/HPC, Kubernetes, data pipelines)
• Driving faster and more reliable incident detection, resolution, and prevention
• Improving capacity planning to avoid emergency scaling and unexpected degradation
• Continuously improving tooling, automation, and operational practices to make the platform easier to operate and scale
• What We Look For In You:
• Experience operating large-scale storage systems in production (distributed or vendor-based)
• Strong understanding of Linux, storage performance, and system behavior under load
• Ability to troubleshoot complex issues and drive them to resolution
• Practical approach to automation and system reliability
• Ownership mindset — ability to take responsibility for critical systems and improve them over time
• Nice-to-have:
• Experience working with high-performance or distributed storage systems
• Understanding of networking in high-throughput environments
• Experience in environments with high reliability and performance requirements (finance, HFT, etc.)