patsnap - Site Reliability Engineering (SRE) Leader
Responsibilities
• Build, lead and develop the UK SRE team, establishing operational standards, best practices, and reliability goals. • Ensure the high availability, stability, security, and performance of business-critical platforms and services. • Define and drive the operational strategy for our global SaaS platform, ensuring exceptional reliability, availability and performance. • Lead major incident management, acting as the senior escalation point during critical production events. • Establish and monitor reliability metrics, including SLIs, SLOs and operational KPIs. • Drive automation across infrastructure, deployments, monitoring and operational workflows to improve efficiency and reduce manual effort. • Champion the adoption of AI-powered operations, leveraging modern AI technologies to enhance engineering productivity and operational excellence. • Partner with Engineering, Product, Security and Infrastructure teams to improve platform architecture, scalability and operational readiness. • Lead disaster recovery planning, operational resilience initiatives and risk management across the platform. • Continuously evaluate emerging cloud, AI and platform technologies to keep PatSnap at the forefront of engineering excellence. • Stay current with emerging cloud, AI, and SRE technologies, driving continuous improvement across the organization. • ## What We'd Love From You: • Bachelor’s degree in Computer Science or a related field, with at least 8 years of experience in DevOps, SRE, or infrastructure operations. • Proven experience leading technical teams and managing production environments at scale. • Strong expertise in cloud platforms (AWS preferred), Kubernetes, Docker, CI/CD pipelines, Infrastructure as Code, and observability platforms. • Deep understanding of distributed systems, high-availability architectures, and large-scale SaaS environments. • Experience driving automation and operational excellence initiatives. • Hands-on experience using AI tools such as ChatGPT, Claude, GitHub Copilot, Codex, or similar technologies to improve engineering productivity. • Strong problem-solving, leadership, communication, and stakeholder management skills. • Fluent in English; Mandarin is highly desirable to facilitate collaboration with teams across multiple regions.
Benefits
• Build technology that powers global innovation. Join a company trusted by over 15,000 organisations worldwide, helping some of the world's most innovative businesses solve complex IP and R&D challenges with AI. • Lead AI-powered engineering. Shape the future of our Site Reliability Engineering function by driving automation, resilience and AI-native operational practices across our global platform. • Own a business-critical function. Lead a talented SRE team, influence engineering strategy and make a direct impact on the performance and reliability of products used around the world. • Work with modern technologies. Solve complex challenges using AWS, Kubernetes, cloud-native architecture and the latest AI-powered engineering tools. • Collaborate globally. Partner with experienced engineers and technology leaders across the UK, Singapore and China in a collaborative, high-performing engineering culture.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT