OKX - Data Operations Engineer, X-Layer
Responsibilities
• Operate and support core data platform components, including Alibaba Cloud DataWorks, MaxCompute / ODPS, Hologres, VVP / Flink, and AWS-based platforms such as Databricks and StarRocks. • Build and maintain monitoring, alerting, and SLA / SLO metrics for data platform services. • Respond to production incidents, troubleshoot issues, participate in post-incident reviews, and drive long-term fixes. • Analyze compute, storage, job, and cluster resource usage to improve performance and optimize cloud costs. • Develop or integrate automation scripts and operational tools to improve health checks, releases, scaling, and troubleshooting efficiency. • Collaborate with data engineering, data warehouse, BI, platform teams, and cloud vendors to continuously improve platform reliability and operational efficiency. • What We Look For In You • 3+ years of experience in data platform operations, big data engineering, platform engineering, SRE, or cloud infrastructure operations. • Experience with either AWS or Alibaba Cloud • Hands-on experience with at least one big data or cloud data platform, such as MaxCompute / ODPS, Hologres, Databricks, StarRocks, Flink, Spark, Hive, Presto, or Trino. • Strong SQL skills, with the ability to troubleshoot data issues, job failures, and performance problems. • Solid Linux fundamentals and scripting experience with Shell, Python, or similar languages. • Good understanding of monitoring, alerting, capacity, access control, and resource management. • Strong incident response and troubleshooting skills, with the ability to drive recovery under pressure. • Good communication and collaboration skills in a cross-functional environment. • Nice-To-Haves • Nice-To-Haves • Experience with Alibaba Cloud DataWorks, MaxCompute / ODPS, Hologres, or VVP / Flink. • Experience with AWS, Databricks, EMR, Glue, Athena, Redshift, or similar cloud data platforms. • Experience operating or tuning OLAP engines such as StarRocks, ClickHouse, Doris, Trino, or Presto. • Experience with monitoring tools such as Prometheus, Grafana, CloudWatch, Datadog, ELK, or OpenSearch. • Experience with Kubernetes, Terraform, Ansible, GitLab CI/CD, Jenkins, or Infrastructure as Code. • Experience with multi-cloud operations, cloud cost optimization, resource governance, or capacity planning.
Benefits
• L&D programs and education subsidy for employees' growth and development • Various team building programs and company events • Wellness and meal allowances • Comprehensive healthcare schemes for employees and dependants • More that we love to tell you along the process! • All official OKX vacancies are published on this website. While roles may appear on selected third-party platforms from time to time, information on other sites may be inaccurate or outdated. If in doubt, please apply directly through our official careers website. • If in doubt, please apply directly through our official careers website. • Information collected and processed as part of the recruitment process of any job application you choose to submit is subject to OKX's Candidate Privacy Notice.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT