OKX - Staff/Senior Staff Web3 Big Data Engineer
Requirements
• Big Data Engineering • Bachelor's degree or above in Computer Science, Software Engineering, or related field; 7+ years of big data development experience. • Proficiency with Hadoop, Spark, and Flink, with experience building batch/real-time data warehouses. • Familiarity with Hive, Kafka, HBase, ClickHouse, and Doris, with large-scale cluster tuning experience. • Strong big data development skills in Java/Scala/Python, with solid SQL tuning ability. • Experience with data governance, data quality monitoring, or metadata management is a plus. • AI Capabilities • Understanding of LLM fundamentals and application patterns, with hands-on experience in prompt engineering and RAG. • Experience with vector databases (e.g. Milvus, Pinecone, Weaviate, pgvector) or Embedding data processing. • Familiarity with emerging AI application architectures such as AI Agents or MCP (Model Context Protocol) is a plus. • Experience connecting big data platforms with AI/ML training and inference pipelines (e.g. feature platforms, real-time feature serving). • Experience using LLMs to accelerate data engineering (automated data quality checks, intelligent ETL generation, Text2SQL) is a plus. • Basic ML/deep learning knowledge and ability to collaborate effectively with algorithm teams is a plus. • Web3 Industry Knowledge • Understanding of blockchain fundamentals (account model/UTXO, consensus mechanisms, gas mechanics). • Understanding of on-chain data structures for major public chains (Ethereum, BSC, Solana, Bitcoin, etc.) — blocks, transactions, event logs, token transfers (ERC-20/ERC-721). • Experience with on-chain data collection/parsing is a plus — e.g. building or integrating with The Graph, Dune, Etherscan API, or node RPCs. • Understanding of DeFi, NFT, or cross-chain bridge business models and their data characteristics is a plus. • Data engineering experience at an exchange, wallet, DeFi protocol, or on-chain analytics company is a plus. • Fluent English reading/writing; able to independently read English technical documentation (e.g. Ethereum Yellow Paper, RFCs/EIPs, AI papers) and write technical proposals and weekly reports in English. • Able to communicate in English with overseas teams, communities, or partners, with strong cross-cultural communication skills. • Clear logical thinking, strong problem decomposition and troubleshooting ability, and a strong appetite for learning new AI technology. • Strong team player, able to deliver consistently in a fast-paced, high-uncertainty Web3 environment. • Nice-to-Haves • Nice-to-Haves • Basic smart contract knowledge (Solidity); able to read contract code logic. • Contribution experience in open-source big data / blockchain data / AI infrastructure projects. • Relevant technical certifications (e.g. AWS/GCP big data or AI certifications). • Experience using AI coding agents (e.g. Claude Code) to improve data engineering efficiency.
Responsibilities
• Platform Architecture: Own the architecture, development, and optimization of our big data platform, supporting large-scale collection, processing, and analysis of on-chain data, trading behavior, and user profiles. • Platform Architecture: • On-Chain Data Pipelines: Design and build real-time/batch data pipelines for on-chain data — blockchain transactions, smart contract events, wallet address behavior. • On-Chain Data Pipelines: • Data Warehouse & Lake: Build and maintain the data warehouse/lake, define data-layering standards, and ensure data quality, consistency, and timeliness. • Data Warehouse & Lake: • AI-Driven Data Applications: Combine big data with LLM capabilities to build intelligent applications — on-chain anomaly/fraud detection, smart risk models, user behavior prediction, and natural-language data querying (Text2SQL). • AI-Driven Data Applications: • AI Infrastructure: Build data infrastructure for AI use cases — vector databases, feature platforms, and Embedding pipelines supporting RAG retrieval and Agent data supply. • AI Infrastructure: • Cross-Team Collaboration: Partner with Algorithm/AI teams on large-scale data processing and pipelines for model training data and feature engineering; partner with Product, Risk, and Growth teams on data needs for trading analytics, anti-fraud, growth, and operations. • Cross-Team Collaboration: • Performance & Reliability: Optimize performance and resource efficiency of big data jobs, ensuring stability and SLAs on core pipelines. • Performance & Reliability: • Technical Direction: Track industry developments in big data, Web3 data infrastructure, and AI, driving technology selection and architecture evolution. • Technical Direction: • Global Collaboration: Communicate and document in English with overseas colleagues and partners (exchanges, public chain teams, etc.). • Global Collaboration:
Benefits
• Competitive total compensation • Comprehensive insurance coverage for employees and their dependants • More that we love to tell you along the process! • All official OKX vacancies are published on this website. While roles may appear on selected third-party platforms from time to time, information on other sites may be inaccurate or outdated. If in doubt, please apply directly through our official careers website. • If in doubt, please apply directly through our official careers website. • Information collected and processed as part of the recruitment process of any job application you choose to submit is subject to OKX's Candidate Privacy Notice.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT