gradera - Data Scientist
Requirements
• Strong ability to interrogate unfamiliar datasets and quickly develop a working understanding of their structure, semantics, and quirks • Experience working with messy, incomplete, or poorly documented real-world data • Skilled in identifying hidden patterns, trends, seasonality, and anomalies through visual and statistical exploration • Ability to ask the right questions about data — challenging assumptions, validating sources, and understanding the context in which data was collected • Proficiency in data profiling, descriptive statistics, and summary reporting to communicate the shape and health of a dataset • Experience creating data dictionaries, documentation, and data quality reports to support team-wide data understanding • Comfort working across structured (relational tables), semi-structured (JSON, XML), and unstructured (text, logs, sensor streams) data formats • Proficiency in Python (pandas, NumPy, scikit-learn, PyTorch or TensorFlow) and/or R • Strong SQL skills with hands-on experience in DB2 and SQL Server • Experience with Databricks for large-scale data processing, feature engineering, and model training • Familiarity with cloud platforms: Azure or AWS • Experience with data warehouses and big data platforms (Databricks, Snowflake, or Redshift) • Knowledge of MLOps tools such as MLflow, Kubeflow, or Airflow • Experience with streaming data technologies such as Kafka or Spark • Solid foundation in probability, statistics, linear algebra, and experimental design • Experience with deep learning, NLP, computer vision, or Bayesian methods • Familiarity with real-time or streaming data pipelines • Open-source contributions or published research
Responsibilities
• Collect, clean, and analyze large structured and unstructured datasets from multiple internal and external sources • Conduct thorough exploratory data analysis (EDA) to understand data distributions, relationships, outliers, and missing value patterns • Profile and audit datasets to assess data quality, completeness, consistency, and fitness for modeling • Investigate and document data lineage — understanding where data originates, how it flows, and how it transforms across systems • Identify and resolve data anomalies, inconsistencies, and integrity issues in collaboration with data engineering teams • Develop a deep understanding of the business domain and the underlying data that represents it — including what each field means, how it is captured, and what its limitations are • Translate raw, messy, real-world data into clean, well-understood analytical datasets ready for modeling and reporting • Apply statistical techniques such as correlation analysis, hypothesis testing, variance analysis, and distribution fitting to extract meaningful signals from noise • Build and deploy machine learning models including regression, classification, clustering, NLP, and time-series analysis • Design, evaluate, and analyze A/B experiments and controlled tests using causal inference techniques • Develop data-driven recommendations backed by rigorous statistical reasoning • Write clean, production-ready code in Python or R • Collaborate with data engineers to build reliable data pipelines and feature stores • Deploy and monitor ML models using MLOps best practices on cloud infrastructure • Build dashboards and self-serve analytics tools to support stakeholder decision-making
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT