Moxie - Staff Platform Engineer (IC-4)
Requirements
• 5+ years of experience in platform, DevOps, or SRE-focused roles • Strong experience operating production systems on AWS (certifications strongly preferred), Vercel, and/or GCP • AWS (certifications strongly preferred), Vercel, and/or GCP • Experience building and maintaining CI/CD pipelines (GitHub Actions or similar) • CI/CD pipelines • Strong understanding of cloud networking, security fundamentals, and IAM • Experience with observability tooling (Datadog preferred)Ability to troubleshoot production issues calmly and systematically • Datadog preferred • Excellent written and verbal communication skills in English (C1 or higher). • Experience with Cloudflare (DNS, WAF, edge configuration) • Cloudflare • Experience with Agentic and LLM based tooling and automations (Cursor, Codex, Claude Code, etc) • Experience improving local development tooling (ie Docker, Husky, bash) • Familiarity with infrastructure-as-code (Terraform or similar) • Experience supporting regulated or compliance-sensitive environment • Experience working with PII, PHI, and sensitive data systems in general. • Our Stack • Our Stack • Cloud & Infrastructure: AWS ECS/Fargate, GCP, and Vercel • Cloud & Infrastructure: • Edge & Security: Cloudflare • Edge & Security: • CI/CD: GitHub Actions • CI/CD: • Observability: Datadog • Observability: • Database: AWS RDS (postgres) • Database: • Version Control: Git, GitHub • Version Control: • AI / LLM Tooling: Claude Code, Gemini, Cursor, CodeRabbit, Glean, Codex • AI / LLM Tooling: • Backend Services: Python, Django (operational ownership, not feature dev) • Backend Services:
Responsibilities
• Observability & Reliability • Participate in incident response as needed • Own and improve monitoring, logging, and alerting using Datadog, AWS, Vercel and related tools • monitoring, logging, and alerting • Ensure systems are observable and failure modes are well understood • Help teams learn from incidents through postmortems and follow-ups • Balance reliability with delivery speed through pragmatic SRE practices • Platform & Infrastructure • Own and operate core platform systems across AWS, GCP, Vercel, Github, and Cloudflare • AWS, GCP, Vercel, Github, and Cloudflare • Improve reliability, scalability, and security of production and non-production environments • Maintain and evolve infrastructure supporting multiple services and teams • CI/CD & Deployments • CI/CD & Deployments • Own and improve CI/CD pipelines (GitHub Actions), focusing on speed, reliability, and clarity • CI/CD pipelines • Improve deployment workflows, rollbacks, and environment consistency • Reduce deployment-related risk and manual intervention • Partner with engineers and our QA team to improve release confidence and velocity • Improve local development environments and onboarding experience for engineers • Reduce friction in common workflows (setup, testing, debugging) • Maintain tooling and documentation that helps engineers move faster with confidence • Cross-Team Collaboration • Work closely with the product engineering team to understand platform pain points and improve local development experience. • Provide guidance and support on infrastructure, deployments, and operational best practices • Contribute to platform standards and shared tooling through collaboration, not mandates
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT