Fundraise Up - Senior DevOps Engineer
Requirements
• 6+ years as a DevOps Engineer / SRE (or very close responsibilities). • 6+ years • Track record of owning technical initiatives end-to-end — from requirements and technical design through production delivery. You can showcase initiatives that were yours, not just tasks you completed. • Track record of owning technical initiatives end-to-end • Confident Linux skills (we use Ubuntu). • Linux • Working knowledge of the Prometheus stack: metric types, exporters, and how alerting works — enough to navigate and extend an existing setup. • Working knowledge of the Prometheus stack • Hands-on experience with CI/CD: pipeline design, build orchestration, artifact delivery. • CI/CD • Containers: Docker, image building, registries. • Containers • Ansible. • Experience with Bash or Python scripting for automation and observability (writing exporters, eliminating routine work). • Python • Production/on-call experience: diagnosing incidents, restoring service, leading post-mortems. • Experience mentoring less experienced engineers. • Ownership and attention to detail. Downtime is expensive: during busy events 10 minutes of downtime can cost us around $500k. • Ownership and attention to detail. • We understand it’s impossible to be an expert in everything, but it’s important to have solid hands-on experience in two or more of the areas below: • two or more • VictoriaMetrics / Prometheus stack at scale: architecture, cardinality control, exporters, alerting infrastructure. • VictoriaMetrics / Prometheus stack at scale • Log pipelines at scale: Graylog / VictoriaLogs / ELK — collection (fluent bit or similar), retention, sharding, performance. • Log pipelines at scale • Jenkins scripted pipelines: shared libraries, pipeline infrastructure, build agent fleets. • Jenkins scripted pipelines • Container registries and artifact management: Harbor, Nexus, base images, image policies. • Container registries and artifact management • Operating applications on Kubernetes: Helm, workload monitoring and log delivery, deploy troubleshooting. • Operating applications on Kubernetes • Grafana: dashboards as code, alerting, performance at scale. • Grafana • Great if you’ve worked with any of the following: • Analytics & DS platforms: JupyterHub, Airflow, Tableau, MLflow, Airbyte — deployment, maintenance, resource limits. Building platform around these tools to improve Quality of Life for Analytics. • Analytics & DS platforms • Remote development environments and AI agent execution environments. E.g. Coder/Telepresence. • Remote development environments • AI agent execution environments. • Bare-metal Kubernetes: provisioning, networking, scaling. • Bare-metal Kubernetes • Flux and GitOps. • Terraform • Sentry on-premise: operating self-hosted error tracking. • Sentry on-premise • ClickHouse, MongoDB. • ClickHouse, MongoDB
Responsibilities
• Own one of our core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you drive its architecture, reliability, and roadmap. • Own one of our core platform areas end-to-end • Drive technical initiatives end-to-end: gather requirements, write the design doc, decompose into tasks, implement, deliver to production, and own the operational health afterwards. • Drive technical initiatives end-to-end • Drive clarity in ambiguous situations by defining requirements, assumptions, and next steps. • Drive clarity in ambiguous situations • Design for reliability and scale: evolve the architecture of our platforms — topology, integration points, scaling approach, and reliability model. • Design for reliability and scale • Support developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deploys, help teams with metrics, alerts, and logs; participate in chat duty in developer support channels. • Support developers • Automate away toil: repetitive operations, provisioning, and maintenance should be codified, not performed by hand. • Automate away toil • Investigate production incidents as the senior escalation point for your area: drive resolution, lead post-mortems, implement systemic fixes. Participate in on-call rotations and raise the bar for how on-call works. • Investigate production incidents • on-call • Mentor less experienced engineers through design discussions, reviews, and pairing; catch debt-inducing shortcuts at the review stage. • Mentor • Use AI in all aspects of day-to-day work: researching, troubleshooting, developing.
Benefits
• A strong, collaborative product team that owns what it builds • Clear product vision and access to real customer feedback from global nonprofit leaders • Flat structure: no politics, just great work with great people • Transparent company culture-we share how we’re growing, where revenue comes from, and what’s next • Long-term focus: we offer equity options and value sustained, meaningful contribution • Final compensation will be determined based on relevant experience, skills, qualifications, and alignment with the role's requirements. • Private medical insurance for the employee and their family • 23 paid vacation days per year • 11 paid public holidays per year • 5 company-paid sick leave days • English learning courses • Relevant professional education • Gym or swimming pool • *Please note: All official correspondence from Fundraise Up will exclusively originate from the @fundraiseup.com domain. Exercise caution and ensure the authenticity of emails claiming to be from our company.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT