buffer - Senior Infrastructure Engineer
Requirements
• You've worked as an Infrastructure Engineer, SRE, "DevOps" engineer, or adjacent role for long enough to be considered senior. • You have hands-on experience operating production Kubernetes at scale on a managed offering (GKE, EKS, AKS), including authoring and maintaining Helm charts, and you're fluent with autoscaling primitives driving KEDA and the cluster auto scaler. • You have AWS depth across IAM, EC2, S3, SQS, ECR, and ALBs. You may have also used Cloudflare (WAF, Workers, etc.) and GCP (BigQuery). • You have strong Terraform skills. You default to modules for structure, and keep the code adaptable, readable, and self-contained. Bonus points if you contributed an OSS module. • You've operated production CI/CD with GitHub Actions (or equivalent) and GitOps via ArgoCD (or similar). You've authored ArgoCD pipelines and Helm configuration yourself, including canary or progressive delivery systems you'd trust to roll back safely. • You've built internal developer tools (CLIs, dev environments, per-PR environments) and you think about them as products with users, not scripts. • You have a track record of pragmatic build-vs-buy decisions on infrastructure tooling. You can defend a choice and revisit it when conditions change. • You've worked with DataDog, Sentry, or similar observability stacks, and you design logs and metrics with cost in mind. You know observability and cloud spend can grow faster than the company if no one is watching. • You're comfortable with the Cloudflare across Workers, Zero Trust, DNS, and the rest of their platform. • You read and modify TypeScript or Node services well enough to upgrade runtimes and unblock teams (legacy PHP and Python show up too). • You're fluent with modern AI tools. You use them to debug, document, and reduce toil, not just to generate code, and you bring those patterns into how infra runs. • You're proactive and you follow through. You spot what needs doing before you're asked, and you close the loop without being chased. • You turn ambiguity into proofs of concept. You take fuzzy asks, ship something rough teammates can react to, and iterate with them until it lands. • You thrive in remote, asynchronous environments. You're clear in your thinking, generous with context. • You don't wait for perfect information to start, and you don't wait for perfect to ship. • You see infra as a force multiplier for engineering, not a gatekeeper. • You care about Buffer's customers. When things are slow for them it's painful for you to see. When errors are flaky you find the root cause and try to eliminate the entire class of problem, because you see the system, not the bug. • You care about performance. If it's too slow to use, it shouldn't exist. You'd rather make it fast than work around it. • You're a generalist engineer with strong spikes: T-shaped folks with depth in infrastructure and the flexibility to pivot as priorities shift. • You create, not just consume. Open source contributions, a technical blog, conference talks, side projects, or active accounts on the platforms Buffer serves, your pick. We're a Team of Creators ourselves, and the closer infra is to the creator's experience, the better the platform becomes. • You think about infrastructure as a platform with users. APIs, SDKs, CLIs, MCP servers, or developer-facing tooling you've shipped where adoption, not just deployment, was the success metric. You've felt the difference between code that ships and code that gets used. • You play the long game. You'd rather invest in compounding fundamentals than chase the platform-of-the-month. • Bonus points if you're already a Buffer user or familiar with social media management tools. • Cloud and IaC. AWS, GCP, Cloudflare. Products you'll find in use here: EC2, EKS, S3, SQS, SNS, ECR, IAM, ALB, BigQuery, mostly managed in Terraform. • Container orchestration and delivery. Kubernetes on EKS, Helm for charting, KEDA for SQS-driven autoscaling. ArgoCD for deploys from git. An in-house canary system that we want to augment with Argo Rollouts. BIBEs and frontend branch deployments for per-PR previews. • CI/CD. GitHub Actions, with self-hosted AWS runners (including KVM-capable instances for Android UI tests). Our monorepo is the consolidation target, and we're moving more services into it over time. • Observability and incidents. Datadog for logs, metrics, APM, and spans. Sentry for errors. Incident.io https://incident.io (and a bit of PagerDuty) for the incident lifecycle and postmortems. • Networking and DNS. Cloudflare across Workers, Zero Trust, DNS, and more. CloudFront for AWS-side distribution. VPC peering for cross-account connectivity. OctoDNS for DNS-as-code. • Data stores. MongoDB fronted by GraphQL, Elasticsearch, Redis. • Local development. Hermes (our in-house environment that mirrors production) running in OrbStack, pre-built production containers from ECR. • Languages and runtimes. Node.js and TypeScript across most services, Python in selected services and tooling, and PHP for legacy services that are still load-bearing and being gradually replaced. • By submitting the application, you consent to Buffer collecting and processing your personal data for recruiting purposes, find more details in our Privacy Policy https://buffer.notion.site/Privacy-Policy-for-Applicants-230a7f5a00fd80e48478f09171d18a24?pvs=143.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT