percona - Sr. Software Engineer - Go/MongoDB (Remote)
Requirements
• Prior work on database internals, CDC pipelines, ETL, or data migration tooling. • Familiarity with the Prometheus style of metrics and observability. • Experience with golangci-lint, vulnerability scanning (for example Trivy), and deb and rpm packaging. • Background maintaining or contributing to an open source project with an external community. • Exposure to MongoDB sharded clusters at large scale, where balancer behavior and backup interaction stop being theoretical. • Both projects live on GitHub, contributions go through pull requests and code review, and we track work in JIRA. We care about keeping open source open, so the default is that the work you do here is public and stays that way.
Responsibilities
• PRIMARY, ON PCSM • The core replication engine: initial collection cloning followed by continuous change capture over MongoDB Change Streams, with correct handling of resume tokens, ordering, and resumability after failures. • Correctness and fault tolerance at scale: recovering cleanly from network drops, primary elections, and restarts without losing or duplicating changes, and reasoning carefully about the delivery guarantees we can honestly promise. • Sharded cluster support: replicating across shards, dealing with the realities of chunk migrations and balancer activity, and keeping the target consistent. • Namespace filtering and automatic index management, plus the edge cases that show up with DDL, TTL, and index differences between source and target. • Performance and throughput: parallelizing the clone, applying backpressure, and keeping memory and connection use sane against large clusters with great change volume. • The CLI and HTTP API that drive and observe a sync, and the metrics and logging that let an operator trust what is happening. • ALSO ACROSS PBM • Consistent backup, restore, and point-in-time recovery across replica sets and sharded clusters, using physical or logical type of the backup. • Backup storage: integrating reliably with main cloud object storage (S3, GCS, Azure Blob Storage...) and remote filesystems, and handling the throughput and failure modes that show up at scale. • The pbm-agent and pbm CLI, and the control-collection state in MongoDB that coordinates them across the cluster. • Working in the open: pull requests, code review, JIRA, and the community forum. • Release quality: tests, packaging, and the CI and security scanning that gate every change. • WHAT HAVE YOU DONE: • Strong Go experience in production, with real fluency in concurrency: goroutines, channels, context cancellation, worker pools, and backpressure. You have debugged a race condition that only showed up under load, and you know how you found it. • Solid grounding in distributed systems and data consistency. You can talk clearly about at-least-once versus exactly-once, idempotency, ordering, and what it takes to make a stateful process resumable. • Hands-on MongoDB knowledge: change streams, the oplog, resume tokens, replica sets, and sharding. You do not need to have built replication before, but you should understand why it is hard. • Comfort building and operating command-line tools and HTTP APIs, and instrumenting them with metrics and structured logs. • A habit of writing tests that catch real problems, and comfort working across a mixed toolchain. • Experience working in the open: Git and pull request workflows, giving and taking code review well, and communicating clearly in writing with contributors you have never met in person.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT