wagey.ggwagey.gg
31,365  jobs31,365  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(31,365)/Solutions Architect Role(596)/andromeda (7) - HPC Architect
Pro members applied to this job 36 hours before you saw itGet Pro ›
andromeda

andromeda - HPC Architect

North America Remote / San Francisco, CA - Hybrid+ Equity2d ago
In OfficePrincipalNADiagnosticsArtificial IntelligenceSolutions ArchitectProspectingCustomer TrainingKubernetesLinuxProcurementDue DiligenceDocumentation

Requirements

• Deep HPC experience: you've designed, built, or operated GPU clusters at meaningful scale and understand what makes them fast, stable, and debuggable. • Strong fabric knowledge, ideally both InfiniBand and RoCE. You can evaluate a topology, interpret fabric diagnostics, and identify why a fabric underperforms, not just that it does. • Experience with distributed orchestration and the HPC software stack: Slurm, Kubernetes, OpenMPI or equivalent, and the Linux systems engineering underneath all of it. • Data-center literacy: power, cooling, cabling, and physical-layer realities. You can walk a provider's facility and know what questions to ask. • Benchmarking judgment. You know which numbers matter for large-scale training workloads, how providers game them, and how to design tests that can't be gamed. • The ability to write standards others can build against: precise, testable, and usable by a provider's engineering team without you in the room. • Credibility in the room with external engineering teams, including the ability to deliver a failing grade to a provider who wants your business, and keep the relationship intact. • Comfort with ambiguity. The qualification framework you'll apply mostly doesn't exist yet; you'll write it. • Strong Candidates May Have • Experience inside a neocloud, hyperscaler, colocation, or data-center provider. • NVIDIA data-center GPU depth: DGX/HGX platforms, NVLink/NVSwitch, GPU health and RAS behavior at fleet scale. • Experience supporting AI research labs or other large-scale training customers, and familiarity with what their workloads punish: stragglers, fabric jitter, storage stalls. • Background in cluster acceptance testing or site bring-up. Ideally you've taken a cluster from delivery to production before. • What Success Looks Like • Within your first year, you'll have delivered: • A qualification bar that's written down and trusted. Andromeda's acceptance criteria, benchmark suite, and quality thresholds exist as documents and tooling — not tribal knowledge. SREs, procurement, and providers all work from the same bar, and a passing grade from you means the cluster performs in production. • Providers who build to our standards before we ever test them. Prospective providers get our technical requirements upfront and arrive closer to qualified, because the standards are clear enough to engineer against. Time-from-sourcing-to-qualified drops with each provider. • Onboarding that's technically boring. Clusters that join the network have been validated the same way every time, and the surprises that used to appear in a customer's training run get caught in acceptance instead. • A technical relationship providers value. Provider engineering teams treat you as the authority on what Andromeda needs and the first call when they're planning a build-out so that we hear about architecture decisions early enough to influence them. • A procurement partnership that changes what we buy. Procurement's provider pipeline reflects your technical due diligence, and deals are shaped by a realistic view of remediation cost.

Responsibilities

• Vet prospective compute providers: assess cluster architecture, GPU hardware, network fabric, storage, and orchestration against Andromeda's quality metrics, and make the call on whether a cluster qualifies. • Define the qualification bar itself. Build the acceptance test suite, benchmark methodology, and quality thresholds from scratch, formalizing what currently exists only as SRE tribal knowledge. Own and evolve these standards as the fleet and the market change. • Run validation hands-on where it matters: burn-in testing, fabric validation (InfiniBand/RoCE), NCCL and application-level benchmarks, storage performance testing. For routine or repeat validation, define the methodology and review results rather than executing everything yourself. • Guide providers through technical onboarding: work directly with their data-center and platform engineers to remediate gaps, tune configurations, and bring clusters up to the bar on a predictable path. • Partner with compute procurement to identify and qualify new providers. Perform technical due diligence during sourcing, and a clear read on how much remediation a candidate cluster needs before procurement commits. • Build and maintain technical relationships with existing providers: you're the engineer their engineers call, and the one who spots architectural or quality drift before it becomes a customer problem. • Feed what you learn back into provider-facing standards and internal documentation, so each qualification is faster and more consistent than the last.

Benefits

• High-growth environment: Get in early at a company at the center of the AI infrastructure boom • Ownership: First HPC Architect for the solutions engineering team, you’ll get to build this function from the ground up • Competitive compensation: + meaningful equity • Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

zipzip - Solutions Architect - Workday Financials1mo ago
·United States·$120k - $180k/year + Equity
RemoteNAPrincipalArchitectureNonprofitSolutions ArchitectWorkdayProspectingProcurementRESTDocumentation
ChainguardChainguard - Senior Partner Solutions Architect - East (Professional Services)3w ago
·Remote - United States·Equity
RemoteNAPrincipalSoftwareSolutions ArchitectKubernetesDocumentationClose
OXIO CorporationOXIO Corporation - Solutions Architect1mo ago
·Remote - United States, Hybrid·Equity
In OfficeNAPrincipalTelecommunicationsSolutions ArchitectProspecting
ZscalerZscaler - Transformation Architect - Healthcare2mo ago
·Remote - California, USA; Remote - Texas, USA - Hybrid·$171k - $171k/year + Equity
In OfficeNAPrincipalCybersecuritySolutions ArchitectDocumentation
Reliable Robotics CorporationReliable Robotics Corporation - Systems Architect4mo ago
·Remote - Mountain View, California, United States·$190k - $300k/year
RemoteNAPrincipalArchitectureRoboticsAerospaceAirlinesSolutions ArchitectDocumentation
zipzip - Solutions Integration Architect1mo ago
·United States·$135k - $200k/year + Equity
RemoteNAPrincipalArchitectureNonprofitSolutions ArchitectProcurementSupplier ManagementProspectingRESTSOAP
Grafana LabsGrafana Labs - Observability Architect5mo ago
·Remote - PT (Pacific)·$182k - $217k/year
RemoteNAPrincipalSolutions ArchitectCustomer OnboardingReportingGrafanaKubernetesDocumentation
Grafana LabsGrafana Labs - Observability Architect | Washington, DC | Remote5mo ago
·Remote - United States (Remote)·$182k - $217k/year
RemoteNAPrincipalSolutions ArchitectCustomer OnboardingReportingGrafanaDocumentationKubernetes
Grafana LabsGrafana Labs - Observability Architect | USA EST| Remote5mo ago
·Remote - ET (Eastern)·$182k - $217k/year
RemoteNAPrincipalSolutions ArchitectCustomer OnboardingReportingGrafanaKubernetesDocumentation

Browse more by category

Show 596 moreSolutions ArchitectShow 2,370 moreProspectingShow 226 moreCustomer TrainingShow 1,672 moreKubernetesShow 806 moreLinuxShow 943 moreProcurementShow 244 moreDue DiligenceShow 5,261 moreDocumentation
Privacy·Terms··Contact·FAQ·Wagey on X